bit worried about it. But you're probably thinking, David, you're just having a bad day. You've not, like, you've not, like, like, adopted some of the practices that have been kind of, like, done so far. And that that might be fair to say. Right? So it might be it's just how I feel about this stuff at the moment. So that's not really good enough in terms of, like, talking to you a lot today. So did a bit of digging, had a look at a lot of reports that are going around. There's things like this one from CodeRabbit. But I've centered on one that's from, Pharos AI. So they're report includes, telemetry from about 22,000 developers across 4,000 teams. And so I'm gonna take you, Seuss, through some of their initial findings that they've had. Right? So their initial findings is that for developers using, like, AI workflows and really going heavily adopting it, that their tasks completed go up 33%. That's pretty good. Right? That's pretty good. Yeah. Dave's like, yep. That's pretty good. 66% of epics gone up as well. How good is that? Right? That's the big chunky stuff getting over the line. Brilliant. 210% increase in tasks involving code. How good is that? Right? So maybe I was, like, a bit delusional, a bit sad and stuff. And I should end the talk there. Is anybody on the leadership track? Anybody? Yeah. You'd probably like these numbers. Right? How good are they? Like, it looks like productivity. Awesome. Right? So should we go for drinks now? Alright. So there's more to this report. There's more to this report. The report also talked about what it was to get the changes out that have been produced. So first thing, 157% increase in median time to first review. So that's that's a developer or an AI getting a change up and having a first review. To get the code all the way over the line, up by 442% median time for review. And then this next one, this puzzles me, right? I've, like, interrogated this one a bit. 31% increase in PRs going out with no review. And I looked at this. I thought, it's not just like agentic reviews or something like that? But it's not. It's like no reviews. This is freaking Yolo Town. Do you know what I mean? Like, so I don't know what the base number was on that, but I wanted to include it because I'm just like, what the actual f, right? So, okay, so that's code reviews telling the story that that these teams starting to feel a bit of pressure in terms of, like, the code that's coming through, and, review times are starting to become the bottle. Bottleneck, which has been covered a couple of in a couple of the other talks. So but what happens next, right, so that's about getting the code in for review, getting it approved, all of that kind of stuff. 243% increase in incidents per PR. I don't like that number because it represents having to act on stuff. And, like, hidden work coming back through, you know. It's not good. 54% increase in bugs per developer. That's like the worst KPI that you could possibly offer people. Right? And then the last one, so my favorite number tends to be like lead time. So lead times, like how long does it take stuff to go from, like, idea to the getting it over the line, that number up by 480% So those numbers that we first looked at where everything was looking pretty good, new were excited, I could see it on your face, they actually show that teams do look very productive initially. But there's this thing called a whiplash from that report, and that's what the report's called. About the fact that lots of hidden work starts to build up and the team starts to struggle a little bit in terms of the influx of code and the fact that these bottlenecks are really kind of like starting to rear their head quite badly. So I still feel like that a little bit. That report didn't help me, did it? It looked pretty good at the start, but then it went really downhill quite fast. So there are options. Right? There's lots of options for us to take. I've picked out three. Now, the first one I'm gonna give you is that you can pretend that imbalance isn't there, that the influx of code and the fact that you need to put more energy in to verify it, it's not gonna be a thing Claude five will come in, rewrite your code base, you're sweet. Or you might get a SaaS tool, bring that in, It's making the reviews kinda fly and that kind of stuff. You can bank on all of that stuff. Doesn't feel like it's quite there yet at the moment, so kind of gambling that those things kinda, like, can catch up and kinda, like, deliver you. There is some boring stuff that you can do as well. So you can make verification cheaper. And what I mean by cheaper is you make it kind of fast and reliable for people. Like in like years ago, this would have been a manual process. Right? Now, there's lots of automation in place. So you're starting to invest more into that to make sure that those steps, those things that developers need and and AIs need to verify things quickly is, like, super sharp, super fast to to to run. These are linters. These are unit tests. These are visual regression tests, all that kind of stuff. The other thing that you can do as well is that if you think about the balance getting out, you can actually put more pressure back on the beginning of the process. So this is the idea that you're trying to make writing the code a little bit like not slower, but putting a bit more rigor into it earlier. You're trying to make sure that when the code does get up to that review stage, that it's actually done at the standards that you expect them to be done at. So these are things like making sure that commits well, like, labeled. You've got why on your PRs, all that kind of stuff to make sure that when it does go through review, it's nice and easy to actually do it. And teams can kind of, like, take action quite easily. There's also things around speculative driven development. We saw a great talk on that earlier today, and touched on a couple of times. It's also making things like ADRs are in the code base. That kind of stuff. So putting pressure back on the process that's earlier in the process. So you've got those so, yeah, so the second two, really about catching problems like before review and then the other one. You're really just hoping. Right? That's okay. I think a lot of us will be we're doing a bit of that right. We've got a lot of faith in how this kind of is going to shake out. Now, I've got your three options there. Thinking about standards now? Like, standards are good. Right? I love a good standard. So starting to think about how you break those standards down into, like, So the first area that you want to be thinking about is, like, the hard rules, the yes and no things, the things that you would die on a hill for. So this is things like, if your security scanner says that you've introduced a vulnerability, you're like, that is not going anywhere near the code base. It's a hard no. Right? Have measurable thresholds. These are things where the performance of your application can be measured in certain ways, and you want to make sure those standards are high. These are things like the response time of the API maybe, or I'm a client side developer back in the day. You were talking about, like, Sid CSS. I'm like a I'm a front end developer until I die. Alright? So Lighthouse scores are really important to me. I love a good lighthouse score, which aggregates a whole bunch of metrics. To say how good the performance is is of a client side application. Also thinking about rubrics. So Tanya, she was in the other track yesterday. She had a great talk on rubrics. These are the things where you define the standard and then you can kind of, like, mark how close you are to that standard and also include when you don't reach it at all. Right? These are things again, these are useful for people as well as LLMs nowadays, right? You've also got the shipping criteria. I was actually really surprised. We're doing a lot of like pilot programs at the moment at nine, and I was actually surprised that teams struggle to define what it was to get something out into production tends to be something which is more like tribal knowledge. For teams, and no one's actually kind of, like, written down what the expectation is. So getting that written down and making that as something that's easily available to both people and LLMs is is good too. Right? So then you got to start you got these standards emerging now. All good. Now you can start to think to verify your standards as well. So you've got a few options in this space as well. So deterministic checks, really bore in really old Right? Just like running the tests as we've always done. If you're a node developer, you've probably got NPM run test already there. You want your test here to not really be about the AI, not about the LLM, You want them to be nice and fast, reliable, and always like, reliable in terms of, if I run it this way, one plus one equals two. You know what I mean. And they got the other side as well. So the advisory stuff. We're seeing more of this this stuff come through. You can use those rubrics that we talked about on the previous slide and you can start to think about how AI might start to run those and flag with a human to say that something is just a little bit off. If you've got certain tasks which are lower risk, then, then you can start to think about, just kind of like letting the judgment of the AI kick kick in at that point? So with these ones, with the deterministic checks, the machine checks, those should never be optional, and the AI review ones should should never be blocking. We're starting to see things come through around, like Copilot can do reviews. They're getting pretty good now. And, I think a few folks have talked to me about things like CodeRabbit as well doing good. Quality reviews on that front. So we're not just limited to people doing this stuff anymore. We can start to, like, augment the workflow with, the AI. So I've got this quote, like, we've been spending a lot of time in this space at Nine, and I really like this quote from one of our principals. Like, early on when we're starting to look into this, there were very concerned about what's gonna happen about this kind of, like, influx of the code that we might get. So, Mitch captured this quite nicely, I think, which which is that in an age of repeated output, as in things increase and the volume of code increasing, that we need to balance that with having repeated quality as well. So what happens when you apply this stuff I got quite overzealous using Claude at one point. I actually destroyed my own website, like, haven't we all done that? Have we all done that? Is it just me? I'm one I'm one of those people, you know. Anyway, I got a bit carried away. Destroyed the website. This was on their Mother's Day weekend as well, so that was really good. Like, I'm thinking, oh, man, I really need to kinda, like, get this thing back. Anyway, I thought about restoring from backup and I thought, no, f it. I'm just gonna rebuild the whole thing with Claude. Like, because I was having some ideas on what it might be bring in standards and have, like, a workflow that means that the quality of the final output might be better if I put certain standards in place. So with that in mind, again, I'm client side developer until I die, right? So measurable threshold that I picked. Was lighthouse. Wanted the standard to be 95 plus. Right? Wrote the standard down. In that case, I just had a a standard markdown file. I think I probably hooked it in through an agent. Markdown. I just kinda, like, pointed it to. We saw an example of that in one of the talks earlier. So this is just about me making sure that the LLM knows the standards I'm trying to reach. Had a deterministic check, which, again, just npm run lighthouse, runs the lighthouse tests for me unlike a sample of the URLs that I've got. And then it came to the results. So I did it with the standards, and then I reran the whole exercise of creating this thing without the standards. So with the standards, it hit 100. So it actually exceeded what I expected it to do, which is awesome. Right? Rerunning again without the standards, It was still decent, but the lighthouse score was 69 on that on that front. So some things to take home and do at home. Right? Pick one dimension that you feel passionately about, or your team feels passionate about. It might be something on an OKR whatever. Right? Like, pick one dimension to focus on. Pick one document that you can write that clarifies what that standard should be, and then implement one check for that standard as well. If you bring all these things together, you do it for the first time, you'll feel the benefits and you'll be able to kind of repeat that process again and again and again until you've got, like, better coverage. Because if you set the standard, the AI will help you reach it. Thank you. No no questions. Oh, okay. You did ask, especially. I did. No. Time. Oh, are we over? Oh, do you want a question? I can I can have a question or two? I was talking to Dave yesterday. I'm like, oh my god. I hate questions so much. I get really kinda, like, stressed, obviously. You can either take a question or everyone can get to the bar soon. Oh, yes. Do you wanna go to the do you wanna get drinks, people? Alright. That's why I still alright. You're off the hook. Yes. Thank you. Thank you. Okay. So before oh, yeah. Alright. So that just about wraps it up for AI Engineer Melbourne for 2026. Let's have one final round of applause for all our speakers today. So it's not quite over yet. So just across Federation Square, there's the transport hotel. So in the glasshouse, there's closing drinks, which are courtesy of our wonderful sponsor, SonaSource. You just gotta bring your, your badge, your lanyard, across to get access. Keep an eye on your inbox. There'll be plenty of news about the conference, including where you can watch the recordings and more. Also, one other thing I've been asked to say is the code check has been moved, so it's now downstairs under the countdown sign that you might have seen on the way up. Make sure you bring your numbers from the cloak check at Zinc. And, apparently, a few of the speakers had left some things as well, so they've also been moved up. Thank you so much. Hope everyone had a great day. And we'll see you at the after drinks. Testing the stream audio in software engineering. One, two, three, four, five, six, seven, eight, nine, 10.