# The 2026 State of AI Engineering — Barr Yaron — session 2026-07-02T16:00:00.000Z → 2026-07-02T16:20:00.000Z

_72 transcript lines · 0 slides · source: full recording_

## Transcript

Launch control. We have a go, Richard. Ladies and gentlemen, welcome to the AI Engineer World's Fair. We're delighted to have you with us as we continue exploring the ideas, technologies. And people shaping the future of AI. Please join me in welcoming your emcee for today's program, developer relations engineer at Replit, Ralph Chabry. Good morning, San Francisco. How's it going, guys? Welcome to AI Engineer day four. Wow. I'm so excited to be here with you. Guys today, so, this is our biggest AI engineer event ever. We have about 7,000 attendees. We also my god. When I look at this room, yeah. It's you guys, you made it to day four. So, let's give it up to you guys. All right. I live in a tiny village in Switzerland, and I can't believe that I can fit my entire village right here in this room. Speaking of Switzerland, who's here from Europe? Wow. So many of you guys and I see that you're strategically placed under AC. So, please guys, white in America, stay cool. Alright. Anybody local to the Bay Area today? Alright. So many of you also? Nice. Okay. To be honest, guys, I'm super excited to be in San Francisco. I think there's no place to build for AI and talk about AI. And what I think I think you guys are the coolest people in the world because you're constantly In the future. You guys drive self driving cars, but you also have food delivered by drones, and what I noticed nowadays is that nobody closes their laptops anymore. I don't know. Everybody seems to be Asian maxing. Speaking of maxing, you guys enjoyed yesterday's talks? Yeah? Let's give it up for our speakers from yesterday. Alright. There was so much to talk about and to think about. Yesterday. So we had Tariq from Entropic talking about Fable five and, the process of discovering new models capabilities. We had another Tariq from Sonar who talked about code verification. But what I noticed and what I heard from speakers is that it's still very important to stay in the loop while working with Frontier models. And today guys it's going to be about harness engine. Engineering. And we have a great lineup of speakers for you today. So we have speakers from leading organizations such as Anthropic. We have also, Stanford. We have the folks from DSPI also who are going to talk to you about new protocols. They're going to talk to you about how to separate the tasks from the model. And, and also they're gonna talk to you about how to scale AI systems. And for having seen some of these talks, I can tell you guys that you've been that you're here for a treat. Alright. And after the keynote, we have breakout sessions and we have so Any tracks for you guys to choose from? So we have tracks such as, software factories, generative media, memory, and so on. And please take also the time to go and check out our expo. We have so many cool swags. Our partners are there. Go talk to them and please if you need to take any swag it's gonna be this is today, this is the last day so, it's now or never. Speaking of partners Please, a big shout out to our presenting sponsor, Microsoft. Alright. Let's give it up for Microsoft, guys. Also, let's give it up for all our lab and platinum sponsors. Please keep it up For also for our gold sponsors. For our silver and bronze sponsors. This is super important, guys. Without the support of each one of these companies here, this event wouldn't be possible. Alright. So, again, we're gonna talk about harness engineering. But before we introduce our first speaker, what I want you to know guys is that you as These you have huge power over over the quality of the talk that you're gonna be attending today. Okay? So, our speakers have been doing this forever. They know their stuff inside out, but what we want them to do is to feel the energy in the room the second they hit the stage. Alright? Okay? So, we're gonna practice that for a second. So, at the count of three, I want you to imagine as if I was a speaker and I want you to welcome me as a rock star. Okay? So, one, two, three. Hi. Nice. So, let's keep it up for our next speaker. Okay? So, without further ado, please let's introduce our first speaker of the day. She is a partner at Amplify. Please join me what, welcoming to the stage Bar Yaron. Now joining us on stage is the partner at Amplify, Bar Yerin. Fantastic. You did a great job practicing. I feel very, very loved. Let's get started. So like you just heard, my name is Barr. I run a survey every year on the state of AI. Engineering. And the funny thing about running a survey on the state of AI engineering is that the field changes as you make the slides. Just in the past week, we've had frontier releases treated like national security events, Meta reportedly exploring selling AI compute. By the time I get off stage, maybe something else will happen, so if I miss a major announcement while I'm up here, please come find me after. But that's exactly why we run the survey every year, to cut through The noise, take a moment, step back, and understand what AI engineers are actually doing. For the first time this year, we were thrilled to partner with Notion and Vercel to run this survey. Very quickly on me, this is the least interesting slide. I'm an investment partner at Amplify, very lucky to invest in companies built by and for AI engineers. And I'll make the same promise that I make every single year, which is short time on Our long time on bar charts, so let's get right into it with lots of bar charts. First, let's talk about well, maybe raise your hand. Did you fill out the survey? This is a very large group. Okay. Yes. I see you in the front. If the answer is you, thank you so much. If the answer is not you, I will find you in 2027. But genuinely, this only exists because a thousand of you gave your time, so thank you. We had one Ten forty eight respondents this year, which is a lot of AI engineers. And to be precise, this is not just AI engineers. As I'm sure you see at the conference, every year we see that AI engineering is more of a discipline than a job title. It touches founders, CTOs, engineers, product people, folks across company sizes, and experience levels. And that range shows up in experience too. For the third year running, we see the same pattern, which is skewed towards senior engineers but new To AI. Of those with over ten years of software experience, over half have three years or less of AI experience, which tracks. These are very experienced engineers learning a new paradigm in real time. And the newest cohort, the ones who just started engineering, the median new engineer has nearly as much AI experience as the median ten year software veteran. Newest engineers have never known software without this. But doing AI doesn't mean one thing. We talked about all these different titles, all these different roles. Before we get into models and agents, I have a more basic question, which is: when people say they're doing AI at work, what are they actually doing? So first up, like to start with the modalities. We asked, Which modalities are you actively building with at work? Can Can anyone take a guess? Text dominates. I know. Hold your applause. I always look at is the ratio of nope, I'm not using this modality to I'm not using it, but I do plan to. I call this the intent to adopt ratio. Of the people who are not building with a modality today, how many say they plan to use it? And audio has the strongest intent to adopt this year. Among AI engineers who are not building with audio today, a whopping 56% say they plan to adopt it in the AI app. But one piece of this chart that I always find very interesting Applications they build. And this is not a brand new signal. Last year, audio also had the highest intent to adopt across modalities, but 37%. So audio continues to take the lead and have high interest, but that interest is accelerating. Now, there has been an audio swing, but if we look at what changed most from the last year in the survey, the biggest jump is actually in people using image generation. The share of respondents using generative AI for images and feeling really good about it doubled from 18% last year to 36% this year. Makes sense if you look at what we launched in the same window. Over the past year, plus survey time, we've had models nano banana nano banana two, Chachi Peti images two point zero. The products have gotten much better. What used to feel like an efficient way to generate cursed hands is just increasingly becoming a part of real work. Audio may have the strongest intent To adopt, but image generation shows us what happens when a modality crosses that threshold. So I'm excited to continue watching these adoption curves every single year. I think we're gonna see a lot this year. Now, models. Who here spends time on Twitter? Alright, yes. I imagine this is a very Twitter pilled crowd. If you spend any time on Twitter in this circle, you've seen a lot written about open weight models these past few months, and I think we'll see it even more. Or in the next year. So we asked what models are you actually using in production? 94% use closed models. 45% are using open weight models. But here's the thing, you know, open weight models are not replacing closed models for the most part, at least not yet. The respondents using open weight models, over 90% of them are also using closed models. So they're looking like an augmentation. Teams are mixing and matching. We also asked, just to double click on this, for the top three considerations when choosing a model. If you're choosing a model, what is important to you? And despite the airtime of the open versus closed, it's not what drives model choice. It was a top three consideration for only 5% of the respondents. What matters is actually more straightforward. It's quality. Quality dominates. Followed by agentic capabilities like tool calling and cost I'd write with it. We'll money, money, money. We'll get back to that. One thing that I found very interesting is that reliability is not near the top. Only one in five named reliability. That doesn't mean teams stopped caring about reliability. There are different ways to interpret this data. My guess is that it's more likely to become a threshold requirement and the models they're choosing are reliable enough so the decision moves up the stack outside of certain circumstances to quality, capability, cost. But we could talk after. Alright, so here's where the model story all comes together. Like I said, teams are not choosing one model and calling it a day. Earlier, I showed that 87% of teams are using more than one model. The model that's the opposite of standardization. And the way that they choose models for given tasks varies. Most popular is By task type, some run multiple models compare outputs, some route based on cost. But models are good at different things. What was interesting was that more than half of respondents said that their organization is starting to standardize on fewer AI tools. They're trading flexibility for standardization. A share of those are mixed. They say they're standardizing on some layers while staying flexible on others. But the headline here is that we're in the early great standardization of the platform and tools, not the models. Alright. This This is the slide where anyone who's opened an AI bill in the last year starts nodding. So, it turns out that infinite intelligence still comes with a usage based bill. Once teams are managing many models and AI workflows, the next question becomes cost. Cost is now a first class engineering constraint. We see this in the data. 40% of respondents say that cost regularly Shapes how ambitiously they use AI. And another 36% say that it sometimes does. Well, this is pretty straightforward. So all in, about, three out of four respondents are adjusting their AI usage based on cost, and maybe the fourth has a company card. That might be surprising, or maybe it's obvious, but twelve months ago, it was not. Token maxing is cool Being able to find real use cases is amazing, but cost is becoming a real big part of the product decision today. And it shows up in monitoring too. What folks are monitoring in production includes cost and token usage as the number two thing they watch for. It's being monitored like an SLA right under quality itself. Which brings us to the biggest line item of them all: agents. We've been talking about agents for a while. This year, as you've seen, as you'll see today, as you've seen in previous days, you're going to talk a lot about harness engineering, the escaping demo world. So we asked respondents what level of tool permissions their agents typically have. And this is where agents start to look more real. There are two things happening at once. First and I don't think this is surprising relative to last year, there are far more teams using agents. This year, 95% this seems high to me. 95% say they're using agents, roughly double last year. Second Amongst the teams that are using agents, those agents are much more likely to have write access. Last year, 52% of folks building with agents said their agents could actually write data. This year, that number is 89%. So when you combine these two shifts, more teams using agents and more of those agents having write permissions, the share of all the respondents and again, it's a survey using write enabled agents is up more than three Times relative to last year. So this is really the big shift. Agents are no longer reading, summarizing, drafting they're taking actions inside of systems. And that raises the obvious question: how are we controlling all of this? With pretty blunt instruments, there are many ways that folks are controlling agents today. The top two are human in the loop approvals and gating permissions, which are the right instincts but kind of the same toolkit you'd use to manage an intern. Below that, the results scatter. Task decomposition, retrieval, memory, sandboxing people are trying everything. Nobody has settled the control layer for agents. Memory and persistent context is one that I'm watching very carefully right now. I think it's going to evolve a lot in the next year. And when agents fail, or when people complain about agents failing, to be more precise, it's usually the thinking, not the plumbing. So, you know, like two thirds say that Hallucination or losing context mid task is what frustrates them the most. Alright, so agents are out in the wild, which makes it a good time to look at what everyone's actually running underneath. So let's take a peek at the stack. We asked, What is the biggest challenge in your stack? Every single year that I ask this, the answer the number one answer is evals. So evals lead here, same as always, but by a very thin margin. Like, that margin is getting smaller. And I'll say the quiet Part here, which is that 96% of the people in the survey in this room have a problem with the stack. You just can't agree on which one. So if you're deciding what to build next, if you're interested in infrastructure, that scatter is the map. And the leading challenge, how to evaluate your AI outputs, requires many different methods, but as always, the Vibe review is number one. So there are some consistent things that we'll see if they change over time, but they have not changed. Okay. This is interesting. So across eight layers of the stack, we asked, what do people build versus buy? Again, maybe the corporate card is is gonna play a part in this. But there is a wide range and mix for every layer of the stack and a few clear takeaways. So, the first is that inference and model serving is the layer that people buy the most. Many people don't want to build inference infrastructure, and fair enough. Prompt management is the opposite. 61% build it themselves. Apparently, everyone's prompts are special. And this is true of a lot of the product logic, prompts, rag, evals, they tend to stay in house on a relative basis.

## Slides
