# Get Out of the Model's Way — Kevin Hou — session 2026-06-30T20:30:00.000Z → 2026-06-30T20:50:00.000Z

_70 transcript lines · 3 slides · source: full recording_

## Transcript

This is great. Thank you. Infrastructure for the meta superintelligence lab and the infrastructure organization. Today, we are going to be talking about production of ads for our Gentech systems. When most people hear the word valuation, they think about benchmarks. A model scores 90% on a benchmark. A new version scores 92%. The team celebrates. But agent systems have fundamentally changed what the evaluation means. Today, the systems don't simply generate answers. The plan, the call tools, territory information, they execute workflows, they interact with the production infrastructure. The The question is no longer did the model generate the right answer. The question is did the system behave correctly? I would like to discuss how evaluation is evolving from model benchmarking into production infrastructure. This is the problem almost every AI organization is encountering today. Offline benchmarks continue improving, yet production reliability often remains unpredictable. Why is that? Because benchmarks measure model capability. Today Production measures system Here we are. A benchmark doesn't capture tool failure, API outage, context changes, user variability, long running workflows. And as systems become more autonomous, the gap between the benchmark performance and production performance grows. The result is what many teams. Alright. Hello, everyone. My name is Kevin. I'm gonna be talking about anti gravity. So are there any World Cup fans out there? Woo. Imagine you are coaching Argentina, and you're in the eighty ninth minute, and you have Messi on your team. What play are you running? It's called Messy the ball and get the heck out of the way. LLMs aren't just role players anymore. They can be your star player if you build the right product around them. And to let your star player cook, you have to get out of the model's way. We might want to get the slide. Are the slides up? Oh, they are. Great. So anti gravity is Google's agentic coding product for technical and non technical users. We launched back in November 2025 and have been accelerating devs both within Google and Externally ever since. My name is Kevin Howe, and I lead part of the engineering team on anti gravity. So let's talk a little bit more about what anti gravity is. We have and always will be unapologetically agent first. So we debuted the anti gravity I d e last year with a brand new agent manager concept, and it was a platform to manage and orchestrate many agents. Since then, we've actually extracted our agent and launched our own anti gravity c l I. And last month at Google I O, we had the pleasure of launching anti gravity two point. Oh, in the theme of getting the model out of the way, we actually decoupled the I D from the agent manager. So now you have two separate applications. And now you can use the agent manager in a stand alone app. And since pictures are worth a thousand words, here's a screenshot of anti gravity two point o in action. As you can see, not Only is it your own dedicated mission control for your agents and projects. You have sub agents. You have all the new models. You have work trees, scheduled tasks, voice mode. There are so many things to unpack with the product. But I don't want to spend today telling you about the product. I want to tell you a little bit more about the behind the scenes, some of the principles that went into it and notably some of the things that led to its roadmap. So as some of you for the long time a I inch fans, this is actually my fifth time speaking at AI eng. And I've been building developers tools since 2022. And the one thing that has stood above all other lessons that I've talked talked about is the idea of scaling with intelligence. This means that as the model gets better, so should your product. And the frontier edge of whatever model you are serving should be apparent inside of your users product experience. So let's get into more concrete examples of what this means. So for those of you that follow me on X or hear me just yap generally for the last four. For years, you'll know that I've been working on a number of these sort of transformations year over year over year. In 2022, I was working on auto complete and chat sidebars. This was based on embeddings, rules, files, A. S. T syntax tree parsing. Basically everything inside of that app is deterministic because that's all that the model could really handle. And in 2024 when agents came onto the scene, it completely changed how developers were going to do work. With it came new primitives like M. C. P. With 2025, we introduced anti gravity's agent manager with many other products following suit in that similar form factor with users managing many agents at once in parallel. And this led to things like skills, hooks, artifacts, and a couple other primitives. And that sort of defined the 2025 era. So let's talk a little bit about 2026 and what those primitives might be. Before we answer S, custom tools and permission systems. Question, I want to take you back to some of these battle scars that are a little bit closer to home. Scaling with intelligence really is not easy. It's really hard to take away something that users love and are familiar with to lead them down potentially, and that's a big key word, a better path. We aren't right a 100% of the time, but there are two that jumped to mind when I was putting together the slides for this talk. The first one is giving AI a terminal. We all remember fears about Son of Anton deleting your entire code base and doing catastrophic things to both, you know, your startup, your company, etcetera, etcetera. But as models got Better and people invested in primitives such as permission systems, users ended up building faster. They ended up shipping more and they did so safely. So we're able to overcome this. And as models got smarter, they were able to make better decisions about what they should and should not run-in your terminal. The second instance is, this tweet, which is very representative of sort of the yelling that I got. When we removed chat From Windsurf. So a lot of users were yelling at our team because we took away something that was very dear to them, the chat sidebar, and replaced it with only an agent. Now at the time, this is something that was familiar and rather difficult to swallow. But when we look back, models have advanced. Multi step research, agentic research and execution became the new paradigm. And here we are today using and loving all these agentic products. And so now I bring you to today's battle. What is going on today? So we decoupled the agent manager from the I D E and with anti gravity two point zero, we split them Into separate applications. We believe that the I D E is to the agent manager what the debugger was to the I D E. You don't always need a debugger, but it definitely is helpful to have it if you need to go a layer beneath and go one step deeper into that abstraction stack. And our prediction is that this idea of agent orchestration, you can call it agent teams, you can call it swarms, you could call it software factories, is the future. And we're willing to bet In that future. So here are the primitives for what we're calling the agent teams 2026 era. These are things like sub agents, generative you I and sidecars, and we'll talk more concretely about what those things are and some examples of how they manifest inside of the product. But it's really important to first understand the why. What brought about these changes and what model changes, what model properties actually led to the development of these new things? And as a product team, do you force the new era of primitives? Or is it something that comes to you? By using the model and experiencing the model. The answer is kind of both. Right? And the privilege of being inside of Google DeepMind is that we do have that relationship between the product and the model. So you remember the crux of anti gravity one point o is to manage agents in parallel. To put the human in the driver's seat. And if you remember my last talk, I talked a lot more about this research product flywheel. And now As promised, because of the anti gravity product, Gemini has now learned a thing or two about how to manage a team of agents. There's still a lot of headroom to make multi agent systems better, more collaborative, better at deconstructing tasks into smaller tasks, but we've got a really good head start with Gemini. And all the basics have been imbued to the model so that we can build a product like anti gravity two point o. Gemini 3.5 flash, yeah Gemini 3.5 flash was launched Back in April. And this brought to market a lot of those capabilities that we had been working on in the background with anti gravity. And Flash now isn't just good at executing tasks. It's actually really good at leading teams. It's faster and cheaper, pushing the Pareto curve of what is intelligent versus the speed and the cost at which you run those things. And putting this all together, we were really excited to announce agent teams in public preview inside of anti gravity. All you have to do is simply type the slash command slash teamwork and you'll see a new mode where you can enter and unleash a swarm of agents onto the task at hand. So we'll talk a little bit about how this works. You as a user will specify your task. The more specific you are, the better, though the nature of these agentic communication styles is that if it needs something more, it can actually ask you for more Until everything is basically clear. You'll work with that lead agent, and it will manage a team of arbitrary size to get that work done. And what I like to say, it's kind of like the avengers. Right? It'll take A bunch of specialized roles. It may front end engineers, back end engineers, infrastructure specialists, QA design. The list goes on and on and on. And there are infinite possibilities for what each of those sub agents could take on. Each sub agent is dynamically generated and can operate independently. And it can even actually select a different model from what the main agent is using. And this is done so by that main agent. Again, we are scaling with intelligence. And one of the coolest aspects of this is that it can use generative U I. With a model that is as fast as flash, things can happen nearly Instantaneously if you ask, hey, what is the status of my task? Show me a Kanban of what's going on. Or maybe, you know, you prefer something a little bit more like, the the Chrome debugger tool. It can show you a timeline like that. And all these things are generated on the fly because it's able to generate UI on demand. So some of the projects that the system has implemented, we've built a photo editor. You can actually edit raw photo. Directly inside of your browser. We've also built a messaging app that might look a little bit familiar to those in the room. And each of these took hundreds of sub agents, and took almost half a day to run. But to really put it through its paces, one of the hero runs that we did was actually building an entire OS kernel. This is something that we got to show off at Google IO, but we built a complete OS kernel from scratch and actually played Doom on it. And my colleague Varun was able to demo this at Google IO. We were super proud of this particular milestone because it really demonstrated that if you throw more intelligence, you throw more Sub agents, at this sort of problem. A model like Gemini 3.5 Flash could do this in a way that was not only very, very powerful, but also scalable and, you know, mildly affordable. Obviously, we're not gonna spend thousands and thousands of dollars to build an OS kernel every day, though it is possible. And some of the stats out of this, it took 93 sub agents over the course of twelve hours, made 15,000 requests, 2,000,000,000 tokens. And it was under a thousand dollars, which was one of the really cool aspects of this project. And so as you can see with this particular example, sub agent primitives are one of the defining parts about building 2026 era of agent teams. So agent teams are just that first example, and I want to show you another example that our team uses internally that sort of demonstrates some of these new primitives. The second one is about automating research tasks. So we work inside of Gemini. We help sort of make Gemini better at coding related tasks, agentic related tasks. And this is where the real magic starts happening with the We have an internal version of anti gravity that researchers, engineers, non technical folks can use. And when they understand the primitives that anti gravity offers, it becomes a very, very powerful way to automate your own workflows. So we'll take the example of side by side eval analysis. So this is a very common workflow, not only at DeepMind, but just generally in the industry. You essentially will take multiple rollouts, one, two, three, four, etcetera. And you want to compare them. So you'll take a set of tasks, you'll do some rollouts, you'll get some results, and they'll essentially be in two different tables. Now you'll look at the control, you'll look at the experiment, and then you'll have to figure out not only what the difference was, but perhaps what are the reasons for those differences and how can we actually iterate from there and make a better version of for the next experiment. Now traditionally, this was a lot of Jupyter notebook elbow grease essentially. But when you start working with the new primitives in 2026, you end up with a lot cleaner of a work. Woah. So research arrears were able to automate 90% of this workflow by simply asking the agent about the evals in question using natural language. Then the agent, that is now primed with skills and an understanding of Google's massive monorepo code base, is able to crunch the numbers and get back to you with a delta. Now what's really cool here is instead of just taking that delta then handing it back to the user, it went the extra step. It spun up for a research agent specialist that proposes a 100 different hypotheses over why those deltas might occur. And then it uses sub agents to then spit up one sub agent for each hypothesis and basically drills into that particular case in parallel, mapping back to a single response and then telling the researcher, hey, here are some areas that I found. Now, here's a report that you can review. And what's really cool is that it doesn't stop at just the report. It actually puts together a generative u I for you to look through, interact, select drop downs, filter, segment, slice, and actually interact richly with that data. And internally, we care a lot about this sort of workflow. Improving the model, improving the product, and understanding the ways that users find success and failure internally at Google. So what used to be a very manual process now takes minutes. So what used to be hand engineering, you'd have to build your own async pool of agents. You'd have to set up your judges. You'd have to tape together data pipelines. All of this now starts becoming grounded in these new primitives that we've established earlier in the slide show. You have a sub agent graph that is completely dynamic. The generative UI comes in at the end to richly convey Findings in a way that the user best understands or maybe caters to their learning style. And all of these things can be regenerated and redone on the fly. All the user had to do was load up a skills file and ask away. So with teamwork and this eval example, we start arriving at these twenty twenty six primitives that I keep talking about. And these model characteristics really change the way that we have to think about the product and how we have to develop the product. So the three examples that we've talked about. First, we have the dynamic sub agent. And to provide a little bit more color here, basically no two sub agents are the same. The main agent is the one that is orchestrating this entirely on its own. It's configuring and prompting and seeding these sub agents on the fly. They can operate in parallel. They can operate in different types of secure environments, be it a sandbox, be it a remote execution system. And they can all take on an infinite number of specialized roles. So the scaling story here is quite obvious. And from the last two examples, you can probably tell. Smarter. Your team will become more specialized. It will become more collaborative. And ultimately, that means it will be capable of getting more complex work done for you. And now the second is this new concept. As the model gets We've alluded to it slightly in the past, but it's called sidecars. This This is a new plug in protocol that we're bringing to anti gravity. A sidecar process is essentially it is a sidecar process. The naming sort of reflects what's going on under the hood, but it's Long lived utility and it's responsible for listening. It allows the model to listen to the outside world and set up its own triggers for things that might happen. For example, this could be SMS messages. This could be web hooks, cron jobs, hooking it up to GitHub PRs. The list goes on and on. But this is a generic plug in primitive. Antigravity already uses sidecars for things that are time based. This is where the scheduled task cron concept comes from. But under the hood, this is all this new sidecar primitive. So we'll be releasing the spec for this so that you all can build on top of this new primitive, later this summer. Internally have been using this sort of concept. And the third and final primitive is generative u I. So we hypothesize that human written specialized u i's are kind of dead. Gemini flash on anti gravity clocks in at almost 900 tokens a second. This is 10 x faster than a lot of the other frontier model experiences. But there are some really really cool ways that people And in a matter of seconds, you're able to go from whatever you were thinking inside of your head into a Pumped into a use case that is designed and embedded inside of your conversation view perfectly. And rather than rely on templates or even HTML files, antigravity can render your generated UI inline. So you can do things like this and play doom, but this also extends to things like bar charts, graphs, tables, anything that you would want to interact with and maybe inspect a little bit further than just a markdown file or just a conversation. In generative u I in many ways reminds me of the quote that the late Steve Jobs said when unveiling the iPhone. He justifies the Removal of the keyboard and says they all have these keyboards. They are there whether you need them or not. And they all have these control buttons that are fixed in plastic and are the same for every application. In an analogous way, we built our product to dynamically scale with the needs of the agent. We skipped the heavy infrastructure and mechanical UIs in favor of sidecars and generative UI. And that creates a product experience That is not fixed in plastic. So sub agents, sidecar triggers, and generative UI are the latest primitives that are powering anti gravity. We've tried our best to stay out of the way and let the model cook. And if you're building a product around an agent, you should consider what are the primitives that are in my product and how might they scale with the model's intelligence? We all are familiar with shipping features is now quite easy with all of these new tools. And it's about deciding what features to actually add. So that the model, so that your product can scale with the next release of the next model, which will inevitably be faster, better, and cheaper. And so with the right primitives, you as a builder or you as a product owner, you might be surprised at what the models can do. And in classic fashion, I'm going to keep using this slide until we've actually conquered the TPU crunch. So you can find me on Twitter. You can DM me for feedback.

## Slides

### 00:00:42

# AI Engineer
## World's Fair

### 00:01:17

# Production Evals for Agentic Systems
Measuring reliability beyond accuracy. Building evaluation systems for autonomous AI workflows.

Nishant Gupta
Tech Lead @ Meta

[Diagram showing a complex network

### 00:01:42

# AI systems evolved faster than our evaluation methods

## The Illusion
90%
Benchmark Accuracy

## The Reality
Invisible Failure Modes
Degraded Production Behavior
Unpredictable User Reliability Gaps

[Gauge showing 90% Benchmark Accuracy]
[Line graph showing performance over time, with three distinct drops labeled as Invisible Failure Modes, Degraded Production Behavior, and Unpredictable User Reliability Gaps]
