# How Anthropic Builds: Lessons from Labs — Mike Krieger, swyx — session 2026-07-02T17:00:00.000Z → 2026-07-02T17:20:00.000Z

_80 transcript lines · 0 slides · source: full recording_

## Transcript

We've gone from evolving few shots to prompts to now harnesses, and now evolving your Emails over time too. But what you need to ask for any of these new techniques is how do they help you solve harder problems or solve your own problems better? And you should ask this question in a data driven manner. You should look at this new technique, say, how can I apply this to the business problem that I have? You should define your problem, and you should hold your prompts, models, and code accountable to the problem that you need them to solve. And what's awesome about when you build in this way, where you have flexible implementations? What you unlock is you unlock the ecosystem of all the techniques that anyone in this room is constantly inventing. You unlock access to the collective intelligence of everyone here, all sharing techniques together. So, if you wanna build reliable AI software, I encourage you to come check out DSPI. We're completely open source, open research, and we're here to help you solve your problems by building reliable software. We have a discord that you should come join, and when you come up with the next technique, you should come contribute it to DSPI, and we can help you distribute it and make this awesome technique available for everyone. Thank you. Joining us on stage is the cofounder of Instagram. And a member of technical staff at Anthropic, Mike Krieger. How's everyone doing? I mean, good morning. Nice. Mike, thank you for releasing Fable, just in time for us. Exactly for the conference. Timed it. We're we're so glad to have you. You are one of the preeminent builders and you're leading labs at Anthropic. How has your model usage changed as as you've, you know, seen models internally grow? Yeah. I mean, for me, it's been, like, both the model shift and then my role shift. So I'd for, like, the first two years I was at Anthropic, I was chief product officer. And then I kept seeing people build with the models. And the FOMO just kept increasing because I was you know, use the As much as possible. But, for example, on product strategy, I would write a strategy doc and then have cloud critique it. And maybe you can use a workflow, but it's not quite the same as, like, building in that pure way. And I was, like, spending all my weekends trying to build with it, and I realized, okay, I actually just need to shift. It's, like, way too interesting a time. And it's actually an interesting trend I've seen now. Like, several people that were CTOs at other places are, like, now joining as I see Zedentropic in other places. This but I made a role shift and it was actually right around the time where we started getting sort of internal snapshots of what became Mythos and Fable. And what was really interesting watching that sort of shift was, that kind of change between I have an idea. I'm gonna, like, sort of break it down in my head much more how I do the engineering normally and then kind of iterate through these different steps to moving to much more of the paradigm of, I'm gonna describe the goal. Like, go off and work on it, and then, like, we could talk about what trade offs you you know, surface some questions along the way, but then figure out what where you landed and where we can go. From there. I find it's hard. I don't know if people have this experience where and I know Fable's only been re enabled for a couple of days. Fable's definitely way, way smarter than me. So sometimes it'll finish work and be like, here's the trade offs I made. I'm like, can you explain it to me like I'm a little dumber than you are because I need you to, like, sort of break this down for me. But that's been one sort of big change is sort of moving from that task delegation to, like, express the end state and then have it go and and cook on it. Yeah. We're all learning how to delegate better. Tariq did us a huge favor yesterday. We do I wanna read it in the newspaper? It's, you know, that we have we have write ups of talks now in in, like, the next day's newspaper. He said be unreasonable. In what ways, you know, have you been more ambitious? Yeah. You're prompting. I love that. I mean, I love that framing. We actually just hit this today. I'm one of the labs initiatives I have is internal product. And, somebody was like, hey. It doesn't work the way I want it to, and can you make some changes? And I've realized I'm just gonna go ask Claude to do this. Like, why don't you ask Claude? And this is a nontechnical person. So I actually think as an industry or even as a product team, we have to teach people to be more unreasonable in their usage, and it's sort of hard to imagine. I think that that if I can digress for a second on product design, I think right now, the, like, kind of first generation of AI products, we put them too much in a box and constrain their their sort of access to tools or kind of degrees of freedom, which means it was much harder to be unreasonable. Right? When you say, do this thing for me, and then it would Oh, I I can't. I can barely like, I can write code, but I can't really run it. Or I can kind of introspect my environment, but not really. And I think as you see our own, like, product progression, even the things like coworker, like, you know, does every single, like, knowledge worker need a virtual machine that can write bash? Like, on the face of it, no. But then when you realize, oh, actually, that way it can remediate an issue where, oh, I tried to parse a PDF using our built in PDF parser. I hit this yesterday, and it was like, I can't parse it this way. Well, okay. Well, I can probably write a script that can do this. As well. So I think that's it. My most unreasonable thing, though, was, one of our labs projects I wrote in Python, like, near and dear to my heart. All of Instagram was in Python. But I think they're finally converting it to PHP now that they have, like, models that they put. I know. Let's give you some tokens. And, for deployment, I realized that Cloud Code had, like, figured out a better deployment story with BUN. And I was like, okay. I need to port this whole thing from Python to TypeScript. Like, as a you know, if I put on my, like, twenty ten's engineering hat or even my early twenty ten twenties, like, that's Dumb idea. Like, who would ever port, like, at that point, you know, a couple 100 thousands of lines of code? But I was like, I think this is doable now. And I basically created this dynamic workflow, set up it over the weekend, had it port the whole thing, like, verify it, double check it, then read both code. Like, basically, churn and churn and churn and then came back Monday to a completed workflow that was a ported version of that thing. So that probably ranks on, like, the more unreasonable things like, yeah, just port this entire Python code base to TypeScript, get it working, get it deployable in, you know, a weekend. Yeah. I mean, a lot of people are talking about the the bun zig to rust Yeah. Version. I think a lot of people are also like, well, it's a compiler. It's a it's a runtime. It's got lots of tests. Easy to do. Can you port Instagram, which you would know very well, to PHP like that, like a like a product? Yeah. I mean, I think the product side, it's even I don't know if it's easier or harder. One of the things we did at Instagram, this is when Python three came out, and we were able to add type hints for the first time. And it was people had a lot of internal conversations, like, are we gonna run out of steam? Python. And my perspective was always, like, I think we can take this way further than we think we can, but I think types are gonna help us not sort of be in our own way. And we built this thing called monkey type where we basically, like, captured runtime type sig like, basically, the types that were actually getting used in production and then map those back to to the types in the code base. And I think because of that sort of pattern, I think there's really interesting ways in which if you're doing sort of conversion Or sort of cross compiling using LMS. You can also lean on production data a lot more or run sort of, like, segmented tests. I think that, like, there's a lot of, things you can do there. But, yeah, I think it's I mean, the sky's the limit there as well. I think the hardest part is always finding the boundary around where you can start doing incrementally without trying to boil the whole ocean and, like, swap it overnight. Yeah. I mean, your users are your test ultimately. And, you know, we I also read another article in the newspaper about how you can just use rollouts, and sometimes you don't really know, what you're gonna need it for. But when that infrastructure exists for your experiment and to roll things out, it enables so much. Yeah. Mean, I always found this was advice we got. It was like we launched Instagram and the happened to be the first week everything melted because we didn't really know what we were doing on the back end side of things. And, coincidentally, that week, there was, like, a lunch that one of our investors just scheduled. Like, not even for us. It was just a a, infrastructure lunch. And we ended up spending we totally, like, monopolized that conversation because everybody had their own opinion about how we could fix our scaling. And, like, the two pieces of advice I got there is, like, 2,010 that I Like, we'll forever retain is, like, like, basically, like, pre measure everything that you think you might even remotely need because the worst thing is an outage where you're like, well, is this, like, number normal or is it high? And, like, oh, I don't know because I don't have data until I just added this metric. And the other one is being, like, really thoughtful about knobs and feature flags. So even, you know, early Instagram, we had, like, a very, simple but really effective, like, way in which you could do, like, ramp ups and rollouts. And dynamic config too where, you know, a lot of our run time configurations had to be changed in a matter of seconds so that we could handle load. Really important. I'm seeing that definitely in in AI as well where, you know, we're making all sorts of different trade outs and having that kind of runtime configuration is super key. Yeah. My my favorite scaling story for Instagram, by the way, I think it's like your launch day when you you DDoS yourself with email. Yes. And being able to, like, do that on a first class way was Which which people should look up that story if, if you haven't seen it. I wanted to go into tags. Very, very major ship. It's it's how 60 something percent of your code is written today. Yep. How did you square that with everything you just said where it's, like, very dynamic, like, you Actually ship one app. You ship one app with 3,000 flags. Yeah. And, like, well, what are you working on today? I don't know. Like, it's it's for this segment of the population. Yeah. Yeah. I mean, I think it's a bunch of things. So, like, with I was really excited. I was talking to Switzer. They're like, I'm really excited that we have tag out there because it is, how we've been working for a while. And I would get up on stages and people be like, how do you work at Anthropic? And I'd be like, oh, yeah. We use these things, like, that are not quite cloud code, but, you know, but it's hard to describe it. But, I mean, if you, like, got to poke into Anthropic, like, you Could see, of course, cloud code usage for things that are, like, more interactive or if you're kind of iterating on a particular, sort of sort of specific thing where you want a lot of, like, the high sort of bandwidth back and forth. But most usage is actually much more delegating, via tagging and via tagging. You can say, like, here's the and the reason it's really interesting is how multiplayer it is. And it reminds me sort of of, like, actually, like, mid journey. Like, the fact that everybody was on Discord seeing how other people were using it. I think it actually, to your earlier question, really helps with that unreasonable. Or ambition where the first time you see somebody tag Claude and be like, hey. You know, don't just fix this bug, but, like, now you are responsible for this part of the code base, and I want you to monitor this feedback channel and proactively take on tasks and then fix them. And then also take like, you know, if this API changes, do that. Like, I saw somebody do that. I was like, oh, wait. I've been totally underutilizing this thing. I've just been using it as, like, a glorified cloud code in Slack. Like, that's definitely, totally, like, sort of new version of it. Right? The more advanced version is really trying to sorry. Think of it as a teammate that is actually sort of how old's context, how's memory, and it can be proactive. And that's just really changed how we operate internally. It's much more like this multiplayer async proactive way than it is a, you know, most people off in their own CLIs. Are you bottlenecked by code review and Git? Obviously, there is cloud review, but someone usually still looks at it. Is there a world in which you just merge it in? Yeah. We're it's a really good question. We are definitely still bottlenecked on review. Especially for things that are, like, touching some architecture pieces. And it's actually more subtle than just being bottlenecked on review because that's you know, okay. We can carve out time differently. It's, like, bottlenecked on human ability to even, like, fully conceptualize what we're doing. So one of the reasons we built Cloud Code artifacts that we shipped a couple weeks ago was partially for that, which is, you would send somebody a PR, and then they'd be like, I don't know, man. This is, like, 2,000 lines of code. Like, it looks like code to me. And what we started doing instead is sharing much more like Here's a cloud code artifact. Like, here's the explanation. Here's the intention of the of the change. Here's the trade offs that were made. And, like, I think that's gonna be much more be the trend by which we communicate, which is the code is ultimately, you know, verifiable using some things. But, actually, like, discussing intent and trade offs and then measuring in production is, I think, the at least the direction of travel we've we've gone. I don't review when I get a pull request, I wish wish I could say I've reviewed every line of code. I definitely do not. I've, like, actually talked to Claude about the the code and say Like, these are the questions that I would have. Can you go investigate it? So it is kind of cloud powered code review, but still human driven and and for the really important ones. And for the ones that are, like, cosmetic visual changes, it's much more like, look, like, we'll fix Ford if we need to fix Ford. Yeah. Totally. I think a lot of people are here are trying to figure that out too. I wanted to talk also a little bit about, Enthropic Labs in general. Nilay Patel, who you've probably met. Before loves to ask ask the question, like, draw the org chart. Yeah. Like like, how like, people you know, you ship your org chart. Like, I think it's important. Like, everyone knows plot code. Now you've got tags. How are you structuring the labs? Yeah. It's a good question. Again, what we were trying to wrestle with was you want sort of people to be supported. Like, you know, I think the the death of the engineering manager discipline has been greatly exaggerated. Like, I think there's still a lot of coaching and interpersonal pieces and personal development that I think is still really, really important. But especially in a labs type group where, like, our whole cadence It's two week reviews where every project goes up for, we call it persevere or pivot. So, basically every project is up for review and either it's time to, you know, keep going persevering or, you know, it's time to pivot it or even shut down. And, you know, we've shut down projects basically every single one of those cycles, and it's like, the more you do it, the less it's just like, oh, no. My project has shut down. I failed. It's like, no. That is definitely the intention of the labs team is to prototype quickly, try to ship internally, maybe get it to early access, and if it doesn't work, wind it down. But because of that kind of, like, Iteration. It means that if you align the org chart too much to the individual projects, you're gonna end up, like, reorging every two weeks, which would be a total nightmare. And so we've actually ended up with this interesting setup where, like, the the pod or the team that is working on a given, we call them bets within labs, Definitely just draws upon, like, alright, somebody from product, somebody from the eng team. You know, I'll jump in when it's a product. I'm particularly interested, and I'll come in and work together with the team on it. And that's the unit for that time. And there is the concept of a bet lead or a directly response Individual. But the interesting thing is that they don't manage usually any of the other people, which kind of breaks the that kind of previous way in which a lot of these things were done. But I think it leads it leads us to be really flexible when you say, okay. Actually, this project is not gonna work out. Let's disband and keep going, and it's not a big deal. And the engine manager is much more playing the, like, make sure every individual is assigned to the thing that they're most excited about and that they're working in the best way possible. Now what we do sort of solidifies when there's a product that has, like, legs. Like, cloud design, for example, started in this sort of ad hoc sort of Group way. And then now that, like, we've shifted, it's gotten traction. We've done, like, a big second release, in June. Like, it's becoming like we've hired people for that specific team, and it has more of a of a structure. So it's, like, loose until it gets solidified down the line. What's the future of cloud design? I think a lot of people are very interested in it's one of your biggest launches this year. Where does this go? I think for me I mean, the things that are holding back cloud design for being even better is better interaction with our other Surfaces. So, you know, I was designing something where I was talking to to Claude Coe the other day. I'm like, I want a really much more seamless, like, what I'm talking about, the design for it, you know, interactive design back to that. And I think in general, it's I mean, this goes back again to, kind of unconstraining Claude. Like, the fact that our services don't talk to each other as well as they could, I think, really holds back a lot of interesting ideas around what we could do. So I think that's one, like, kind of major area that we're looking at. And then the other one is people like, the lines between Mean, a cloud design and an app get blurrier and blurrier over time. Like, I've seen people of course, there's no, like, persistence, but build, like, fully functional, like, even games, which is definitely not what we designed cloud design for. But you can do it. It's just HTML and JavaScript. So blurring those lines even further and thinking through, like, what is the path from a, like, fully featured design that looks really well to really good to something that is maybe more like an artifact where you're actually able to go and you know, persist data and share it with others and build from there. So I think that those lines get really interesting. Over time too. Yeah. A big part of design is having taste. I actually asked Fable what Fable wants to ask you. And this this this is what Fable came up with. You deleted almost all of Burbn to get to Instagram, which is like you had a whole, you know, solo mo whatever thing, and you went to Instagram. What would you delete in AI? Or more spicily, what would you delete in Claude? Oh, I like the spice. I think I mean, we have it's interesting. We have a a one of our Slack channels, like project. Which is, like, what is in the product right now? It's and it and, I mean, this is hard at Instagram. The Instagram, we well, some things that had, like, four to 5% usage. You're like, oh, that's really not very many. But then you have, like, 20 features that each have four to 5% usage. It's like the classic Microsoft Word problem of, like, everybody uses some disjoint subset of the of the functionality. So that that's always the challenge. Now I think, we're a younger product. So hopefully, we have Of those things. Like, we unshipped styles, I think, recently where it was, like, used by a small percentage of people and was not really AGI built in a lot of ways. It was, like, very sort of prescriptive in the way that it worked and skills were a much better, application, something like that. So I think you have to be willing to take the primitives of, like, one generation of AI and, like, un ship them or at least, like, supplement them or, supplant them with the next one as well. I think the biggest thing is I look at it. I've been spending some time, like, outside of labs on some of this is, like, man, like, we're asking people to make, like, code versus co work versus, like, chat distinctions. And, like, one, they don't interop Well, and they can't delegate to each other. And two, I think the average person off the street could not explain to you why those surfaces are all different. So I think deleting some of the product complexity within our our code or our product, I think, is a a thing that would would serve all. Also, because then Cloud can do what it needs to do and and do well. Like, there's nothing more frustrating than having a co work session where you're like, great. I've mapped out exactly what I want you to build and then be like, can you please, like, create a Graph that I can paste into cloud code. Like, that is some 2020, you know, kind of workflow there that really shouldn't exist anymore. Yeah. I think drawing lines on what you don't want to do and also sort of leaving room for others is interesting. A lot of people today is, like, the start ups day for AI.

## Slides
