# Getting the most out of Codex — Jason Liu — session 2026-06-30T17:45:00.000Z → 2026-06-30T18:05:00.000Z

_86 transcript lines · 11 slides · source: full recording_

## Transcript

No. As per the standard in gene link response, your gut will tell you to pull the raw prompt from the telemetry logs pass. To the same model using the same prop and run it locally to isolate the bug, which we'll all do. And, surprisingly, it will work as well. Run it again, it will run again. You run it 10 more times. It will be just perfect every time. But now let's talk about that one run which costed you, and that will be gone. You can reproduce it. And if you can reproduce it, you can debug it. And if you can debug it, you can promise it won't happen to your Next customer or user. Right? Now I am Tisha. I have Sushin with me as my co presenter. We both run agents against real production back ends. You know, the kind of place where a bad right isn't. Oh, well. Hey, everyone. I am Kushan. I worked at Sarvam as a founding engineer for two years. Let's talk about what I'm interested in right now, and that is browser agents. Browser agents as an idea are so cool. Right? Browser agents should go crazy. Right? I personally have not seen that adoption. I mean, me myself, I don't use browser agents that much. I've been exploring that for some Time. I'm trying to understand why that is. So on my screen right now, we have the browser challenge. But, this is a very interesting benchmark for browser agents because there are so many things that you have to do long rise in sequencing of your tasks. And this actually reveals, you know, why browser agents suck. If you saw at the beginning of the video, the browser, this agent took, like, maybe ten, twenty seconds just to click the start button. And now we're on step one. There are 30 steps, and it has taken so long just to click one button. So enough of this. I wanna show you what I've been building. So same website. I've tried to sort of replicate the feeling of seeing what's happening. You know, you can see what the agent is thinking. But as you can see, it is so much faster and so much quicker, and I'm using a much cheaper model. Right? The hypothesis here is models are pretty smart, but it's the infrared around them that sucks. If you noticed in the video earlier, maybe I put a screenshot. The agent is trying to debug what's going on. It's trying to click something, but it doesn't understand what's going on. So my core thesis here has been give a nice environment for the To use. Right? So it can plan long sequences. It can figure out where it failed, what is going on, and it can plan the click correctly. I figured out it's a is scattered across not just a code base or a single repo, but across things like Slack and email, meetings and documents, a bunch of other stuff. Right? And then half of it is actually just stuck in someone else's head. And so I have to have my AI reach out, have them answer questions, and move forward. And, you know, when I was a individual contributor, maybe it's very easy to just write code. But as you sort of scale yourself, you find that a lot of the knowledge work that is happening today is just figuring out, like, what has Happened, who is waiting on what, and which loose threads needs attention. And so if you wanna take any away from this talk, you can just leave right after this if you want, is these six things. One, compaction works now. You can really feel comfortable pinning a very long thread and knowing that it's generally gonna remember, what you've told it. You should also start getting really comfortable with talking to your computer. Voice is really good, and voice is only gonna get better through time, and you're gonna be able to do even more with, dictation rather than a keyboard. Computer use and app shots are very great. It's a great way of bringing in context, and I'll talk a little bit more about what that looks like. And then, most importantly, really start investing in what your personal memory looks like. Right? What are the skills that you want to use and what are the skills you wanna automate? But also, what are the plugins you want to share with the rest of the team? And once you have that set up, you can start thinking about these pin threads and automations like teammates. And more interestingly, I don't know if many people know this, but every thread in codecs can talk to each other. So maybe today you can work on having a pin thread feel like a teammate, but very quickly you might have threads that run automations that communicate other threads. Now you could have this idea of a manager, and these are some of the really big concepts I wanna bring to you today. And so if you've used something like chat to BT, in the past, maybe you're taught. Okay, you know what? You have to really have these very short threads. You know, they only happen every once in a while. After a couple of messages, the context is gonna run out and you really need to create a new thread to capture context. Now all of my work happens in a pinned thread. You pin it. You give it a name. It runs over time. And with automations and triggers, you can start waking these things up. And so codex really happens in three acts. Right? You bring the contact into the system. You do work on this pinned thread, and then you figure out how the computer can then act and write back to the rest of the world. And we can talk about these things pretty quickly. How many people here actually use dictation? Just show a show of hands. I'm curious, like, how many people yeah. You know, maybe Like, 30%, 40%. I I I can almost ensure you that by the end of the year, more and more people will do so. Right? Tony Stark is not like typing into a keyboard talking to Jarvis, and I really want you to sort of recognize that by the end of the year, you're gonna feel like Tony Stark. You're gonna be automating a lot of your life through voice. Voice is, like, three times faster than most people type, and I have, like, a hand injury. I can't even type anymore. I just use a foot pedal to control dictation, and it's pretty awesome. And I'm also really lazy. So if I had to type, I would just send a message, like, fix this. Right? And and I hope that the AI can figure out what's going on. And then when it doesn't, I complain on Twitter that, you know, the AI sucks. But in reality, when I use dictation, I do just blabber. Right? Take a look at this issue, check the browser. There's some white space I'm not happy with. Maybe it's a screenshot from the website. Maybe it's on Figma. Then, like, make a pull request, edit only the CSS, and then also, like, when you're done, message the guy on Slack and then wait for the preview link and then send them the preview link too. I would never type this, but this basically is very easy to say. And the The model can figure out how to actually take these actions. Once you have all this input in place, you just have to connect it with the rest of the world, and this is where you wanna install plug ins. Right? Most of my work is just Slack, Gmail, Calendar, Notion, Linear, Obsidian. By having these plugins, you can really just give your AI the ability to reach out into the world and read more information. Once you do this, I just want you to try out a very simple prompt. Right? Triage what's changed around my projects. Bring me the important things, and you're really gonna be surprised at how much these newer models can just figure out what you do for work and how you can become more productive and what's important. Once you trust the plug ins work, you can start also thinking about your own plug ins. So I work on the developer experience team a lot of the time. I'm also just dealing with all of the messiness of the feedback on slack and on twitter. And so I built out some skills that Just do delegation and triage. It knows, given some feature, who's worked on it and what slack should I need to share this in. Right? And I can invest in this for myself and then share this with the team and you can kind of become the plug in hero. Right? You build these automations and you can scale not only yourself, but the entire organization. But sometimes, you know, the context doesn't really come in by asking codecs to read Slack and figure out what I just saw or read Twitter. There's this really cool feature called app shots. It's one of my favorite features of all time. All you have to do is press both command keys beside the space bar. It'll take a screenshot of the application plus all the context and the accessibility tree. And then half my conversation with codex is just I take a snapshot of something. I'm having a conversation on slack and I just go like answer this question, reply back. Also make a pull request if it's important and turn this into a skill. I've actually invested so much in my triaging skills that usually I just feel like a manager now. Right? It's just like I take an app shot. I send the question mark. And the AI has to figure out, oh, someone asked for, like, credits from Twitter based on some outage and let me go handle this. Right? So if you just do anything today, you try try an app shot workflow. Right? Maybe you're reading some feedback on Twitter and Slack. You take an app shot. Codec will then trigger some kind of triaging skills or some calm skills to communicate between the internal Slack versus the external Twitter, and and then you can just pin that thread. Okay. This is some outage or this is some rate limit issue that people are facing on Twitter. Let's figure out how we can manage this. And that brings us to part two working on the work. So now basically my sidebar is just every single project I manage. It's a chief of staff thread that just wakes up every once in a while. I have a single thread that runs all the documentation for open a I one for these slides. I I maintain the open source community, and also I just have a thread to monitor x feedback. You can pin this and work with it over time. And if you want to do something like keep an eye on it every 30 minutes, you just ask it to. And just by doing that, it will just wake up every 30, check for any update. And update the context that's around you. You can do this to babysit a pull request. Right? You can just say, make sure the tests are green. Check every hour. Address all the feedback. If c I is broken, fix it, and make sure it's always mergeable into main. And I'll do that. Right? I also do things like coordinating support. Check that someone on Slack has answered a question. If you can answer the question, then report back on Twitter. And then also my chief of staff thread. It just wakes up every thirty minutes. I don't pay for my tokens. I do thirty minutes. If you are using Your pro plans, like, doing it, like, twice a day is very, very productive. But not only can you run these automations on your computer, you don't have to be glued to your computer. Right? These are where this is where features like remote control really come in. So remote control is the ability to view every local and remote thread on your computer in the cloud from your phone, And so a lot of the work now it can be done, like, at a park or riding a bike. It's actually pretty convenient. Last week, I was working on a launch video. This is not like a code basis. Launch video. Someone gave me some feedback, and so I pulled out my phone, and I told my computer to go find the video, make some changes in iMovie, share the video on Slack, and just monitor that Slack thread and say, okay, every thirty minutes, if someone says the wording is off or the timing is weird, try to fix it and send a revision. And when I got back home to work on this video, there was no more feedback to execute on. But threads really just work across a single work stream. One of the big things I really also recommend is just setting up, like The memory base. Right? Like, I just use an Obsidian Vault. It generally looks like this. I have a people directory. I have a project's directory and some agent notes. And if you scan this QR code, you're gonna get a template of how I set up my my system. And so you can just pull this repo, telco does it set up for you, and it'll install all the skills that I use and introduce the same structures that I really recommend. And then lastly, you can sort of choose the forms that you wanna interact over. For example, the same plug ins that can read context, right? You can read an email. You can also draft things. I don't know if people here really trust ai to send emails and slack messages on your behalf. But generally, most of my automation is now also will prepare drafts. So every email that I ever has when I review it, there's already a draft. Every slack message where someone's asking me a question, I already have a draft set up there. You can also do a lot of progress, make a lot of progress with artifacts. You can make pdfs and slides and excels and documents, and they all Open within the codecs app, and that way you can just annotate and give feedback directly in the application. And then recently, we've launched a sites feature. So now you can just, like, buy put your little applications. You have logged in with chat GBT, and you can connect your databases, And you can effectively have applications that you share across the org. And so more and more, instead of sharing a Google Doc, I might actually just share a website with the database to communicate with the rest of my team and track the progress of some kind of work. Right? Maybe it's launch readiness, or maybe it's some kind of migration. And if there's no integration, if there's no plug in to do something, this is where things like computers comes in. Right. Codex has a very powerful browser built off the same technology we built Atlas with. And so you can do a lot with things on the chrome extension side. Right? You can have, authentication. You can have multiple tabs running in the background. And if it's a native app, then you can use just general computer use. Right? With computer use, it doesn't Take over your screen. It still uses the accessibility tree and allows you to control your computer in the background. And so oftentimes, I'll just be doing work and I'll switch it to a tab or I'll switch to some other application, and I realize that codecs is just doing what it needs to do to get the job done. And these things are also incredibly powerful. In the past week, I've had codecs do a bunch of really fun things. It's gotten me refunds on a plane ticket. Right? It just checks the customer support channel every five minutes, and it it just wants me to get my money back. Every time I see a really long form, I just hit an app shot, and I say, fill out this form for me. You know, recently, I had and I don't recommend this at all, but recently I had, Codecs use DocuSign and sign a document and then send a fax message to my doctor. And I was like, oh, I need to send a fax message. This is, like, you know, faxpdf.com. I just need your attention to put in the credit card information, and then we can go forward from that. So a lot of these things are automated now. Really fun. You can also use this to test native applications. Right? Most of the time when I'm working on the codecs app, I just spin up another version of the codecs app, and the main codecs is just testing the future it's building. And I can watch it, do this automation, or I can just go Something else and, you know, ride a bike. And recently, I didn't know this until maybe, like, you know, two days ago, it somehow can also control the iPhone through screen mirroring. I don't know what you can do with that just yet, but I think it's a very cool experience to not only be able to control your computer, but effectively control your iPhone just through technology, like, through screen mirroring. But then you might ask, okay, well, can can can computer use control codex? And the answer is that you don't really have to. Because like I said before, these codec threads can already talk to each other, right? Which means you already have managers. So in the earlier example where I said maybe someone is giving me model feedback or on Twitter or on slack, I might take a screenshot, take a app shot. I might have the agent gather some context, triage of the right engineer, figure out how to communicate this outage, and then I might pin the thread. Right? But this Still requires me taking the app shot or taking having the agency to take some kind of action. But because threads can now talk to each other, you can really change the way you do your work. Now I have a single monitor threat, and that threat wakes up every hour and just using computer use or using the CLI reads all of slack and all of Twitter, and then it will try to figure out what are the big issues, that we have. Then instead of doing the triage in the comms itself, it will make a new thread. It will rename it. It'll pin it. And then that threads job is to figure out how to triage. That drops. That threads job is how to do communications. And then that thread is the one that does this automation. And then maybe in a in a day from now, maybe someone else also has some feedback about this issue. The main agent can say, okay. This is still an ongoing issue. It doesn't seem like anyone's addressed this. Let me send a message to this thread about the, like, rate limit outage. Right? And that thread might decide to post a message on Slack or ask a question or maybe read the docs or just see if any pull requests are Review. And these are the kind of really cool things that you can do once you understand these basic concepts around pinning these threads, enabling heartbeats and computer use, and then just allowing these agents to talk to each other. Right? So I really hope that in this quick session, I've really convinced you that the the nature of work has changed. Compaction works. Right? Generally speaking, I feel like we're pretty close to things like continual learning by just having a memory base and a pin thread. And by taking actions using plug ins and Computer use and making sure you write things down in your memory vault, you can get a lot of results in terms of how you can improve your productivity. And by just having something like a heartbeat, just by telling codecs, keep an eye on this, you can now run these automations that carry your work stream across, different sessions. And so I think if you're a busy developer or a manager and executive, this is a really great way of using codex, not just for coding. You know, if you want to keep track of everything in your little, obsidian vault, if you want to use a chief of staff Thread to manage your daily briefs or your meetings research or if you just want to use loops to basically automate this work, check out the codecs out, try some of these ideas out. And then today at 02:50PM at Track 4, we're gonna do a longer, like hour long workshop where we could actually have a longer conversation and deeper conversation about what we actually do. And I can give you guys some feedback on the specifics of how I've been using these kind of And, lastly, come say hello at the OpenAI booth. We're right in the middle. You can't miss it. It's a lot of fun. And, that's it. Thank you. And now suddenly your team is on on call rotation to Out what actually went wrong. Pretty common. Right? No. As per standard in gene response, your gut will tell you to pull the raw prompt from the telemetry logs, pass it to the same model using the same prompt, and run it locally to isolate the bug, which we'll all do. And, surprisingly, it will work as well. Run it again, it will run again. You run it 10 more times, it will be just perfect every time. But Now let's talk about that one run which costed you, and that will be gone. You can reproduce it. And if you can reproduce it, you can debug it. And if you can debug it, you can promise it won't happen to your next customer or user. Right? Now I am Tisha. I have Sushin with me as my co presenter. We both run agents against real production backends. You know, the kind of place where about right? Isn't over done it again. It's you on a call with a Customer explaining where the data actually went. This whole talk is going to be about that one thing to lose the second an agent goes haywire in production, but just being able to reproduce it. That will be a not start for the next ten minutes to follow. Now let's look at how this actually blows up. You've got an agent hooked to a broker API, which is the scenario I'm taking. The user says, hey, sell a thousand dollars of Now comes the interesting part. Instead of doing the math, the agent stealth the raw number 1,000 and dumps it straight into the quantity field. Guess what? It says 1,000 shares instead. Now at a $190 a share, a thousand dollar intent will become how much? A $190,000 disaster. Right? And the very side part is that the API or my infrared returned. A clean 2 100 okay in thirty milliseconds. We got zero exceptions, zero alerts. If you see the trade, it's completely wrong, but your dashboards are sitting there perfectly green, perfectly flawless. Then such a scenario as we last discussed comes up, what's the first thing which you will do to try and fix this? The reflex here is to, you know, just turn the model temperature down to absolute zero. Assuming Greedy decoding will make everything that the mystic. Right. But that's a complete misconception setting the temperature to zero doesn't fix a broken reasoning part. It just means the model is going to make the exact same logical error, the exact same day at the exact same time, and honestly, even worse than that. Back up the scenario we just discussed, look at the engineering threads on Reddit and Hacker News.

## Slides

### 00:14:54

### BRING CONTEXT IN VERSION 2
## Example: Feedback with Loops

1.  **Capture**
    Single Monitor thread creates new threads
2.  **Triage Skills**
    Gather evidence and decide where it belongs
3.  **Comms Skills**
    Prepare the reply, issue or proposed change
4.  **Pin the Thread**
    The report now has a durable working context
5.  **Self Maintenance**
    Heartbeats handle communication, memory maintenance, implementing feedback

### 00:15:26

BRING CONTEXT IN VERSION 2

## Example: Feedback with Loops

1.  **Capture**: Single Monitor thread creates new threads
2.  **Triage Skills**: Gather evidence and decide where it belongs
3.  **Comms Skills**: Prepare the reply, issue or proposed change
4.  **Pin the Thread**: The report now has a durable working context
5.  **Self Maintenance**: Heartbeats handle communication, memory maintenance, implementing feedback

### 00:15:55

# Chat -> Workstream
## The unit of work has changed

- Pin a thread
  - Give the workstream a durable home
- Take actions, log things
  - Save decisions, owners, blockers, and next actions
- Wake it up
  - Schedule the next check instead of starting over

### Engineering the future of AI
[Three cards, each with an icon: a pin, a pencil, and a clock]

### 00:16:27

# Chat -> Workstream
## The unit of work has changed

- **Pin a thread**
  Give the workstream a durable home
- **Take actions, log things**

### 00:16:48

# Check out the Codex desktop app
- Try out some of these ideas
- Come to the workshop on Track 4 later today at 2:50 p.m
## Come say hello by the OpenAI Booth!

### 00:17:23

# AI Engineer World's Fair
AIE

### 00:17:56

### AI Engineer World's Fair

# Your agent failed in prod.
_good luck reproducing it._

Tisha Chawla · Susheem Koul · Microsoft

### 00:18:21

# AI Engineer World's Fair
## Your agent failed in prod.
_good luck reproducing it._

Tisha Chawla · Susheem Koul · Microsoft

### 00:18:56

# Your agent failed in prod. *good luck reproducing it.*

### 00:19:20

asked: sell $1,000

### 00:19:52

temperature = 0
