# Trends in AI — George Cameron, Micah Hill-Smith — session 2026-07-01T23:50:00.000Z → 2026-07-02T00:10:00.000Z

_6 transcript lines · 0 slides_

## Transcript

Also gives us a perspective on the progress that's been made over the last couple of years. On this task, which is a commercial due diligence task, GPT four o presents a pretty basic slide. O three, a breakthrough model that was released early last year. Thinking about that o three was only last year is crazy to me. You can see that o three produces a few bullet points, helpful, but not what we would expect of ourselves in complete This kind of task. And so this shows us the progress that's been made when we look at Opus four point eight's output and Fable five's output, which goes a lot more in-depth depth in terms of analytical rigor and presentation quality. So let's look at how models completed this task and what it cost. If you remember Micah's slide, He showed that some models are taking are using over $20 worth of tokens, to complete these tasks. And so let's look at the drivers to learn a bit about the costs of agentic tasks. Four drivers to look at, and the key drivers here are token price, the number of turns in the agent trajectory, the token efficiency and usage of models, and last, but potentially most important, the impact of prompt caching. Taking a look to start with the prompt with the token prices. What we can see as a first takeaway here when looking at the cash hit rate, token price, the input not considering a cash hit or without a cash hit price, and the output token price, firstly, is that there's orders of magnitude differences between the model. This is a critical driver. There's order of there's two orders of magnitude Difference in terms of the token price between frontier models like Claude Fable five and still good, very usable workhorse models like deep seek v four Flash and GPT OSS one twenty b. The second takeaway here is the difference between the individual token or the types of token prices. You can see that there's vast differences in the cash hit price and the input token Without a cash hit price and the output token price, and we'll get to that impact later when we look at token usage. Next, these are long running agentic tasks that we are now asking of models, especially in realistic environments where they need to navigate all of these thousands of files to get to an answer. And models are doing that. They're starting to really explore the environment, actually, similar to humans when we search Slack.

## Slides
