# Closing Keynote — Wei-Lin Chiang — session 2026-07-02T00:10:00.000Z → 2026-07-02T00:30:00.000Z

_51 transcript lines · 0 slides_

## Transcript

And and do similar tasks like that. You can see here with the breakdown of two calls of models is that they're doing hundreds of calls and they're exploring their environment, they're viewing images, they're reading files, their writing files to do ad hoc analysis that's going to feed into the the slide output that we just saw. And this cost, each turn is output tokens, and then those output tokens flow into Input tokens in the agent trajectory and we pay for that. When we look at the output tokens to complete a task, we can see there's vast differences. You can see that Claude Sonnet five released only yesterday used over 200,000 output tokens per task. Compare that to your ChatGPT query, a couple of years ago where you might have been doing Couple of 100 tokens? Couple of thousand tokens, maybe? 200,000 tokens to complete a task. And you can see here that models vary orders of magnitude, and this is driven by two things. This is the number of turns that we just looked at, and secondly, it's the output verbosity of the model, both in terms of how much reasoning they're doing, how many reasoning tokens they're outputting to complete a task and also in completing their answer. It needs to put together that slide and all Of that detail. That takes tokens, and we pay for those tokens. But stepping back, not just at output tokens that the models output, but to total tokens that we're paying for. We have that on the left hand chart here, AA briefcase token breakdown, answer tokens, reasoning tokens, input tokens. Can anybody see any in output tokens here? They're all input tokens. Vast majority of tokens to complete long running agentic tasks are input tokens. You can barely see any output tokens there. And so, therefore, the two token prices that we wanna look at first is the input token price without a cache hit and the input token price with a cache hit. And if we remember that slide, there's vast differences between those models, and you can see that on the right chart here, which is the cash discount. For a cache hit of an input token. It's usually around 90% here, but it's also different for models and providers, whereby some models here, 99% and now others around 80%. And if we think about all the the vast majority of tokens being input tokens, you can understand that this can change by, multiples, a difference in a cash discount or a cash hit rate. The total amount of an agentic task. And so I think we're used to thinking about output tokens, but I'd ask us, let's start with the cash hit price when thinking about the cost of an agentic task and tokens. I think the last perspective we wanna share with you and wrap up with is the most important chart for understanding the AI landscape in 2026. In 2025, it was simpler. Is our intelligence index bar chart. Now we start with the intelligence versus cost per task, as we are now wrestling with these trade offs of the cost of intelligence. And a helpful archetype to understand this and to reason about how to think about cost per task, whether we should just use the most intelligent model or the cheapest model, is to break down tasks into two archetypes. The first archetype is a task whereby There's not a ceiling on how much intelligence you could want to complete the task. More intelligent equals better outputs, and this is the case for most knowledge work today in professional tasks. Not everybody agrees with that, but that's something that artificial analysis, we believe quite strongly. Think about analysis that you might do on strategy or on how we can save costs or on even writing a job description. Always be better. We can always do a better job as humans, and that's the case for models. So there's not a ceiling on that in terms of what level of intelligence we need, but we do need to trade off costs, and so the question therefore is how much are we willing to pay for the extra intelligence, and you want to look at the line here in making that decision. The second archetype of task is whereby there's a ceiling. An example is, how much did I spend on Stripe fees last month? A smarter model doesn't necessarily give you a different or a better answer to that. There's a ceiling on the task and then you want to think about what is the level of intelligence, the minimum level of intelligence that can complete the task, and then you want to choose the cheapest model, that which is to the left on this chart. So that is the cost of intelligence. We're artificial analysis. We're hiring. Thanks very much. Thanks. Please join me in welcoming the cofounder and chief technology officer at Arena, Weilyn Chiang. Hello, everyone. Excited to be, here sharing our experience, building agentic evals in arena. My name is Wei Lin. I'm the cofounder and CTO at arena. Quick intro on me. I did my PhD in AI research at UC Berkeley, where my focus was Robust, scalable evaluations for AI systems, and that will eventually become the foundation for what we're building today at Arena, to measure intelligence in the real world. Some of you, some of you may have heard, our earlier work, like LMS Judge back in, 2023, we did, some of the early study, as well as building a chatbot arena, which and some of the, evaluation. Research I was fortunate to contribute. So what is Arena? Simply put it, Arena is a AI evaluation company. Our mission is to measure intelligence in the real world beyond just static benchmark, but, the intelligence actually delivering real values to the users, the customers. And over the past couple years, we have been tracking, you know, All the major AI breakthrough. Obviously, after, you know, the Charge GPD moment in 2022, after that, it was GPD four turbo, GPT four o, having the breakthrough in chat and multimodal capability, and then evolving to, the reasoning model, thinking model with, OpenAI o one. And in 2025, we saw the image Generation breakthrough of NANA banana, which was originally, started testing in arena as a codename, before its public release. And we are also seeing GROC catching up, GPT images two recently released, to become, you know, the current frontier of image, models. As well as, you know, the video AI. Operations, Beal, and recently, ByteDance, CDENCE. So towards the end of twenty twenty five, when Opus 4.5, 4.6, went from being a great coding model to a generally agentic coding model that can do longer horizon, tasks. That also showed up, in Arena 2, that where we measure in core arena. We see, you know, An improvement over the past generation model. And the most recent FableFi breakthrough, where we measure in Asian arena, we will talk a little bit more later. As well as the most recent GLM 5.2 release, which is, like, really a big milestone, for the open source model community. So, we have at Arena, we have done this with scale. We now see 10,000,000 monthly visitors going To, our product, arena .ai, and we have collected 700,000,000 conversations across all the modalities, text, vision, image, video, coding, these days, agentic. And we have hit a huge milestone. Very excited to share that just we just recently announced we hit a 100,000,000, annualized revenue in just eight months after we first released our evaluation product. We are also ranked among the top Gen AI product globally by a unique number of monthly visitors according to a ACES in the analysis. So, the, topic I want to cover today, and the core of what we are offering, is life leaderboard, which is based on real world evaluations, powered by the 10,000,000 users. 700,000,000, traces to rank all the top AI models from tier models, for the past couple years. And we cover text, image, video, code, agent. So really wanted to build a, leaderboard that can help everyone to find the best model for their use cases. And it's free. It's available for anyone to see to use at arena.ai/leaderboard. You can see all the other Takes their Pareto frontier comparing cost, performance, you know, use cases, different category, different modality of these models capability. So, yeah. So the real problem today I want to talk about is to share the experience how we how do we evaluate agents. I wanted to share our first hand experience. In the past couple of months, we've been building, the agentic eval, which is very, very different from the, you know, past in the past, we evaluate chatbots. Wanted to share some lesson here. Before we dive in into, the details, first, why does this matter? I wanted to talk about the trend. So we have been seeing, the very rapid shift from, the chatbot to agent, paradigm shift. And if you look at the OpenAI's data on codecs traffic, the share of the output token coming from agent has just skyrocketed. And you can see inside OpenAI essentially 100% of the, output token strong agent from codecs. And for other organizations, you know, average is like above 60% now, and individual also climbing very fast. So, there's no question that the token flow is now driven by agents. And we also see that agents are not just for engineers, right? It's not just for software engineering. If you look at codecs adoptions by Department at, OpenAI. Engineering, obviously, 99%. But also finance, recruiting, legal, and so on, they are all, like, almost, like, 90%. And as so as well as you can see, you know, the studies from Goldman Sachs estimates the monthly token usage is also skyrocketing towards, like, you know, 60 quadra a quadrillion tokens in the next couple years. So really, you know, the economy Also tell the same story. If you look at the RAM data, the AI spending is getting closer to people spend. Right? So if you see, like, you know, the top 1% of the company's monthly AI spend is per employee is actually already, like, 7.4 k, roughly half of the salary of software engineer. So this is really, like, you know, historical shift that, meaning also the stack of, like, choosing the best. Model the right model, and optimizing your agent take AI workflow is, you know, more it has never been more important. So the key question here is, like, we give agents lots of autonomy. We spend a lot, we invest a lot. And the key question here is, like, how do we actually measure agents' outcome? So that's really the bottleneck, right? You want to understand the value of these agentic output. Actions. And this turned out to be a pretty hard technical problem for a few reasons. First, agents are multi component systems. Right? You have the model, the agentic loop, the tool, the harness, you know, any of these pieces can break the system. You also, have agents operate through complex workflow now in a real environment. You're building a device. Doing research, producing document, slide deck, and so on. So it's, like, more involved task. And third, the signals that we can collect, you know, in this trajectory are also becoming sparse, spread across longer horizon. You know, a task may take 100 toll calls to finish, right, before you know if it's succeeding or failing, or you give any feedback of a chance to steer it. To deeply understand the problem, at Arena, we decided to actually, first hand building real world, you know, agentic product and app to actually source the organic traces and feedback from the actual users for us to, you know, do research and deeply understand that. So last month, we launched, agent mode in arena, to allow anyone to go to, you know, arena to experience and evaluate agentic capabilities. It's right now available for everyone to use. And wanted to show you a very quick demo, if if I can start the, is the video moving? Yeah. Okay. So this is agent arena. You go to agent. You go to arena.ai, you you choose the agent mode, and this is a real world, you know, agentic product. You can go and evaluate model. You come in and type any question you want. In this case, it's like I ask. Download Google's q one earning report, and create a slide deck summarizing these all output in PowerPoint. And you can see the agent goes off and and doing work, searching the web, pulling the right website, start structuring the deck, and then using some of the bash tool, writing Python code to, generate the the slide deck. Right? And you can see that at the end, there's like an artifact. Generated by the model, that user can download and see. And this is, like, a, you know, a real PowerPoint, that outputted by the model. And then user can at the end, we ask every turn, like, we ask, was this task successful or not, and user can provide feedback that way. And that's one of the signals that we use to evaluate and understand whether agent actually delivers the outcome. So, yeah, this is just to highlight the panel. And Under the hood, how we build the agent arena, it you know, we give model set of tools, file systems tools, rewrite, edit, and so on, and search, web fetching, image, generation, speech as well, recently added. So just really giving the model tools similar to, like, a cloud co work like harness, and also terminal access to run code to to to to, you know, do work. And we also are adding More and more, connectors soon, like GitHub, which can connect to your repo to, you know, do more serious software engineering task. And you can see this plot is the the usage of these tools, in a in a time in a one week time frame. You see 5,700,000 tool calls. You know, Bash was the, you know, the number one used around 46%. And these agents are actually using these tools to do real. Real work for users. So we also, you know, dig into the data and seeing users are, you know, pushing really hard to, trying to do more harder and complex task. So real session, we've been seeing, like, you know, users are building, you know, a movie, watch list app, debugging a control systems for autonomous, you know, vehicle and, architecting, building a rack pipeline. No. Implementing features in microbe and so on. So these are the sessions, like, go over hundreds. Some of them go hundreds of turns and a couple hundreds of toll calls, very serious stuff. And you can from this, you can tell that the, the agent that we built, at Arena is actually doing real work with users and giving user real value. And we believe the best evaluation should be, grounded and measured in real world use cases like this. So we launched AgentReliance. Just a month ago. And the first month over, we collected over a million agentic traces, and these are in task spending, coding, research, document, brainstorming, planning. And we see more than half of these, traces fall into work related category, more like towards professional use and complex tasks. And we have seen Asian also written, more than 50,000,000 lines of code on ARENA. Python, Markdown, HTML, JavaScript, and so on. Oh, this is the tool distributions that you can see. The coding is the number one. And some of these, tasks you can see is some of them are more complex using more tool. Some of them use less.

## Slides
