# Evals Track Intro — Laurie Voss, Aparna Dhinakaran — session 2026-07-01T17:25:00.000Z → 2026-07-01T17:30:00.000Z

_17 transcript lines · 9 slides · source: full recording_

## Transcript

Humans are incredibly biased in what we feel is a correct solution. I mean, we're the result of an evolutionary training. That help us survive in the jungle, right, not doing quantum computing. So I think that even though we can be brilliant and innovative, there are a whole bunch of progress and breakthrough that can be done which we just cannot see or perceive. If I had more time, I would give some examples, but I think that's one of the thing where ML is such a different viewpoint on many of those problems that we're going to get the, oh my god, this was in front of us the whole time and we could not see. So exciting times ahead. Thank you very much. Ladies and gentlemen, as we continue today's program, please welcome back your MC, developer advocate at IBM, Tayos Kumar. What an incredible start to the day. Everybody's leaving. This looks amazing from here. Before we break off, or after, after, let's take a moment and acknowledge the sponsors. Honestly, this would not be possible without them. We're gonna get the slides up. Listen, you need to give them your biggest round of applause. I mean, it is so cool. Thank you. Thank you. Thank you. Thank you. Microsoft, thank you to all the other sponsors here. This event would not Possible without them. There's plenty of other things happening, in the other stages but there's no doubt that, evals are a huge deal in AI. In fact, they're the gate of quality. Right? We can ship a lot of things but if they're not evaled well, we ship a lot of slop. And so, our next discussion, our next session is gonna be from Aparna Dinakaran from Arise who's gonna talk to us a little bit about eval's. Please your biggest round of applause for Aparna. Please welcome to the stage cofounder and chief product officer at Arise, Aparna Dinakaran. Hey, everyone. Can you all hear me? Alright. Let's go. Oh, let me go one back here. Awesome. Well, hey, everyone. My name is Aparna, one of the founders of Arise. We work with some amazing teams to help them build Evals, and we have an incredible lineup of talks for you all today at the Evals track. It's happening in Room 2005, and there's gonna be amazing speakers from Turnbench and Uber and Snorkel kind of all happening after this. But today, I'm here to talk to you about the future of evals. Evals have gone from the new skill that every PM and every AI Has to learn to the thing that every serious AI team is betting on. We've been really fortunate to get to work with some of the best AI teams in the world. So we get a front row seat into not just what's happening when they're building their actual agents and before they actually ship, but actually the evals that teams are running on their live production agent via their traces. Little bit of some stats for you guys. We run over a 100,000,000 evals every month. The average Team runs about 12 different eval jobs with the top teams running over 3,800 different evaluators. And offline evals, online evals, they each have their own place. But today, what I'm actually gonna talk to you about is the teams that are running evals on their traces. This is actually what's helping teams figure out what's working, catch their failures, and that's the type of data you need to fuel your continual learning loops. And the industry kind of Agrees. I mean, all the CPOs of Anthropic, OpenAI, all you know, GDB, you have Gary Tan saying, evals are everything you need. And the whole industry kind of agrees. So we added evals. They catch all the failures. Right? Here's the problem. While we were building all of these first gen evals, the thing that we were actually evaluating has changed underneath us. In 2023, it was about just answering a prompt. In 2020 Four, we started to see all different tier models. They've added tool calls. They've added reasoning. They've added deep research. Now what we have is teams running loops on real world data with sub agents kicked off on, long horizon tasks.

## Slides

### 00:00:20

# The gold we cannot see

### Research to Reality with Google DeepMind
Benoit Schillings / Vice President of Research Google DeepMind

### 00:01:14

# IBM
[IBM logo]
[Partial text "AI" visible on the right side of the screen]

### 00:01:51

# Participating Organizations / Sponsors
- AGI Lab
- World's Fair
- Microsoft
- OpenAI
- together.ai
- Akamai

[A grid of company logos and

### 00:02:17

- Amazon AGI Lab
- AI Engineer World's Fair
- together.ai
- Open...
- W...
[Colorful square logo, possibly Microsoft]
[Small dog

### 00:02:50

# APARNA DHINAKARAN
## CO-FOUNDER & CPO
### arize
[Ornate compass or astrolabe-like object in the background]

### 00:03:25

- AI s Fair
- ther.ai
- C AI
- AI Worl
- OG
[Grid of company logos or event names, including "AI s Fair", "ther.ai

### 00:03:50

# AI Engineer World's Fair
- OpenAI
- Akamai
- Datadog

[Logos for OpenAI, Akamai, and Datadog]

### 00:04:20

# Evals are growing quickly

*   **100M** evals run every month
*   **3,800+** eval jobs run by the top AI teams

### 00:04:49

# Agents got more complex

- **2023**
    - Answer a prompt (GPT-4)
    - Call functions & run code (Function calling, Code Interpreter)
