# Notion's Token Town — Sarah Sachs — session 2026-06-30T21:50:00.000Z → 2026-06-30T22:10:00.000Z

_82 transcript lines · 24 slides · source: full recording_

## Transcript

Before I get started, you guys, this is a huge keynote room. Can everyone, like, come forward? Because I'm talking to, like, four empty rows and dispersible Me a favor. I'm spending thirty minutes telling you all of our secrets. I can see you still. Thank you. Thank you. Thank you. Thank you. We're just gonna chat. It's a giant room and there's 500 of us. This room is way larger than that. Thank you. Honestly, I knew you guys had it in you. It's really not so hard. Thank you. I also sit in the back. I also work during talks. I get it. I totally get it. I did all day. Not for me. Okay. I'm gonna start, but I'm gonna still point at you if you're in the back like you. Okay. I'm Sarah. I'm, lead our engineering teams for AI at Notion. Welcome to my talk. It's about token town. How do you go from not go from AI build to AI poor? Okay. I know that today is all about software factories. We're gonna talk about that, but we're gonna talk about how to do it sustainably. This is me. This is on my first day at Notion in a very sweaty subway. Like I said, I lead Our AI teams at Notion, and I negotiate AI contracts for a living. My team jokes I act like Anna Wintour, so this is a nice a nice image of me with AI Anna Wintour hair after a press article referred to me that externally. And that's kind of the idea. Right? How do you think about negotiating between different vendors, making sure that you maintain taste for your company? I don't do it Alone. This is launch day at one of our recent launches. This is just a subset. Any good engineering manager points out that we have a whole company of people building this. I'm just the one that gets to come talk to you about it. So we've been building a lot. This is, an example of our AI usage, just in 2026. And we've been really proud of how we've been able to grow that usage, and I'm gonna talk to you about how you can build an AI native product and an AI native company. But this is just to give me some credit that that we're doing it kind of well. Okay. It's always been that durable system of record. It's always been the place where you can collaborate with your peers. But today, that point of collaboration is a little bit different. It's not just humans. Notion's always been the place for collaboration, and today, that collaboration happens between humans and agents, humans and humans, agents and agents. So for those of you that don't know, Notion And we like to think about AI transformations going through this journey, and I'm sure some If you are looking at the slide and wondering where you are, AI as a thought partner is when we all started tinkering. We all started just going to the very first version of ChatGPT on Thanksgiving when it came out three years ago, four years ago. And we started saying, like, how can I send this email to my landlord to say that I shouldn't pay for repainting? Right? Then we'd copy paste it, enter it into our email. Eventually, we started getting to a place where we could use AI like an assistant. AI was able to maybe execute individual tasks. That's how Notion AI really took off in the beginning. And it was able to save employee time, but functionally was limited in its capabilities based on what humans asked it to do. AI as teammates is what we were really excited to launch almost a year ago now, but this is true in many products where you can do repetitive work and think about a process and have AI do that process. What I think is really interesting is when AI actually becomes that critical workflow where processes are interfacing With each other and you have entire systems running. How many of you guys feel like you have AI as a system down? Aren't you sad you came up now? I'm kidding. Great. None of you. Exactly. We have found that no one has figured out how to do this well. Eighty eight percent of people can't even get past AI as an assistant. And why is that? We have a thesis at Notion. It's because there's too much siloed data and not a durable system of record for that point of collaboration. And we believe that for your soft For a factory to work, for your company to work, and for your systems to work, you need that durable system of record, and that is Notion's mission. So doing that is expensive. You see a lot of companies that try and commit themselves to this vision, and these are just a series of headlines all within a week, of how that's painful. So you can put all of your money into a process to try and make a system, and you end up feeling like this. Right? You end up using a blowtorch to light what is actually a large cigar, but you kinda get the idea. Cost is a structural barrier to entry. It makes it hard for you to serve products. It makes it hard for you to build factories. And it is ultimately, I would posit, one of the largest reasons why things do not happen at scale successfully today. And I would argue for anyone working in an applied AI company, it's something for them to be really familiar with to understand the trade offs that they're making to build Durable and exciting and enlightening product for their customers. But that's not really how the market is today. Right? I'm not gonna name names here, but you guys have search engines. You can figure it out. Exhibit a, a reasoning model gets upgraded. Amazing. The per token pricing is the same. What's not to love? You try it out. It uses three times as many output tokens. Right? Exhibit b, a model gets upgraded, but it has an entire new digit. Right? Whatever Markation system that model family likes, it's brand New. It's 40% more than its predecessor which is being deprecated in the next four months. These are real scenarios that we face at Notion. All of you are nodding because these are common pretty much monthly now. But here's the problem, are you growing 40% in that time period? Are you making 30 three x more revenue? No. So how do you navigate the system? If you just Auto upgrade your model and everything that you're doing, you're you're giving someone a bad deal, either your customers or your investors, depending on how you charge and where you get your money. Neither are good. Fortune 5,000,000 companies have the capability to navigate this. They can hire large consulting teams, have durable teams on their own, and build expertise on how to navigate these trade offs. Most people don't. Everyone else has no ability to negotiate with leverage and they're stuck in these scenarios. Right? Part of my job is that Anna Wintour joke. Is to think about advocating for the fortune 5,000,000. The non fortune 500 companies that don't have the mass to have leverage and negotiate but need to think about how and I'm gonna share some of the lessons that I've learned when I have kind of large amounts of traffic behind me that I think scale to those who don't. This is probably less of a secret now than it was When I started giving talks like this, maybe four months ago, your supplier is your competitor. I know very few people who have convinced me that that's not true. You will always be getting a bad deal on tokens with someone who builds them natively. Right? Sometimes the cost of goods served is extremely different. You're basically they're serving a first party product and then you're buying those tokens at a huge surcharge and then selling them again at another surcharge. That's not really value you can defend. You're getting a really bad deal and if you tie yourself to one provider you have no exit. If you build an AI product that you're selling with this structure, you are crossing your fingers and hoping that you are a viable business. I do not encourage that. This is really interesting. Dylan in semi analysis posted this. I think it's it says eight hours ago it wasn't at this point. It was probably a month ago. They purchased a subscription plan and they just highlighted, right, how Different what Frontier Labs charge customers for first party products are versus what they sell. It's a bad deal. Don't play this game. Or try and let me know how you win. I don't recommend. Think about everyone else. Think about what that structure means and where you have expertise. I don't think that that's winning on the token economics. I think it's about product. It's about building data flywheels and understand Your customers better than anyone else, understanding when you need capability, when you need low price, when you need latency improvements. I promise you, you don't always need what is usually the slowest but the most capable model out there. And then build compelling UI and orchestration, and I'll show you some examples of that, to justify the cost on the bad deal tokens that you do resell. The job is not to train I mean, some of you might be training the best model. I'd love to serve it and come talk to me afterwards. But most of you are not doing that. Stop trying to win that game and think about the best product that uses many models. Help your customers, help your team bet on the frontier, not on the lab, and we'll talk about what it looks like to do that. This cost per capability per second trade off is actually really intense. Citadel came out with this memo, a while ago, maybe two weeks ago. I loved it. The idea Is that for the economy at large, simpler models might be the most cost effective productivity augmenting pathway. They talk about this bifurcation on frontier versus everyday usage. I really believe that. And for every product, the definition of frontier versus everyday, the definition of saturated capabilities or model capability overhangs depends on your expertise on your product. No one can replace that. And not all traffic is equal. It is a huge miss to send all of these to the latest Opus model. Some of these absolutely. Data analysis when you do it on Notion, we'll recommend Opus. Right? When you triage an email inbox, if we're charging you to do that on Opus, we're ripping you off and ourselves. Think about where your traffic patterns are, And then think about how Frontier Lab model providers are structured today. I mean, it's functionally an oligopoly, right, and that's fine because they're racing to the top. And I think the top is really hard and really important. This is not to say that products don't have a place for frontier difficult tasks. I want everyone to nod and understand that's not what this talk is about. Understand when you need those tasks, and it's not everything. The problem with those tasks are is keep in mind how pricing is incentivized. You can figure out who these players are. Either you are the best model. Everything above what AI can't do today is your market, you can basically price it as high as you kind of want. If you're slightly behind that best model, all you need to be is like A dollar per million tokens cheaper, and you have the rest of the market. You know that economic theory about gas stations where the best gas stations are the ones that are right next to each other because they cover east and west the most? Yeah. It's the same with model pricing, which means that price does not correlate with capability growth. So for this complex task, understand what capabilities you need. But be the expert on what complexity is. And keep in mind that who handles complexity changes. Oftentimes, you'll see applied AI companies really be super outspoken on marketing with a specific lab. That's always kind of a red flag for me when they're not model agnostic because if you look at this graph, it basically shows that they're behind every month. Right? The new model and the new model provider of the best frontier capabilities change. And if you hit your ride with one particular provider in exchange for, for instance, a larger discount You're doing a disservice to your customers, like, half of the time. Right? So really think about if that discount is worth not actually having a Frontier product. And remember that that optionality is your leverage. If you don't have the capability to walk at any point, you are stuck. And again, I think that's probably the most expensive decision you'll make regardless of what discount you get or the engineering Work to have model interoperability. One option to navigate this is to stay model agnostic. Have different models and capabilities in your system so that at any point, if pricing seems unfair or untenable, you are not out of business. Notion's auto model does this really well. We have state of the art models available always, but we also have an auto model there at the top that handles about 75% of our traffic. Right? We have the ability to switch between models in our product And we also offer it to our customers that they have access to these models without vendor lock in. That's part of our AI Switzerland approach. You guys love taking photos of slides. This is the slide. Okay. Model agnostic playbook. This is how you do it. Build for multimodal. It is hard to kill the cash and switch models mid transcript. I understand that. We invest in that technology. It does Even have to be per thread. Just think about your harness as model interoperability. Think about the cost per capability per second, not just the tokens. Here's a great example. We posted this review when we, announced our partnership with Parallel as our web search provider. If you were to look at just latency of a single call or just cost, Parallel might not be the cheapest. But if you have expertise in entire web search trajectories, you'll see how it differs. The granularity of this eval is what lets us The best decisions for our customers because we understand all of the trade offs on entire trajectories, not just single calls. Switch fast and often. I think we talked about that. And give them something back. That expertise on use cases is also very valuable to Frontier Labs. We find that our evals and our early access program partnerships actually help us a lot with Frontier Labs and is something that we can exchange instead of extraordinarily large commits, and I don't think the discount is ever worth the lost in optionality. That's a perspective you can choose to keep or not. The second option is Moderate tasks. Understanding open weights place there. Open weight models are really strong enough to handle these tasks, and the possibility to RL on top of them has also kind of expanded the upmarket growth that they can cover. I view open weight models as basically lowering the barrier to entry on cost for our customers, and they also give you negotiation leverage. So it's kind of a Alternative that's putting that downward pressure on pricing that if there's an oligopoly of two or three providers at the top is unavailable right now otherwise. I think CHEMI two six is probably the first time that we really saw a model that outperformed five two GPT five two, GLM five two now is another five two. Bombshell in the villa that also probably does best here. But it's no longer the case where open way models are good for just SFT on small tasks. Really think about, without RL, if they're capable enough for what you need. Don't just think about external benchmarks. Be able to have expertise on your system. What are your tool errors? What's the actual latency that you need? Right? And again Here's an example of a benchmark that we posted. It's a little bit stale on purpose. Right? But you get the idea. Philip at base ten showed this slide once, and I've stolen it ever since. Thank you. Are you here? Buddy. Okay. We'll chat. Hi. Well, he could come up and say it better, but the idea is that you don't have to be at the top. Right? I'm not trying to make a case that open weight is the best model out there. The case being made, however, is that, the gap gets covered eventually. So if the tasks that you're having today are good enough, then in six months, they're probably covered by open weight. So be prepared now. And the last thing is CPUs over GPUs. We've we've recently launched Thing at Notion called workers. I don't think that the GPU was necessary for every job. A lot of the jobs that we have are actually serving, discrete pieces of code. Like, you don't need an LLM to turn a CSV into a PDF. You don't need an LLM to talk to Notion tool calls if we have a CLI. You definitely don't need an LLM to do deterministic SQL queries. This is where people become token poor very quick. And I think the last option here besides open weight CPUs and optionality is actually governance. There's a lot of AI governance. One is visibility, understanding who's using the data, understanding its maintainability and control. When you have model optionality, you can offer a lot more to your customers. Here's an example of how that governance works in Notion. So final tips again: think about architecture, think about open weight, and build value that transcends tokens. So we're gonna depart Token Town. I know I said welcome to Token Town. We're going to spend the next ten minutes really thinking about what to do next. So I think the challenge of the next six months doesn't have to do with capabilities. I think it has to do with security. Let's start there. There's this concept called the lethal trifecta. Simon Wilson, I think, crafted this. If you have access to private data, exposure to untrusted content, whether it be through ingestion MCP email. Right? And the ability to ex to communicate externally, and that can include, like, payloads in a web search. The second you have that system, you're exposing risk. And in fact, the more autonomous your system is, the more unsupervised this risk is. I think that this is what builds valuable product, not just capability. Same with sandboxes and computers. We talked about this, but it really is something that builds better determinism in your product and also better token For your customers. In multi agent orchestration, understanding what agents see and do and what persists. I think persistence of enterprise knowledge is something that's actually really not discussed enough. It's starting to be with some recent launches. You know oh, there is audio. So don't have your workflows look like this, and I think this is where most software factories are today. Right? It's like actually your entire engineering time. Just spends time babysitting the factory. Right? I mean, I get it. Ours started off like this. Agent orchestration is one of the most difficult tasks of making factories work. So, okay, this is me telling tPain to tell people to buy Notion AI. And the reason I included this slide is I am going to sell Notion for a second. It's my job. Always be closing, always be selling, always be hiring. Come find me. But I'm going to talk for a second about how Notion does this. Today, we already have the ability to inspect tasks. And you can imagine any task that you look at, in a Notion document, you can have Claude actually go ahead and scope out what you need. We've launched this manage agent capability today. So if I go ahead to the top of this task, I can actually ask Cloud Agent to scope out the task. Right? Ideally, it's working, and you'll see it'll actually populate, an entire spec of what needs to be done.

## Slides

### 00:04:51

# Engineering the future of AI

### Uber burned through its entire 2026 AI budget in four months. Now its COO is questioning whether it's worth it

### One company spent half a billion dollars on Claude in a single month: Report comes as AI costs climb

### Microsoft reports are exposing AI's real cost problem: Using the tech is more expensive than paying human employees

[Screenshot of three news articles discussing the high costs associated with AI implementation]

### 00:05:15

## Cost is a structural barrier to entry.
[Line art illustration of a coin with a face]

### Engineering the future of AI

### 00:05:52

# The token market feels structured against buyers

## Exhibit A
- A reasoning model gets upgraded at identical per-token pricing.
- **Upgrade uses ~3x more output tokens** for certain tasks.

## Exhibit B
- A successor model makes significant steps in reasoning.
- **Costs 40% more than its predecessor.**

# Engineering the future of AI

### 00:06:15

# The token market feels structured against buyers

## Exhibit A
- A reasoning model gets upgraded at identical per-token pricing.
- Upgrade uses ~3x more output tokens for certain tasks.

## Exhibit B
- A successor model makes significant steps in reasoning.
- Costs **40% more** than its predecessor.

### 00:06:48

# The Fortune 500 has dedicated AI teams
Everyone else negotiates alone, with no leverage

### Engineering the future of AI

[Image of a Fortune magazine cover, featuring "FORTUNE 500 GLOBAL"]

### 00:07:14

# The Fortune 500 has dedicated AI teams
Everyone else negotiates alone, with no leverage

### Engineering the future of AI

[Magazine cover for Fortune 500 Global]

### 00:07:45

# Your supplier is also your **competitor**

- Lock in with no exit
  Tie yourself to cheaper provider today. Prices change and your left with nowhere to go.
- Value you can't defend
  You must provide additional value to supersede the obvious "bad deal" you have on tokens.

Engineering the future of AI

[Illustration of a red folder or file icon, broken in half]

### 00:08:15

# Engineering the future of AI

> SemiAnalysis @SemiAnalysis_ · 8h
> Recently, we purchased one of each Anthropic/OpenAI subscription plan and randomly ran long horizon coding tasks

### 00:08:40

# Negotiate for the **Fortune 5 Million**
## Engineering **the future** of AI
[Illustration of a person holding up a block with the letter 'N' on it, radiating light]

### 00:09:12

# Win on the product, not the token

## Data flywheels
- Reinforcement fine-tuning on open weight.
- Increase your product's own intelligence
- Reduce dependence on frontier pricing.

## Product moats
- Compelling UI, orchestration, architecture, and integrations to justify the cost.

Engineering the future of AI

[Illustration of two browser windows: one with a wavy pattern and a gear icon, the other with a bar chart and a bouncing blue circle]

### 00:09:47

# Engineering the future of AI
- ~~Train the best model~~
- Build the best product that uses many models

[Illustration of a person holding an open book, gesturing towards the text]

### 00:10:20

# Citadel Securities: Tokenomics

> "We do not think this implies that the frontier of inference-intensive AI will be abandoned, only that it is **likely to be concentrated** among

### 00:10:48

## Providers are incentivized to pursue one of two paths

Best reasoning model

Worse model, priced as close to the top as possible

### Engineering the future of AI

[Diagram illustrating two

### 00:11:20

# Providers are incentivized to pursue one of two paths

Best reasoning model
Worse model, priced as close to the top as possible

## Engineering the future of AI

[Diagram with a central vertical axis, showing two diverging paths: a white arrow curving left towards "Best reasoning model" and a red wavy line with an arrow curving right towards "Worse model, priced as close to the top as possible"]

### 00:11:47

# Providers are incentivized to pursue one of two paths
- Best reasoning model
- Worse model, priced as close to the top as possible

## Engineering the future of AI

[Diagram showing two contrasting paths: a simple upward curving white arrow for "Best reasoning model" and a complex, wavy, spiraling path with red arrows for "Worse model, priced as close to the top as possible"]

### 00:12:13

# The "best" frontier model changes fast

## Frontier Language Model Intelligence, Over Time

Artificial Analysis Intelligence Index v4.0 incorporates 10 evaluations: GDPval-AA, r*-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity's Last Exam, GPQA Diamond, CritPt

- Alibaba
- Anthropic
- DeepSeek
- Google
- Kimi
- KwaiKAT
- LG AI Research
- MBZUAI Institute of Foundation Models
- Meta
- MiniMax
- Mistral
- OpenAI
- TII UAE
- xAI
- Xiaomi
- Z AI

### Engineering the future of AI

[Line graph showing Artificial Analysis Intelligence Index over time for various language models]

### 00:12:42

## Optionality is leverage
[Illustration of a stack of blue and white building blocks, with a cursor-like arrow placing another blue block on top]

### 00:13:20

# Option #1: Model Agnostic

- **Assignee:** Codex Agent, **Task:** Implement empty-state
- **Assignee:** Cursor Agent, **Task:** Bugfix for homepage

### 00:13:43

# The Model Agnostic Playbook

1.  Build for multi-model
    Run infrastructure across all major providers. If you can't walk away, you have no leverage.
2.  Evaluate on value, not tokens
    Evaluate based on cost-per-capability-per-second basis.
3.  Switch fast, switch often
    Switch every 2-3 weeks as new models drop and tools require different functionality.
4.  Give frontier labs something back
    Provide detailed eval scorecards and feedback on why we switched. Uniquely valuable.
5.  Forgo discounts for optionality
    Short-term margin isn't long-term lock in. Optionality leads to more customer trust and growth.

## Engineering the future of AI

[Stylized icon resembling a face or a letter 'T' in a circle]

### 00:14:20

### Here's a high-level summary of the findings from [SOT] Web Search Eval Readout.

| Provider | Agent prod-log Recall@10 (% queries w

### 00:14:49

# The Model Agnostic Playbook

1.  **Build for multi-model**
    Run infrastructure across all major providers. If you can't walk away, you have no leverage.
2.  **Evaluate on value, not tokens**
    Evaluate based on cost-per-capability-per-second basis.
3.  **Switch fast, switch often**
    Switch every 2-3 weeks as new models drop and tools require different functionality.
4.  **Give frontier labs something back**
    Provide detailed eval scorecards and feedback on why we switched. Uniquely valuable.
5.  **Forgo discounts for optionality**
    Short-term margin isn't worth long-term lock in. Optionality leads to more customer trust and growth.

## Engineering the future of AI

[Stylized icon resembling a face profile or intertwined letters]

### 00:15:16

# Open-weight models are now strong enough to handle these workloads

## Cost lever
Stop overpaying for capabilities you don't need on routine tasks

## Negotiating leverage
A credible alternative puts downward pressure on frontier pricing

### Engineering the future of AI

### 00:15:49

## Kimi 2.6
### Gigantic industry shift

> Akshay Kothari (@akothari) tweeted:
> "Kimi K2.6 just landed in @Not

### 00:16:13

# Engineering the future of AI

eli(as) @earlierism · Apr 25
We've often been asked about how we eval each new model at Notion, since we'
