# First Steps Toward Automated AI Research — Richard Socher — session 2026-07-01T17:45:00.000Z → 2026-07-01T18:05:00.000Z

_52 transcript lines · 39 slides · source: full recording_

## Transcript

Ever an agent tries to execute code or tries to create files. It does so in a sandbox. So we gave our agents. Alright. Alright. Hello, everyone. Really excited to be here. It's a big room. Very, very cool conference so far. I want to talk to you today about something that's been on my mind for many, many years. This is actually the first time I talk about it, sort of my version of going to Mars, and that is the Eureka machine, a machine that will eventually invent pretty much All future inventions for humanity. And the way we're gonna get there is, by taking a step back and thinking about what else has given us a lot of really incredible inventions, namely evolution, and how that leads us to automating research and pushing the scientific frontier forward. And this is joint work with a lot of amazing folks, at recursive,u.com, and even some folks at AIX Ventures. And some of these slides are Actually inspired by, and taken, partially from one of my co founders at Recursive Tim Rochteschel. So, why do I talk about evolution and why is it so important? I think, basically, evolution is this, like, open ended process that has gotten us to a lot of different things that we really like. It started in biology, it's moving to science, technology, and eventually AI. And I think it can inspire us in a lot of different ways to build better AI systems. As well. In fact, whenever we take out and there's this famous saying, whenever I fire a linguist, my accuracy goes up. I think that's true for machine translation back in the day. And it may be true that we should fire all the AI engineers, and that are here, and have them mostly manage an actual AI engineer that is AI and works on AI. And so that may be, one of the conclusions of this talk. And I think most of us are gonna be excited about it because it means that we'll all become managers of such an AI rather than having to do the nitty gritty ourselves. Alright. So let's start with evolution. Right? The really, really big picture, three and a half billion years or so. This is kind of the incredible process, that has led from, you know, simple bacteria and plants and fish and amphibians and so on to, after many billions of years, us. Right. That's that's a good starting point. That gives us some indication that evolutionary processes can do pretty amazing things. Right. But now let's zoom in and, go maybe down to a few million years. There, we can also see how in the very first primitive ways technological evolution has basically increased the world's sort of product in terms of monetary value. It's a little bit harder to estimate in the beginning, but we can see these Of exponentials and most exponentials eventually become s curves. They flatten out. But humanity has done pretty well by basically developing many of these very basic technologies, hunting, farming, but then also thinking about science, the scientific method, in the early days of the enlightenment and of course the industrial revolution. So now we can zoom even further, and no worries, we're eventually gonna get to nanochat and actual auto research and what we're doing. It's a very, very Zoom. And now we can zoom down to the last few thousands of years and what we're seeing there is that with more technology we were able to sustain more people. Right? So, when we're working on pushing that frontier forward, we're very certain that that will lead to more human flourishing. Right? And especially in the last few, hundred years we're seeing this incredible explosion in the population of people because of technology and the evolution. That it brings. And in many cases, that evolutionary process is run by us, so it's sort of conscious, but there are sort of interesting, inspirations that we can take from that as we're thinking about the evolution of AI in the next cycles. In fact, and I might not agree with everything with Marc Andreessen, but he is very smart and we agree on a lot of things. And so I think he wrote this really great techno optimist manifesto in which he, I think, correctly points out that the only perpetual source of growth. For the entire economy, a lot of people worry about AI taking jobs and things like that, but the truth is it will very, very likely increase the economy massively and that will benefit benefit a lot of us. And so the perpetual source of growth is technology. In fact, we can go even further and say that there's no material problem, and again, it's not sort of psychological problems and things like that, but no material problems, that cannot be solved with even more technology. Right? If you have a problem of starvation Into the green revolution, darkness, light, cold, indoor heating, heat, air conditioning, and the list goes on. So I think we can kind of realize that this evolutionary process has been going on for a very long time and continues to make a huge amount of progress. In fact, the progress is so fast that there can within one lifetime be a major, major shift. Right? If you were born in 1900, then three years When you're three years old, the first human ever was able to, thanks to the Wright brothers, kind of have sustained motored flight. And then about sixty ish years later, in 1969, humans flew all the way to the moon. Right? So that within one lifetime, humanity went from, like, no one can fly for a very long time other than sort of gliding downhill or something, no one can really fly, to we all fly to the moon. Right? And so for us, I think what That means is we are probably, and I sometimes say this, we're, like, too late to explore Earth. We're too early to explore the stars, but we're right on time to build an AI that could actually do what flying did for some in one lifetime due to intelligence. We can build and move from AI being worse at everything that we do to possibly being better at any specific task that we do. Right? And that that will probably be our our sixty year time frame. Meant because everything was faster, it might only be thirty years or so. So then, there's an interesting connection between technology and science and theory. Right? Like, sometimes the application comes first and then we develop the theory later and then improve the technology. Sometimes the theory comes first and from that we can build new kinds of technologies. And so it's very helpful to think a little bit about the philosophy of science and know better to be inspired there than Karl Popper wrote that. Just like in other types of evolution, when we choose a theory, we also choose one that is best in competition with other theories. Of course, you need, if you wanted LLMs to do that, they need to find them, you need web search, for instance. But, in the theory that best holds its own, it's one that, just like evolution, has a certain natural selection process. Right? It proves itself and there is also a sort of survival of the fittest going on in scientific theories. And, in fact, a lot of science, according to Popper, is Basically us proposing a new theory, hypothesis, or explanation or description and then subjecting it to rigorous empirical testing. That is the, essentially evolution, evolutionary pressure of scientific theories. And basically, that was a very short run through, sort of the history of open ended evolution, which Hopefully makes us all realize that more science will lead to more technology, which will lead to more growth, which will lead to more human flourishing. And so that then begs the question, does it make sense for us to try to just scale up and spend a lot of our resources as humanity to scale up scientific discovery in order to lead to this flourishing. When when you double click into that you kind of realize, which Sanus Aflaem already realized a long time ago, that the exponential growth of science will To be at some point halted by the lack of people working on it. Right? There are so many niche subfields now in all the different areas of science that it's very hard to get a million people to work on that particular thing. And so as a result of this incredible widening of the scope, he says, the number of people focusing on any single section of it has decreased. And that then leads us to really thinking about how could we automate this and automate scientific discovery. That then leads us to what I call the Eureka machine. This is basically, our attempt at trying to build a machine that automates the process of scientific discoveries. And, in fact, I like in a couple months, I'll have a book coming out on this exact idea. And so I'll just give you a super high level highlight of how such a Eureka machine could be built for basically everything from physics, chemistry, biology, neuroscience, medicine, economics, astrophysics. So on. And there are essentially four pillars that are all extremely important to this machine. One is, of course, you have to understand what knowledge is already out there, what, things humanity has already invented. You have to get all the scientific measurement data into, as in the second pillar, this machine, then for things that you cannot yet measure, we don't yet know, you should try to then build simulations. Anything you can simulate. You can verify and you can then solve with AI. And if all else fails or at the very end of these processes, you still need to have some kind of physical industrial, like, lab that actually can run real experiments in the real world. And on top of all of this, you'll have basically an agent swarm that will deal with all of these different sources of knowledge and data and experimentations and rewards. And in terms of, you know, the foundational model of knowledge, of course, we also, you know, basically is is a good example of how every single technology we've built so far, especially in AI but also before that the internet, browsers, GPUs, and so on, we can rethink and there are a lot of startups possible in rethinking every single one of the layers of technology as infrastructure for superintelligence. At U dot conf, for instance, we work on web servers. For LMs, right, and agents and so on. And that actually is quite different. Right? Agents can read thousands of very long snippets, rather than just 10 blue links with, like, a very short snippet. And so you can rethink each of these different layers of technology that we've built for people and rebuild them for AI in order to use them as tools to then build superintelligence. Now, that is essentially The sort of why. Like like, we wanna build superintelligence in order to automate science. And to me, that will be the next big step function change, in in humanity and technology as we know it. Now, how do we actually build it? I think the best way to build it is to have it built itself. Right? We've moved as a field and especially natural language processing, for instance, which I've worked on for many years, we've moved from Not having linguists, this feels like ancient, you know, BC history, but before chat GBT, we we moved from having linguists tell us a bunch of things about language and then training statistical models on top of that. And when we allowed neural networks to actually automate learning those features with word vectors and, other neural network architectures and back to back, end to end learning and back propagation, we basically, were able to get much bigger improvements than We did a bunch of architecture engineering. Now a bunch of people at least are working on a unified architecture, but even that unified architecture has a lot of manual processes. And so it's clear over and over again in AI that when we take out a manual process and we replace it with a learned system, improvements will follow. And so that's why I think we should try to build a secure machine by having an RSI that builds itself. And the beauty is that only Now, AI can actually do this because AI is code and AI can code now. This this ability to really code in longer and longer time horizons has really only happened in the last, like, six to eight months. And that now enables such an RSI to work on itself, to develop almost a certain sense of self awareness of its own shortcomings and then fix those shortcomings. And then once we have that machine, that it has gotten really, really good at doing research in the AI. Self, we can then use it to do AI research for a lot of other things, in in other scientific fields. And so at a high level, it's quite easy. Right? We have three steps, ideation, implementation, and validation of ideas. That's true for basically almost every scientific field. And so, to end maybe on some very specific examples, we have built this first kind of version of such a Eureka machine. And we wanted to just show that it works on some small samples that a lot of people know and are aware of. And so we basically started with three things that show you and give you a very first glimpse of and sort of simple proof points, of what such a machinery can do. And that was basically better training, faster training, and and better kernels, for for NVIDIA GPUs. The first one, nano chat. I'm sure many of you have heard of it. A lot People think that's already recursive self improvement and it is kind of a weak form in the sense that usually when you do auto research, it's it's not recursive self improvement. Right? True recursive self improvement is when you have an AI that has a sense of self awareness of its own shortcomings, full access over everything, in its arsenal from pretraining to RL training and harnesses and everything, and then actually updates that entire system in the next version of itself. Now, you can also take such The system and just ask it to improve some other process, some other AI, like a small nano chat run where you can train something in five minutes. And that is really exciting. It's an important milestone, but it's not actual RSI. So here basically showed three examples of such an auto research system and what it can do. And after a very, very short time, it essentially was able to outperform many different teams and teams that also use other For AI research. So let's double click into some of these. Nanochat is a really exciting example. Basically, you train a very small, chat model, in less than, five minutes and you basically want to have it get to the best possible bits per byte number. And so the whole community had worked on this for quite some time and got to 0.93. And after This for a little more than a day or two, we basically got it down to 0.91, which is pretty exciting. Now, it wouldn't be that exciting if all it did was just find a couple of hyperparameters, and tune them carefully. But it actually did find truly interesting novel ideas like hash bigrams and trigram embeddings and tables for those, and mixing that into various value paths of the intention. Through a variety of learned gates. So it actually started to doing more and more interesting things rather than just kind of tuning hyperparameters. Another one, a nano gbt speed run. Obviously, speed's very important. So here we're able to work on this again, apply the system and after a very short amount of time it got better than, people working often together with AI for over a year, on on this very on this benchmark and made the whole thing another two Seconds over two seconds faster, at seventy seconds. And again, discovering, very interesting ideas in the process. And then the third one is CUDA kernels. Of course, we all care about not burning through our GPU budgets too quickly, and trying to be very efficient. I think in general, it's actually kind of shocking how inefficient a lot of mixture of expert models still are run-in very large clusters that cost billions of dollars and only have, like, 30% or so utilization. There's a lot of work that's ongoing in the world to improve that and different Fields or different groups of people are various different stages of that. But long story short, lots of different CUDA kernels are used during training and testing. And here, we basically, again, took that system and after, a couple days, it discovered better kernels, than the leaderboards best, on the NVIDIA, benchmark website. By, again, quite quite a sizable margin across all the different, categories of those kernels. And while we are pretty good at AI and, like, we actually and the team didn't have any particular CUDA kernel experts who just spent their entire careers writing good kernels. But still, you know, we do just enough to make sure and work together with NVIDIA to make sure that there are no reward hacks here and and other issues, but actually found, that Eventually, these all checked out and or indeed, pretty much all the different kernels, found the best solutions there. And so with that, I hope I could convince you, that indeed RSI could be that next big, s curve, an exponential that gets layered, on top of previous exponentials. And, that should help us, with not just AI, but eventually science and then all of technology and then, allowing many

## Slides

### 00:00:17

# Humans Remain In Control

## Humans in the Loop

> Please get my current entity and my current context

### Tool approval required
OG Assist is trying to make a tool call and needs your approval.
Tool: `getEntityDetails`
Buttons: Reject, Accept

Tool Calls can be gated behind tool call approval UI requiring manual intervention and explicit permission from a user prior to the agent taking action.

**humans are in the driver seat.**

[Screenshot of a user interface demonstrating an AI assistant requesting approval for a tool call (`getEntityDetails`) from the user.]

### 00:01:23

# AI Engineer World's Fair

- ORACLE
- arize
- bright data
- BrowserBase
- paper compute co.
- extend
- vast.ai
-

### 00:01:50

# AI Engineer World's Fair
- Z.AI
- qodo
- PayPal
- neo4j
- Google DeepMind
- arize
- reduc
- Amazon A

### 00:02:18

## The Eureka Machine:
Evolution and Automated Research

Richard Socher
Recursive, You.com, and AIX Ventures
July 1 2026

## Engineering the future of AI

[Abstract background image with a network of connected circles, some glowing orange and purple]

### 00:02:41

# Open-Ended Evolution
## From Biology to Science, Technology, and AI
Engineering the future of AI

[Abstract illustration of a glowing, tree-like neural network or brain structure against a dark, starry background]

### 00:03:16

# Open-Ended Evolution
From Biology to Science, Technology, and AI

Engineering the future of AI

[Abstract illustration of a glowing, tree-like network of dots and lines against a dark, starry background]

### 00:03:55

# Biological evolution

## Engineering the future of AI

[Radial phylogenetic tree illustrating biological evolution from Earth's birth to the present, showing major life forms like Bacteria, Archaea, Eukaryotes, Plants, Fungi, and various animal groups, along with significant geological and extinction events.]

### 00:04:15

# Technological evolution
### Fig. 1 World Product, Data vs. Models
- CES Combined Exp Model
- Hyperbolic Model
- Sum of Exp Model
- World Product Estimates

Hanson, R. (2020). Long-Term Growth As A Sequence of Exponential Modes.

# Engineering the future of AI

[Log-log line graph showing world product estimates over time, with phases labeled Hunting, Farming, Enlightenment, and Industrial Revolution, and three different exponential growth models.]

### 00:04:47

# Technological evolution
### Fig. 1 World Product, Data vs. Models

> Hanson, R. (2020). Long-Term Growth As A Sequence of Exponential Modes.

## Engineering the future of AI

[Line graph showing World Product Estimates over time, with three different exponential models (CES Combined Exp Model, Hyperbolic Model, Sum of Exp Model) plotted against the estimates. The x-axis represents time (logarithmic scale, decreasing from left to right), and the y-axis represents World Product (logarithmic scale). Key historical periods are labeled on the curve: Hunting, Farming, Enlightenment, and Industrial Revolution, indicating accelerating growth.]

### 00:05:17

# Technological evolution -> More people

[Line graph showing global population (in millions) over time (from -9000 to 2000+ AD), with key historical and technological events marked on the timeline. The graph shows slow population growth until around 1700 AD, followed by a sharp exponential increase. Events marked include: Beginning of 1st Agricultural Revolution, Beginning of Pottery, Invention of Plow, 1st Irrigation Works, 1st Cities, Beginning of Metallurgy, Beginning of Writing, Beginning of Mathematics, Peak of Rome, Peak of Greece, Discovery of New World, Black Plague, Beginning of 2nd Agricultural Revolution, Beginning of 2nd Industrial Revolution, Invention of Watt Engine, Beginning of Railroads, Germ Theory, Invention of Telephone Electrification, Invention of Automobile, Invention of Airplane, Penicillin, War on Malaria, Discovery of DNA, Nuclear Energy, High Speed Computers, Man on Moon, PCs, and Genome Project.]

Fogel, R. W. (1999). Catching Up with the Economy. American Economic Review, 89(1), 1–21. https://doi.org/10.1257/aer.89.1.1

## Engineering the future of AI

### 00:05:51

# Technological evolution -> Growth

> **The only perpetual source of growth is technology.**
>
> In fact, technology – new knowledge, new tools, what the Greeks called techne – has always been the main source of growth, and perhaps the only cause of growth, as technology made both population growth and natural resource utilization possible.
>
> — Marc Andreessen (2023)
>
> https://a16z.com/the-techno-optimist-manifesto/

### Engineering the future of AI

### 00:06:20

# Technological evolution -> Growth

> **"The only perpetual source of growth is technology."**
>
> In fact, technology – new knowledge, new tools, what the Greeks called techne – has always been the main source of growth, and perhaps the only cause of growth, as technology made both population growth and natural resource utilization possible.
>
> — Marc Andreessen (2023)
>
> https://a16z.com/the-techno-optimist-manifesto/

### Engineering the future of AI

### 00:06:49

# Technological evolution -> Growth

> "We believe that there is no material problem – whether created by nature or by technology – that cannot be solved with more technology.
> We had a problem of starvation, so we invented the Green Revolution.
> We had a problem of darkness, so we invented electric lighting.
> We had a problem of cold, so we invented indoor heating.
> We had a problem of heat, so we invented air conditioning.
> We had a problem of isolation, so we invented the Internet.
> We had a problem of pandemics, so we invented vaccines.
> **We have a problem of poverty, so we invent technology to create abundance.**"
> – Marc Andreessen (2023)

[https://a16z.com/the-techno-optimist-manifesto/](https://a16z.com/the-techno-optimist-manifesto/)

## Engineering the future of AI

### 00:07:22

# Change in one lifetime.

- Beginning of 1st Agricultural Revolution
- Beginning of Pottery
- Invention of Plow
- 1st Irrigation Works
- 1st Cities
- Beginning

### 00:07:46

# Change in one lifetime.

## Historical Milestones and Population Growth

-   **Early Innovations (approx. -9,000 to -2,000):** Beginning

### 00:08:14

# Technology <-> Science and Theory

> We choose the theory which best holds its own in competition with other theories*; the one which, by natural selection, proves itself the fittest to survive [...] the one which is also testable in the most rigorous way.
>
> — Karl Raimund Popper (1959)
> *LLMs need search to find them

## Engineering the future of AI

[Black and white photo of Karl Popper]

### 00:08:56

# Technology <-> Science and Theory

> We choose the theory which best holds its own in competition with other theories*; the one which, by natural selection, proves itself the fittest to survive [...] the one which is also testable in the most rigorous way.
>
> — Karl Raimund Popper (1959)
>
> *LLMs need search to find them

## Engineering the future of AI

[Black and white photo of Karl Raimund Popper, an older man with glasses, looking forward and gesturing with his hands]

### 00:09:14

# What is Science (according to Popper)?

- **Variation**: We propose a new theory / hypothesis / explanation.
- **Selection**: We subject it to rigorous (empirical) testing and criticism for falsification.

If it survives, it's provisionally accepted, but still open to future challenge.

### Engineering the future of AI

### 00:09:48

# Engineering the future of AI
- More Science -> More Technology -> More Growth -> More human flourishing
- Could we scale scientific discovery?

### 00:10:18

# Engineering the future of AI**
> "The exponential growth of science will be halted by the lack of human resources. As a result of the immense widening of the scope of scientific research, the

### 00:10:49

# THE EUREKA MACHINE
WHY **AI** IS THE KEY TO UNLOCKING A NEW ERA OF SCIENTIFIC DISCOVERIES

## Engineering the future of AI

[The word "E

### 00:11:16

# The Eureka Machine
## Full-Stack Scientific Superintelligence

- **01** Foundation Model of Knowledge
- **02** Physical Reality Grounding
- **03** High-Fidelity Simulation
- **04** Autonomous Physical Labs

### Engineering the future of AI
[Four abstract icons representing the concepts of knowledge, physical reality, simulation, and autonomous labs]

### 00:11:49

# The Eureka Machine
## Full-Stack Scientific Superintelligence

1.  Foundation Model of Knowledge
2.  Physical Reality Grounding
3.  High-Fidelity Simulation
4.  Autonomous Physical Labs

## Engineering the future of AI

[Icon representing a network of connected nodes]
[Icon representing an atom with orbiting electrons]
[Icon representing a 3D cube outline]
[Icon representing a branching, tree-like network]

### 00:12:20

# GPUs, The Internet, Browsers, Search will become infrastructure for scientific superintelligence and part of pillar 1 of the Eureka Machine.

1.  Even the most powerful models can't know everything – they need search for current, accurate, real-world knowledge.
2.  Every AI system needs trusted knowledge from the web to act.
3.  The knowledge layer is the prerequisite for pillars 02, 03, and 04. You cannot simulate what you do not know.

## Engineering the future of AI

### 00:12:50

# GPUs, The Internet, Browsers, Search will become infrastructure for scientific superintelligence and part of pillar 1 of the Eureka Machine.

1.  Even the most powerful models can't know everything – they need search for current, accurate, real-world knowledge.
2.  Every AI system needs trusted knowledge from the web to act.
3.  The knowledge

### 00:13:20

# RSI: The next step for AI and technology

## What? Automate Science
To build the ultimate invention generating machine we need to automate the process of science itself, starting with the science of intelligence.

## How? Recursive Superintelligence
That requires a minimal sense of self awareness of that AI system and naturally leads to recursive self-improving superintelligence.

### Engineering the future of AI

### 00:13:45

# How to build the Eureka Machine?
It has to build itself.

### Engineering the future of AI

### 00:14:16

# How to build the Eureka Machine?
It has to build itself.

## Engineering the future of AI

### 00:14:51

# AI is Code... and AI can Code ... for a longer time
## Time horizon of software tasks different LLMs can complete 80% of the time

https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/

# Engineering the future of AI

[Line chart showing the task duration for humans (Y-axis, log scale) versus year (X-axis) for various LLMs (GPT-2, GPT-3, GPT-3.5, GPT-4, Claude models) completing software tasks like training a classifier, finding facts, counting words, and answering questions, with a dashed line indicating a 1:1 time horizon.]

### 00:15:20

# Ideation, Implementation, Validation

1.  **Proposes an idea**
    Uses what it learns to choose the next experiment
2.  **Implements it**
    Comb

### 00:15:47

## First simple proof points
# Better training, faster training, better kernels

-   **NanoChat Autoresearch**
    *   Several dozens of humans and hundreds of their agents
    *   SOTA: 0.9372 BPB → **0.9109 BPB**
    *   1.3x speedup to reach the same loss

-   **NanoGPT Speedrun**
    *   83 human record-setting contributions to the leaderboard
    *   SOTA: 79.7 s → **77.5 s**
    *   Similar or larger improvement than recent human contributions

-   **SOL-ExecBench**
    *   235 kernel-writing tasks derived from real workloads
    *   SOTA: 0.699 SOL → **0.754 SOL**
    *   18% reduction in gap to the optimal performance estimate of 1.0

## Engineering the future of AI

### 00:16:19

# First simple proof points
## Better training, faster training, better kernels

- **NanoChat Autoresearch**
  - Several dozens of humans and hundreds of their agents
  - SOTA: 0.9372 BPB -> **0.9109 BPB**
  - 1.3x speedup to reach the same loss
- **NanoGPT Speedrun**
  - 83 human record-setting contributions to the leaderboard
  - SOTA: 79.7 s -> **77.5 s**
  - Similar or larger improvement than recent human contributions
- **SOL-ExecBench**
  - 235 kernel-writing tasks derived from real workloads
  - SOTA: 0.699 SOL -> **0.754 SOL**
  - 18% reduction in gap to the optimal performance estimate of 1.0

### Engineering the future of AI

### 00:16:49

### First simple proof points
# Better training, faster training, better kernels

- **NanoChat Autoresearch**
  - Several dozens of humans and hundreds of their agents
  -

### 00:17:22

# NANOCHAT AUTORESEARCH - 01 / 03

## NanoChat Autoresearch: best validation BPB over wall-clock time
Recursive vs. Karpathy, SkyPilot, and autoresearch@home - single-seed evaluation - lower is better

0.9109
BPB - beats community best 0.9372

### WHAT THE SYSTEM FOUND
- Hashed bigram and trigram embedding tables, mixed into the attention value path through learned gates.
- A cheap way to use local n-gram information without paying the time cost of slower convolutional or attention-heavy alternatives.

0.0263 BPB improvement - evaluated on 10 random seeds

Engineering the future of AI
[Line graph showing validation BPB over cumulative runtime, comparing Recursive, autoresearch@home, Autoresearch, and SkyPilot, with annotations for various optimization steps.]

### 00:17:50

# NANOCHAT AUTORESEARCH

## NanoChat Autoresearch: best validation BPB over wall-clock time

**0.9109**
BPB - beats community best 0.9372

0.0263 BPB improvement - evaluated on 10 random seeds

### WHAT THE SYSTEM FOUND

- Hashed bigram and trigram embedding tables, mixed into the attention value path through learned gates.
- A cheap way to use local n-gram information without paying the time cost of slower convolutional or attention-heavy alternatives.

[Line graph showing validation BPB over cumulative runtime for different autoresearch methods, with annotations for key events.]

### 00:18:21

## NANOGPT SPEEDRUN · 02 / 03

### NanoGPT Speedrun: training time to a 3.28 validation loss
Track 1: GPT-2

### 00:18:54

## SOL-EXECBENCH - 03 / 03

### SOL-ExecBench: SOL score per kernel
### SOL-ExecBench: mean SOL score by kernel category

- **

### 00:19:17

# SOL-EXECBENCH - 03 / 03

## SOL-ExecBench: SOL score per kernel
Recursive vs. leaderboard best - higher is better (0.5 =

### 00:19:49

# SOL-EXECBENCH - 03 / 03

## SOL-ExecBench: SOL score per kernel
Recursive vs. leaderboard best: higher is better (zoomed to 0.5-1.0)

## SOL-ExecBench: mean SOL score by kernel category
Recursive vs. leaderboard best, doubleAI, and Cursor - higher is better (zoomed to 0.5-1.0)

- **0.564** Cursor
- **0.690** doubleAI
- **0.699** Leaderboard best
- **0.754** Recursive

0.5 = optimized PyTorch baseline · 1.0 = analytical optimal performance estimate · 235 kernels

# Engineering the future of AI

[Scatter plot showing SOL scores per kernel for Recursive vs. Leaderboard best]
[Bar chart comparing mean SOL scores by kernel category for Recursive, Leaderboard best, doubleAI, and Cursor]
