# GLM-5.2: Frontier Intelligence, Open Weights. — Zixuan Li — session 2026-06-30T16:45:00.000Z → 2026-06-30T17:05:00.000Z

_41 transcript lines · 10 slides · source: full recording_

## Transcript

My constraint shifted to token compute. All these threads run at the same time and my MacBook starts sounding like a jet engine. That's mostly fixed by using test boxes, so agents can run tests on a separate machine. Now, I'm primarily constrained by tension. And unlike tokens or compute, I can't simply add more of it. So the Most important skill is today is deciding where to spend it. Are you still staring at the agent while the code flies by? I know, I know it's it it feels cool, but With the earlier models, this was necessary, you know, you you you see the agent go in a direction you don't like, you hit escape, you steer it, you steer it back. But the latest generation of models is so good at understanding intent that it's a little bit of a waste of time. To watch the agent generate code. Imagine someone files an issue on one of my open source projects. The manager wakes up, reads it against the project's goals, notes and vision, and decides whether it might be a fit. If it does, it creates a worker. That worker investigates. Implements the change, runs the tests, and another agent can review the result. I don't need to watch those agents work or consume every intermediary message. When the manager needs me, it returns a PR, the original issue, the proposed diff, maybe a video or even a running build I can VNC into. I review once, I leave a note, I maybe approve, the loop continues, and it can land after the checks pass. The agent runs the inner execution loop. I set the direction and I make decisions in the outer loop. You know, Paul, Paul Salt is already running a version of this. He pinned his chief of staff, it wakes up every ten minutes, and it coordinates his GitHub work. The agent creates threads in the sidebar. So Paul can jump in whenever the work needs additional steering. And, you know, once the manager is long lived, tying it to a laptop just feels wrong. Codecs can already move work between hosts. OpenClaw has a gateway and nodes. But neither feels like the final form. I don't even want to think where I work. My agent should be able to connect to any of my machines. They should know which work can be done in the cloud or which work requires my local machine. The manager shouldn't be a session trapped inside your app. It should be an agent that I can text, steal from Slack, or hear from wherever I am. Really? Why can't I talk to my agent and have it Please join me in welcoming cofounder of AI Engineer and editor at latent space, Swix. Hi. So we actually were supposed to have Tishan from Zee, joining us today. He unfortunately couldn't make it in, to The United States. But we have him dialed in. Tishan, Be live. The the I hope I hope he's where where the the team backstage is is running him live. One thing I also wanted to, thank people for doing is this is a very, very big community effort. I love that. Thank you so much to the open OpenAI team, for and to Microsoft for bringing the energy in the opening things. But even, all you guys have been organizing all these side events. And one of them, I just wanted to shout out in particular, the AI Engineer Kids Day. You know, we're very sort of Professional, serious, but we, you know, we also care about the next generation of AI engineers. So I wanted to play the video. We're here in San Francisco at AI Engineer World's Fair, which is the biggest collection of AI engineers and entrepreneurs in the world, and we're holding the first ever AI Engineer Kids Day. And I really love this workshop series because we're Firing the next generation of AI technologists, these kids are the people who will be doing AI in the future. They're very AI native. They'll be the AI experts. Something that makes this work workshop stand apart is traditionally when you work with three d games, if you don't use AI, it takes a really long time and you won't get as far. But in this workshop, because we use codecs, we were able to get a lot further than you normally would in a workshop and be able to explore a lot more advanced concepts than normal as well. Come here. It's really fun. Blah blah blah blah. And you can train AIs and create AIs and ask AIs to create the game. It's really fun. Give it up for the kids, events. Thank you to Steven Chin for running that, and basically generating the funnel for us for ten, fifteen years from now, for our next attendees, we have to grow our TAM somehow. Showing up in person, but I have a team whole team coming to the town, and we have a booth. So if you have any questions, feel free to reach out to me on x, LinkedIn, and also you can look for my team and our booth. And since I cannot see my slides, so I will need Swix to help me flip through all the slides for me. Yeah. You're good. Can you share, like, what which slides we are in right now? Yeah. So we're we're on the opening slide. We're talking about intelligence and z AI. Okay. K. So yeah. So you can you can see that it's the first time for us to introduce Gian four point two, five point two to the role. And, also, we are gonna share something about z.ai and GLM because maybe people will think z.a and GM, they're irrelevant. Right? Your your company is not g.ai or your model is called z one or z two. And you can find my x and our company's x account here, so you can just search for My name is Xichuan Li and the z a I org, and you can follow them for the follow ups. So, yeah, we can go to the second slide. Yeah. Actually, I can all see the slide, so I'll try to So we we Yeah. Try to make sure that it is correct. So the company x actually is called Jupu. Maybe some people have heard of it. And the mall The models, it's it's called GLM. Actually, it's not a a brand name. It's a generic term. So GLM actually represent general language model for training with autoregressive blank filling. That paper was published back in 2021. So actually, we were the one of the first labs to do explorations on large language models. So at the same time with OpenAI and Anthropic and DeepMind. And even today, we we no Use GLM as the architecture. We still use the name GLM as our brand name. So we use j GLM 5.1, 5.2, and it become, like, one of our, like, most prod productist product and model. And the second thing that we look for is intelligence upper bound. So in terms of intelligence, we we may feel that it's represent IQ or something something like that. And when DeepSeek launched one O one lunch, people are talking about the model's capability to solve math problems, physics problems. But what actually intelligence mean is not just, like, IQ or ANE or other other physics problem. So from GLM 4.5 to GLM 4.7, we are exploring, like, several things like reasoning, coding, and genetic capabilities. So as you can see from the slides, so we add, like, yeah, the last slide. Yeah. We were I have a I haven't, like, finished that slot. Yeah. So okay. Yeah. I need to go back to the to to the slides. I don't have a Back I don't have a back button. Thank you. I would I don't understand kickers that don't have back buttons. Like, why okay. You know? Anyway, go ahead. Yeah. Mine. Yeah. Because people wanna see the GRN 5.2. They they don't wanna see, like, GRN 5.1 or five. But, like, GRN 5.2 actually specialize in coding and genitive task as you can see from the graph because there are a lot of rumors whether your model is close to miss those fable, but, actually, I want to share these slides to all of you. So you can see it's somewhere between Opus four point seven and four point eight, and we use the hardest problems like DeepSuite, Terminal Bands. Two one one, which was mentioned by the OpenEye team, several minutes ago. And all the, like, long horizon task and benchmark shows that the the capability is on par with at least Ovis 4.7, and it it shows, significant improvements over 1.1. Also, for GRF 1.2, we add Thinking, level called high. So because we also noticed as we move to the harder task, it might consume more tokens, and also we we care a lot about the token efficiency. So it's the first time we add the high level for thinking budget, but even without thinking, the non thinking model is better than the 5.1 thinking model. So I think it's a huge improvement for the open wave model. That's what really impressed the world and why people are talking about GLM 5.2 lately. Okay. The next slide. And one one thing that I wanna mention is that GLM is more more than a coding model because people use it inside clock code, codex, open code. But, actually, we have trained a lot of things outside coding. For example, we improve a lot in GDP valve and also math problems. We also care about math problems, frankly speaking. And also we train a lot of thing, Related to role play, general chat, we want to improve every aspect of the model. So you can see from the artificial analysis intelligence index, actually leads the other OpenWay model a lot and close to the Frontier model. So I want you if you you haven't experienced GLM yet, you can use GLM to do general chat, use it to process your daily workflow, not just for coding, but you can For the model, like, beyond the coding scope. Next slide. And GLM 4.2 is a open weight model. So people always ask me, why do you open weight? So do you care about your business or, like, do you care about losing market to some inference providers? But actually we open the weights for several things because they are users' needs and they are ours' needs. If we can Meet their needs, I think it's okay. It's definitely okay for us to to open the model. For example, if our users wants security and control and we want to build trust, we can open with the model. For for some cases, if an enterprise or government, especially in the Western world, want to use the model, we open way we upload to the hugging phase so that they can use the model on premise. I think it's very

## Slides

### 00:06:32

# Engineering the future of AI
- Tokens
- **Compute**

### 00:07:32

- Issue filed
- Worker created

### Engineering the future of AI
[Icon of an exclamation mark in a circle next to "Issue filed"]
[Icon of a person with a plus sign in a circle next to "Worker created"]

### 00:08:06

- Issue filed
- Worker created
- Agents work
- Review and approve

## Engineering the future of AI
[List item with an exclamation mark icon: Issue filed]
[List item with a person and plus icon: Worker created]
[List item with a pencil icon: Agents work]
[List item with a thumbs up icon: Review and approve]

### 00:08:32

# Engineering the future of AI
## Paul Solt's Tweet
My NEW Codex workflow is better than I expected.
8 new features ready for release in my app.
Took

### 00:09:06

## The thread is no longer bound to the machine
Engineering **the future of AI**

### 00:09:35

# Engineering the future of AI
## One agent. Every surface.

### 00:10:40

# SWYYX
## CO-FOUNDER, EDITOR
- AI Engineer
- LATENT SPACE

[Abstract dark, textured, organic-looking shape in the background]

### 00:16:55

# More than a "coding" model
## GLM-5.2: Frontier Intelligence, Open Weights.
Zixuan Li / Head of Z.ai Z.AI

[Bar chart comparing various AI models (e.g., Claude Fable 5, GPT-5.5, GLM-5.2, Gemini, Grok, Mistral, Gema, Solar Pro 3) by a numerical score, categorized by proprietary, open weights, or commercially restricted open weights.]

### 00:17:22

# GLM-5.2 specializes in coding and agentic tasks

## Agentic Coding Performance by Effort Level
Average over Terminal-Bench 2.1, DeepSWE and SWE

### 00:17:46

# GLM-5.2 specializes in coding and agentic tasks

## Agentic Coding Performance by Effort Level
Average over Terminal-Bench 2.1, DeepSWE and SWE-Atlas QnA, evaluated on Claude Code 2.1.167

[Line graph comparing Agentic Coding Performance (Score %) against Average Output Tokens (Per Task) for GLM-5.2, GLM-5.1, Claude Opus 4.8, and Claude Opus 4.7, with data points indicating different effort levels like Low, High, Non-Thinking, and Max.]
