# Thom Wolf keynote — Thom Wolf, Olive Song — session 2026-06-30T17:05:00.000Z → 2026-06-30T17:25:00.000Z

_66 transcript lines · 1 slides · source: full recording_

## Transcript

Official for the whole ecosystem to explore the model, especially when the capabilities is close to the frontier model. And second, if people want diversity. So in terms of diversity, I mean, the capabilities in legal, finance, security. They can fine tune the model, so we need to open the model to let them fine tune. For example, Harvey is fine tuning g I 4.1, and maybe they're thinking about Tuning GLM 5.2 afterwards. And I've I have, like, talked to a lot of other companies. They are also thinking about fine tuning GLM as their, like, next step or next strategy to differentiate themselves from other application problem. And third, if our customer or individual want to co design and predict the future, sometimes they need to See the architecture of the model. They need to see the recipe of how you train the model. So we want to make the norm. We want to make right bets. So we wanna co shape the future with our customers, with the open source, community. So I think our needs and our ecosystem pretty, like, fits into each other. And g I five point two couldn't succeed without you, all the open source community. Players like Oniflaz, NVIDIA, and and some, like, individual super developers, they are part of it, and it affects a lot to application builders like Peter. Right? Because open source doesn't just include open source model, but also open source softwares and the open source other sorts of support. So All your support and what are you you're doing right now push us to to make better models. I think, you you are the true hero. And the last, I wanna share a a great resource for you to go through GIS 5.2. It's our tech blog. So, actually, in that tech blog, we share several things, like our Hugging Face, repo, how how you can try the GLM 5.2. So you can call the API, and, also, we have a coding plan like the codex or CloudCo subscription for you to use your tokens as individuals. And, also, we share something about our training pipeline training recipe, which you can understand why it's a good model. So there are a lot of things behind the behind the model, not just, a a a model that had great data. We also have fantastic technologies behind the model. So You can try it inside the chatbot, agent, and You can explore the model yourself. We can see what difficulties we have gone through. And the last slides, actually, it's kind of a one more thing. So it's the first time we share z code to the whole community. So that one more thing is we actually have our own harness. The z code, actually, it's built for g I 4.2, but also support all Frontier. Models. You can bring your own key. You can connect to z code. Actually, the I think the operations is similar to codex. Actually, you you can try some techniques like Go or other, like, compact technique techniques like what you've done in in codecs and call code. And this harness, I think it's the perfect one for GLM. Welcome to. If you haven't experienced it, you can just search z code or Or you can go to our booth. We have our team members showing this harness to you, and welcome to our booth, and welcome to, like, talk to me in the future. And next time, I'll definitely be in SF talking to everyone. Yeah. Thanks. Thank you very much. So, trust me, the very first World's Fair, I wasn't allowed to come back in the country. For for my own conference, so I know exactly this feeling, but it's all good. The Z AI team, I really appreciate them, making the effort. They really wanna meet you. They're here to meet you. And this is the whole point of the World's Fair, to bring all the top world's AI companies and labs, all in one place so you can do business together, meet the people behind the models that you use, ask the questions that he cannot answer in public, but you can ask in private. I'm also very proud, he showed So she showed that that list of, hugging face, you know, top, contributors. I think we are four for eight, in that list, present at Willows Fair. We're working with NVIDIA and Unslaugh and Ollama and and, all those sort of fine tuners as well to, to make a sort of the inference and the local tracks that you're gonna see over the next few days. With that, thank you so much, Xin. I'm gonna invite on the next couple folks. I think I might be doing Ali's job here. So, We'll talk to Hugging Face, next in in Mini Max. Thank you. Joining us on stage is the cofounder and chief science officer at Hugging Face, Thomas Wolfe. Hello, everyone. Hello, Olive. Nice to have you on stage. Hi. Nice to thanks for having me. Yeah. So I think you're on for a treat today because you just saw, GLM, which is, current number, two on the artificial analysis table. I take that out because nobody can use it. And now we have number four. So, basically, you will have all the top models, at least the top open source model in a row. Was a pretty amazing path in life. So she came to The US, Pennsylvania. She was studying, doing PhD at NYU, in the lab of, working on JPA, but we decided we won't talk about JPA today. Right? Something for another day. And we are very lucky to have, a Liv. And then, instead of joining Hugging Face, which was, in New York also at that time, she, decided to go Join, MiniMax. So, for those who who maybe don't know all the all the NeoLabs, around the world, and you're you're forgiven because I think there's, like, 64 NeoLabs right now. Minimax is one of the top of what we call the AI dragons in China. So these are the new there's, there's DeepSeek, which is very well known now, Moonshot with SKIMI, z, and GLM that you just saw, and now we have, a MiniMax. They're all extremely good, extremely talented team, fighting for the first spot. Release of MiniMax was m three, just earlier earlier in June, which was the the top model at that time, top open source model. Very impressive. There's a lot of very interesting things about this model, so we'll quickly dive in them and then talk a little bit about, what's what's what's specific about minimax, what's what's great there. So, maybe, Olif, to to start a little bit, can you can you give us, you know, a little bit of your your view of of m three? What you're So the the the latest Like about this model? How is the release? Mhmm. Yeah. M three we released m three earlier this month, and it is a smaller model with 400 around 400 billions total parameters and 20,000,000,000 activated, but it is very capable in terms of both coding performances and also it understands vision. So, that's, what open source models don't usually have, is that they can The model can only deal with coding, but it can also understand videos, images, and it has a super, long context of 1,000,000, with our new architecture called MSA, minimax sparse attention. So we we really put these three things together, because we know that they are they will be very important in future AI applications, coding capabilities, agentic capabilities, longer context, and multimodal. Understanding. Yeah. I think that would be very interesting about the model. Yeah. So so there's a lot to impact impact in this model, and it's, it's it's still, I think, the only top five model open source model that is actually multimodal, so we need to talk about that. But maybe first about the long context because there was also the first one that really had this real 1,000,000 token, long context that's actually functional. And you guys had also the the minimax pass attention, which is this one technique To make that efficient that you also published and and share extensively. So can you can you talk a little bit about this, maybe how the project went from from the attention, how to make this long context? Yeah. I would say the story about long context went back to either minimax m one and minimax zero one, where the model was actually was able to perform tasks of 10,000,000, token context. 10,000,010, yes. But then it was not an And take a model. Right? It was just, 10 for example, dumping a book, it would be able to give reviews on it, stuff like that. So, what we realized was that, you know, longer context actually unlocks a lot of capabilities, especially when interacting with users. And now when, you know, the agent is interacting with the whole environment and getting all the tool responses, getting multi rounds, the, like, shorter context wouldn't be Enough to perform the complex tasks. So for this version, we said, oh, we have to have our longer context backs. So what we pursued was with our minimax sparse attention, which, you know, was the architecture that was scalable and had a simple design. So, I would say from a higher level, right, it has an index branch, that You know, selects on a higher level what is what matters more in the context, and then we have a sparse attention branch that calculates performs the calculation on the selected blocks, to actually perform the tasks. And so, yeah, like that we really designed, an elegant architecture so that we can scale the lens and then scale the Model size in the future with that. That's beautiful. I like how for for those who've been in the field for quite some time, we we had a lot of work and attention. Right? This n square, and there was a lot of linear attention. Yeah. And then some that somehow all of this disappeared at some point. When flash attention came around, we discovered we just needed more efficient kernel. Now I like how we come back to thinking, you know, first principle, what is attention? How can we make that more efficient? So 1,000,000 token is crazy. Right? GPT 2 was a thousand 24, and and everyone was like, oh, that's really big. We we we never need more. Where do you see this coming? Like, going in the future, like, Jeff Dean was pitching me the other day a trillion token attention. Do you think we should go to your token attention? That's definitely something we can explore towards. Right? Ultra lens of the context, definitely. That's something that's very exciting to explore with and something that architecture design along with hardware, would require a lot of research onto that. Yeah. You think there's still a lot of low hanging fruits? So typically today, we saw OpenAI really reducing Mean, we don't know how as a firm, but, like, reducing their their inference bill by half by probably having some more efficient processing around tensions or anything like that. You think there is still a lot of low hanging fruit that can be getting, how we can process that. So so one one thing is still very interesting about m three is how cheap it is in particular because of this part of attention or in part because of its small one, but it's also very efficient. Right. Do you think we can go even way further? And maybe how did you guys invented, min max pass attention? Was it An agent coming up with the idea? Was it a human still coming up with the idea? Tell us a little bit about Yeah. So we do think there's still a lot of work that can get into architecture and inference optimization so that the model can be more efficient, especially if there are tasks that are very task sensitive but require very strong capabilities, right? And for that, those kind of tasks, we really want the model to be efficient. And who came up with this part? Actually, I think An intern from our team worked on that. Yeah, an intern. That doesn't usually happen in a lot of labs, because I think in some labs, interns don't have access to the data, the work, and stuff. But, yeah, we are open to anyone who would like to contribute to our model, so, the architecture was actually designed by an intern. That's very good. Still some work for interns here. Good, no good news. That's also a good segue. Also how min max is working internally. So so we were discussing before coming on stage, they were saying everyone can propose a project. Can you tell us a little bit about how you are organized, how you do research? Mhmm. I think that is very different from, even in school or even in earlier, you know, the earlier tech companies is pretty, pretty different. It's that, what we what we make sure is that we have good foundation and good Infrastructure so that anyone can play with the model and can think of what they can improve with the model, and then after model releases, when they are free, right, they can play with the model, they can think of their own evaluations, they can find their own weaknesses, and propose a thing that they want to improve on the model. And then other people who are interested in that would, you know, propose to join the project, and they will work on for a couple of weeks or even a couple of months, and When they work out, the final thing is shipped to our model. It is, you know, we use that in our final training, and it's shipped out to the audience. Interesting. So you can have people working for a really long time on project? When you say a couple of months, it can be like really deep exploration, I mean, if possible. Yes. I would say, for example, architecture might require a longer time of investigation, research, experiments, even redoing the evaluations for pre training. Yes, so it might require a longer time. It's really nice. Yeah. And I know you're also very big on evaluation. I agree. We could talk about that. I think one one thing probably related to that is this unique specificity that m three and your team has, around multimodality. So not just text, but this model can also understand image and video. And as I understand, but but please explain explain better. When we read the model card on Hugging Face, it say the model was trained from the first step as a multimodal, not just The, like, user one is after source. Right? Yes. Can you tell us a little bit more about that and why you think it's important? Mhmm. And and and why starting from the first step on multimodal training and not just just training? So we call it native multi modality. And so it is somehow typical for model labs to train the multi model, let's say, vision understanding capabilities after the text pretraining is done. They put adapters and then train that part, but what we found out was that that would actually On the text performance, and the vision vision understanding performance wouldn't converge that well because the model is kind of converges towards the model, the text understanding, and it's just not the most optimal, and also not the most scalable if you think about it. We want to scale the data, right? And also we can also some labs, train this capability from Halfway through the pre training, for example, continued pre training, but what we found is that this would be very, you know, recipe sensitive. It is different for the recipe would be different for different architectures, different, you know, data mixtures, different learning rates. It's hard to control, hard to, you know, scale to you can't really scale your experiment results and conclusions to a larger model. And so, you know, what we thought was why not just training? And from the very first step, that comes the most natural. We know that a lot of labs run into problems doing that. The model would collapse after a couple of steps of training, you know, both text and vision understanding, but we managed to solve that problem. We did a lot of work on, the IT, and we did a lot of work on the data that we're actually training. For example, we do interleave the data, what we call interleave the data. It's actually natural data, but we keep the, images and videos in instead of masking it out, and we do some pretty good cleaning and masking on the data, and we do very good reward modeling so that we train it from the first step and scales up a lot. Yeah. It does does not collapse. That's really impressive. Impressive. Should do should we expect much larger model in the future? So this one is still fairly small. Right? It's It's 428,000,000,000 parameters, 23 active billion. Well, do do you think you will go past the trillion? Definitely. Yeah. Definitely in the future, there are many tasks that wouldn't be able to the more model wouldn't be able to perform what we're good at with smaller parameters. We are definitely going more ambitious than this. That's great. Looking forward. Another interesting thing, I always I find fascinating about minimax is how how you also have this whole range of of apps and product. Right? So I remember already so so minimax started to open source things on the on the Hugging Face platform in January last year, so that eighteen months ago. And and we were chatting a little bit about the team to understand what you were doing, and I remember you so you were already having a huge usage on some of these, of some of these apps. Can you tell us a little bit how this started?

## Slides

### 00:14:24

[Dark background with a repeating pattern of white, outlined partial company names and logos, including the Microsoft logo.]
