# TCP and RDMA are Killing Inference Throughput; Homa can Fix It — John Ousterhout — session 2026-07-02T16:20:00.000Z → 2026-07-02T16:40:00.000Z

_70 transcript lines · 0 slides · source: full recording_

## Transcript

Fine tuning is the clearest not yet, like most people don't have it at all. And, folks are pretty locked in. So, those who bought aren't looking as much to build. Those who built aren't looking As much to buy. But those are those are the core takeaways from the usage in our stack. So many of you work on teams, and like we said at the start, these range from solo founders to large enterprises. What is this doing to teams? And remember, this is a builder heavy sample. But among builders, the vibes are good, which, you know, I'm sure if you look to your left and right, you're feeling that. The vibes are pretty good. Percent report a net positive effect on their organization. The top effect isn't really just speed. It's cheaper failure, more experimentation, more prototypes, more bets. It didn't just make engineers faster, but it made trying things nearly free. And so there's some happy campers as a result of that. But it's not free free. You know, there's no free lunch as nothing is. So the same tool that increases experimentation also increases Review burden. Both can be true. And, you know, over nine in ten respondents are feeling negative downstream effects in some way, the most common ones being widely discussed at this conference, online and anywhere that you see AI engineers, erosion of deep technical skills and understanding of the codebase. And these are consequences of cheap code generation. And the org chart is really feeling it. So, many folks, 81%, are saying that AI is blurring the line between their role as engineers and product design and marketing. These stats shocked me. Where you feel it the most is shipping software once exclusively the engineers domain. I know folks talk about vibe coding and how that's accessible to more folks than ever before in different roles. Today, over a third of teams have non developers shipping features, which was pretty wild to me. Mostly smaller, mostly internal, but 17% say that non developers are regularly shipping customer facing features across the stack. And even when non developers aren't shipping, a third of teams see them building really useful things: prototypes, front end mocks, and more. So shipping software is not gated on being an engineer. We knew this, but The extent to which it's being pushed is is higher than I expected. Alright, so where does all of this go? We always ask people to place bets rapid fire. So let's talk about those results. So present tense first, 76% say AI boosted their job satisfaction. So that's good for most of this crowd. I hope you're, as, Elphaba and Glenda say, I hope you're happy. Now, that's great. But 59% fear today's AI code creates long term liabilities. Only a third call software engineering a solved problem. Although, when I have conversations with folks, sometimes the way in which they define software engineering is different. So you can read into that stat as you will. Happier, faster, but embracing the maintenance build is the TLDR. And people are unsure what's going to happen with hiring. And for the five year bets, we have 67% expect a leading lab will declare AGI in the next five years. Note the wording. We said will we asked about the press release, not the achievement. So will they declare it? Yes. What does that mean? Not sure. Only 9% bet on transformers being state of the art in five years. Most are unsure, but that was interesting. And then my favorite: Will there be more AI compute in space or on land? 36, yes. 38, no. The most divisive question in the survey is about outer space. I promised you a lot of bar charts, and that was a lot of information, so a review or our 2026 wrapped. Impact is overwhelmingly positive. ImageGen doubled, or happy ImageGen doubled, while audio has the highest adoption intent, the same as last Here. Cost really became a first class constraint, and we see that everywhere in monitoring and how ambitious folks that are going out and building AI products are behaving. Open weights augment, but they don't replace. So, we're seeing a multimodel future with a consolidation of the stack. Agents got right access more than ever before. Trips Relative to last year. While the guardrail stayed pretty primitive and inference is the by market, everything closer to product logic tends to relatively stay more in house. It is a very exciting time to be an AI engineer. I cannot wait to see how the next year unfolds. So you can find the full report in the link up here, every chart plus some cuts that we didn't have time for today. I won't ask you to fill out a survey about the survey, but if there's something that you want on the books for 2027, something you're curious about, you can come find me here on the internet. I'm easy The spot. Thank you so much. We will see you next year or per 36% of you, maybe in orbit. Thank you. Please welcome to the stage the Professor Emeritus at Stanford University, John Oosterhout. Good morning. It's really great to be here to talk about the network side of AI applications, and in particular, to make the case that latency matters, and it's probably gonna be mattering more in the future. But But I just wanna say this is a talk is unusual for me. I've never before given a talk where there are fog generators. In the auditorium, just a really San Francisco experience, I guess. So it's well known that AI workloads depend on really great networking performance in order to achieve their own performance. And of course, that's because the workloads are so large that they have to be distributed across machines, and then you have to communicate between the machines. But what I want to talk about today is that it seems that those workloads are changing. And so I hope to do three things over the next fifteen or twenty minutes. First, to convince you that in fact Workloads are changing and that whereas the workloads used to be completely dominated by large transfers where throughput is the key metric that matters. That we're seeing more and more smaller transfers where the latency is crucial. The second thing I hope to do is to convince you that legacy protocols like TCP and RDMA are poorly suited to this environment. They weren't designed for this environment and unfortunately they suffer from very high tail latency when you mix small Messages with large ones. I'll talk a little bit about why that's the case. And third, I'd like to introduce HOMA, which is a new protocol we've developed at Stanford that actually was designed in a clean slate redesign to handle data center workloads like these. And in fact, it does quite well on those workloads and can reduce tail latency by an order of magnitude or more. So I'll tell you a little bit about Houma. So let's dive in. First, workloads. Historically, AI workloads have consisted of enormous Transfers between machines. That's all that really mattered. Gigabytes of data for things like weight gradients and and so on. In these workloads, what you really care about is throughput. How many gigabits per second you can pump through the pipes. And these are relatively easy workloads for networks because if it takes a while to set up the connection and start the transfer, it doesn't matter. The transfers go on for so long that all that really matters is the throughput. And so when these Environments, TCP and RDMA perform pretty well. By the way, when I say RDMA, what I really mean is RoCE, RDMA over converged ethernet, which is the underlying transport that's used by RDMA for most purposes today. So anyhow, the old workloads, big transfers, throughput matters. The legacy protocols work pretty well. However, it appears that the workloads are changing. They're becoming more granular with smaller chunks of computation and smaller exchanges of Data. And this seems to be particularly true in the world of inference and also in agentic workloads. Not so much for training workloads are still massive transfers. And so what's happening is that more and more there are small message exchanges, typically for things like metadata and coordination, such as checking to see if a particular entry is present in a kv cache that's distributed, or doing barrier synchronization at the end of periods of compute. And for these workloads, what really matters is latency. That is, what's the round trip time to send some small piece of data across the network, do a little bit of computation, and get a small result back again? And in fact, it isn't just just latency or average latency that matters. What really matters is tail latency. That is you'd like to know that if we send a whole lot of small messages, all of them will complete quickly. So for example, we typically measure things like ninety ninth percentile latency. And if we have high tail latency that can limit the overall throughput of the system. So here's an example. Suppose a common thing is to take a workload and split up across several nodes which do intensive computation using their GPU's for some period of time. And then once they've all finished their computation, you do some small exchange between the nodes, exchange data, metadata, and then it'll go on to the next round of computation. And while that exchange is happening, that synchronization is happening, the GPU's are sitting idle. So if even one of those Those exchanges takes a long time. It turns out the whole process stalls. You need all of those exchanges to complete before you can go on to the next phase of computation. Now if the computation phase is, say, five seconds and it takes a few milliseconds for the exchange, you know, not a problem. And that's historically what it's been. But now with the agentic workloads, we are trying to pump out tokens relatively rapidly at a regular rate. The periods of computation are getting down into sort of the millisecond time scale. And if it also takes milliseconds to do that synchronization, then you're wasting a significant fraction of your GPU's resources waiting for the the synchronization to occur. So I'm curious. I'd like to just do a quick audience poll here. Is there anybody here where you have reason to believe that the latency of small messages is impacting the overall throughput of your applications? If so, can you just raise your hand? See is there anybody out there today? Actually, more hands than I expected. So quite a few people. They are raising their hands. I think this problem is likely to get worse as the trends continue. So what's going on? Why is tail latency bad? Well, typically, the cause is congestion resulting from in cast. So in cast is when several nodes all decide simultaneously to transfer data to some destination node. And if they all send large messages, well, the links are the same everywhere in the network. So three nodes can Transfer three times as fast as one node can possibly receive. And so what happens is that packets accumulate at the last hop going to that destination in the top of rack switch at its egress port for the destination node. Then if some other node decides it wants to send a short message to that same destination, the short message gets stuck behind the long ones in the queue there. And actually, that causes delay. And in the worst case, so many packets are I've that the switch runs out of buffer space and it has to drop packets and then there are there are timeouts and retransmissions that make everything even worse. So somehow we need some way to reduce the congestion in those queues. Somehow we have to get the sending nodes to stop sending so fast, so the queues don't just build up without limit. So the way this is done historically, virtually all network protocols before HOMA, including TCP and RDMA. Congestion control is the responsibility of the sender. So senders somehow have to figure out that congestion is happening and they have to slow down their rate of transmission. Now you might wonder why are senders doing it because the congestion is way over at the other end of the data center network. How does the the sender find out? Well, in the in the old, really old days, the way they would find out is the queues would overflow and packets would get dropped. The sender would detect the packets got lost because it wouldn't get acknowledgements back and it would assume that means there's And then slow down its rate of transfer. That's really expensive. So today there are better techniques that mostly involve the switches providing information. So a top of rack switch, when it sees that the queue length for an egress port has reached some threshold starting to fill long before the queue overflows, it starts marking all of the packets to pass through with what's called early congestion notification, ECN. Arcing. And so when those packets pass through to the receiver, the receiver sees the marking in the packets. And then when it communicates back to the sender next, for example, to send an acknowledgment, then it includes that marking that goes back to the sender. And now the sender sees that the sender realizes, oh, there's congestion someplace. I've got to slow down my rate of transmission. So that's the basic idea. Unfortunately, getting this right is really hard. Really hard. It's very hard for the congestion to figure out exactly how to set its rates because it gets one bit of information. There's congestion someplace. There are multiple senders all sending to the same destination. They're all trying to make adjustments simultaneously. How much do you cut back? And how do I know when I can ramp up again? And even worse, it's really hard to do this in a way that's stable because there's control lag. And That is, it takes time before the sender finds out that there's congestion. And in fact, using this process, it typically takes several round trips for the sender to gradually This rate to get just the right rate to match the available bandwidth. But by the time you do that, in a network, things have changed. New transmissions have started or old ones have finished. And so these systems tend to never stabilize. They're constantly oscillating between sending too much and sending too little. Now, this problem has been around for a long time. It's been known in the research community for more than twenty years now. There have been tons of papers published on it. There have been some improvements made. That's undeniable. But we're still a long ways from anything that works well. And the problem is with the fundamental nature of it doing the the congestion control. Underside. It just doesn't work very well. So you end up with a lot of queue build up. And in fact, you can see the only way to find out that there's congestion is if there's queues. And so by that point, we're already experiencing delays. So that's a problem. There's one other problem with TCP and RDMA also is that their their basic data model is a byte stream, just a stream of bytes with no differentiation in it. So if you send a series of messages, say, through a TCP socket, they get serialized into That stream. And on this slide, I've, you know, I've shown the messages appear like they have different colors in the stream. Well, there are no colors in real life. TCP has no idea where the message boundaries are. And that also makes life hard. For example, you don't know how much more data is coming. If you knew how big the message was, you know how much more is coming. And you can't prioritize short messages, which we'd really like to do, get the short messages through faster. And you could end up with what's called head of line blocking, where somebody sends a series of messages To the same destination, and they send two really large ones and then a small one after that that gets stuck behind them in that stream. And so it gets delayed. And again, you have tail latency issues. So all in all, TCP and RDMA are just not well suited to this environment. So what do we do? Well, what I'd like to do next is tell you about a new protocol called Homa that we developed at Stanford, which was based on a completely clean slate redesign for network transport. If you could start from scratch and rethink how You do transport for data centers. How would you do it? And it turns out, in Houma, virtually every major design decision is different from TCP and RDMA. TCP, for all the amazing things that's done, is just not a good match to today's data centers, nor RDMA. So what Homa does particularly well is to manage a combination of large and small messages, and to make sure that the messages, short messages, have really low latency. So this started off as a PhD dissertation for one of my students, Ben I'm Montessori. And then the results were so great that I decided to make it my personal project to see if we could get it out of the lab and into production. As you may know, I'm not like most professors in that I love to code. And so I turned this into my own programming project. I mention three things. First, it's Message based, not stream based. In fact, the fundamental unit at HOMA is a remote procedure call, which consists of two things a request message sent from a client to a server, and then a response message returned back from the server to the client. So the key thing here is that Homa knows about message lengths. They're buried in the transport all the way down to the bottom. And this has a bunch of advantages. First, it allows us to predict the future. As soon as a receiver gets the first packet of a message, it knows exactly how much more data the sender wants to send. And that's so so much more information for doing congestion control. Second, Houma prioritizes shorter messages. It uses SRPT, shortest remaining processing time first to try and prioritize shorter messages. And third, because messages are all independent, they're not serialized into a stream. Every message is independent. Shorter messages can Bypass long ones so they don't get queued behind long messages. The second thing about Houma that's different is that it controls congestion from the receiver. Now, when you think about it, this makes sense because the congestion happens primarily at that last downlink to the receiver. And so the receiver has way more information. In fact, with Homa, as soon as it gets the first packet of a message, it knows exactly how much more is So it has essentially complete information about congestion, and it can therefore respond to congestion much more quickly and much more precisely. The way things work with HOMA is that when a sender has a message to send, it breaks it up into packets, but it only transmits the first few packets, those are called unscheduled packets, to the receiver. Packets after that are called scheduled packets, and they only get transmitted when the receiver asks for them. So the receiver will send grant packets back. It'll paste them out and send those back to the

## Slides
