Back
Stanford professor pushes Homa as a replacement for TCP in AI datacenters
SiTech AI Team3 min read

Stanford professor pushes Homa as a replacement for TCP in AI datacenters

John Ousterhout, a retired Stanford computer science professor, argues TCP is a poor fit for AI datacenter workloads and is promoting Homa, a message-based protocol he says cuts short-message latency by 13 times.

TCP, the protocol the web and most cloud computing run on, is ill-suited for emerging AI workloads, argues John Ousterhout, professor emeritus of computer science at Stanford University. His answer: Homa, which he has promoted as his "life's mission" since retiring.

"TCP, for all it has done, is not a good match for datacenters," Ousterhout said at the AI Engineer World's Fair.

Why TCP doesn't fit AI datacenters

TCP brought order to networks: flow control, guaranteed delivery, handshakes, congestion control. Its data model, though, is byte streams: messages are serialized as one stream with no priority, so receivers cannot tell longer transfers from shorter ones, and senders can only guess how much to slow down.

That is tolerable for ordinary traffic, but not for latency-sensitive AI workloads: AI labs move model weights, KV caches and checkpoints between expensive GPUs, while large transfers share bandwidth with short bursts from agents and control tasks. "For these workloads, what really matters is latency," Ousterhout said; even a millisecond of delay leaves an expensive GPU idle.

How Homa works

Homa is a clean-slate rethink of congestion control. It is message-based, not stream-based: as with RPC, message lengths are explicitly defined. Unlike TCP, congestion control sits with the receiver: it knows from the first packet how much data is coming and schedules packets accordingly, prioritizing short messages via SRPT.

The result: p99 latency for short messages is 92 microseconds on Homa versus 1.2 milliseconds on TCP, 13 times faster, on a 100 Gbps network at 80% utilization. Even the longest messages are about twice as fast.

Standardization and adoption

Work on Homa began as a PhD dissertation published in 2019 by Behnam Montazeri, now a staff engineer at Google. Ousterhout is drafting an IETF standardization document and working to upstream it into the Linux kernel; in March it was backported to Red Hat Enterprise Linux 8 and 9.5.

Adding it is simple: compile Homa from its GitHub source, then install the module into Linux kernels on clients and servers, no reboot required, he told The Register. "Homa works side by side with TCP, so you can gradually move applications from TCP to Homa," he wrote; running Homa even makes remaining TCP applications faster. He is now prototyping with a large financial services company.

Skeptics and alternatives

Not everyone is convinced: network architect Ivan Pepelnjak's scathing 2023 paper questioned Ousterhout's TCP performance claims and called Homa a solution looking for a problem.

Others bypass TCP too: databases use DPDK, storage networks turned to NVMe over Fabrics, and Google built QUIC, the basis for HTTP/3. Specialized RDMA fabrics and AWS's Scalable Reliable Datagram offer their own fixes, but for now, TCP remains the champion.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.