6/9/2026 at 8:37:32 PM
Oh man, James Hamilton blog posts, I love these things! (Edit: for more concrete details, the Arxiv paper linked from the blog post is here https://arxiv.org/pdf/2604.15261 and the amazon.science link has some higher level view of the details https://www.amazon.science/blog/how-flat-is-replacing-fat-in... )> The results were striking: compared to traditional fat-tree networks, RNG (Resilient Network Graphs) uses 69% fewer routers, delivers 33% higher throughput, cuts network power by 40%, and lowers operating costs by 27. In early 2026, RNG became the default design for most newly built Amazon data centers globally.
> For cabling, they developed the ShuffleBox—a passive optical device whose internal wiring combined with randomized ShuffleBox-to-ShuffleBox cabling yields “quasi-random” graphs that behave like truly random graphs.
This is pretty incredible, random layouts of networks that have on-average better properties...
I'm really curious about the long tail of performance though. What is the worst case scenario here? And are there some better case scenarios? Uniformity in Clos networks is pretty great, but many loads don't need uniformity, and if these RNG-based networks have non-uniformity, perhaps that has operational characteristics that can be helpful or harmful.
by epistasis
6/9/2026 at 9:51:50 PM
"Performance guarantees are stochastic rather than deterministic. The worst case performance (for metrics such as number of hops and oversubscription) is known, but for RNG our models are stochastic (i.e., the worst case performance is known with high probability). This is a weaker limitation than it might appear. Fat-tree guarantees are also effectively stochastic once you account for real-world failures, which are frequent at scale. RNG simply makes the stochastic nature explicit and designs for it from the start."by UltraSane
6/9/2026 at 10:50:03 PM
Well I guess I'd like to see those guarantees, but more specifically, the variance of them.I think Section 9, and Figures 13/14 in the Arxiv preprint sort of address this, but it doesn't mention anything about accounting for real-world failures in fat trees. I haven't had a chance to read it all, though...
by epistasis