AI Workloads Are Rewriting the Network Engineer's Job
At this year's @Scale: Networking event—the largest yet—engineers from Meta, ByteDance, Google, Microsoft, Oracle, AMD, Broadcom, Cisco, and NVIDIA gathered to compare notes on building and operating AI networks. The sessions made one thing clear: networking is no longer just a supporting layer for AI, but the critical substrate that determines whether massive GPU fleets actually deliver on their promise.
The Pressure Points: Scale and Speed
Physical Infrastructure Is Being Built at Unprecedented Scale
AI companies are planning hundreds of billions of dollars in infrastructure over the next several years. Meta's own trajectory illustrates the trend: gigawatt-scale clusters like Prometheus and Hyperion, clean power partnerships, and the largest transoceanic fiber cable systems in the world. In the short term, the company has even resorted to "sprung structures" to get capacity online faster.
Workloads Are Shifting Faster Than Ever
The AI training landscape has moved from straightforward large-scale pretraining to a much more varied mix. In just the last 9-12 months, the industry has seen rapid adoption of mixture-of-experts models, reasoning models, reinforcement learning, post-training, synthetic data generation, and distributed inference—each with distinct network requirements. Meta's own journey from 4K to 24K to 129K-GPU clusters (all Ethernet/RoCE-based) in under two years shows how quickly the goalposts move on both performance and reliability.
Network as the Great Abstraction Layer
From the model's perspective, AI infrastructure should present itself as one gigantic GPU. That abstraction work falls on the network. The challenge is multi-faceted:
- Co-design with the AI stack: Bridging varying distances and bandwidths between scale-up and scale-out domains, while accommodating different accelerators, NICs, and fabrics, demands tight tuning of NICs, routing, and congestion control in concert with GPU-based frameworks.
- Reliability: The network must not only deliver performance, but also operate with high reliability—identifying and reacting to failures seamlessly.
- Innovation and optionality: With constant change both above (models, workloads) and below (infrastructure), the network stack needs to blend high-performance computing capabilities with open, scalable distributed systems principles.
Two Threads Running Through This Year's Talks
The session program coalesced around two major themes. The first half focused on the physical layer: switch topologies and control planes, NIC and host networking, and operational approaches for scalability and high reliability. The second half was model-centric: parallelism design, job-level debuggability, scaling pretraining, and adapting to new use cases like reinforcement learning, mixture-of-experts, and inference.
The event also featured keynotes from Meta and Microsoft on what's next for AI and networking, along with a vendor panel that included key GPU and network ASIC makers.
All talks from this year's event, along with other @Scale: Networking sessions, are available on the @Scale YouTube channel. Meta continues to organize these events across Systems & Reliability, AI & Data, and Product (the latter coming in October) to keep the community sharing challenges and innovations.



