Meta’s Open Hardware Push for AI Infrastructure

Meta’s AI infrastructure is scaling at a pace that demands new hardware approaches, and at the Open Compute Project (OCP) Global Summit 2024 the company detailed its latest contributions. A central design principle is simple: the scale of AI training compute can only be sustained on open, shared building blocks. Meta’s largest model, Llama 3.1 405B, required training across more than 16,000 NVIDIA H100 GPUs on 15 trillion tokens, and today’s models already run on two 24K-GPU clusters. Acknowledging an expected upward curve for compute demand, Meta lays out open hardware—rack architecture, platform designs, and network fabrics—as the avenue to meet growing workloads.

Networking and Rack Challenges

AI clusters are more than GPU fleets. Networking weighs heavily on training performance, with systems built on tightly integrated HPC compute and isolated, high-bandwidth networks connecting accelerators. Meta expects injection bandwidth to grow to around a terabyte per second per accelerator, with matching normalized bisection bandwidth, representing more than an order-of-magnitude jump over today’s network capacity.

These increases require a multi-tier, non-blocking fabric that can handle congestion predictably when pushed hard. Meta emphasizes that continued advancement depends on openness: new architecture, network fabrics, and system designs are most effective when built on community standards.

Catalina: A Modular High-Power Rack for GB200

Among the new contributions to OCP is Catalina, a rack-scale solution based on the NVIDIA Blackwell platform design and tailored for the GB200 Grace Blackwell Superchip. Catalina targets modularity with the goal of letting others customize their rack configurations around existing and emerging industry standards.

Its key new element is the ORv3 high-power rack, supporting up to 140kW, addressing the rising power draw of accelerators. The Catalina solution is fully liquid cooled and integrates a power shelf along with a compute tray, switch tray, Wedge 400 fabric switch, management switch, battery backup unit, and rack management controller.

Grand Teton Gains AMD MI300X Support

First announced in 2022, the Grand Teton AI platform has been expanded to support AMD Instinct MI300X accelerators, and this new version heads to OCP. The platform retains a monolithic system design that consolidates power, control, compute, and fabric interfaces—a configuration Meta says simplifies deployment and improves reliability at scale for large inference workloads.

The MI300X variant carries the platform's general upgrades, including more compute capacity for convergence on larger model weight sets, added memory for hosting larger models locally, and higher network bandwidth, enabling larger and more efficient scale-up clusters.

Networking: DSF, New Fabric Switches, and FBNIC

Supporting this scale, Meta’s Disaggregated Scheduled Fabric (DSF) reworks the network backend to be a vendor-neutral option rather than bridging proprietary switch designs. DSF’s advantages, Meta states, include improved capacity, supplier diversity, and power density. Its core consists of the OCP-SAI standard and FBOSS, Meta’s open network operating system. It also exposes RoCE endpoints to accelerators from partners such as NVIDIA, Broadcom, and AMD.

Two related hardware pieces come with DSF:

  • New 51T fabric switches built on Broadcom and Cisco ASICs.
  • FBNIC, a NIC module with Meta's first in-house network ASIC.

Partnership on Mount Diablo Disaggregated Power

Within OCP, Meta and Microsoft have deepened their collaboration. The two started co-developing the Switch Abstraction Interface (SAI) in 2018, then contributed to the Open Accelerator Module standard and SSD standardization. The current focus is Mount Diablo, a disaggregated power rack design built on a 400 VDC unit. Its scalable architecture is meant to support more AI accelerators per IT rack by decoupling power delivery from the compute racks.

Meta’s position remains that AI hardware must be peer to the open software runtime, placing the full stack within the OCP's remit. By inviting others to work on these systems, the aim is to turn shared infrastructure problems into faster, cheaper, and adaptable hardware cycles for the size of training jobs now in orbit.