OCP 2025: Meta's Next-Generation AI Network Fabrics Go Open

At the Open Compute Project (OCP) Summit 2025, Meta is detailing the evolution of its network fabric architectures for AI training clusters. Since co-founding OCP in 2011, Meta has shared its data center designs and open-sourced its FBOSS network operating system. Continuing that trajectory, the company is introducing several new network hardware platforms and architectural concepts aimed at supporting ever-larger AI workloads. Following are the key announcements, each representing a milestone in Meta's ongoing commitment to disaggregated, open infrastructure at hyperscale.

Scaling Scheduled Fabrics with a Dual-Stage DSF

Meta's Disaggregated Scheduled Fabric (DSF), first introduced at OCP 2024, is a VOQ-based system built on the OCP-SAI standard and Meta's FBOSS software. It uses a standard Ethernet RoCE interface to connect to endpooints and accelerators from vendors including Meta's own MTIA. This year, Meta details an evolution of DSF to a 2-stage architecture. This dual-stage design supports a non-blocking fabric that can interconnect up to 18,432 XPUs, providing a building block for clusters that span single or multiple regions.

The new dual-stage DSF architecture supports non-blocking fabric, enabling interconnect between a larger number of GPUs in a cluster. At Meta, we’ve used it to build out clusters of 18k GPUs at the scale of entire data center buildings.

Shallow-Buffer NSF for Gigawatt-Scale Clusters

Alongside DSF, Meta has been working on a separate architecture called Non-Scheduled Fabric (NSF). This new design uses shallow-buffer OCP Ethernet switches and is engineered to deliver low round-trip latency. It supports adaptive routing for load balancing and congestion avoidance, serving as a foundational component for gigawatt-scale AI clusters.

NSF — Three-tier Non-Scheduled Fabrics for building scale AI clusters.

Expanding the 51T OCP Switch Portfolio

Last year, Meta joined OCP with the Minipack3 and Cisco 8501 switch platforms, both 51T Ethernet switches offering 51.2 Tbps across 64x OSFP ports. These designs, based on Broadcom Tomahawk5 and Cisco Silicon One G200 respectively, require no retimers and run FBOSS. They now form the basis for Meta's next-generation frontend and backend fabrics.

This year, Meta adds the Minipack3N to this family. The new OCP switch leverages the same system design as Minipack3 but is based on NVIDIA's Ethernet Spectrum-4 switching ASIC.

The Minipack3N, a 51.2 Tbps switch (designed by Meta and manufactured by Accton) based on the NVIDIA Spectrum-4 Ethernet switching ASIC.

Evolving FBOSS and SAI for New Fabrics

Meta is continuing to develop OCP-SAI, using it to onboard new fabrics, switch hardware, and optical transceivers into FBOSS. In collaboration with vendors and the OCP community, SAI has been extended to support DSF, NSF, and enhanced routing schemes for modern data center and AI workloads.

Meta has widely deployed the 2x400G FR4 BASE (3-km) optics introduced last year to support its 51T platforms, both across backend and frontend networks and the DSF. This year, Meta expands its optical portfolio with both a long- and shorter-reach solution:

  • 2x400G FR4 LITE (500-m) optics: Optimized for the majority of intra-data center use cases, this variant supports fiber links up to 500 meters and is intended to help accelerate optics cost reduction for shorter-reach applications.
  • 400G DR4 OSFP-RHS optics: Meta's first-generation DR4 solution for AI host-side NIC connectivity.
  • 2x400G DR4 OSFP optics: Deployed on the switch side to provide host-to-switch connectivity.
The 400G DR4 (left), 2x400G DR4 (center), and the 2x400G FR4 LITE (right).

ESUN: Bringing Ethernet to the Scale-Up Domain

Meta is also a founding participant in the Ethernet for Scale-Up Networking (ESUN) initiative, a new workstream that launched within the OCP Networking Project at the 2025 OCP Global Summit. ESUN is designed as an open technical forum for operators and vendors to address the high-performance interconnect demands of scale-up networks, which are the connections among AI accelerators within a cluster.

ESUN focuses specifically on the network functionality aspect of scale-up systems, addressing technical challenges related to managing and transmitting data traffic across switches. This includes defining best practices for:

  • Protocol headers
  • Error handling mechanisms
  • Achieving lossless data transfer across a network

Work will focus on Ethernet framing and switching layers to ensure robust multi-hop topologies, and the group intends to align with other standards bodies including UEC and IEEE. Founding members alongside Meta include AMD, Arista, ARM, Broadcom, Cisco, HPE, Marvell, Microsoft, NVIDIA, OpenAI, and Oracle. Meta's contributions include providing technical leadership on ESUN requirements and sharing best practices from its own commercial Ethernet fabrics, with the goal of ensuring interoperable rather than proprietary solutions.