The Metro Problem

Meta’s backbone is split into two networks: Classic Backbone (CBB) and Express Backbone (EBB). CBB handles the flexible, global reach from data centers (DCs) out to points of presence (POPs), running traditional IP/MPLS-TE. EBB is the high-capacity DC-to-DC network, running a heavily customized stack with the Open/R routing protocol and in-house traffic engineering. It is EBB that poses the hardest scaling problems, driven by relentless traffic growth since its first deployment around 2015.

The core difficulty is physics and logistics: interconnecting growing DCs requires enormous quantities of long-haul fiber, often spanning hundreds of miles. Site selection for new DCs considers many factors beyond network ease, and building out dedicated long-haul fiber to each new location on demand is slow and painful.

Meta's answer is a three-pronged evolution it calls 10X Backbone, which reshapes the metro architecture, scales the IP platforms, and integrates IP with optical transport to handle a tenfold increase in capacity demands.

Pre-Building the Metro

The first technique attacks the timeline problem of connecting new DCs. Instead of planning fiber from scratch for every campus, Meta pre-builds key metro components.

The design uses two pre-built rings of fiber to provide scalable metro capacity. Long-haul fibers are connected to these rings, and two POPs are established to provide connectivity toward remote sites. New DCs are then attached directly to the rings, enabling or increasing capacity between the DC and the POPs quickly.

This architecture offers three advantages: it simplifies the construction of DC connectivity and the WAN topology, it provides a standardized and scalable physical design, and it cleanly separates the metro and long-haul network layers.

Scaling Platforms: Up and Out

IP platform scaling for EBB operates along two axes: scaling up and scaling out. These are not mutually exclusive, and Meta has used both in its 10X journey.

Scaling up relies primarily on vendor technology in two forms:

  • Larger chassis: A 12-slot chassis offers 50% more capacity than an 8-slot unit. This brings significant trade-offs, including tougher mechanical and thermal designs, higher power density per rack, a larger number of ASICs with associated control-plane programming complexity, higher interface and cabling counts, increased network OS complexity, and greater infrastructure design challenges—though it simplifies NPI when the ASIC and line-card technology remain unchanged.
  • Faster interfaces: Moving from 400G to 800G line cards doubles capacity by leveraging modern ASICs. This path introduces similar challenges: complex thermal designs, higher power requirements, difficult NPI with a new forwarding pipeline, cabling density issues, and specific support for 800G-ZR+ transceivers.

Scaling out is more within Meta's control and has two historical flavors in EBB:

  • Adding more backbone planes: Going from four to eight planes doubles global capacity. However, implementation is highly disruptive—fiber restriping must be coordinated across many locations simultaneously—and adds power and space requirements globally without altering per-rack power density. It can complicate routing across planes with uneven capacity and may require stronger interconnects, though it does not introduce new technology.
  • Multiple devices per plane: This more surgical approach scales capacity only at chosen locations. It is disruptive at the target site and requires careful planning. The interconnect full mesh must be extended to Nx devices, introducing new failure modes where a single device failure impacts a subset of the backbone in that plane. Network operations also become more complex around device sets for software upgrades and maintenance.

The ZR Power Play

The third scaling technique fundamentally alters space and power economics. By adopting ZR pluggable transceivers, Meta has eliminated standalone transponders.

Before ZR, each transponder consumed up to 2kW to deliver 4.8–6.4Tb, and a clear operational demarcation separated IP and optical layers. With ZR, that function is embedded in the router's pluggables, which consume only 10–15W of incremental power each.

The result is a dramatic aggregate power reduction of 80 to 90%, despite the added draw on the routers. The impact is seen across several dimensions:

  • Space and power efficiency: The same backbone capacity fits in a much smaller space and power envelope. Rack allocation shifts from 90/10 (optical/IP) to 60/40 in favor of IP. Fiber landing capacity improves fourfold, from one to four fiber pairs per rack, since standalone transponders are gone.
  • Simplified deployments and operations: Installing pluggables is easier and more predictable than wiring up transponders, and fewer active devices simplify network operations.
  • Vendor diversity: ZR enables interoperability across vendors.
  • New complexity: The optical channel now terminates inside IP devices, blurring domain ownership. A clear IP/optical demarcation is harder to maintain, and optical telemetry collection is bound to IP devices, adding CPU load on the routers.

Beyond the Campus: AI Backbone

Growing demand for GPU clusters has pushed megawatts footprints beyond what any single campus can support, even with adjacent land. Because cluster performance is latency-sensitive, Meta must find expansion sites within bounded geographic proximity, expanding outward until regional scale is achieved.

Once candidate sites are identified, fiber-sourcing teams assess the feasibility of very high-scale connections. Significant construction is often required to lay new fiber. The distance to a site dictates one of three connectivity solutions:

  1. FR plugs: For buildings up to 3 kilometers away—extending the 2-kilometer standard spec under adjusted loss and connector assumptions.
  2. LR plugs: Longer reach optics stretching to 10 kilometers.
  3. ZR + DWDM: Beyond 10 kilometers, active optical components multiplex and amplify signals. Multiplexing cuts fiber count by a factor of 64 versus FR/LR, making this the path to real scale.

These longer-reach links run over a design with optical-protection switching and C+L-Band 800G ZR technology to reduce port consumption on IP platforms versus IP-layer protection. Each fiber pair carries 64x 800G channels, yielding 51.2T per pair, and capacity scales horizontally by adding fiber pairs.

Current requirements sit at the lower end of the distance range, avoiding the need for intermediate amplification sites that become necessary beyond roughly 150 kilometers—fortunate, since those sites would be substantial given the fiber volume requiring amplification. The protection-switching design does add operational work: external tooling must track whether underlying connectivity is in a protected or unprotected state. Yet the primary benefit remains in reducing required IP platform ports.

The scale of these interconnects is stark: a single AI Backbone site-pair today is twice the size of the entire global backbone built over the last decade. That volume of equipment and connections now drives Meta's focus on streamlining deployment processes and completing the physical fiber build-out.

Key Takeaways from Scaling EBB

The evolution of Meta's Express Backbone (EBB) over the past eight to nine years has been defined by unexpected acceleration. Plans initially targeting 2028 were pulled forward to 2024, driven by the need to support rapid network growth. Several core lessons emerged from this process:

  • The 10x Backbone is achievable through a combination of scaling up and scaling out innovations.
  • Pre-built, scalable metro designs allow for quicker adaptation to changing network demands.
  • Integrating IP and optical layers reduces active device counts, space, and power requirements, enabling further scaling.
  • Technology developed for the 10x Backbone was directly reusable in building the AI Backbone.

Preparing for Next-Generation Data Centers

With plans to construct city-sized data centers, the backbone architecture must evolve to keep pace. Two key directions will define this next phase.

First, leaf-and-spine architecture is identified as the next step for scaling out existing platforms. This design provides the required scale while minimizing disruptive scaling steps. Second, the initial AI Backbone rollout will proceed on schedule, with iterative refinements as more sites come online and operational maturity grows. The goal is to understand the unique traffic patterns and requirements of AI workloads as they traverse the optical network in real-world conditions.