Why Dropbox jumped to 400G ethernet

Dropbox’s infrastructure teams have been preparing for a bandwidth crunch driven by new AI features and denser server designs. Products like Dropbox Dash and Capture, along with rising video and image uploads, are pushing throughput demands past what 100G ethernet can economically handle. Rather than scale out its existing 100G architecture with more devices and optics—a stopgap the team estimated would need replacing within 24 months—Dropbox designed its first data center around 400G ethernet.

The decision was heavily influenced by sustainability goals. A single 400G port with optics is more cost-efficient and consumes less power than four separate 100G ports delivering the same aggregate bandwidth. The company also calculated that deploying a temporary 100G architecture would create large amounts of e-waste in short order. The move to 400G delivers a fourfold increase in per-link data rate using digital signal processing chips that support pulse-amplitude modulation and forward error correction.

The new design targets efficiency gains in four network areas: the fabric core, top-of-rack switch connections, data center interconnect (DI) routers, and optical transport shelves.

Fabric core: Zero-optic design reduces power

The fabric core retains Dropbox’s production-proven quad-plane topology but swaps 3.2T 32x100G switches for 12.8T 32x400G devices. This preserves desirable characteristics of the previous design—non-blocking oversubscription rates, small failure domains, and scale-on-demand modularity—while quadrupling speed.

Power requirements did not grow, thanks to the use of 400G direct attach copper (DAC) cabling for spine-leaf interconnections. 400G-DAC is electrically passive, requiring virtually no additional power or cooling. That choice fully offset the increased energy draw of the faster switch chips. Measured against the legacy 100G data center, the new 400G fabric core is 3x more energy efficient per gigabit.

A high-level overview of our 400G network architecture

The main constraints of 400G-DAC are its three-meter range and thicker cable profile. Dropbox addressed this with careful planning and lab mock-ups of device placement, port assignments, and cable management strategies. The result is an “odd-even split” main distribution frame (MDF) design optimized for the physical layout.

A simplified version of our 400G data center MDF racks using 400G-DAC interconnects. Spine switches are stacked in the center rack, connected to leaf switches that are striped evenly between the adjacent racks. Only DAC cables to the first leaf switch in each of the odd (left) and even (right) racks are pictured. This design was repeated four times for each of the data center’s four parallel fabric planes

Top-of-rack interconnect: Backward compatibility

The optical fiber plant connecting top-of-rack switches to the fabric core had to meet three requirements: support for both existing 100G and next-generation 400G top-of-rack switches, runs up to 500 meters for multi-megawatt-scale deployments, and maximum reliability while minimizing power and material costs.

After testing various transceivers, Dropbox selected the 400G-DR4 optic. It supports existing 100G top-of-rack switches via fan-out to 4x100G-DR links, with the built-in digital signal processor handling signal conversion between 400G and 100G without adding computational overhead to the switches. The optic’s 500-meter maximum range covers even the largest facilities. At 8 watts of max power draw, the 400G-DR4 is more efficient than four 100G-SR4 optics at 2.5 watts each (10W total). It also runs over single mode fiber, which requires 30% less energy and materials to manufacture than the multi-mode fiber used in prior 100G deployments.

Consolidated data center interconnect

The DI layer underwent a complete redesign. Previously, the network used three separate device roles: one tier for cross-datacenter traffic and another tier for external traffic between data centers and POPs. With 400G technology, Dropbox consolidated these three devices into a single data center interconnect.

Our old data center interconnect design

Feature advancements made this consolidation possible. Class-based forwarding, which wasn’t available when the tiered design was first built, allows quality-of-service markings to logically separate traffic over different label-switched paths with appropriate priorities. The new DI tier also uses MPLS RSVP TE in place of ECMP, making the data center edge bandwidth-aware and improving resiliency and efficiency. Route aggregation, community tags, and advertising only the default route down to the fabric streamline routing further.

Our new data center interconnect design

The consolidated DI layer delivers several advantages:

  • 60% reduction in the number of devices at the tier, improving space utilization, energy efficiency, and device cost.
  • Backward compatibility with 100G hardware, allowing partial upgrades while preserving existing investments.
  • Scalability up to eight times current maximum capacity for future expansion.

The optical transport tier is a dense wavelength division multiplexing (DWDM) system handling all data plane connectivity between the data center and the backbone. Using two strands of fiber between the data center and each backbone POP in the metro, the architecture provides two 6.4 Tb/s tranches of diverse capacity, totaling 12.8 Tb/s. The system can scale to 76.8 Tb/s (38.4 Tb/s diverse) before additional dark fiber is required. Without the DWDM system, a pair of fiber would carry at most 400 Gb/s.

One of the two 6.4 Tb/s diverse data center uplinks spans

This generation uses 800 Gb/s tuned waves compared to 250 Gb/s in the previous generation, providing greater density and lower cost per gigabit. The optical tier was also engineered for flexible deployment of 100G/400G client links, which helped Dropbox adapt to equipment delivery delays caused by commodity shortages and still bring the 400G data center online on schedule.

Field lessons from the first 400G rollout

Dropbox’s first 400G data center has been serving customer traffic since December 2022, with additional facilities planned before the end of 2023. The deployment was not without friction. The routers, switches, cables, and optics were all among the first of their kind, so the team built a dedicated 400G test lab with a packet generator capable of emulating future-scale workloads. Every component was physically and logically stress-tested for performance and multi-vendor interoperability.

That testing surfaced a critical gap before production: a widely deployed 100G top-of-rack switch was missing support for the 100G-DR optic chosen to link existing gear to the new 400G fabric. Because the issue was caught early, the vendor was able to ship a patch adding support.

Supply chain volatility also shaped the design phase. The team qualified multiple sources for every component. When the vendor supplying 400G DI devices backed out a month before launch because of a chip shortage, a contingency plan was already in motion. Since 400G QSFP-DD ports accept 100G QSFP28 optics, the team used 100G devices in the DI role temporarily and swapped in permanent 400G units later.

Extending 400G across the network

The success of the initial rollout has accelerated plans elsewhere in the production network. Data centers based on the same design are scheduled for US-CENTRAL and US-EAST before the end of 2023. Test racks of seventh-generation servers with 400G top-of-rack switches are already live in US-WEST, with production-scale deployment expected in early 2024. The 400G rollout will extend to Dropbox’s backbone during 2024 and 2025.

On the horizon is 400G-ZR+, an emerging long-haul optical technology that could replace 12-foot-tall optical transport shelves with a pluggable transceiver roughly the size of a stick of gum.

Daniel King, one of our data center operations technicians, holds a pluggable transceiver in front of the equipment it will eventually replace.