Designing for the exabyte era

Dropbox's journey from a handful of servers to tens of thousands of machines with millions of drives has been marked by deliberate, incremental hardware evolution. The company's approach—pioneered during the 2015 Magic Pocket migration that brought over 90% of customer data onto custom-built infrastructure—has always treated hardware as a strategic advantage rather than a commodity. Now, with the launch of its seventh-generation platform, the company is making its most significant architectural leap yet.

The new generation introduces three tiers for traditional workloads—Crush for compute, Dexter for databases, and Sonic for storage—alongside two new GPU classes: Gumby and Godzilla. Underpinning this refresh are three major infrastructure changes: substantially higher storage bandwidth, roughly double the available rack power, and a redesigned storage chassis engineered to minimize vibration and thermal issues.

Dropbox seventh-generation hardware

Three pillars of the hardware strategy

Before committing to specific designs, the Dropbox team established a set of guiding principles derived from lessons learned across previous generations and emerging industry trends.

Embracing rapid hardware evolution

The hardware landscape is shifting quickly. CPUs continue gaining cores and per-core efficiency, networks are migrating toward 200G and 400G speeds, and storage areal density has finally reached a new plateau. Dropbox's strategy aimed to harness these trends rather than chase every specification. The team focused on what actually mattered for its workloads: performance per watt, rack-level efficiency, and enabling features that would support AI-driven products.

Deep supplier partnerships

Rather than proceeding as a passive buyer, Dropbox co-developed its storage platforms directly with suppliers. This collaboration provided early access to higher-density drives and high-performance controllers, plus the opportunity to tune firmware for specific workloads. These partnerships proved essential for addressing the acoustic, vibration, and thermal challenges inherent in dense system designs.

Software-first co-design

Software teams were brought in early to identify what would genuinely move the needle for their services. Their goals—packing more compute into each rack, enabling GPU support for AI and video processing, and improving database responsiveness—shaped the platform's direction. This wasn't just about designing servers; it was about building platforms that elevated the services running on them.

Platform-by-platform breakdown

Compute: from Cartman to Crush

The CPU selection process started with a broad field of over 100 processors, narrowed down through rigorous testing. Criteria included maximizing system throughput at server and rack level, reducing process latency, improving price-performance for Dropbox workloads, and ensuring balanced I/O and memory bandwidth. Using SPECintrate benchmarks, the team compared performance per watt and per core, ultimately selecting a chip that delivered 40% better performance than the previous generation.

The leap from the sixth-generation Cartman platform to Crush centered on a processor change: the 48-core AMD EPYC 7642 "Rome" chip gave way to the 84-core AMD EPYC 9634 "Genoa." The upgrade delivered substantial gains across the board:

  • 75% more cores per socket (48 to 84), improving bin packing for containerized services
  • 2x memory capacity (256GB to 512GB), benefiting memory-heavy workloads
  • DDR4 to DDR5 transition, providing higher bandwidth
  • 25Gb to 100Gb networking, keeping pace with growing internal traffic
  • NVMe gen5 support, accelerating local disk access and boot times

All of this was achieved while retaining the compact 1U "pizza box" form factor, allowing 46 servers per rack without additional space requirements.

Legend: Red dot indicates processor selected

Databases: single-socket efficiency with Dexter

For database workloads, the focus was on CPU performance rather than core count. The Dexter platform retains the same number of cores as its predecessor but gains 30% more instructions per cycle (IPC) and a higher base frequency, which increased from 2.1GHz to 3.25GHz. A notable architectural change was the shift from dual-socket to single-socket design, eliminating inter-socket communication delays.

These changes produced measurable results: up to 3.57x less replication lag—the delay between writing data to a primary system and copying it to a secondary one. That improvement had a significant impact on high-demand workloads like Dynovault and Edgestore.

During development, the team recognized substantial overlap between the database and compute platform requirements. Consolidating both onto a single vendor platform simplified the support stack, making component, firmware, driver, and OS updates easier to manage. The resulting dual-use design removed key scaling bottlenecks without compromising density or efficiency.

Storage: preparing for 200Gbps throughput

Storage design goals had to evolve alongside increasing drive capacities, with some drives now exceeding 30TB. Dropbox's internal performance benchmark historically required 30Gbps per PB of data, but anticipated future systems pushing past 100Gbps prompted a more ambitious target: 200Gbps throughput.

To reach this, the SAS topology was reworked to distribute bandwidth evenly across each drive while delivering total system bandwidth beyond 200Gbps. This level of scale also demanded stronger networking, so the team collaborated with Dropbox's network engineering group to design a 400G-ready data center architecture supporting the new storage systems.

Power and cooling: reworking the rack

Moving to higher-core CPUs forced a redesign of Dropbox's thermal and power strategy. Processor demands grew across the board, so the team capped thermal design power (TDP) to maximize cores per rack without exceeding cooling or power limits. Instead of relying on manufacturer "nameplate" maximums—which tend to overstate real usage—they modeled actual system workloads.

Those models showed servers drawing more than 16kW per cabinet, which would have violated the existing 15kW per-rack budget. Rather than overhauling the power infrastructure entirely, Dropbox partnered with its data center engineering team and switched from two power distribution units (PDUs) per rack to four, working with existing busways and adding receptacles.

Seventh-generation compute and storage (left); quad PDUs to meet higher power requirements (right)

That change effectively doubled available rack power, creating room for current loads and future accelerator cards. The company also collaborated with suppliers on airflow improvements, heatsink redesigns, optimized fan curves, and full-load thermal testing.

Storage density and the vibration problem

Dropbox's storage approach centers on packing more capacity into the same physical footprint. Traditional 3.5" hard drives have grown from roughly 14TB to over 30TB in a few years, lowering both cost and power per terabyte. But higher densities amplify sensitivity to acoustic and vibrational interference—particularly for Dropbox, where over 99% of the storage fleet uses shingled magnetic recording (SMR).

Dropbox continues to leap up to the next available areal density curve

The read/write head inside these drives operates with nanometer precision, so even minor vibration can knock it off track. Cooling fans spinning above 10,000 RPM in dense servers generate exactly that kind of disturbance, producing position error signals (PES) and, in severe cases, write faults that force retries—hurting latency and IOPS.

At the same time, the drives need airflow to stay near their optimal operating temperature of around 40°C; excessive heat accelerates aging and raises error rates. Balancing those opposing needs required co-developing a next-generation storage chassis with system and drive suppliers, focusing on:

  • Vibration control: acoustical isolation and damping
  • Thermals: improved fan control and airflow redirection
  • Future-proofing: compatibility with next-generation large-capacity drives

Redirecting high-velocity air to augment cooling

That effort made Dropbox one of the first to adopt Western Digital's Ultrastar HC690, a 32TB SMR drive packing 11 platters into a standard 3.5" casing—a capacity bump of more than 10% over the prior generation.

Introducing GPU-based server tiers

Supporting Dash, Dropbox's universal search and knowledge management product, required bringing GPUs into the data center. Workloads like intelligent previews, document understanding, fast search, and video processing demand high parallelism, large memory bandwidth, and low-latency interconnects that CPU-only servers can't economically provide. The seventh-generation rollout therefore introduced two GPU-enabled server tiers: Gumby and Godzilla.

GPU generations

  • Gumby builds on the Crush compute platform but accommodates a broad range of GPU accelerators, with TDP support from 75W to 600W and both half-height half-length (HHHL) and full-height full-length (FHFL) PCIe form factors. It targets lightweight inference tasks such as video transcoding, embedding generation, and service-side machine learning enhancements.
  • Godzilla handles heavier lifting, supporting up to 8 interconnected GPUs for LLM testing, fine-tuning, and other high-throughput machine learning workflows.

Together, the two tiers let Dropbox scale AI across products while keeping control over performance, cost, and energy efficiency.

Key lessons from the rollout

Power and thermals now drive design decisions. Across compute, storage, and GPU platforms, power demands are rising. The shift from two to four PDUs per rack allowed higher-density systems without compromising stability. Although total rack power increased, power consumption per petabyte and per core decreased, supporting sustainability targets.

Early supplier collaboration pays off. Several of the biggest wins came from working with vendors early—redesigning hard drive acoustics, fine-tuning chassis layouts, and gaining early access to components. Co-development produced vibration-optimized enclosures and systems ready for next-generation storage.

Hardware strategy followed software requirements. The GPU tier emerged from early engagement with machine learning and video teams as AI tools like Dash came online. That input shaped platforms built specifically for inference, video processing, and hosting large language models, ensuring infrastructure was ready when the software needed it.

Looking ahead

The seventh-generation rollout positions Dropbox's in-house server hardware for continued evolution. The latest Crush and Dexter platforms use newer CPUs that deliver significant gains in instructions per cycle and transaction speed. On the storage side, the Sonic system supports higher capacities thanks to vendor collaboration. And the dedicated GPU hardware tier addresses growing AI demand.

Future infrastructure changes are already on the horizon. Heat-assisted magnetic recording (HAMR) promises large capacity gains but will demand even greater precision in acoustic and thermal management. Liquid cooling is moving from niche to necessity as compute densities climb. The current generation serves as a launchpad for that work—tightly integrated platforms, stronger supplier partnerships, and the flexibility to keep evolving.