Numbers That Put Scale In Perspective

Each edition of the roundup tends to surface a handful of figures that help frame just how large modern infrastructure and platforms have become. This week’s crop spans stellar astronomy through app-store economics, and the numbers are, as usual, a mixed bag of the absurd, the impressive, and the sobering.

  • Starting with the cosmos: roughly 1 septillion stars, 40 quintillion black holes, and 10 septillion planets are estimated to exist in the observable universe—figures that, as one observer noted, come in just a tad under Apple’s yearly revenue.
  • Speaking of Apple, the company reported $123.9 billion in revenue for the quarter ending Dec. 25, up 11% year over year, with services contributing $19.52 billion, a 24% YoY increase. In parallel, Microsoft’s revenue hit $51.7 billion with cloud revenues up 46%.
  • On the platform side, Apple paid out $60 billion to App Store developers in 2021, while users spent 3.8 trillion hours on mobile apps during the year, up 30% from 2019. Instagram now claims 2 billion monthly users, and 8 mobile games each crossed $1 billion in annual revenue in 2021.

The music industry continues its long-tail trend: the top 200 new tracks account for just 5% of total music streams, with older music making up 70% of the US music market and growing at the expense of new releases—a key reason catalogs are being acquired.

Not all the numbers are rosy. 77TB of research data was lost at a university due to a backup error, a reminder that backups are only as good as the restore process they support. On the security front, Cloudflare blocked an almost 2 Tbps multi-vector DDoS attack, and network-layer DDoS attacks increased by 44%.

The Scale Of The Platforms We Run On

Cloud providers and large-scale platforms continue to publish eye-opening operational metrics:

  • 60 million EC2 launches occur each day, double the rate seen in 2019. IAM handles half a billion API calls per second using a hierarchical edge cache. Lambda sees over 150 million invocations per minute, API Gateway handles over 200 million calls per minute, and ElastiCache processes over 275 million hits per minute.
  • Shopify’s peak Black Friday sales on Google Cloud reached $3.6 million per minute, with 79% of traffic coming from mobile and 30TB/min of egress traffic across their infrastructure.
  • Netflix’s EVCache spans ~18,000 servers holding roughly 14 petabytes of data.
  • On the developer front, **45 million** developers are expected globally by 2030, up from 26.9 million in 2021. O’Reilly’s survey shows 90% of respondents’ organizations use the cloud (up 2%), with AWS at 62%, Azure at 48%, and Google Cloud at 33%. Amazon’s usage actually dipped from 67%, while 20% plan to migrate all applications and 47% pursue a cloud-first strategy.

Life, the Universe, and Everything In Between

A few more numbers round out the picture. Patreon creators have earned $3.5 billion lifetime, with $1.5 billion in 2021 alone, a 50% uptick over 2020. Reddit saw 366 million posts, 2.3 billion comments, and 46 billion upvotes in 2021. SpaceX plans a record 52 launches in 2022, up from 31 last year and 26 in 2020. And as ever, PHP still powers 78% of the web.

Quotable Engineering Wisdom

Microservices, once the darling of architecture diagrams, are getting a sober second look. One Hacker News commenter likened the shift to “getting drunk: a way to briefly push all your problems out of your mind and just focus on what's in front of you. But your problems didn't really go away, and in fact you just made them worse.” The hangover, it seems, is the operational complexity that follows the initial thrill of decoupling.

The debate over data models continues, with a sharp critique of the NoSQL marketing narrative. “All data is relational. Developers know this. What developers need to know is how to model those relationships efficiently in NoSQL,” said one observer, arguing that the term “non-relational” was invented to explain the difference to non-engineers. Sentiment against document stores is growing, with one engineer suggesting “the document data model is dying a slow death,” while “SQL is having a resurgence… tackling ease of scalability & replication & analytics.”

On the infrastructure front, real-world patterns reveal a mix of pragmatism and skepticism toward the cloud. A DynamoDB Global Table setup was praised as a solid pattern: deploy active application stacks in each region, route clients via Route53, and let DynamoDB replicate data cross-region. However, others question the cost and complexity of managed services. “Running a MySQL instance on a dedicated server with humongous amounts of RAM and speedy NVME drives for $100/month or so is not a bad deal,” argued one commenter, suggesting that 99% of services don't need migrations at 3am.

That skepticism extends to the container ecosystem. A recurring complaint is that new DevOps engineers learn Kubernetes as “the base truth,” leading to an industry “full of kubernetes experts who nail every service with k8s hammer and then drive insane amounts of cloud infra bills.” The simpler vision of the cloud is being lost, where offerings were “friendly to indie devs as much as they were for BigCorps.” Containers remain popular precisely because they hit a sweet spot: “They are familiar enough but push the envelope in interesting ways,” without requiring a whole new mental model, even when a serverless design would be more efficient.

The pressure of operating at scale was a recurring theme. One comment described the reality of “throwing bodies” at a problem: “A small bunch of people will be overworked, stressed, constantly fighting fires and struggling to fight technical debt… Production is always a hair away from falling over but luck and grit keeps it running. To the team it's a nightmare, to the business everything is fine.” Another, ex-Amazon employee Ben Adam, noted that at Amazon’s scale, “centralization is the enemy of efficiency… Being efficient on a macro level requires being (very) inefficient at the micro level,” pointing to the duplication of tooling across orgs—including 56 different internal design systems—and the startling reliance on Excel spreadsheets for critical processes.

Cost predictability emerged as a bigger concern than absolute price. “My primary concern with AWS service pricing is how hard it is to predict, not that they're too expensive,” said one expert, adding that a couple of services might even be underpriced. The narrative around dev environments also skews perception: “When you're building something, your spend is ALL dev environments. That narrative is sticky. Once you start scaling, you still think of dev environments as 'expensive' despite the fact that it's now sub-5% of your bill.”

The long tail of cloud adoption is clearly in “Establishment IT,” not the startups that led the charge for the last fifteen years. With global IT spend estimated north of $4 trillion, roughly 95% of IT isn’t on the cloud yet. The revenue potential is at least 10x today’s, and it’s not coming from cloud-native scenarios.

A counterpoint came from a user in a boring industry far from Silicon Valley, who finds the cloud indispensable for bursty compute. Having a scriptable “cluster” of six 128GB RAM machines that only runs 200 hours a year, and being able to scale up to 256GB or switch to 100 single-core instances on demand, is a godsend—and impossible to justify as physical hardware. “I find myself coming back to the cloud. Why? It costs more and you have less control… But, in my experience, not dealing with an IT department is the main reason,” said another sysadmin, capturing the enduring appeal of abstraction.

The quote about “stateless” architecture summarized a key tautology: “'Stateless' is just another way of saying you've left maintenance of state to someone more competent.” Perhaps the most resonant, practical advice came from Lee Atchison: “You can’t solve scaling and availability by code… It’s a combination of developers and management doing the right things, and putting processes and procedures in place.” Or, in the pithier words of the quotable @MissAmyTobey, proud of her boring architecture: “'that architecture is fine I guess, but it's so... boooring' 'yes. I am very proud of that aspect in particular.’”

Event-Driven Architecture at Scale: What Pokémon GO and Twitter Reveal

Niantic’s Pokémon GO is a case study in managed cloud services absorbing extreme load. During community events, transaction rates jump from 400K per second to nearly a million in minutes as regions come online. The infrastructure that absorbs this demand is built on Google Kubernetes Engine, Cloud Spanner, and Bigtable.

At steady state, the game runs on roughly 5,000 Spanner nodes and thousands of GKE nodes. Despite this massive distribution, all players share a single realm and game state. Traffic flows through Cloud Load Balancing into an NGINX reverse proxy, with static assets served from Cloud Storage and Cloud CDN. A third tier, the Spatial Query Backend, maintains a location-sharded cache that determines what appears on the map and which Pokémon spawn where. The frontend manages the player; the spatial backend manages the map.

Every user action is written as protobuf logs into Bigtable with strict retention policies, and simultaneously published to a Pub/Sub topic for downstream analytics. Server logic is fully deterministic, which means two players on different machines in the same location receive identical results. The operational model relies heavily on Google Cloud Monitoring for dashboards and alerts; the Niantic SRE team primarily manages quota ahead of event spikes.

Twitter’s real-time pipeline has moved partially to Google Cloud Platform. The company processes roughly 400 billion events per day, generating petabyte-scale data. The new architecture splits processing between on-prem and cloud: preprocessing and relay services on Twitter's own hardware convert Kafka events to Pub/Sub topics with at-least-once semantics, while Google Cloud Dataflow jobs handle deduplication and real-time aggregation before writing to Bigtable. The overall system processes millions of events per second with latency up to ~10 seconds.

The move has a clear payoff: no separate batch pipelines to build and maintain, higher aggregation accuracy with stable low latency, and no need to operate duplicate real-time aggregations across multiple data centers.

The Multi-Region Resilience Debate

The December 2021 AWS outages reignited the argument over whether multi-region architectures are the right resilience strategy. Forrest Brazeal offered a blunt counterpoint: stateful multi-region workloads are hard distributed systems problems, and treating the pattern as a silver bullet is misguided.

A concrete failure case raised on Hacker News showed why: during one outage, the STS service in us-east-1 failed, and other regions depend on it. Customers who had built around Amazon's published reliability model saw services in every region impacted by a failure in one availability zone.

The counterpoints are economic. One commenter managing a large SaaS noted their Kafka broker failed and had to be manually replaced during the outage — their multi-AZ setup meant customers noticed nothing. Another described the cost pressure: achieving margin targets while adding multi-AZ capacity is nearly impossible at smaller scales. And a DR drill skeptic warned that "middle-of-nowhere-west-1c" often lacks the core services an entire platform depends on — and that during real outages, everyone races to spin up capacity at once, with no room left.

Roblox's 73-Hour Outage: A Postmortem in Streaming and Freelists

Roblox published a detailed account of its October 2021 outage, which lasted 73 hours and affected over 50 million players. The platform runs on more than 18,000 servers and 170,000 containers in Roblox's own data centers, orchestrated with HashiCorp's Nomad and Consul.

The failure chain started with a caching system that routinely handles 1 billion requests per second across multiple layers. The breakthrough came when engineers disabled the streaming feature for all Consul systems — the 50th percentile for Consul KV writes dropped to 300ms. HashiCorp explained that streaming uses fewer concurrency control elements (Go channels) than long polling; under combined high read and high write load, the design creates contention on a single Go channel, blocking writes.

But one issue remained. BoltDB, Consul's underlying key-value store, tracks free pages in a "freelist." Roblox's workload exposed pathological performance degradation in freelist maintenance. Ben Johnson, BoltDB's author, acknowledged the design on Hacker News: "The project was never intended to go to production but rather it was a port of LMDB so I could understand the internals... my poor design stuck around."

Recovery required careful load management. With caches cold after restart, engineers used DNS steering to admit a controlled percentage of players while directing the rest to a maintenance page. Observability tooling also complicated diagnosis — commenters noted the outage was extended by circular dependencies in the observability stack itself.

Lessons From Building a Production Database

Mahesh Balakrishnan, who spent four years leading the Delos storage system team at Facebook (described as Facebook's version of Chubby), compiled 42 lessons from building production infrastructure without a single severe outage. Several are counterintuitive:

  • Design for re-orgs. A management hierarchy is inherently fragile — a tree is a 1-connected graph. Socialize projects continuously with future managers so churn doesn't produce unfair outcomes for individual contributors.
  • Avoid performance arms races. Competing on raw speed with other teams escalates into wasted optimization for point workloads and apples-to-oranges comparisons. Compete on fundamental design characteristics instead.
  • Prioritize consistency and durability over availability early. These properties are harder to measure and harder to fix once broken. Availability gets the operational pressure because it's measurable; push back.

Avoiding Fallback in Distributed Systems

AWS's builders library argues for a clear principle: favor code paths exercised in production continuously over rarely used fallbacks. The advice is to improve availability of primary systems directly — for example, pushing data to systems that need it rather than pulling, which risks a remote call at a critical moment. Subtle behaviors like excessive retries can flip a system into fallback-like modes.

When fallback is unavoidable, exercise it as often as production allows. A rarely tested fallback fails exactly when it is needed most.

The Economics of Cloud vs. Dedicated

Polyhaven's bill breakdown shows how far a creative caching strategy can stretch: 80TB of bandwidth and 5 million page views a month for under $400 total. The math is revealing:

  • Serving that traffic from S3 alone would cost around $4,000 per month.
  • Cloudflare acts as a massive caching layer, absorbing roughly 85% of asset downloads and 93% of other traffic for $40 per month.
  • Backblaze's B2 handles the long tail, and the Bandwidth Alliance means Cloudflare customers aren't charged for Backblaze download traffic.
  • The frontend runs on Next.js via Vercel ($20/month), with a separate $5 Vultr server running the API to control caching and avoid Vercel usage charges.
  • Google Firestore accounts for about half the budget at $100/month — a deliberate spend to avoid self-managing reliability.
  • Images live on Bunny.net's CDN at roughly $27/month, including dynamic resizing and compression.

Hacker News commenters offered alternative math. One runs 6 Hetzner dedicated servers handling 150TB and 26M page views for roughly $500, with a share-nothing architecture tolerating five server failures. Another pointed out that Netflix pushes close to 400Gbps on 1U commodity hardware.

Practical Tools and Papers From the Scalability Community

Timestamping and Data Integrity

freeTSA.org offers a free Time Stamp Authority service. Adding a trusted timestamp to code or an electronic signature creates a digital seal of data integrity, along with verifiable date and time information for when a given transaction took place.

New Open Source Projects

pinterest/memq is a new PubSub system designed to augment Kafka at Pinterest. It relies on a decoupled storage and serving architecture similar to Apache Pulsar and Facebook Logdevice. According to the accompanying article, MemQ is 90% more cost-effective than the company's existing Kafka footprint.

Netflix has released cachemover, a project that dumps memcached data (all key-value pairs) to disk and then populates a memcached process on a different server. As detailed in the accompanying write-up, Netflix reduced total warm-up times by roughly 90% compared to its previous architecture.

Spotify open-sourced XCRemoteCache, a remote caching tool for Xcode projects. It reuses target artifacts generated on a remote machine, served from a simple REST server, and cuts clean build times by 70%, as covered in the Spotify engineering blog.

Videos

Recent Papers and Publications

Version 2.0 of Architecting for Scale: How to Maintain High Availability and Manage Risk in the Cloud is now available.

Log-structured Protocols in Delos presents experiments and production data showing that log-structured protocols impose low overhead while enabling optimizations that improve latency by up to 100X (for example, via leasing) and throughput by up to 2X (via batching).

Research on PI-Terminal Planetary Defense suggests a surprisingly effective approach to asteroid deflection: launch an array of rods (such as a 10x10 grid of multi-meter long hardened penetrators, some carrying explosives) into the asteroid's path. The relative velocity would pulverize the asteroid into a cloud of fragments that could then burn up in the atmosphere, producing many uncorrelated smaller airbursts rather than a single catastrophic impact. Existing launchers could deliver 100 penetrators, or 10 tons, into an asteroid's path, while Starship with refueling could deliver 100 tons — enough to "take on asteroids well in excess of 100 m diameter" with a goal of mitigating an Apophis-class (370m diameter) asteroid.

The Hybrid Networking Lens AWS Well-Architected Framework whitepaper helps customers review and improve cloud-based architectures and understand the business impact of their design decisions. It describes general design principles as well as specific best practices and guidance for the five pillars of the Well-Architected Framework.