Notable Numbers from the Week in Scale

A birthday service pushing existential reflection? Apparently so. The latest round of data points from around the distributed world spans everything from cloud revenue milestones to the rising cost of attribute names in DynamoDB. Here's what stood out.

The Cloud and the Datacenter

The shift away from traditional infrastructure continues to accelerate. Gartner projects that by 2025, 80% of enterprises will have shuttered their conventional datacenters, up from just 10% today. The momentum behind that transition is visible in the earnings of the major cloud providers. AWS posted quarterly revenue of $10 billion, representing 33% year-over-year growth and making the division larger than Oracle as a whole. Microsoft's Azure, meanwhile, grew 59% in the same period.

Even as enterprises move workloads off-premises, the economics at the edge of the network are reshaping other businesses. Dropbox reported its first quarterly profit ever, a notable inflection point for the file-sync pioneer. The New York Times, by contrast, is bracing for ad revenue to fall at least 50% in the second quarter, even as it added 587,000 net new digital subscriptions relative to Q4 2019.

Consumer Platforms in the Spotlight

TikTok's growth curve remains steep. The app has now surpassed 2 billion total downloads, having added 500 million in just five months. Its most recent quarter generated more than 315 million installs across the App Store and Google Play, the most any app has ever accumulated in a single quarter. The scale of engagement on the platform is similarly outsized: Drake's "Toosie Slide" has been played 4.8 billion times on TikTok, compared with 1.6 million views on YouTube as of May 9.

The advertising economy is showing strain, however. Roughly 50% of online ad spending now flows to middlemen; out of 267 million ads placed online, researchers could match the end-to-end process for only 31 million. That squeeze is hitting creators directly. A YouTube channel with 2 million subscribers, TierZoo, saw monthly revenue fall from $11,000 to $4,800, with two-thirds of income now coming from sponsorship deals negotiated separately from platform advertising.

Apple's services business continues to be a profit engine, with gross margins around 65%. The company reports 515 million paid subscriptions and over 1 billion active iPhones. Meanwhile, the mobile games market shows a curious contraction: Apple's App Store saw only 22,000 new game releases, down sharply from 285,000 in 2016.

Security and Infrastructure Realities

On the security front, the numbers illustrate both the scale of the problem and the sophistication of the attackers. Microsoft, with 47,000 programmers, generates roughly 30,000 bugs per month, yet its AI-based detection system can identify 97% of critical and non-critical security flaws. Zero-day exploits were more prevalent in 2019 than in any of the prior three years, and private companies appear to be supplying a larger share of those vulnerabilities than in the past. The FBI pegs total US cybercrime losses at $2.7 billion for 2018, with investment scams the most common vector. Industry estimates suggest global cybercrime damages could reach $27 billion by 2025.

Google released a 29-day Borg cluster trace covering every job submission, scheduling decision, and resource usage event from May 2011, offering a rare window into the operation of its internal compute orchestrator. Elsewhere in the data layer, one operator found that attribute names consume about 70% of DynamoDB IOPS and storage; switching to shorter attribute names would eliminate 14K WCU at peak and reduce table size by 3.5 TB.

The Macro Picture

The economic fallout from COVID-19 appears in nearly every dataset. The US unemployment rate hit 14.7% in April, with 20.5 million jobs lost. A PwC survey of US CFOs found 32% anticipate layoffs in the next six months, while 49% plan to make remote work a permanent option. The robotics angle adds another layer: each additional robot displaces roughly 3.3 workers and lowers wages by about 0.4%.

Not everything is grim. Venture funding for AI startups reached $6.9 billion in Q1 2020, a record-setting pace before the pandemic took hold. Shopify has seen traffic roughly double. And PyPI, the Python package index, now delivers 200 million packages daily.

Even the astronomical data points feel relevant. The closest black hole to Earth sits about 1,000 light-years away.

The Week in Distributed Systems

Scale, Cost, and Cloud Realities

Efficiency gains continue to mask the explosive growth in cloud demand. Google's Urs Hölzl highlighted data showing compute capacity up 550% globally while energy use rose only 6%, crediting efficiency improvements for keeping data center energy consumption nearly flat. The economic picture is more complicated: ad-dependent businesses are seeing traffic rise while revenue falls, and publishers who once commanded premium ad rates are watching their pricing power evaporate.

The Serverless Paradox

A recurring theme this week: serverless platforms promise to abstract away infrastructure but often deliver the opposite. As one engineer put it, "The goal was to abstract away the infrastructure. We instead converted every application developer into an infrastructure developer." Another observation: "Serverless will be the biggest LOL f*uck you on developers since full stack. They think they're buying into only having to focus on code and not infrastructure, but instead we are secretly turning them into distributed systems engineers more than developers."

Cost realities vary widely depending on architecture. One insurance startup running a serverless stack on AWS reported an April bill of just under $740, with DynamoDB ($202), CodeBuild ($116), CloudWatch ($100), and S3 ($66) leading the way. Their AppSync, Cognito, and Lambda costs combined came to only $32. Another practitioner noted there's an inflection point where EC2 becomes more cost effective, but argues it's much higher than people realize: "Lambda I can deploy some code and forget about it for a year. For EC2 I have to plan a 30 day machine re-image/rebuild maintenance schedule."

The enthusiasm isn't universal. Developers working from home report that "cloud first development" is painful when upload bandwidth is limited. "This model seems fundamentally broken for the new normal of working from home as it assumes you're in an office with a fast upload connection so you can push images. Give me local dev any day."

NoSQL Data Modeling and Operational Patterns

DynamoDB continues to attract attention for both performance and cost reasons. A Kindle Collection Rights Service migrated from RDBMS to DynamoDB, noting that running nine distinct queries millions of times a day on a relational database was "a massive waste of CPU" and the switch saved significant money.

Practical modeling advice circulated: "Data modeling in #NoSQL is about replacing JOINs with indexes. Store all objects in one table, decorate them with common attributes, and index those attributes to produce groupings of objects needed by the app." For operational analytics, the suggested pattern is maintaining time-bound rollups in DynamoDB with Streams/Lambda, then querying relevant items by date range and aggregating client side for fast reporting. Another user pointed out the full text of Moby-Dick fits in a single DynamoDB row with xz compression.

DAX (DynamoDB Accelerator) received praise: "Regular caching is much different from DAX. Improving response from milliseconds to microseconds. Now that is caching on steroids."

Microservices at Scale: The Reckoning

Uber's experience with thousands of microservices drew cautionary notes. An engineer who joined Uber in 2016 recalled conference talks about scaling to thousands of services — talks that stopped after a year or two. "Turns out, having thousands of microservices is something to flex about, and make good conference talks. But the cons start to weigh after a while."

The downsides cited include hard integration testing, outages traced to parallel deployments of two services, ownership problems when a person who owned a critical-path microservice leaves, and on-call loads where engineers manage four to five services they launched. The counter-movement: tiering to ensure critical services have proper monitoring and alerting, merging related services, and consolidating toward fewer, bigger services.

One commentator argued that creating new microservices for a new project is "the epitome of and posterboy for premature optimization," recommending instead a modulith approach: "Make your code loosely coupled enough so any functionality can be easily extracted out into a microservice if need be."

Scale Facts and Figures

  • Google Meet and Microsoft Teams together handle close to 5 billion minutes of calling per day. The entire UK telecoms system handles about 600 million.
  • A single SQS queue accumulated 253,477,653,099 unread messages; the owner was reading at only 1.56 million messages per second in a backfill operation. Everything worked fine.
  • SQLite alternative: One application receives more than 10 million hits a day through Kong, which uses Redis for its rate limit plugin, running on a t3.micro without issues.
  • Cloudflare testing found HTTP/3 loads a 15KB test page in 443ms versus 458ms for HTTP/2, but the advantage disappears at 1MB pages, where HTTP/3 is slightly slower (2.33s versus 2.30s).

Architecture Humility

"Any critique of a company's architectural choices needs to also include their business metrics too," one commenter observed. A counterpoint to technical purity: if the revenue, growth, and stock price look fine, the architecture may be working well enough.

A related sentiment on pricing: "At Acme, we're really good at spending \$2 million+ on something. Less than \$500k not so much. Less than \$100k is credit card money almost." The implication for enterprise SaaS startups: if prospects aren't paying attention, maybe you're not charging enough for them to care.

Scalability advice was blunt: "2000 customers @ \$39/month is almost \$1M/year. If your software can handle 2000 customers, you can worry about scalability once the \$1M/year is flowing in. Scalability doesn't get you customers. First have customers, then worry about scalability. The order matters."

Spot Instances and Cost Optimization

With the economy moving online, spot instance termination rates have risen significantly. One team switched their autoscaler's SpotAllocationStrategy from default to "capacity-optimized" and added more instance types, which seemed to reduce termination rates. The setting didn't exist when their autoscalers were first configured.

An SRE review by Mike Julian and Corey Quinn at Duckbill Group identified EBS volumes that could switch from GP2 to ST1, roughly halving storage costs. Another cost observation: "Although there are significant technological risks to data stored for the long term, its most important vulnerability is to interruptions in the money supply."

Cloud Provider Lock-in and Incentives

Cloud credits and sweetheart deals are reshaping vendor selection. A former AWS practice lead recounted prospects saying, "'Google is offering us a \$2 million credit to choose their cloud. What do you say to that?'" His answer: "You should take them up on that offer." Cloud providers will "literally buy your business for years on the bet that at some point they will make it back."

A startup acquired by Cisco chose Oracle's bare metal cloud for similar reasons: "AWS and GCP were great, but also fairly expensive until you get locked in." Because they relied on open source technology rather than proprietary cloud services, transitioning was trivial. Another cautionary tale claims AWS has used data from third-party hosted services to build competing products and poach customers.

Corey Quinn critiqued Azure's region-pairing strategy: "You can think of each Azure region as an AWS or GCP Availability Zone. That's great right up until it isn't, and you wind up with a bunch of small regions rather than fewer more robust ones, and a sudden influx of demand causes those regions to run out of headroom."

Rust is gaining ground as a default choice for new data plane software, with one engineer observing "the tide has turned" in recent design reviews. Go, meanwhile, has solidified its position: "Docker, Kubernetes, and many other components of cloud computing written in Go... Go has indeed become the language of cloud infrastructure." At Google, the SRE team mandated "no new Python projects" and shifted automation to Go.

Chiplets were framed as a reversal of the historical integration trend — breaking monolithic chips into functional blocks. The most interesting possibility: not extending Moore's Law by packing more transistors, but enabling experimentation with new materials and processes across specific blocks rather than an entire SoC at once.

Optical computing drew skepticism: "Wavelength of the light emitted by these devices: ~4000nm. Latest generation commodity CPU transistor structure size: 7nm." Photons don't trap well, requiring delay lines and optical amplifiers to hold bits, making electrons far more practical for storage.

On Protocol and APIs

A debate emerged over whether SQL or GraphQL makes a better API language. The pro-SQL argument: you don't have to expose your full schema, only carefully designed views, so you can refactor tables without breaking the API. Time limits should cut off expensive queries in either approach. GraphQL advocates were challenged: "SQL is a better API language than GraphQL. Convince me otherwise!"

Remote Work and the Office

Working from home is proving sticky. A CEO of a 400+ person company said work-from-home is going well enough that he may not renew his San Francisco lease, saving \$10 million per year on office and lunch costs, and instead run a few all-hands offsites annually. Commentators noted many executives are likely thinking similarly, potentially permanently altering commercial real estate.

Engineering Wisdom

Several pithy observations circulated:

  • "Good design means when I make a change, it's as if the entire program was crafted in anticipation of it. I can solve a task with just a few choice function calls that slot in perfectly, leaving not the slightest ripple on the placid surface of the code."
  • "Tools always break. She should have remembered that."
  • "Every computation must produce heat, but both energy consumption and heat production can be outsourced — by, say, human leadership, distributed computing, or a simple virus."
  • On complexity: an engineer spent two days charting every path and pattern of timer events, then wrote exhaustive code to handle each case. "It never had to be addressed again. But it did have to be addressed. So many folks are unwilling to face the music with complexity."

Notable Outages and Observations

John Conway reflected on his late twenties "black period" before finding international prominence: "I started to wonder whether it was all nonsense... I made a vow to myself... I was going to study whatever I thought was interesting and not worry whether this was serious enough. And most of the time I've kept to that."

The Naikon group continues its long-running cyber operation, having "updated its new cyberweapon time and time again, built an extensive offensive infrastructure and worked to penetrate many governments across Asia and the Pacific."

HaveIBeenPwned saw significant uptick since COVID-19 restrictions: daily unique users up from about 150k to 200k, and monthly users up 41% on the previous 12-month high.

A 10-year-old was discovered playing with her dog during a Zoom class. She took a screenshot of herself "paying attention," cut her video, and replaced it with the picture: "It's a gallery view of 20 kids, mom. They can't tell."

William Gibson was wrong about the internet, according to Lev Grossman: "He thought it was a place that we would all leave the world and go to. Whereas in fact, it came here."

Scaling Through a Pandemic: Zoom’s Architecture Under the Microscope

Zoom’s leap from 20 million to 300 million users almost overnight has drawn intense scrutiny. The absence of visible growing pains from the outside, however, masks the internal chaos. Critics have called some of Zoom’s design choices “bad architecture,” but these decisions—made when the company was a small startup—are better understood as the natural evolution of a product forced to scale in weeks rather than years.

What is known about Zoom’s setup? The traffic mix tells an interesting story. Real-time video conferencing for paid customers historically stayed in Zoom’s own data centers. As the pandemic drove daily records, those centers couldn’t cope. AWS began spinning up thousands of servers daily, with Oracle Cloud providing additional capacity. The company now runs on a three-part foundation: its own data centers, AWS as the primary public cloud, and Oracle to a lesser extent.

Architecturally, Zoom’s competitive edge comes from treating video differently than its rivals. Many competitors route traffic through a data center, transcode it into a standard view, and mix video before sending it to participants. That creates latency, consumes CPU, and complicates scaling. Zoom instead uses the SVC (Scalable Video Codec) codec. Unlike AVC, which requires multiple streams to support multiple bitrates, SVC packs all resolutions and bitrates into a single layered 1.2 Mbps stream. Once requiring an ASIC, SVC now runs in software thanks to Moore’s law.

Around this, Zoom developed “Multimedia Routing.” Content enters the cloud, but nothing is transcoded or mixed. Clients subscribe to different layers of a participant’s resolution and pull streams directly from routing with zero processing. An application-layer QoS system between cloud and client handles network quality: telemetry on CPU, jitter, and packet loss determines which stream gets switched to a client, and clients can downgrade their own outbound video in poor conditions. The best experience is attempted first over UDP, falling back to HTTPS, then HTTP. The rationale was to avoid inconsistent user experiences at all costs, which explains some seemingly odd early decisions.

The commercial model also drove adoption: free 40-minute meetings, free dial-in, and giving away what competitors sold. VOIP adoption is at 89% for Zoom, versus an industry average under 30%. The strategy aims to tear down paywalls between collaboration tools and build what the company calls “the largest network of connected collaboration.”

One persistent question—whether Zoom runs its own data centers—draws an industry clarification. Running a “datacenter” for most large players content means renting colocation space and buying servers. Equinix dominates the high-end with premium rents (2–3x cost) and ~$300 per cross connect per month, making it useful for network POPs but uneconomical for compute. CoreSite and Digital Realty serve the mid-market; AWS famously does leasebacks (owning equipment in a shell owned by another). Oracle Cloud historically leased massive retail and wholesale space from CoreSite and Digital Realty to move quickly. Zoom likely follows a similar mix of providers, optimizing for each market it needs a PoP in.

The Internet’s Report Card: Resilient, Sometimes Warm

While individual services wobble, the internet itself has held up remarkably well. Global traffic data from ISPs and IXPs tells a consistent story: growth in breadth, not spikes in peaks.

BT’s network, for example, saw higher average traffic but peaks never exceeded a pre-crisis football match’s 17 Tbps. Networks that used to be quiet during the day are now busy around the clock. Netflix watching is less peaky—traffic rises and stays up. Enterprise traffic dropped dramatically (everyone is at home), but VPN traffic is way up.

The EU, seeing problems in Italy and Spain early, asked Netflix to reduce video traffic. Netflix complied by serving 4K at a lower bitrate. ISPs built capacity ahead; Netflix pulled forward planned capacity augments, aided by having data center space already secured. The only real frictions: datacenter remote-hands services slowed (staffing shortages) and supply-chain uncertainty loomed for future hardware. A decade ago, BGP-related outages might have plagued this moment; cleaner routing tables and faster operator cooperation meant networks adjusted quickly with a "light touch" from regulators.

Consider the statistics from the OECD:

  • 60%: increase in internet traffic seen by some operators.
  • DE-CIX Frankfurt now peaks over 9.1 Tbps, with video conferencing up 120% and gaming up 30%.
  • BT: 35–60% increase in daytime fixed broadband usage; Telecom Italia: 63% fixed, 36% mobile; Orange: 80% of traffic from France goes to the US.
  • AT&T: mobile voice up 33%, Wi-Fi calling up 75%, fixed-line voice up 64%, core network traffic up 23%.
  • Cisco Webex: peaking at 24 times higher volume; Facebook: voice calls up 100%, messaging up 50%, and group calls in Italy up tenfold.

Ironically, the legacy PSTN now shows busy signals—it was never upgraded. Mobile traffic is down slightly, as people aren't in cities and are using Wi-Fi calling, FaceTime, and WhatsApp. BT worried most about the mobile network, but it performed. The consensus: capacity isn't binary. The internet was built over decades to absorb contention, and average consumers won't notice slowdowns as failures. Yet the warning is equally clear: don’t take this resilience for granted.

Cost, Architecture, and Operational Lessons

Several engineering posts from the week crystallize dominant themes in scaling and systems design.

Serverless economics shift the calculus. Netflix, long known for building sophisticated infrastructure on AWS, found that Lambda beat EC2 for image processing. After using Lambda in production since early 2020, they saw faster cold-start handling and mixed warm results (300ms slower), but the economics were decisive: under $100 per day for Lambda versus $1,000 per day for EC2. Lambda also handled 15x–20x load spikes gracefully.

Error handling in serverless requires explicit strategy. Jeremy Daly makes a case for “failing up the stack.” Don’t swallow exceptions inside a Lambda function; fail it and let the cloud’s built-in retry mechanics and dead-letter queues (DLQs) capture errors. Lambda-native architecture—many small functions handling single business units—beats monolithic “Lambdaliths,” enabling synchronous/asynchronous invocation, easy reuse in step functions, and least-privilege IAM roles. Cloud services guarantee at-least-once delivery, so idempotency is mandatory. Lambda Destinations surpass DLQs when you need failure context, not just the payload.

Edge rendering closes the latency gap. Lambda@Edge lets you run SSR logic from CloudFront edges, splitting page requests (SSR) from asset requests (S3). A case study reported 88ms responses in Paris, 44ms in Frankfurt, with average 168ms, p95 of 449ms, and p99 of 671ms.

At extreme scale, theory bends. Building a 10,000-node Akka Cluster came down to fixing bugs one at a time. The Rapid membership protocol, designed to beat consensus protocols, forms a 2000-node cluster 2–5.8x faster than Memberlist or Zookeeper. Gossip alone risks tail latency; Rapid's multi-node cut detector adds stability. Practical pain points: AWS autoscaling costs rival EC2 itself, and CloudWatch is nearly as expensive—direct `RunInstances` calls were cheaper and faster. A key design win is that Rapid’s gossip broadcast organizes nodes in an expander graph topology, where convergence is provable by relay count.

The queue is where latency lives. Adrian Cockcroft distills performance basics: measure response time at the user and service time at the service. Queues are why response times spike. For spikes, use one-second averages; for rare requests, average at least 20 together. Little's Law (Average Queue = Average Throughput * Average Residence) governs it all. Load tests with constant rates simulate a conveyor belt, not reality—they miss the bursty request patterns that create real queues. Keep network utilization below 50%, and expect multiprocessor systems to "hit the wall" at high utilization. Timeouts should not be uniform: long at the edge, short deep inside. Set load-shedding rules before a sustained 100% utilization that causes queues and unresponsiveness.

Kubernetes is robust, but only to a point. A production cluster scaled to 4x nodes and workloads without issue. User experience degraded only when performance tuning was ignored. The control plane handled everything thrown at it.

Trace costs to specific resources. Riskified cut ~60% of its DynamoDB bill by adding a simple in-memory cache. Index table write throughput was the culprit behind skyrocketing costs. With only 250MB, cache hit-rate reached 75–80%, DynamoDB write throughput dropped over 80%, and the bill fell for the first time in two years.

Layer removal pays twice. LinkedIn merged two layers in its identity services handling over half a million QPS. The result: 10% lower latency, and improvements across all percentiles—p50, p90, and p99 improved 14%, 6.9%, and 9.6%. Memory allocation rates dropped 28.6% per host. Decommissioning the data service cluster freed over 12,000 cores and 13,000 GB of memory.

Configuration distribution at internet scale. Cloudflare serves 14 million HTTP requests per second across 200 cities in 90 countries. Their custom Quicksilver system replaced Kyoto Tycoon with LMDB for key-value distribution. LMDB’s strengths: snapshotting with minimal read degradation (99th percentile read latency dropped by two orders of magnitude for DNS), support for multi-process concurrent access (enabling zero-downtime upgrades), and an append-only design that is crash-proof. Distribution uses a fan-out topology where edges query masters, which query top-masters. A monotonically increasing sequence number detects lost updates—an old trick, but simply a log at heart. In three years of running over 90,000 database instances serving 2.5 trillion reads and 30 million writes daily, they experienced one bug and zero corruption.

Simplification often just moves complexity. As Fred Hebert argues, complexity can't be removed—only shifted. Move it out of your code, and it appears in people's heads, workflows, and processes. The better design approach acknowledges complexity, gives it a named place, and builds systems around it.

Several stories trace the same motivational arc: scale requires a deliberate architectural stance, not just more servers.

  • Databases at Quora are sharded in MySQL. It is a massive operational effort, but the design works at scale.
  • Multi-cloud is nuanced. Flexera reports 93% of enterprises have a multi-cloud strategy; 87% use hybrid. The interesting change, per PlanetScale, is Kubernetes making true provider-agnostic workloads possible. Vitess's sharding keeps cross-cloud data transfer minimal, making migrations feasible without a data fire hose.
  • A microservices discipline at Monzo runs 1,600 services in Docker containers. A shared core library (stripped of unused code at build) avoids reimplementing abstractions like data marshalling, and built-in metrics surface every deployment on dashboards immediately.
  • Event-driven glue using EventBridge and Zendesk builds automated workflows that coordinate with state machines—a powerful pattern, but complexity rears its head in such multi-step choreography.
  • Configuration and consistency at Slack: ~12 scheduled deploys daily, only during North America business hours, with a “deploy commander” rolling builds out gradually, allowing hotfixes and rollbacks if error rates spike.

The producer edge case: Dropbox, YouTube creators, and gaming companies all face the same mathematical reality: advertising-based income is volatile in a crisis (one creator went from $11,000 to $4,800 a month). Sponsorships are steadier. Netflix's approach to production is fundamentally different: global distribution means that when one region locks down, another can resume production. Movie studios without that footprint can’t simply continue.

The business of gaming is shifting. E-sports differ from traditional sports because one entity owns the game and controls its economics like a marketing program. Streaming services won't swallow games the same way they did music or video—latency and compression limits are real. The future is hybrid: on-device, edge, and central cloud all engaged depending on content. Expect the “Netflix of games” not to emerge from a single streaming model but from a platform brand like Google, Microsoft, or Netflix working through a blend of delivery mechanisms.

Resilience Beyond Code

Complexity’s natural home: Stripping complexity from code just relocates it. Legacy systems and social structures internalize it; when the original engineers leave, understanding disappears. Systems built with that knowledge embedded become more resilient.

Resilience via wide distribution: Airbnb’s off-grid craftsman and the 51-year-old temporary building serve as reminders that longevity is not accidental. “Simplicity is difficult,” said the architect. “It is easy to make things complex.”

Resilient event formats: The NFL's remote draft worked because AWS hosted always-on streaming across 150+ smartphone feeds. American Idol kits with iPhones yielded excellent pictures. The Voice's stack—Microsoft Teams and Surface Pros—looked unprofessional by comparison. The lesson: pick technology based on result quality, not vendor sponsorship.

A network analogy from immunology: Rather than trying to predict a novel pathogen’s signature, our immune system generates random antibodies, then deletes any that attack self. What remains is a negative complement—a response to anything foreign. For computer security, this suggests generating random possible interventions, then weeding out those that attack your own code. Anonymizing the self removes the target.

A final reminder from failure: “Nobody cares about backups; what they care about are restores.” Backups must be complete. Test environment changes rigorously, no matter how small. For example, a routine update in a power grid AGC system failed exactly because control systems weren’t thoroughly validated first. And when Scaleway failed, the path forward required spare reserved instances, verified backups, and a clear restore runbook. On the flip side, video poker glitch hunters found that reproducing bugs often depends on quirks, and the best defense is a simple one: don’t be greedy.

This week’s roundup spans storage, networking, and post-quantum theory. Here’s what stood out.

  • facebookincubator/ntp — A collection of Facebook's NTP libraries.
  • Liftbridge — Adds Kafka-style durable, fault-tolerant streams to NATS. It offers a highly available, horizontally scalable publish-subscribe log, with a stated goal of simplicity.
  • Database schema templates — The DrawSQL template gallery collects real-world schemas from open-source projects to use as architectural inspiration.
  • HSE — An embeddable key-value store from Micron, tailored for NAND flash and persistent memory. It maps data across DRAM and multiple SSD tiers to boost performance and endurance. In referenced YCSB benchmarks, HSE delivered up to nearly 8x the throughput of MongoDB/WiredTiger and nearly 6x that of RocksDB.
  • Serverless Redis — A managed offering from Lambda Store with ~2ms latency and costs ~6x lower than Elasticache or RedisLabs. The project started with its most-used commands, with missing features slated for gradual rollout.
  • Netflix on TLS 1.3 — As detailed on their tech blog, modernizing encrypted transport improved play delay by 3.5–8.2% and cut media rebuffers by 7.4%, while also lowering CPU load.
  • TerminusDB — A GPLv3 in-memory graph database with a query language, WOQL.
  • AsyncAPI — A specification family for defining asynchronous APIs via machine-readable schemas.

From the Research Shelf

Systems papers in this batch tackle stream processing management, metadata curation, and hardware-level primitives.

Turbine: Facebook’s Stream Management Layer

Turbine closes the gap between generic cluster schedulers and Facebook’s internal stream processing. Its three core features are a low-latency task scheduler, a predictive auto scaler, and a transactional application update path (with fault-tolerance and consistency guarantees). In production for over three years, it manages thousands of pipelines across tens of thousands of machines, processing terabytes per second. Facebook reports that Turbine’s scheduler smooths load fluctuation, its auto scaler absorbs unexpected spikes, and its updater completes large changes within minutes.

Goods: Organizing Google’s Datasets

Goods is Google’s post-hoc dataset discovery system. Rather than forcing an upfront schema, it scrapes metadata from datasets after pipelines create or access them. This passive approach requires no change to how engineers work; Goods harvests metadata in the background and then powers search and organization over otherwise siloed data.

The Cloud’s Blind Spots

On gray failures, the authors describe a class of partial faults that evade standard failure detectors. Their key insight is differential observability—one component’s view of “healthy” differs from another’s suffering from the fault. The suggested remedy, then, is aligning failure perceptions across systems. For a practitioner-readable summary, this post walks through the paper.

Scalog and the Shared Log

Scalog targets reconfiguration and total order in a scalable shared log. It supports application-custom data placement, offers availability during reconfiguration, and accelerates failure recovery.

Hardware: Copy and Logic Inside DRAM

ComputeDRAM achieves in-memory computation on stock, unmodified DRAM chips. By intentionally violating the memory’s timing specification and activating rows in fast succession, multiple rows can be left open, enabling charge sharing across bit lines. This approach yields row copy, plus logical OR and AND primitives, forming a basis for massive parallel compute without custom silicon.

WormSpace and an Abstraction for Consensus

In WormSpace, the proposed star is a Write-Once Register (WOR). With a data-centric API backing single-shot consensus, the WOR abstraction can underpin durability, concurrency control, and failure atomicity in verifiable distributed systems.

A Theoretical Interlude

For a broader view, a paper by Freeman Dyson explains why Maxwell’s equations are hard to grasp: it’s the first field theory, and the same two-layer structure of simple dynamics under abstract mathematical descriptions recurs in relativity, quantum mechanics, and the Standard Model. A better grip on that pattern, Dyson argues, might clarify the interpretive fog around quantum mechanics.

Also in the Queue

  • CSE138 — UC Santa Cruz’s undergraduate distributed systems lectures, now available as video.