Measurement Is the High-ROI Work Engineers Skip

Engineers often treat "building" as the only worthwhile activity and "measuring" as a lesser pursuit. That view gets the ROI backwards. Because so few people bother to rigorously measure things, there is far more low-hanging fruit in measurement than in building. A well-executed measurement project can expose flaws that years of development and marketing missed.

The canonical example of measurement's power is Kyle Kingsbury's Jepsen work. Before Jepsen, only a handful of huge tech companies had seriously tested distributed systems, and their methods didn't spread. Almost everyone else shipped systems that were—by any reasonable standard—poorly tested. The typical online defense of a distributed database was "it works for me; it's never corrupted my data," with no methodology behind that claim.

Jepsen's early tests, despite being far less sophisticated than the current framework, found catastrophic problems in nearly every system examined. Redis lost 56% of writes during a partition. MongoDB dropped a "phenomenal" amount of data. Riak's last-write-wins semantics led to unbounded data loss. RabbitMQ lost roughly 35% of acknowledged writes, and the author noted this wasn't theoretical—he knew of two production deployments that had hit it. etcd's registers weren't linearizable; Consul allowed stale reads from any node that considered itself leader. Elasticsearch's health endpoint would report green during split-brain, and 645 of 1961 acknowledged writes were lost.

The bugs often weren't new. The Elasticsearch issue had been reported almost two years before Jepsen tested it, with no documentation, no public discussion, and no mention in training classes. When catastrophic bugs were reported prior to Jepsen, the typical response chain ran through predictable defenses: the report must be wrong; if a repro exists, the behavior is actually fine; if the bug is real, it's too rare to fix; if it needs fixing, it's a fluke and no methodology change is warranted. Jepsen blew through each defense. Without that kind of independent measurement, the industry would likely still be having those arguments instead of adopting test methodologies that actually produce reliable systems.

Measurement Changes Markets

Published measurements force improvements in areas where marketing buzzwords had previously substituted for engineering. A few documented examples:

  • Keyboards: After latency measurements were published, at least one major manufacturer advertising high-speed gaming devices started actively optimizing input latency. Before that, almost no one measured keyboard latency—only one other measurement could be found, and it confirmed the implausibly high results. Today every major gaming keyboard and mouse maker ships low-latency devices; before, they focused on things like higher USB polling rates that had little real impact.
  • Computers: Published latency measurements led an engineer at a major software company to start measuring and optimizing UI latency, and the author of alacritty filed a ticket on reducing its latency.
  • Vehicle headlights: When Consumer Reports began testing headlights, auto engineers thanked the publication for giving them ammunition to win arguments against designers who preferred better-looking but less effective lights. There was no business case for safer headlights until a third-party metric existed; now designers sometimes lose that argument.
  • Vehicle braking: When Consumer Reports and Car and Driver found the Tesla Model 3 had extremely long braking distances—152 feet from 60 mph, 196 feet from 70 mph—Tesla updated its brake modulation algorithms and went from worst in class to better than average.
  • Impact safety: Apart from Volvo, car makers generally engineer to the highest possible score on published crash tests. Safety improvements appear when new tests are published, not before.

None of these were exotic projects. Anybody with a budget for test equipment—or even a car rental—could have done them. Measurement projects are easy to find precisely because so few people do them. If you want to change the world, there's no shortage of high-impact things to measure; only if your goal is making money is building probably the easier path.

Curiosity Is the Usual Motivation

The real impetus for most measurement work is simple curiosity—wanting to know the answer to a question. The proof is in the projects themselves. Chris Fenton's and Oona Räisänen's projects, which frequently hit the top of Hacker News, were clearly not motivated by the fame; both were doing interesting work long before their blogs were popular. The suggestion that someone builds a high-quality measurement project primarily to get to the top of HN is inconsistent with the most obvious explanations: they're having fun, they're curious, and they want to know how something actually works.

Most measurement posts stem from a desire to resolve a personal question, not to gain attention:

  • Car safety: investigating whether there are significant safety differences between manufacturers given that most cars get top marks on US tests
  • Input lag: whether modern computers genuinely feel higher latency than older ones
  • Keyboard latency: determining the latency contribution of the keyboard, since display latency was already well-tested
  • Terminal latency: whether iTerm2's slowness was real, given that existing benchmarks mostly measured throughput
  • Keyboard vs. mouse: testing non-bogus productivity claims in place of the widely-cited but obviously flawed numbers
  • Web bloat: quantifying how unusable the web was on a road trip without fast internet

Reviewing such a list reveals that the data-driven habit predates both a job and a blog. It's a mode of thinking, not a career stage.

Why Most Published Reviews Can't Be Trusted

Bad measurements both increase the value of good ones and blunt their impact. People generally can't distinguish between valid and invalid measurements, so they choose reviews based on things other than quality. A compounding problem is that many reviews are compromised by manufacturer relationships from the start.

Car reviews are an extreme example. Consumer Reports is the only major reviewer that independently purchases its test cars, which is why its findings often diverge from others. Independent sourcing lets them buy the trim level actual consumers purchase, rather than the manufacturer-provided one, and avoids receiving cars that have been hand-picked to avoid cosmetic issues or tuned with software that would void a consumer's warranty. It also lets them write genuinely negative reviews: explicitly criticizing a car costs a reviewer access to future manufacturer-provided vehicles, which is why most car publications require reading between the lines.

Camera lenses have a documented history of reviewers receiving unusually good copies. Copy-to-copy variation is tremendous, and vendors select good copies for reviewers to borrow. Based on how many copies people return before getting a good one, it appears the median lens may have noticeable manufacturing defects, with perhaps one in ten lenses being defect-free. LensRentals' data finally quantified this and found that manufacturers have very different levels of copy-to-copy variation.

SSD reviews have similar documented problems. ExtremeTech reported multiple times that Adata, Crucial, and Western Digital provided review samples that didn't match retail units. The publication's framing—that a reviewer trusts the manufacturer to provide a representative sample—absolves the reviewer of the responsibility to obtain and verify representative hardware. Trusting vendors is not a strategy. Vendors will lie and cheat to look better in benchmarks, and blaming them afterward doesn't make reviews accurate or useful.

The review problem is compounded by the fact that most highly-ranked online reviews are SEO affiliate farms, and even respected review sites suffer from a deeper issue: people can't tell good reviews from bad ones. Wirecutter, despite being the default recommendation engine for a large swath of tech workers, uses methodology that is, on inspection, quite poor. In one test of video call equipment, its recommended webcam performed no better than the camera in a 2014 iMac and produced white balance and autofocus issues that users commonly encounter. Its recommended microphone was roughly comparable to a laptop's built-in mic. This is typical rather than exceptional for the site.

The Market Doesn't Automatically Fix Bad Products

The belief that free markets naturally reward good products and eliminate bad ones doesn't survive contact with evidence. People who espouse that view often fill their homes with products whose flaws are only visible through third-party testing. Measurement changes this dynamic, and the effects show up in unexpected places:

  • Electronic stability control: Multiple manufacturers (Toyota RAV4 and Hilux, Nissan Rogue, Jeep Grand Cherokee) made major system improvements after independent reviews exposed problems.
  • Tires: Most manufacturers other than Michelin see severely degraded wet, snow, and ice performance as tires wear. One technical reason: sipes that improve grip aren't cut to full depth because it significantly increases manufacturing cost. A non-technical reason: most published tire tests use new tires, so partial-depth siping costs nothing in marketing benchmarks. As TireRack's all-around winter handling tests grew prominent, manufacturers adopted more multi-directional siping—Consumer Reports' snow and ice scores only test straight-line acceleration and braking. Measurement's impact is bounded, though: Michelin, designing the successor to the excellent X-ICE Xi3, made visual appearance a primary design criterion because customers chose worse tires that looked like winter tires. They also renamed the product to X-ICE SNOW to make its capabilities clear.
  • Camlink and cheap HDMI-to-USB converters: Documented USB chipset issues mean some converters simply don't work on certain Windows machines. Cheap alternatives generally work with at least one computer and one piece of software—but fail to work or produce distorted video in other cases. There is no benchmark for converter quality.
  • HDMI-to-VGA converters: Many run hot and stop working after 15 minutes to two hours; some don't warm up at all. There's no way to tell them apart without testing.
  • Water filtration: Two Amazon reviewers measured lead levels in contaminated water before and after Brita "longlast" filtration and found no reduction. The filters previously had a slow flow problem; now some filter faster but don't meet Brita's claimed filtration levels.
  • Storage containers: The Rubbermaid brand famous for durable Roughneck and Toughneck containers was bought by a firm that cut materials and strength; the similar-looking containers now buckle under stacking. No one benchmarks them for load.

Measurement is so underrated that almost everything you interact with is sub-optimal in ways that are relatively straightforward to fix once data exists. The bottleneck isn't engineering talent or capital—it's the scarcity of people willing to do the unglamorous work of finding out how things actually behave.