Crash Tests as Benchmarks: When Automakers Game the Safety Rating System

Any benchmark that gets taken seriously eventually attracts people trying to game it. In computing, this is well documented. Compiler tweaks have re-written benchmark kernels to inflate scores dramatically; GPU vendors have added benchmark-detecting code to drivers that lowers image quality during testing to produce higher scores. Gaming a benchmark is cheaper and offers a higher return on investment than actually improving the product.

Vehicle crash tests are physical benchmarks that are hard for outsiders to reproduce. They're highly specific, well-known protocols that destroy a car in the process. Because manufacturers could theoretically optimize for the test rather than real-world safety, the question becomes: what happens when a new, unexpected test is introduced? The IIHS added the driver-side small overlap test in 2012 and the passenger-side version in 2018. This gives us a rare look at how well automakers were doing before they had a specific incentive to optimize for that particular crash scenario.

The small overlap tests are significant because they address a real hazard. The IIHS estimated that small overlap crashes accounted for about 25% of vehicle fatalities at the time the driver-side test was added. And this wasn't novel information — small overlap crashes have been implicated in a significant fraction of fatalities since at least the 1990s.

Who Passed the Out-of-Sample Test

Looking at 12 automakers and their scores before and after the small overlap tests were introduced, a clear pattern emerges. The results are grouped into rough tiers:

  • Tier 1: Good without modifications — Volvo
  • Tier 2: Mediocre without modifications; good with modifications — None
  • Tier 3: Poor without modifications; good with modifications — Mercedes, BMW
  • Tier 4: Poor without modifications; mediocre with modifications — Honda, Toyota, Subaru, Chevrolet, Tesla, Ford
  • Tier 5: Poor with modifications, or no modifications made — Hyundai, Dodge, Nissan, Jeep, Volkswagen

These classifications are approximations; Honda, Ford, and Tesla are the poorest fits, with Ford arguably straddling Tier 4 and Tier 5.

Volvo was the only automaker that consistently scored Good on the new tests without needing structural changes. The reason appears to be straightforward: Volvo was already running small overlap crash tests internally before the IIHS made it an official benchmark. They'd identified the risk and addressed it. Most other manufacturers reacted only after public test scores threatened their reputations.

When the driver-side test was added in 2012, most automakers modified their cars to improve scores. But when the passenger-side test was added in 2018, it became clear that those modifications were often asymmetrical. Many cars that scored Good on the driver-side test scored significantly worse on the passenger side, indicating that automakers had engineered for the specific test, not for passenger safety generally. Mercedes, BMW, and Tesla were noted exceptions that didn't need further passenger-side modifications.

The Dummy Problem

Crash test dummies represent another form of overfitting. For years, NHTSA and IIHS tests used a single male dummy based on 1970s anthropometric data: 5'9" and 171 lbs. Regulators called for a female dummy in 1980, but budget cutbacks shelved the plan until 2003. The female dummy is a scaled-down version of the male, representing a 5th-percentile 1970s woman at 4'11" and 108 lbs. When female dummies appear in frontal tests, they're always in the passenger seat. Over the same period, real U.S. adults have gotten substantially heavier — by 2019, the average man weighed 198 lbs and the average woman 171 lbs.

The consequences of this male-dummy overfitting appear in real-world safety data. Systems designed to reduce whiplash, for example, tend to be more effective for men than women. Volvo and Toyota use systems that reduce whiplash for both sexes (with slightly more benefit for women), but most automakers use systems that help men while having little impact on women's whiplash injuries. Similarly, until female dummies entered some crash tests around 2003, most manufacturers — Volvo excluded — showed a significant gender-based fatality differential in side crashes.

Other Crashes and Other Markets

Volvo claims to run crash tests well beyond what regulatory agencies require, including rollovers, rear collisions with children in the third row of SUVs, and run-off-road scenarios at a "standard" ditch facility. It's difficult to confirm whether other automakers conduct similar testing. The crash test evidence suggests they weren't even considering small overlap crashes before 2012, making an extensive suite of internal, non-standard tests seem unlikely.

The same benchmark-gaming behavior appears across markets. For example, a Nissan NP300 sold in Europe (where EuroNCAP testing applies) showed dramatically different crash performance than a superficially identical NP300 sold in Africa, where no such benchmark pressure existed until recently. The African-market version, like many vehicles sold in India before recent testing began, scored catastrophically poorly — the natural result of optimizing purely for cost in the absence of a safety benchmark.

Caveats and Limitations

Several factors should temper any conclusions drawn from these results:

Sample size is tiny. The IIHS typically tests one example of each model. Even nominally identical cars can behave differently under identical test conditions. In the case of the Dodge Dart, two tests were conducted because electrical power to cameras was interrupted during the first. In the second test, the driver's door opened when the hinges tore away from the door frame — in the first test, the door stayed shut despite severe hinge damage. Had the cameras kept working, this variation would never have been documented.

Test scores are shared between similar models. The Kia Stinger shares a score with the Genesis G70, but the G70 was the model actually tested. The Stinger is 6 inches longer and up to 500 lbs heavier. Given that ostensibly identical Dodge Darts produced different results, this is a meaningful gap.

Quality changes over time. Volvo's recent vehicles on the P3 and SPA platforms score outstandingly, but some models on the older Ford C1 platform (the 2004-2012 S40, for example) score merely Acceptable in some categories. Past or future performance isn't guaranteed.

Markets differ. Vehicles with similar names sold in different regions can have different safety optimizations, driven by the benchmarks and regulations of each market.

Weights aren't comparable across scores. The small overlap test drives the vehicle into a fixed obstacle, which is effectively crashing it into a vehicle of the same weight. A 2,700 lb subcompact scoring Good has been tested against another 2,700 lb car; a 5,000 lb SUV scoring only Acceptable is being compared against a 5,000 lb collision partner. The heavier vehicle may still be safer in a real-world accident with another heavy vehicle.

The Limits of Real-World Data

Public crash fatality data might seem like a potential way to validate crash test scores, but this data has its own problems. The IIHS's published fatality rates show considerable noise; confidence intervals sometimes run from zero upward, which implies a vehicle so safe that the probability of a fatality is effectively zero — an absurd conclusion for any 2014 vehicle. Many confounding factors aren't controlled for, such as miles driven and rural versus urban road usage. These omissions tend to make trucks look worse and luxury vehicles look better. AWD versions of vehicles often show wildly different fatality rates than 2WD versions, suggesting that who buys a car and how they use it matters as much as the car's intrinsic safety.

The IIHS data only tracks driver fatalities, not passengers. This makes it impossible to observe the impact of the asymmetrical safety improvements that were clearly made to game the driver-side small overlap test.

Beyond Crash Tests

As crash avoidance technology becomes increasingly important, crash mitigation testing is where progress lags. These tests remain relatively primitive, and the same rigor issues that plagued crash testing likely apply. Driver assistance systems' software quality appears to be a serious concern across the industry. On the other hand, well-designed driver assistance — such as GM's Super Cruise, which Consumer Reports has repeatedly rated highly — is an area where safety improvements are achievable and measurable in ways that go beyond crash test scores.

The 2021 introduction of the IIHS's tougher side-impact test (which covers small SUVs) shows that manufacturers' ability to game tests is an ongoing issue. As of the first results release, only the Mazda CX-5 scored Good, while Volvo — which excelled so dramatically on the small overlap tests — only achieved Acceptable. A 2024 analysis of fatalities per mile driven from 2018-2022 adds more context: the worst manufacturers were Tesla, Kia, Buick, Dodge, and Hyundai, while the three top-ranked manufacturers in this post's analysis (Volvo, Mercedes, BMW) had no models on the list of highest fatality rates — though all three are luxury brands, and luxury cars tend to be heavier and more expensive, factors that correlate with lower fatality rates.

The fundamental takeaway is that when safety is measured by a fixed protocol, products will be engineered to that protocol. The small overlap test results showed most automakers neglected passenger-side safety until forced to address it. It's reasonable to suspect that other unmeasured crash scenarios are similarly neglected, and that most manufacturers — with Volvo as a notable exception — are optimizing for test scores, not for safety in the broad range of accidents that actually happen on the road. The distinction only becomes clear when a new test is introduced, exposing which manufacturers were already addressing the problem and which were just checking boxes.