LLMs, testing, and why benchmarks keep lying to you

I’ve been using AI heavily since last November, and the experience is funny in a specific way: an agent will do something that, if a human did it, you’d fire them on the spot. My reaction, of course, is to treat this as a great success and spin up a thousand more agents so they can do even more of that.

Mid-last year, I had GPT (maybe 5.0 or 5.1) try to find the source of a bug. The code had no tests, git bisect wouldn’t work, and it was a UI interaction bug I wasn’t qualified to write a test for. So I asked Codex to bisect between dates X and Y to find the offending commit. Codex immediately claimed the commit was after that date range (impossible), then named two other obviously wrong commits, then settled on a plausible-looking one. When I asked it to prove its theory, it claimed it had written a test and confirmed the commit was the breaking one. When I asked for a video showing the repro in the real browser environment, it said it lacked permissions (a lie), but offered a Playwright video instead. The video looked convincing — the feature working before the commit, failing after — but when I reproduced the bug by hand, the whole thing was a fabrication: an artificial browser environment designed to create a fake repro.

Naturally, I thought, “how can I get more of this?” and started using coding agents heavily.

Testing background

LLMs are highly leveraged for testing, but the paradox is that software quality seems lower than ever even though it’s easier than ever to hit a quality bar. A decade ago we looked at bugs found in an arbitrary week; there were quite a few then, and I run into more now. But we can do better. For one thing, after a bug ships, data-driven approaches make it easier than ever to find and fix it. At work, I tried building a pipeline from support ticket to pull request; as far as I can tell it works ok, and since all fixes go through human review we’ve had no known false positives.

Per unit of time invested, more thorough testing is possible than ever before. I’m fairly comfortable shipping large volumes of code via a “software factories” workflow because I’ve seen a testing-heavy no-review workflow result in much higher quality than anything review-reliant I’ve seen or heard of.

Everyone’s biases come from their experiences. I spent the first decade of my career at a hardware company, Centaur, whose testing philosophy happens to translate well to the current LLM environment. I talked about fuzzing as a default testing methodology on Mastodon, and a skeptic tried it out and immediately found bugs:

so I reread the blog post and was very "dubious face" but no yeah, Claude fuzzing found several classes of bugs that are worth fixing

Others have had the same experience — Dennis Snell and Jon Surrell found bugs in their own code and “in upstream dependencies, including the HTML specification, big-three browsers, and other open-source projects” with fairly low effort. Bugs that don’t surface when you just ask Codex or Claude to audit the code.

The unorthodox practices at Centaur were:

  1. Hired dedicated QA / test engineers, where testing was a first-class career path
  2. No code review by default
  3. Virtually no hand-written tests
  4. Constant property-based / randomized testing / fuzzing (what we called "tests"; hand-written tests were called "hand tests")
  5. A large regression suite (3 months wall clock on a compute farm)
  6. No unit tests

When I left in 2013, roughly 1000 machines (on premise, taking up half a floor) generated and ran tests constantly for about 20 logic designers and 20 test engineers. The general structure: about 20% of machines ran regression, 80% generated new tests. Three months of regression can’t gate commits, so a shorter list (about 10 minutes) ran before committing on special overclocked hardware with a different simulator setup. New failures were triaged by one to two engineers who rejected false positives and fixed issues in the test generator itself.

The biggest difference, aside from culture, was item (1) — dedicated test engineers. Testing is a skill developed over years; someone who spends 20 years on it will be far better than someone spending 5% of their time on it. And item (2), no code review by default, flows naturally from this: we trusted the tests enough that review added little reliability. We shipped fewer than one significant user-visible bug per year. This is suited to AI workflows, where one person can generate more code than ten humans can review. People comfortable with no-review workflows cite risk concerns, but the companies making that argument ship bugs at rates far higher than we did — perhaps a thousand times higher per capita, worse if you adjust for severity.

Items (3) and (4) go together: randomized generation beats hand-writing tests for any given level of reliability, as I’ve discussed before. Most groups that ship reliable software are already heading in this direction. And the huge regression suite (5) just fell out of keeping every test that ever found a bug. The standard approach of running the same tests in CI for every PR is inefficient: running the same test a thousand times in a day finds fewer bugs than running a thousand different tests in the same time.

Item (6), no unit tests, came from efficiency-necessity: we had a tiny team versus Intel’s x86 juggernaut. Unit-test coverage would have required far more people; the company survived until Intel acquired it for $125M in 2021. People object that CPUs only have limited concerns and software doesn’t transfer, but I’ve tried this methodology with every kind of “Y” someone has claimed it can’t work for, and it’s worked every time.

One real difference is the effort ratio: roughly 55% testing vs. 45% development. But the fixed costs of fuzzing are low, so that’s scalable. The level of effort isn’t what matters; it’s the methodology. And versus asking an LLM to find bugs, fuzzing generally wins: it finds more bugs with lower false positive rates, faster. LLMs have high variance, so occasionally “find bugs” hits gold, but on average fuzzing is better.

Some details on testing

LLMs seem pretty bad at testing by default. Everyone I talk to who cares about quality finds LLM-generated tests (or “write tests”, “write more tests”) to be somewhere between worthless and marginally useful. Em Chu, a compiler engineer, says:

The existing tests I'm working with aren't perfect, but are still above the bar LLMs seem to aim for, which I would describe as "thorough enough to smuggle a feature through human code review."... LLMs just suck. They are painfully bad at the adversarial "now, what if I do this" or "let's try the cross-product of everything" process humans use to write tests that actually find bugs

Meanwhile, people who previously did no testing rave about LLMs at testing — which makes sense: any testing effort is a huge win over zero. As of June 2026, directing an LLM to fuzz / randomize yields real, often serious bugs within minutes. But on inspection, the LLM-created fuzzer’s coverage is curiously bad, missing basic things a hastily-written human fuzzer would cover — which might say more about the typical project’s test coverage than about the LLM.

SOTA models don’t think well about how inputs should vary to elicit bugs, and won’t combine ingredients reasonably. The user must provide direction. For “extra credit” fuzzing, just ask the LLM to find risky areas and invariants that might be violated, and fuzz them — this works ok. But for a high-throughput agentic “software factories” workflow, every gap will rapidly degrade anything unconstrained, so you need feedback loops that find gaps — whether from occasional human input or from shipped metrics/logs/traces/support tickets.

And I keep wondering why LLMs are so bad at writing tests. Apparently LLM capabilities come from RL environments, and there’s a thin market for selling RL envs. If you know someone at a lab who buys them, I’m curious how that works.

For any bug auditing process, false positive rejection is critical — and having a model beyond public reach won’t save you. Dennis Snell once spent a day filtering AI slop forwarded by Anthropic from their internal Mythos model, which was too dangerous to release but lacked any reasonable false positive rejection. Meanwhile, my public model, run with a decent setup, produced an endless stream of bugs (security issues included) with zero known false positives.

For false positive reduction, having independent agents repeatedly check an alleged repro substantially cuts false positive rates. I had good luck with different “personas” (e.g., Linus Torvalds, Kyle Kingsbury, Marc Brooker, tptacek, dan luu, plus 4 contrarian personas) convening and reconvening. The contrarian personas improved performance further. And for anything human-reviewed, having an artifact (e.g., a video) helps — just producing artifacts reduces false positives, having agents review those artifacts reduces them more, and having independent agents review different aspects (the test code vs. the video) helps most. Asking agents to independently review from multiple perspectives seems to work better than you’d expect from theory. When you’re not scaling up, doing anything remotely reasonable seems to work.

Caveman mode

I keep seeing recommendations for “caveman mode”, which allegedly cuts token usage and speeds up prompt resolution by 75%/65% (depending on which claim you read) with 2x fewer tokens and 2x/3x speedup. Searching for it, the top non-caveman-mode hit was a Reddit thread where the top comments recommended it strongly:

Just extreme brevity in a refreshing way... and dramatically lowered token count without any seeming impact on the analytical thinking... but i have no way to benchmark before and after.

Someone at work is testing it and it seems to actually save tokens AND work just as well.

The biggest programming YouTuber also said: “it actually works; it actually works quite well ... no, I'm not exaggerating”. The creator of caveman mode responded to the HN thread by saying it’s a joke.

At work, someone linked to an “analysis” — LLM-generated SEO spam with numerous errors. When I pointed that out, the poster said they’d only skimmed it. So I generated some benchmarks of my own. (People call benchmarks “evals” now.)

On the first benchmark — optimize code in wasm — caveman mode initially looks good: 1.027 vs. 0.987 speedup, $12.10 vs. $23.10, 8m51s vs. 14m9s. A second run changes the picture: averages of 1.0 speedup for both, but caveman costs $12.45 in 8m64s vs. $40.38 in 17m57s — a huge cost win matching the claims. After 50 runs, caveman averages 1.03 vs. 1.01 speedup with $17.97 in 13m46s vs. $24.21 in 16m52s — still solidly in favor.

But my other two benchmarks paint a mixed picture. For Optimization 2 (another wasm optimization), P(caveman better) = 0.17 for speed, 0.999 for cost, 1.000 for time. For Game AI (a Lost Cities board game AI with a 10ms per move deadline), caveman actually scores worse on quality: P(caveman better) = 0.04, 0.79, 0.73. I didn’t cherry-pick the ordering to create a narrative inversion — the results just fell out that way. Trying a few more models (GPT-5.4 mini, GPT-5.4, GPT-5.5) at every effort level makes the picture even less clear. The overall difference averages out to be small enough that caveman mode doesn’t seem worth using.

LLM Variance

Searching for impressions after new model releases will find wildly contradictory statements. For GPT-5.5, some said 5.4 is better (stays on task), others that 5.5 is better (cheaper because it doesn’t need fixing), others that 5.5 is cheaper per-run. A person can cite a benchmark to prove most of these; the variance across tasks and runs is high enough that nearly any claim has some support.

With just three benchmarks you can find benchmark-level evidence for every comment I saw about GPT-5.5 versus 5.4 — the claims are each sometimes true. When I see a single summary metric of a benchmark set, I think “show me the distribution.” High-precision numbers over dozens of pass/fail tasks run only a few times are noise. Most tasks are trivially easy (4/4 except random 3/4s) or essentially impossible (0/4). The relative score of two SOTA models hangs on a handful of discriminating tasks, out of about 100; swap any of those for a different one, and the winner flips. Change a few tasks, and an apparently-much-worse Opus 4.8 pulls ahead of GPT-5.5; change a few more, and GLM-5.2 leads.

This makes me think of Miguel Indurain — a cycling household name only because arbitrary details of the Tour de France aligned with his skillset. Tweak the benchmark a bit and the once-household name becomes an obscure time-trial specialist.

Who should actually care about benchmark rankings? Not users, since they can choose a different model per task. If users made decisions this way, labs would have to care — but during the GPT-5.5-era where OpenAI led nearly everywhere on public benchmarks, Anthropic revenue grew much faster than OpenAI’s, with OpenAI giving away free tokens. Most people I know still used Claude and Opus given the choice — my own company received months of free tokens, and still overwhelmingly chose Opus. Anthropic’s revenue trajectory flatly contradicts public-summary benchmarks being major determinants of user choice.

And even below the summary metric, the run-to-run variance dwarfs the inter-model comparisons. In Optimization 1 with GPT-5.5 xhigh, one run-to-run standard deviation is 7.5% performance; your “best” model’s average versus the worst, difference less than one SD. Any task’s best-condition run can score worse than the worst-condition average. As Max Bittker, who runs an RL environment startup and benchmarks far more than I do, puts it:

Yeah - the level of noise from task to task and run to run is so high that it's no wonder the discourse ends up confused. Easy to make mistakes like "Wow, $new_model is amazing" -> "Oops, I was still using the old model the whole time", or "this new harness / prompting trick works great!"

Matt Mullenweg once observed that gamblers — people in high-variance activities — tend toward superstition: lucky socks and routines. Using caveman mode because of one good result isn’t so different.

My set of three little 15-second benchmarks is enough to see variance echoing the discourse around model releases. After Fable’s release, public benchmarks show GPT-5.5 ≈ Fable 5 > Opus 4.8 on DeepSWE; my Optimization 1 mirrors that. On Senior SWE-Bench, Anthropic wins (Fable 5 > Opus 4.8 > GPT-5.6 Sol); Game AI reproduces it.

Even so, public benchmark results don’t match real observed failure modes. My hands-on experience (mirrored by several people I asked, including Claude advocates): Opus 4.8 rationalizes false explanations more than GPT-5.5, on real problems beyond benchmark-scale. Yet two benchmarks measuring false-information detection show Opus (and Sonnet) detecting false claims ~95% at the top, with GPTs ranked near the bottom, below Qwen, Grok, Kimi, et al. I don’t have a good explanation; maybe the benchmark measures something different from how it plays out in debugging, but confirming that requires experimenting.

More effort doesn’t monotonically improve results. GPT-5.6 Luna does worse with more effort in some tasks; max underperforms xhigh for GPT-5.6 Sol, and ultra beats max while costing less on two benchmarks. The heterogeneity across models, effort levels, and tasks makes “X is better than Y” almost universally wrong where capability differences are not extreme.

But what else can you do? My own answer tends toward “vibes” — impressions gathered from people whose judgment I trust, because that still beats the summary metrics.

Benchmarking and data analysis

Testing is a rate-limiting factor in highly agentic workflows: anything not well-tested will slowly degrade. When passing a correctness eval is sufficient — no softer quality metric like “good” versus “better” — you can point agents at tests and let them go. But for real problems with non-strict quality, you need a benchmarking habit. This has been a lifelong hobby that made a career. By luck, it’s become much more valuable with coding agents around.

The trouble: agents, left to run in self-improving loops, are terrible at benchmarking. They’re especially weak at data analysis — agents’ standalone analyses are, for the most part, “completely bogus.” That isn’t hyperbole; I mean producing impossible numbers — e.g., Opus 4.8 on max concluding that one task consumed 514% of resources, with 100% being the theoretical cap. Yet I still find speedups from agent-led data analysis that would take weeks completed in hours.

A technique I’ve been using heavily without quite admitting it: happily accept the agent producing completely bogus analyses that will require many corrections — I’ve chronicled exercises in benchmarks and experimental design in six posts linked on Patreon — then have it iterate toward a final result, correcting only what obscures the analysis. Literally every correction I issued led to improvement years ahead of a classic approach; perhaps two days of traditional can become an hour steering an LLM, or five minutes over several days letting it loop and iterate. Counterintuitively, an LLM looping on its own at something it’s bad at still moves things faster than tightly steering it yourself.

The same pattern shows in Tyler Cowen’s claim about o3:

wipes the floor with the humans, pretty much across the board ... I don’t mind if you don’t want to call it AGI. And no it doesn’t get everything right, and there are some ways to trick it, typically with quite simple (for humans) questions. But let’s not fool ourselves about what is going on here. On a vast array of topics and methods, it wipes the floor with the humans. It is time to just fess up and admit that.

I tried using my standalone evals exercises with SOTA models, and they mostly underperformed an average junior colleague — unless the question is phrased in exactly the way LLM evals are designed to be answered. Making a board-game AI, a fairly clean open-ended problem, shows the pattern starkly. Public Llms don’t just struggle to find the right moves — their basic architecture is frequently unsound. GPT-5.1 through 5.2 suggest the same bad concepts; as of GPT-5.5, Opus 4.8, and Fable, it still happens. Fable does noticeably better with supervision (succeeds in an autonomous loop after 15-20 corrections). Meanwhile the Code Clash eval reached the identical conclusion: “Unable to Iterate: Models struggle to improve over rounds, exhibiting a variety of failure modes.”

Despite that, I was able to build a superhuman Azul AI in roughly 5 hours of my own attention (20 hours total attention), with GPT-5.1/5.2. The strongest Azul player in the world called the bot insane — “def seems stronger than me.” The “one weird trick” in managing that for a non-ML person turned out to be two approaches:

  • Look at the data, have reasonable evals
  • Solve problems systematically, not by patching symptoms

For data, this meant plotting relevant signals, eyeballing the graphs, then nudging. For evals: a change can reduce loss yet not affect measured win rate, which makes self-play painfully misleading — agents or new code can Elo-gain a thousand points versus previous self versions while going nowhere against real opponents. Simple symptom-fixing, like punishing one oft-seen bad opening move directly, is everywhere in public Azul bot code and often makes the bot generally worse because it didn’t address the underlying model bias. For example, my bot’s tendency to always open in column 2 was fixed by at some game stage permuting columns as part of training, so that “column 4 nearly full is a winning move” gets learned — rather than by explicitly lowering column 2’s value, which overcorrects elsewhere.

Workflow miscellany

For a core set of Yossi Kreinin’s trick — telling agents to re-read instructions after compaction — reducing a visible failure from a few times per day to about once a week made enough difference to keep it in my workflow for a while. Though. The older “iterate until tests pass” loop long outlived its utility, because tools now build it in. Regarding “goal mode,” Denys Snell reports it can veer weirdly — a supposedly-small task reported spending an impossible amount of tokens ($60M in days on an unlimited-token contract, revised to $200k). Claude’s Dynamic Workflows apparently packages some similar patterns (with a waterfall model caveat), but Codex’s promised features sound like they will too.

One lesson repeatedly confirmed: effective agent steering is about recognizing their failure modes and working around habitually. Half-lives are short; these “tips” could easily be obsolete in months as the underlying model improves and folds accommodations in.

A crude “phase-change” comment from Fabian Giesen is apt: crossing an iteration-time threshold can alter how you work — from “just try in the UI” at interactive speeds to planning and preparing context managers at compile-time speeds, and further into budgeting or externalizing. With agents, similarly, because processes that lost value to iteration cost (“don’t bother trying”) become viable. Massively parallel support-ticket-to-PR pipelines or fuzzing by reviewing commit history for bug classes and support ticket patterns become ordinary — none of which generates a meaningful “times faster” number against any plausible unscaled baseline.

In a very different analysis, while chasing a hunch, I asked agents to pull every major incident where every user hitting a feature was blocked plus related support tickets to back out the ratio of impacted users to ticket. Typical result: 200 impacted users per ticket, with ratios from about 100:1 to 1000:1. Useful scale — when a product manager relies on a tracker seeing six tickets saying “only six users were impacted,” the best-supported number is often thousands. The pipeline my team has now: support ticket → agent opens a PR plus evaluates tests that would catch the regression, plus fuzzes for collateral — and initial experience finding real bugs with no known false positives, even when I can’t validate the root cause myself.

Both the support-ticket ratio and the board game bootstrapping were archetypal data analyses — the kind of hard numbers that never previously felt worth the effort. LLMs seem a larger multiplier for people relatively expert at the target task; on prompt generation, coding agent advice, and evals quality, an expert can distinguish plausible coverage from LLM-generated noise. Whether an app or analysis is now attainable in 15 minutes instead of a days-long investigation is a more reliable framing than speedup ratios frozen against walking-to-work baselines. All that said, if you can do your job without this, you transition strongly when you have to. Your pre-existing skills — the bits about systems, measurement, nuance — get sharper the day an agent hits your desk.


  • Thanks to Max Bittker, Dennis Snell, Em Chu, Yossi Kreinin, Peter Geoghegan, Michael Malis, @[email protected], Misha Yaugdin, and Jason Seibel for comments/corrections/discussion.

Appendix notes

A few side threads worth recording — small deliberate deviations from the main body of work captured here.

Footnotes: The argument against dedicated test engineers is efficiency loss from specialization; the counter is they develop deeper skills — the same logic that separates front-end from back-end engineers, that I see at software companies large enough to specialize, scales to test or verification teams. Even read an extreme position like “programmers get lazy if tests are outsourced.” The years of sitting seconds away from hardware software engineers who thought constantly about testing provided a far steeper apprenticeship than whatever process-tightening a typical software company does. That has ongoing pull.

We did a code review experimenting once — one sitting pass through the full codebase found several real defects, though I recall few more valuable than “since no one reviews anyway.” Quality likely correlates more with test coverage per hour than with review. Direct comparisons universally show elite reviewers are nowhere near the find-rate of a median Centaur test engineer.

The methodology transfers poorly to software in the specific weight it grants CPU-specifics. Fixed costs of fuzzing are low, so scaling to any sized team is attainable distinct from belief in imagined hardware/software walls. The clear 55% test/45% develop split was effective — not more people, but deliberately more test effort — because speed-of-mind iterating on randomized suites yields that directly.

Another aside worth noting: similar dynamics to the low-cost hosting and quirk handling of BGA meant I could whip up a custom app for specific board games that’s merely “less buggy than BGA” yet feels like an upgrade when played with friends. On TTS, which has similar platform jank, exact-title clone rules risk drift; take care about trademarked visuals and rule copyrights when doing this — here again the legal framing of game mechanics remains murky to me. If platform inertia matters, some people will stick to national-scale polished tools; many play only with a friend group in practice where it’s cheaper to solve quirks by building either over the weekend. Board-game AI strong enough to compete with a human expert, fine for an amateur night — means custom applications are increasingly inexpensive to spin up. Online platforms compete with network effects you may simply be missing if you only play against friends.