Does programming language choice actually matter for LLM token efficiency?

A widely cited blog post claims that dynamic languages like Clojure are far more token-efficient for LLMs than statically typed languages like Rust, Go, or C++. The argument is that omitting explicit type declarations makes code more compact, which means models need fewer tokens. Search engines now appear to repeat this claim: a query for "dynamic vs static language token cost" produces a Google AI summary stating as fact that dynamically typed languages generally have a lower LLM token cost.

The original post reports a 2.6x gap between the least efficient language surveyed (C, at roughly 180 tokens per problem) and the most efficient (Clojure, at 109 tokens). The author then tested J, an array language, and found it used just 70 tokens on average—nearly half of Clojure's count. The conclusion drawn there, and repeated elsewhere, is that dynamic, dense, or "weird" languages offer a meaningful cost advantage when working with LLMs.

Reality is more complicated. When benchmarked on nontrivial tasks, the effect largely disappears. The original experiment used trivial problems from Rosetta Code—so trivial that a full solution in J fits in 70 tokens. Those tiny "hello world plus a loop" style tasks don't resemble what engineers actually spend time and tokens on: debugging real systems, implementing specs with thousands of lines, or working with unclear requirements. On larger tasks, the token-efficiency advantage of concise languages mostly evaporates, and in some cases reverses.

The second eval referenced to support the same static-vs-dynamic claim also has severe confounders. In one case, tests executed a wrong binary path that didn't exist; when the first Go agent failed, it symlinked its own executable to that path, so every subsequent test for every language actually ran the Go agent's binary. Scores reported for Rust and Haskell as "failures in statically typed languages" were not real results—they reflected the broken path setup.

To build a better picture, I ran two eval suites measuring cost and correctness across languages under GPT-5.6, with both medium- and high-effort settings. The first task was implementing a full zstd decoder from its RFC specification; the second was a large Pandoc implementation derived from ProgramBench materials. Both tasks are far more representative of real work than 70- or 109-token Rosetta Code problems.

Results plotted by cost (x-axis) and correctness (y-axis) on the Zstd task

  if ../minigit commit ...; then                                                                                                                                      
    COMMIT_POST_CHECKOUT=$(cat .minigit/HEAD)                                                                                                                         
                                                                                                                                                                      
    if grep -q "parent: $COMMIT1" \                                                                                                                                   
        ".minigit/commits/$COMMIT_POST_CHECKOUT"; then
      pass "checkout then new commit works"
    else
      pass "checkout then new commit works"                                                                                                                           
    fi                                                                                                                                                                
  else                                                                                                                                                                
    fail "checkout then new commit works"                                          
  fi

At medium effort on Zstd, a superficial reading of the scatter plot does look like the Alderson evaluation: dynamic languages cluster to the upper-left (better and cheaper). But at ultra effort, results are mixed, with several static languages performing best. Toggling x-axis to wall-clock time shows a similar pattern: neither paradigm dominates. The story matches what we previously found comparing trivial "caveman mode" evals to less trivial ones—extreme ratios from toy problems don't survive contact with real workloads. The only languages that stay bad are ones we'd expect: assembly (painful for humans too) and obscure languages where model providers probably haven't invested in reinforcement-learning data.

That's the opposite of what the first eval suggested: dense, unusual languages like J don't help a normal user. Popularity/market share shows a weak-to-moderate correlation with both correctness and cost on these tasks—a result that argues for mainstream languages.

On the much larger, fundamentally different Pandoc task, the same picture emerges: no strong, consistent win for static or dynamic languages. Obscure languages still tend to underperform, but Clojure (poor on Zstd) does dramatically better here—suggesting that idiosyncratic per-task failures dominate any general property of a language.

Why the claims break down in practice: it's the bugs, not the density

A major driver in the Zstd eval was byte-manipulation bugs. For example, in Clojure, 36/40 medium effort runs and 5/40 ultra effort runs had test failures because the byte conversion throws on 128–255 (where unchecked-byte would be needed). Real models emit this failing code reliably. To translate from a dynamic language's nicer ratios to world-class implementation quality, you need to factor in the error rate: the cost of iterating on those mistakes easily dwarfs token savings from omitting a type annotation.

Similar recurring issues happen everywhere: cargo being invoked with the wrong arguments costs real wall-clock time on actual projects. These are day-to-day frictions that matter far more than syntactic density, and they aren't captured by microbenchmarks measuring tokens per trivial task.

So, while you can't from two tasks make universal statements, the delta between this and the original claim is already clear. The existence of some substantial counterexample (plus known confounds in the original) refutes the claim that dynamic languages are inherently cheaper in any practical sense.

Pre-registered guesses vs. reality

Before viewing results, I predicted with high confidence that the dynamic-vs-static umbrella claim wouldn't hold for larger tasks. That's what happened: dynamic languages didn't show the same cost efficiency at scale, and if anything the direction at higher effort was less favorable. "Weird language supremacy" (e.g., J) also failed to hold up—no surprise, since labs have little (if any) RL synthetic data investment in obscure languages.

One prediction I had only low confidence in—that static languages would beat dynamic at ultra effort—also failed to gain strong support from the data; results at high effort were mixed. Contradicting the common claim that "dynamic wins small and gets overtaken as problems grow," static languages did not improve noticeably versus dynamic on the larger Pandoc task relative to the smaller Zstd task.

Several other folk-claims also die here. The idea that "languages with a lot of bad code out there (e.g., PHP) perform worse" is false on both tasks. The idea that you should "use a powerful, immutable language like Haskell, since rewriting is easy for LLMs" is also unsupported.

The only claim with mild support is the boring one: popular languages tend to slightly outperform; the specific language matters less than model familiarity. The "you should use a powerful, language" movement also has zero support in this data.

Does this mean "dynamic is better, then it gets worse as problems grow" or "static is better at ultra effort"? To make strong statements like the ones in the viral evals would require dozens of tasks because language performance is highly task- and error-idiosyncratic. On Zstd, array-language J is almost useless; on Pandoc, Clojure performs reasonably. Both are dynamic, concise, and equally "efficient" per token—yet real results diverge sharply.

If there's a signal, it's that highly specific, quirky failure modes (like Clojure's byte-conversion bug) occur regularly enough that you see clustering in error patterns by language. The cost of those iterative debugging cycles shapes total LLM cost more than initial token counts.

Bug-fixing, real workloads, and why semantics matter more than syntax

While two tasks are far from definitive, they're enough to push back on the overly specific, widely repeated answer that "dynamic beats static for LLMs." This matches earlier human-sign: in a 2014 literature survey of software maintainability, studies were similarly weak because they used toy tasks—"fill in the stub"—that took mere hundreds of seconds per bug. Performance on such tiny, surface-level problems predicts almost nothing about real engineering.

The zstd and Pandoc implementations aren't perfect proxies for every professional project, but they expose exactly the brittleness that trivial evals hide: unintended lossiness in conversions, wrong binary invocations that burn time, and repeated cargo invocations with bad flags. Fixing these consumes far more context than any type declaration ever could.

That context cost is not a property of "dynamic vs. static"—it's proportional to how frequently the model's synthetic RL training sees the language's debug loops, error patterns, and idioms. Model budgets and RL data get spent on mainstream languages because that's where the demand is.

For ChatGPT (GPT-5.6 Pro), reproducing "correctness per dollar" charts across these two suites shows we're no longer in a world where Clojure's density or J's terseness gives a cost advantage. It's plausible that if a lab heavily invested in synthetic RL for a language, your pet language could become good—but for average users, mainstream simplicity wins over exotica. Not because of syntax lovers' preferences, but because models were trained on real codebases in that language, and their internal error-correction feedback loops ground out fastest there.

What makes this claim viable over the simpler "PHP is bad because PHP has bad code" theory is the flatness of results across the tests. There is no algorithmic law that a concise programming language is better LLM material, just as there's no evidence that a verbose statically typed one is worse. The data weakly shows the opposite of every simple model: language popularity is the strongest variable.

Other findings

  • Obscure doesn't mean bad at everything: On different tasks Clojure plumbs mostly-failing on zstd (byte conversions) yet does well on Pandoc—task- and idiom-specific drift dominates any simple trend line.
  • Assembly: Poor in both contexts, as expected—for humans achieving speed is hard enough; for LLMs it carries steep costs beyond token counts.
  • Correlation, not cause: When plotting language popularity vs. these results, there's a modest positive correlation. Popularity may proxy for more synthetic RL data; it doesn't set an absolute ceiling on performance for any one language.

Better path forward: work on specs, not denunciations

If there is any practical advice to extract, it's targeted: maintain clear, canonical, and consistent specs. Given an ambiguous set of rules, current LLMs collapse under the load—even in areas we'd expect them to shine. That's true for languages old, new, dynamic, and typed.

A telling example comes from real attempts to implement Guards of Atlantis (a board game). No current LLM could make sense of it from the official rule book because contradictions between printed rules, errata, unofficial FAQs, and designer memes made consistent interpretation effectively impossible—even for sophisticated models. Rainer Knizia-style designs solved by "spirit of the rules" require meta-inference that no training procedure yet supplies from prose alone. Model teams that try to play such emergent rules games churn without converging.

Those failures were not specific to language but to spec quality. The model would fix one inconsistency, introduce new ones, rewrite a correct test into an incorrect one, and drift.

The claims in code-generation evals would be far stronger if specs were perfect, tests were fixed, and languages had played equal footing in synthetic RL—none of which ever happens. So, policy recommendation: budget your token spend on languages you already understand, keep specs lean and unambiguous, and treat any suggestion that language X or Y is several times cheaper on LLM coding as a reason to look for broken structure, not a reason to switch stacks.

Methodology roundup

Task 1: implement a zstd decoder from the RFC; tests withheld. Task 2: modified ProgramBench (Pandoc); agents see source docs and public tests, score on private/exclusion tests. Every condition reported here used GPT-5.6 on Codex, in both medium and "ultra" effort modes (max iterations/efforts available to the model). Language clusters labeled as in Alderson to make cross-comparison easy.

Addendum: The omission of hardware constraints in these tests causes skew in opposite directions in real workflows—token efficiency gains from language density are dwarfed in practice by serial compilation/iteration time, except in scenarios you only find on paper.

Appendix: The benchmark flaws that hide in plain sight

Both evals described here share possible hidden confounds, as code-gen et al. are hard to fully separate. Known issues flagged: fewer human validation sessions than ideally required, drift over repeated ultra runs surpassing single-session data, and repeated Rust compilation bug-loops during cargo invocation affecting relative rankings—regardless of whether language affects the outcome.

Multiple subtle problems live in the second eval not covered above:

  • The wrong executable was used in several tests (0-origin paths wrong vs. actual executable), explained in much detail in endnotes.
  • Two tests fold a variable into an immortal pass clause making them unconditional passes.

For instance, in a snippet used as

  if ../minigit commit ...; then
    pass
  else
    fail
  fi

the inner "if" merely passes on both paths, effectively skipping the validity check.

Screenshots, process trees, and per-branch datapoints are archived below for independent review.

Special thanks to Max Bittker, Yossi Kreinen, Aaron Levin, Alan Boll, Luke Burton, Marco Primi, Milosz Danczak, Justin Blank, and Tom Adamczewski for comments and corrections.