LLMs Make Workload-Specific Optimization Practical
For years, the cost of performance optimization—especially for specialized workloads—has been prohibitive for most projects. A native code compiler for a regex engine, a custom JIT, or a workload-fitted search index each required rare expertise and substantial time. Recent experiments with LLM-driven development suggest that threshold has changed dramatically: optimizations that once required a specialist team can now be prototyped with minutes of human effort and a few sentences of instructions.
Marc Brooker of AWS observes that "dynamic custom software, fitted to a particular workload rather than a class of workloads, seems like a very likely outcome." As he notes, this mirrors the ethos of FFTW and of demoscene techniques—tightly optimized code targeting specific problems and hardware. Michael Malis, working on pgrust, makes a related point: JIT compilers historically helped many workloads, yet their rarity suggests they were simply too hard to justify. "LLMs have lowered the barrier to entry and made it much easier to write a JIT compiler," he argues. "Now, with AI, we can be more ambitious about the type of software we build."
Compiling Regexes on the Fly
These claims have become testable in practice. Prior work produced FRE, a regex engine built by an agent loop that ran for a month against the rebar benchmark suite—though that engine remained heavily overfit to those benchmarks until warned about a holdout set. The FRE engine's native AOT-compiled version performed well on longer searches, which raised an obvious experiment: run normal matching while compiling native code in another thread, then cut over when ready. Even accounting for the cost of losing a thread to compilation, longer queries should win.
Actually trying that experiment took minutes of human time. With a few typed sentences, an agent implemented the compiler handoff and ran benchmarks against real ripgrep queries. A few very simple queries showed 2x–4x speedups from the AOT path. For representative holdout queries where AOT made sense, the overall result was roughly a 7% improvement. That is not a revolutionary outcome, but it came from a few minutes of instructions, and the agent was still optimizing.
The Cost of Exotic Work
Admittedly, a compiler for a regex engine may be the wrong tool when an index would serve better—the author has hands-on experience with BitFunnel, the Bing search index specialized for fast ingestion. Same idea, different scale: custom indexers with multiple JITs were once major undertakings. Now the "weekend project" label applies to some of them. For someone with modest API access, a faster ripgrep plus an off-the-shelf index is probably sufficient; the whole-machine, fast-ingesting index is left as an exercise for people with access to heavier accelerators and proportionally heavier search loads.
The changing economics apply well beyond text search. A roughly month-long side project produced the strongest known game AI for the game Azul. Structured to answer the question of whether AI, not engineering, was the limiting factor, this AI dominated a published second-best on the strength of engineering: multi-threaded matching, a native code version, a shared wasm-memory-plus-JavaScript variant, and multiple search architectures paired with distinct multithreading algorithms. Multithreading was re-written several times after LLM-picked approaches initially proved wrong, and an agent handled nondeterministic debug logging with replay and reproduction—each piece tedious and weeks of hand work on its own.
For Elo-like games each doubling of search speed is worth roughly 100 Elo points, and accumulating just a few optimizations ends any reasonable race against a hand-written engine. Optimizations that change the result have an extra wrinkle—experimental design—and publicly available SOTA models are still weak at that, so human setup of the evaluation framework was necessary. Once in place, the loop was like any other optimization problem.
Similar reports come from other practitioners. Jamie Brandon faced Anthropic's (now public) performance take-home; a model picked up where he left off and improved substantially on the same problem. Brandon's assessment: many changes were plausible but he had not reached them, and "[o]thers were just crazy shit that I would never try unless I was working on this for weeks." Put directly: on a well-defined optimization problem, an experienced performance engineer does not beat a capable model under comparable time constraints.
A Practical Experiment
A regex engine built specifically around one user's grep workloads demonstrates the pattern. An agent received the training queries, optimized against them, and had to be held out against later queries. After one optimization pass, the workload-specific version was 2% faster than standard ripgrep on the held-out set and still improving. Again, a few minutes of human time, small win—like a benchmark with long-tail payoff rather than a headline gain. Note the contrast with FRE's general holdout performance: without a constant stream of feedback about holdout sets, an agent had no good way to generalize. But when the goal is optimization for one user's actual queries, plenty of data already exists and is being generated continuously.
Workload-specific optimization raises concerns about overfitting if regimes shift and old data stops applying. That is inherent and warrants care. But it improves on the old baseline where workload fitting was possible only for large-scale, well-funded projects. For individuals, the path is now trivial to start down: minutes typed, and an agent applies optimizations to either general-purpose or workload-specific targets.
Brooker's and Malis's point extends, then, beyond engineers at cloud providers and database companies. They and their colleagues are well positioned to pilot customer-specific optimization, using their users' data to guide per-customer tuning at scale. For someone outside such a company, the same practice on personal projects is viable. Because elapsed time for running an experiment may be nanoseconds for a few lines of code or hours for a full test, the act of "just trying" has been transformed.
At the scale of everything from startup prototypes to consumer software, what does not change is the need for judgement about whether an optimization is worth doing—but that judgement can be applied with far more candidates in the queue than before. "It would help by 2%" was previously paired with "but it will take N person-days to verify." Now N is not zero, but it can be tiny. And saying "no one is to blame for slow software" remains true. But user-grade LLM competence—optimizer as well as code author—should remove slow software as an inevitability for anyone who bothers to ask.
Appendix: How Codex Actually Uses ripgrep
Actual query distributions on one developer's machine show a pattern likely informed by agent behavior data (not held out to be representative wider-world): the p50 regex length is 55 characters, the p90 is 119, far longer than typical hand-written grep patterns. The simplest single patterns coexist with very complex ones; the longest are mostly long alternations of function or test names, like:
fn (hot_byte_compiler_is_generic_only_and_anonymous_count_uses_auto_count|
one_pattern_count_spans_uses_the_retained_complete_span_session|
formal_compact_state_byte_visitors_coexist_with_native_count|
fixed_boundary_record_visit_matches_line_relative_reference_and_is_atomic|
unbounded_languages_refuse_finite_extraction_before_allocation|
formal_single_raw_span_sweep_preflight|
assert_exact_fixture_uses_formal_large_continuation_sweep|
url_only_compile_identity_binds_language_and_owner_mode|
url_only_compile_exact_limits_and_runtime_refusals_close|
url_only_compile_post_plan_allocation_faults_close|
url_only_owner_discriminator_is_stable_and_precharged|
url_only_compile_owner_is_strategy_and_operation_scoped|
formal_rebar_url_owner_is_compile_only_and_matches_oracle|
formal_rebar_url_exact_fixture_uses_certified_execution|
formal_fixed_schema_materialization_matches_both_record_oracles_and_controls|
formal_single_count_selects_compact_state_byte_complete_bound_visitors|
authenticated_bound_line_total_lf_free_domain_opportunity_exceeds_five_percent|
prepared_absolute_onepass_fuses_slots_and_preserves_pre_source_fallback|
authenticated_word_boundary_russian_compact_lowering_public_canary|
ordered_nfa_x86_epsilon_edges_bypass_the_assertion_call|
ordered_nfa_aarch64_epsilon_edges_bypass_the_assertion_call|
ordered_edge_dispatch_v2_is_target_neutral_deterministic_and_relocation_free|
ordered_edge_dispatch_v2_copies_canonical_tables_and_cap_falls_back_to_v1|
ordered_nfa_v3_composes_terminal_range_and_dispatch_without_data_relocations|
ordered_nfa_x86_terminal_range_emits_authenticated_reverse_scan|
ordered_nfa_aarch64_terminal_range_emits_authenticated_reverse_scan|
ordered_nfa_x86_boundary_assertion_cache_is_lazy_and_boundary_scoped|
ordered_nfa_aarch64_caches_repeated_assertions_once_per_boundary|
boundary_assertion_cache_requires_dense_exact_kind_reuse|
boundary_assertion_cache_selection_is_compiler_only_and_deterministic)
Some are perversely numerological constructions:
:(13[0-9]|14[0-9]|15[0-9]|16[0-9]|17[0-9]|18[0-9]|19[0-9]|20[0-9]|21[0-9]|22[0-9]|23[0-9]|24[0-9]|25[0-9]|26[0-9]|27[0-9]|28[0-9]|29[0-9]|30[0-9]|31[0-9]|32[0-9]|33[0-9]|34[0-9]|35[0-9]|36[0-9]|37[0-9]|38[0-9]|39[0-9]|40[0-9]|41[0-9]|42[0-9]|43[0-9]|44[0-9]|45[0-9]|46[0-9]|47[0-9]|48[0-9]|49[0-9]|50[0-9]|51[0-9]|52[0-9]|53[0-9]|54[0-9]|55[0-9]|56[0-9]|57[0-9]|58[0-9]|59[0-9]|60[0-9]|61[0-9]|62[0-9]|63[0-9]|64[0-9]|65[0-9]|66[0-9]|67[0-9]|68[0-9]|69[0-9]|70[0-9]|71[0-9]|72[0-9]|73[0-9]|74[0-9]|75[0-9]|76[0-9]|77[0-9]|78[0-9]|79[0-9]|80[0-9]|81[0-9]|82[0-9]|83[0-9]|84[0-9]|85[0-9]|86[0-9]|87[0-9]|88[0-9]|89[0-9]|90[0-9]|91[0-9]|92[0-9]|93[0-9]|94[0-9]|95[0-9]|96[0-9]|97[0-9]|98[0-9]|99[0-9])[0-9]:
Equivalent to :(?:1[3-9]|[2-9][0-9])[0-9]{2}: and nearly identical in real-query performance. Hard to imagine a human producing that by hand; but it came from a pipeline like:
cargo clippy … | rg 'crates/fre-aot-regex/src/module.rs:' | rg NUMBER_REGEX | head -250
Unsurprisingly, agents behave differently from humans here.
Query duration has a heavy tail: p99 approaches one minute, p999 approaches ten minutes, and over a month one query ran for nearly two hours. Line numbers are often requested; PCRE2 is rare but not absent. Among regexes, about 94% of search patterns appear only once, but files searched show high locality—searches tend to repeat against recently-referenced files. Around 99% are regex searches, and about 55% of files searched contain some Unicode (only about 1% of search expressions have non-ASCII).
The entire learning loop is cyclical: informative optimizations generate new near-misses for the next optimization while repeated searches against near-identical workloads pattern-index data for agents following lines like the "compiler-side" shortcut. Fast realizations about slow red-pilled "grep everything" answers, as on a machine hammered by powerful searches, tip more engineering toward checking the "cheap first wins" that a longer session can purchase.



