Measuring an AI reviewer is awkward because "good" depends on what you want from it. One system flags more issues, another stays quiet; one is strong on critical defects, another also reports smaller improvements. What teams actually need is a way to see what a reviewer catches, what it misses and which tradeoff it made — and, for those building review agents, an offline signal that reliably points the same direction as production.
ReviewBench is a code review offline benchmark built to close that gap. It mirrors the language, repository size and pull request size distribution observed across more than 100 million GitHub pull requests, draws its ground truth from multiple sources judged under one rubric, and has been independently validated by senior engineers. The research preview is available now.
103.9M GitHub pull requests analyzed for language, repository size and change shape
The corpus and how findings are established
ReviewBench holds 219 public pull requests from 187 public open source licensed repositories across 19 languages. Language and repository-size distributions track GitHub overall, with one deliberate deviation: pull request size is weighted toward the reviewable middle and tail, so tiny single-file changes are less overrepresented and substantive multi-file reviews are preserved. The complete dataset is public.
No single producer — human or model — finds everything worth finding in a PR. The golden set is therefore assembled in three stages:
- Collect candidates from diverse sources. Findings come from real human reviewers, issues inferred from author follow-up commits, deterministic analysis tools, and multiple frontier LLMs across model families.
- Deduplicate semantically. Findings describing the same underlying issue are merged, so coverage broadens without letting agreement between producers inflate the golden set or letting one source's blind spots define it.
- Validate under one rubric. Provenance carries no weight: a finding counts as a true positive only if it is true, relevant and non-trivial.
Claude Sonnet 5acts as the LLM grader, applying that rubric uniformly across every submission. Both the rubric and the judge configuration are published.
Six metrics, and why the split matters
Most benchmarks report precision and recall against a fixed golden set. ReviewBench reports six metrics in two families.
- Grounded precision, recall and F1 use only existing gold-set labels — the strict comparison of how many known issues an agent found and how many of its findings matched something known.
- Augmented precision, recall and F1 also score findings that match nothing in the golden set. Here the judge decides independently whether each unmatched finding is a true or false positive, so a reviewer can get credit for a valid issue no golden-set producer surfaced.
The distinction grows more important as agents improve: any fixed golden set becomes incomplete once systems find issues its builders did not anticipate, and augmented metrics recognize that instead of penalizing it. Because augmented recall's denominator expands with whatever an agent discovers, grounded recall is the headline cross-system comparison and augmented metrics serve as a per-system diagnostic.
Tuning the benchmark to your preferences
There is no universally optimal review experience. Some developers want only critical issues; others value lower-severity, non-breaking findings. Some want broad coverage, others minimal noise, and some have specialized needs such as security- or privacy-focused review. ReviewBench results can be sliced by severity — critical, medium, low — and by category, including correctness, security, reliability, maintainability and testing. Adjusting β in the Fβ score shifts weight toward recall for coverage or precision for lower noise, and the leaderboard re-ranks accordingly.
Reliability checks behind the numbers
Senior engineers who had no part in building the dataset independently re-labeled every ground-truth finding from scratch before release. Their true/false-positive judgments agreed with ReviewBench 96.6% of the time. The dataset, judge and matcher used in each evaluation are versioned, so results can be compared under identical configuration and revalidated when the benchmark changes. The validation methodology, agreement measurements and known threats to validity are published as well, along with the published rubric, a human-labeled dev set, a grader calibrated to human judgment and uniform labeling across all sources.
Benchmark movement is also checked against online experiments: improvements and regressions observed offline have tended to appear online too.
Using ReviewBench
The full dataset — pull requests, findings, labels, severity and category annotations — is public, so you can inspect exactly what systems are evaluated on and reproduce results. Evaluated agents are published on a common leaderboard with views by overall performance, severity, category and precision–recall preference. To evaluate your own agent, the benchmark dataset, evaluation methodology, LLM judge prompt, judge model configuration and self-serve runner are all available publicly, letting you inspect strengths and gaps and iterate against the same configuration. Onboarding and result submission are handled through the ReviewBench website.
Copilot code review as a working example
Copilot code review (CCR) has been evaluated with ReviewBench across successive iterations, giving a consistent measure of progress and a way to catch regressions and prioritize changes. The most useful property is the early offline signal: across experiments assessed with ReviewBench before A/B testing, offline movements have consistently pointed in the same direction as production results.
A recent lite-tier experiment illustrates this. A multi-model ensemble review, which combines several independent model runs into one review instead of relying on a single run, was predicted by ReviewBench to raise precision, recall and comment volume while lowering cost per review. The online counterparts used for comparison are addressed rate — the share of CCR comments that an LLM judges to have prompted a corresponding code change, determined from the diff, thread, reactions, resolution state and post-review code — as the analogue of precision, and the amount of additional human review still required as the measure of recall.
The A/B test followed the predicted direction relative to the production control:
- Addressed rate (precision): +8.0%
- Recall: +13.6%
- Comment volume: +61%
- Cost per review: −8.0%
Volume alone says nothing about quality, since a critical finding is not a low-severity nit. Severity-level evaluation captured the composition shift as well: ReviewBench predicted a 227% increase in critical comments against 262% online, alongside the same broader movement toward more moderate comments and fewer nits.
Online experiments remain the definitive measure of user impact. ReviewBench's contribution is a fast, repeatable signal about which changes are worth taking that far.
Submitting a run
Entries are registered and scored through the ReviewBench site, starting with a GitHub sign-in. Each submission needs a container image, a configuration, and your own model key — the judge is supplied by the benchmark, not by the entrant.
Tuning happens against a smaller set before anything reaches the leaderboard:
- Sign in with GitHub on the ReviewBench website.
- Register your agent. Provide a container image, your configuration, and your own model key. We provide the judge.
- Try it on the test set. Run against a 25-PR test set with per-PR detail and repeat as you tune your configuration.
- Do a final run. When you’re ready, run the full set of 219 pull requests (three rounds), scored by the same judge as every other entry.
- Publish to the leaderboard. Your scores remain private until a maintainer reviews and approves the submission. Scores are published to the leaderboard only if they outperform the agent’s current leaderboard score, or if this is the agent’s first leaderboard entry.
The 25-PR test set exists purely for iteration; the 219-PR full set, run three times, produces the numbers that count. Visibility is gated twice over: a maintainer approval step before publication, and a rule that a new score appears on the leaderboard only when it beats the agent's existing entry — otherwise an agent keeps its first published result.
Who built it
ReviewBench came out of a joint effort between GitHub and Microsoft, spanning methodology design, pull request curation, the golden set, and the evaluation pipeline — work that was done so that the benchmark can be run by anyone.



