GitHub Releases ReviewBench to Compare What AI Code Reviewers Catch and Miss
The public research preview lets teams rank reviewers by coverage or lower noise. Its scoring also distinguishes known issues from valid findings outside the answer set.
Loading page…
The public research preview lets teams rank reviewers by coverage or lower noise. Its scoring also distinguishes known issues from valid findings outside the answer set.
Listen to this story
GitHub’s October 5, 2026 research preview of ReviewBench gives teams a shared way to compare AI code-review agents while tuning whether they prioritize finding more issues or avoiding false alarms. Grounded scores support cross-agent comparisons against fixed labels; augmented scores can reward valid findings missing from those labels, but their recall is not directly comparable because each agent’s discoveries change its denominator. The benchmark’s published methods and evaluation tools make it easier for developers to reproduce results and assess their own reviewers.
ReviewBench covers 219 pull requests from 187 open-source repositories across 19 programming languages.
GitHub shaped the benchmark’s workload using analysis of 103.9 million pull requests, while weighting substantive, multi-file changes more heavily.
Claude Sonnet 5 judges findings under a published rubric; an internal audit found 96.6% agreement between its reference judgments and independent senior engineers.
Teams comparing AI code reviewers can now adjust a public leaderboard to favor broader issue coverage or fewer false alarms. GitHub released ReviewBench on October 5, 2026, as a research preview, with common test cases, scoring rules and tools for evaluating an agent. The benchmark also gives reviewers credit for valid problems missing from its reference answer set.
Teams can filter results by severity and issue category, then adjust the balance between precision and recall. Precision measures the share of surfaced issues that are valid; recall measures the share of known valid issues found. F1 balances both equally. The leaderboard reorders to reflect those preferences.
There is a comparison limit: augmented recall changes its denominator according to each agent's discoveries. GitHub therefore uses grounded recall as the headline cross-system measure and augmented results as diagnostics for individual systems.
ReviewBench contains 219 pull requests—proposed code changes—from 187 public, open-source-licensed repositories across 19 languages. GitHub says it analyzed 103.9 million pull requests to model the workload, aligning the benchmark's language and repository-size distributions with GitHub overall.
The sample deliberately gives more weight to substantive, multi-file changes, reducing the dominance of tiny, single-file edits. Its reference findings come from human reviews, author follow-up commits, deterministic analysis tools and multiple model families.
Overlapping findings are merged before validation. A finding must be true, relevant and non-trivial to count; GitHub uses Claude Sonnet 5 to judge it under a published rubric.
GitHub says senior engineers who did not build the dataset independently relabeled every ground-truth finding in an internal audit, agreeing with its true-or-false-positive judgments 96.6% of the time.
GitHub says ReviewBench's offline trends consistently matched the direction of later Copilot code-review production experiments. It versions the dataset, judge and finding-matching system, and publishes its validation methodology and known threats to validity.
The public research preview includes the complete dataset, evaluation methodology, judge prompt and configuration, plus a self-serve runner. Developers can inspect the test cases, reproduce results and evaluate their own agents under the same benchmark configuration.
Loading discussion...
Join the conversation
Explain which kind of mistake would bother you more.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.