ReviewBench: An open benchmark for AI code review
We all added AI review bots to our PRs, and then measured them by vibes. GitHub’s ReviewBench scores them on 219 real pull requests from 187 repos across 19 languages, deliberately weighted toward multi-file changes using distributions drawn from 103.9M PRs. Senior engineers agreed with its golden true-positives 96.6% of the time, and the runner is self-serve. Precision and recall per severity, in the open, is a good trade for the noise.