Read the original: GitHub Blog · AI & ML· Michelle Zhou· Published · added · 4d agoPick63/100AI score63/100
GitHub releases ReviewBench, an open benchmark for AI code review agents
Original titleReviewBench: An open benchmark for AI code review
AISummary
GitHub has released ReviewBench, an open benchmark for evaluating AI code review agents on 219 public pull requests across 19 languages.
The benchmark reports grounded and augmented precision, recall, and F1 metrics, and its dataset, rubric, and judge are publicly available.
GitHub says ReviewBench predicted the direction of a Copilot code review ensemble experiment's production results before A/B testing.
AIWhy it matters
The post explains how ReviewBench was built and validated, and reports an offline-to-production comparison that shows how well a benchmark predicts real experiment outcomes.
Source: GitHub Blog · AI & ML · github.blogPublished · added here