Skip to content
Read the original: GitHub Blog · AI & ML· Published Pick63/100AI score63/100

GitHub releases ReviewBench, an open benchmark for AI code review agents

Original titleReviewBench: An open benchmark for AI code review

AISummary

GitHub has released ReviewBench, an open benchmark for evaluating AI code review agents on 219 public pull requests across 19 languages.

The benchmark reports grounded and augmented precision, recall, and F1 metrics, and its dataset, rubric, and judge are publicly available.

GitHub says ReviewBench predicted the direction of a Copilot code review ensemble experiment's production results before A/B testing.

AIWhy it matters

The post explains how ReviewBench was built and validated, and reports an offline-to-production comparison that shows how well a benchmark predicts real experiment outcomes.

Read the original github.blog

Source: GitHub Blog · AI & ML · github.blogPublished · added here