All posts
2 min readby Romiel Inolino

GitHub's ReviewBench shows how to test an AI step before you ship it. Build a golden set for your automations

AI EvaluationGitHubAutomationCRMQuality Assurance

GitHub built a test set for AI code review. The same method works for the AI steps in your CRM and inbox workflows.

What happened

GitHub launched ReviewBench, an open benchmark for AI code review agents. It contains 219 pull requests from 187 open source repositories across 19 languages, picked so language and repository size match GitHub overall.

The interesting part is how GitHub built its answer key, the golden set. It gathered candidate findings from human reviewers, issues inferred from follow-up commits, deterministic tools and several frontier models, merged duplicates, then validated each one against a shared rubric. Senior engineers who had not built the dataset re-labelled every finding and agreed with it 96.6% of the time.

ReviewBench scores precision (how much of what a reviewer flags is valid) and recall (how much of the known issues it finds), and lets users weight one over the other depending on whether they want fewer false alarms or broader coverage.

GitHub says offline results have consistently pointed the same way as later production A/B tests. In one experiment combining several model runs into one review, the live test showed precision up 8.0%, recall up 13.6% and cost per review down 8.0%.

My take

Most business automations now have at least one AI step: classify this lead, tag this ticket, extract fields from this invoice. Almost none of them have a test set. Changes get judged on a couple of examples and a gut feeling.

The ReviewBench method scales down nicely:

  1. Pull 50 to 100 real records from the client's CRM or inbox, covering the messy cases.
  2. Label the right answer for each, with the person who owns the process. That is your golden set.
  3. Score precision and recall separately. For lead routing, a missed hot lead costs more than a false alarm. For auto replies, a wrong send costs more than a skipped one.
  4. Rerun the set before every prompt or model change. If the score drops, the change does not ship.
  5. Check live results against it so you know the test set still reflects reality.

This takes an afternoon to set up and saves you from silently breaking a workflow that already works.

More posts