Template

AI code review benchmark scorecard

A scorecard records method before results. Preserve the corpus, expected findings, reviewer versions, failures, and unknown cost coverage.

Direct answer

A scorecard records method before results. Preserve the corpus, expected findings, reviewer versions, failures, and unknown cost coverage.

01Dataset fields
02Adjudication
03Coverage

How it works

Start from a pinned, least-privilege workflow.

The action reference uses a released commit SHA. Store provider keys in GitHub secrets and inspect the first review before you require it.

name: Juror review
on:
  pull_request:
    types: [opened, synchronize, reopened]

permissions:
  contents: read
  pull-requests: write

jobs:
  review:
    if: github.event.pull_request.head.repo.fork == false
    runs-on: ubuntu-latest
    steps:
      - uses: Juror-AI/juror@3eb0c88ce1931dd6227d554b2d9707ee4bca123f
        with:
          github-token: ${{ github.token }}
          preset: balanced
          cost-target-usd: "4.00"

How it works

A scorecard records method before results. Preserve the corpus, expected findings, reviewer versions, failures, and unknown cost coverage.

For ai code review benchmark scorecard, use the published configuration and source repository as the product record. Keep the workflow small enough to inspect, and record any exception in the pull request rather than assuming a model result is final.

Limits to keep in view

Models, providers, and benchmark conditions change. Juror does not replace code ownership, test suites, static analysis, or a human decision to merge. Treat unknown cost and unevaluated compatibility as explicit unknowns.