Documentation

Benchmark Juror

A benchmark is a repeatable adjudication process, not a single score. Preserve all failed or skipped reviews and report unknowns alongside results.

Direct answer

A benchmark is a repeatable adjudication process, not a single score. Preserve all failed or skipped reviews and report unknowns alongside results.

01Corpus
02Command
03Limitations

How it works

Start from a pinned, least-privilege workflow.

The action reference uses a released commit SHA. Store provider keys in GitHub secrets and inspect the first review before you require it.

name: Juror review
on:
  pull_request:
    types: [opened, synchronize, reopened]

permissions:
  contents: read
  pull-requests: write

jobs:
  review:
    if: github.event.pull_request.head.repo.fork == false
    runs-on: ubuntu-latest
    steps:
      - uses: Juror-AI/juror@3eb0c88ce1931dd6227d554b2d9707ee4bca123f
        with:
          github-token: ${{ github.token }}
          preset: balanced
          cost-target-usd: "4.00"

How it works

A benchmark is a repeatable adjudication process, not a single score. Preserve all failed or skipped reviews and report unknowns alongside results.

For benchmark juror, use the published configuration and source repository as the product record. Keep the workflow small enough to inspect, and record any exception in the pull request rather than assuming a model result is final.

Limits to keep in view

Models, providers, and benchmark conditions change. Juror does not replace code ownership, test suites, static analysis, or a human decision to merge. Treat unknown cost and unevaluated compatibility as explicit unknowns.