How it works
A benchmark is a repeatable adjudication process, not a single score. Preserve all failed or skipped reviews and report unknowns alongside results.
For benchmark juror, use the published configuration and source repository as the product record. Keep the workflow small enough to inspect, and record any exception in the pull request rather than assuming a model result is final.
Limits to keep in view
Models, providers, and benchmark conditions change. Juror does not replace code ownership, test suites, static analysis, or a human decision to merge. Treat unknown cost and unevaluated compatibility as explicit unknowns.