How it works
A scorecard records method before results. Preserve the corpus, expected findings, reviewer versions, failures, and unknown cost coverage.
For ai code review benchmark scorecard, use the published configuration and source repository as the product record. Keep the workflow small enough to inspect, and record any exception in the pull request rather than assuming a model result is final.
Limits to keep in view
Models, providers, and benchmark conditions change. Juror does not replace code ownership, test suites, static analysis, or a human decision to merge. Treat unknown cost and unevaluated compatibility as explicit unknowns.