How it works
Review can question intent and integration; tests can repeatedly check specified behavior. Both need ownership and neither catches every defect.
For code review vs testing: where each fails, use the published configuration and source repository as the product record. Keep the workflow small enough to inspect, and record any exception in the pull request rather than assuming a model result is final.
Limits to keep in view
Models, providers, and benchmark conditions change. Juror does not replace code ownership, test suites, static analysis, or a human decision to merge. Treat unknown cost and unevaluated compatibility as explicit unknowns.