Plug and Play Reviewer
Plug and Play ReviewerPlug and Play Reviewer

Scorecard

Measured quality, or the real refusal. Nothing below is a placeholder.

Measurement

Model
gpt-4o-mini
Measured
2026-09-10
Sample
7 holdout cases, all from the Zod repository

Quality

Precision, per finding
0.6666666666666666
Precision, per case
0.5714285714285714
Recall, per finding
0.5714285714285714
Recall, per case
0.5714285714285714
False findings per PR
0.2857142857142857
Cost (USD)
0.00460575
PRs reviewed
7

Feature flags

FeatureStatusMeasurement
RetrievalOffholdout ready; retrieval comparison not run without configured reviewers
Code graphOffholdout ready; context-source comparison not run without configured reviewers
SpecialistsOffholdout ready; specialist comparison not run without configured reviewers
LangGraphOffholdout ready; langgraph comparison not run without configured reviewers