Scorecard
Measured quality, or the real refusal. Nothing below is a placeholder.
Measurement
- Model
- gpt-4o-mini
- Measured
- 2026-09-10
- Sample
- 7 holdout cases, all from the Zod repository
Quality
- Precision, per finding
- 0.6666666666666666
- Precision, per case
- 0.5714285714285714
- Recall, per finding
- 0.5714285714285714
- Recall, per case
- 0.5714285714285714
- False findings per PR
- 0.2857142857142857
- Cost (USD)
- 0.00460575
- PRs reviewed
- 7
Feature flags
| Feature | Status | Measurement |
|---|---|---|
| Retrieval | Off | holdout ready; retrieval comparison not run without configured reviewers |
| Code graph | Off | holdout ready; context-source comparison not run without configured reviewers |
| Specialists | Off | holdout ready; specialist comparison not run without configured reviewers |
| LangGraph | Off | holdout ready; langgraph comparison not run without configured reviewers |