Veridict eval scorecard
Weekly system-wide pass rates and calibration metrics for the AttributionOps deliberation suite. Auto-syncs every Sunday 02:00 UTC.
Pass rate
0.0%
0 passed of 1 decidable (N = 1)
Grounded rate
0.0%
Findings supported by retrievable evidence
Run promptfoo-af5211c1-20260613 · Jun 13, 2026 · 0.2s mean elapsed
By difficulty
Pass rates split by scenario difficulty tier. The hard tier is the load-bearing reliability signal — easy-tier saturation is expected.
Calibration
No calibration snapshot has been persisted yet. Calibration metrics (ECE, Brier, selective accuracy) appear once the first weekly snapshot completes.
26-week trend
Trend appears once at least two weekly calibration snapshots have accumulated.
Methodology
Veridict is the in-house deliberation eval suite that exercises every AttributionOps reasoning surface against a frozen scenario set. Each weekly run replays the full suite and records pass/fail per scenario, plus calibration on the model’s self-reported confidence. We publish the rolled-up aggregates here for the same reason HELM, MT-Bench and the HuggingFace Open LLM Leaderboard publish theirs: a closed eval is not a credibility signal.
Methodology version: v1
Citations
- Liang, P. et al. 2023. Holistic Evaluation of Language Models (HELM). Stanford CRFM. arXiv:2211.09110.
- Zheng, L. et al. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023.
- Papineau, K. et al. 2002. BLEU: a Method for Automatic Evaluation of Machine Translation. ACL 2002.
- Guo, C. et al. 2017. On Calibration of Modern Neural Networks. ICML 2017.
- Gibbs, I. & Candès, E. 2021. Adaptive Conformal Inference Under Distribution Shift. arXiv:2106.00170.