Veridict eval scorecard

Weekly system-wide pass rates and calibration metrics for the AttributionOps deliberation suite. Auto-syncs every Sunday 02:00 UTC.

Pass rate

0.0%

0 passed of 1 decidable (N = 1)

Grounded rate

0.0%

Findings supported by retrievable evidence

Run promptfoo-af5211c1-20260613 · Jun 13, 2026 · 0.2s mean elapsed

By difficulty

Pass rates split by scenario difficulty tier. The hard tier is the load-bearing reliability signal — easy-tier saturation is expected.

easy0.0% · 0/1
medium0.0% · 0/0
hard0.0% · 0/0

Calibration

No calibration snapshot has been persisted yet. Calibration metrics (ECE, Brier, selective accuracy) appear once the first weekly snapshot completes.

26-week trend

Trend appears once at least two weekly calibration snapshots have accumulated.

Methodology

Veridict is the in-house deliberation eval suite that exercises every AttributionOps reasoning surface against a frozen scenario set. Each weekly run replays the full suite and records pass/fail per scenario, plus calibration on the model’s self-reported confidence. We publish the rolled-up aggregates here for the same reason HELM, MT-Bench and the HuggingFace Open LLM Leaderboard publish theirs: a closed eval is not a credibility signal.

Methodology version: v1

Citations

  • Liang, P. et al. 2023. Holistic Evaluation of Language Models (HELM). Stanford CRFM. arXiv:2211.09110.
  • Zheng, L. et al. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023.
  • Papineau, K. et al. 2002. BLEU: a Method for Automatic Evaluation of Machine Translation. ACL 2002.
  • Guo, C. et al. 2017. On Calibration of Modern Neural Networks. ICML 2017.
  • Gibbs, I. & Candès, E. 2021. Adaptive Conformal Inference Under Distribution Shift. arXiv:2106.00170.