Positioning spread: every benchmark, one model
| Benchmark | Task acc | Confidence | F₁ | Leans |
|---|---|---|---|---|
| SQuAD (factual recall) | 0.45 | 0.61 | 0.67 | +0.16 overconfident |
| MMLU-Pro (knowledge) | 0.46 | 0.92 | 0.64 | +0.46 overconfident |
| LegalBench (legal reasoning) | 0.84 | 0.78 | 0.81 | -0.06 cautious |
| MathBench (competition math) | 0.95 | 1.00 | 0.97 | +0.05 overconfident |
| OmniMath (advanced math) | 0.64 | 0.99 | 0.79 | +0.35 overconfident |
| SciCode (scientific code) | 0.56 | 0.47 | 0.52 | -0.08 cautious |
In the full cloud
gemini-2.5-flash conditions all other model/condition points equal relative confidence and pass rate
Pairwise signal: pairs involving gemini-2.5-flash
Match accuracy controls for the performance base-rate gap
gemini-2.5-flash pairs
18/ 171
gemini-2.5-flash mean tau
+0.006
All-pairs mean
+0.037
gemini-2.5-flash p<0.05
6(33.3%)
all model pairs (observed) base-rate-matched null calibration-preserving null gemini-2.5-flash pair (filled = p<0.05) gemini-2.5-flash mean all-pairs mean
The four metacognitive outcomes
No curated cases for this selection yet — outcome-matrix extraction currently covers a sample of MMLU-Pro trials.