Does the machine know when it is guessing? · 机器是否知道自己在猜
A reliability diagram plots, for each confidence level, how often the model was actually right. On the dotted diagonal, confidence equals accuracy — the model's certainty is honest. A specialist restorer hugs the line; a general LLM floats high and to the right: near-total confidence, near-zero accuracy. The two right-hand diagrams are the fair pairing — the same 2–3-character (v2) benchmark for both models.
Latin restoration: recover the masked characters an editor once supplied. exact = recovered them; CER = character error rate; ECE = expected calibration error(how far stated confidence is from real accuracy — lower is more honest). 中文对照:exact=完全命中;CER=字符错误率;ECE=期望校准误差(自报把握与真实准确率之差,越低越诚实)。Two distinct benchmarks: single-character masks (clean3_lat) and the harder 2–3-character spans (clean3_lat_v2) — numbers are only comparable within a benchmark.
| Model | Benchmark | n | exact | CER | ECE |
|---|