Coverage measures ANALYSIS, not accuracy. Structural defaults (EVENT_ANCHOR/PARTICIPANT role scaffold, PREDICATION frame, given/topic info, unscored rhetoric) are excluded from coverage.
| layer | honest coverage | note |
|---|---|---|
| thematic role | 0.55% (gross incl. EVENT_ANCHOR scaffold: 28.62%) | honest = real thematic roles (agent/patient/...); gross counts EVENT_ANCHOR/PARTICIPANT structural scaffold |
| event frame | 11.77% | excludes PREDICATION/CONTEXT/EVENT filler |
| information structure | 14.3% | excludes given + definitional clause-head topic |
| rhetoric | 14.92% | categorical priors, NOT measured accuracy |
1834 open review items of 1873 total ({'already-reviewed': 39, 'review-needed': 1834}). This is the published error budget the project owes the reader.
| § | tier | content | role | frame | rhetoric | queue open |
|---|---|---|---|---|---|---|
| §1 | reviewed-unit | 57 | 32 | 34 | 57 | 18 |
| §2 | candidate-agreement | 82 | 0 | 7 | 15 | 37 |
| §3 | candidate-agreement | 76 | 0 | 15 | 8 | 21 |
| §4 | exploration | 91 | 0 | 5 | 20 | 44 |
| §5 | exploration | 85 | 0 | 17 | 18 | 34 |
| §6 | exploration | 82 | 0 | 0 | 17 | 37 |
| §7 | exploration | 58 | 0 | 0 | 7 | 19 |
| §8 | exploration | 79 | 0 | 15 | 11 | 25 |
| §9 | exploration | 49 | 0 | 1 | 4 | 13 |
| §10 | exploration | 70 | 0 | 5 | 8 | 18 |
| §11 | exploration | 99 | 0 | 20 | 21 | 39 |
| §12 | exploration | 71 | 0 | 15 | 14 | 31 |
Status computed by compute_edition_status.py — data-driven, cannot be hand-set.
How much six LLM specialists (one base model + shared Watson commentary) agreed with each other. Shared cause → correlated errors → an UPPER-ish-bound proxy, not reliability. Human agreement: not yet measured (humanCalibrationOnFile = false). Data-driven from iaa_report.json / iaa_s2s3_report.json / iaa_compare.json; content-tokens only.
| slice / tagset | tier | full-tag κ_agent | base-label κ_agent | n | one-line reading |
|---|---|---|---|---|---|
| §1 round-1 v1.1 | macro | 0.754 | 0.803 | 57 | In-sample pilot; small full/base gap = minor boundary noise. |
| §1 round-1 v1.1 | micro | 0.389 WEAK | 0.905 | 51 | Category was never in dispute (base 0.905); the collapse is purely the free B/I/C span prefix — a definitional artifact, not disagreement on hard judgments. |
| §1 round-1 v1.1 | anchor | 0.629 | 0.629 | 57 | Full-content κ; STRICT anchored-only κ = 0.491 over n=27 anchor-bearing tokens — report the strict ~0.49. |
| §1 round-2 v1.1 freeze | macro | 0.853 | 0.853 | 57 | Full==base by construction once the B/I/C prefix is rule-derived. IN-SAMPLE (tagset tuned on §1) — not evidence of reproducibility. |
| §1 round-2 v1.1 freeze | micro | 0.974 | 0.993 | 51 | 0.389→0.974 is the deterministic prefix rule removing an annotator degree of freedom — DEFINITIONAL, not a gain on harder judgments. (0.974 is FULL-tag; base 0.993.) |
| §1 round-2 v1.1 freeze | anchor | 0.683 | 0.683 | 57 | Hardest tier; interpretive variance persists even in-sample. |
| §§2-3 pooled v1.2 | macro | 0.57 WEAK OUT-OF-SAMPLE | 0.921 | 158 | THE UNDISCLOSED WEAK TIER. The full/base gap that 'closed by construction' on §1 REOPENS on unseen text → genuine clause-boundary disagreement, not a prefix artifact. |
| §§2-3 pooled v1.2 | micro | 0.897 OUT-OF-SAMPLE | 0.924 | 145 | Fine-grained category self-consistency HOLDS out-of-sample. |
| §§2-3 pooled v1.2 | anchor | 0.721 OUT-OF-SAMPLE | 0.721 | 152 | Full-content band ~0.63-0.74; strict anchor-bearing ~0.49. Irreducible interpretive variance (abl-abs head, periphrastic esse) → MANDATORY human adjudication. |
| — | human–human (all tiers) | PENDING — no human annotation on file. This hole is shown as a hole, never filled by a model. | |||
Reading. The deployed 0.85/0.97 are §1, in-sample, after a tagset re-freeze. On unseen text the macro FULL-tag self-consistency falls to κ_agent 0.57 — genuine boundary disagreement, not a definitional artifact. Micro was 0.389 before the prefix freeze. These are agreements among LLMs, not humans.
A real human / Watson-blind calibration is prepared and waiting: gate5_iaa/human_kit/RUNBOOK.md. Build unstamped.