Provenance. §1 = reviewed unit (candidate gold, Watson-grounded). §§2–3 = candidate (agent-panel self-consistency, not human-calibrated). §§4–73 = machine-assisted, unreviewed heuristic exploration. Every agreement figure here is agent-panel self-consistency (κ_agent) — a same-base-model proxy CONTROL, NOT human inter-annotator agreement. No human has annotated: humanCalibrationOnFile = false. On unseen text the hardest tier is weak (macro full-tag κ_agent = 0.57, §§2–3). Interpretive layers are valid-but-heuristic rule-sets — coverage ≠ accuracy, not a digital edition. Editorial statement · Error budget

Honest Error Budget

Coverage measures ANALYSIS, not accuracy. Structural defaults (EVENT_ANCHOR/PARTICIPANT role scaffold, PREDICATION frame, given/topic info, unscored rhetoric) are excluded from coverage.

Per-layer coverage (content tokens = 5818)

layerhonest coveragenote
thematic role0.55% (gross incl. EVENT_ANCHOR scaffold: 28.62%)honest = real thematic roles (agent/patient/...); gross counts EVENT_ANCHOR/PARTICIPANT structural scaffold
event frame11.77%excludes PREDICATION/CONTEXT/EVENT filler
information structure14.3%excludes given + definitional clause-head topic
rhetoric14.92%categorical priors, NOT measured accuracy

Open review queue (the deficit)

1834 open review items of 1873 total ({'already-reviewed': 39, 'review-needed': 1834}). This is the published error budget the project owes the reader.

Section status (first 12; full in edition_status.json)

§tiercontentroleframerhetoricqueue open
§1reviewed-unit5732345718
§2candidate-agreement82071537
§3candidate-agreement76015821
§4exploration91052044
§5exploration850171834
§6exploration82001737
§7exploration5800719
§8exploration790151125
§9exploration4901413
§10exploration7005818
§11exploration990202139
§12exploration710151431

Status computed by compute_edition_status.py — data-driven, cannot be hand-set.

Agent-panel self-consistency by tier (κ_agent) — NOT human IAA

How much six LLM specialists (one base model + shared Watson commentary) agreed with each other. Shared cause → correlated errors → an UPPER-ish-bound proxy, not reliability. Human agreement: not yet measured (humanCalibrationOnFile = false). Data-driven from iaa_report.json / iaa_s2s3_report.json / iaa_compare.json; content-tokens only.

slice / tagsettierfull-tag κ_agentbase-label κ_agentnone-line reading
§1 round-1
v1.1
macro0.7540.80357In-sample pilot; small full/base gap = minor boundary noise.
§1 round-1
v1.1
micro0.389 WEAK0.90551Category was never in dispute (base 0.905); the collapse is purely the free B/I/C span prefix — a definitional artifact, not disagreement on hard judgments.
§1 round-1
v1.1
anchor0.6290.62957Full-content κ; STRICT anchored-only κ = 0.491 over n=27 anchor-bearing tokens — report the strict ~0.49.
§1 round-2
v1.1 freeze
macro0.8530.85357Full==base by construction once the B/I/C prefix is rule-derived. IN-SAMPLE (tagset tuned on §1) — not evidence of reproducibility.
§1 round-2
v1.1 freeze
micro0.9740.993510.389→0.974 is the deterministic prefix rule removing an annotator degree of freedom — DEFINITIONAL, not a gain on harder judgments. (0.974 is FULL-tag; base 0.993.)
§1 round-2
v1.1 freeze
anchor0.6830.68357Hardest tier; interpretive variance persists even in-sample.
§§2-3 pooled
v1.2
macro0.57 WEAK OUT-OF-SAMPLE0.921158THE UNDISCLOSED WEAK TIER. The full/base gap that 'closed by construction' on §1 REOPENS on unseen text → genuine clause-boundary disagreement, not a prefix artifact.
§§2-3 pooled
v1.2
micro0.897 OUT-OF-SAMPLE0.924145Fine-grained category self-consistency HOLDS out-of-sample.
§§2-3 pooled
v1.2
anchor0.721 OUT-OF-SAMPLE0.721152Full-content band ~0.63-0.74; strict anchor-bearing ~0.49. Irreducible interpretive variance (abl-abs head, periphrastic esse) → MANDATORY human adjudication.
human–human (all tiers)PENDING — no human annotation on file. This hole is shown as a hole, never filled by a model.

Reading. The deployed 0.85/0.97 are §1, in-sample, after a tagset re-freeze. On unseen text the macro FULL-tag self-consistency falls to κ_agent 0.57 — genuine boundary disagreement, not a definitional artifact. Micro was 0.389 before the prefix freeze. These are agreements among LLMs, not humans.

You MAY cite
§1 as a reviewed proof-of-concept annotation; the tool + method + data; agent-panel self-consistency figures labeled as such.
You MAY NOT cite
any section as a digital edition; these κ as human IAA/reliability; coverage as accuracy; §§2-73 text as settled.

A real human / Watson-blind calibration is prepared and waiting: gate5_iaa/human_kit/RUNBOOK.md. Build unstamped.