Cicero, Divinatio in Q. Caecilium · urn:cts:latinLit:phi0474.phi004 · 73 §§, 6903 tokens.
Base text: cicero-reader phi004.json (Perseus phi0474.phi004-lat2). Tokenizer: deterministic, reproducible (GATE 7). Tagset: frozen coarse v1.2 (GATE 5).
Interpretive layers. valid-but-heuristic: the semantic/pragmatic/rhetoric/move layers are transparent RULE-SETS, not learned models (GATE 6). Coverage measures analysis, not accuracy.
Constitutio scope. §1 = reviewed unit (candidate gold + 26 Watson-model corrections + §§1-3 agent-panel self-consistency pass). §§2-3 = candidate (agent-panel self-consistency consensus). §§4-73 = machine-assisted, unreviewed heuristic exploration.
Agreement caveat (load-bearing) — agent-panel self-consistency, NOT human IAA. Cross-model label agreement produced by a panel of LLM specialist agents that SHARED the Watson commentary and one base model. This is agent-panel self-consistency, a same-base-model PROXY, NOT inter-annotator agreement or reliability (those terms are reserved for humans; Artstein & Poesio 2008). humanCalibrationOnFile=false. A human / Watson-blind calibration is REQUIRED before any 'edition' claim. See CALIBRATION-PROTOCOL.md and the turn-key human kit in gate5_iaa/human_kit/.
Human calibration on file: False — until true, no section may be branded an edition.
| multidimQuality | siglum | resp | cert | gloss |
|---|---|---|---|---|
| seeded-gold | [G] | #human-source-aligned | high | gold, source-aligned (§1) |
| watson-reviewed | [R] | #panel-watson | high | reviewed against Watson 2025 |
| watson-enriched | [E] | #panel-watson | medium | enriched, Watson-grounded |
| watson-contested | [C] | #panel-watson | low | contested — editors disagree |
| seeded | [s] | #machine-seed | low | machine seed, §1 only |
| heuristic | ∇ | #machine | unknown | machine-assisted, unreviewed |
| watson-fullfix | [F] | #machine-rule | low | deterministic corpus rule-fix (not human gold) |
Divinatio Wave Lab (build
Divinatio Wave Lab (build unstamped). Cicero, Divinatio in Q. Caecilium §1 — reviewed unit (candidate gold; 26 Watson-model corrections; agent-panel self-consistency — six LLM specialists sharing one base model + the Watson commentary; a same-base-model proxy, NOT human IAA/reliability). Accessed [date].
Divinatio Wave Lab (build unstamped). Cicero, Div. Caec. §§2-3 — candidate annotation; agent-panel self-consistency consensus (model concordance, NOT human IAA; not human-calibrated). NOT a critical edition. Accessed [date].
Divinatio Wave Lab (build unstamped). Cicero, Div. Caec. §§N — machine-assisted, unreviewed heuristic exploration; transparent rule-sets, no human review. NOT a critical edition. Accessed [date].
Versioned by build_provenance.json. Gates applied: 0 honesty, 1 rhetoric-provenance, 2 punctuation, 3 scope-labels, 4 pipeline+orchestrator, 5 self-consistency §§1-3, 6 layer-validity, 7 scaling, 8 publishability.
How much six LLM specialists (one base model + shared Watson commentary) agreed with each other. Shared cause → correlated errors → an UPPER-ish-bound proxy, not reliability. Human agreement: not yet measured (humanCalibrationOnFile = false). Data-driven from iaa_report.json / iaa_s2s3_report.json / iaa_compare.json; content-tokens only.
| slice / tagset | tier | full-tag κ_agent | base-label κ_agent | n | one-line reading |
|---|---|---|---|---|---|
| §1 round-1 v1.1 | macro | 0.754 | 0.803 | 57 | In-sample pilot; small full/base gap = minor boundary noise. |
| §1 round-1 v1.1 | micro | 0.389 WEAK | 0.905 | 51 | Category was never in dispute (base 0.905); the collapse is purely the free B/I/C span prefix — a definitional artifact, not disagreement on hard judgments. |
| §1 round-1 v1.1 | anchor | 0.629 | 0.629 | 57 | Full-content κ; STRICT anchored-only κ = 0.491 over n=27 anchor-bearing tokens — report the strict ~0.49. |
| §1 round-2 v1.1 freeze | macro | 0.853 | 0.853 | 57 | Full==base by construction once the B/I/C prefix is rule-derived. IN-SAMPLE (tagset tuned on §1) — not evidence of reproducibility. |
| §1 round-2 v1.1 freeze | micro | 0.974 | 0.993 | 51 | 0.389→0.974 is the deterministic prefix rule removing an annotator degree of freedom — DEFINITIONAL, not a gain on harder judgments. (0.974 is FULL-tag; base 0.993.) |
| §1 round-2 v1.1 freeze | anchor | 0.683 | 0.683 | 57 | Hardest tier; interpretive variance persists even in-sample. |
| §§2-3 pooled v1.2 | macro | 0.57 WEAK OUT-OF-SAMPLE | 0.921 | 158 | THE UNDISCLOSED WEAK TIER. The full/base gap that 'closed by construction' on §1 REOPENS on unseen text → genuine clause-boundary disagreement, not a prefix artifact. |
| §§2-3 pooled v1.2 | micro | 0.897 OUT-OF-SAMPLE | 0.924 | 145 | Fine-grained category self-consistency HOLDS out-of-sample. |
| §§2-3 pooled v1.2 | anchor | 0.721 OUT-OF-SAMPLE | 0.721 | 152 | Full-content band ~0.63-0.74; strict anchor-bearing ~0.49. Irreducible interpretive variance (abl-abs head, periphrastic esse) → MANDATORY human adjudication. |
| — | human–human (all tiers) | PENDING — no human annotation on file. This hole is shown as a hole, never filled by a model. | |||
Reading. The deployed 0.85/0.97 are §1, in-sample, after a tagset re-freeze. On unseen text the macro FULL-tag self-consistency falls to κ_agent 0.57 — genuine boundary disagreement, not a definitional artifact. Micro was 0.389 before the prefix freeze. These are agreements among LLMs, not humans.
A real human / Watson-blind calibration is prepared and waiting: gate5_iaa/human_kit/RUNBOOK.md. Build unstamped.