Provenance. §1 = reviewed unit (candidate gold, Watson-grounded). §§2–3 = candidate (agent-panel self-consistency, not human-calibrated). §§4–73 = machine-assisted, unreviewed heuristic exploration. Every agreement figure here is agent-panel self-consistency (κ_agent) — a same-base-model proxy CONTROL, NOT human inter-annotator agreement. No human has annotated: humanCalibrationOnFile = false. On unseen text the hardest tier is weak (macro full-tag κ_agent = 0.57, §§2–3). Interpretive layers are valid-but-heuristic rule-sets — coverage ≠ accuracy, not a digital edition. Editorial statement · Error budget

Editorial Statement (ratio edendi)

Cicero, Divinatio in Q. Caecilium · urn:cts:latinLit:phi0474.phi004 · 73 §§, 6903 tokens.

Base text: cicero-reader phi004.json (Perseus phi0474.phi004-lat2). Tokenizer: deterministic, reproducible (GATE 7). Tagset: frozen coarse v1.2 (GATE 5).

Interpretive layers. valid-but-heuristic: the semantic/pragmatic/rhetoric/move layers are transparent RULE-SETS, not learned models (GATE 6). Coverage measures analysis, not accuracy.

Constitutio scope. §1 = reviewed unit (candidate gold + 26 Watson-model corrections + §§1-3 agent-panel self-consistency pass). §§2-3 = candidate (agent-panel self-consistency consensus). §§4-73 = machine-assisted, unreviewed heuristic exploration.

Agreement caveat (load-bearing) — agent-panel self-consistency, NOT human IAA. Cross-model label agreement produced by a panel of LLM specialist agents that SHARED the Watson commentary and one base model. This is agent-panel self-consistency, a same-base-model PROXY, NOT inter-annotator agreement or reliability (those terms are reserved for humans; Artstein & Poesio 2008). humanCalibrationOnFile=false. A human / Watson-blind calibration is REQUIRED before any 'edition' claim. See CALIBRATION-PROTOCOL.md and the turn-key human kit in gate5_iaa/human_kit/.

Human calibration on file: False — until true, no section may be branded an edition.

Sigla

multidimQualitysiglumrespcertgloss
seeded-gold[G]#human-source-alignedhighgold, source-aligned (§1)
watson-reviewed[R]#panel-watsonhighreviewed against Watson 2025
watson-enriched[E]#panel-watsonmediumenriched, Watson-grounded
watson-contested[C]#panel-watsonlowcontested — editors disagree
seeded[s]#machine-seedlowmachine seed, §1 only
heuristic#machineunknownmachine-assisted, unreviewed
watson-fullfix[F]#machine-rulelowdeterministic corpus rule-fix (not human gold)

Defensible citation

Divinatio Wave Lab (build ): a reviewed proof-of-concept annotation of Cicero, Divinatio in Q. Caecilium §1, with a transparent machine-assisted heuristic scaffold and an open review queue over §§2-73. A methods/tooling/data contribution, NOT a digital critical edition.

Cite this (the brand travels with the quotation)

§1 (reviewed unit)
Divinatio Wave Lab (build unstamped). Cicero, Divinatio in Q. Caecilium §1 — reviewed unit (candidate gold; 26 Watson-model corrections; agent-panel self-consistency — six LLM specialists sharing one base model + the Watson commentary; a same-base-model proxy, NOT human IAA/reliability). Accessed [date].
§§2-3 (candidate)
Divinatio Wave Lab (build unstamped). Cicero, Div. Caec. §§2-3 — candidate annotation; agent-panel self-consistency consensus (model concordance, NOT human IAA; not human-calibrated). NOT a critical edition. Accessed [date].
§§4-73 (machine-assisted exploration)
Divinatio Wave Lab (build unstamped). Cicero, Div. Caec. §§N — machine-assisted, unreviewed heuristic exploration; transparent rule-sets, no human review. NOT a critical edition. Accessed [date].

NOT defensible

Versioned by build_provenance.json. Gates applied: 0 honesty, 1 rhetoric-provenance, 2 punctuation, 3 scope-labels, 4 pipeline+orchestrator, 5 self-consistency §§1-3, 6 layer-validity, 7 scaling, 8 publishability.

Agent-panel self-consistency by tier (κ_agent) — NOT human IAA

How much six LLM specialists (one base model + shared Watson commentary) agreed with each other. Shared cause → correlated errors → an UPPER-ish-bound proxy, not reliability. Human agreement: not yet measured (humanCalibrationOnFile = false). Data-driven from iaa_report.json / iaa_s2s3_report.json / iaa_compare.json; content-tokens only.

slice / tagsettierfull-tag κ_agentbase-label κ_agentnone-line reading
§1 round-1
v1.1
macro0.7540.80357In-sample pilot; small full/base gap = minor boundary noise.
§1 round-1
v1.1
micro0.389 WEAK0.90551Category was never in dispute (base 0.905); the collapse is purely the free B/I/C span prefix — a definitional artifact, not disagreement on hard judgments.
§1 round-1
v1.1
anchor0.6290.62957Full-content κ; STRICT anchored-only κ = 0.491 over n=27 anchor-bearing tokens — report the strict ~0.49.
§1 round-2
v1.1 freeze
macro0.8530.85357Full==base by construction once the B/I/C prefix is rule-derived. IN-SAMPLE (tagset tuned on §1) — not evidence of reproducibility.
§1 round-2
v1.1 freeze
micro0.9740.993510.389→0.974 is the deterministic prefix rule removing an annotator degree of freedom — DEFINITIONAL, not a gain on harder judgments. (0.974 is FULL-tag; base 0.993.)
§1 round-2
v1.1 freeze
anchor0.6830.68357Hardest tier; interpretive variance persists even in-sample.
§§2-3 pooled
v1.2
macro0.57 WEAK OUT-OF-SAMPLE0.921158THE UNDISCLOSED WEAK TIER. The full/base gap that 'closed by construction' on §1 REOPENS on unseen text → genuine clause-boundary disagreement, not a prefix artifact.
§§2-3 pooled
v1.2
micro0.897 OUT-OF-SAMPLE0.924145Fine-grained category self-consistency HOLDS out-of-sample.
§§2-3 pooled
v1.2
anchor0.721 OUT-OF-SAMPLE0.721152Full-content band ~0.63-0.74; strict anchor-bearing ~0.49. Irreducible interpretive variance (abl-abs head, periphrastic esse) → MANDATORY human adjudication.
human–human (all tiers)PENDING — no human annotation on file. This hole is shown as a hole, never filled by a model.

Reading. The deployed 0.85/0.97 are §1, in-sample, after a tagset re-freeze. On unseen text the macro FULL-tag self-consistency falls to κ_agent 0.57 — genuine boundary disagreement, not a definitional artifact. Micro was 0.389 before the prefix freeze. These are agreements among LLMs, not humans.

You MAY cite
§1 as a reviewed proof-of-concept annotation; the tool + method + data; agent-panel self-consistency figures labeled as such.
You MAY NOT cite
any section as a digital edition; these κ as human IAA/reliability; coverage as accuracy; §§2-73 text as settled.

A real human / Watson-blind calibration is prepared and waiting: gate5_iaa/human_kit/RUNBOOK.md. Build unstamped.