两方石碑,两种方括号,两个模型。编者的判断与模型的预测,逐处并置。
本页是课程注释幻灯第 2.3 节的实验档案;幻灯 s2-17a 至 s2-17n 各页给出课堂语境与讲者注,页内导航右端可跳转。
本地实跑 2026-08-26 · 门户复测 2026-08-28 · Ithaca 与 Aeneas 本地权重 · beam 100 · 含连书对照八组
两种括号里的字,石头上都没有;补出它们的历来是编者。下面把同样的问题交给模型:希腊碑给 Ithaca,拉丁碑给 Aeneas,每一个数字都来自本次实跑。
先说前提:这两方石碑都在模型的训练语料里,而且编者补的字以明文在内。这一点须在读所有结果之前先行说明。
IG I³ 36 = ML 71 = OR 156 · 雅典娜胜利女神庙女祭司决议(二) · 公元前 424/3 年 · Ithaca
第 1 步 · 认识这块石头
课件第 8 讲第 26 页,图 2(ML 71 = IG I³ 36 = OR 156)
决议规定:司库们每年塔尔盖利昂月,付给雅典娜胜利女神庙的女祭司五十德拉克马。
石头不完整,文本却不长。要紧的在笔迹:第一位刻工刻完前五行和第六行开头,第六行第四格被刮去一个字母之后,换了第二位刻工接手。两人拼法不同:前一位用阿提卡老拼法,不写 η 和 ω;后一位用爱奥尼亚拼法,两个字母都用(Rhodes 2018, 155;OR 156 注)。
一块石头,两只手,一条看得见的分界线。
第 2 步 · 编者印出的文本
ἔδοχσεν τε͂ι βολε͂ι καὶ το͂ι δέμοι· … τε͂ι hιερέαι τε͂ς Ἀθενάας τε͂ς Νίκες πεντήκοντα δραχμὰς τὰς γεγραμμένας ἐν τῆι στήλ[ηι] ἀποδιδόναι τὸς κωλακρ[έτας], οἳ ἂν κωλακρετῶσι το͂ Θ[αργηλιω̑]νος μηνός, τῆι ἱερ[έαι τῆς Ἀθην]αίας τῆς Νίκη[ς]
据 Osborne 与 Rhodes 2017 年校勘本第 156 号(第 339 页,馆藏 PDF 逐页核对)。红色即编者补文。注意后半段的 τῆι στήλ[ηι]、κωλακρετῶσι 已经带着第二位刻工的 η 与 ω。
第 3 步 · 把补文抹掉,问模型
前三问是编者真实补过的位置。后两问不同:那两处字母石上还在,遮住它们,是要看模型在两位刻工各自的区段里会选哪一种拼法。
…γεγραμμενας εν τηι στηλ?ι αποδιδοναι τος κωλακρ???? οι αν κωλακρετωσι το θ???????νος μηνος τ?? ιερεαι…
输入模型的是与训练语料同款的小写无重音文本,问号逐字符标出缺口长度。
第 4 步 · 模型的回答
δραχμας τας γεγραμμενας εν τηι στηληι αποδιδοναι τος κωλακρ???? οι αν κωλακρετωσι το θαργηλιωνος μηνος τηι ιερεαι της αθηνα
τηι στηληι αποδιδοναι τος κωλακρετας οι αν κωλακρετωσι το θ???????νος μηνος τηι ιερεαι της αθηναιας της νικης
αας τες νικες πεντηκοντα δραχμας τας γεγραμμενας εν τηι στηλ?ι αποδιδοναι τος κωλακρετας οι αν κωλακρετωσι το θαργηλιωνος
ευε νεοκλειδες εγραμματευε αγνοδεμος επεστατε καλλιας ειπε τ?? ιερεαι τες αθεναας τες νικες πεντηκοντα δραχμας τας γεγραμμ
οναι τος κωλακρετας οι αν κωλακρετωσι το θαργηλιωνος μηνος τ?? ιερεαι της αθηναιας της νικης
第 5 步 · 裁决
补词方面模型表现很好:司库、月名、石碑,全在候选前列。拼法方面有两处偏差:爱奥尼亚区的三个探针,两个被预测成阿提卡拼法。编者以换手后的拼法作断代证据;这类少数拼法,在语料统计中被多数拼法压低。
它见过这块石头和石上的 ω,仍按整个语料的多数拼法作答。模型的输出反映语料的统计平均;换手拼法这类少数证据,正是统计平均容易压低的信息。
CIL 14.4393 · 奥斯提亚,献给凯撒迪阿杜梅尼亚努斯 · 公元 217/8 年 · Aeneas
第 1 步 · 认识这块石头
课件第 8 讲第 35 页(inv. 19787),第 1、3、6 至 7 行的凿除带清晰可见
奥斯提亚的消防守望队为小凯撒迪阿杜梅尼亚努斯立碑,他是皇帝马克里努斯之子。一年后父子兵败被处死,元老院下令除名,两人的名字被从各地铭文中凿去。这方碑上有四处凿痕:家族名两处、凯撒名一处、皇帝名一处。
照片上的凹带就是凿除的痕迹。编者在双括号里复原的,正是这四处。
第 2 步 · 编者印出的文本
M(arco) [[Opellio]] / Antonino / [[Diadumeniano]] / nobilissimo Caes(ari) / principi iuventutis / Imp(eratoris) Caes(aris) M(arci) [[Opelli]] Severi / [[Macrini]] Pii Felicis Aug(usti) / trib(unicia) potest(ate) co(n)s(ulis) design(ati) / II p(atris) p(atriae) proco(n)s(ulis) filio / Valerio Titaniano / praef(ecto) vig(ilum) em(inentissimo) v(iro) / curante / Flavio Lupo subpraef(ecto)
献给马尔库斯·奥佩利乌斯·安东尼努斯·迪阿杜梅尼亚努斯,至尊贵的凯撒,青年之首;他是皇帝凯撒马尔库斯·奥佩利乌斯·塞维鲁·马克里努斯(虔敬、幸运的奥古斯都,享保民官权力,第二次获指定执政官,国父,代行执政官)之子。时任消防守望队长官为卓越之人瓦莱里乌斯·提塔尼亚努斯;副长官弗拉维乌斯·卢普斯督办。
这里的编者依据与第一章不同:确定这四个名字的主要是官衔与年代,语言证据只起次要作用。谁在 218 年被除名,史料有明确记载。
第 3 步 · 按石面实际残损提问
先逐一遮,再来一次最难的:四处凿痕一起遮住,模型看到的就是除名令执行之后的石面。
m opellio antonino diadumeniano nobilissimo caes principi iuventutis imp caes m opelli severi macrini pii felicis aug trib potest cosulis design ii p p procosulis filio valerio titaniano praef vig em v curante flavio lupo subpraef
第 4 步 · 模型的回答
m ??????? antonino diadumeniano nobilissimo caes principi iuventutis
m opellio antonino ???????????? nobilissimo caes principi iuventutis imp caes m opelli seve
bilissimo caes principi iuventutis imp caes m opelli severi ??????? pii felicis aug trib potest cosulis design ii p p procosuli
第 5 步 · 裁决
对照实验:把从未被凿、就刻在这方碑上的长官名 titaniano 遮住,其余全部保留。
被凿的名字能恢复,原因是除名令并未清除各地所有石刻,这个名字在语料里还剩十六处。titaniano 不能恢复,原因是它在全库只出现这一次,即便本碑就在训练语料内也是如此;模型给出的 plautiano 是另一位更常见的禁卫长官。模型能否恢复一个词,取决于它在语料中的出现次数,与它是否刻在本碑上无关。恢复被凿的名字与替换仅见一次的名字,用的是同一个统计机制。
第 6 步 · 门户复测(2026-08-28)
课堂首轮(2026-08-27)在门户滑杆上限 20 字符下提问,四处凿痕实需 32 字符,前二十个候选全部是错误的名字组合,且无任何提示。复测先核实上限:输入区说明(每次最多恢复 20 个 ? 或 # 字符)、滑杆元件(最小 1,最大 20)、服务器校验(以脚本提交 32,答复为请求未通过校验),三处一致;更长的补文,门户说明指向 Colab 笔记本。
分辑复测:四处凿痕按已知长度拆成两问,每问不超过 20 字符;另两处凿痕以等长横线(缺失、不求补)占位,防止答案进入上下文。
前二十个候选全部为这两个名字的拼写变体(diatumeniano、diadumediano 之类);错误的名字组合不再出现。
其余候选为拼写变体(maurini、opelii、upelli)。平行文本面板第一条仍是本碑(TM 254776 = EDR106390),训练语料是否含本碑由此可见。
给定字符数低于实际缺损时,模型给出的候选全部错误,且没有提示;给定字符数与实际缺损相符时,四个名字全部首选恢复,其余候选只是拼写变体。使用这类门户,应先依石面测量确定缺损的字符数,再决定如何拆分提问;归属与断代即使正确,也不能据以推断补文可靠,两类输出要分开核验。完整记录见课程档案 portal-rerun-20260828.md。
同一块石头,两种写法 · 由课堂上的门户实验引出 · Ithaca,八次实跑
第 1 步 · 问题从课堂来
课堂实验把 IG I³ 36 去掉部分空格输入门户,模型给出了几处希腊语里不可能的答案。这引出一个更准的问题:古典雅典的列刻式铭文是连书,字母在网格上一个挨一个,词与词之间没有空格;而模型的训练语料来自 PHI 的编者转写,词是分开的。清洗管线剥掉了方括号,却留下了空格。
那么模型学会的是石头的写法,还是编者的写法?把同一段文本各写一遍:一遍照语料分词(分书),一遍照石头连书,同样的缺口各问一次。
分书:εν τηι στηλ?? αποδιδοναι τος κωλακρ???? οι αν…
连书:…εντηιστηλ??αποδιδοναιτοςκωλακρ?????ιαν…
第 2 步 · 两列结果
| 缺口 | 分书(语料写法) | 连书(石头写法) |
|---|---|---|
| στηλ[ηι](两格) | ηι 第 1 · 0.81 | ηι 第 2 首选是 η 加空格(0.48):在 στηληι 词内造出一个词界 |
| κωλακρ[έτας] | ετας 第 1 · 0.84 | 五格实为 ΕΤΑΣΟ,编者读法在一百条候选中一条不见 首选是 ετας 加空格(0.10) |
| 连书四格(只求 ετας) | (不适用) | 跌至全 beam 第 34 首选 εται |
| 分书六格(求 ετας 空格 ο) | ετας ο 第 1 · 0.70 连空格一起生成 | (不适用) |
四组分书全部第一,置信度 0.70 至 0.95。同样的缺口换成连书:一处在词内插入词界,一处让编者读法从整个 beam 里消失。连书条件下,空格进入全部四组的前十,两组排第一。
第 3 步 · 被凿的那一格
第六行第四格被古人刮去,网格上留一个空位。这一格问了三次:
模型在被凿的一格上没有固定答案,输出随问法而变:旁边有空格时给字母,顶替空格位置时给空格。它给出的空格是对语料分词层的重建,与石面的留白无关。石面上的空位是刻工的安排,转写中的空格是编者的分词记号;两者在字符串里是同一个字符,模型无法区分。
第 4 步 · 裁决
训练语料去掉了方括号,保留了空格。模型学到的文本既非石面原状,也非完整的校勘本,而是一份半编辑状态的文本。输入编者分好词的文本,预测与编者一致;输入石面样式的连书,它在缺口处优先补出编者的分词记号。词间空格出自编者的分词,模型把它当作普通字符学了进去。第 6 节讲编码要保存文字与编辑层的分界,依据在此。
十例常规方括号对照 · 语料污染核查 · 一个空格的控制案例
语料污染核查(38 例实跑)。Ithaca 与 Aeneas 的训练语料就是 PHI 与 LED,也正是各家校勘本的来源。逐例回查:
全部 38 例首选一致 28 例。污染样本一致率 74%,未污染样本 75%,就这批数据看记忆没有明显抬高成绩;样本仅 8 例干净,不足以下定论。第一版查漏曾把 12 例漏判为干净,改用 PHI 编号权威比对后更正。查不动的要记为未检验,不能记为干净。
SEG IX 8 = IRCyr2020 C.101 (FIRA I2 68; Oliver, Greek Constitutions 8-12); EpiDoc file magalia matrix-hub gove
αικην επαρχηαν υπεξειρημενων των υποδικων κεφαλης υπερ ων ος αν την επαρχηαν διακατεχη αυτος διαγεινωσκειν κ?? ισταναι η συνβουλιον κριτων παρεχειν οφειλει υπερ δε των λοιπων πραγματων παντων ελληνας κριτας διδοσθαι αρεσκει ει μ
编者依据:No apparatus note at this point in the edition. The supplement completes καί at the end of l. 66; grounds are word-completion plus the required coordination of the two infinitives διαγεινώσκειν ... ἱστάναι.
读法:Function word in a fully formulaic legal coordination. Ithaca gives it 0.873 with a clean decay to the runner-up (0.032). The easiest class of bracket: the editor and the model are answering the same low-entropy question.
CIL VI 1527 + 31670 + 37053 = ILS 8393; EpiDoc file governance/laudatio-turiae_epidoc.xml
lepidi consulis collegae praesentis pedibus aduoluta humi non modo non adleuata sed tracta et seruil?? in modum rapsata liuoribus corporis repleta firmissimo animo eum admoneres edicti caesaris cum gratulatione restitu
编者依据:The edition prints the supplement without a dedicated apparatus note at l. 55 (the listApp entry at loc. 30 documents a different, virtue-catalogue slot, restored "on parallel-pool grounds from the other laudationes mulierum (Mantzilas 2017)"). Grounds here are idiomatic: seruilem in modum is a fixed adverbial phrase requiring the accusative singular.
读法:DIVERGENCE. Aeneas ranks the editor 5th at 0.050, preferring seruil[ia] (0.353) and seruil[is] (0.332). Every one of its top four is a morphologically possible Latin ending; what it does not command is the frozen idiom X in modum, which fixes the case. The editor is not out-guessing the model on letters but applying a phraseological constraint.
CIL VIII 11451 (+ 270 + add.); EpiDoc file governance/sc-beguensis_epidoc.xml
icorum lucili africani c larissimi v iri qui petunt ut ei permittatur in provincia afr ica regione beguensi territorio musulami??um ad casas nundinas iiii nonas novemb res et xii k alendas dec embres ex eo omnibus mensibus iiii non as et xii k alendas su
编者依据:No specific apparatus note; the edition supplies the genitive plural of the ethnic Musulamii, the tribe named as the territory-holder in this document.
读法:Proper-noun morphology. Aeneas returns or at 0.910. The name is common enough in African epigraphy that the model has the pattern; agreement here is cheap.
AE 1976, 653 (SEG 26, 1392; EDCS-09300547); EpiDoc file governance/sotidius-strabo-edict_epidoc.xml
cormasa et conanam. neque tamen omnibu s huius rei ius erit sed procuratori principis optimi filioque eius usu dato ??que ad carra decem aut pro singulis carris mulorum trium aut pro singulis mulis asinorum binorum quibus eodem te
编者依据:No apparatus note at this point (the edition's listApp documents other lines, e.g. the engraver's omission at l. 5 and Frisch's correction at l. 7). Grounds: usque ad + accusative governs the following carra decem.
读法:Aeneas agrees at 0.477, but with a genuinely divided field: [ne] 0.233, [is] 0.117, [at] 0.104. The lowest-confidence agreement in the set, and a good illustration that a correct top-1 can still rest on a thin margin.
CIL XI 1420-1421 (ILS 139-140; Ehrenberg & Jones 68-69); EpiDoc file governance/decreta-pisana_epidoc.xml
augusti caesaris patris patriae pontificis maximi tribuniciae potestatis xxv fili auguris consulis designati princip?? iuventutis patroni coloniae nostrae q. d. e. r. f. p. d. e. r. i. c. cum senatus populi romani inter c
编者依据:No dedicated apparatus note. The supplement completes the imperial title princeps iuventutis in the genitive, in agreement with the surrounding titulature of Lucius Caesar.
读法:Titulature is the most predictable material in Latin epigraphy. Aeneas returns is at 0.914 with the runner-up three orders of magnitude down (0.0025). Note the corpus status: the inscription is in the Aeneas training corpus but that corpus does NOT carry these restored letters, so the model is completing the title, not reciting it.
C. B. Welles, Royal Correspondence in the Hellenistic Period (1934), nos. 18-20; EpiDoc file governance/welles
βασιλευς αντιοχος μητροφανει χαιρειν πεπ??καμεν λαοδικηι παννου κωμην και τημ βαριν και την προσουσαν χωραν τηι κωμηι ορος τηι τε ζελειτιδι χωραι και τηι κυζικηνηι και τηι
编者依据:The edition prints πεπ[ρά]καμεν, the perfect of πιπράσκω, "we have sold" - the operative verb of the whole dossier, which is a royal land-sale. The word is split across ll. 1/2 (lb break="no").
读法:DIVERGENCE, and the most instructive one. Ithaca prefers πεπ[οη]καμεν (0.884) over the editor's πεπ[ρά]καμεν (0.776) - a near-tie between two real Greek perfects. πεποιήκαμεν "we have made" is by far the commoner verb in the corpus; πεπράκαμεν is what the document is FOR. The editor is reading the legal act, the model is reading letter frequency. Note also that this text, restoration included, is in Ithaca's training corpus and the model still misses it.
I.Sicily ISic002987 (TM 645339; DOI 10.5281/zenodo.4387873); the file's apparatus states "Text of Manganaro 19
ιλια ταλαντα λοιπον επτα ενενηκοντα λιτραι πεντε ικοσι διακοσια τρισχιλια μυρια ταλαντα ??μιαις εσοδος τεσσαρακοντα λιτραι οκτακισ χιλια δισμυρια ταλαντα εξοδος μια τρια κοντα λ
编者依据:No per-line apparatus note beyond the edition statement. The supplement yields ταμίαις, "to the treasurers", the dative plural that heads the account block.
读法:Ithaca returns τα at 0.912 with the runner-up at 0.0014. A financial document with heavy internal repetition gives the model exactly the redundancy it exploits best.
I.Sicily ISic003610 (TM 645683; DOI 10.5281/zenodo.4388821); the file's apparatus states "Text of Rizzone as a
καρδαμαν χρηστος και αμεμπτος ετων ογδοηκοντα μηνας δε ??
编者依据:The EpiDoc encodes the whole sequence as <num value="10">: the editor reads δέ + [κα] as the single numeral δέκα, split across ll. 5/6 (lb break="no") with an ivy-leaf ornament standing INSIDE the word.
读法:DIVERGENCE, with a control that explains it. As fed, Ithaca ranks the editor 2nd (κα 0.066) behind ημ (0.127). The ivy-leaf ornament canonicalises to a space, so the model saw "μηνας δε ??" and read δέ as a separate particle. Re-fed WITHOUT that space ("μηνας δε??"), Ithaca returns κα at 0.712 - top-1, agreeing with the editor. The editor's advantage here is not linguistic but material: knowing that an ornament can sit inside a word, and that the layout does not mark a word boundary.
I.Sicily ISic030278 (Decree in honour of Nemenios, copy B); the file's apparatus states "Text from autopsy" (c
ανδρειας και καλοκαγαθιας και τας των προγονων αρετας δικαιον δε εστι τους αγαθους των ανδρων και ταν αυτων ε?????? ενδεικνυμενων τιμαν και προεδριαν τυνχανειν και αθανατον αυτων μναμαν παρα τοις ευ πα θοντ
编者依据:No per-line note; the edition rests on autopsy. The supplement gives εὔνοιαν, "goodwill", in the standard honorific collocation εὔνοιαν ἐνδείκνυσθαι.
读法:The longest supplement in the set (6 characters) and still a confident hit: Ithaca returns υνοιαν at 0.726, with υνοιας and υνοιαι further down - the model has the lexeme and is choosing among its endings. Decree formulae are the model's strongest ground.
I.Sicily ISic001516 (TM 645016; PHI 333765; DOI 10.5281/zenodo.4354259); the file's apparatus states "Text of
ιος ουσα ευχα ριστουσα τω ειδιω αν δρι πολλας ευχαρισ τιας α ω ευομει?????
编者依据:No per-line note. The supplement completes the epithet εὐομείλητος, "affable", the last word of the epitaph; the preceding μεί is carried by a ligature (<hi rend="ligature">), which is why the editor can be sure of the stem.
读法:DIVERGENCE, and a total one: the editor's reading is absent from Ithaca's entire top-10, and the model's best guess is the non-word " και " at 0.009 - two orders of magnitude below its confidence on the other cases. The supplement sits at the very END of the text, so there is no right-hand context at all, and εὐομείλητος is a rare compound. When the model has nothing, its own confidence says so; this is the clearest case in the set of a bracket the editor fills from lexical knowledge the model does not have.
控制案例:一个空格改变了答案。I.Sicily ISic003610:编者把 δέ 与补出的 κα 读为一个数词 δέκα(十个月),这个词被一处常春藤叶装饰断在词内又跨了行。规范化把装饰变成空格后,模型把 δέ 当成独立小品词,编者读法只排第二(0.066);去掉那个空格重跑,模型给出 κα,0.712,首选。编者在这里的优势与语言无关:他知道装饰可以立在词中间。所谓模型出错,有多少其实是文本送进模型时的序列化造成的。第 6 节讲编码保住的正是这类信息。
| 模型 | DeepMind predictingthepast:Ithaca(希腊语)、Aeneas(拉丁语),本地权重实跑,beam 100 |
|---|---|
| 两方主碑 | IG I³ 36:括号位置逐一核对 Osborne 与 Rhodes 2017 校勘本(馆藏 PDF 第 339 页);CIL 14.4393:以课件第 8 讲第 35 页所印双括号文本为准,LED 存文(TM 254776 = EDR106390)按其原样输入 |
| 附录十例 | 取自 82 件治理文书 EpiDoc 与 4,823 件 I.Sicily 铭文的编者补文;补文前后 130 字符内无其他编辑成分 |
| 掩码 | 问号逐字符标长度;文本按各模型训练语料同款规范化 |