XIII · Embedding Outliers · 嵌入离群点
↑ Matrix Hub ↑ magalia.wiki
Finding 13 · model-only discovery queue分析十三 · 纯模型发现队列

Embedding Outliers嵌入离群点

Most Greek inscriptions resemble their neighbours: an Attic decree reads like other Attic decrees of its century. This queue lists the texts that don't — the 100 texts (of 18,387 scorable) that sit farthest from their own (region × century) neighbourhood in embedding space. An outlier can be a misdating, a misattributed region, a rare genre, heavy damage — or nothing at all. That is exactly what makes it a discovery queue: a ranked list of places worth a scholar's look, each shown with the neighbours it failed to resemble. 大多数希腊铭文与其邻域相似:一件阿提卡政令读来如同其世纪的其他阿提卡政令。此队列所列,恰是不相似者 —— 在 18,387 件可评分文本中,于嵌入空间最偏离其(地区 × 世纪)邻域的 100 件。离群可能意味着误定年代、误归地区、罕见文类、严重残损 —— 也可能毫无意义。这正是其为发现队列之义:一份值得学者过目之处的排序清单,每条均展示其未能相似之邻居。

model-only Evidence class. Scores come from the in-house joint model's embeddings (phase-16 semantic index, a 30,000-item sample of the Greek record — not the corpus). Every lead is unreviewed. Nothing here asserts an error in any corpus; absence from this list proves nothing. Dates and regions are copied from the index metadata, never inferred. Method and data: datasets · re-run deltas: discovery log · regenerate with scripts/build_embedding_outliers.py. 证据等级。 分数出自自建联合模型之嵌入(phase-16 语义索引,希腊记录之 30,000 件抽样 —— 并非全库)。每条线索均未经审核。此处不断言任何语料有误;不在此列亦不证明无事。年代与地区照录索引元数据,绝不推断。方法与数据:数据集 · 以 scripts/build_embedding_outliers.py 再生成。

The queue队列

Score = 1 − mean cosine similarity to the 10 nearest same-cell neighbours (median across all scored texts: ). Click an id to read the text in its source database via the resolver.分数 = 1 − 与同格 10 个最近邻之平均余弦相似度(全体已评分文本之中位数:)。点击编号可经解析器往源数据库读取全文。