
settled Beinecke MS 408 is a genuine early-15th-century codex: vellum radiocarbon-dated 1404–1438, iron-gall ink, period pigments, five scribal hands, provenance from Prague c.1600 (section 09). Modern forgery is dead.
excluded Not a plaintext, a one-to-one substitution, a letter transposition, an abjad rendering or a null-padded rendering of any of eleven European languages in ordinary word segmentation (sections 03, 04 and 05). Conditional character entropy sits at least 1.0 bit below the natural-language floor in raw EVA and at least 0.4 bit below it under three independent glyph segmentations.
live A verbose or slot-template encoding, one plaintext unit becoming a rigid two- or three-glyph group, against slot-preserving meaningless generation by copy-and-modify (sections 06 and 08). Every statistic that separates Voynichese from language is reproduced by a member of each class. Judgement*: roughly even, about 45:50, the remainder on an unenciphered exotic language.
01 The question and the method
This is not a decipherment. The question was whether the text is meaningless, enciphered or a language, and how far public data can push that in one session. Five literature sweeps, one per hypothesis or evidence class, produced SOURCES.md. Six computational tests, a1 to a6, ran on the ZL transliteration against natural-language corpora and implemented hoax generators. A separate agent re-ran every test adversarially (v_a1 to v_a6) and corrected several; corrected figures are quoted. Two baseline scripts written before the agents ran reproduce the a1 and a2 headlines: h2 of 2.32 bits in raw EVA without spaces, Zipf slope −1.04, adjacent repeats 0.81%.
(S) marks a number computed this session by a named script; (L) one quoted from the literature, with source and year. Judgement carries an asterisk.
Data (S). ZL3b-n, the Zandbergen–Landini EVA transliteration (voynich.nu, 13 May 2025), paragraph text only: 34,116 tokens after dropping words with unread or extended-EVA glyphs, about 10,750 in Currier A and 22,900 in B, 7,236 word types. Cross-checked on RF1b-er, IT2a, GC2a (v101) and CD2a (Currier). The 2,463 uncertain spaces were treated as breaks; joining them changes nothing material (v_a2). Language controls are single Gutenberg books and an OCR’d 16th-century Italian corpus. Entropies are plug-in; Miller–Madow moves h2 by at most 0.007.
02 Strongest counterargument first
Every character statistic assumes the transliteration, and that EVA spaces mark plaintext units. Latin re-segmented after every e, s and m moves 60% of the way to Voynichese slot rigidity, and Latin cut into vowel-final syllables is more rigid than Voynichese at equal length (S: v_a3). The exclusion therefore covers word-boundary-preserving encodings of eleven languages in modern orthography. Rule-based re-segmentation, or a language with rigid consonant-vowel syllables written syllabically, was not tested.
Second, the verbose control that matches raw EVA on h1, h2 and their ratio at once produces words 2.3 times too long, 11.52 symbols against 5.07, while carrying Latin word entropy (S: v_a1). That refutes a letter-level verbose cipher of whole Latin words. Both points narrow the verdict. Neither overturns it.
03 Character level
The firmest result is conditional character entropy h2, the uncertainty of a glyph given the one before it. Bennett (1976) put Voynichese at 2.22 bits against 3.01 to 3.37 for European languages (L); Lindemann & Bowern (2021) found it below all 250 languages in a Wikipedia corpus (L). The gap survives glyph merging and the Currier and v101 alphabets. All rows (S), no spaces, a1, v_a1 and v_a5.
| text | h1 | h2 | h1−h2 | h2/h1 | mean word length |
|---|---|---|---|---|---|
| ZL raw EVA, Currier A | 3.832 | 2.353 | 1.479 | 0.614 | 4.98 |
| ZL raw EVA, Currier B | 3.859 | 2.182 | 1.677 | 0.565 | 5.12 |
| ZL raw EVA, all | 3.865 | 2.311 | 1.554 | 0.598 | 5.07 |
| ZL merged (benches, i/e groups, aiin, qo) | 4.039 | 2.870 | 1.170 | 0.710 | 3.7–3.9 |
| GC v101 (68 symbols) | 4.144 | 2.893 | 1.25 | 0.698 | 3.7–4.0 |
| Currier CD2a (35 symbols) | 3.857 | 2.660 | 1.197 | 0.690 | |
| Latin, Aeneid | 4.025 | 3.510 | 0.515 | 0.872 | 5.77 |
| German (diacritics folded 3.335) | 4.160 | 3.391 | 0.769 | 0.815 | 4.88 |
| Range of 11 languages (floor 3.32 folded) | 3.94–4.52 | 3.32–3.88 | 0.49–0.77 | 0.82–0.88 | 4.2–6.6 |
| Latin, vowels removed | 3.659 | 3.515 | 0.144 | 0.961 | 3.11 |
| Verbose cipher of Latin, every letter 2–3 symbols, 24 symbols | 3.901 | 2.325 | 1.576 | 0.596 | 11.52 |
| Verbose cipher, 8 letters kept single | 4.102 | 2.916 | 1.186 | 0.711 | 7.64 |
Raw EVA gives 2.18 to 2.35 bits; ZL merged, v101 and Currier give 2.66 to 2.89; the eleven-language floor is 3.32 with diacritics folded. Any transposition of Latin gives at least 3.44. Vowel deletion leaves Latin at 3.52 with h2/h1 of 0.961, the wrong direction for an abjad.
Where the redundancy sits. Inside words. Within-word h2 is 1.92 to 2.13 bits raw against 2.94 to 3.41 for languages. Mutual information between glyph and word position is 0.69 to 0.73 bits across ZL, Currier and v101 against 0.17 to 0.27; 21% to 31% of glyphs are locked to one word position against 0% to 8% of letters (S: a5, v_a5).
Pair merging. Greedy merging of frequent glyph pairs closes the h1−h2 gap only after about 40 merges, so an encoding of two- to three-glyph units drawn from 50 to 60 unit types is compatible with the character statistics and a plain letter cipher is not (S: a5). Greedy merging tests one heuristic, not the whole verbose class (S: v_a5).
Information budget. The word stream carries 425,468 bits, 12.47 bits per token, 0.82 times a Latin prose text of equal length under the same model (S: a5, v_a5).
04 Word level
At word level Voynichese looks like a language, as Landini (2001) and Reddy & Knight (2011) reported (L). Table (S), a2 and v_a2, 34,116 tokens.
| corpus | Zipf slope, ranks 1–1000 | hapax % of types | Heaps beta | top-word share |
|---|---|---|---|---|
| Voynich all | −1.040 | 69.7 | 0.710 | 2.25% |
| Voynich A / B | −0.994 / −1.068 | 72.2 / 68.6 | 0.743 / 0.679 | 4.27% / 2.13% |
| Latin (Aeneid) | −0.820 | 67.0 | 0.749 | 4.9% |
| German | −1.045 | 65.1 | 0.779 | 3.46% |
| English | −1.056 | 55.1 | 0.675 | 4.68% |
Word-type entropy is 9.8 to 10.3 bits against 10.3 to 10.8 for Latin prose and 9.4 to 9.8 for Italian, French and German (S: v_a1). The low pooled top-word share is an A+B mixture effect (S: v_a2); against single-book samples the hapax rate is above every language (S: v_a5). None of this discriminates: Zipf and hapax are reproduced by human gibberish (Gaskell & Bowern 2022), by generators (Rugg & Taylor 2016; Timm & Schinner 2020) and by a hand cipher from Latin (Greshko 2025) (L).
Word length
Merged EVA word lengths (S: a2, v_a2):
| corpus | mean | variance | variance / mean of (length−1) |
|---|---|---|---|
| Voynich all (raw EVA 5.07 / 3.73 / 0.92) | 4.11 | 2.48 | 0.80 |
| Voynich A | 3.93 | 2.89 | 0.99 |
| Voynich B | 4.20 | 2.27 | 0.71 |
| Aeneid alone / Confessions alone | 4.97 / | 1.04 / 1.74 | |
| Italian / German / English / Danish | 4.53 / 4.99 / 4.23 / 4.67 | 6.52 / 7.04 / 5.41 / 7.40 | 1.85 / 1.77 / 1.67 / 2.02 |
The famous under-dispersion, “binomial word lengths”, is a Currier-B property. B has variance-to-mean 0.71, A has 0.99 and single-text Latin verse 1.04, so Currier A matches Latin verse and the binomial-fit contrast in a2 was a search-cap artefact. B and the pooled text stay far from the 1.67 to 2.02 of modern prose.
05 Word grammar and repetition
Tiltman (1951), Stolfi (2000) and Zattera (2022) described a rigid positional grammar inside Voynichese words (L). Measured here as the fraction of glyph pairs whose order violates the single best glyph ordering, length-matched; a shuffle null gives 0.49. All rows (S): a3, a2, v_a2.
| statistic | Voynich | Latin | Italian | German | English |
|---|---|---|---|---|---|
| Pair-order violations against the best glyph order, length-matched (shuffle null 0.49) | 0.160 (v101 0.140) | 0.363 | 0.318 | 0.278 | 0.308 |
| Adjacent identical tokens, % and observed:shuffled | 0.907 / 2.5 | 0.047 / 0.1 | 0.047 / 0.1 | 0.157 / 0.3 | 0.137 / 0.2 |
| Line-initial first-glyph effect, Cramér’s V, paragraph-first lines excluded | 0.277 | pseudo-lined prose 0.02–0.03; hexameter 0.148 | |||
Voynichese is 1.7 to 2.3 times more rigid than any tested language. The glyph order derived from the data with no assumptions, q, then p f sh, then ch, then o t k, then e ee, then d s a, then l r y, then g n iin m, is the Stolfi and Zandbergen layout recovered blind (S: a3).
Line position matters, as Currier (1976) said (L). The line-initial effect exceeds hexameter verse and is an order of magnitude above pseudo-lined prose; the line-final effect a2 reported collapses to V 0.07 to 0.16 once the final m and g allographs are removed (S: v_a2).
06 Self-citation and the generator bake-off
Timm (2014, 2016) argued that similar words cluster on the page, the fingerprint of copying an earlier word and modifying it; Timm & Schinner (2020) published an algorithm that does this (L). The test finds, for each token, the nearest earlier token on the same page within edit distance 1, with every corpus poured into the identical page and line skeleton. Two Rugg-style generators and a re-implemented autocopist ran through the same code. All rows (S), a4 and v_a4.
| corpus | median distance | within 10 tokens | within-page shuffle | global shuffle | long tokens (5+) within 10 |
|---|---|---|---|---|---|
| Voynich ZL | 14 | 0.269 | 0.238 | 0.167 | 0.225 |
| Latin (Aeneid) | 41 | 0.035 | 0.034 | 0.030 | 0.006 |
| English (Italian, German similar) | 17 | 0.176 | 0.177 | 0.164 | 0.019 |
| Danish (verse) | 15 | 0.196 | 0.172 | 0.138 | 0.035 |
| Autocopist (Timm–Schinner re-implementation) | 13 | 0.329 | 0.275 | 0.079 | 0.257 |
| Rugg grille | 36 | 0.067 | 0.249 | 0.066 | 0.033 |
| Rugg sorted table | 34 | 0.329 | 0.162 | 0.112 | 0.341 |
| Order-3 character Markov model trained on Voynichese | 19 | 0.194 | 0.194 | 0.194 | 0.119 |
Sequential excess, real minus within-page shuffle, measures copying order: Voynich +0.031, Danish verse +0.024, autocopist +0.054, prose about zero. Page-vocabulary excess, real minus global shuffle: Voynich +0.102, languages at most +0.058, autocopist +0.250. For tokens of five or more glyphs the language baselines collapse to 0.035 or less while Voynichese stays at 0.225, but an order-3 character Markov model that copies nothing reaches 0.119: about half the long-token density is word-internal structure, not copying (S: v_a4).
The bake-off
| text | types | h2 raw | h2 merged | Zipf slope | hapax % | adjacent repeat | pair-violation |
|---|---|---|---|---|---|---|---|
| Voynich ZL | 7,236 | 2.311 | 2.646 | −1.040 | 69.7 | 0.0082 | 0.160 |
| Autocopist, two seeds | 11,660 / 11,330 | 2.709 / 2.723 | 3.274 / 3.005 | −0.825 / −0.859 | 72.2 / 71.1 | 0.0013 / 0.0018 | free-edit 0.41–0.46 |
| Rugg grille / sorted table | 3,747 / 1,757 | 2.798 / 2.629 | 3.176 / 2.912 | −0.429 / −0.622 | 5.7 / 13.9 | 0.0003 / 0.0893 | fitted table 0.122 |
No generator matches on every column. The autocopist yields only 37% real Voynich word types by token, 2% qo-initial words against 15%, and h2 0.4 to 0.6 bits too high; it over-produces locality and page vocabulary where the manuscript’s own sequential excess is only the size of Danish verse. It is a re-implementation with documented deviations, not Timm’s reference program. The grilles fail Zipf and hapax outright. The one table that reproduces the slot rigidity had its columns read off Voynichese: sufficiency only (S: v_a3).
07 Currier A and B
Currier (1976) found two statistical languages in the manuscript (L). What separates them (S: a6, v_a6):
| statistic | A | B |
|---|---|---|
| Tokens ending -edy / containing ed | 0.20% / 0.29% | 17.34% / 20.84% |
| Containing cho / starting qo- / ending -ol | 16.1 / 10.0 / 14.5% | 2.7 / 17.7 / 7.1% |
| Glyph after ch: e / o | 24.5 / 46.1% | 60.1 / 10.1% |
| daiin per thousand tokens | 42.5 | 13.0 |
Unsupervised clustering of 197 folios recovers Currier’s labels at 95.4% to 99.5%. A single -edy threshold separates them at 99.5% leave-one-out: A a spike, 112 of 114 folios below 2%, B a continuum from 3% to 42%; the same in IT2a. Parisel’s 2026 preprint reaches the same single-switch description independently (L).
Hand determines language. With Davis’s (2020) five hands from the IVTFF headers, H(language | hand) is 0.09 bits and 98.2% of lines are predicted from hand alone; the only mixed cell is hand 3 in the stars section. The scribe, not the subject matter, sets the language.
Vocabulary sharing, calibrated. Size-matched A against B gives Jaccard 0.147 and top-100 overlap 0.95, against 0.207 and 1.00 for two halves of A, 0.172 and 0.91 for two English novels, 0.116 and 0.67 for Danish against Norwegian. A word-final pattern moving from 0.2% to 17% has no parallel between books in one language but does between languages: French against Italian, final -o, 0.19% to 17.5% (S: v_a6).
Section words. Section-specific vocabulary exists, mean chi-square per token 0.654 against 0.121 shuffled and 0.40 for Italian, and persists within one scribe and one language. But the section words belong to the same sub-word families, qokain, qol, qokeedy, okeol, not to anything resembling content vocabulary (S: v_a6).
08 Hypothesis status
excluded
- Plaintext or monoalphabetic substitution of the eleven tested languages: h2 gap at least 1.0 bit raw, at least 0.4 bit merged (S: a1, v_a1; L: Bennett 1976, Zandbergen, Lindemann & Bowern 2021).
- Substitution plus vowel removal or conventional abbreviation (S: a1, a5; L: Lindemann & Bowern 2021).
- Letter-level transposition of a real plaintext, including grille rearrangement (S: v_a5; L: Parisel 2026).
- Vigenère-type polyalphabetics (L: D’Imperio 1978, Stolfi 2000).
- Single-glyph nulls as the source of the redundancy (S: a5).
- Named decipherments: Cheshire 2019 (refuted by its reviewers), Hauer & Kondrak 2016 read as a decipherment (the authors’ own caveat), Bax 2014, Gibbs 2017, Ardic, Gladyseva 2023, Schechter 2025 and the 2025–2026 LLM-assisted readings: no reproducible key, none survives blind application (L).
- Modern forgery (L: radiocarbon, ink, provenance).
disfavoured*
- An unenciphered exotic language: the nearest published h2 is 2.77, for Hawaiian (L), and the slot grammar has no parallel.
- Rugg’s table-and-grille as the mechanism: Cardan grilles date from about 1550 against 1404–1438 vellum, the special roles of f and p cannot arise from the table (L: Zandbergen 2021), and the session’s grilles fail Zipf and hapax (S: a4).
- Free-edit self-citation as implemented here: it over-produces locality and vocabulary and under-produces character redundancy (S: a4, v_a4; section 06).
- Dee and Kelly glossolalia (L: Daruka 2020), which conflicts with the vellum date.
live
- (a) A verbose or slot-template cipher of short units, a nomenclator, or a constructed language with meaning. Compatible with sections 03, 05 and 07 and with a hand-executable 15th-century design that reproduces several headline statistics (L: Greshko 2025, the Naibbe cipher; Bowern & Gaskell 2022). It must also meet the word-length constraint (verbose controls give 7.6 to 11.5 symbols; S: v_a1), the near-zero token-to-token predictability (L: Rozanova & Temerev 2026, preprint) and the line-initial effect (S: v_a2).
- (b) Slot-preserving copy-and-modify generation (L: Timm & Schinner 2020, 2023; Timm 2026). Compatible with everything measured, but the slot grammar is an input to that model, not a product (S: a3), and no implementation has been run at manuscript length against page, section and scribe structure (L: all three sweeps).
- (c) A real language under a non-word segmentation. Untested (S: v_a3).
09 What the physical evidence adds
All (L). Radiocarbon on four folios: 1404–1438 (Hodgins, University of Arizona, 2009). Iron-gall ink and period pigments throughout (McCrone 2009; Yale 2014). Five hands, three writing 188 of the 227 pages, share one writing system (Davis 2020). Romance-dialect month names and a German-hand note on f116v put it in German- and French/Occitan-speaking hands within decades. Multispectral images of ten folios (taken 2014, released 2024) recovered an erased alphabet trial on f1r and Tepenecz’s ex libris, not new Voynichese. Provenance runs from Prague c.1600 through Baresch (1639), Marci (1665) and Kircher to 1912.
The consequence is symmetric. A meaningless-text account must be a 1420s production on fresh calfskin by several collaborating scribes who kept one slot grammar and one A/B switch consistent across 200-plus pages, a cost no statistic addresses. A cipher account must explain why the scribe rather than the subject matter sets the language (section 07).
10 What would settle the rest
- A pre-registered generator bake-off at manuscript length: self-citation, table-and-grille, Naibbe-enciphered Latin and human gibberish under one battery including the Montemurro–Zanette long-range keyword peak (L: 2013), topic-by-illustration-by-scribe co-clustering, page-adjacency similarity, A/B separability, the f and p rules and token predictability. Scripts a1 to a6 are a partial battery for any text in the TSV format.
- Manuscript-length human gibberish written over weeks with illustrations and sections, the open test named by Gaskell & Bowern (2022) (L).
- Illustration-to-text mutual information with hand, quire and adjacency partialled out.
- Re-tokenisation on certain spaces only, and the same statistics on a syllabic segmentation of Latin and Italian, to close the gap left in section 02.
- For any claimed key: blind application by a third party to five A and five B folios, with a shuffled-glyph control.
11 Checked, not checked, must verify
Checked. Every headline number in a1 to a6 was reproduced by an independent re-implementation, v_a1 to v_a6, including the Levenshtein code, entropy formulas, estimator bias, corpus contents and Gutenberg boilerplate. Robustness across ZL, RF1b, IT2a, GC2a and CD2a. Corrections applied: single-book samples; Latin word entropy from prose, not verse; a folded floor of 3.32; the rigidity ratio at 1.7 to 2.3; under-dispersion as a Currier-B property; the line-final effect as the m/g allograph; the A/B rule at 99.5% leave-one-out.
Not checked. Paywalled Cryptologia full texts (Rugg 2004, Schinner 2007, Timm & Schinner 2020 and 2023, Daruka 2020, Greshko 2025) were characterised from abstracts and secondary summaries. Higher-order and per-hand entropies, significance tests on A against B and bootstrap intervals were not computed. No higher-level literature statistic (Montemurro–Zanette peak, topic models, adjacent-page similarity) was recomputed. Davis’s hand assignments were taken from the IVTFF headers. Physical evidence beyond radiocarbon, ink, hands and the 1639 letter, including ink layering, pigment order and quire order, was not examined.
User must verify before quoting. The corpus choices (single Gutenberg books; an OCR Italian corpus with about 8% Latin; verse against prose Latin), the glyph merge lists, the design of the 24-symbol verbose controls, the treatment of uncertain spaces as breaks, and the labels excluded, disfavoured and live. Two figures appear in more than one form: merged h2 is 2.870 in the entropy table and 2.646 in the bake-off table; the adjacent-repeat rate is 0.907% in the rigidity table, 0.0082 in the bake-off table and 0.81% in the baselines. The notes do not reconcile them; different merge lists and token filters are the likely cause*, which the per-script JSON in results/ would confirm. All output is unvalidated until review.
12 Files and reproduction
Everything lives in voynich/. parse_ivtff.py converts IVTFF to TSV; data/ZL3b-n.words.tsv is the parsed main transliteration. The raw files come from voynich.nu/data with browser headers; the server refuses plain curl. baseline.py and baseline2.py are the independent baselines. Tests: a1_charstats.py, a2_wordstats.py, a3_wordgrammar.py with a3_supplement.py, a4_selfcitation.py, a5_cipher.py, a6_structure.py; re-runs v_a1.py to v_a6.py. results/ holds per-script JSON and Markdown plus generator samples. NOTES.md is the adjudication this page condenses; SOURCES.md the five literature sweeps.