Working notes · September 2026

The Voynich manuscript — probably elaborate nonsense, possibly a verbose cipher

A real fifteenth-century book, but not a language in disguise: every simple cipher of every European language tested is excluded. What remains is a rigid glyph-group encoding or structured meaningless text, and the evidence leans slightly to meaningless.

Daniel Bourdeau · attempted

Voynich manuscript, folio 34r: a herbal page with four paragraphs of Voynichese
Voynich manuscript, folio 34r: a herbal page with four paragraphs of Voynichese. Beinecke MS 408, public domain, via Wikimedia Commons.
The verdict in three tiers.

settled Beinecke MS 408 is a genuine early-15th-century codex: vellum radiocarbon-dated 1404–1438, iron-gall ink, period pigments, five scribal hands, provenance from Prague c.1600 (section 09). Modern forgery is dead.

excluded Not a plaintext, a one-to-one substitution, a letter transposition, an abjad rendering or a null-padded rendering of any of eleven European languages in ordinary word segmentation (sections 03, 04 and 05). Conditional character entropy sits at least 1.0 bit below the natural-language floor in raw EVA and at least 0.4 bit below it under three independent glyph segmentations.

live A verbose or slot-template encoding, one plaintext unit becoming a rigid two- or three-glyph group, against slot-preserving meaningless generation by copy-and-modify (sections 06 and 08). Every statistic that separates Voynichese from language is reproduced by a member of each class. Judgement*: roughly even, about 45:50, the remainder on an unenciphered exotic language.

01 The question and the method

This is not a decipherment. The question was whether the text is meaningless, enciphered or a language, and how far public data can push that in one session. Five literature sweeps, one per hypothesis or evidence class, produced SOURCES.md. Six computational tests, a1 to a6, ran on the ZL transliteration against natural-language corpora and implemented hoax generators. A separate agent re-ran every test adversarially (v_a1 to v_a6) and corrected several; corrected figures are quoted. Two baseline scripts written before the agents ran reproduce the a1 and a2 headlines: h2 of 2.32 bits in raw EVA without spaces, Zipf slope −1.04, adjacent repeats 0.81%.

(S) marks a number computed this session by a named script; (L) one quoted from the literature, with source and year. Judgement carries an asterisk.

Data (S). ZL3b-n, the Zandbergen–Landini EVA transliteration (voynich.nu, 13 May 2025), paragraph text only: 34,116 tokens after dropping words with unread or extended-EVA glyphs, about 10,750 in Currier A and 22,900 in B, 7,236 word types. Cross-checked on RF1b-er, IT2a, GC2a (v101) and CD2a (Currier). The 2,463 uncertain spaces were treated as breaks; joining them changes nothing material (v_a2). Language controls are single Gutenberg books and an OCR’d 16th-century Italian corpus. Entropies are plug-in; Miller–Madow moves h2 by at most 0.007.

02 Strongest counterargument first

Every character statistic assumes the transliteration, and that EVA spaces mark plaintext units. Latin re-segmented after every e, s and m moves 60% of the way to Voynichese slot rigidity, and Latin cut into vowel-final syllables is more rigid than Voynichese at equal length (S: v_a3). The exclusion therefore covers word-boundary-preserving encodings of eleven languages in modern orthography. Rule-based re-segmentation, or a language with rigid consonant-vowel syllables written syllabically, was not tested.

Second, the verbose control that matches raw EVA on h1, h2 and their ratio at once produces words 2.3 times too long, 11.52 symbols against 5.07, while carrying Latin word entropy (S: v_a1). That refutes a letter-level verbose cipher of whole Latin words. Both points narrow the verdict. Neither overturns it.

03 Character level

The firmest result is conditional character entropy h2, the uncertainty of a glyph given the one before it. Bennett (1976) put Voynichese at 2.22 bits against 3.01 to 3.37 for European languages (L); Lindemann & Bowern (2021) found it below all 250 languages in a Wikipedia corpus (L). The gap survives glyph merging and the Currier and v101 alphabets. All rows (S), no spaces, a1, v_a1 and v_a5.

texth1h2h1−h2h2/h1mean word length
ZL raw EVA, Currier A3.8322.3531.4790.6144.98
ZL raw EVA, Currier B3.8592.1821.6770.5655.12
ZL raw EVA, all3.8652.3111.5540.5985.07
ZL merged (benches, i/e groups, aiin, qo)4.0392.8701.1700.7103.7–3.9
GC v101 (68 symbols)4.1442.8931.250.6983.7–4.0
Currier CD2a (35 symbols)3.8572.6601.1970.690
Latin, Aeneid4.0253.5100.5150.8725.77
German (diacritics folded 3.335)4.1603.3910.7690.8154.88
Range of 11 languages (floor 3.32 folded)3.94–4.523.32–3.880.49–0.770.82–0.884.2–6.6
Latin, vowels removed3.6593.5150.1440.9613.11
Verbose cipher of Latin, every letter 2–3 symbols, 24 symbols3.9012.3251.5760.59611.52
Verbose cipher, 8 letters kept single4.1022.9161.1860.7117.64

Raw EVA gives 2.18 to 2.35 bits; ZL merged, v101 and Currier give 2.66 to 2.89; the eleven-language floor is 3.32 with diacritics folded. Any transposition of Latin gives at least 3.44. Vowel deletion leaves Latin at 3.52 with h2/h1 of 0.961, the wrong direction for an abjad.

Where the redundancy sits. Inside words. Within-word h2 is 1.92 to 2.13 bits raw against 2.94 to 3.41 for languages. Mutual information between glyph and word position is 0.69 to 0.73 bits across ZL, Currier and v101 against 0.17 to 0.27; 21% to 31% of glyphs are locked to one word position against 0% to 8% of letters (S: a5, v_a5).

Pair merging. Greedy merging of frequent glyph pairs closes the h1−h2 gap only after about 40 merges, so an encoding of two- to three-glyph units drawn from 50 to 60 unit types is compatible with the character statistics and a plain letter cipher is not (S: a5). Greedy merging tests one heuristic, not the whole verbose class (S: v_a5).

Information budget. The word stream carries 425,468 bits, 12.47 bits per token, 0.82 times a Latin prose text of equal length under the same model (S: a5, v_a5).

04 Word level

At word level Voynichese looks like a language, as Landini (2001) and Reddy & Knight (2011) reported (L). Table (S), a2 and v_a2, 34,116 tokens.

corpusZipf slope, ranks 1–1000hapax % of typesHeaps betatop-word share
Voynich all−1.04069.70.7102.25%
Voynich A / B−0.994 / −1.06872.2 / 68.60.743 / 0.6794.27% / 2.13%
Latin (Aeneid)−0.82067.00.7494.9%
German−1.04565.10.7793.46%
English−1.05655.10.6754.68%

Word-type entropy is 9.8 to 10.3 bits against 10.3 to 10.8 for Latin prose and 9.4 to 9.8 for Italian, French and German (S: v_a1). The low pooled top-word share is an A+B mixture effect (S: v_a2); against single-book samples the hapax rate is above every language (S: v_a5). None of this discriminates: Zipf and hapax are reproduced by human gibberish (Gaskell & Bowern 2022), by generators (Rugg & Taylor 2016; Timm & Schinner 2020) and by a hand cipher from Latin (Greshko 2025) (L).

Word length

Merged EVA word lengths (S: a2, v_a2):

corpusmeanvariancevariance / mean of (length−1)
Voynich all (raw EVA 5.07 / 3.73 / 0.92)4.112.480.80
Voynich A3.932.890.99
Voynich B4.202.270.71
Aeneid alone / Confessions alone4.97 /1.04 / 1.74
Italian / German / English / Danish4.53 / 4.99 / 4.23 / 4.676.52 / 7.04 / 5.41 / 7.401.85 / 1.77 / 1.67 / 2.02

The famous under-dispersion, “binomial word lengths”, is a Currier-B property. B has variance-to-mean 0.71, A has 0.99 and single-text Latin verse 1.04, so Currier A matches Latin verse and the binomial-fit contrast in a2 was a search-cap artefact. B and the pooled text stay far from the 1.67 to 2.02 of modern prose.

05 Word grammar and repetition

Tiltman (1951), Stolfi (2000) and Zattera (2022) described a rigid positional grammar inside Voynichese words (L). Measured here as the fraction of glyph pairs whose order violates the single best glyph ordering, length-matched; a shuffle null gives 0.49. All rows (S): a3, a2, v_a2.

statisticVoynichLatinItalianGermanEnglish
Pair-order violations against the best glyph order, length-matched (shuffle null 0.49)0.160 (v101 0.140)0.3630.3180.2780.308
Adjacent identical tokens, % and observed:shuffled0.907 / 2.50.047 / 0.10.047 / 0.10.157 / 0.30.137 / 0.2
Line-initial first-glyph effect, Cramér’s V, paragraph-first lines excluded0.277pseudo-lined prose 0.02–0.03; hexameter 0.148

Voynichese is 1.7 to 2.3 times more rigid than any tested language. The glyph order derived from the data with no assumptions, q, then p f sh, then ch, then o t k, then e ee, then d s a, then l r y, then g n iin m, is the Stolfi and Zandbergen layout recovered blind (S: a3).

Line position matters, as Currier (1976) said (L). The line-initial effect exceeds hexameter verse and is an order of magnitude above pseudo-lined prose; the line-final effect a2 reported collapses to V 0.07 to 0.16 once the final m and g allographs are removed (S: v_a2).

06 Self-citation and the generator bake-off

Timm (2014, 2016) argued that similar words cluster on the page, the fingerprint of copying an earlier word and modifying it; Timm & Schinner (2020) published an algorithm that does this (L). The test finds, for each token, the nearest earlier token on the same page within edit distance 1, with every corpus poured into the identical page and line skeleton. Two Rugg-style generators and a re-implemented autocopist ran through the same code. All rows (S), a4 and v_a4.

corpusmedian distancewithin 10 tokenswithin-page shuffleglobal shufflelong tokens (5+) within 10
Voynich ZL140.2690.2380.1670.225
Latin (Aeneid)410.0350.0340.0300.006
English (Italian, German similar)170.1760.1770.1640.019
Danish (verse)150.1960.1720.1380.035
Autocopist (Timm–Schinner re-implementation)130.3290.2750.0790.257
Rugg grille360.0670.2490.0660.033
Rugg sorted table340.3290.1620.1120.341
Order-3 character Markov model trained on Voynichese190.1940.1940.1940.119

Sequential excess, real minus within-page shuffle, measures copying order: Voynich +0.031, Danish verse +0.024, autocopist +0.054, prose about zero. Page-vocabulary excess, real minus global shuffle: Voynich +0.102, languages at most +0.058, autocopist +0.250. For tokens of five or more glyphs the language baselines collapse to 0.035 or less while Voynichese stays at 0.225, but an order-3 character Markov model that copies nothing reaches 0.119: about half the long-token density is word-internal structure, not copying (S: v_a4).

The bake-off

texttypesh2 rawh2 mergedZipf slopehapax %adjacent repeatpair-violation
Voynich ZL7,2362.3112.646−1.04069.70.00820.160
Autocopist, two seeds11,660 / 11,3302.709 / 2.7233.274 / 3.005−0.825 / −0.85972.2 / 71.10.0013 / 0.0018free-edit 0.41–0.46
Rugg grille / sorted table3,747 / 1,7572.798 / 2.6293.176 / 2.912−0.429 / −0.6225.7 / 13.90.0003 / 0.0893fitted table 0.122

No generator matches on every column. The autocopist yields only 37% real Voynich word types by token, 2% qo-initial words against 15%, and h2 0.4 to 0.6 bits too high; it over-produces locality and page vocabulary where the manuscript’s own sequential excess is only the size of Danish verse. It is a re-implementation with documented deviations, not Timm’s reference program. The grilles fail Zipf and hapax outright. The one table that reproduces the slot rigidity had its columns read off Voynichese: sufficiency only (S: v_a3).

07 Currier A and B

Currier (1976) found two statistical languages in the manuscript (L). What separates them (S: a6, v_a6):

statisticAB
Tokens ending -edy / containing ed0.20% / 0.29%17.34% / 20.84%
Containing cho / starting qo- / ending -ol16.1 / 10.0 / 14.5%2.7 / 17.7 / 7.1%
Glyph after ch: e / o24.5 / 46.1%60.1 / 10.1%
daiin per thousand tokens42.513.0

Unsupervised clustering of 197 folios recovers Currier’s labels at 95.4% to 99.5%. A single -edy threshold separates them at 99.5% leave-one-out: A a spike, 112 of 114 folios below 2%, B a continuum from 3% to 42%; the same in IT2a. Parisel’s 2026 preprint reaches the same single-switch description independently (L).

Hand determines language. With Davis’s (2020) five hands from the IVTFF headers, H(language | hand) is 0.09 bits and 98.2% of lines are predicted from hand alone; the only mixed cell is hand 3 in the stars section. The scribe, not the subject matter, sets the language.

Vocabulary sharing, calibrated. Size-matched A against B gives Jaccard 0.147 and top-100 overlap 0.95, against 0.207 and 1.00 for two halves of A, 0.172 and 0.91 for two English novels, 0.116 and 0.67 for Danish against Norwegian. A word-final pattern moving from 0.2% to 17% has no parallel between books in one language but does between languages: French against Italian, final -o, 0.19% to 17.5% (S: v_a6).

Section words. Section-specific vocabulary exists, mean chi-square per token 0.654 against 0.121 shuffled and 0.40 for Italian, and persists within one scribe and one language. But the section words belong to the same sub-word families, qokain, qol, qokeedy, okeol, not to anything resembling content vocabulary (S: v_a6).

08 Hypothesis status

excluded

  1. Plaintext or monoalphabetic substitution of the eleven tested languages: h2 gap at least 1.0 bit raw, at least 0.4 bit merged (S: a1, v_a1; L: Bennett 1976, Zandbergen, Lindemann & Bowern 2021).
  2. Substitution plus vowel removal or conventional abbreviation (S: a1, a5; L: Lindemann & Bowern 2021).
  3. Letter-level transposition of a real plaintext, including grille rearrangement (S: v_a5; L: Parisel 2026).
  4. Vigenère-type polyalphabetics (L: D’Imperio 1978, Stolfi 2000).
  5. Single-glyph nulls as the source of the redundancy (S: a5).
  6. Named decipherments: Cheshire 2019 (refuted by its reviewers), Hauer & Kondrak 2016 read as a decipherment (the authors’ own caveat), Bax 2014, Gibbs 2017, Ardic, Gladyseva 2023, Schechter 2025 and the 2025–2026 LLM-assisted readings: no reproducible key, none survives blind application (L).
  7. Modern forgery (L: radiocarbon, ink, provenance).

disfavoured*

live

09 What the physical evidence adds

All (L). Radiocarbon on four folios: 1404–1438 (Hodgins, University of Arizona, 2009). Iron-gall ink and period pigments throughout (McCrone 2009; Yale 2014). Five hands, three writing 188 of the 227 pages, share one writing system (Davis 2020). Romance-dialect month names and a German-hand note on f116v put it in German- and French/Occitan-speaking hands within decades. Multispectral images of ten folios (taken 2014, released 2024) recovered an erased alphabet trial on f1r and Tepenecz’s ex libris, not new Voynichese. Provenance runs from Prague c.1600 through Baresch (1639), Marci (1665) and Kircher to 1912.

The consequence is symmetric. A meaningless-text account must be a 1420s production on fresh calfskin by several collaborating scribes who kept one slot grammar and one A/B switch consistent across 200-plus pages, a cost no statistic addresses. A cipher account must explain why the scribe rather than the subject matter sets the language (section 07).

10 What would settle the rest

  1. A pre-registered generator bake-off at manuscript length: self-citation, table-and-grille, Naibbe-enciphered Latin and human gibberish under one battery including the Montemurro–Zanette long-range keyword peak (L: 2013), topic-by-illustration-by-scribe co-clustering, page-adjacency similarity, A/B separability, the f and p rules and token predictability. Scripts a1 to a6 are a partial battery for any text in the TSV format.
  2. Manuscript-length human gibberish written over weeks with illustrations and sections, the open test named by Gaskell & Bowern (2022) (L).
  3. Illustration-to-text mutual information with hand, quire and adjacency partialled out.
  4. Re-tokenisation on certain spaces only, and the same statistics on a syllabic segmentation of Latin and Italian, to close the gap left in section 02.
  5. For any claimed key: blind application by a third party to five A and five B folios, with a shuffled-glyph control.

11 Checked, not checked, must verify

Checked. Every headline number in a1 to a6 was reproduced by an independent re-implementation, v_a1 to v_a6, including the Levenshtein code, entropy formulas, estimator bias, corpus contents and Gutenberg boilerplate. Robustness across ZL, RF1b, IT2a, GC2a and CD2a. Corrections applied: single-book samples; Latin word entropy from prose, not verse; a folded floor of 3.32; the rigidity ratio at 1.7 to 2.3; under-dispersion as a Currier-B property; the line-final effect as the m/g allograph; the A/B rule at 99.5% leave-one-out.

Not checked. Paywalled Cryptologia full texts (Rugg 2004, Schinner 2007, Timm & Schinner 2020 and 2023, Daruka 2020, Greshko 2025) were characterised from abstracts and secondary summaries. Higher-order and per-hand entropies, significance tests on A against B and bootstrap intervals were not computed. No higher-level literature statistic (Montemurro–Zanette peak, topic models, adjacent-page similarity) was recomputed. Davis’s hand assignments were taken from the IVTFF headers. Physical evidence beyond radiocarbon, ink, hands and the 1639 letter, including ink layering, pigment order and quire order, was not examined.

User must verify before quoting. The corpus choices (single Gutenberg books; an OCR Italian corpus with about 8% Latin; verse against prose Latin), the glyph merge lists, the design of the 24-symbol verbose controls, the treatment of uncertain spaces as breaks, and the labels excluded, disfavoured and live. Two figures appear in more than one form: merged h2 is 2.870 in the entropy table and 2.646 in the bake-off table; the adjacent-repeat rate is 0.907% in the rigidity table, 0.0082 in the bake-off table and 0.81% in the baselines. The notes do not reconcile them; different merge lists and token filters are the likely cause*, which the per-script JSON in results/ would confirm. All output is unvalidated until review.

12 Files and reproduction

Everything lives in voynich/. parse_ivtff.py converts IVTFF to TSV; data/ZL3b-n.words.tsv is the parsed main transliteration. The raw files come from voynich.nu/data with browser headers; the server refuses plain curl. baseline.py and baseline2.py are the independent baselines. Tests: a1_charstats.py, a2_wordstats.py, a3_wordgrammar.py with a3_supplement.py, a4_selfcitation.py, a5_cipher.py, a6_structure.py; re-runs v_a1.py to v_a6.py. results/ holds per-script JSON and Markdown plus generator samples. NOTES.md is the adjudication this page condenses; SOURCES.md the five literature sweeps.