Working notes · September 2026

The Scorpion letters (1991) — two ciphers below the unicity distance

Two of the five cipher passages sent to John Walsh in 1991 are public, and in both the key holds more information than the English redundancy of the text, so no ciphertext-only attack can single out a plaintext.

Daniel Bourdeau · attempted · sources: Schmeh, Top 50 no. 12; Oranchak’s scans; Pelling, Cipher Mysteries 2014–2018; the Cipher Foundation; zodiackillerciphers.com forum transcription of S5

Scorpion cipher S1, 70 symbols in a 10 by 7 grid, 1991
Scorpion cipher S1, 70 symbols in a 10 by 7 grid, 1991. FBI release via Oranchak and Schmeh.

Status   not solvable from the published text

S1 has 53 distinct symbols in 70 positions; S5 has 145 in 180. At those ratios more than one fluent English plaintext fits each ciphertext, and a search that returns fluent English has done only what the arithmetic guarantees. The one claimed reading tested here fails the single hard test a homophonic key imposes, on one securely identified symbol in each cipher.

01 What the letters are, and what is public

In 1991 a writer signing as Scorpion sent letters to John Walsh, host of America’s Most Wanted, containing five cipher passages. Two are public. S1 came with the first letter; S5 came with the letter that opens “Hi! Remember me?”. S2 to S4 are held by law enforcement and have not been released. Klaus Schmeh lists the pair as no. 12 in his Top 50 unsolved ciphers, using scans from David Oranchak’s site.

cipherlayoutsymbolsdistinctrepeat pairstranscription used here
S110 × 7705322made here from the Cipherbrain scan (454 px, upscaled 3×)
S512 rows of 1518014543forum numeric transcription, symbols 1–145 in order of first appearance
S2–S4withheld

The published material is Schmeh’s article, Nick Pelling’s Cipher Mysteries posts of 2014 to 2018, the Cipher Foundation page, and the zodiackillerciphers.com thread “Scorpion = Zodiac?”, which carries the numeric S5. The S1 transcription in s1.txt uses descriptive symbol names and flags four identifications as uncertain: circL, sqnotchBR, sqbarT and flag. Its count of 70 symbols and 53 distinct matches the published figures. The letter writes S5 at 15 symbols per row; the forum poster re-wrapped the same row-major sequence at 16, for the reason given next.

02 The structure measured here

S5: every repeat at a multiple of 16

S5 has 43 pairs of positions carrying the same symbol. Every one of the 43 lies at a distance that is a multiple of 16.

distance163248648096112128144160
pairs9564532621

Read in 16 columns, no symbol ever appears in two columns. Each column holds 11 or 12 symbols, of which 6 to 11 are distinct. That is the behaviour of 16 substitution alphabets used in strict rotation, which is what Pelling and the forum poster “Teddy” concluded between 2007 and 2014. The per-column counts, computed from s5.txt:

column12345678910111213141516
symbols12121212111111111111111111111111
distinct711101011810696109109910
repeats5122031525121221

Two numbers go further. A simulation enciphers English with 16 independent 26-letter alphabets and counts what S5 counts.

testobserved16 monoalphabetic English alphabets (simulated)
within-column repeats3547.7 ± 4.9   (z = −2.6)
spread of repeats across columns (variance)2.281.38,   P(≥ observed) = 0.05

S5 has too few repeats for 16 plain alphabets, as GeoffLaT also found. Giving the seven commonest letters two homophones each in every alphabet fits the total (32.6 ± 4.5) but makes the uneven spread across columns less likely still (P = 0.014). Column 1, the first symbol of each 16-run, repeats far more than the rest: 7 distinct symbols in 12, against 10 or 11 in most columns. Columns 1, 8 and 10 carry 5 repeats each, 15 of the 35 between them.

S1: a weak period-5 signal

S1 has 22 repeat pairs. Eight of them lie at distances that are multiples of 5, against 4.1 expected when the positions are shuffled (P = 0.036). Four lie at multiples of 10, against 1.9 expected (P = 0.12). That is the whole of the evidence for Pelling’s five cycling alphabets. It depends partly on the identifications of the filled-square and half-circle glyph families, which are the uncertain ones. S, lambda and K each recur at positions that break a strict period of 5.

03 Why they cannot be solved from what is published

A homophonic key assigns each distinct symbol one of 26 letters, so it carries about log2 26 = 4.70 bits per symbol. English supplies about 3.2 bits of redundancy per letter, taking its entropy as 1.5 bits against the 4.70 of a uniform alphabet. When the key holds more information than the text’s redundancy, the ciphertext does not determine the key.

cipherkey informationredundancy of the textratio
S153 × 4.70 = 249 bits70 × 3.2 = 224 bits1.11
S5145 × 4.70 = 682 bits180 × 3.2 = 576 bits1.18

Both sit below the unicity distance. More than one fluent English plaintext is consistent with each ciphertext, so no ciphertext-only method can single out the right one. The 16-alphabet structure of S5 makes this worse, not better: 16 independent alphabets is far more key. It would help only if the alphabets were tied to each other by a rule, and since no symbol recurs across two columns the ciphertext offers no such tie.

Matched control

unicity.py enciphers random English passages of 70 and 180 letters with random homophonic keys of exactly 53 and 145 symbols, then attacks them with the same 5-gram English annealer used elsewhere on this site, eight restarts of 15,000 steps. The per-letter figures divide the script’s totals by the text length.

controlletter accuracy of best solutionscore of found textscore of true text
S1-sized, trial 10.13−112 (−1.60 per letter)−107 (−1.53)
S1-sized, trial 20.13found: snotconcetorasestayalongdispointent…
S5-sized, trial 10.03−330 (−1.83 per letter)−326 (−1.81)
S5-sized, trial 20.12−334 (−1.86 per letter)−301 (−1.67)

The solver returns fluent English every time, and it is never the plaintext. Letter accuracy runs from 3% to 13%, and in two of the three scored trials the false solution scores within 5 points of the true text. Any solution of S1 or S5 produced by AZdecrypt-style search should be read in that light, including the three on record: Farmer 2007, Roberts 2016, and the 2018 forum reading tested next. Fluent output is guaranteed at this multiplicity and proves nothing.

04 The claimed solution of 2018, tested

Forum user Rubislaw32 posted a reading of S1 and S5 on zodiackillermystery.freeforums.net, in the thread “Scorpion’s Ciphers — The Zodiac, 1991 and beyond”, first on 19 October 2018. The plaintexts as claimed:

S1. “A picture in collection of people: Bagel Bob’s Old Dairy Frothy Late Cofee. Pour action.”
S5. “I am sending other picture of people for the collection of recent hybrid genders with edge, kept from artistic fee, if person of age has to purchase pick of university uglies. Espresso NYP cofee forces a cold enema review.”

A homophonic key maps each symbol to one letter, so the one hard test is that every repeated symbol decodes to the same letter wherever it occurs. claimed.py lays each plaintext over the transcription and checks this; results_claimed.txt holds the output.

cipherrepeated symbolsconsistentconflicts
S11310crosshair: N in COLLECTION, L in OLD. sqnotchBR: T, D, O. circL: I, L, E. The last two are flagged uncertain in the transcription.
S52726symbol 41: C in COLLECTION, G in AGE, C in COFEE

The two S1 conflicts on uncertain glyphs could be transcription error. The crosshair is not flagged: the scan shows the same glyph at row 2 column 10 and row 4 column 9, and the claimed text needs N at one and L at the other. Symbol 41 in S5 comes from the forum’s numeric transcription, which the same reading otherwise fits 26 times out of 27.

Consistency at that rate is cheap. S1 has 70 positions and 53 distinct symbols, so only 17 positions are tied to an earlier one; S5 ties 35 of 180. Every other position is free choice, and the claimed key spends it freely: in S5, 22 different symbols stand for E, 14 for O and 12 for I. A text assembled word by word under so few equalities can be almost anything. The matched controls showed that with real keys; the table at the end of this section shows it with dictionary words.

The language-model score is the second measure. It is the mean log-probability the model assigns to each letter given the four before it, on the same model that scored the controls.

text5-gram score, nats per letterdictionary coverage
S1 as claimed−2.740.83
S5 as claimed−2.610.77
reference English passage−1.580.76
control false solutions, 70 letters−1.60
control false solutions, 180 letters−1.83 to −1.86

Genuine English sits near −1.6 on this model, and so do the annealer’s false solutions, because that score is what the annealer maximises. The claimed texts sit 1.0 to 1.2 nats per letter lower, which means the model finds each letter roughly three times less likely than in ordinary English. Proper names and the spelling “cofee” cost part of that; word choice costs the rest. Dictionary coverage does not separate the three texts. So the claimed reading is less English-like than what a blind search produces on a control, and it fails the equality test once in each cipher on a securely read symbol. Neither fact says what the plaintext is. The unicity arithmetic says that nothing in the ciphertext can.

Constructively: alternatives.py fills S1 with real English words under exactly the same-symbol-same-letter rule (beam search over a Gutenberg vocabulary, word-frequency scoring). The first alternatives it returns, and their scores under the same 5-gram letter model used above:

alternative reading of S1 (70 letters, every repeat honoured)nats / letter
to the whale and in the whale and the whale and the through the highly a lay of jonah her−1.32
to the whale and of the whale and the whale and the in which the highly a lay of the ah her−1.59
the 2018 claim: a picture in collection of people bagel bobs old dairy frothy late cofee pour action−2.74

They are nonsense, they are English, and they fit the ciphertext at least as well as the claim. The corpus bias (Moby Dick) is visible and beside the point: any English word list yields such fillings. The S5 run was still in progress when this page was built; its output goes to scorpion/results_alternatives.txt.

05 Status, and what would change it

Not solvable from the published material, for a reason that is quantitative rather than a matter of effort. Three things would change it.

whatwhy it would matter
Release of S2 to S4More text under the same 16 alphabets, if the writer kept the system, adds redundancy at every letter while adding key only for symbols not yet seen.
A glyph-feature transcription of S5It would test whether shape families run in diagonals across the 16 alphabets. That is the one hypothesis that would tie the alphabets to each other and shrink the key.
The writer’s own statement“All of my ciphers can be decoded simply, once the limited patterns and systems are discovered.” If true, it points to a systematic key table rather than a random one, and a systematic table is far less key than the unicity table assumes.

Until one of those arrives, further search on the two public passages has a known result: fluent English, at a letter accuracy the controls put between 3% and 13%.

06 Files

filecontents
scorpion/s1.txtS1, 70 descriptive symbol names in 7 rows of 10, uncertain identifications marked in the notes
scorpion/s5.txtS5, the forum numeric sequence, symbols 1–145, wrapped at 16
scorpion/unicity.pythe unicity table and the matched controls; needs the English model built by copenhagen/solve.py
scorpion/claimed.pythe consistency test and 5-gram scores of the 2018 reading; output in results_claimed.txt
scorpion/alternatives.pybeam search filling each cipher with dictionary words under the same-symbol-same-letter rule, scored with the same model