4 · pot · 97 · 400   |   806 845 61 407 850 900 740 = the horned archer
255 435 690 740 · 705 33 520 · name + ending, never split by a line break

Mohenjo-daro, Harappa and 50 other sites · logo-syllabic script, no accepted decipherment · c. 2600–1900 BCE · attempted, not read

The Indus script: Parpola’s decipherment and 2,364 registered tests

Parpola’s Dravidian readings, and then 2,364 hypotheses registered before testing, run on two independent corpora with Linear B and Ur III seals as controls: the grammar a reading must fit is set out, and no sign value is established.

Daniel Bourdeau · posted

Method: not solved · Extent: text not obtained

At a glanceNot solved · 1,224 of 2,364 registered hypotheses held

The text
The starting point was Asko Parpola’s lecture “Study of the Indus Script” (ICES Tokyo 2005, on harappa.com). The inscriptions behind it are some 4,500 seals, tablets and sherds, most with fewer than five signs.
Where
Two machine-readable transcriptions made independently: a database derived from the Interactive Corpus of Indus Texts (ICIT; Wells and Fuls), and Mahadevan’s 1977 concordance (M77). They were tied together sign by sign here. A digitisation of Parpola’s own Corpus of Indus Seals and Inscriptions served as a third check.
The script
About 400–600 signs, mostly written right to left, with no word dividers, no bilingual and no known language.
How it was tried
First, every count in Parpola’s lecture and in his 1994 table of readings was rerun, and his Tamil check was given a chance baseline. Then 175 sets of hypotheses, 2,364 in all, were written down and committed before the data were looked at. Each was tested and recorded whether it held or failed, and the main results were rerun on held-out samples and against two scripts that can be read (Linear B, Ur III cuneiform).
What came of it
A detailed grammar with no sound values. A name ends in a slot whose form (740, 520 or a closer) is chosen by the name’s last sign. Scribes never split a bound pair or a name from its ending across a line. Numerals are fixed with the signs they count. Seal names behave like personal names, as Linear B names of the same length do. Parpola’s star-name readings rest on one or two seals each.
Still open
Everything phonetic, and the language. Several of this project’s own earlier conclusions were withdrawn by later tests; they are listed in section 13.

01 The corpora and the signs

Parpola’s arguments are counts: how often the commonest sign occurs, which signs stand together, which never do. To rerun them the project used two transcriptions made independently of each other. The first is a database derived from the Interactive Corpus of Indus Texts (Wells and Fuls), published with the indus-website project: 2,543 objects, 11,253 signs, with CISI object numbers, sites and field motifs. A fuller export of the same database (5,680 records, with find spots, levels, sizes and materials) was used for most of the later tests. Its sign totals closely match those printed in Fuls’s Catalog of Indus Signs (2023) for the signs checked. The ICIT signs are numbers only, so they were identified by drawing each one in the database’s own font and matching Parpola’s descriptions. The second transcription is Mahadevan’s 1977 concordance (2,906 texts). The two sign lists were tied together by aligning the texts they share: 991 lines, with 96% of signs on a single counterpart. The 1,664 M77 texts missing from the first corpus served as an independent replication sample throughout.

The signs discussed on this page with their ICIT numbers
The signs this page talks about, with their ICIT numbers and counts: the jar 740, the fish series (220 plain, 235 ‘roof’, 233, 231, 240), the crab 798, the fig 772–786, the fig + crab ligature 777, the eye, the double curve, the pot 700, the numerals, and the copper-tablet signs 806 845 61 407 850. Image: Drawn here in the indus-website font (ICIT sign numbers).

02 The lecture’s claims, checked

Impression of the bar seal M-414 from Mohenjo-daro
M-414, a bar seal from Mohenjo-daro, as impressed. Read from right to left: an unclear crossed sign, the U-shaped ‘fig’ with barred branches, and the plain ‘fish’. This is Parpola’s second ‘fig + fish’, vaṭa-mīn ‘north star’. Image: Corpus of Indus Seals and Inscriptions, vol. 1 (Joshi and Parpola 1987), p. 100, M-414 a.

03 The 1994 readings and the Tamil check

Parpola’s book of 1994 closes with a table of 24 readings and a list of the 99 Tamil compounds in mīn ‘fish, star’ that he drew on. In both corpora these sign pairs are real, recurring units: 3 + fish, 6 + fish, eye + eye, hearth + rings, rings + ‘space’, and ‘space’ + fish. Fish + fish, 7 + fish, fig + fish, fig + ‘space’ and crab + plain fish are at or below chance. The book says the numbers before the fish “are restricted to 3, 4, 6 and 7”. The corpora have 1, 2, 3, 4, 6, 7 and 12, and never 5, although Tamil has a star name for 5.

The check behind every reading is that the compound it gives is a real Tamil word. A chance baseline shows how weak that is. Take 98 picture concepts fixed in advance. Some Tamil word for the picture begins one of Parpola’s Tamil star names 15% of the time, and 30% with the sound latitude the readings use.

04 The copper tablets: seven sign = image equations

The copper tablets of Mohenjo-daro come in sets of identical copies. Each has an inscription on one side and an animal or figure on the other, and in some sets a single sign takes the place of the picture, with the same inscription. Matching Parpola’s typology of the 46 sets to the transcriptions gives seven equations:

These are the firmest meanings the script offers. The picture signs 749, 341 and 753 occur only on copper tablets, and 777 on one seal besides: no seal name uses them. The heading that opens many seal texts never opens a copper-tablet text (0 of 198). The copper tablets behave differently from the seals. In the capture–recapture test of section 9, their texts form a small closed set of labels, repeated across the city like titles, while seal texts are individual. The tablets’ pictures also cluster by quarter of Mohenjo-daro: text-only tablets in DK-G South, the animal series in VS-A, figures and composites on the citadel mounds.

The commonest copper-tablet inscriptions with their counts and images
The commonest copper-tablet inscriptions, as written (right to left), with the number of copies, the reading order and the picture on the other side: the hare, the archer, the elephant, the fig + crab ligature standing alone. Image: Drawn here in the indus-website font from the ICIT-derived corpus.

05 Which language? Sanskrit and Sumerian as controls

The same checks were run on Sanskrit (Monier-Williams) and Sumerian (the ePSD2 glossary). A picture gives a star name in 19% of cases in Sanskrit and 10% in Sumerian, against 15% for Tamil. The fish = star pun is not Dravidian only either: Sumerian mul is ‘star’ and also a fish, and Sanskrit has ṛṣi (a fish; the Seven Sages of the Great Bear). Across 2,895 languages (CLICS4), though, the ordinary words for ‘fish’ and ‘star’ coincide in only three, none in South Asia, and Old Tamil mīn is one of them. So the first step of every fish reading is genuinely distinctive of Dravidian, even though the Tamil compound check built on it is not. The numbers before the fish do not follow the Tamil star names: only 6 + fish, the Pleiades, is enriched, and 5 + fish never occurs.

06 How the rest was done: hypotheses registered before testing

Everything after the checks above was run as registered predictions. For each set, the hypotheses and their thresholds were written into indus/PREDICTIONS.md and committed to the repository before any script looked at the data. The test was then run, and the result recorded as held or failed without changing the threshold. 175 sets were run this way, with 2,364 hypotheses: 1,224 held and 1,140 failed. Where a test turned out to be ill-posed, or a script had a bug, the record says so and gives the corrected result beside the original.

Counts are over distinct texts. The same seal impressed on many sealings, or a batch of identical tablets, counts once, because repeated copies inflate any pattern. Main results were rerun on samples not used to find them: Mahadevan’s M77 additions, the smaller sites, the fuller corpus without copper tablets, and Parpola’s CISI transcription. Two scripts that can be read served as controls: Linear B (Mycenaean Greek, from the DAMOS database) and the seal legends of the Ur III period (Sumerian, 22,402 impressions from ORACC).

07 The grammar a reading must fit

08 A second transcription, and other people’s claims

09 What kind of names?

A seal text is short, sits under an animal and ends in a grammatical slot. Is it a personal name, a title or office, a god? Capture–recapture, the method ecologists use to estimate a population from two overlapping samples, gives a way to ask this without sound values. Take the names on Mohenjo-daro seals and the names on Harappa seals as two catches from one population. How many names appear in both tells how large the whole population was.

10 Place and time

11 The language, as far as order shows it

Without sound values, only word order and the shape of the grammar speak to the language. Four features point one way:

That fits Dravidian and Indo-Aryan, and counts against Sumerian and Elamite, which put the head first. Two further points narrow it. The endings split people from the fish names (Parpola’s stars), which Munda grammar would not do. And the fish = star word is Dravidian-only in South Asia. Together they lean to Dravidian, but only if the fish names are stars, and nothing here separates Dravidian from Indo-Aryan on order alone. The family expectations are standard typology, not a checked source for each language.

Name lists in the candidate languages add two results. Sumerian personal names from Ur III seals fix their first element (Ur-, Lu-, Nin-, the divine sign), while Indus names fix their last, as Linear B and Sanskrit names do. That rules out a Sumerian-type name structure a second time. Between Dravidian and Indo-Aryan the lists do not decide. Indus name lengths sit nearer Sanskrit names. The dominance of one ending (740 ends 82% of names) sits nearer Old Tamil personal names, 67% of which end in the masculine -aṉ, the suffix Mahadevan reads 740 as. A closer Indo-Aryan comparator then turned the argument around. Lüders’s list of early Brahmi inscriptions (1912) gives 374 Prakrit donor names from Sanchi and Bharhut, c. 200 BCE–400 CE. These are personal names on donated objects, near the Indus seals in time, place and genre. Indus names are built like them:

The Prakrit donor names also carry one near-universal ending, the genitive -sa of the gift formula. So a dominant ending like 740 fits an Indo-Aryan donor-style formula as well as Tamil -aṉ. The matching Dravidian test used Tamil names from donation records: 1,396 names from the English translations of the DHARMA Tamil inscriptions, 7th century onwards. It splits the evidence:

Name structure therefore rules out a Sumerian type but does not decide between Indo-Aryan and Dravidian. The Tamil records are also much later than the Prakrit list, and many of their names are Sanskrit-derived.

A contemporary Dravidian list was then built from the page scans of Mahadevan’s Early Tamil Epigraphy: 104 donor names from the Tamil-Brahmi cave inscriptions, c. 2nd century BCE – 4th century CE. That matches the Prakrit list in period and genre.

So on structure the Indus names look more like early Indo-Aryan donor names than contemporary Tamil ones. That is not a language identification. The Tamil-Brahmi list is small, syllables are not Indus signs, several Tamil-Brahmi donors bear Prakrit names themselves, and the Indus texts are some two thousand years older than either list.

12 What remains open

13 What was withdrawn along the way

Registering predictions first means some of this project’s own readings were overturned by later tests. They are kept in the record, and listed here:

14 Sources

The scripts, the registered hypotheses (PREDICTIONS.md), the results of every set (results/) and the working notes (NOTES.md) are in the repository folder indus/.