At a glanceNot solved · 1,224 of 2,364 registered hypotheses held
- The text
- The starting point was Asko Parpola’s lecture “Study of the Indus Script” (ICES Tokyo 2005, on harappa.com). The inscriptions behind it are some 4,500 seals, tablets and sherds, most with fewer than five signs.
- Where
- Two machine-readable transcriptions made independently: a database derived from the Interactive Corpus of Indus Texts (ICIT; Wells and Fuls), and Mahadevan’s 1977 concordance (M77). They were tied together sign by sign here. A digitisation of Parpola’s own Corpus of Indus Seals and Inscriptions served as a third check.
- The script
- About 400–600 signs, mostly written right to left, with no word dividers, no bilingual and no known language.
- How it was tried
- First, every count in Parpola’s lecture and in his 1994 table of readings was rerun, and his Tamil check was given a chance baseline. Then 175 sets of hypotheses, 2,364 in all, were written down and committed before the data were looked at. Each was tested and recorded whether it held or failed, and the main results were rerun on held-out samples and against two scripts that can be read (Linear B, Ur III cuneiform).
- What came of it
- A detailed grammar with no sound values. A name ends in a slot whose form (740, 520 or a closer) is chosen by the name’s last sign. Scribes never split a bound pair or a name from its ending across a line. Numerals are fixed with the signs they count. Seal names behave like personal names, as Linear B names of the same length do. Parpola’s star-name readings rest on one or two seals each.
- Still open
- Everything phonetic, and the language. Several of this project’s own earlier conclusions were withdrawn by later tests; they are listed in section 13.
01 The corpora and the signs
Parpola’s arguments are counts: how often the commonest sign occurs, which signs stand together, which never do. To rerun them the project used two transcriptions made independently of each other. The first is a database derived from the Interactive Corpus of Indus Texts (Wells and Fuls), published with the indus-website project: 2,543 objects, 11,253 signs, with CISI object numbers, sites and field motifs. A fuller export of the same database (5,680 records, with find spots, levels, sizes and materials) was used for most of the later tests. Its sign totals closely match those printed in Fuls’s Catalog of Indus Signs (2023) for the signs checked. The ICIT signs are numbers only, so they were identified by drawing each one in the database’s own font and matching Parpola’s descriptions. The second transcription is Mahadevan’s 1977 concordance (2,906 texts). The two sign lists were tied together by aligning the texts they share: 991 lines, with 96% of signs on a single counterpart. The 1,664 M77 texts missing from the first corpus served as an independent replication sample throughout.

02 The lecture’s claims, checked
- The commonest sign (the ‘jar’, 740) is almost 10% of all signs: 9.9% in M77, 11.3% in ICIT. It never stands beside itself in South Asia (0, against 4.0 expected if the order within each line were shuffled, p about 0.007). It does so once in each corpus, on the same unprovenanced round seal, as Parpola says.
- Seals found in the Gulf and Mesopotamia use common signs in unusual order. One square seal (Salut, Oman) is also unusual, against the lecture.
- Fish signs are 9.7% of seal signs. The crab stands beside a fish sign far more often than chance, but beside the marked fish, not the plain one.
- ‘7 + fish’, read as Ursa Major, occurs once, on a single seal from Harappa, in both corpora. ‘Fig + fish’, the North Star, occurs twice (M-172, and M-414, checked on its photograph in CISI vol. 1).
- The numbers before the fish. Parpola calls ‘6 + fish’ (the Pleiades) the commonest numeral + fish sequence. Counting each distinct text once, the commonest is ‘2 + fish’ (113 texts), then ‘3 + fish’ (28); ‘6 + fish’ has 8. Parpola’s own transcription in the CISI digitisation gives the same answer on the seals it covers: 2 + fish thirteen times, 6 + fish never. Counting every copy, or using Mahadevan’s transcription instead, 2 + fish is still the commonest (117 and 123 cases, against 9 and 12 for 6 + fish). Three transcriptions and two ways of counting agree, so the difference is not a matter of how ICIT identifies the signs. Part of the ‘2’ is the long stroke pair (51 of 113 lines), which Parpola reads not as a number but as vēḷ, long pair + fish being his ‘Venus’. Even without it, short 2 + fish (60) far outnumbers 6 + fish (8).

03 The 1994 readings and the Tamil check
Parpola’s book of 1994 closes with a table of 24 readings and a list of the 99 Tamil compounds in mīn ‘fish, star’ that he drew on. In both corpora these sign pairs are real, recurring units: 3 + fish, 6 + fish, eye + eye, hearth + rings, rings + ‘space’, and ‘space’ + fish. Fish + fish, 7 + fish, fig + fish, fig + ‘space’ and crab + plain fish are at or below chance. The book says the numbers before the fish “are restricted to 3, 4, 6 and 7”. The corpora have 1, 2, 3, 4, 6, 7 and 12, and never 5, although Tamil has a star name for 5.
The check behind every reading is that the compound it gives is a real Tamil word. A chance baseline shows how weak that is. Take 98 picture concepts fixed in advance. Some Tamil word for the picture begins one of Parpola’s Tamil star names 15% of the time, and 30% with the sound latitude the readings use.
04 The copper tablets: seven sign = image equations
The copper tablets of Mohenjo-daro come in sets of identical copies. Each has an inscription on one side and an animal or figure on the other, and in some sets a single sign takes the place of the picture, with the same inscription. Matching Parpola’s typology of the 46 sets to the transcriptions gives seven equations:
- the fig + crab ligature 777 = the markhor goat and the horned archer;
- sign 749 = the markhor goat;
- 341 = the rhinoceros;
- 753 = the hare;
- a lens-shaped sign = the long-horned bull;
- “3 700 900 165” = a spotted bull.
These are the firmest meanings the script offers. The picture signs 749, 341 and 753 occur only on copper tablets, and 777 on one seal besides: no seal name uses them. The heading that opens many seal texts never opens a copper-tablet text (0 of 198). The copper tablets behave differently from the seals. In the capture–recapture test of section 9, their texts form a small closed set of labels, repeated across the city like titles, while seal texts are individual. The tablets’ pictures also cluster by quarter of Mohenjo-daro: text-only tablets in DK-G South, the animal series in VS-A, figures and composites on the citadel mounds.

05 Which language? Sanskrit and Sumerian as controls
The same checks were run on Sanskrit (Monier-Williams) and Sumerian (the ePSD2 glossary). A picture gives a star name in 19% of cases in Sanskrit and 10% in Sumerian, against 15% for Tamil. The fish = star pun is not Dravidian only either: Sumerian mul is ‘star’ and also a fish, and Sanskrit has ṛṣi (a fish; the Seven Sages of the Great Bear). Across 2,895 languages (CLICS4), though, the ordinary words for ‘fish’ and ‘star’ coincide in only three, none in South Asia, and Old Tamil mīn is one of them. So the first step of every fish reading is genuinely distinctive of Dravidian, even though the Tamil compound check built on it is not. The numbers before the fish do not follow the Tamil star names: only 6 + fish, the Pleiades, is enriched, and 5 + fish never occurs.
06 How the rest was done: hypotheses registered before testing
Everything after the checks above was run as registered predictions. For each set, the hypotheses and their thresholds were written into indus/PREDICTIONS.md and committed to the repository before any script looked at the data. The test was then run, and the result recorded as held or failed without changing the threshold. 175 sets were run this way, with 2,364 hypotheses: 1,224 held and 1,140 failed. Where a test turned out to be ill-posed, or a script had a bug, the record says so and gives the corrected result beside the original.
Counts are over distinct texts. The same seal impressed on many sealings, or a batch of identical tablets, counts once, because repeated copies inflate any pattern. Main results were rerun on samples not used to find them: Mahadevan’s M77 additions, the smaller sites, the fuller corpus without copper tablets, and Parpola’s CISI transcription. Two scripts that can be read served as controls: Linear B (Mycenaean Greek, from the DAMOS database) and the seal legends of the Ur III period (Sumerian, 22,402 impressions from ORACC).
07 The grammar a reading must fit
- Four kinds of text on one vocabulary. Almost every line is a name with its ending, a line ending in a ‘closer’, a count (numbers with the thing counted), or a bare sequence without an ending. Only 3 of 2,722 distinct lines fit none of the four. Each sign keeps its position and its company across all four frames.
- The ending slot is a paradigm chosen by the name’s last sign. A name ends in the jar 740, the ‘arrow’ 520, one of a dozen closers, or 740 followed by a stacking closer. Then comes an optional 400, or 90 (after 740 only). The last sign of the name predicts which ending it takes (a model of this predicts 95% of the endings in Mahadevan’s separate transcription). Only 5 of 881 names ever appear with both 740 and 520. 740 and 520 never touch.
- The two endings are not alike. 740 is the open ending: every name built on a main word first seen in the late levels takes it (34 of 34). The 520 names use only 26 distinct main words in 163 names, mostly fish signs, and include fixed expressions such as the seal closing formula ‘705 (or 706) + long 3 + 520’.
- Names are head-final. The ending follows the last sign of a name, not the first. Names grow at the front from existing names, and numerals stand before what they count (472 against 106 cases). That is the order of Dravidian and Indo-Aryan, and the reverse of Sumerian and Elamite. There is no number agreement, and no long chains of suffixes: the ending slot is short and closed.
- Line breaks respect the units. On multi-line seals and tablets, in either line order, the scribes never split one of the 30 most tightly bound sign pairs across a line (0 against about 5 by chance). They almost never split a recurring name unit, or a name from its ending. Each line is a complete text of one of the four kinds, and where one line is a name, the other is a separate field. Signs 1 and 2 do not act as word dividers, as Fuls (2024) argues against Wells (2011).
- Numerals are fixed with the signs they count. One number system serves every kind of text, written with short, long and tiered strokes. Inside names, the number before a given sign is fixed, as if numeral + sign were a word. The commonest value before the same sign agrees between Mohenjo-daro and Harappa for only 39% of signs, with ‘2 + fish’ the shared exception. But real counts are about as local: in Linear B the commonest quantity of a commodity differs between Knossos and Pylos for 8 of 17 commodities. So this looks like counting, not a naming custom.
- Predictability. A model of sign position and neighbours predicts held-out texts at 4.66 bits per sign. Half the signs can be guessed from their two neighbours. And the script has forbidden pairs: 20 pairs of common signs that chance would put together five or more times never occur, 17 times more such pairs than in shuffled lines. For example, the fish 220 and 240 never stand directly before 390.
08 A second transcription, and other people’s claims
- Parpola’s own transcription agrees. The open digitisation of the Corpus of Indus Seals and Inscriptions covers the Mohenjo-daro seals M-1 to M-184. On those seals it matches the ICIT transcription on length for 96% of objects, on 92% of sign positions and on 98% of last signs. The disagreements fall on rare signs, and the ending rules hold in Parpola’s numbering.
- Mahadevan’s ‘merchant of the city’ (2014), the phrase 255 435 690 740, is a real unit. It appears in 22 distinct texts, at three sites, after nine different signs. But its first half takes either ‘690 740’ or a bearer closer after it, rather than qualifying other nouns. And the ‘bearer’ signs he reads as names behave as closers: they are the main word of a name 1 time in 196.
- The Soviet team’s inflection claims, as Parpola reports them, fail as stated. 740 is no more often final on seals than on tablets, and a stroke after it rarely ends the text.
- What kind of writing. Compared at the same sample size (5,000 tokens), the Indus signs fall between Linear B’s signs and its words, nearer the signs: 495 types against 198 and 2,128, with the share of rare types and the predictability also between. That fits a logo-syllabic script of a few hundred signs, the class Fuls (2023) reaches from entropy. On repetitiveness measures the Indus signs resemble the Ur III seal legends instead, which is the mark of a short formulaic genre.
- ‘Too little repetition to be writing’ (Farmer, Sproat and Witzel 2004) fails on a control of the same genre. Indus seal lines of 5–7 signs repeat a sign in 9.5% of cases. Ur III seal legends of 5–7 words repeat a word in 13.4%. Both are far below shuffled text (28% and 34%), which is ordinary for real texts.
- Published keys. Yajnadevam’s Sanskrit key, Fairservis’s Dravidian values and Parpola’s 30 values were scored on how much of the corpus each key turns into dictionary words, against copies of the same key with its values shuffled. None reads its own language better than its shuffles. A key fitted blindly to the corpus reads about 93% of it as Sanskrit, Dravidian or Sumerian alike, so a reading rate on its own is no evidence.
09 What kind of names?
A seal text is short, sits under an animal and ends in a grammatical slot. Is it a personal name, a title or office, a god? Capture–recapture, the method ecologists use to estimate a population from two overlapping samples, gives a way to ask this without sound values. Take the names on Mohenjo-daro seals and the names on Harappa seals as two catches from one population. How many names appear in both tells how large the whole population was.
- Names are a large, open population; the main words are a closed one. 1,279 seals carry 585 distinct names, and only 20 appear in both cities. That puts the population at roughly 3,000–5,500 names (Lincoln–Petersen and Chao1 estimates), and 92% of names are on a single seal. The main words (the last sign of a name) are nearly all seen already: about 130 are known, and about 150–220 are estimated. The same holds inside one city, comparing two areas of Mohenjo-daro (7.4 times the observed number) or its early and late levels, and in Mahadevan’s separate transcription.
- The control on Linear B. The same test on Mycenaean Greek, with Knossos and Pylos as the two catches, separates the word classes that are known. Personal names give a large population (3.6 times the number observed, 18% shared, 71% on one tablet). Titles and trades give a small shared one (1.6 times, 50% shared). The comparison was redone at matched length, because longer units are more varied whatever they mean. At three signs, Indus names give 3.0 against 3.7 for three-syllable Linear B names; at four, 3.0 against 3.4. And real names recur between sites 25 times more often than random sign combinations of the same lengths would, as Linear B names do.
- The control on Ur III seals sharpens what the test can say. Sumerian personal names come from a limited stock borne by many people, so on their own they look like titles (2.0 times observed, 31% shared). Only the whole seal legend, name + title + father’s name, is as individual as an Indus name (17.5 times, 3% shared). So the test measures how individual a whole text is, not whether it names a person. Indus seal texts are individual compositions, like Linear B personal names and whole Mesopotamian seal legends. Unlike the Mesopotamian legends, they almost never carry a second name (1.6% of seals against 65%) or a separate title line (13% against 71%).
- A name is not tied to its seal’s animal. Seals with the same name carry the same animal no more often than any two seals (12% against 11%), and no sign goes with the elephant, the zebu or the cult standard beyond chance. Fifteen names occur at three or more sites. ‘590 390 740’ is at five: Allahdino, Chanhu-daro, Dholavira, Harappa and Mohenjo-daro.
10 Place and time
- Text does not follow find spot. The fuller corpus records where in Mohenjo-daro each object was found (area, block, house). What a seal says does not depend on where it was found. Seals with the same text are not significantly clustered, and seals from one house share no special vocabulary. Kind of text, ending, animal and text length do not vary by quarter, and seals found in streets are names as often as those found in houses.
- The grammar stands still. Between the early and late levels of both cities, the ending slot (the 520 share, the closers, 400 and 90, the stacking closers) and the count texts do not change. The numerals do, and at Mohenjo-daro the way names begin shifts. Long names that open with one of the 33 frequent opening signs rise from 39% to 66%, on square seals as well as rectangular ones. But the main words and the overall mix of signs at Mohenjo-daro do not move towards Harappa’s, so this is a local change in how names begin, not a merging of the two cities’ usage.
- Graphic variants follow the object and the maker. Two forms of the same sign agree on one object and across copies of one text more than chance, and depend on the medium. There is no trace of a household habit and no drift over time.
- Foreign names break the grammar. The seals from the Gulf and Mesopotamia use ordinary Indus signs but lack the ending slot, the name units and the usual openings. They score 6.5 bits per sign under the Indus model, and they use stroke signs heavily inside the line. That may hint at strokes used for their sound when spelling foreign names.
11 The language, as far as order shows it
Without sound values, only word order and the shape of the grammar speak to the language. Four features point one way:
- numerals come before what they count;
- names are head-final;
- the ending is a short class suffix fixed by the noun, with no number or class agreement elsewhere;
- the first element of a name comes from an open set, with no set of prefixes.
That fits Dravidian and Indo-Aryan, and counts against Sumerian and Elamite, which put the head first. Two further points narrow it. The endings split people from the fish names (Parpola’s stars), which Munda grammar would not do. And the fish = star word is Dravidian-only in South Asia. Together they lean to Dravidian, but only if the fish names are stars, and nothing here separates Dravidian from Indo-Aryan on order alone. The family expectations are standard typology, not a checked source for each language.
Name lists in the candidate languages add two results. Sumerian personal names from Ur III seals fix their first element (Ur-, Lu-, Nin-, the divine sign), while Indus names fix their last, as Linear B and Sanskrit names do. That rules out a Sumerian-type name structure a second time. Between Dravidian and Indo-Aryan the lists do not decide. Indus name lengths sit nearer Sanskrit names. The dominance of one ending (740 ends 82% of names) sits nearer Old Tamil personal names, 67% of which end in the masculine -aṉ, the suffix Mahadevan reads 740 as. A closer Indo-Aryan comparator then turned the argument around. Lüders’s list of early Brahmi inscriptions (1912) gives 374 Prakrit donor names from Sanchi and Bharhut, c. 200 BCE–400 CE. These are personal names on donated objects, near the Indus seals in time, place and genre. Indus names are built like them:
- they are about as open (5.1 against 5.9 times the number observed);
- they are nearer in length and in closure than the Old Tamil names;
- they end in a stock of compound heads. The ten commonest Prakrit heads (-rakhita, -guta, -dina, -mita …) cover 41% of names, and the ten commonest Indus heads cover 43%.
The Prakrit donor names also carry one near-universal ending, the genitive -sa of the gift formula. So a dominant ending like 740 fits an Indo-Aryan donor-style formula as well as Tamil -aṉ. The matching Dravidian test used Tamil names from donation records: 1,396 names from the English translations of the DHARMA Tamil inscriptions, 7th century onwards. It splits the evidence:
- Nearer the Prakrit donor names: the stock of compound heads (Indus 43%, Prakrit 35%, Tamil 16%) and how strongly the last element is fixed.
- Nearer the Tamil names: name length (almost identical) and the single dominant suffix (740 on 82% of names, Tamil -aṉ on 47%, the commonest Prakrit stem ending on 23%).
Name structure therefore rules out a Sumerian type but does not decide between Indo-Aryan and Dravidian. The Tamil records are also much later than the Prakrit list, and many of their names are Sanskrit-derived.
A contemporary Dravidian list was then built from the page scans of Mahadevan’s Early Tamil Epigraphy: 104 donor names from the Tamil-Brahmi cave inscriptions, c. 2nd century BCE – 4th century CE. That matches the Prakrit list in period and genre.
- Nearer the Prakrit names: all three structure measures. How fixed the last element is: Indus 0.83, Prakrit 0.90, Tamil-Brahmi 1.45. Length. And the stock of heads: 43%, 35% and 30%.
- Nearer the Tamil-Brahmi names: the dominant suffix (82% against 73% for -aṉ). But that contrast is partly made by the sources, because Lüders gives the Prakrit names without their case endings.
So on structure the Indus names look more like early Indo-Aryan donor names than contemporary Tamil ones. That is not a language identification. The Tamil-Brahmi list is small, syllables are not Indus signs, several Tamil-Brahmi donors bear Prakrit names themselves, and the Indus texts are some two thousand years older than either list.
12 What remains open
- No sign has a sound value that the evidence forces, and the language is not identified. Beyond the structure above, internal testing has run out. The last rounds returned mostly nulls and corrections. What is left needs new material: a bilingual inscription, the full ICIT corpus (4,537 objects, by account from Andreas Fuls), or digitised CISI volumes 2 and 3.
- The seven sign = image equations of the copper tablets (749, 341, 753, 777, the lens sign) are the firmest meanings. A reading in any language must fit them.
- Whether an Indus name packs a personal name and an office or lineage into one unit, as a Mesopotamian legend spreads them over lines, is open (see section 13).
13 What was withdrawn along the way
Registering predictions first means some of this project’s own readings were overturned by later tests. They are kept in the record, and listed here:
- A shortlist of sound signs (46 signs that combine freely with neighbours, like Linear B sound signs) failed all three of its registered tests. The related finding that rare names use freer signs appeared in Linear B too, so it is a general effect of rare words.
- ‘Two numeral systems’ (an early reading) became one number system in several stroke forms. Long strokes go with the measure sign 700 on the Harappa tablets.
- ‘Office + name’. A contrast between a title-like first sign and a person-like rest of the name turned out to compare single signs with multi-sign units. Single signs are a closed set almost by definition. Once length was controlled, only a weak link between the first sign and the seal’s animal remained.
- ‘Numeral compounds are local naming customs’. Linear B counts turned out to be just as local, so this was withdrawn to ‘local, as counts are’.
- ‘Mohenjo-daro converges on Harappa’. This was corrected to a local shift in how names begin (section 10).
- ‘Indus names are personal names’. This was narrowed by the Ur III control to ‘individual compositions’ (section 9).
14 Sources
- A. Parpola, “Study of the Indus Script”, Transactions of the International Conference of Eastern Studies 50 (2005), 28–66, harappa.com; Deciphering the Indus Script (Cambridge, 1994).
- J. P. Joshi and A. Parpola, Corpus of Indus Seals and Inscriptions 1 (Helsinki, 1987); its open digitisation, mayig/indus-valley-script-corpus (MIT).
- The ICIT-derived corpus of the indus-website project (github.com/yajnadevam/indus-website); A. Fuls, Corpus of Indus Inscriptions (Berlin, 2022) and A Catalog of Indus Signs (Berlin, 2023); A. Fuls, “Comparison of Linear Elamite and Indus Writing Systems”, Iranian Journal of Archaeological Studies 14.1 (2024), 3–16.
- I. Mahadevan, The Indus Script: Texts, Concordance and Tables (1977), machine-readable via github.com/joyboseroy/indus_decipher; I. Mahadevan, Dravidian Proof of the Indus Script via the Rig Veda: A Case Study (Bulletin of the Indus Research Centre 4, 2014).
- Language comparators: H. Lüders, A List of Brahmi Inscriptions (Epigraphia Indica X appendix, 1912; Internet Archive); DHARMA, Early Inscriptions of Āndhradeśa and the Tamil corpora (erc-dharma tfa-*, CC BY 4.0); Wikipedia, List of Sangam poets; Monier-Williams; I. Mahadevan, Early Tamil Epigraphy (Harvard Oriental Series 62, 2003), Appendix II, from the Internet Archive scan.
- Controls: DAMOS, Database of Mycenaean at Oslo (Linear B), with the Tiripode lexicon; ORACC epsd2/admin/ur3 (Ur III seal legends, CC0); Monier-Williams (Cologne Digital Sanskrit Lexicon); ePSD2; CLICS4; L. L. Merriam and A. Fuls, The Dravidian Database (Zenodo, 2026).
The scripts, the registered hypotheses (PREDICTIONS.md), the results of every set (results/) and the working notes (NOTES.md) are in the repository folder indus/.