← alle innlegg

Thayer's Greek Lexicon, rebuilt from the 1889 page scans — verified entry by entry against the page

Short answer: Thayer's Greek–English Lexicon (1889 Corrected Edition) is public domain, but every open digital copy of it is lossy. Most descend from one defective transcription; the ones that do not are lossy in their own ways. Publifye rebuilt it from the page scans: 716 pages, 5,870 entries, transcribed twice and collated at 98.91% agreement across 856,793 tokens, contested spots settled on the page images, and a second, unrelated model family run across it as an independent witness. It is live in Darash.

Thayer's Greek–English Lexicon has been public domain for over a century, and we could not find a first-class digital text of it anywhere in the open. We say it that way deliberately — it is a claim about a documented search, not a proof of absence. We swept every natural habitat before starting — the Bible-app ecosystem, commercial Bible software, the open-data community on GitHub, continental scholarship, Perseus, the research libraries, Wikisource — and recorded why each candidate failed.

The failures have a fingerprint. We test every candidate at one entry, βάπτισμα, against the printed page, which reads …Acts i. 22; x. 37; xviii. 25; [xix. 3]…. Every open, downloadable Thayer dataset we tested renders it Acts 18:25; (); — the bracketed reference dropped, its punctuation left orphaned. That mistake was made once, decades ago, and the datasets carrying it forward are the ones people download today. It is not a formatting quibble: Thayer uses square brackets to mark his own additions to Grimm, so a text that drops them erases the line between the two men's work — one open dataset we measured keeps 726 brackets where the page has 17,540.

The copies served on the web are not one family either, and it is worth saying so plainly rather than letting the tidier story stand: Blue Letter Bible keeps that bracket. It loses other things — Thayer's sq., and his terminal asterisk, the mark by which he tells you every New Testament occurrence of the word is cited in the entry. No digital witness we tested preserves everything, which is exactly why collating them cannot reconstruct the page. And several datasets are not Thayer at all but gloss-sets and outlines wearing his name. In machine-readable form, the only complete Thayer that existed was a scan of ink on paper: the Harper & Brothers 1889 Corrected Edition, 716 body pages of dense, two-column Greek and English.

So we rebuilt it from the page scans. All of it. Here is how, because the how is the reason you can trust the result — and it is the same discipline behind everything Darash serves.

The model kept remembering the wrong Thayer

Transcription started the obvious way: show a vision model a page, get the text back. It failed in an unexpected and telling way — the model's memorisation guard kept firing. The widely-replicated defective digitisation of Thayer is all over every model's training data, so when the model saw a real 1889 page, it recognised something close enough to a memorised text and refused to transcribe, to avoid reciting.

Think about what that means: the corrupted version of this book is so pervasive that an AI cannot look at the original without being haunted by the copy. That is the problem in one image. Suggestive rather than proven — but the refusals cluster inside κύριος and χάρις, among the most-quoted entries in the book and so the most copied.

The escape was tiling: cut each column into as many as sixteen pieces, because a shorter span matches a memorised one less often. Four transcription routes in escalating order — whole page, then column, then tiles, then a second, independent model family entirely — each page taking the cheapest route that worked. Pages no engine could deliver were transcribed by hand, each hand file carrying the exact crop command that produced the view, so a reviewer can regenerate precisely what the transcriber saw.

Two passes, a third witness, and judges paid to break it

Every page was then transcribed twice by different routes — whole page against column-by-column — and the two passes collated: 98.91% agreement across 856,793 tokens over 685 pages. Both passes are the same model on a different prompt shape, and we are careful about what that buys: it is corroboration, not independence. Two runs of one model can be confidently wrong together, and one of them was — it wrote [where where the page reads [here, because that scans better as English. So a genuinely unrelated model family was run as a third witness; it smooths toward different mistakes, and each one's smoothing shows up as a disagreement. Agreement is not truth, but disagreement is a finding — the collation cut the review queue from 716 pages to 16 pages plus a ranked list of individual spots.

Then came the part most digitisation projects skip. No single model did this job, and none of them was asked to mark its own work. The pipeline ran as a set of separate agents with conflicting jobs: one model family transcribing, an unrelated one reading the same pages as a witness, executors applying repairs, a production QA agent that fetched all 5,870 entries off the live service and audited what was actually being served, and five rounds of adversarial judges — instructed to break the claims, not confirm them, and to end with a declaration of where they had not looked or have their verdict discarded. One refused to accept our own derived data as evidence that no dictionary entries had been lost, on the correct grounds that you cannot audit a thing with an artifact built from it. It went instead to three signals the transcription never touched — the printed running heads, an alphabetical-gap test whose own recall it measured first, and a third model as an independent reader — and could not find a lost entry by any of them.

The judges found real defects. The best one: a page where the scanner's ink read ἐνπ- for ἐντ-, which caused a matching cursor to jump — and seven entries silently vanished into their neighbour, one short article swelling to 8,295 characters. A summary line had reported this as out of order: 33, which reads like a sorting nit. It was data loss. A count is not a check — that sentence is now written on the project.

The governing rule

One rule governed every repair: never write a character that is not on the page. A fix was allowed only where the mapping had no judgement in it, or where the correct reading was taken off the page image by eye and recorded as such. Everything else stayed flagged. There is one digit in a Josephus reference that no eye could resolve on the scan — it is preserved as "?", because an honest question mark outranks a confident guess.

Nothing was spared on this. Pages no engine would return were cut into as many as sixteen pieces and then read by hand. A second physical copy of the same stereotype plates was rendered at 600dpi so that four contested digits could be settled by comparing the inking on two separate printings — because on identical plates a glyph difference can only be damage, never a variant. Judges were run until they stopped finding things, and then one more time. And the reason the paragraph above ends in a question mark rather than a digit is that the effort was spent on being right, not on looking finished.

So not 100%, and we will not say it is. When the operator asked exactly that — is it 100%? — the QA agent auditing the live service answered no, and listed six defects. They were fixed. The residue that remains is measured and published rather than rounded away: 0.19% token disagreement on the pages read twice, one Josephus digit still "?", and a verification pass that is thorough but not uniform. A corpus that claims perfection is telling you it has stopped looking.

Finally, every one of the 5,870 entries was read in full against the transcription of its own pages — giant entries made whole, block scrambles on two pages relocated to where the printer actually put them — until the invariants held: zero duplicated entries, zero empty ones, zero silent drops, text conservation proven end to end, every entry provably a verbatim, in-order, non-overlapping slice of the pages it was cut from, and every entry carrying the sha256 of its own printed text so one definition can be checked against one page.

What that is not is 716 pages collated line by line against the images by eye, and the distinction matters enough to state: eyes went to the disagreements, to the spots the judges sent us back to, and to sampled pages — not to every line of the book. On the thirteen pages the scanner happened to shoot twice, and which were therefore transcribed independently twice, Greek-bearing token disagreement runs at 0.19% — about two per page, mostly accents, iota subscript and συν-/συγ- assimilation. That is the honest error bar.

Why a Bible company does this

Not for the love of digital humanities. We had to do this simply to be able to provide Thayer to people at all — we found no first-class text to license, buy, or borrow. When what the open ecosystem offers is lossy texts and gloss-sets wearing the name, the only way to hand readers the real book is to make it exist.

And because trust in a research engine is not a marketing claim — it is a property of the pipeline that filled it. When your AI assistant reads the Greek with you through Darash and cites Thayer, the citation has to be Thayer — the 1889 text, not a gloss-set wearing his name. As far as our documented source sweep found, this is the only first-class digital Thayer in the open — and if someone shows us a better one, that is a good day, not a bad one. It went into Darash on 2 September 2026, held back until every entry had survived verification, because a Thayer that isn't Thayer helps nobody.

The same standard runs through the rest of the estate — the lexicons, the morphology, the 4,200+ published editions — and it is the standard we bring when we build for churches, too. Rigor is not a feature we added. It is the company.

By the numbers

Source Harper & Brothers 1889 Corrected Edition, page scans
Body pages transcribed 716
Entries in the finished lexicon 5,870
Transcription passes collated 2 (same model, different prompt shape) + a third from an unrelated model family
Token-level agreement between passes 98.91% of 856,793 tokens over 685 pages
Measured token disagreement, pages transcribed independently twice 0.19%
Transcription routes 4 (page, column, tiles, second model family) + hand
Silent entry losses found by adversarial review 7 (one page, restored)
Adversarial judging rounds 5, plus a live-service QA audit
Fabrications caught by review and removed 13 (one invented letter, twelve fabricated asterisks)
Characters added that are not on the page, in the shipped text 0

Frequently asked

Is Thayer's Greek Lexicon public domain?

Yes. Joseph Henry Thayer's Greek–English Lexicon of the New Testament was first published by Harper in 1887 — its preface is dated 1886, which is where the commonly cited date comes from — and corrected in 1889. Both are long out of copyright. The problem is not rights but the quality of the digital copies.

Why are the free online versions of Thayer unreliable?

Most of the downloadable ones descend from a single early digitisation with provable shared defects (dropped bracketed references, orphaned punctuation, merged entries). The ones that do not are lossy in other ways — dropped abbreviations, dropped terminal asterisks — and none of the copies we tested preserves everything the printed page carries. Several are abridgements or gloss lists rather than Thayer's text at all.

Where can I use the rebuilt Thayer?

Inside Darash, alongside BDB, LSJ and Abbott-Smith, through any MCP-capable AI assistant such as Claude or Cursor. darash.publifye.pro offers a free trial on the full corpus.

How was accuracy verified?

Two transcription passes by different routes collated token by token, a third pass from an unrelated model family as an independent witness, adversarial reviewers instructed to break the no-lost-entries claim, running-head and alphabetical-gap tests sharing no machinery with the transcription, and a final read of every entry — with the contested spots settled on the page image itself.


Read the Greek with the real Thayer: connect your AI assistant to Darash — free trial, nothing to install.

Written by Jørn André Halseth, founder of Publifye AS — ten years publishing Bibles and building the presses they run on, publisher of 4,200+ Bible editions, and the author of Darash, Junifye, Lexifye and Tringine. About · Contact