AI-Driven Decipherment: Apollo, the First Foundation LLM for an Ancient Language

Updated

AI-Driven Decipherment: Apollo, the First Foundation LLM for an Ancient Language

Update (2026-09-28): the pipeline gets its first purpose-built foundation model — and a corporate patron. On September 23, 2026, the Austrian Academy of Sciences (ÖAW) unveiled Apollo, "the world's first foundation Large Language Model (LLM) for Ancient Greek," developed with European AI companies Mistral AI and SAIL Reply in cooperation with David Smith (Northeastern University) — the same Northeastern orbit that produced the Vesuvius virtual-unwrapping work. The partners funded it with roughly €400,000; the ÖAW owns it and it is free and open-access (apollo.vbc.ac.at).

The numbers that matter:

  • Trained on only ~600 million words of Ancient Greek (vs ~40 billion words typical for LLMs), yet achieves "an accuracy of around 80 percent for masked—that is, deliberately hidden—passages" on papyrus (Anna Dolganov, ÖAW).
  • In a blind study, international experts "rated Apollo's restorations at least as good as human-generated completions in 77 percent of cases, and in 16–20 percent of cases even better."
  • Launch demos: a Roman-era Greek birth certificate (149 CE, Austrian National Library) fully reconstructed; "key passages in a philosophical text carbonized during the eruption of Vesuvius were made fully legible"; and a Black Sea inscription whose restoration "provides evidence that Roman law was in effect in that region."
  • Context for the market size: "To this day, 92 percent of papyri remain undeciphered."
  • Roadmap: extend Apollo to decipher previously unread papyri, add thematic/semantic search across the whole Greek corpus, and build a Latin model (~12 billion words available).

Same week, the physical-pipeline side advanced too. C&EN (Sep 24, 2026) covered the UC Berkeley + NIST PLOS ONE study validating a two-step triage for the Herculaneum scrolls: screen each scroll with a handheld XRF instrument for lead in the ink (lead absorbs X-rays far more strongly than carbon-based papyrus), then send lead-positive scrolls to X-ray tomography + virtual unrolling. On kiln-carbonized model scrolls, "The handheld XRF instrument returned a lead signal from every sample" down to 25 µg/cm² — below the lowest concentration found in ancient fragments. Vito Mocella (CNR Italy, not involved): "This study can only be seen as a first step, but it is a good workflow." Limits: the method is blind to lead-free inks, and Mocella's own XRF work has found metals other than lead in Herculaneum inks — "I would try to look for these other metals in the scans as well."

What it means. Two shifts in one week: (1) the model paradigm moves from task-specific restorers (Ithaca, Aeneas) to a foundation model with register awareness — Apollo can tell Homeric Greek from documentary Greek and restore gaps in the right register; (2) the decisive funder of decipherment is now corporate AI patronage, not a humanities grant line — a frontier lab bankrolled a national academy's flagship tool for the price of a rounding error, with open-access output. The Vesuvius threads converge: Apollo demonstrated on a carbonized scroll, while the XRF triage determines which physical scrolls can economically enter the CT pipeline. Funding implications tracked in Commercial CRM and Funding: Corporate AI Patronage Arrives — a National Academy's Greek Model for €400K; the open question of how many scrolls test lead-positive and whether Italy permits systematic screening remains open.

Part of

This finding is an example of a pattern recurring across your work:

Revision history

  • Update with Apollo launch (Sep 23) and C&EN XRF triage validation (Sep 24)
    · by the agent
  • Weekly update: PLOS ONE lead-ink replica-scroll method (Sept 16) and npj Heritage Science ENCI portable CT for sealed cuneiform (Sept 2026) — decipherment enters a screening/portability phase.
    · by the agent
  • Weekly update: PLOS ONE lead-ink replica-scroll method (Sept 16) and npj Heritage Science ENCI portable CT for sealed cuneiform (Sept 2026) — decipherment enters a screening/portability phase.
    · by the agent
  • Updated without a stated reason.
    · by the agent
  • Updated without a stated reason.
    · by the agent
  • Update the Vesuvius Challenge finding with the historic June 25, 2026, end-to-end reading of PHerc. 1667, the Philodemus attribution, the BM18 ESRF 3D ink breakthrough, and the newly launched $1M 2027 automation prize.
    · by the agent
  • Update the Vesuvius Challenge and text decipherment note with the June 2026 Scientific Reports breakthrough proving the morphological hypothesis via 3D optical profilometry.
    · by the agent
  • Update the AI text decipherment note with the historic June 2026 breakthrough where the Vesuvius Challenge successfully read an entire Herculaneum scroll (PHerc. 1667) end to end.
    · by the agent
  • Update the note with the June 2026 Vesuvius Challenge full-scroll reading, the July 2025 Google DeepMind Aeneas Latin translation tool, and LMU Munich's Fragmentarium project for cuneiform.
    · by the agent
  • Update with June 25, 2026 Vesuvius Challenge breakthrough on PHerc. 1667 and Google DeepMind's Aeneas model published in Nature.
    · by the agent
  • Updated without a stated reason.
    · by the agent