Academic papers

Papers brief: ARI — when Joseon archives need more than the page in front of you

arXiv ARI pairs LLMs with retrieval-augmented generation to restore damaged Korean historical documents, especially proper nouns masked language models miss.

  • academic papers
  • Korean history
  • digital heritage

Source: arXiv

Paper

Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models — Gabeen Kim, Kyeongpil Kang (Kangwon National University; submitted Jul 2026)
ID: arXiv:2607.21936

What it claims

Historical manuscripts are knowledge archives, but fading ink, brittle fiber, and centuries of handling leave illegible gaps — blank squares in digitized Joseon-era corpora that OCR and human readers alike struggle to fill. Masked-language-model restorers work well on local context (the sentence around a hole) but stumble on named entities — people, places, dates, titles — where the answer lives in external history, not on the damaged line.

The authors propose ARI (Archive Restoration Intelligence): an LLM framework with retrieval-augmented generation (RAG) that blends a pre-trained model’s implicit web-scale knowledge with explicitly retrieved passages from related historical documents. On Korean historical corpora — including materials tied to UNESCO-listed archives such as the Annals of the Joseon Dynasty and the Journal of the Royal Secretariat — they report improvements over baselines on both random character masks and named-entity masks, plus expert assessments positioning ARI as a practical assist for domain specialists.

The breakdown

The failure mode is intuitive. A model reading only a damaged Joseon line that names a royal office or place — say a blank where a 승정원 (Royal Secretariat) official’s title should sit — may invent a fluent Hanja guess from local grammar; the right proper noun often lives in another court record, not neighboring brush strokes. Prior MLM-only restorers, per the abstract, never leave that single document — fine for ordinary characters, weak for context-dependent proper nouns.

ARI’s approach combines three moves. First, a causal LLM pretrained on broad text supplies implicit background. Second, RAG pulls related documents from the archive into the prompt so the model can draw on era-appropriate names and titles. Third, fine-tuning on ancient Korean corpora (Hanja-heavy Joseon records with metadata from the National Institute of Korean History) adapts the system to brush-transcription noise and real damage patterns. The abstract emphasizes evaluations on both random corruptions and named-entity-focused masks, where external knowledge should matter most.

Why readers outside the lab should care

Korea’s flagship historical databases — the ones overseas researchers bookmark for Joseon politics, diplomacy, and court ritual — can contain damaged tokens in digitized scans, as the authors describe. If you are an expat historian or policy analyst mining English portals built on those texts, every unfilled □ is a silent error in your timeline. ARI-class tooling is infrastructure for making UNESCO-grade archives more machine-readable with retrieval evidence meant to make proposed restorations checkable — the difference between “plausible Hanja” and a claim you can cross-check against cited related documents.

What travelers and expats should watch

  • Do treat restored characters in digital portals as hypotheses until human experts or parallel sources confirm — RAG reduces guesswork but does not certify truth.
  • Do prefer tools that show which related documents supported a fill-in; that provenance signal is the operational signature of external-knowledge methods like ARI.
  • Don’t assume English summaries of Joseon records are independent of fill-in-the-blank restoration — damaged-token rates in major corpora remain high per the authors’ framing.
  • Expect named entities (personal names, place names, office titles) to be the hard cases — the abstract’s core claim is gains precisely there.
  • Do check the arXiv page / accompanying materials for any code or model release before assuming a public toolkit ships with the paper.

Context

Read this as a Korean historical NLP engineering report, not as a museum exhibit or genealogy product launch. Korelay frame: the Annals of the Joseon Dynasty and Journal of the Royal Secretariat are globally famous archives; overseas readers encounter them through English gateways and citation databases long before they visit Seoul. ARI argues that unlocking those records requires retrieval outside the damaged leaf, not just smarter autocomplete on the damaged leaf itself.

Source

arXiv:2607.21936 — abstract and framing cited; open the OA PDF for ARI architecture, corpora statistics, and expert evaluations. Do not republish the PDF.