Academic papers

Papers brief: Korean journal abstracts show a late-2024 LLM register shift

arXiv study of 398,296 KCI abstracts finds morphology-aware excess vocabulary rising from late 2024 — sisahada-style hedging up, plain verbs down, with lower-bound LLM prevalence estimates through 2026.

  • academic papers
  • Korean NLP
  • LLM detection

Source: arXiv

Paper

An LLM-Associated Register Shift in Korean Journal Abstracts: A Morphology-Aware Excess-Vocabulary Study, 2018-2026 — Aron Lee (INTFRAME Research; arXiv 2609.07447, Sep 2026).

What it claims

English scholarship already tracks excess vocabulary — words above their pre-2023 trend — as a post-2022 signal. Lee adapts that to Korean with morphological units on 398,296 KCI abstracts (2018–August 2026), plus 47,165 Vietnamese abstracts for comparison.

Placebo floors stay small (0.1–2.2 single-word; at most 2.9 split-half). Korean abstracts show nothing in 2023, onset in late 2024, a rise through 2025, then flattening mid-2026. Sisahada (“suggest”) hits 21.4% of 2026 abstracts vs 5.3% expected; plain verbs like araboda (“look into”) fall to roughly a quarter of trend.

Single-word conditional lower bounds on LLM-processed abstracts are 3.5%, 10.5%, and 16.1% for 2024–2026; split-half bounds are 7.8%, 20.6%, and 33.0%. Holzwarth et al.’s estimator gives 41.9% and 72.1% for 2025–2026 — wider, same direction.

Controls shrink but do not erase the signal. Style-only lemmas leave 14.7 of 33.0 points; journal-pair matching leaves 34.1. Translation routes do not explain the rise. In paired English abstracts the excess appears a year earlier; where English shows none, Korean persists at 30–66% of the rate where English carries it. Control abstracts from three providers reproduce the rising words.

The breakdown

This is corpus forensics on Korea’s national journal index, not a product demo. KCI abstracts sit where Korean researchers meet grant panels and overseas readers who never open the PDF. Morphology-aware units matter because Korean hedging verbs conjugate — raw string counts miss part of the register move.

The timeline is the Korea headline: no 2023 blip, then late-2024 onset lagging English in the same articles by about a year — consistent with Korean-side polishing arriving after English-first drafting workflows.

Why readers outside the lab should care

If you cite Korean science from English-only workflows — systematic reviews, policy memos, VC diligence — you may read abstracts whose register shifted after 2024 even when the English side looks older. A fluent English abstract plus a suddenly hedged Korean sibling is a new integrity edge case for bilingual teams, not just a plagiarism-scanner problem.

Editors and funders already watching English preprint bots should extend suspicion to KCI metadata and Korean grant summaries. The paper shows population-level vocabulary turnover control generations can mimic — integrity policy needs corpus baselines, not keyword blocklists.

What builders and Korea-touching teams should watch

  • Do treat late-2024+ KCI abstracts as a distinct register stratum in Korean literature RAG.
  • Don’t assume English abstract timestamps proxy Korean-side cleanliness in the same record.
  • Expect hedging lemmas like sisahada to over-index and plain discovery verbs like araboda to under-index vs pre-2024 trend.
  • Re-check integrity playbooks: bounds run ~3–33% method-dependent — policy needs ranges, not one headline percentage.

Context

Read this as a morphology-aware alarm on Korean scholarly register, not proof any single abstract is machine-written. Korelay’s frame: Korea’s national abstract archive acquired an LLM-adjacent voice after English moved first — and the gap between paired languages is now measurable.

Source

Primary: arXiv:2609.07447 (abstract and framing cited; open the OA PDF for methods and full results). Do not republish the PDF.