Beyond Mere Words: Tracing Large-Language-Model Style in Korean Academic Writing through Excess Vocabulary. A study of 316,538 Korean journal abstracts, 2018–2026, with a Vietnamese comparison on 47,165 abstracts
Abstract
Excess vocabulary, a word's frequency above its pre-2023 trend, has measured how large language models changed English scholarly writing. We adapt it to Korean and Vietnamese with morphological units, applying it to 316,538 Korean abstracts from 2,258 KCI journals (2018 to August 2026) and, as an external comparison, 47,165 Vietnamese abstracts. Placebo runs on pre-ChatGPT years put the false-positive level at 0.1 to 2.3 points. Korean abstracts show nothing in 2023, an onset in late 2024 and a rise through 2025 flattening in mid-2026: 시사하다 "suggest" appears in 21.4% of 2026 abstracts against 5.2% expected, 단순하다 "mere" in 12.3% against 1.3%; plain verbs like 알아보다 "look into" fall to a quarter of trend. The single-word lower bound on LLM-processed abstracts (excess ratio at least 1.5) is 3.7% in 2024, 9.4% in 2025 and 16.2% in 2026 (14.1% density-normalised); a set bound chosen on half the journals and measured on the other half is 7.5%, 19.3% and 31.1%. Control abstracts from three providers reproduce the rising words; markers sort by model generation, not provider. The set indicator fires on 60–73% of model-drafted abstracts, 38–52% of model-rewritten and 28–35% of model-polished ones, so the floor lies far below the share: drafting puts it at 69–98% (propagated ranges from 51%); rewriting by the two API models reproduces the two headline words at 66–90% but not the broader 2026 vocabulary. Vietnamese shows the same words a year later, on a corpus too noisy for a bound. Code, tables and a public dictionary are released. Working paper, version 4 (26 August 2026). The version history is given in Appendix J of the paper. The attached reproducibility package contains the analysis code, the per-year document-frequency tables for Korean and Vietnamese, the generated positive-control abstracts (1,180 abstracts in thirteen conditions from OpenAI, Anthropic and EXAONE models) and the manuscript consistency gate. A public Korean AI-style dictionary and checker built from the same measurements are at https://os.intframe.com/report/ai-style-dictionary-ko and https://os.intframe.com/report/ai-style-check-ko.
// Source
Authors: Ahron Lee
Institutions: CellaMedic Biotechnology (South Korea)