Society & Economicsarticle2026-08-09

Why Language Models Work: A Philosophical Investigation

Open access0 citations

Abstract

The people who build large language models do not know why they work. "For all its runaway success, nobody knows exactly how, or why, it works," reports the MIT Technology Review (Heaven, 2024), summarising the state of knowledge about deep learning. Boaz Barak, a computer scientist at Harvard seconded to OpenAI, likens the situation to physics at the start of the twentieth century, a pile of experimental results that surprise the experimenters and resist explanation. Anthropic, which built Claude, states plainly in its own research that it does not understand how its models do most of what they do. Hattie Zhou, an AI researcher at Montreal, recalls arriving in the field, asking her teachers why the systems work, and being told there were no good answers. The admission is genuine, and the response to it has been to treat the gap as an engineering puzzle awaiting more interpretability tools and more scale. This paper treats it as what it is, a question about language rather than about machines, and defends an answer the field does not want to hear: the model understands nothing, and it produces fluent text only because it harnesses a structure already present in human writing, the residue left in text by the millions of people who did understand. The empirical surprise that something so simple as next-token prediction can generate publishable prose is a fact about language, not about the machine. The achievement of the engineers in building such systems is real and not contested here. The question this paper takes up is narrower: why does the operation work at all? The case is built from the ground up, assuming no prior acquaintance with the philosophy of language and teaching each idea from ordinary examples before naming the thinker who found it. It converges on a single property of language that explains, at once, why fifty years of rule-based artificial intelligence came to nothing and why statistical approximation works: language is regular without being systematic, caught by patterns yet beyond the reach of any rules.

// Source

View paper (DOI)Open access versionOpenAlexKnowledge Commons (Lakehead University)Published 2026-08-09

Authors: Moreno Nourizadeh