Membership or Understanding? What the Imitation Game Measures When the Respondent Is a Language Model
Abstract
Benchmarks measure whether a language model produces the right output. That is areasonable test of factual knowledge, but expertise is not mainly a stock of facts. Itis the ability to move fluently inside a working community. This review sets out analternative drawn from the sociology of expertise. Collins and Evans separate the abilityto talk about a practice from the ability to perform it, calling the first interactionalexpertise, and they test for it with the imitation game, in which practitioners themselvesdecide who is one of them. The review then asks what that test actually picks up.When Collins passed as a gravitational wave physicist, his protocol banned personalquestions and edited markers of style out of the dialogues. Two of the three cues thatexposed a language model in a recent Turing test with rock climbers fall inside thoseexclusions: its prose was too polished, and its choice of a climbing hero too famous.The game may therefore be picking up social membership rather than understanding, aconcern Plaisance and Kennedy raised about human participants. Two things follow forhow such studies should be designed. They should say which of the two they are testing,since neither choice is wrong and it is the silence that causes trouble. And they shouldrecord why judges decided as they did, not only what they decided. An outcome tellsyou someone passed. Only the reasoning tells you what the judges were responding to.
// Source
Authors: Soumya Roy