A review of language models and safety certification finds a structural mismatch between the two. Certification regimes require behavior that can be bounded, traced and supported by evidence, while language models retain a nonzero error floor and can produce plausible but incorrect outputs.

The authors argue that common safeguards—including retrieval tools, guardrails, formal checks, uncertainty estimates and hybrid systems—can reduce risks but do not eliminate the gap. They identify two possible paths: set statistical acceptance criteria for narrowly bounded tasks, or place the language model in an untrusted proposer role inside a deterministic, independently verifiable control system.