Researchers examined 274 source texts, including recent and older news articles and historical articles. They created controlled versions that differed in factual accuracy and presentation, then tested how the language model judged them.

The model’s overall reliability score stayed relatively high as accuracy fell, and completely fabricated texts received mean scores above 3 on a 1–5 scale. When researchers separately tested factual correspondence and the credibility conveyed by the writing, factual judgments were more accurate, but performance varied substantially with how familiar the information was.