Trial finds AI access raised physician scores on clinical cases
A three-country randomized study found different-sized gains when physicians could use GPT-4o, but real-world effects remain uncertain.
Moderate evidenceHuman studyInterpret with caution
Medical disclaimer: This article summarizes research findings and is for informational purposes only. It is not medical advice.
Editorial illustration — not from the study.
Researchers randomly assigned physicians either to a control group or to a group that could use GPT-4o while working through standardized clinical vignettes. Reported score gains were largest in Kenya, followed by Indonesia and the Netherlands, and the differences in scores between Kenya and the Netherlands became smaller.
The results also varied within countries. Score variation increased in Indonesia and decreased in Kenya, and greater use of the language model was associated with larger gains. Some physicians without access still scored higher than some physicians with access, indicating that performance was not uniform among users.
What the study looked at
The researchers examined whether access to a large language model, or LLM, affected physician performance and whether it changed differences between countries. In a parallel-group randomized controlled trial, 249 physicians from Indonesia, Kenya, and the Netherlands were assigned either to a control group or to an intervention group with access to GPT-4o. They completed standardized clinical vignettes; the abstract does not provide further details about the participants, tasks, or study duration.
What the trial found
The abstract reports higher clinical-performance scores among physicians with LLM access. The reported gain was 18% in Kenya, with a 95% confidence interval of 12.7% to 23.2%; 10.7% in Indonesia, with a 95% confidence interval of 5.7% to 15.7%; and 7.2% in the Netherlands, with a 95% confidence interval of 3.7% to 10.7%. All three results had p values below 0.001. Access was also reported to reduce performance differences between Kenya and the Netherlands. However, score patterns differed within countries, and some physicians without access outperformed physicians with access. Greater use was associated with greater gains, but this association does not by itself show why some people benefited more.
Who this is relevant to
These findings may be most relevant to physicians similar to those who took part and to tasks resembling the standardized clinical cases used in the trial. They do not directly establish effects for patients, other healthcare workers, other countries, different language models, or routine clinical practice. The study provides no basis for assuming that access would produce the same results in every setting or improve patient outcomes.
Why this matters
The findings suggest that access to an LLM may affect how physicians perform on written clinical cases, and that the size of the effect may differ across settings. They also raise the possibility that such tools could influence differences in performance between countries. However, the study measured scores on standardized vignettes rather than diagnoses, treatments, safety, or health outcomes for actual patients, so it does not establish that LLM access improves patient care or reduces health inequalities in practice.
Limitations & evidence assessment
The study included 249 physicians, so its results may not represent all physicians or healthcare systems. Standardized vignettes are not the same as caring for patients, and the abstract does not report whether the findings persist over time or affect patient outcomes. It is also unclear from the abstract how participants were recruited, how the tasks were scored, what instructions or safeguards were used, and whether GPT-4o's performance was assessed separately from the physicians' performance. The results concern access to a particular model in the study setting and may not apply to other models or clinical environments.
Why this evidence level: Randomized trial with a smaller detected sample (n≈249).
Evidence levels are editorial estimates derived from study metadata — they are not clinical appraisals.
// Source
npj Digital Medicine · 2026 · DOI: 10.1038/s41746-026-03111-5
A longitudinal study of 1,056 U.S. Catholic priests found that earlier flourishing and stronger support from fellow priests, bishops and lay networks were associated with less loneliness three years later. Psychological distress and burnout were associated with more loneliness, while private prayer frequency was not a significant predictor.
A study of 11 Iranian university students finds that unclear institutional rules around AI-mediated writing can leave students caught between pressure to produce polished work and expectations of independent authorship. Students described reduced ownership of their writing, secrecy, anxiety relief and moral uncertainty.
A focused review links prolonged, close-up smartphone use with reports of acute acquired comitant esotropia, a sudden inward turning of the eyes that can cause double vision. The evidence suggests a possible contribution in some people but does not establish that smartphone use causes the condition across the population.