A review found strong results in laboratory images, but limited outside testing makes it unclear how widely these systems will work.
Researchers reviewed 4,323 records and included 77 studies of artificial intelligence for diagnosing human gastrointestinal parasitic infections. Helminths, such as worms, were the focus of 42 studies, while 28 examined protozoa and seven covered both groups.
The review found promising results in controlled settings, especially for helminth detection. However, only 17.9% of protozoan studies used external validation—testing on data from outside the development setting—and those studies showed performance declines of 15 to 30 percentage points.
What the review found
Artificial intelligence methods were more often studied for helminths than for gastrointestinal protozoa. Among externally validated helminth studies, reported sensitivity—the share of infections correctly detected—ranged from 85% to 98%, and several systems had higher sensitivity than expert microscopists.
For protozoa, internal testing often reported accuracy above 90%, but those results did not consistently hold up on outside data. Only 17.9% of protozoan studies performed external validation, and performance fell by 15 to 30 percentage points in the studies that did. Overall, 71.4% of datasets came from a single institution. Convolutional neural networks, a type of image-analysis system, were used in 66.2% of the studies.
Why wider testing matters
More than 3.5 billion people are affected by human parasitic gastrointestinal infections, while diagnosis still relies heavily on conventional microscopy. The review suggests that artificial intelligence could assist parasite detection, but results from carefully controlled images may not transfer reliably to other laboratories, populations or imaging conditions.
The authors argue that wider use will require shared imaging standards, more diverse open-access image collections and further development of machine-learning methods. The review also highlights protozoan diagnosis as an area where the evidence base and outside testing remain limited.
Evidence and limits
This study is a systematic review of 77 published studies, rather than a new test of one artificial intelligence system. Its findings are based on the performance measures reported by the included studies, which used different datasets and testing approaches.
The main limitation identified by the review is limited generalizability: 71.4% of datasets came from a single institution. Internal validation can make systems appear more accurate than they are on new data, and only 17.9% of protozoan studies included external validation. The reported 15-to-30-percentage-point decline in those studies shows why results from controlled settings should not automatically be treated as evidence of performance in routine clinical use.
// Source
Parasitology · 2026 · DOI: 10.1017/s0031182026102522
Authors: Jorge González, Sebastián Zambrano, Kurt Montoya, Víctor Torres González, Juan San Francisco, Bessy Gutiérrez, Camila Gutiérrez, Isidora Ahumada, José Luis Vega
Institutions: Universidad de Antofagasta, University of Concepción, San Sebastián University