Researchers report strong performance from an AI model for stomach biopsy images
The model was tested on biopsy images from thousands of patients, but its effect on clinical decisions and patient outcomes remains unknown.
Moderate evidenceHuman studySome caution advised
Medical disclaimer: This article summarizes research findings and is for informational purposes only. It is not medical advice.
Editorial illustration — not from the study.
Researchers developed an artificial intelligence model that analyzes whole-slide images from gastric biopsies. The model was designed to classify biopsy findings into six categories and included versions intended to distinguish early from advanced stomach cancer and to estimate the likelihood of lymph-node spread before treatment.
The model was developed using images from more than 17,000 patients at six centers and evaluated in external datasets and a prospective group of nearly 3,000 patients. In testing, it showed high sensitivity and specificity, and it improved the accuracy and speed of nine pathologists in an auxiliary experiment. These results describe performance in the study settings; they do not by themselves show improved outcomes for patients.
What was studied
Researchers retrospectively collected 20,711 whole-slide pathology images from 17,086 patients across six centers to develop and test a gastric biopsy artificial intelligence model. They also prospectively enrolled 3,698 images from 2,965 patients for validation. The model performed six-category classification on gastric biopsy images, and separately fine-tuned versions assessed whether images could distinguish early-stage from advanced-stage stomach cancer and predict lymph-node metastasis. Nine pathologists also took part in an auxiliary experiment evaluating performance with and without model assistance.
What researchers observed
The six-class model, called GBAIM, had 96.1% sensitivity and 95.0% specificity in external cohorts, and 93.4% sensitivity and 99.0% specificity in the prospective cohort. In auxiliary testing, it increased the accuracy of all nine participating pathologists by 1.7% to 39.0% and reduced their diagnostic time by 29.4% to 50.5%.
A version designed to distinguish early-stage from advanced-stage stomach cancer had an area under the curve of 0.907 in an internal test set and 0.826 in an external test set. A version intended to predict lymph-node metastasis had areas under the curve of 0.814 and 0.706 in the corresponding internal and external test sets. These are measures of model performance, not proof that the system improves diagnosis or treatment for patients in routine care.
Who this may apply to
The findings may be most relevant to biopsy images and patient populations similar to those represented across the six study centers. They do not establish how the model would perform in every hospital, population, scanner system, or clinical workflow, and the study did not show that using the model improves patient outcomes.
Why this matters
Pathology review of gastric biopsies is important for identifying stomach disease and helping characterize cancer. A tool that performs consistently and supports pathologists could potentially help with workload and interpretation, particularly where biopsy-image review is demanding. However, this study measured image-classification performance and performance in an auxiliary pathologist experiment; it did not show that the model improves patient outcomes, reduces missed diagnoses in routine care, or is ready for use in all hospitals. Further evaluation in clinical workflows would be needed to establish those questions without assuming that the reported associations and performance measures will translate directly to better care.
Limitations & evidence assessment
The study used retrospectively collected images to develop the model, which can introduce selection and information biases, although a separate prospective cohort was also assessed. The abstract does not describe how patients were selected, how the six diagnostic classes were balanced, or whether performance varied across centers, scanners, or patient groups. It reports diagnostic accuracy and pathologist assistance, but not whether use of the model changes treatment decisions, complications, survival, or other patient outcomes. The available information is also insufficient to determine how the model compares with existing clinical workflows outside the study settings.
Why this evidence level: Cohort study on humans; observational but structurally stronger than cross-sectional designs.
Evidence levels are editorial estimates derived from study metadata — they are not clinical appraisals.
// Source
npj Digital Medicine · 2026 · DOI: 10.1038/s41746-026-03102-6
A survey of human-AI collaboration argues that more capable AI systems do not automatically produce reliable partnerships. It identifies human-centered design, governance and attention to high-stakes uses as central to making these systems trustworthy and beneficial.
Deep-learning tools for predicting RNA splicing can rely on genomic clues unrelated to splicing and miss important effects of RNA structure, a study finds. These weaknesses were linked to systematic errors, especially when the tools evaluated sequences that differed from those used for training.
AI vulnerability detectors often perform less well on software projects outside their training data. A study finds that cleaner, more varied training data and an encoder-based design improved detection, including a reported 6.8% gain in recall on the BigVul benchmark.