Health & Medicinearticle2026-08-25

Guideline-Based Accuracy Scores and Educational Quality of Artificial Intelligence-Based Language Model Responses to Search-Result-Derived Korean Pediatric Dental Trauma Queries: A Two-Time-Point Exploratory Study

Open access0 citations

Abstract

This exploratory study described guideline-based accuracy scores, educational quality, and temporal variation in ChatGPT and Gemini responses to eight searchresult-derived Korean pediatric dental trauma queries. Four tooth-avulsion and four crown-fracture prompts were submitted once to each service in January 2026 and one month later. The queries were treated as a convenience sample of recurrent searchresult content, not a representative set of caregiver questions. Two pediatric dentists scored 32 responses using query-specific key points, predefined major errors, and the Global Quality Score. Inter-rater quadratic weighted kappa was 0.910 for the guideline-based accuracy score and 0.835 for the Global Quality Score. ChatGPT accuracy scores were unchanged in five prompts, increased in one, and decreased in two; Global Quality Score values were unchanged in six and decreased in two. Gemini accuracy scores were unchanged in four, increased in one, and decreased in three; all Global Quality Score values were unchanged. No predefined major errors were identified, but avulsion-related terminology appeared in three of four crown-fragment reattachment responses. Within the sampled queries, clinically important omissions, injury-concept mixing, and temporal variation were observed. The findings do not represent caregiver information needs or establish real-world educational effectiveness, repeatability, or service superiority.

// Source

View paper (DOI)Open access versionOpenAlexTHE JOURNAL OF THE KOREAN ACADEMY OF PEDTATRIC DENTISTRYPublished 2026-08-25

Authors: Doyeon Won, Seonmi Kim, Hyewon Shin

Institutions: Chonnam National University, Yang Hospital