Hlaing NHMM, Park K, Wan Q, Lee SJ, Gallucci GO, Lee JH* (Corresponding author). Accuracy and hallucination behaviors of artificial intelligence-based large language models for dental implant identification on radiographs. J Dent. 2026 May 15:106754. Epub ahead of print.
Abstract
Objectives: This study aimed to evaluate the accuracy and error patterns of four artificial intelligence-based large language models (LLMs) in identifying dental implant characteristics including implant level, connection type, brand, and manufacturer, on radiographic images.
Methods: In total, 120 standardized radiographic images (60 periapical and 60 panoramic) representing six implant systems (Straumann BL, Straumann TL, Osstem TS [Hiossen ET], Osstem [Hiossen] US, Osstem [Hiossen] SS, and Dentium SuperLine) were retrospectively collected and analyzed. For each image, four multimodal LLMs (ChatGPT-5, ChatGPT-4o, Gemini 2.5 Pro, and Grok 4) were queried six times with a standardized prompt that requested the implant level, connection type, brand, and manufacturer. The primary outcome was response-level accuracy. Responses were classified as successful or unsuccessful identification; unsuccessful responses were further categorized as definitive incorrect (hallucination), tentative incorrect, ambiguous, or abstention. Statistical analysis used generalized estimating equations for accuracy and Chi-square tests for error types (α=.05).
Results: Structural characteristics (implant level and connection type) were identified with accuracies ranging from 51.67% to 82.5%, whereas brand and manufacturer recognition remained consistently poor (4.17%-32.78%). Hallucinations represented the predominant error type (>87% of errors in all models), while ChatGPT-5 showed a comparatively higher rate of abstention and lower hallucination frequency.
Conclusions: Current LLMs have limited capability for the definitive radiographic identification of dental implant systems, particularly for specific brands. Recent model iterations show a promising trend toward reduced hallucinations and increased reliability through abstention; however, their high error rates indicate that they can only be considered as supplementary tools requiring clinician verification.
Clinical significance: Current LLMs are not sufficiently reliable for definitive radiographic implant identification, and their outputs should be interpreted with clinician verification. However, the increasing tendency of recent models to acknowledge uncertainty indicates promising progress toward safer clinical decision support systems.
Keywords: Artificial intelligence; Computer-assisted diagnosis; Dental implants; Dental radiography; Hallucinations; Large language models.
