LATEST VISIBLE BATCH
Performance of multimodal large language models in interpreting lateral cephalometric superimpositions: A comparative observer-performance study.
Observation date: 2026-07-28. The primary-source endpoint did not yield a verifiable new batch in this run; this issue uses the latest visible record from 2026-07-27, not a publication-date claim for the observation date.
Multimodal large language models (LLMs) can generate free-text interpretations of clinical images, but their performance on orthodontic cephalometric superimpositions is unknown. This study compared zero-shot interpretations from three LLMs with those of a second-year orthodontic resident. Ninety lateral cephalometric superimposition images from a private orthodontic practice were analyzed, including 30 nongrowing, 30 growing, and 30 orthognathic cases. Each image included overall maxillary regional, and mandibular regional superimpositions. ChatGPT 5.5, Gemini 3.1 Pro, Claude Opus 4.8, and the resident interpreted the same images using the same prompt, with no case context provided. Two senior orthodontists scored each interpretation against adjudicated reference interpretations using a 16-item rubric, yielding total scores from 0 to 32 and four domain scores. Friedman tests compared methods; Wilcoxon signed-rank tests with Holm adjustment compared each LLM with the resident. Total scores differed significantly among methods (Friedman P<0.001; Kendall W=0.592). Median total scores were 30.5 (interquartile range [IQR]: 28-32) for the resident, 17 (IQR: 14-22) for ChatGPT, 12 (IQR: 9-16) for Gemini, and 9.5 (IQR: 4.25-14) for Claude. All LLMs scored significantly lower than the resident overall, by domain, and within each case type (adjusted P<0.001). In zero-shot cephalometric superimposition interpretation, all tested multimodal LLMs performed substantially below a second-year orthodontic resident. These models should not be used as stand-alone interpreters without expert review.
Evidence boundary: 单中心、无病例背景的观察不能外推为全部影像判读或临床决策能力。
Community pulse · verified public sources
Social Network Radar
2026-07-28—2026-07-28. Items come from public APIs and pass original-link and Beijing-time window checks. They indicate public attention and discussion, not paper quality, clinical effectiveness, or scientific consensus.
No social item passed the quality gate in this window; the system does not generate filler.