ribaoo Research Daily

Primary papers first · explicit date boundaries · study-design-aware claims

27 July 2026 Medical AI | 27 July primary-source batch
No. 2026-07-27 · date-scoped primary-source batch

PRIMARY SOURCE

Performance of multimodal large language models in interpreting lateral cephalometric superimpositions: A comparative observer-performance study.

Observation and source date: 2026-07-27. This issue uses PubMed Date-Publication records; the same-day arXiv submitted-date window returned no relevant records.

Multimodal large language models (LLMs) can generate free-text interpretations of clinical images, but their performance on orthodontic cephalometric superimpositions is unknown. This study compared zero-shot interpretations from three LLMs with those of a second-year orthodontic resident. Ninety lateral cephalometric superimposition images from a private orthodontic practice were analyzed, including 30 nongrowing, 30 growing, and 30 orthognathic cases. Each image included overall maxillary regional, and mandibular regional superimpositions. ChatGPT 5.5, Gemini 3.1 Pro, Claude Opus 4.8, and the resident interpreted the same images using the same prompt, with no case context provided. Two senior orthodontists scored each interpretation against adjudicated reference interpretations using a 16-item rubric, yielding total scores from 0 to 32 and four domain scores. Friedman tests compared methods; Wilcoxon signed-rank tests with Holm adjustment compared each LLM with the resident. Total scores differed significantly among methods (Friedman P<0.001; Kendall W=0.592). Median total scores were 30.5 (interquartile range [IQR]: 28-32) for the resident, 17 (IQR: 14-22) for ChatGPT, 12 (IQR: 9-16) for Gemini, and 9.5 (IQR: 4.25-14) for Claude. All LLMs scored significantly lower than the resident overall, by domain, and within each case type (adjusted P<0.001). In zero-shot cephalometric superimposition interpretation, all tested multimodal LLMs performed substantially below a second-year orthodontic resident. These models should not be used as stand-alone interpreters without expert review.

Evidence boundary: 来自一家私人正畸诊所、且没有病例背景;不能把该结果外推为所有影像或临床决策性能。

Observation

2026-07-27

Asia/Shanghai

Routine batch

10

PubMed Date-Publication

Selected

3

primary-source and boundary screen

Same-day arXiv

0

no older batch relabelled