ribaoo Research Daily

Weekend observation · primary papers first · source and observation dates kept distinct

25 July 2026 Medical AI | 25 July weekend observation primary-source batch
No. 2026-07-25 · weekend observation / latest visible primary-source batch

DAILY LEAD

Performance of Large Language Models for Oncology Nursing Decision Support: Cross-Sectional Study.

Weekend observation (25 July 2026, Asia/Shanghai): no usable routine new-paper batch was visible. This issue only summarizes the latest visible PubMed Date-Publication records from 2026-07-24 (24 July 2026); their source date is not a new-publication claim for the observation date.

Large language models (LLMs) are increasingly used in health care, with emerging applications in clinical decision support and nursing education. However, evidence on their performance in nursing contexts, particularly in oncology nursing, remains limited. Given the complexity and high-risk nature of oncology care, it is important to evaluate the performance and clinical relevance of LLM-generated responses in oncology nursing contexts. This study aimed to compare the performance of LLMs in oncology nursing decision support tasks using standardized examination questions and case-based clinical scenarios and explore LLMs' potential applicability and current limitations in oncology nursing practice. A total of 33 case-based questions derived from 10 oncology nursing clinical scenarios in a nationally used training manual, along with standardized examination-oriented questions from a commercially published preparation book for the Chinese Nursing (Intermediate) Qualification Examination, were used to evaluate the performance of 5 LLMs (DeepSeek, Qwen, Spark-Desk, WiseDiag, and ChatGPT). All models generated responses using a standardized prompt. Two oncology nurses with more than 5 years of clinical experience independently rated the case-based responses using 3 evaluation dimensions: correctness, clarity, and conciseness. Interrater reliability was assessed using the quadratic weighted Cohen κ, intraclass correlation coefficient, and Spearman rank correlation coefficient. Differences among models were analyzed using the Kruskal-Wallis test with the Dunn post hoc test. In addition, examination performance was evaluated based on total score, accuracy rate, and completion efficiency. Interrater reliability analyses indicated moderate agreement between evaluators. The median correctness, clarity, and conciseness scores were as follows: 11.50 (IQR 10.50-12.00) for DeepSeek, 11.00 (IQR 10.50-12.00) for Qwen, 10.50 (IQR 9.50-11.50) for Spark-Desk, 10.00 (IQR 9.50-11.50) for WiseDiag, and 10.00 (IQR 9.00-11.50) for ChatGPT. The Kruskal-Wallis test indicated statistically significant differences among models (H=11.416; P<.05), with post hoc analysis showing a significant difference only between DeepSeek and ChatGPT (P<.05). In examination-based tasks, all models achieved passing performance, with accuracy rates ranging from 77% (77/100) to 93% (93/100). In terms of response completion, DeepSeek and ChatGPT completed all tasks in a single interaction, whereas other models required multiple interactions due to output interruptions. LLMs showed relatively strong performance on structured knowledge and examination-based tasks but remained limited in complex oncology nursing scenarios requiring individualized assessment and dynamic clinical judgment. Their potential use may be most relevant to information retrieval, knowledge organization, and patient education. Because the correctness, clarity, and conciseness rubric showed only moderate interrater reliability, the case-based comparisons should be interpreted as preliminary signals rather than definitive evidence of between-model differences. LLM outputs should therefore be used as supportive information and interpreted alongside professional clinical judgment.

Evidence boundary: 横断面比较且案例评分者一致性为中等;不能将考试表现视为临床决策有效性。

Observation

2026-07-25

Asia/Shanghai · Saturday

Routine batch

not visible

PubMed E-utilities server 500; no old record relabelled

Latest visible

30

2026-07-24 PubMed Date-Publication

Selected

3

latest visible primary sources