本文へ移動

論文 ·日本語 ·未確認

GPT-4 Turbo with Vision fails to outperform text-only GPT-4 Turbo in the Japan Diagnostic Radiology Board Examination

Yuichiro Hirano Shouhei Hanaoka Takahiro Nakao Soichiro Miki Tomohiro Kikuchi Yuta Nakamura Yukihiro Nomura T. Yoshikawa Osamu Abe

刊行年
2024-05-11
収録
『Japanese Journal of Radiology』 42(8) pp. 918-926
出版
Springer Science+Business Media
言語
英語
OpenAlex
W4396831262
DOI
10.1007/s11604-024-01561-z
PubMed
38733472
ISSN
1867-1071
URL
https://link.springer.com/content/pdf/10.1007/s11604-024-01561-z.pdf

要旨

PURPOSE: To assess the performance of GPT-4 Turbo with Vision (GPT-4TV), OpenAI's latest multimodal large language model, by comparing its ability to process both text and image inputs with that of the text-only GPT-4 Turbo (GPT-4 T) in the context of the Japan Diagnostic Radiology Board Examination (JDRBE). MATERIALS AND METHODS: The dataset comprised questions from JDRBE 2021 and 2023. A total of six board-certified diagnostic radiologists discussed the questions and provided ground-truth answers by consulting relevant literature as necessary. The following questions were excluded: those lacking associated images, those with no unanimous agreement on answers, and those including images rejected by the OpenAI application programming interface. The inputs for GPT-4TV included both text and images, whereas those for GPT-4 T were entirely text. Both models were deployed on the dataset, and their performance was compared using McNemar's exact test. The radiological credibility of the responses was assessed by two diagnostic radiologists through the assignment of legitimacy scores on a five-point Likert scale. These scores were subsequently used to compare model performance using Wilcoxon's signed-rank test. RESULTS: The dataset comprised 139 questions. GPT-4TV correctly answered 62 questions (45%), whereas GPT-4 T correctly answered 57 questions (41%). A statistical analysis found no significant performance difference between the two models (P = 0.44). The GPT-4TV responses received significantly lower legitimacy scores from both radiologists than the GPT-4 T responses. CONCLUSION: No significant enhancement in accuracy was observed when using GPT-4TV with image input compared with that of using text-only GPT-4 T for JDRBE questions.

主題

この書誌の出所

  • openalex— W4396831262(2026-08-14取得)

引用

Yuichiro Hirano・Shouhei Hanaoka・Takahiro Nakao・Soichiro Miki・Tomohiro Kikuchi・Yuta Nakamura・Yukihiro Nomura・T. Yoshikawa・Osamu Abe(2024-05-11) GPT-4 Turbo with Vision fails to outperform text-only GPT-4 Turbo in the Japan Diagnostic Radiology Board Examination 『Japanese Journal of Radiology』 42(8) pp. 918-926 Springer Science+Business Media

Hirano2024GPTTurboVision
書誌 67,320件 語別索引 17,251件 資源 113件 研究者 303名 JSON