Volume 4,Issue 5
A Comparative Study on the Multidimensional Performance of Large Language Models in Long-text Medical Education in Hematology
Background: As large language models (LLMs) surpass the million-token context window, their potential in medical long-text education has become increasingly evident. However, systematic evaluation of performance differences among various models in hematology long-text teaching scenarios remains lacking. It is an urgent need to introduce large language models to revolutionize education. Objective: To compare the comprehensive performance of large language models in long-text teaching tasks in Hematology. Methods: Four categories of long-text teaching tasks (n = 24) were constructed, including instructional guidelines, progressive analysis of long cases, multiple rounds of Socratic questioning, and cross-document knowledge integration. Representative LLMs (GPT-5.5 and Kimi K2.6 as examples) were selected to complete these tasks. Nine hematology teaching invited for blind scoring to evaluate six dimensions: medical accuracy, teaching logic, long-text information integration ability, language adaptability, educational inspiration, and response completeness. Each dimension was scored from 0 to 20, with a total score ranging from 0 to 120. Long-text specific metrics, including information omission rate and cross-document consistency error were also quantified. Results: No significant difference was found in overall scores between the two models (100.01 ± 1.45 vs. 100.30 ± 1.52, P = 0.623), demonstrating a pattern of "convergent total scores with divergent dimensional performance." Dimensional analysis showed that GPT-5.5 outperformed Kimi K2.6 in medical accuracy (17.11 ± 0.32 vs. 16.89 ± 0.28, P = 0.089), linguistic appropriateness (16.71 ± 0.29 vs. 16.15 ± 0.41, P = 0.018), and educational heuristics (17.25 ± 0.31 vs. 15.58 ± 0.22, P < 0.001), whereas Kimi K2.6 performed better in pedagogical logic (17.05 ± 0.30 vs. 16.47 ± 0.28, P = 0.042), long-text information integration (17.42 ± 0.33 vs. 15.83 ± 0.35, P < 0.001), and response completeness (17.21 ± 0.35 vs. 16.64 ± 0.38, P = 0.003). Kimi K2.6 demonstrated a significantly lower information omission rate (4.2 ± 1.3% vs. 8.5 ± 2.1%, P < 0.001), and GPT-5.5 exhibited a notable "Lost in the Middle" phenomenon. Conclusion: In the long text teaching in the Hematology Department, models with strong long-text integration capabilities are suitable for systematic lesson preparation and long-term case teaching, which can effectively reduce the rate of information omission and ensure the completeness of teaching content and the coherence of the time axis. Models with high educational heuristics are more suitable for clinical thinking training and multi-round interactive teaching. It is recommended to establish a three-layer application framework of "education demand-driven, model capability-matched, and teacher quality control as the bottom line". This will promote AI-assisted teaching to move from "available" to "trustworthy" and "optimal use".
[1] Zhang ZQ, Qi YH, 2025, Application status, challenges, and prospects of artificial intelligence large language models in the medical field. China Development, 25(6): 59-64.
[2] Yao Q, Wang Z, Guo YJ, et al., 2026, Principle of Large Language Model and Its Applications in Healthcare. Computer Systems and Applications, 35(1): 102-116.
[3] Wang YN, Jiang ZX, He HY, et al., 2025, Evolution of large language models and their applications in clinical medical education. West China Medical Publishers, 40(5): 777-782.
[4] Tong B, Tian C, Li YM, et al., 2026, Feasibility of Large Language Models Assisting in the Creation of Health Science Popularization Works. Medical Journal of Peking Union Medical College Hospital, 17(4): 1063-1072.
[5] Li Y, Li Z, Zhang K, et al., 2023, ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain Knowledge. Cureus, 15(6): e40895.
[6] Umucu E, Solis G, Garza L, et al., 2025, Empathy by Design: Aligning Large Language Models for Healthcare Dialogue. 2025 IEEE International Conference on Big Data (BigData), 4693-4702.
[7] Tang CX, Pang QG, 2026, How LLMS and MCP reshape the future of medical education. China Medical Education Technology, 40(2): 177-186.
[8] Chen SF, Alyakin A, Seas A, et al., 2026, LLM-assisted systematic review of large language models in clinical medicine. Nature Medicine, 32(3): 1152-1159.
[9] Yin MZ, Li Q, Xiong WF, et al., 2020, A preliminary study on the application of PBL combined with CBL under multidisciplinary team collaboration mode in geriatric medicine teaching. Continuing Medical Education, 34(3): 22-24.
[10] Zhang LL, 2015, Application of PBL combined with CBL in clinical teaching of metabolic diseases. Journal of Modern Medicine and Health, (3): 461-463.
[11] Sun L, Liu XR, Lu XS, 2020, Problems and countermeasures of case-based teaching in early basic medical education. Basic Medical Education, 22(5): 327-329.
[12] Mennin S, 2021, Ten Global Challenges in Medical Education: Wicked Issues and Options for Action. Medical Science Educator, 31(1): 17-20.
[13] Eysenbach G, 2023, The Role of ChatGPT, Generative Language Models, and Artificial Intelligence in Medical Education: A Conversation With ChatGPT and a Call for Papers. JMIR Med Educ, 9: e46885.
[14] Shi Y, Yu K, Dong Y, et al., 2026, Large language models in education: a systematic review of empirical applications, benefits, and challenges. Computers and Education: Artificial Intelligence, 10: 100529.
[15] Qu X, Yang JM, Chen T, et al., 2023, Reflections on the Implications of the Developments in ChatGPT for Changes in Medical Education Models. Journal of Sichuan University (Medical Sciences), 54(5): 937-940.
[16] Wei MH, Liang JP, Li ZY, et al., 2026, A comparative study on the multidimensional performance of GPT-4o and DeepSeek-R1 in medical question answering in Hematology. China Medical Education Technology, 40(2): 222-227.
[17] Kimi K2: Open Agentic Intelligence. visited on May 5, 2026, https://arxiv.org/abs/2507.20534.
[18] Yang YY, Nie ZW, Zhou ZQ, 2026, Research on the Information Hallucination in Large Language Model Applications for Medical Education and Its Mitigation Strategies. Medical Education Research and Practice, 34(2): 293-302.
[19] Lin YH, Ye ZC, Chen YR, et al., 2026, Analysis of Dilemmas in the Integration of Generative Artificial Intelligence into Medical Education from the Perspective of Symbiotic Agency Theory. Medicine and Society, 39(3): 114-122.
[20] Ding XM, 2026, Research Progress on the Application of ChatGPT Technology in Medical Education. China Health Industry, 23(5): 128-131.
[21] Cong S, Bai WH, Chen Z, et al., 2026, Research Progress of Large Language Models in Case Generation, Personalized Teaching and Assessment. Journal of Medical Molecular Biology, 23(1): 97-106.