从视觉—语言模型到临床转化:放射报告自动生成的技术演进、事实性与评价

吴 哲洋
上海理工大学

摘要


自动放射报告生成(automated radiology report generation,ARRG)已从图像描述问题进入医学视觉—语言建模。本文按技术演进、事实性、评价和临床转化四条主线整理代表性研究。方法由卷积神经网络—循环神经网络编码器—解码器发展到融合记忆与知识的Transformer,并延伸至医学多模态大模型;文本更流畅并不意味着影像依据更充分,无依据生成、异常遗漏和部位错误仍反复出现。检索增强、知识图谱、事实一致性目标和视觉定位分别作用于不同环节,效果取决于检索库、视觉表征与机构外验证。BLEU、ROUGE只反映表面文字相似度,CheXbert、RadGraph和GREEN补充实体、关系与错误严重度,但不能替代医师评价。现阶段更现实的定位,是在适用范围与医生监督明确的前提下辅助提高报告质量与效率。

关键词


放射报告生成;医学视觉—语言模型;多模态大模型;事实一致性;临床评价

全文:

PDF


参考


[1]BRADY A P. Error and discrepancy in radiology: inevitable or avoidable?[J]. Insights into Imaging, 2017, 8(1): 171-182. DOI:10.1007/s13244-016-0534-1.

[2]JING B, XIE P, XING E. On the automatic generation of medical imaging reports[C]//Proceedings of ACL. 2018: 2577-2586. DOI:10.18653/v1/P18-1240.

[3]WANG X, PENG Y, LU L, et al. TieNet: Text-image embedding network for common thorax disease classification and reporting in chest X-rays[C]//Proceedings of CVPR. 2018: 9049-9058. DOI:10.1109/CVPR.2018.00943.

[4]CHEN Z, SONG Y, CHANG T H, et al. Generating radiology reports via memory-driven transformer[C]//Proceedings of EMNLP. 2020: 1439-1449. DOI:10.18653/v1/2020.emnlp-main.112.

[5]LI C, WONG C, ZHANG S, et al. LLaVA-Med: Training a large language-and-vision assistant for biomedicine in one day[C]//Advances in Neural Information Processing Systems. 2023, 36.

[6]XIA P, ZHU K, LI H, et al. MMed-RAG: Versatile multimodal RAG system for medical vision language models[EB/OL]. arXiv:2410.13085, 2024. DOI:10.48550/arXiv.2410.13085.

[7]ZHANG Y, WANG X, XU Z, et al. When radiology report generation meets knowledge graph[C]//Proceedings of AAAI. 2020, 34(7): 12910-12917. DOI:10.1609/aaai.v34i07.6989.

[8]MIURA Y, ZHANG Y, TSAI E, et al. Improving factual completeness and consistency of image-to-text radiology report generation[C]//Proceedings of NAACL-HLT. 2021: 5288-5304. DOI:10.18653/v1/2021.naacl-main.416.

[9]YU F, ENDO M, KRISHNAN R, et al. Evaluating progress in automatic chest X-ray radiology report generation[J]. Patterns, 2023, 4(12): 100802. DOI:10.1016/j.patter.2023.100802.

[10]SMIT A, JAIN S, RAJPOOT P, et al. Combining automatic labelers and expert annotations for accurate radiology report labeling using BERT[C]//Proceedings of EMNLP. 2020: 1500-1519. DOI:10.18653/v1/2020.emnlp-main.117.

[11]JAIN S, AGRAWAL A, SAPORTA A, et al. RadGraph: Extracting clinical entities and relations from radiology reports[C]//Advances in Neural Information Processing Systems: Datasets and Benchmarks. 2021.

[12]OSTMEIER S, XU J, CHEN Z, et al. GREEN: Generative radiology report evaluation and error notation[C]//Findings of EMNLP. 2024: 374-390. DOI:10.18653/v1/2024.findings-emnlp.21.

[13]LI C Y, CHANG K J, YANG C F, et al. Towards a holistic framework for multimodal LLM in 3D brain CT radiology report generation[J]. Nature Communications, 2025, 16: 2278. DOI:10.1038/s41467-025-57426-0.

[14]BLANKEMEIER L, KUMAR A, COHEN J P, et al. Merlin: A computed tomography vision-language foundation model and dataset[J]. Nature, 2026. DOI:10.1038/s41586-026-10181-8.

[15]TRAN C, PAL B, STIRRAT T, et al. Recent advances in artificial intelligence for radiology report generation: A brief review[J]. BJR Artificial Intelligence, 2026, 3(1): ubag003. DOI:10.1093/bjrai/ubag003.




DOI: http://dx.doi.org/10.12361/2661-3506-08-09-163725

Refbacks

  • 当前没有refback。