面向多模态内容的情感识别大模型研究及应用
摘要
足和跨场景鲁棒性弱等问题,本文提出面向多模态内容的开放词表情感识别框架EMER-OV,将任务重构为“标
签—证据—置信”联合生成,并引入生成式情感推理验证机制和模态异步提示机制。基于GoEmotions数据集的文本
原型实验表明,该框架在证据重合度和解释一致性方面优于通用零样本方法,并在缺模态模拟条件下表现出较稳定
的推理能力。研究为开放词表、可解释和跨场景泛化的情感识别提供了可实现路径。
关键词
全文:
PDF参考
[1]吴敖,王海龙,柳林,史文韬.情感识别大模型研
究综述[J].计算机科学与探索,2026,20(3):625-649.
[2]Busso C, Bulut M, Lee C C, et al. IEMOCAP:
Interactive emotional dyadic motion capture database[J].
Language Resources and Evaluation, 2008, 42(4): 335-359.
[3]Bagher Zadeh A, Liang P P, Poria S, et al. Multimodal
Language Analysis in the Wild: CMU-MOSEI Dataset and
Interpretable Dynamic Fusion Graph[C].Proceedings of ACL.
2018: 2236-2246.
[4]Zhao J, Zhang T, Hu J, et al. M3ED: Multi-modal
Multi-scene Multi-label Emotional Dialogue Database and
Benchmark[C].Proceedings of ACL. 2022: 5690-5705.
[5]Jiang X, Zong Y, Zheng W, et al. DFEW: A LargeScale Database for Recognizing Dynamic Facial Expressions
in the Wild[C].Proceedings of ACM MM. 2020: 2881-2889.
[6]Liu Y, Dai W, Feng C, et al. MAFW: A Large-scale,
Multi-modal, Compound Affective Database for Dynamic
Facial Expression Recognition in the Wild[C].Proceedings of
ACM MM. 2022.
[7]Demszky D, Movshovitz-Attias D, Ko J, et al.
GoEmotions: A Dataset of Fine-Grained Emotions[C].
Proceedings of ACL. 2020: 4040-4054.
[8]Tsai Y H H, Bai S, Liang P P, et al. Multimodal
Transformer for Unaligned Multimodal Language
Sequences[C].Proceedings of ACL. 2019: 6558-6569.
[9]Radford A, Kim J W, Hallacy C, et al. Learning
Transferable Visual Models From Natural Language
Supervision[C].Proceedings of ICML. 2021: 8748-8763.
[10]Liu H, Li C, Wu Q, Lee Y J. Visual Instruction
Tuning[C].Advances in Neural Information Processing
Systems. 2023.
Refbacks
- 当前没有refback。
