认知扰动标注(CPA):一种基于认知冲突的数据标注方法

罗 书锟
清远同胜教育咨询有限公司

摘要


传统基于客观事实的精确标注数据,在提升模型逻辑推理和抗幻觉能力方面遭遇了显著瓶颈[3]随着大语言
模型和多模态模型进入“规模报酬递减”阶段。现有的众包标记理论过分追求标注者之间的一致性[4-5],实际上削弱
了模型适应自然语言固有歧义性的能力[6]。针对此悖论,提出认知扰动标记(CPA)范式,建构标注熵(AE)量化
扰动价值,透过建构「语义地雷」与「逻辑陷阱」,生成认知冲突样本,并设计认知时薪框架,促使标记人员转型为
资料训练角色,提供资料层面的可行方案,以纾解大模型的幻觉。

关键词


认知扰动;标注熵;对抗性标注;红队测试

全文:

PDF


参考


[1]Kaplan J, McCandlish S, Henighan T, et al.

Scaling laws for neural language models[J]. arXiv preprint

arXiv:2001.08361, 2020.

[2]Hoffmann J, Borgeaud S, Mensch A, et al. Training

compute-optimal large language models[J]. arXiv preprint

arXiv:2203.15556, 2022.

[3]Ji Z, Lee N, Frieske R, et al. Survey of hallucination

in natural language generation[J]. ACM Computing Surveys,

2023, 55(12): 1-38.

[4]Snow R, O‘Connor B, Jurafsky D, et al. Cheap

and fast—but is it good? Evaluating non-expert annotations

for natural language tasks[C]//Proceedings of the 2008

Conference on Empirical Methods in Natural Language

Processing. 2008: 254-263.

[5]Artstein R, Poesio M. Inter-coder agreement for

computational linguistics[J]. Computational Linguistics, 2008,

34(4): 555-596.

[6]Bender E M, Gebru T, McMillan-Major A, et al.

On the dangers of stochastic parrots: Can language models be

too big?[C]//Proceedings of the 2021 ACM Conference on

Fairness, Accountability, and Transparency. 2021: 610-623.

[7]Christiano P F, Leike J, Brown T, et al. Deep

reinforcement learning from human preferences[C]//Advances

in Neural Information Processing Systems (NeurIPS). 2017, 30.

[8]Ouyang L, Wu J, Jiang X, et al. Training language

models to follow instructions with human feedback[C]//

Advances in Neural Information Processing Systems

(NeurIPS). 2022, 35: 27730-27744.

[9]Casper S, Davies X, Shi C, et al. Open problems

and fundamental limitations of reinforcement learning from

human feedback[J]. arXiv preprint arXiv:2307.15217, 2023.

[10]Perez E, Huang S, Song F, et al. Red teaming

language models with language models[C]//Proceedings

of the 2022 Conference on Empirical Methods in Natural

Language Processing. 2022: 341-362.

[11]Hendrycks D, Gimpel K. A baseline for detecting

misclassified and out-of-distribution examples in neural

networks[C]//Proceedings of the International Conference

on Learning Representations (ICLR). 2016.

[12]Guo C, Pleiss G, Sun Y, et al. On calibration of

modern neural networks[C]//International Conference on

Machine Learning (ICML). PMLR, 2017: 1321-1330.

[13]张琦,刘奕群,马少平.众包标注质量控制研究

综述[J].计算机研究与发展,2022,59(10):2157-2176.

[14]车万翔,刘挺,冯骁骋.大语言模型对齐技术研

究进展[J].中文信息学报,2024,38(1):1-18.


Refbacks

  • 当前没有refback。