清华大学研究团队在 ACM MM 2026(第 34 届 ACM 国际多媒体会议)发表论文,提出多模态大模型物体幻觉的新解释机制——"视觉源幻觉"(visual-origin hallucination)。
传统观点认为,模型幻觉主要源于"语言先验"——训练语料中某些概念的共现统计使模型产生偏见。但清华团队发现,当模型输出很短(如仅回答 Yes/No)时,语言先验的作用大幅削弱,此时幻觉根植于视觉特征提取本身。研究人员据此提出 ACFT 方法,仅用 COCO 数据集 0.9% 的数据量、且不增加任何推理开销,就在 POPE、MME 及四个描述级幻觉基准上,于 LLaVA、MiniGPT-4、Qwen2.5-VL 三个模型上取得优异表现。
论文链接:http://arxiv.org/abs/2609.00231
【AICOR 点评】"视觉源幻觉"的发现具有范式意义:它证明多模态模型的可靠性瓶颈不仅在于"怎么说",更在于"怎么看"。对于部署 AI 视觉系统的企业——尤其是自动驾驶、医疗影像等高风险场景——这一研究提醒我们:幻觉治理不能只在文本侧做文章,视觉编码器的质量同样决定系统安全边界。AICOR 认为,可信 AI 的建设需要跨模态的全链路审视,而非单点优化。
Copyright Notice This article is AI-assisted, rewritten from public reports. Copyright of the information belongs to the original authors and media; content is for industry sharing only and does not constitute investment or business advice.
For copyright concerns, contact AICOR (400-601-8080 / WeChat: aicor-ai); we will handle it promptly upon notification.