Paper: 2604.20665 Authors: Karan Goyal, Dikshant Kukreja (KDD) Categories: cs.AI, cs.CV
Abstract
The rapid proliferation of Vision-Language Models (VLMs) operates on a dangerous, unquestioned axiom: that current VLMs faithfully synthesise multimodal data. This paper argues they do not. Rather than extracting grounded knowledge from visual inputs, state-of-the-art models frequently exhibit functional blindness, exploiting strong language priors to bypass severe visual representation bottlenecks.
Key Problem: Functional Blindness
VLMs often:
- Rely on language priors rather than visual evidence
- Bypass visual representation bottlenecks via language priors
- Fail to truly “see” and integrate visual information
Proposed Framework: Modality Translation Protocol
Novel information-theoretic approach that translates semantic payloads rather than ablating them, formulating three metrics:
- Toll of Seeing (ToS): Cost of relying on visual vs. language information
- Curse of Seeing (CoS): Performance degradation when forced to use visual info
- Fallacy of Seeing (FoS): Mismatch between claimed visual understanding and actual behavior
Semantic Sufficiency Criterion (SSC): Combined metric to evaluate multimodal trustworthiness
Divergence Law of Multimodal Scaling
As language engines scale to unprecedented reasoning capabilities, the mathematical penalty of the visual knowledge bottleneck paradoxically increases.
Takeaways
- Current VLM evaluation confounds dataset biases with architectural incapacity
- Need to shift from passive diagnostic to active architectural blueprint
- KDD community should abandon “multimodal gain” pursuit
- Trustworthy multimodal reasoning requires fundamentally different architecture
论文: 2604.20665 作者: Karan Goyal, Dikshant Kukreja (KDD) 分类: cs.AI, cs.CV
摘要
视觉-语言模型(VLM)的快速普及建立在一个危险的、未被质疑的公理之上:当前VLM忠实地综合多模态数据。本文认为它们并非如此。与其从视觉输入中提取有根据的知识,SOTA模型经常表现出功能性失明,利用强大的语言先验来绕过严重的视觉表示瓶颈。
核心问题:功能性失明
VLM经常:
- 依赖语言先验而非视觉证据
- 通过语言先验绕过视觉表示瓶颈
- 未能真正”看到”并整合视觉信息
提出的框架:模态翻译协议
一种新颖的信息论方法,翻译语义负载而非删除它们,制定三个指标:
- 视觉之税(Toll of Seeing, ToS):依赖视觉vs语言信息的代价
- 视觉之咒(Curse of Seeing, CoS):被迫使用视觉信息时的性能下降
- 视觉之谬(Fallacy of Seeing, FoS):声称的视觉理解与实际行为之间的不匹配
语义充分性准则(SSC):评估多模态可信度的综合指标
多模态缩放的散度定律
随着语言引擎扩展到前所未有的推理能力,视觉知识瓶颈的数学惩罚反而增加。
要点总结
- 当前VLM评估混淆了数据集偏差与架构能力不足
- 需要从被动诊断转向主动架构蓝图
- KDD社区应放弃对”多模态增益”的追求
- 可信赖的多模态推理需要根本不同的架构