Paper: 2604.20665 Authors: Karan Goyal, Dikshant Kukreja (KDD) Categories: cs.AI, cs.CV

Abstract

The rapid proliferation of Vision-Language Models (VLMs) operates on a dangerous, unquestioned axiom: that current VLMs faithfully synthesise multimodal data. This paper argues they do not. Rather than extracting grounded knowledge from visual inputs, state-of-the-art models frequently exhibit functional blindness, exploiting strong language priors to bypass severe visual representation bottlenecks.

Key Problem: Functional Blindness

VLMs often:

  • Rely on language priors rather than visual evidence
  • Bypass visual representation bottlenecks via language priors
  • Fail to truly “see” and integrate visual information

Proposed Framework: Modality Translation Protocol

Novel information-theoretic approach that translates semantic payloads rather than ablating them, formulating three metrics:

  1. Toll of Seeing (ToS): Cost of relying on visual vs. language information
  2. Curse of Seeing (CoS): Performance degradation when forced to use visual info
  3. Fallacy of Seeing (FoS): Mismatch between claimed visual understanding and actual behavior

Semantic Sufficiency Criterion (SSC): Combined metric to evaluate multimodal trustworthiness

Divergence Law of Multimodal Scaling

As language engines scale to unprecedented reasoning capabilities, the mathematical penalty of the visual knowledge bottleneck paradoxically increases.

Takeaways

  • Current VLM evaluation confounds dataset biases with architectural incapacity
  • Need to shift from passive diagnostic to active architectural blueprint
  • KDD community should abandon “multimodal gain” pursuit
  • Trustworthy multimodal reasoning requires fundamentally different architecture

论文: 2604.20665 作者: Karan Goyal, Dikshant Kukreja (KDD) 分类: cs.AI, cs.CV

摘要

视觉-语言模型(VLM)的快速普及建立在一个危险的、未被质疑的公理之上:当前VLM忠实地综合多模态数据。本文认为它们并非如此。与其从视觉输入中提取有根据的知识,SOTA模型经常表现出功能性失明,利用强大的语言先验来绕过严重的视觉表示瓶颈。

核心问题:功能性失明

VLM经常:

  • 依赖语言先验而非视觉证据
  • 通过语言先验绕过视觉表示瓶颈
  • 未能真正”看到”并整合视觉信息

提出的框架:模态翻译协议

一种新颖的信息论方法,翻译语义负载而非删除它们,制定三个指标:

  1. 视觉之税(Toll of Seeing, ToS):依赖视觉vs语言信息的代价
  2. 视觉之咒(Curse of Seeing, CoS):被迫使用视觉信息时的性能下降
  3. 视觉之谬(Fallacy of Seeing, FoS):声称的视觉理解与实际行为之间的不匹配

语义充分性准则(SSC):评估多模态可信度的综合指标

多模态缩放的散度定律

随着语言引擎扩展到前所未有的推理能力,视觉知识瓶颈的数学惩罚反而增加

要点总结

  • 当前VLM评估混淆了数据集偏差与架构能力不足
  • 需要从被动诊断转向主动架构蓝图
  • KDD社区应放弃对”多模态增益”的追求
  • 可信赖的多模态推理需要根本不同的架构