Paper: 2609.05399 Authors: Julien Colin, Nuria Oliver, Thomas Serre Categories: cs.CV, cs.HC
The Gap
Over the past decade, explainable artificial intelligence (XAI) in computer vision has assembled an enormous technical toolbox: gradient attribution maps (Grad-CAM, Integrated Gradients), feature visualization via activation maximization, concept-based testing (TCAV), and mechanistic circuit analysis.
Yet despite thousands of published papers, the field has fallen into a self-referential trap. Researchers spend virtually all their energy comparing methods against methods: debating whether attribution technique A produces sharper saliency edges than attribution technique B on a frozen ResNet-50.
Meanwhile, the actual core scientific question — are modern neural network architectures (such as Vision Transformers, ConvNeXts, and Vision-Language Foundation Models) becoming more intrinsically interpretable than their predecessors? — has been almost completely neglected. The field behaves like an astronomy department that spends ten years arguing about the lens coatings of five different telescopes, without ever pointing a single telescope at the stars to measure how galaxies evolve.
[CURRENT XAI PARADIGM: METHOD-CENTRIC]
Method A (Grad-CAM) vs Method B (Integrated Gradients) vs Method C (TCAV)
|
v (Evaluated endlessly on frozen legacy backbones)
Outcome: 10,000 papers arguing about visualization tools,
while real models grow into trillions of unmapped weights!
|
+-------------------------------+
v
[THE PROPOSED MODEL-CENTRIC SHIFT]
Point mature XAI tools directly at model architectures!
Measure: Does shifting from ConvNets to ViTs to VLMs
make internal circuits more aligned with human cognition?
The Increment
One sentence: Before this paper, explainable AI was stuck in a meta-debate comparing visualization techniques; after it, Brown University and ELLIS researchers draw a decisive parallel to systems neuroscience, establishing an actionable, model-centric XAI agenda focused on measuring representational alignment and human mental predictability across model generations.
Core Mechanism
The authors argue that the XAI toolbox is now sufficiently mature to cease inventing new post-hoc attribution masks. Instead, the community must redeploy these instruments toward Model-Centric Explainability along two empirical axes:
- Mechanistic Representational Characterization: Using representational similarity analysis (RSA), sparse autoencoders, and linear probes to systematically track how different model architectures (ResNet vs ViT vs ConvNeXt vs Multimodal Transformers) structure their latent spaces. Do self-attention heads form more human-aligned semantic concepts than convolutional kernels, or do they distribute reasoning across uninterpretable global superposition?
- Empirical Human Understandability (Mental Model Predictability): Moving beyond expert confirmation bias. Instead of showing an AI engineer a heatmap to see if they nod approvingly, evaluate non-expert end users (such as radiologists, drivers, or loan officers). Can an operator accurately predict when a model will fail after observing its explanations on 20 examples? This is an empirical behavioral property that must be directly measured via psychophysical experiments, not deduced from mathematical axioms.
- The Systems Neuroscience Analogy: Neuroscience does not exist to debate fMRI versus EEG; it uses both instruments to map cortical representations in biological brains. XAI must treat artificial neural networks as computational organisms to be mapped and understood over evolutionary generations.
MODEL-CENTRIC XAI FRAMEWORK
[Mature XAI Toolbox]
(Attribution, Circuit Probing, Sparse Autoencoders, Concept Vectors)
|
v
+-----------------------------------------------------------+
| MODEL-CENTRIC AUDITING |
| |
| [Axis 1: Representational Mapping] |
| Track feature geometry across ResNet -> ViT -> VLM |
| |
| [Axis 2: Human Operator Mental Models] |
| Psychophysical tests: Can humans anticipate model bugs? |
+-----------------------------+-----------------------------+
|
v
[Result: Scientific Metric of Architectural Interpretability]
Select model architectures that are fundamentally safer to deploy.
To illustrate this shift, consider a structural metaphor of deep-sea marine biology. For twenty years, oceanographers argued endlessly in conference halls over whether sonar, underwater lidar, or fluoroscopic cameras had the best color contrast. While they argued, strange new abyssal creatures were swimming past uncataloged. A model-centric marine biologist steps onto the boat, packs all three instruments, dives into the trench, and focuses 100% of their attention on dissecting the giant squid to understand its organs and circulatory system.
Key Concepts
- Model-Centric XAI: An empirical paradigm shift that evaluates and benchmarks machine learning architectures on their intrinsic transparency and cognitive alignment, rather than evaluating explainability methods.
- Mental Model Predictability: An objective, human-in-the-loop metric evaluating whether a user can correctly anticipate whether a model will succeed or fail on a novel input after reviewing its explanations.
- Representational Alignment: The degree to which an artificial network’s internal latent activations correspond to natural human perceptual concepts and task invariants.
Framework Shift
Before (Method-Centric Echo Chamber):
Model (Fixed) ---> [Compare Method 1 vs Method 2 vs Method 3] ---> Paper: "Our heatmap is 3% sharper"
(The actual AI model remains an uninspected black box)
After (Model-Centric Scientific Investigation):
Methods (Fixed) -> [Instrument Suite] -> [Audit ResNet vs ViT vs VLM] -> Finding: Architectural Insight!
(Empirically identifies which model designs are safer and more legible to society)
From inventing endless cosmetic heatmap variations to conducting rigorous comparative anatomy across neural network architectures, the core shift is transforming XAI into an observational science.
Expert Assessment
Problem choice: Profound and clarifying. The XAI literature has suffered from severe diminishing returns, publishing thousands of redundant saliency map papers that fail to make deployed models one bit safer.
Method maturity: Grounding the argument in systems neuroscience and psychophysics provides the exact theoretical rigor that post-hoc machine learning explainability has desperately lacked.
Experimental integrity: The paper provides a comprehensive taxonomic survey of the few pioneering papers that have actually compared model architectures, proving that the gap is real and the tools are already ready to use.
Writing quality: Eloquent, provocative, and deeply scholarly. The critique is incisive without being dismissive of foundational work.
Verdict: strong accept — A visionary position paper that will serve as a guiding compass for the next generation of explainable AI research.
Takeaways
- Stop writing papers that introduce minor variants of Grad-CAM or integrated gradients; the existing toolbox is already good enough.
- When choosing an architecture for high-stakes deployment (e.g., healthcare or autonomy), benchmark its mental model predictability across human operators before looking at test-set accuracy.
- Treat artificial neural networks as empirical systems to be dissected and mapped, following the rigor of cognitive neuroscience.
论文: 2609.05399 作者: Julien Colin, Nuria Oliver, Thomas Serre 分类: cs.CV, cs.HC
缺口
在过去的十余年间,计算机视觉领域的可解释性人工智能(XAI)构建了一套庞大而成熟的工具箱: 从梯度热力图归因(如 Grad-CAM、集成梯度 Integrated Gradients),到特征最大化可视化,再到基于概念的测试(TCAV)以及神经回路归因。
然而,在发表了数万篇学术论文之后,整个可解释性领域却陷入了一场荒谬的自我循环。 绝大多数研究者的精力,全部耗费在了**“解释方法之间的横向互踩与对比”**: 学者们在一个冻结不动的 ResNet-50 老旧模型上,激烈争辩归因算法 A 的热力图边缘是否比归因算法 B 更加锐利细腻。
而真正关乎人类文明安全的科学本源问题——随着视觉神经网络从卷积网络(CNN)演进到视觉 Transformer(ViT),再演化到多模态基座大模型(VLM),我们的模型内部究竟是变得更透明、更契合人类认知,还是变得更加晦涩不可控了?——几乎无人问津。 这就好比一个天文学系花了整整十年时间,关在实验室里激烈争论五款天文望远镜的镀膜工艺谁更透光,却从来没有人把望远镜对准星空,去记录银河系真实的演化轨迹。
[既有 XAI 困境: 以解释方法为中心]
归因算法 A (Grad-CAM) vs 算法 B (集成梯度) vs 算法 C (TCAV)
|
v (在早已过时的陈旧冻结基座上无休止对比微调)
行业结局: 产出上万篇论文争夺“哪种热力图更顺眼”,
而产业界的真实大模型却在万亿参数黑盒里彻底脱轨失控!
|
+-------------------------------+
v
[破局之策: 以模型为中心的范式跃迁]
将成熟的 XAI 工具箱直接调转枪口,对准模型本身!
系统度量: 从 CNN 到 ViT 再到视觉大模型,
其内部表征与人类认知的对齐度到底是在进步还是在倒退?
增量
一句话: 在这篇论文之前,可解释性 AI 长期困死在可视化调优的自娱自乐之中;在这篇论文之后,布朗大学与 ELLIS 团队将认知系统神经科学与现代 AI 深度绑定,正式吹响了“以模型为中心(Model-Centric XAI)”的研究号角,推动学界从评测工具转向为不同模型架构建立认知解剖学。
核心机制
作者旗帜鲜明地指出:XAI 现有的工具箱已经完全足够成熟,无需再发明更多五花八门的热力图修饰技巧。 当务之急是将这些工具转化为“显微镜”,从两个关键维度全面审视模型架构本身(Model-Centric XAI):
- 机制化表征解剖(Mechanistic Representational Characterization):利用表征相似性分析(RSA)、稀疏自编码器(SAE)以及线性探针,跨代系追踪不同视觉架构(ResNet vs ViT vs ConvNeXt vs 多模态大模型)的内部几何演进。 Transformer 的自注意力头究竟是形成了更具人类解剖学意义的语义概念,还是把特征打散到了不可解释的全局高维叠加态之中?
- 人类心理模型可预测性(Empirical Mental Model Predictability):彻底破除开发者的确认偏差。 可解释性绝不是让算法工程师看着热力图心领神会地自圆其说,而是必须通过严格的人机心理物理学双盲实验:当放射科医生、自动驾驶安全员连续观察了模型在 20 个样本上的解释之后,能否在第 21 个样本上准确预测模型是会成功还是会犯错? 这种心理预测力只能通过实验测量,绝不能在数学上凭空假设。
- 系统神经科学对照:神经科学绝不会耗费十年去争执核磁共振(fMRI)与脑电图(EEG)谁更高级;神经科学家用这些工具去解构大脑皮层的视觉流。 XAI 必须将人工神经网络视为人造生物大脑,去记录其进化史。
以模型为中心的 XAI 科学架构
[成熟的 XAI 工具箱 (充当探针与显微镜)]
(梯度归因、回路追踪、稀疏自编码器、概念激活向量)
|
v
+-----------------------------------------------------------+
| 以模型为中心的大脑解剖 |
| |
| [维度一: 内部表征空间几何测量] |
| 追踪不同架构 (ResNet -> ViT -> VLM) 的语义聚集规律 |
| |
| [维度二: 终端人类操作员心理模型评测] |
| 心理物理学实测: 人类能否精准预判该架构的致命盲区与 Bug? |
+-----------------------------+-----------------------------+
|
v
[终极产出: 模型架构内在可解释性与安全度的大一统评估指标]
直接指导工业界选择在本质上更易受监管、更安全的神经网络架构。
可以用一个深海海洋生物学家的核喻来理解这场范式转向: 二十年来,深海研究者们坐在岸上的会议室里争吵不休,争论声呐、水下激光雷达和荧光显微镜哪个画质更饱满。 就在他们吵得不可开交时,未知的巨型深海生物正静悄悄地游过黑暗的海沟。 一位务实的海洋生物学家背起全部三套仪器跳上科考船,直潜马里亚纳海沟。 他不再给仪器写测评,而是全神贯注地用仪器解剖巨型大王乌贼,记录它的内脏构造、神经系统与循环机制。
关键概念
- 以模型为中心的可解释性(Model-Centric XAI):一种不再纠结于解释工具本身优劣,而是将解释工具作为客观测量尺度的科学范式,专注于解构不同神经网络架构的内在可读性。
- 心理模型可预测性(Mental Model Predictability):衡量人类操作员在接触模型解释后,能否在未知复杂场景下准确推断模型决策行为与潜在漏洞的量化心理学指标。
- 表征对齐度(Representational Alignment):人工神经网络的高维隐藏层特征与人类认知常识概念之间的数学同构程度。
框架转变
之前(以方法为中心的论文内卷):
固定老旧模型 ---> [对比归因方法 A vs 方法 B vs 方法 C] ---> 论文宣称:“我的热力图平滑度提升了 3%!”
(真实世界里越来越庞大的 AI 架构依然是绝对的暗黑深渊)
之后(以模型为中心的自然科学求索):
固定成熟方法 -> [标准化诊断仪器组] -> [跨代系解剖 ResNet vs ViT vs VLM] -> 提炼:架构安全性发现!
(建立客观标准,指引学界设计出在结构上天生更易被人类监管理解的架构)
从无休止地粉饰热力图的美观度,转向对神经网络不同物种展开严谨的“比较解剖学”,核心转变在于让可解释性 AI 摆脱工程内卷,回归认知实证科学。
专家评审
选题眼光: 振聋发聩,如惊雷般击碎了计算机视觉可解释性领域的集体怠惰。 在视觉大模型席卷一切的今天,传统的局部归因论文早已产出递减,亟需这样的宏观宣言重振军心。
方法成熟度: 巧妙引入系统神经科学与实验心理物理学的经典框架,填补了计算机科学在“测量人类主观认知”上的方法论空白。
实验诚意: 全面梳理了过去数年间少为人知的跨模型横向比对工作,论据翔实有力,证明“以模型为中心”绝非空想,而是完全具备即刻落地的技术可行性。
写作功力: 笔锋犀利、格局恢弘,既有对行业积弊的辛辣讽刺,又有对未来蓝图的条分缕析。
Verdict: 强接收(Strong Accept) — 必将载入可解释性 AI 发展史册的纲领性指南。
要点总结
- 停止发表对经典归因算法进行微小修补的灌水论文;既有的 Grad-CAM、TCAV 和稀疏探针已经足够胜任测量仪器的角色。
- 医疗、自动驾驶等高危场景在选型模型时,务必引入“人类操作员心理可预测性测试”,绝不能仅看准确率榜单。
- 将神经网络视为由硅基突触构成的新物种,以认知神经科学的求真精神去描摹其内部表征地图。