Paper: 2608.12299 Authors: AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini, AmirMohsen Eshghi, Siavash Arjomand Bigdel Categories: cs.AI, cs.CV
The Gap
Here’s the situation this paper walks into. CAM (2016) needed a global-average-pooling head, so it only worked on a narrow class of CNNs. Grad-CAM removed that constraint by using gradients as channel weights. Grad-CAM++ and XGrad-CAM patched gradient saturation and averaging artifacts. Score-CAM and Ablation-CAM threw gradients out entirely and scored channels by measuring what happens to the output when you mask or ablate them. LayerCAM and HiResCAM attacked the resolution problem — last-layer feature maps are 7x7 and upsampling them to 224x224 is a lie dressed as a heatmap. Then ViTs arrived and the whole “last convolutional layer” premise stopped existing, so token-attribution methods appeared. Then CLIP, DINO, and SAM arrived and the premise of “a target class score” also weakened, because now you might be explaining a text-image similarity or a self-supervised feature cluster.
So the literature is not a list, it’s a chain of patches. But it’s read as a list. Existing surveys tend to either group everything under “XAI methods” at a level too abstract to be useful, or benchmark a handful of CAM variants on one dataset with one faithfulness metric. Neither tells you why Score-CAM exists or what it failed to fix. The specific gap this paper targets: nobody has organized the CAM family by the structural properties that actually determine when a method is applicable — what signal it extracts, what architecture it assumes, and what it optimizes for — and nobody has traced the gap-to-gap dependency between methods.
Problem: 10 years of CAM variants read as a flat list.
"Which one should I use?" has no principled answer.
|
v
Assumption: every method can be located on 3 readable axes
(attribution mechanism / architectural dependence
/ evaluation objective), and each one exists
because it names a flaw in an earlier one.
|
v
Method: strict corpus of 57 method-centered papers, 2016+
- build the 3-axis taxonomy
- review gradient CAMs, hybrid CAMs, model-aware CAMs
- for each: contribution AND the gap it leaves open
|
v
Evidence: the corpus-wide trend is one-directional
one class score --> comparative / contrastive
one conv layer --> multi-layer aggregation
deterministic --> probabilistic, causal
CNN-only --> tokens, CLIP, DINO, SAM
... but faithfulness, localization, robustness,
cost, and human trust each use different
protocols across papers.
|
v
Conclusion: mechanism innovation is healthy;
evaluation is the actual bottleneck.
You cannot rank these methods because
the field never agreed on the ruler.
The Increment
One sentence: Before this paper, the CAM literature was a decade-long pile of variants you had to read chronologically to understand; after it, you have a three-axis coordinate system that tells you which methods are even comparable to each other, plus an explicit map of which gap each method opened and which later method closed it.
Core Mechanism
The “method” here is a classification scheme, so the internals are the axes. Axis one, attribution mechanism: what signal does the method turn into a weight? Gradient-based methods use backprop signal at some layer. Gradient-free score methods use forward passes on perturbed inputs. Ablation methods delete a channel and measure the output drop. Causal methods try to isolate the effect of a region while holding confounders fixed. Distribution-comparison methods (the foundation-model-era group) compare a sample’s features against a reference distribution rather than reading a class logit at all. This axis predicts computational cost almost perfectly — one backward pass versus hundreds of forward passes.
Axis two, architectural dependence: does the method require GAP (original CAM), any convolutional stack (Grad-CAM family), attention plus residual token streams (transformer attribution), or an external foundation model as a prior (CLIP text alignment, DINO features, SAM masks)? This axis is the applicability filter. A practitioner with a ViT-B/16 can immediately discard two-thirds of the corpus.
Axis three, evaluation objective: is the method trying to be faithful (deletion/insertion curves, sanity checks), to localize (pointing game, IoU against boxes, WSOL/WSSS scores), to be robust (stability under noise or adversarial shift), to be cheap, or to be trusted by humans? This is where the paper’s negative finding lives. Methods that win on localization frequently lose on faithfulness, because a sharp, object-shaped mask looks like a good explanation to a human and to an IoU metric whether or not it reflects the model’s actual computation. The review’s per-method template — contribution, remaining gap, successor that addresses it — turns the corpus into a directed graph instead of a bibliography.
[ corpus: 57 method papers, 2016 .. present ]
|
+-------------------+-------------------+
| | |
v v v
AXIS 1: MECHANISM AXIS 2: ARCH DEP. AXIS 3: OBJECTIVE
. gradient . GAP-CNN only . faithfulness
. grad-free score . any CNN . localization
. ablation . ViT tokens . robustness
. causal / debias . CLIP / DINO / SAM . cost
. distributional . model-agnostic . human trust
| | |
+-------------------+-------------------+
|
v
per-method record template:
[ what it adds ] -- [ what it still misses ]
|
v
GAP CHAIN (excerpt)
CAM
| needs GAP head, must retrain
v
Grad-CAM
| gradient saturation, coarse 7x7 maps
v
Grad-CAM++ / XGrad-CAM ....... fix the weighting
| gradients unreliable / shattered
v
Score-CAM / Ablation-CAM ..... drop gradients, pay in FLOPs
| still last-layer, still low resolution
v
LayerCAM / HiRes-CAM ......... multi-layer, sharper
| assumes convolutions exist
v
token attribution (ViT) ...... attention + residual flow
| assumes a labeled class score exists
v
CLIP / DINO / SAM-based ...... open-vocabulary, label-free
| no shared benchmark to prove any of this
v
[ OPEN: unified evaluation protocol ]
Think of it as a relay race where nobody agreed where the finish line is. Each method is a runner. The baton is a specific named defect — “you need to retrain the network,” “your gradients saturate,” “your map is 7x7,” “your model has no convolutions.” A runner accepts the baton, fixes that one thing, and in the act of running creates a new defect, which becomes the baton for the next leg. That’s the gap chain, and it’s why the literature is genuinely sequential rather than parallel.
The three axes are the lanes. Runners in the gradient lane can be timed against each other; a runner in the CLIP-distribution lane is on a different track entirely, and comparing their times is meaningless — which is exactly the mistake papers make when they benchmark Score-CAM against a CLIP-based explainer on one deletion curve. And the finish line is the evaluation objective: some runners are sprinting toward faithfulness, some toward IoU with a bounding box, some toward “a radiologist said it looked right.” They all get medals, because each race declared its own tape. The review’s contribution is not a new runner. It’s drawing the lanes and pointing out that there are five finish lines, so the leaderboard everyone cites is fiction.
Key Concepts
-
Why gradients work as “importance” at all: Take the last convolutional layer. It outputs, say, 512 feature maps, each a 7x7 grid — think of them as 512 detector sheets, one lighting up for stripes, one for round shapes, one for grass texture. To explain “zebra,” you need to know which sheets mattered. The gradient of the zebra score with respect to a sheet answers “if this sheet got slightly brighter everywhere, how much would the zebra score rise?” Average that over the sheet, use it as the sheet’s weight, sum all 512 weighted sheets, and you get a 7x7 heatmap. That’s Grad-CAM in one paragraph. The catch: gradients answer an *infinitesimal question. If the zebra score is already saturated at 0.999, nudging the stripe detector changes nothing, so the gradient is near zero and the stripe sheet gets weight zero — even though stripes are the entire reason for the prediction. This one failure mode generated a whole subfamily of methods.
-
Gradient-free scoring, and what it costs: Instead of asking the infinitesimal question, ask the blunt one. Take feature sheet number 137, turn it into a mask over the image, show the model only that region, and see what the zebra score becomes. High score means sheet 137 was carrying real evidence. Do this for all 512 sheets, use the scores as weights. No gradients, no saturation, no shattered-gradient noise. The price is 512 forward passes instead of one backward pass — a factor of hundreds in latency. This is the paper’s cleanest illustration of the mechanism axis being a cost axis in disguise: the same explanation quality can be bought with math or bought with compute.
-
Faithfulness vs. localization is a real conflict, not a metric detail: Faithfulness asks “does this heatmap reflect what the model actually used?” You test it by deleting the highlighted pixels and checking that the prediction collapses. Localization asks “does this heatmap sit on the object?” You test it against a human-drawn box. These diverge whenever the model is right for the wrong reason. If a classifier detects “boat” from water texture, the faithful explanation highlights water and scores terribly on IoU with the boat box; a method that produces crisp boat-shaped masks scores beautifully on localization while actively hiding the model’s real behavior. Pretty explanations and honest explanations are in tension, and the review’s finding that these two objectives are measured in separate papers with separate protocols means the field has been able to avoid confronting this.
Framework Shift
Before (mainstream approach): After (this paper):
one image one image
| |
v +-----+-----+-----+
one CNN | | | |
| v v v v
last conv layer only layer1 layer2 ... tokens
| | | | |
v +-----+--+--+-----+
one class logit |
| v
v compare: class A vs class B,
weights = mean gradient sample vs reference distribution,
| image vs text prompt (CLIP),
v region vs SAM mask prior
7x7 map --> upsample 224x224 |
| v
v probabilistic / causal weight
"looks about right" |
v
evaluate on: faithfulness?
localization? robustness?
cost? human trust?
|
v
( no shared protocol )
<-- the open gap
One sentence: from explaining one class score in one low-resolution CNN layer to comparative, multi-layer, probabilistic, token- and foundation-model-aware attribution, the core shift is that an explanation stopped being a single readout and became a comparison — and the field’s bottleneck moved from producing heatmaps to agreeing on how to judge them.
Expert Assessment
Problem choice: The gap is real but modest. CAM surveys exist, and general XAI surveys exist by the dozen. What’s genuinely underserved is the seam this paper sits on: the transition from “explain a CNN classifier’s logit” to “explain a CLIP alignment or a DINO feature,” where the original CAM formulation’s assumptions — a class score, a convolutional stack, a GAP layer — have all quietly dissolved. Someone needed to write that down. The three-axis taxonomy is a sensible organizing move and the gap-chain framing is the best thing in the paper, because it captures something a flat taxonomy can’t: this literature is causally sequential. The observation that evaluation is fragmented is correct, well-known, and has been stated in nearly every XAI survey since 2020, so credit for the diagnosis is thin.
Method maturity: A taxonomy is not a method, and this one is defensible rather than surprising. Mechanism / architecture / objective is close to what a thoughtful practitioner would reach for unaided. The 57-paper “strict corpus” is the weak joint. Strict is good for reproducibility, but 57 is small for a decade-plus of one of the most-copied ideas in computer vision, and the abstract doesn’t reveal the inclusion rule. If the filter is “papers proposing a named CAM variant,” it’s defensible. If it’s a keyword search with a citation floor, then the corpus is a convenience sample and the “clear trend” claim is partly an artifact of what got included. Reviews live and die on this and it should be adjudicated in section 2, not asserted.
Experimental integrity: There are no experiments, which is legitimate for a review but sets a ceiling. The paper’s central finding — evaluation is fragmented and protocols aren’t comparable — is exactly the kind of claim that would be devastating if demonstrated and is merely persuasive when asserted. A single table reimplementing eight representative methods under one deletion protocol, one pointing game, and one cost measurement would have shown rank inversions across metrics and converted the argument from “these papers use different rulers” to “here is the specific method whose reported superiority evaporates.” That table is the missing 30% of this paper. Related risk: the taxonomy is unvalidated. If two independent readers can’t place the same method on the same axis coordinates, it’s a narrative device rather than a classification.
Writing quality: The abstract is well-organized and the per-method “contribution plus remaining gap” template is a discipline most surveys lack — it forces the authors to say something evaluative rather than paraphrase. Where I’d expect corners cut: the foundation-model section. CLIP/DINO/SAM-based explanation is the newest and thinnest part of the literature, it’s where the paper claims its novelty, and it’s also where a review is most likely to devolve into three-sentence summaries of ten papers without a shared frame. If one section rewrite would elevate the whole thing, it’s that one — specifically, making explicit what “explanation” even means when there is no class, and how faithfulness is defined when the thing being explained is a similarity in an embedding space. That question is unanswered in the field and the review is well-positioned to sharpen it.
Verdict: weak accept — a useful, well-structured map of a fragmented literature with a genuinely good gap-chain framing, held back by a small unjustified corpus, an unvalidated taxonomy, and a headline finding about evaluation that is asserted rather than demonstrated.
Takeaways
Concrete things worth stealing:
- The gap-chain template as a literature-review tool. “Contribution, remaining gap, successor that addresses it” is a better structure than “contribution” alone, and it generalizes to any field where methods are patches on prior methods — optimizers, PEFT variants, RAG architectures, quantization schemes. It converts a bibliography into a directed graph, and the leaf nodes are your open problems.
- Separate the applicability axis from the quality axis. Before comparing two explanation methods, check whether they even assume the same architecture and the same explanation target. Half the CAM comparisons in the wild are category errors. Same trap exists in agent benchmarks and in retrieval evaluation.
- Mechanism is cost in disguise. Gradient-based versus perturbation-based is not primarily an accuracy tradeoff, it’s one backward pass versus O(channels) forward passes. If you’re picking an explainer for a production inference path, that axis decides for you before quality does.
- Distrust localization scores as explanation quality. A crisp object-shaped mask is evidence about the metric, not about the model. If your model is right for the wrong reason, the honest explanation will score badly on IoU. Always pair a localization metric with a deletion/insertion metric, and treat disagreement between them as information rather than noise.
- The reference-distribution reframe. The foundation-model-era move — explain by comparing a sample’s features against a reference distribution instead of reading a class logit — transfers well beyond vision. It’s how you get explanations for models with no label space at all: embedding models, anomaly detectors, self-supervised encoders.
What not to expect: no new method, no benchmark, no numbers you can cite. Read it to orient yourself in the CAM family and to steal the taxonomy; don’t read it to decide which explainer is best, because the paper’s own conclusion is that the literature can’t answer that yet.
论文: 2608.12299 作者: AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini, AmirMohsen Eshghi, Siavash Arjomand Bigdel 分类: cs.AI, cs.CV
缺口
先说这篇论文面对的局面。
2016 年的 CAM 要求网络必须有全局平均池化头,所以只能用在一小类 CNN 上。
Grad-CAM 用梯度当通道权重,把这个限制拆掉了。
Grad-CAM++ 和 XGrad-CAM 又去补梯度饱和与平均化带来的偽影。
Score-CAM 和 Ablation-CAM 干脆不要梯度了:直接遮住或删掉某个通道,看输出掉多少,用这个当权重。
LayerCAM 和 HiResCAM 攻的是分辨率问题——最后一层特征图只有 7x7,硬上采样到 224x224,那不是热图,那是一种体面的谎言。
然后 ViT 来了,“最后一个卷积层”这个前提直接消失,于是出现了 token 归因方法。
再然后 CLIP、DINO、SAM 来了,“目标类别分数”这个前提也松动了——你现在要解释的可能是一个图文相似度,或者一簇自监督特征。
所以这堆文献本质上不是一个清单,而是一条补丁链。但大家当清单读。
现有综述要么把所有东西塞进”XAI 方法”这个抽象到没有用的层级,要么拿五六个 CAM 变体在一个数据集上跑一个忠实度指标就算 benchmark。
两种都回答不了:Score-CAM 到底为什么存在?它又没修好什么?
这篇论文瞄的具体缺口是:没人按”真正决定一个方法能不能用”的结构属性来组织 CAM 家族——它抽取什么信号、它假设什么架构、它优化什么目标——也没人把方法之间”缺口传缺口”的依赖关系画出来。
问题: 十年 CAM 变体被当成一张平铺清单读。
"我该用哪个"没有有原则的答案。
|
v
假设: 每个方法都能定位在 3 条可读的轴上
(归因机制 / 架构依赖 / 评测目标),
而且每个方法的存在理由,都是它点出了
前一个方法的某个具体缺陷。
|
v
方法: 严格筛出 57 篇方法型论文 (2016 至今)
- 建三轴分类
- 分别评梯度类、混合类、架构感知类
- 每篇都写: 贡献 + 它留下的缺口
|
v
证据: 全语料的趋势是单向的
单一类别分数 --> 比较式 / 对比式
单一卷积层 --> 多层聚合
确定性权重 --> 概率化、因果化
只能 CNN --> token / CLIP / DINO / SAM
... 但忠实度、定位、鲁棒性、算力成本、
人类信任,各家用的协议都不一样。
|
v
结论: 机制创新是健康的;
评测才是真正的瓶颈。
你排不出高下, 因为这个领域
从来没就"尺子"达成一致。
增量
一句话: 在这篇之前,CAM 文献是一堆必须按年份顺读才能理解的变体;在这篇之后,你有了一套三轴坐标,能判断哪些方法之间才具备可比性,还有一张明确的图:每个方法开了什么缺口、后来谁把它补上。
核心机制
这里的”方法”就是一套分类体系,所以内部结构就是那三条轴。
第一条轴,归因机制:这个方法把什么信号变成权重?
梯度类方法用某一层的反传信号。
免梯度打分类方法用扰动输入的前向传播结果。
消融类方法直接删掉一个通道,量输出的下降。
因果类方法试图在控制混杂因素的前提下隔离某个区域的效应。
分布比较类(也就是基础模型时代那一批)根本不读类别 logit,而是把样本特征跟一个参考分布比。
这条轴几乎完美预测算力成本:一次反向传播 vs 几百次前向传播。
第二条轴,架构依赖:这个方法要求 GAP(原始 CAM)、要求任意卷积堆叠(Grad-CAM 家族)、要求注意力加残差 token 流(transformer 归因),还是要求一个外部基础模型当先验(CLIP 文本对齐、DINO 特征、SAM 掩码)?
这条轴是可用性过滤器。
手上是 ViT-B/16 的人,可以当场划掉语料库的三分之二。
第三条轴,评测目标:这个方法想要忠实(删除/插入曲线、sanity check)、想要定位准(pointing game、跟框算 IoU、WSOL/WSSS 分数)、想要鲁棒(噪声或对抗扰动下稳定)、想要便宜,还是想要人类觉得可信?
论文最有价值的负面发现就在这里:定位分数高的方法经常忠实度差,因为一张锐利的、贴着物体轮廓的掩码,无论它是否反映模型的真实计算,在人眼和 IoU 指标看来都像个好解释。
那个”贡献 + 遗留缺口 + 谁来补”的逐方法模板,把语料库从参考文献列表变成了一张有向图。
[ 语料库: 57 篇方法论文, 2016 至今 ]
|
+-------------------+-------------------+
| | |
v v v
轴1: 机制 轴2: 架构依赖 轴3: 目标
. 梯度 . 只能 GAP-CNN . 忠实度
. 免梯度打分 . 任意 CNN . 定位
. 消融 . ViT token . 鲁棒性
. 因果 / 去偏 . CLIP/DINO/SAM . 算力成本
. 分布比较 . 模型无关 . 人类信任
| | |
+-------------------+-------------------+
|
v
逐方法记录模板:
[ 它加了什么 ] -- [ 它还缺什么 ]
|
v
缺口链 (节选)
CAM
| 需要 GAP 头, 必须重训
v
Grad-CAM
| 梯度饱和, 图只有 7x7
v
Grad-CAM++ / XGrad-CAM ...... 修权重公式
| 梯度本身不可靠 / 碎裂
v
Score-CAM / Ablation-CAM .... 扔掉梯度, 用算力换
| 还是最后一层, 还是低分辨率
v
LayerCAM / HiRes-CAM ........ 多层, 更锐利
| 前提是"存在卷积层"
v
token 归因 (ViT) ............ 注意力 + 残差流
| 前提是"存在带标签的类别分数"
v
CLIP / DINO / SAM 类 ........ 开放词表, 无标签
| 没有共享基准能证明上面任何一句
v
[ 未解决: 统一评测协议 ]
可以把它想成一场没人说清终点线在哪的接力赛。
每个方法是一名跑者。
接力棒是一个被具体点名的缺陷——“你得重训网络”、“你的梯度饱和了”、“你的图只有 7x7”、“你的模型里没有卷积”。
跑者接棒,修掉这一个问题,而修的过程本身又制造出一个新缺陷,这个新缺陷就成了下一棒。
这就是缺口链,也是为什么这批文献是真正有先后顺序的,而不是并行铺开的。
三条轴是跑道。
同在梯度跑道上的跑者,成绩可以互相比;而在 CLIP-分布跑道上的跑者根本不在一条道上,比时间毫无意义——这恰好就是那些拿 Score-CAM 和某个 CLIP 解释器在一条删除曲线上对轰的论文犯的错。
终点线就是评测目标:有人朝忠实度冲,有人朝跟框的 IoU 冲,有人朝”一位放射科医生说看着对”冲。
大家都拿到了奖牌,因为每场比赛自己拉的自己的终点带。
这篇综述的贡献不是新增一名跑者,而是把跑道划出来,并指出终点线有五条——所以大家一直在引的那张排行榜是虚构的。
关键概念
-
梯度凭什么能当”重要性”: 看最后一个卷积层。假设它输出 512 张特征图,每张 7x7——把它们当成 512 张探测纸,一张对条纹亮,一张对圆形亮,一张对草地纹理亮。要解释”斑马”,你得知道哪几张纸起了作用。斑马分数对某张纸的梯度回答的是:“如果这张纸整体亮一点点,斑马分数会涨多少?“把它在纸上求平均,当这张纸的权重,512 张加权求和,就得到一张 7x7 热图。这就是一段话版的 Grad-CAM。麻患在于:梯度回答的是无穷小的问题。如果斑马分数已经饱和在 0.999,那么动一下条纹探测器什么都不会变,梯度接近零,条纹那张纸权重就是零——尽管条纹正是这个预测的全部理由。就这一个失效模式,催生了一整个子家族。
-
免梯度打分,以及它的代价: 不问无穷小的问题,改问粗暴的问题。拿第 137 张特征图,把它变成图像上的一个掩码,只给模型看这块区域,看斑马分数变成多少。分数高,说明第 137 张纸真的携带证据。512 张纸全做一遍,用分数当权重。没有梯度,没有饱和,没有碎裂梯度的噪声。代价是 512 次前向传播换 1 次反向传播——延迟差几百倍。这是论文里最干净的一个例子,说明”机制轴”其实是伪装过的”成本轴”:同样的解释质量,可以用数学买,也可以用算力买。
-
忠实度和定位是真冲突,不是指标细节: 忠实度问的是”这张热图是否反映了模型真的用了什么”,做法是把高亮像素删掉,看预测是否崩。定位问的是”这张热图是否落在物体上”,做法是跟人画的框比。只要模型是”因为错的理由而答对”,这两者就会分道扬镳。如果一个分类器是靠水面纹理判断”船”,那么忠实的解释会高亮水面,跟船框的 IoU 惨不忍睹;而一个能产出干净船形掩码的方法定位分数漂亮,却在主动掩盖模型的真实行为。好看的解释和诚实的解释是有张力的。而这篇综述指出这两个目标在不同论文里用不同协议测量——这也正是这个领域一直得以回避这个矛盾的原因。
框架转变
之前(主流方法): 之后(本文方法):
一张图 一张图
| |
v +-----+-----+-----+
一个 CNN | | | |
| v v v v
只取最后卷积层 层1 层2 ... token
| | | | |
v +-----+--+--+-----+
一个类别 logit |
| v
v 比较: A 类 vs B 类,
权重 = 梯度均值 样本 vs 参考分布,
| 图像 vs 文本提示 (CLIP),
v 区域 vs SAM 掩码先验
7x7 图 --> 上采样到 224x224 |
| v
v 概率化 / 因果化权重
"看着差不多对" |
v
评测于: 忠实度? 定位?
鲁棒性? 成本? 人类信任?
|
v
( 没有共享协议 )
<-- 真正的缺口
一句话:从在一个低分辨率 CNN 层里解释一个类别分数,到比较式、多层、概率化、token 与基础模型感知的归因,核心转变是解释不再是一次读数,而变成了一次比较——同时该领域的瓶颈从”生产热图”移到了”就如何评判热图达成一致”。
专家评审
选题眼光: 缺口是真的,但不大。CAM 综述有,泛 XAI 综述几十篇。真正被服务不足的是这篇所处的接缝:从”解释 CNN 分类器的 logit”过渡到”解释 CLIP 对齐或 DINO 特征”,在这个过渡里,原始 CAM 的三个假设——有类别分数、有卷积堆叠、有 GAP 层——已经悄悄全部溶解了。这件事需要有人写下来。三轴分类是个合理的组织动作,而缺口链的框架是全文最好的东西,因为它抓住了平铺分类抓不到的一点:这批文献在因果上是有序的。至于”评测碎片化”这个观察,正确、众所周知、且从 2020 年起几乎每篇 XAI 综述都说过,所以诊断本身能给的分不多。
方法成熟度: 分类体系不算方法,而且这套体系属于站得住脚而非令人惊讶。机制/架构/目标,跟一个有想法的从业者自己会想到的东西很接近。真正的薄弱关节是那个 57 篇的”严格语料库”。严格有利于可复现,但对于计算机视觉里被抄得最多的想法之一、跨十余年的文献来说,57 篇偏少,而摘要没有交代纳入规则。如果筛选标准是”提出了具名 CAM 变体的论文”,那说得通;如果是关键词搜索加引用数门槛,那这就是一个便利样本,“趋势清晰”这个结论有一部分是纳入标准的产物。综述的生死全在这一点上,应该在第 2 节里正面裁决,而不是一句断言了事。
实验诚意: 没有实验,对综述来说合法,但这也设了一个天花板。论文的核心发现——评测碎片化、协议不可比——恰恰是那种”被证明出来就毁灭性、只被断言就仅仅是有说服力”的主张。只要一张表:拿八个代表性方法在统一的删除协议、统一的 pointing game、统一的成本测量下重跑一遍,就能展示出跨指标的排名翻转,把论证从”这些论文用了不同的尺子”升级成”这里就是那个所谓优势会蒸发的具体方法”。那张表就是这篇论文缺掉的 30%。另一个相关风险:分类体系未经验证。如果两位独立读者没法把同一个方法放到同一组轴坐标上,那它就是叙事装置,不是分类法。
写作功力: 摘要组织得不错,“贡献 + 遗留缺口”的逐方法模板是多数综述缺的一种纪律——它逼作者说出评价,而不是复述。我预期偷懒的地方是基础模型那一节。CLIP/DINO/SAM 类解释是这批文献里最新也最薄的部分,是论文自称新意所在,同时也最容易退化成”十篇论文各三句话、没有共同框架”。如果只能重写一节来抬升整篇,就是这一节——具体说,要讲清楚:当根本没有”类别”时,“解释”到底指什么;当被解释的东西是嵌入空间里的一个相似度时,忠实度怎么定义。这个问题整个领域都还没答,而这篇综述有很好的位置去把它磨锐。
判决: 弱接收 —— 一张有用、结构清晰的碎片化文献地图,缺口链框架确实好;但被三点拖住:语料库偏小且未论证、分类体系未经验证、以及关于评测的核心发现是断言而非证明。
要点总结
值得偷走的具体东西:
- 把缺口链当文献综述的模板。“贡献 / 遗留缺口 / 谁来补”比只写”贡献”好得多,而且能迁移到任何”方法是前一个方法的补丁”的领域——优化器、PEFT 变体、RAG 架构、量化方案。它把参考文献列表变成有向图,而叶子节点就是你的开放问题。
- 把可用性轴和质量轴分开。比较两个解释方法之前,先确认它们是否假设了同一种架构、同一种解释对象。野生的 CAM 对比里有一半是范畴错误。同样的坑在 agent benchmark 和检索评测里也有。
- 机制就是伪装的成本。梯度类 vs 扰动类,首要差别不是精度取舍,而是一次反向 vs O(通道数) 次前向。如果你要给生产推理链路挑解释器,这条轴在质量之前就已经替你决定了。
- 别把定位分数当解释质量。一张干净的物体形掩码,是关于指标的证据,不是关于模型的证据。如果你的模型是”因为错的理由答对”,诚实的解释在 IoU 上一定难看。定位指标永远要配一个删除/插入指标,而且要把两者的分歧当信息而不是噪声。
- “参考分布”这个重构视角。基础模型时代的那一招——不读类别 logit,而是拿样本特征去比一个参考分布——能迁移到视觉之外。这是你给”根本没有标签空间的模型”做解释的路子:嵌入模型、异常检测器、自监督编码器。
不要期待的:没有新方法,没有基准,没有可以引用的数字。
读它是为了在 CAM 家族里给自己定位、顺走那套分类;不要读它来决定”哪个解释器最好”,因为这篇论文自己的结论就是:这批文献目前还答不了这个问题。