Paper: 2609.13144 Authors: Anssi Moisio, Mathias Creutz, Mikko Kurimo Categories: cs.CL

The Gap

Compositional generalisation is usually divided into lexical and structural generalisation, and previous work has found that structural generalisation is harder than lexical for Transformers. That finding has been read as a property of the architecture — a claim about what Transformers can and cannot do compositionally.

The paper proposes a different explanation: the difference is not inherent to Transformers, but due to the high diversity of lexical types and low diversity of structural types in the specific datasets of those previous works. So the comparison may have been confounded, not by anything about the model, but by a property of the two test sets that nobody had varied.

And the definition of the property is the load-bearing part: by type diversity we mean the number of different constructors of that type, instead of, for example, the specific word combinations that might populate the structure. So it is not how many examples, and not how many surface strings — it is how many distinct ways of building a thing occur. That precision is what makes the hypothesis testable.

   THE RECEIVED FINDING AND THE PROPOSED EXPLANATION

   compositional generalisation is divided into
     LEXICAL generalisation
     STRUCTURAL generalisation
        |
        v
   PREVIOUS WORK: STRUCTURAL is HARDER than LEXICAL for Transformers
     -> read as a property OF THE ARCHITECTURE
        a claim about what Transformers CAN AND CANNOT DO
        compositionally
        |
        v
   [THE PAPER'S PROPOSAL]
     the difference is NOT INHERENT TO TRANSFORMERS, but due to
       HIGH diversity of LEXICAL TYPES and
       LOW diversity of STRUCTURAL TYPES
     in the SPECIFIC DATASETS of those previous works
     -> the comparison may have been CONFOUNDED, not by anything
        about the MODEL, but by a PROPERTY OF THE TWO TEST SETS
        THAT NOBODY HAD VARIED

   [THE DEFINITION IS THE LOAD-BEARING PART]
     by TYPE DIVERSITY we mean
       THE NUMBER OF DIFFERENT CONSTRUCTORS OF THAT TYPE
     instead of, for example
       THE SPECIFIC WORD COMBINATIONS THAT MIGHT POPULATE THE STRUCTURE
     -> not HOW MANY EXAMPLES, not HOW MANY SURFACE STRINGS
     -> it is HOW MANY DISTINCT WAYS OF BUILDING A THING occur
     <- that precision is what makes the hypothesis TESTABLE

The Increment

One sentence: Before this paper, structural generalisation’s difficulty was attributed to the architecture; after it, varying type diversity in linguistically controlled dataset variants shows that lexical and structural test cases behave alike, contradicting compound divergence as the explanation.

Core Mechanism

The method is to vary the suspected confound while holding everything else controlled, which is exactly the right design for a claim about a confounded comparison.

Linguistically diverse variants of the COGS and SLOG datasets are created using Grammatical Framework. Two properties of that choice matter. The datasets are previously published ones — the same ones behind the received finding — so the comparison is against the literature rather than against a new benchmark. And Grammatical Framework provides systematic linguistic control, so the variants differ in the intended property rather than incidentally.

And the result is that type diversity correlates with compositional generalisation equally in lexical and structural test cases, supporting the hypothesis. “Equally” is the finding: the asymmetry that motivated the original conclusion disappears once diversity is varied. So the difference between the two test cases is attributable to the datasets’ diversity profiles, not to a general architectural limitation on structural generalisation.

The paper also reports a contradiction with prior explanation: it notes a contradiction with the proposition in previous work that compound divergence explains the difficulty in compositional generalisation tasks. So the account that had been offered for why structural generalisation is hard does not survive the controlled variation either. That is a stronger claim than merely proposing an alternative, because it identifies an existing explanation as unsupported.

And the investigation does not stop at the main hypothesis: the effects of other dataset properties on compositional generalisation are also examined, such as the diversity of types other than the novel test structure, and surface properties of the logical semantics format. Two things are worth noting about that list. The first extends the diversity argument beyond the type under test, which is a natural follow-up — if diversity matters, it plausibly matters for surrounding types too. The second concerns the format of the logical semantics — a surface property — and its inclusion signals that the effect is not purely about type counts.

   THE METHOD IS TO VARY THE SUSPECTED CONFOUND WHILE HOLDING
   EVERYTHING ELSE CONTROLLED
     <- exactly the right design for a claim about a CONFOUNDED
        COMPARISON

   LINGUISTICALLY DIVERSE VARIANTS of COGS and SLOG
   created using GRAMMATICAL FRAMEWORK
     <- the datasets are PREVIOUSLY PUBLISHED ones -- THE SAME ONES
        behind the received finding
        -> the comparison is against THE LITERATURE, not a NEW BENCHMARK
     <- GRAMMATICAL FRAMEWORK provides SYSTEMATIC LINGUISTIC CONTROL
        -> the variants differ in THE INTENDED PROPERTY rather than
           INCIDENTALLY

   THE RESULT
     TYPE DIVERSITY CORRELATES WITH COMPOSITIONAL GENERALISATION
     EQUALLY IN LEXICAL AND STRUCTURAL TEST CASES
       <- "EQUALLY" IS THE FINDING: the ASYMMETRY that motivated the
          original conclusion DISAPPEARS once diversity is varied
       -> the difference between the two test cases is attributable to
          the DATASETS' DIVERSITY PROFILES, not to a GENERAL
          ARCHITECTURAL LIMITATION on structural generalisation

   AND A CONTRADICTION WITH THE PRIOR EXPLANATION
     the paper notes A CONTRADICTION WITH THE PROPOSITION IN PREVIOUS
     WORK THAT COMPOUND DIVERGENCE EXPLAINS THE DIFFICULTY in
     compositional generalisation tasks
       <- so the account offered for WHY structural generalisation is
          hard DOES NOT SURVIVE the controlled variation EITHER
       <- a STRONGER claim than merely proposing an ALTERNATIVE: it
          identifies an EXISTING EXPLANATION as UNSUPPORTED

   AND THE INVESTIGATION DOES NOT STOP AT THE MAIN HYPOTHESIS
     the effects of OTHER DATASET PROPERTIES are also examined, such as
       THE DIVERSITY OF TYPES OTHER THAN THE NOVEL TEST STRUCTURE
       SURFACE PROPERTIES OF THE LOGICAL SEMANTICS FORMAT
     <- the FIRST extends the diversity argument BEYOND THE TYPE UNDER
        TEST -- a natural follow-up: if DIVERSITY MATTERS, it plausibly
        matters for SURROUNDING TYPES TOO
     <- the SECOND concerns the FORMAT of the logical semantics, a
        SURFACE property, and its inclusion signals the effect is NOT
        PURELY about TYPE COUNTS

Think of it as testing whether someone is bad at a category, or whether they have simply seen few examples of it. If you assess a learner on kitchen tasks and they do well, then assess them on workshop tasks where they have only ever seen one kind of workshop, the gap measures your test design rather than their ability. What you would do next is generate workshop tasks covering many kinds of construction and re-test — and if the gap closes, the original conclusion was about your test sets. And note the paper’s rule for what counts as a “kind”: not how many different objects appear, but how many different ways of assembling them. That is the difference between showing someone a hundred chairs and showing them a hundred ways to join wood.

Key Concepts

  • Type diversity as the confound: the number of distinct constructors of a type, not the number of examples or surface strings. It is the property the previous comparison did not vary.
  • Architecture claim versus dataset claim: whether structural generalisation is hard for Transformers, or hard on the datasets that tested it. The paper’s controlled variation distinguishes them.
  • Reusing the original datasets: COGS and SLOG with Grammatical Framework variants. It makes the comparison against the literature rather than against a new benchmark.
  • Contradicting compound divergence: an existing explanation does not survive the controlled variation. It is a stronger result than proposing an alternative account.
  • Extending past the tested type: diversity of surrounding types, and surface properties of the semantics format. It tests whether the effect is purely about type counts.

Framework Shift

Before (structural generalisation as harder for Transformers):
  lexical versus structural generalisation compared on published datasets
  -> structural is harder
  -> attributed to the architecture, with compound divergence as the reason
  -> the diversity of the two test sets is not varied

After (type diversity varied under linguistic control):
  COGS and SLOG variants built with Grammatical Framework
  -> type diversity correlates with generalisation EQUALLY in both
  -> the asymmetry was a property of the datasets' diversity profiles
  -> compound divergence does not survive either
  -> other properties, including a surface format effect, also matter

From reading a dataset asymmetry as an architectural limitation, to varying the diversity of constructors in linguistically controlled variants, the core shift is that how many distinct ways of building a structure you are shown may matter more than the structure itself.

Expert Assessment

Problem choice: Excellent, and the reframing targets a claim that shapes how people reason about Transformer compositionality. “Structural generalisation is harder” has been read as a fact about models, so showing it is a fact about the datasets changes what the field should infer from it.

Method maturity: The design is well chosen in three respects: it uses the same published datasets the original finding rested on, it varies the suspected confound systematically rather than by ablation, and the definition of type diversity is stated precisely enough that the hypothesis is falsifiable. Reporting the contradiction with compound divergence is the part that makes the result more than an alternative: an existing explanation failing under controlled variation is stronger evidence than a new one fitting. The follow-up examinations extend the argument beyond the tested type, which is what one would want if the mechanism is genuinely about diversity.

Experimental integrity: Reusing the original datasets is the strongest methodological choice, since a new benchmark would have introduced a different confound — the datasets themselves. The conclusion is scoped to the datasets examined, and the paper reports the additional properties it investigates rather than claiming the diversity account is complete. The limitation is that the variants are generated with one framework, so the linguistic control is only as broad as Grammatical Framework’s coverage, and the finding is about how generalisation behaves under controlled variation rather than a proof that no architectural asymmetry exists.

Writing quality: The definition of type diversity is stated before the hypothesis, which is exactly the right order for a claim that hinges on what the term means. Because the practical consequence is “vary your test sets before concluding about architecture”, a short passage demonstrating the reasoning on one concrete structural type would make the method easy to apply to other benchmarks.

Verdict: strong accept — it shows that a widely read architectural conclusion is a property of the test sets, does so on the original datasets under systematic linguistic control, and contradicts the explanation previously offered for the effect.

Takeaways

  • Count distinct constructors, not examples. How many ways of building a thing appear matters more than how many instances of it.
  • Vary the suspect confound before blaming the model. Compare on the original datasets so the new variation is the only difference.
  • Check whether an existing explanation survives your control. An account that breaks under controlled variation is a stronger result than a new account that fits.
  • Test surrounding types, not just the one under test. If diversity is the mechanism, it should operate beyond the novel structure.

论文: 2609.13144 作者: Anssi Moisio, Mathias Creutz, Mikko Kurimo 分类: cs.CL

缺口

组合泛化通常被分成词汇泛化与结构泛化,而此前的工作发现:对 Transformer 来说,结构泛化比词汇泛化更难。 这个发现被读作架构的属性——一个关于”Transformer 在组合性上能做与不能做什么”的主张。

论文提出了一个不同的解释:这个差别并非 Transformer 固有,而是源于那些既有工作所用「特定数据集」中”词汇类型多样性高、结构类型多样性低”。 所以这次比较可能被混淆了——混淆它的不是模型的任何性质,而是那两个测试集的一个没人变动过的属性。

而这个属性的定义才是承重的那一部分:所谓「类型多样性」,我们指的是”该类型「不同构造子的数量」“,而不是比如”可能填充该结构的那些具体词组合”。 所以它不是有多少样例,也不是有多少表层字符串——而是有多少种不同的「构造方式」出现。正是这份精确性,才让假设可被检验。

   既有发现,与所提出的解释

   组合泛化被分为
     「词汇」泛化
     「结构」泛化
        |
        v
   此前工作:「结构」对 TRANSFORMER 而言比「词汇」更难
     -> 被读作「架构的属性」
        一个关于"TRANSFORMER 在组合性上「能做与不能做什么」"的主张
        |
        v
   [论文提出的解释]
     这个差别「并非 TRANSFORMER 固有」,而是源于
       「词汇类型的多样性高」与「结构类型的多样性低」
     出现在那些既有工作的「特定数据集」里
     -> 这次比较可能被「混淆」了——混淆它的不是模型的任何性质,
        而是那两个测试集的一个「没人变动过的属性」

   [「定义」才是承重的那一部分]
     所谓「类型多样性」,我们指的是
       「该类型「不同构造子的数量」」
     而不是比如
       "可能填充该结构的那些具体词组合"
     -> 不是「有多少样例」,也不是「有多少表层字符串」
     -> 而是「有多少种不同的「构造方式」」出现
     <- 正是这份精确性,才让假设「可被检验」

增量

一句话: 在这篇论文之前,结构泛化的难度被归因于架构;在这篇论文之后,在语言上受控的数据集变体里改变类型多样性,显示词汇与结构两种测试情形表现一致,从而反驳了”复合发散”这一解释。

核心机制

方法就是”在把其他一切都控制住的前提下,变动那个可疑的混淆因素”——对一个关于”被混淆的比较”的主张而言,这正是正确的设计。

COGS 与 SLOG 的”语言上多样的变体”是用 Grammatical Framework 生成的。 这个选择有两点要紧:这些数据集是此前已发表的——正是支撑那个既有发现的那几个——所以比较是对着文献做的,而不是对着一个新基准;而 Grammatical Framework 提供了系统性的语言学控制,因此这些变体是在预期的那个属性上不同,而不是顺带地不同。

结果是:类型多样性与组合泛化的相关性,在词汇与结构两种测试情形上「一样」。 “一样”才是那个发现:一旦把多样性变动起来,当初引出那个结论的不对称就消失了。所以那两个测试情形之间的差别,可归因于数据集各自的多样性画像,而不是”结构泛化存在某种普遍的架构局限”。

论文还报告了与既有解释的一处矛盾:它指出与”此前工作中’复合发散(compound divergence)解释了组合泛化任务中的难度’这一命题存在矛盾”。 也就是说,那个曾被用来解释”结构泛化为何难”的说法,在受控变动之下同样站不住。这比”提出一个替代解释”是更强的主张,因为它指认了一个既有解释缺乏支撑。

而且考察没有停在主假设上:论文还检验了其他数据集属性对组合泛化的影响,比如”除那个新的测试结构之外的其他类型的多样性”,以及”逻辑语义格式的表层属性”。 这份清单有两点值得注意。第一点把多样性论证扩展到了被测类型之外——这是一个自然的后续:如果多样性确实是要紧的,那么它对周边类型很可能也要紧。第二点关乎逻辑语义的格式——一个表层属性——把它包含进来,说明这个效应并不纯粹关于类型计数。

   方法就是"在把其他一切控制住的前提下,变动那个可疑的混淆因素"
     <- 对一个关于"被「混淆」的比较"的主张而言,
        这正是正确的设计

   COGS 与 SLOG 的「语言上多样的变体」用 GRAMMATICAL FRAMEWORK 生成
     <- 这些数据集是「此前已发表的」——正是支撑那个既有发现的那几个
        -> 比较是「对着文献」做的,而不是对着一个新基准
     <- GRAMMATICAL FRAMEWORK 提供「系统性的语言学控制」
        -> 变体是在「预期的那个属性」上不同,
           而不是顺带地不同

   「结果」
     「类型多样性与组合泛化的相关性,在词汇与结构两种情形上一样」
       <- "一样"才是那个发现:一旦把多样性变动起来,
          当初引出那个结论的「不对称就消失了」
       -> 那两个测试情形之间的差别,可归因于
          「数据集各自的多样性画像」,而不是"结构泛化存在
          某种普遍的架构局限"

   「与既有解释的一处矛盾」
     论文指出与"此前工作中「复合发散解释了组合泛化任务中的难度」
     这一命题存在矛盾"
       <- 那个曾被用来解释"结构泛化「为何」难"的说法,
          在受控变动之下「同样站不住」
       <- 这比"提出一个替代解释"是「更强」的主张:
          它指认了一个「既有解释缺乏支撑」

   「考察没有停在主假设上」
     还检验了其他数据集属性的影响,比如
       「除那个新的测试结构之外的其他类型的多样性」
       「逻辑语义格式的表层属性」
     <- 第一点把多样性论证「扩展到了被测类型之外」——
        一个自然的后续:如果多样性确实要紧,
        它对「周边类型」很可能也要紧
     <- 第二点关乎逻辑语义的「格式」(一个「表层」属性),
        把它包含进来,说明该效应「并不纯粹关于类型计数」

可以用**“测试一个人是「不擅长某一类任务」,还是「只是见过太少的该类做法」“来理解这件事: 如果你在厨房类任务上评估一个学习者,他表现不错;然后在车间类任务上评估他,而他只见过一种车间做法**——那么这道差距衡量的是你的测试设计,而不是他的能力。 接下来你会做的是:生成覆盖许多种构造方式的车间任务并重测——如果差距消失,那么原来那个结论是关于你的测试集的。 而注意论文对”一种做法”的判定规则:不是出现了多少不同的物体,而是有多少种不同的组装方式。这就是”给某人看一百把椅子”与”给他看一百种接木头的方法”之间的差别。

关键概念

  • 以类型多样性作为混淆因素: 一个类型不同构造子的数量,而不是样例数或表层字符串数。它正是此前那次比较没有变动的属性。
  • 架构主张 vs 数据集主张: 结构泛化对 Transformer 难,还是在测试它的那些数据集上难。论文的受控变动把两者区分开。
  • 重用原始数据集: 用 Grammatical Framework 生成 COGS 与 SLOG 的变体。它让比较对着文献进行,而不是对着新基准。
  • 反驳复合发散: 一个既有解释在受控变动下不成立。这比提出替代说法是更强的结果。
  • 扩展到被测类型之外: 周边类型的多样性、以及语义格式的表层属性。它检验该效应是否纯粹关于类型计数。

框架转变

之前(把结构泛化当作对 TRANSFORMER 更难):
  在已发表数据集上比较词汇与结构泛化
  -> 结构更难
  -> 归因于架构,并以「复合发散」作为原因
  -> 两个测试集的「多样性」未被变动

之后(在语言学控制下变动类型多样性):
  用 GRAMMATICAL FRAMEWORK 构建 COGS 与 SLOG 的变体
  -> 类型多样性与泛化的相关性在两者上「一样」
  -> 那个不对称是「数据集多样性画像」的属性
  -> 「复合发散」同样站不住
  -> 其他属性(含一个表层格式效应)也有影响

从”把一个数据集层面的不对称读作架构局限”,转变为”在语言上受控的变体里变动构造子的多样性”,核心转变在于:你被展示了多少种不同的构造方式,可能比那个结构本身更要紧。

专家评审

选题眼光: 极好,而这个重构瞄准的是一个塑造人们如何思考 Transformer 组合性的主张。 “结构泛化更难”一直被当作”关于模型的事实”来读,因此表明它是”关于数据集的事实”,会改变这个领域应当从中推出什么。

方法成熟度: 设计在三点上选得好:它用的是原结论所依赖的同一批已发表数据集;它以系统性方式变动可疑的混淆因素、而不是靠消融;而”类型多样性”的定义被陈述得足够精确,使假设可被证伪。 报告与”复合发散”的矛盾,是让它不止于”一个替代说法”的地方:一个既有解释在受控变动下失败,比”一个新解释恰好吻合”是更强的证据。 后续考察把论证扩展到被测类型之外——如果机制确实是关于多样性的,这正是人们会想要的。

实验完整性: 重用原始数据集是最强的方法学选择,因为一个新基准会引入另一个混淆——数据集本身。结论被限定在所考察的这些数据集上;论文报告了它额外检验的属性,而没有声称这个”多样性说法”已经完整。 局限是:变体是用一个框架生成的,因此语言学控制的广度只到 Grammatical Framework 的覆盖面;而这个发现讲的是”泛化在受控变动下如何表现”,并不是”不存在任何架构层面的不对称”的证明。

写作功力: “类型多样性”的定义被放在假设之前——对一个关键在”这个词是什么意思”的主张来说,这正是正确的顺序。 由于实际后果是”在就架构下结论之前先变动你的测试集”,若能给一小段在一项具体结构类型上演示这个推理,会让方法容易迁移到其他基准。

判决: 强接收(Strong Accept) — 它表明一个被广泛引用的架构结论其实是测试集的属性,并且是在原始数据集、系统性语言学控制之下做到的,还反驳了此前为该效应提出的解释。

要点总结

  • 数不同的构造子,而不是样例。出现多少种构造方式,比出现多少个实例更要紧。
  • 在归咎于模型之前,先变动那个可疑的混淆因素。在原始数据集上比较,使新的变动成为唯一的差别。
  • 检查一个既有解释能否扛住你的控制。一个在受控变动下崩掉的解释,比一个恰好吻合的新解释是更强的结果。
  • 测周边类型,而不只是被测的那一个。如果多样性是机制,它应当在那项新结构之外也起作用。