Paper: 2609.11917 Authors: Atindra Jha, Margaret Li, Jure Leskovec, Percy Liang, Luke Zettlemoyer Categories: cs.CL, cs.LG
The Gap
As the supply of human-written text is exhausted, repeating language model training data has become standard practice. So the properties of repetition are not a niche question; they are a property of how models are trained now.
And the question had been answered for the wrong architecture. Prior work studied data repetition for densely activated Transformers, while the effects of data repetition remain largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts — despite their increased compute efficiency. That combination is what makes the gap consequential: MoE’s whole selling point is compute efficiency per unit of quality, and its efficiency comes from activating a subset of parameters per token. If sparse activation interacts badly with repeated data, then the architecture people adopt for efficiency has an unexamined dependence on an increasingly standard training practice.
REPETITION IS STANDARD, AND ITS EFFECTS WERE STUDIED
FOR THE WRONG ARCHITECTURE
supply of human-written text EXHAUSTED
-> REPEATING TRAINING DATA IS STANDARD PRACTICE
-> so repetition's properties are not a NICHE question;
they are a property of HOW MODELS ARE TRAINED NOW
|
v
AND THE QUESTION WAS ANSWERED FOR THE WRONG ARCHITECTURE
prior work: data repetition for DENSELY ACTIVATED TRANSFORMERS
largely unexplored: RECENTLY DOMINANT SPARSE ARCHITECTURES such
as MIXTURE-OF-EXPERTS -- DESPITE their increased compute
efficiency
|
v
[WHY THAT COMBINATION MAKES THE GAP CONSEQUENTIAL]
MoE's whole selling point is COMPUTE EFFICIENCY PER UNIT OF
QUALITY
its efficiency comes from ACTIVATING A SUBSET OF PARAMETERS
PER TOKEN
|
v
-> if SPARSE ACTIVATION interacts badly with REPEATED DATA,
then the architecture people adopt for EFFICIENCY has an
UNEXAMINED DEPENDENCE on an INCREASINGLY STANDARD training
practice
The Increment
One sentence: Before this paper, repetition’s effect on Mixture-of-Experts was unknown; after it, MoEs degrade faster, the effect scales with sparsity via total rather than active parameters, and masking-based regularization recovers the advantage at heavy repetition.
Core Mechanism
The sweep is designed to separate the variables that a single comparison would conflate: repetition rates varied across single- and multi-domain data mixes, and across MoE settings including expert count and granularity. Three axes at once — how much repetition, what the data is, and how the sparse architecture is configured. That breadth is what allows the findings to be attributed rather than merely observed.
The headline result is a comparison across architecture classes at matched scale: for models ranging from 80M to 1B active (8.5B total) parameters, MoEs degrade more rapidly under data repetition. Two things about that sentence matter. The range spans more than an order of magnitude, so the effect is not a small-model artefact. And the phrasing “active (total)” is deliberate — the models are matched on active parameters while differing in total, which is a fair comparison for inference cost but is also where the mechanism finding comes from.
That mechanism finding is the sharpest part: the effect increases with sparsity, dictated by total rather than active parameters. So the relevant quantity is how many parameters exist, not how many fire per token. That is counterintuitive and it is exactly what the architecture’s premise would not predict: MoE efficiency is usually reasoned about in terms of active parameters, and here the fragility tracks the inactive ones. The interpretation is that the extra parameters are where the memorisation lives.
The numbers give the shape of the degradation and of the trade:
- 80M dense models can repeat data over 8× with minimal degradation.
- MoEs instead begin to suffer at 4×, and deteriorate rapidly.
- They cede their performance benefits in all-unique data settings to underperform dense models after 32×.
Read together, those three lines describe a crossover: the sparse architecture wins when data is unique and loses when it is heavily repeated, with the reversal somewhere between 4× and 32× depending on how severe the degradation becomes.
Then the remedy, reported with its own limit. Some existing regularization methods can mitigate overfitting — dropout in particular — and with strong masking-based regularization, MoEs are able to outperform dense models even when data is repeated more than 64 times. But: no method fully matches the performance of all-unique training data. So regularization shifts the crossover rather than removing the penalty, which is the honest form of a mitigation result.
And the mechanism analysis connects the behaviour to internals: MoE routing universally stabilizes early in training, and expert specialization correlates with overfitting to repeated data. Two findings, and the second is the load-bearing one: overfitting is associated with experts specialising, which suggests the failure is not generic capacity excess but a specific division of labour that repeated data encourages. That points the suggested future work in a direction — reducing over-specialization by disrupting memorization patterns — rather than merely asking for better regularizers.
THE SWEEP IS DESIGNED TO SEPARATE VARIABLES THAT A SINGLE COMPARISON
WOULD CONFLATE
repetition rates varied across SINGLE- and MULTI-DOMAIN mixes
AND across MoE settings including EXPERT COUNT and GRANULARITY
<- THREE AXES AT ONCE: how much repetition, what the data is,
and how the sparse architecture is configured
-> breadth allows findings to be ATTRIBUTED rather than merely
OBSERVED
THE HEADLINE RESULT: A COMPARISON ACROSS ARCHITECTURE CLASSES AT
MATCHED SCALE
for models ranging from 80M to 1B ACTIVE (8.5B TOTAL) parameters,
MOES DEGRADE MORE RAPIDLY UNDER DATA REPETITION
<- the range spans MORE THAN AN ORDER OF MAGNITUDE
-> not a SMALL-MODEL ARTEFACT
<- "ACTIVE (TOTAL)" is DELIBERATE: matched on ACTIVE parameters
while differing in TOTAL
-> fair for INFERENCE COST, and also WHERE THE MECHANISM
FINDING COMES FROM
THE MECHANISM FINDING IS THE SHARPEST PART
the effect INCREASES WITH SPARSITY, DICTATED BY TOTAL RATHER THAN
ACTIVE PARAMETERS
-> the relevant quantity is HOW MANY PARAMETERS EXIST, not how
many fire per token
<- COUNTERINTUITIVE, and exactly what the architecture's premise
would NOT predict: MoE efficiency is usually reasoned about in
terms of ACTIVE parameters, and here the fragility tracks the
INACTIVE ONES
-> interpretation: the EXTRA PARAMETERS are where the
MEMORISATION LIVES
THE NUMBERS GIVE THE SHAPE OF THE DEGRADATION AND OF THE TRADE
80M DENSE models can repeat data over 8x with MINIMAL degradation
MOEs instead begin to suffer at 4x, and DETERIORATE RAPIDLY
they CEDE their performance benefits in all-unique data settings to
UNDERPERFORM DENSE MODELS after 32x
<- read together, these describe a CROSSOVER: the sparse
architecture WINS on unique data and LOSES on heavily repeated
data, with the reversal somewhere between 4x and 32x
THEN THE REMEDY, REPORTED WITH ITS OWN LIMIT
some existing regularization can MITIGATE OVERFITTING -- DROPOUT
in particular
with STRONG MASKING-BASED REGULARIZATION, MoEs OUTPERFORM DENSE
MODELS even when data is repeated MORE THAN 64 TIMES
BUT: NO METHOD FULLY MATCHES THE PERFORMANCE OF ALL-UNIQUE
TRAINING DATA
<- regularization SHIFTS THE CROSSOVER rather than REMOVING THE
PENALTY: the honest form of a mitigation result
AND THE MECHANISM ANALYSIS CONNECTS BEHAVIOUR TO INTERNALS
MoE ROUTING UNIVERSALLY STABILIZES EARLY in training
EXPERT SPECIALIZATION CORRELATES WITH OVERFITTING TO REPEATED DATA
<- the LOAD-BEARING one: overfitting is associated with experts
SPECIALISING
-> suggests the failure is NOT GENERIC CAPACITY EXCESS but a
SPECIFIC DIVISION OF LABOUR that repeated data ENCOURAGES
-> points future work in a DIRECTION -- REDUCING
OVER-SPECIALIZATION BY DISRUPTING MEMORIZATION PATTERNS --
rather than merely asking for BETTER REGULARIZERS
Think of it as a large team where each member becomes an expert in one specific customer’s quirks. With a diverse customer base that is valuable: each specialist handles their cases well. Repeat the same few customers over and over and the specialisation hardens around their idiosyncrasies — the team gets better at those exact cases and worse at everything else, which is not what more practice should do. That is the shape of the finding: the fragility grows with the number of specialists asked for (total parameters), not with how many are active on any given case, because the specialisation itself is where the memorisation is stored. And it explains why regularization only shifts the crossover: dropout and masking spread the load, but they do not stop specialists from forming.
Key Concepts
- Repetition as a standard practice with an unexamined interaction: the effect was characterised for dense models and not for sparse ones. The consequence is that the efficiency-motivated architecture depends on a practice nobody checked it against.
- Degradation tracking total rather than active parameters: fragility scales with how many parameters exist, not how many fire. It inverts the quantity MoE reasoning usually centres on.
- The crossover between architecture classes: dense wins under heavy repetition, sparse wins on unique data, with reversal between 4× and 32×. It makes the choice conditional on the data regime rather than absolute.
- Regularization that shifts rather than removes: masking-based methods let MoEs beat dense models at over 64× repetition, without matching all-unique training. The remaining gap is the honest part of the remedy.
- Specialization as the correlate of overfitting: expert specialisation rises with repeated data. It locates the failure in a division of labour rather than in raw capacity.
Framework Shift
Before (repetition characterised for dense models):
repeat data as human text runs out
-> the effect is known for dense transformers
-> sparse architectures assumed to behave comparably
-> no account of the interaction with sparsity
After (the interaction measured, cross-class):
repetition rates swept across mixes and MoE configurations
-> MoEs degrade faster (80M-1B active, up to 8.5B total)
-> the effect scales with TOTAL parameters, not active
-> dense wins under heavy repetition; sparse wins on unique data
-> masking-based regularization shifts the crossover but does not
remove the penalty
-> expert specialization correlates with the overfitting
From assuming that an efficiency-motivated architecture inherits the known properties of repetition, to measuring a crossover between architecture classes and locating it in expert specialisation, the core shift is that sparsity changes how a model responds to repeated data, and it does so through inactive parameters.
Expert Assessment
Problem choice: Excellent, and the practical framing is what makes it matter. Repetition is already standard rather than hypothetical, and the architecture being tested is the one chosen for efficiency — so an unexamined interaction between the two is exactly the kind of dependency worth finding before it is relied on.
Method maturity: The sweep across three axes — repetition level, data mix, and MoE configuration including expert count and granularity — is what permits attribution rather than description. The scale range from 80M to 8.5B total rules out a small-model explanation. And the mechanism analysis is what elevates the paper above a scaling study: finding that routing stabilizes early and that expert specialisation correlates with the overfitting gives the result a cause, and it points the suggested remedy at over-specialisation rather than at regularisation in general.
Experimental integrity: The comparison is honest in a subtle way: models are matched on active parameters, which is the fair basis for inference cost, and the paper then reports that the fragility tracks total parameters, which is the axis on which they are not matched. That is a finding that cuts against the convenience of the setup, and reporting it is what makes the mechanism claim credible. The mitigation result is also reported with its limit intact — no method matches all-unique data — rather than being presented as a solution.
Writing quality: The numbers are arranged to reveal the crossover rather than to show a single degradation figure, which is the right presentation for a trade-off. Because the practical reader wants to know what to do, a short passage translating the mechanism into a diagnostic — for instance whether monitoring expert specialisation could serve as an early warning — would make the finding actionable rather than merely explanatory.
Verdict: strong accept — it measures an unexamined interaction between a standard training practice and the architecture chosen for efficiency, traces the fragility to inactive parameters, and locates the mechanism in expert specialisation.
Takeaways
- Check interactions, not just components. Repetition was characterised for dense models, and the architecture adopted for efficiency inherits an unverified assumption.
- Watch which parameter count a property tracks. Fragility following total rather than active parameters inverts the quantity that sparsity is usually justified by.
- Present trade-offs as crossovers. Dense versus sparse reverses somewhere between 4× and 32× repetition, which makes the choice a regime question rather than a verdict.
- Look for the specialisation behind overfitting. If experts hardening on repeated data is the mechanism, spreading the load is not the same fix as preventing the hardening.
论文: 2609.11917 作者: Atindra Jha, Margaret Li, Jure Leskovec, Percy Liang, Luke Zettlemoyer 分类: cs.CL, cs.LG
缺口
随着人类撰写的文本供给被耗尽,重复训练数据已成常规做法。所以”重复”的性质不是一个边缘问题;它是当下模型如何被训练的一个属性。
而这个问题是在错误的架构上被回答的。此前的工作研究了稠密激活 Transformer 上的数据重复,而对近来占据主导的稀疏架构(如专家混合 MoE)而言,数据重复的影响仍基本未被探索——尽管它们有着更高的算力效率。 正是这个组合让缺口变得有后果:MoE 的全部卖点是单位质量下的算力效率,而它的效率来自每个 token 只激活一部分参数。如果稀疏激活与重复数据相互作用得很糟,那么人们为了效率而采用的这个架构,就对一项日益常规的训练做法有着未经审视的依赖。
重复已是常规,而它的影响是在「错误的架构」上被研究的
人类撰写文本的供给「被耗尽」
-> 「重复训练数据已成常规做法」
-> 所以"重复"的性质不是「边缘问题」;
它是「当下模型如何被训练」的一个属性
|
v
而这个问题是在「错误的架构」上被回答的
此前工作:稠密激活 Transformer 上的数据重复
基本未被探索:「近来占据主导的稀疏架构」,
如专家混合 MoE——「尽管」它们有着更高的算力效率
|
v
[为什么这个组合让缺口有后果]
MoE 的全部卖点是「单位质量下的算力效率」
它的效率来自「每个 token 只激活一部分参数」
|
v
-> 如果「稀疏激活」与「重复数据」相互作用得很糟,
那么人们为了效率而采用的这个架构,就对一项
「日益常规的训练做法」有着「未经审视的依赖」
增量
一句话: 在这篇论文之前,重复对专家混合的影响是未知的;在这篇论文之后,MoE 退化更快、这种效应通过总参数量(而非激活参数量)随稀疏度放大,而基于掩码的正则化能在重度重复下挽回优势。
核心机制
这个扫描的设计,是为了把”单次比较会混淆的那些变量”分开:重复率在单域与多域数据混合上变化,同时在包括专家数与粒度在内的 MoE 配置上变化。 三条轴同时变化——重复多少、数据是什么、以及稀疏架构如何配置。正是这种广度,让结论可以被归因、而不只是被观察。
头条结果是在匹配规模下跨架构类别的比较:在 80M 到 1B 激活(8.5B 总)参数的模型范围内,MoE 在数据重复下退化得更快。 这句话有两点要紧。范围跨越了一个数量级以上,所以这个效应不是小模型的产物。而**“激活(总)“这个措辞是刻意的——模型在激活参数上被匹配、而在总参数上不同;这对推理成本是公平的比较,也恰恰是机制发现的来源**。
那个机制发现是最锋利的:这种效应随稀疏度增强,而由「总参数量」而非「激活参数量」决定。 也就是说,相关的量是存在多少参数,而不是每个 token 有多少被激活。这反直觉,而且正是该架构的前提不会预测的:关于 MoE 效率的推理通常围绕激活参数展开,而这里的脆弱性跟随的是未被激活的那些。解释是:那些多出来的参数,正是记忆化所栖身之处。
数字给出了退化的形状与取舍的形状:
- 80M 的稠密模型可以把数据重复 8 倍以上而退化极小。
- MoE 反而从 4 倍起就开始受损,并迅速恶化。
- 它们会在”数据全不重复”的设定下让出自己的性能优势,在 32 倍之后反而不如稠密模型。
把这三行一起读,它们描述的是一次交叉:稀疏架构在数据不重复时胜出、在重度重复时落败,而反转点大致落在 4 倍到 32 倍之间。
接着是补救,连同它自己的限界一起报告。 一些现有正则化方法能缓解过拟合——尤其是 dropout;而在强力的基于掩码的正则化下,即便数据被重复 64 次以上,MoE 也能胜过稠密模型。 但是:没有任何方法能完全追平”数据全不重复”的表现。 所以正则化移动了交叉点,而不是移除了惩罚——这是缓释类结果诚实的形态。
而机制分析把行为与内部状态连起来:MoE 的路由在训练早期普遍会稳定下来,而「专家分化」与”对重复数据的过拟合”相关。 两条发现,第二条承重:过拟合与专家分化相关联,这提示该失效不是笼统的容量过剩,而是重复数据所鼓励的某种特定分工。这也把建议的后续工作指向一个方向——通过打破记忆化模式来减少过度分化——而不只是”要更好的正则化器”。
扫描的设计,是为了把"单次比较会混淆的变量"分开
重复率在「单域与多域」数据混合上变化
「并且」在包括「专家数与粒度」的 MoE 配置上变化
<- 「三条轴同时」变化:重复多少、数据是什么、
以及稀疏架构如何配置
-> 广度让结论可以被「归因」,而不只是被「观察」
「头条结果:在匹配规模下跨架构类别的比较」
在 80M 到 1B「激活」(8.5B「总」)参数的模型范围内,
「MOE 在数据重复下退化得更快」
<- 范围跨越「一个数量级以上」
-> 不是「小模型的产物」
<- "激活(总)"是「刻意的」:在「激活」参数上匹配、
而在「总」参数上不同
-> 对「推理成本」公平,也恰恰是「机制发现的来源」
「机制发现是最锋利的」
这种效应「随稀疏度增强」,而由「总参数量」而非
「激活参数量」决定
-> 相关的量是「存在多少参数」,而不是每个 token
有多少被激活
<- 「反直觉」,而且正是该架构的前提「不会预测」的:
关于 MoE 效率的推理通常围绕「激活参数」,
而这里的脆弱性跟随「未被激活的那些」
-> 解释:「那些多出来的参数,正是记忆化所栖身之处」
「数字给出了退化与取舍的形状」
80M 的「稠密」模型可以把数据重复 8 倍以上而退化极小
MOE 反而从 4 倍起就开始受损,并「迅速恶化」
它们会在"数据全不重复"的设定下让出优势,
「在 32 倍之后反而不如稠密模型」
<- 一起读,这三行描述的是一次「交叉」:
稀疏架构在「数据不重复」时胜出、在「重度重复」时落败,
而反转点大致落在 4 倍到 32 倍之间
「补救,连同自己的限界一起报告」
一些现有正则化能「缓解过拟合」——尤其是「DROPOUT」
在强力的「基于掩码的正则化」下,即便数据被重复
「64 次以上」,MOE 也能「胜过」稠密模型
但是:「没有任何方法能完全追平"数据全不重复"的表现」
<- 正则化「移动了交叉点」,而不是「移除了惩罚」:
缓释类结果诚实的形态
「机制分析把行为与内部状态连起来」
MOE 的「路由」在训练早期普遍会「稳定下来」
「专家分化」与"对重复数据的过拟合"「相关」
<- 「承重」的那一条:过拟合与专家「分化」相关联
-> 提示该失效「不是笼统的容量过剩」,
而是重复数据所「鼓励的某种特定分工」
-> 把后续工作指向一个「方向」——「通过打破记忆化模式
来减少过度分化」——而不只是"要更好的正则化器"
可以用**“一个很大的团队,其中每位成员都变成某一位特定客户怪癖的专家”来理解这件事: 客户群多样时这很有价值:每位专家都能把自己的案子处理得很好。而把同样那几位客户反复重复,分化就会围绕着他们的怪癖固化下来——团队在这些确切案例上变得更强**、在其他一切上变得更差,而这并不是”更多练习”应当带来的结果。 这就是这个发现的形状:脆弱性随被要求有多少位专家(总参数量)增长,而不是随”任何一次任务上激活了多少位”增长——因为分化本身就是记忆化的存放之处。 这也解释了为什么正则化只是移动了交叉点:dropout 与掩码把负载摊开,却阻止不了专家成形。
关键概念
- 一项常规做法与一个未经审视的交互: 该效应在稠密模型上被刻画过,在稀疏模型上没有。后果是:为效率而选的架构,依赖着一项没人拿它检查过的做法。
- 退化跟随「总参数」而非「激活参数」: 脆弱性随”存在多少参数”放大,而不随”激活多少”放大。它把 MoE 推理通常聚焦的那个量倒转了。
- 架构类别之间的交叉: 重度重复下稠密胜、数据不重复时稀疏胜,反转在 4 倍到 32 倍之间。它让这个选择成为区间问题,而不是定论。
- 移动而非移除的正则化: 基于掩码的方法让 MoE 在 64 倍以上重复时能胜过稠密模型,却仍追不平”全不重复”。剩下的差距是补救中诚实的部分。
- 以「分化」作为过拟合的相关物: 专家分化随重复数据上升。它把失效定位在分工上,而不是原始容量上。
框架转变
之前(重复的性质在稠密模型上被刻画):
人类文本用尽,于是重复数据
-> 该效应在稠密 Transformer 上已知
-> 假定稀疏架构表现类似
-> 对"与稀疏度的交互"没有任何说法
之后(跨类别测出该交互):
重复率在多种数据混合与 MoE 配置上被扫描
-> MoE 退化更快(80M~1B 激活,总参最高 8.5B)
-> 效应随「总参数」放大,而非激活参数
-> 重度重复下稠密胜;数据不重复时稀疏胜
-> 基于掩码的正则化移动了交叉点,但没移除惩罚
-> 专家分化与过拟合相关
从”假定一个为效率而选的架构会继承重复的已知性质”,转变为”测出架构类别之间的交叉、并把定位落到专家分化上”,核心转变在于:稀疏性改变了模型对重复数据的响应方式,而它经由「未被激活的参数」起作用。
专家评审
选题眼光: 极好,而实践性的框定才是让它要紧的原因。 重复已经是常规而非假设;而被测的架构正是为效率而选的那一个——所以”两者之间未经审视的交互”,恰恰是值得在被依赖之前找出来的那种依赖。
方法成熟度: 三条轴的扫描——重复水平、数据混合、以及含专家数与粒度的 MoE 配置——才使归因而非描述成为可能。 80M 到 8.5B 总参数的跨度排除了”小模型解释”。而机制分析把论文提升到”一项扩展研究”之上:发现”路由早期稳定”与”专家分化与过拟合相关”,给了这个结果一个原因,并把建议的补救指向过度分化,而不是泛泛的正则化。
实验诚意: 比较在一种微妙的意义上是诚实的:模型在激活参数上被匹配——这是推理成本的公平基准——而论文随后报告”脆弱性跟随总参数”,而那正是它们未被匹配的那条轴。这是一个与设定便利相逆的发现,把它报出来才让机制主张可信。 缓释结果也保留了限界——没有方法能追平”全不重复”——而不是被呈现为解法。
写作功力: 数字被安排成揭示交叉、而不是展示单一退化数值,对一项取舍而言这是正确的呈现方式。 由于实践型读者想知道”该做什么”,若能补一小段把机制翻译成诊断——比如”监控专家分化能否作为早期预警”——会让这个发现可操作,而不只是可解释。
判决: 强接收(Strong Accept) — 它测出了一项常规训练做法与”为效率而选的架构”之间未经审视的交互,把脆弱性追溯到未被激活的参数,并把机制定位在专家分化上。
要点总结
- 检查交互,而不只是组件。重复的性质在稠密模型上被刻画过,而为效率而采用的架构继承了一个未经核实的假设。
- 留意一个性质跟随的是哪个参数量。在稀疏架构中,脆弱性跟随总参数而非激活参数,这倒转了人们为其效率辩护时所用的那个量。
- 把取舍呈现为交叉。稠密与稀疏在 4 倍到 32 倍重复之间反转,这让选择成为区间问题而不是定论。
- 去找过拟合背后的分化。如果机制是”专家在重复数据上变硬”,那么摊开负载与阻止变硬就不是同一种修法。