Paper: 2609.20807 Authors: Anton Xue, Litu Rout, Aditya Akella, Adam Klivans, Sujay Sanghavi Categories: cs.CL, cs.LG
The Gap
Diffusion Language Models (DLMs) represent one of the most promising alternatives to standard autoregressive (AR) generation. By denoising all tokens iteratively rather than emitting tokens strictly left-to-right, DLMs unlock true parallel decoding, arbitrary text infilling, and controllable controllable planning.
Training a diffusion language model from scratch is prohibitively expensive. The standard industry recipe is AR-to-DLM adaptation: take an existing, well-trained autoregressive model and adapt its attention weights to bidirectional diffusion via masked training.
However, the frontier of foundation models has evolved away from pure full-attention Transformers. Architectures like Qwen3.5 increasingly adopt hybrid backbones that interleave quadratic attention layers with linear recurrent (RNN / SSM) layers to tame long-context KV cache overhead.
This introduces a severe architectural dilemma: linear RNNs are fundamentally causal. Recurrent state updates strictly step forward through time. Diffusion denoising, conversely, demands bidirectional context—token must attend simultaneously to future tokens and past tokens . Conventional wisdom assumed hybrid models were fundamentally incompatible with diffusion adaptation without gutting the recurrent layers.
THE ADAPTATION DILEMMA
Autoregressive Hybrid Model (e.g. Qwen3.5)
[ Attention Layer ] <--- Can be bidirectionalized (unmask attention matrix)
|
[ Recurrent / RNN ] <--- Structurally Causal: h_t = f(h_{t-1}, x_t)
| How can token 3 attend to token 10?
v
Diffusion Language Model Requirement:
Global Bidirectional Context at every denoising timestep
|
v
Conventional Assumption:
"Hybrid backbones cannot be adapted to diffusion without replacing RNNs"
|
v
dQWEN3.5 DISCOVERY:
Hybrid backbones adapt gracefully!
Interleaved attention layers distribute bidirectional signal,
while RNN layers preserve compact local inductive bias.
Reaches target loss in 50% fewer training tokens than full-attention.
The Increment
One sentence: Counter to the belief that causal recurrent layers disqualify hybrid architectures from diffusion modeling, dQwen3.5 demonstrates across 0.8B to 9B parameters that hybrid attention-RNN backbones can be converted into diffusion language models, halving the required adaptation tokens compared to full-attention controls while preserving strong parallel decoding and arbitrary infilling.
Core Mechanism
The dQwen3.5 framework systematically evaluates the adaptation of hybrid Qwen3.5 backbones across four model scales: 0.8B, 2B, 4B, and 9B parameters.
The adaptation mechanism operates under a surprisingly elegant design:
- Partial Bidirectionalization: The full-attention layers in the hybrid stack have their causal triangular masks removed, allowing bidirectional token mixing.
- Causal RNN Preservation: Rather than converting the linear recurrence into an expensive bi-directional scan (which would forfeit inference efficiency and break pretrained recurrent weights), the RNN layers are kept causal.
- Interleaved Representation Diffusion: The alternating architecture acts as a two-phase information exchange. The bidirectional attention layers propagate global foresight across the full context window, while the forward recurrent layers compress and consolidate representations.
dQWEN3.5 LAYER INTERLEAVING PIPELINE
Masked / Noisy Token Input (t)
|
v
+------------------------------------+
| Bidirectional Attention Layer | -> Global cross-token visibility
| (Causal mask removed: i <-> j) | (Propagates future + past context)
+------------------------------------+
|
v
+------------------------------------+
| Causal Recurrent / RNN Layer | -> Compact sequence consolidation
| (Kept strictly forward causal) | (Reuses pretrained linear weights)
+------------------------------------+
|
v
[ Repeated Across 0.8B - 9B Scales ]
|
v
Parallel Denoised Token Predictions
The headline finding is that this hybrid topology is not a compromised second-best: it beats the full-attention control in sample efficiency. Across equal token budgets, dQwen3.5 converges to a target validation perplexity using approximately half the tokens () needed by a standard full-attention model.
The structural metaphor is a high-speed highway with alternating flyovers and one-way tunnels.
- Traditional full-attention diffusion is an open circular grid where every car can see and steer toward any other car in any direction simultaneously. Building and maintaining this grid at scale is immensely costly.
- A naive reading argued that putting a one-way tunnel (causal RNN) onto the highway would make circular two-way travel impossible.
- dQwen3.5 shows that you don’t need every stretch of road to be two-way. As long as you have elevated flyovers (bidirectional attention layers) that let drivers see the entire terrain from above before plunging into each local one-way tunnel, information flows everywhere it needs to go. The cars reach their destinations twice as fast because one-way tunnels prevent traffic gridlock.
Key Concepts
- Diffusion Language Model (DLM): A generative model for discrete text that starts from completely masked or corrupted tokens and refines them simultaneously over multiple denoising steps, enabling parallel output generation.
- Hybrid Attention-RNN Backbone: Modern architectures (such as Qwen3.5, Mamba-Transformer hybrids) that replace a large fraction of self-attention layers with linear state-space or recurrent layers.
- Sample Efficiency in Model Recycling: Rather than spending millions of dollars pre-training diffusion models from scratch, recycling existing pre-trained autoregressive checkpoints is essential for scalable DLM research.
Framework Shift
Before (Architectural Dogma):
Diffusion models strictly demand 100% bidirectional layers everywhere
-> Hybrid models with causal RNNs dismissed as incompatible
-> DLM researchers restricted to pure full-attention backbones (expensive KV)
After (dQwen3.5 Hybrid Proof-of-Concept):
Interleaved bidirectional attention + causal RNNs works seamlessly
-> Reaches target training loss in half the tokens (~2x sample efficiency)
-> Supports any-order masking, text infilling, and parallel decoding
-> Opens up next-generation hybrid foundation models for diffusion adaptation
From “assuming causal RNN layers make hybrid models unsuitable for diffusion,” the core shift is demonstrating that interleaving bidirectional attention with forward-causal recurrence actually accelerates diffusion adaptation.
Expert Assessment
Problem choice: Timely and strategic. As the open-weights ecosystem shifts en masse toward hybrid attention-SSM/RNN models to slash serving memory, understanding how to adapt these models into diffusion generators is critical for the survival of DLMs.
Method maturity: The method is clean because of what it chooses not to do. It avoids complex bidirectional RNN reformulations that would degrade throughput or destroy pretrained weights, testing the simplest possible intervention: unmask the attention layers, keep the RNNs intact.
Experimental integrity: Strong scaling rigor. Testing four discrete scales (0.8B, 2B, 4B, 9B) with explicit full-attention control baselines proves that the 2x sample efficiency gain is a structural property, not a small-scale fluke. Any-order generation and parallel decoding benchmarks confirm that diffusion capabilities were not compromised.
Writing quality: Transparent and crisp. The paper acknowledges the theoretical tension upfront and lets rigorous convergence curves provide the empirical answer.
Verdict: strong accept — A groundbreaking architectural result that bridges hybrid recurrent architectures with discrete diffusion generation.
Takeaways
- Do not discard pretrained hybrid models when exploring diffusion generation; causal RNN layers compose effectively with bidirectional attention.
- When adapting an autoregressive hybrid backbone to a diffusion model, simply remove the causal mask from the attention layers and leave the recurrent layers untouched.
- Expect roughly faster training convergence to target loss compared to adapting an all-attention transformer of equivalent parameter scale.
论文: 2609.20751 作者: Anton Xue, Litu Rout, Aditya Akella, Adam Klivans, Sujay Sanghavi 分类: cs.CL, cs.LG
缺口
扩散语言模型(Diffusion Language Models, DLM)正被视为打破传统自回归(Autoregressive, AR)生成模式的最强候选者。 自回归模型必须从左到右严格逐字吐词,而扩散语言模型通过对整段文本的掩码同时去噪,天然具备任意顺序解码、局部填空(Infilling)以及一次性并行生成的超强能力。
然而,从零开始预训练一个扩散语言模型成本高昂。 业界目前的标准打法是借鸡生蛋(AR-to-DLM Adaptation):直接拿现成训练好的自回归开源大模型,通过掩码扩散任务将其注意力权重改造为双向扩散模型。
问题在于,开源基座模型的底层架构已经发生了剧烈演变。 以 Qwen3.5 为代表的新一代架构,为了遏制超长上下文下的 KV Cache 膨胀,普遍采用了注意力与线性循环(RNN / SSM)交替堆叠的混合架构(Hybrid Architecture)。
这就带来了严重的架构阻碍:线性 RNN 在结构上是严格因果单向的。 循环隐状态的转移方程 严格随时间从前向后流动。 但扩散模型的去噪核心是全向双向感知——位于中间位置的 Token 必须能够同时纵观上文和下文。 学术界此前普遍默认:混合架构中的因果单向 RNN 存在天然结构缺陷,若不彻底拆毁或重写其循环层,根本不可能改造为合格的扩散模型。
改造路线的架构撞墙困境
先进自回归混合架构 (如 Qwen3.5)
[ 注意力层 (Attention) ] <--- 容易双向化(只需拿掉下三角因果 Mask)
|
[ 线性循环层 (RNN/SSM) ] <--- 结构上严格因果单向:h_t = f(h_{t-1}, x_t)
| 第 3 个词怎么可能看到第 10 个词?
v
扩散语言模型(DLM)的硬性要求:
去噪全过程要求全局双向无死角视野
|
v
学术界既有固有成见:
「因果单向的 RNN 结构无法做双向扩散,混合架构模型被排除在 DLM 之外」
|
v
dQWEN3.5 突破性实证发现:
混合架构不仅能完美适配扩散去噪,甚至更强!
双向注意力负责横向扩散全局远见,单向循环负责纵向压缩紧凑归纳。
相比全注意力对照组,仅需一半(50%)的 Token 即可收敛至同等损失。
增量
一句话: 颠覆了「因果循环层无法适配双向扩散」的传统偏见,dQwen3.5 在 0.8B 至 9B 全尺寸上证明了混合注意力-RNN 架构可以无缝转化为高性能扩散语言模型,且相比纯全注意力模型节省了约一半的微调 Token,同时完整保留了任意序与并行解码优势。
核心机制
研究团队在 Qwen3.5 的 4 个经典参数体量(0.8B、2B、4B 与 9B)上全面展开了扩散适配验证。
改造的核心设计出人意料地克制而优雅:
- 局部非对称双向化:仅将混合架构中现存的注意力层解禁,去除其因果下三角掩码,使其具备全向交叉注意力能力。
- 完全保留单向因果循环:不对线性 RNN 层做任何复杂的「双向扫描重写」(那不仅会破坏推理效率,还会摧毁预训练沉淀的线性权重),而是让循环层继续保持纯粹的前向因果流动。
- 交替分工的表征扩散机制:混合交叠结构在深层形成了一种节奏鲜明的双相信息加工循环——双向注意力层在宏观上拉通全序列的全局上下文与前后预见;单向循环层则在局部高效压缩、梳理并固化表征。
dQWEN3.5 分层交替去噪数据流
掩码加噪文本序列 (t)
|
v
+------------------------------------+
| 全向双向注意力层 (Bidirectional Attn)| -> 解禁因果掩码,打破前后边界
| (各 Token 彼此全局可视互通) | (全局传递未来的前瞻信息)
+------------------------------------+
|
v
+------------------------------------+
| 原生前向循环层 (Causal RNN / SSM) | -> 严格保持单向因果压缩
| (直接继承预训练权重,高效固化) | (提供稳固的序列局部归纳偏置)
+------------------------------------+
|
v
[ 在 0.8B 至 9B 各层深度交替循环 ]
|
v
高质量并行去噪预测输出
最震撼的发现是:这种看似在因果性上「妥协」的混合架构,其实际训练样本效率直接碾压了纯全注意力模型。 在相同的训练预算下,dQwen3.5 达到目标验证集困惑度所消耗的训练 Token 仅为全注意力对照组的一半左右(约节省 50%)。
这里的核喻是一条高架立交桥与单向深埋隧道交替穿插的高速交通网。
- 传统的全注意力扩散模型,就像把整座城市建在一个没有任何隔离护栏的无边界大广场上,每一辆车都能同时斜向穿行到任意角落。 在理论上视野拉满,但随着车流增多,维护这种全向通行的管制成本极高。
- 传统教条认为,只要路网里插进了一条只能向前开的单向隧道(单向 RNN),两辆车就无法相互联络,整个广场就瘫痪了。
- dQwen3.5 证明了城市并不需要全向无死角:只要在关键节点设立高架立交观景桥(双向注意力层),让每辆车在进隧道前俯瞰全城路况,随后驶入高效通过的单向隧道(单向循环),车辆就能极其顺畅地抵达目的地。 因为单向隧道极大地规范了局部的交通流向,整座城市的通行效率反而比乱哄哄的空旷广场快了一倍。
关键概念
- 扩散语言模型(DLM):一种用于离散文本生成的非自回归范式。 从完全随机的掩码符号出发,像雕刻大理石一样经过若干步去噪迭代同步还原完整文本,具备突破顺序束缚的并行生成潜力。
- 混合注意力-循环架构(Hybrid Backbone):当代主流开源大模型的新范式。 用计算复杂度为 的状态空间模型(SSM)或线性循环单元替代大部分高消耗的 注意力层,极大压低 KV Cache 显存开销。
- 预训练资产转化效率(Adaptation Sample Efficiency):如何用最少的算力与 Token 预算,将已经成熟的自回归千亿模型资产低成本迁移至扩散等新型生成范式,是 DLM 能否走出象牙塔的生命线。
框架转变
之前(传统架构教条):
扩散模型必须要求网络内 100% 所有层均具备全向双向性
-> 判定混合架构中的因果 RNN 具有原罪,无法用于扩散生成
-> 扩散研究者只能死守高开销的全注意力模型,受制于庞大的显存瓶颈
之后(dQwen3.5 混合解耦):
双向全局注意力 + 单向因果 RNN 混合交替,完全胜任扩散建模
-> 仅需一半 Token 即可达到同等预训练损失(2 倍样本利用率)
-> 原生继承任意序填充、文本重写与并行解码能力
-> 为大批拥有先进混合架构的开源基座模型打开了通往扩散模型的大门
从「认定单向循环组件是扩散去噪的死敌」,核心转变在于:证明了双向注意力与单向因果循环的交叠不仅不会阻断扩散去噪,反而能借助更紧凑的归纳偏置让适配收敛速度翻倍。
专家评审
选题眼光: 极具时效性与战略嗅觉。 开源社区正全面走向以 Qwen3.5 为代表的混合注意力架构,在此节点解答「混合模型能否做扩散」这一关键疑难,为扩散语言模型的生态延续铺平了道路。
方法成熟度: 胜在克制。 团队没有去搞把 RNN 强行双向展开的复杂花架子(那会彻底丢掉推理效率),而是采取了最朴素直接的手术:只拿掉注意力层的掩码,循环层原封不动,却收到了意想不到的奇效。
实验诚意: 尺度跨度扎实(0.8B、2B、4B、9B 四个量级完整对齐)。 设置了控制变量完全一致的全注意力对照组,确认了「2 倍样本效率」是底层结构的系统性收益而非小规模偶然现象,并在并行解码与填空评测中充分验证了功能完整性。
写作功力: 直击矛盾焦点,直面理论冲突并用翔实的实验收敛曲线予以解答。
判决: 强接收 (strong accept) — 语言模型底层架构探索的破局之作,成功拆除了混合架构与扩散生成之间的学术藩篱。
要点总结
- 不要轻易放弃现有的混合架构开源模型资产;因果单向的 RNN 算子完全可以与双向注意力协同胜任离散扩散任务。
- 在把自回归混合模型改造成扩散模型时,只需移除注意力层的下三角因果 Mask,循环层保持原样不动即可。
- 混合架构在扩散适配任务中展现出惊人的样本效率,训练收敛至同等困惑度所需 Token 量相比全注意力模型直接减半。