Paper: 2609.20800 Authors: Taoyong Cui, Zhongyao Wang, Xinyue Xu, Weiyang Liu, Zhaochen Yu Categories: cs.CL, cs.AI, cs.LG

The Gap

World models serve as the mental simulation engines of intelligence: they anticipate future states, forecast the consequences of candidate interventions, and allow agents to plan without burning real-world trials.

Yet existing predictive architectures remain fiercely siloed. Video models (e.g. V-JEPA, Sora) operate over spatiotemporal visual patches; molecular simulators run specialized graph neural networks; climate surrogates rely on spherical harmonics; and biological systems depend on gene-regulatory graphs.

The underlying question has sat unanswered: Can a single, unified predictive learning principle govern world modeling across radically heterogeneous dynamical systems?

Vanilla Joint-Embedding Predictive Architectures (JEPAs) eliminate pixel-level generative waste by predicting in latent abstract feature space. However, when applied to complex physical and scientific systems, vanilla JEPAs suffer from factor entanglement: predictive latents lump fast vs. slow modes, invariant versus intervened factors, and global conservation laws into a single undifferentiated vector. This entanglement leads to cascading error drift over long rollouts and renders targeted interventions unpredictable.

   THE CHALLENGE: FRAGMENTED DYNAMICS PREDICTION

   Visual Worlds      Molecular Dynamics     Biological Cells      Celestial Orbits
   (Pixels/Patches)   (Atoms & Forces)       (Gene Regulatory)     (Gravity/Kepler)
         \                  |                      |                     /
          \                 |                      |                    /
           v                v                      v                   v
     [ Domain-Specific Models ] [ Custom Simulators ] [ Statistical Trajectories ]
                             |
                             v
   UNIFIED QUESTION: Can one predictive formulation model them all?
                             |
                             v
   LIMITATION OF VANILLA JEPA:
     Single monolithic latent space -> Fast/slow modes entangled
     -> Long-horizon rollout collapses -> Interventions unpredictable
                             |
                             v
   SOLUTION: JEPA-ANYTHING + Orthogonal Predictive Factorization (OPF)
     Latent Target Decomposed into Orthogonal Dedicated Pathways

The Increment

One sentence: JEPA-Anything extends Joint-Embedding Predictive Architectures into a domain-agnostic foundation via Orthogonal Predictive Factorization (OPF), decomposing complex state transitions into complementary latent factors that yield lower rollout error across 7 distinct scientific domains and accurately recover Keplerian orbital laws.

Core Mechanism

Rather than predicting the entire subsequent state embedding as a monolithic vector zt+1z_{t+1}, Orthogonal Predictive Factorization (OPF) dissects the target latent manifold into mutually complementary sub-spaces:

  1. Orthogonal Pathway Decomposition: The target state representation is factored into orthogonal subspaces: z=⨁k=1Kz(k),subject to ⟨z(i),z(j)⟩≈0(∀i≠j)z = \bigoplus_{k=1}^K z^{(k)}, \quad \text{subject to } \langle z^{(i)}, z^{(j)} \rangle \approx 0 \quad (\forall i \neq j) Each subspace captures distinct dynamics: e.g., conserved geometric invariants, high-frequency localized vibrations, and controllable action-sensitive transitions.
  2. Dedicated Predictor Branches: Each factor is predicted along its own parameterized path, conditioned on prior history and external interventions.
  3. Recombinant Co-Attention: The factorized predictions are recombined through a shared interaction layer that enforces joint consistency without allowing cross-factor collapse.
   ORTHOGONAL PREDICTIVE FACTORIZATION (OPF)

   Input Context (t) -----------------------------> Factor Predictor
          |                                               |
          v                                               v
   +---------------+                              +---------------+
   | Shared Trunk  |                              | Dedicated     |
   | State Encoder |                              | Pathways      |
   +---------------+                              +---------------+
          |                                        |      |      |
          v                                        v      v      v
   Latent Factors: [ Conservation ] [ Fast Fluctuation ] [ Action Shift ]
          \                  |                  /
           \                 |                 /
            v                v                v
      +----------------------------------------------+
      |        Orthogonal Alignment Constraint       |
      +----------------------------------------------+
                             |
                             v
               Predicted Next Latent z_{t+1}

The structural metaphor is a symphony orchestra’s multi-track mixing console. A vanilla predictive model is like recording an entire orchestra with a single cheap omnidirectional microphone. If you try to predict how the piece will sound two measures later, the booming timpani drowns out the quiet oboe melody, and you cannot adjust or “intervene” on the violin section without distorting the whole track. JEPA-Anything equips the conductor with isolated directional microphones (dedicated orthogonal pathways) for each instrumental section—strings, brass, woodwinds, percussion. Each track is predicted along its own acoustic logic (timpani follow rhythm; flutes follow melodic phrasing) and then recombined at the master mixing board. Because the tracks are cleanly isolated, you can change a single violin note (an intervention) without breaking the tempo of the drums.

Key Concepts

  • Joint-Embedding Predictive Architecture (JEPA): A paradigm proposed by Yann LeCun that learns representations by predicting representations of masked or future states rather than reconstructing raw pixels or sensor noise.
  • Orthogonal Predictive Factorization (OPF): Decomposing high-dimensional dynamical transitions into linearly or geometrically independent latent sub-factors to avoid feature entanglement during multi-step auto-regressive rollout.
  • Intervention Generalization: The ability of a world model to accurately predict system response when a variable is actively manipulated (counterfactual/do-calculus), tested from classic Atari games to biological organoids.

Framework Shift

Before (Domain-Specific or Entangled Latent Models):
  Domain X -> Specialized Architecture -> Monolithic Latent Transition
  -> Invariant properties bleed into transient fluctuations
  -> Multi-step rollouts drift into nonsense after 10-20 steps
  -> Interventions perturb un-targeted physical variables

After (Domain-Agnostic Factorized Dynamics):
  Any Dynamical System -> Generic Transformer Trunk -> Orthogonal Factor Pathways
  -> Conserved quantities preserved separately from high-frequency noise
  -> Stable 100-step molecular dynamics rollouts
  -> Verified physical constants: fitted Kepler exponent of -1.4991 (theory: -1.5)

From “designing custom neural PDE simulators or video generators for each scientific vertical,” the core shift is unifying heterogeneous dynamics under an orthogonal, factorized joint-embedding objective.

Expert Assessment

Problem choice: Audacious and fundamental. Tackling world modeling not as a narrow visual forecasting trick but as a universal scientific representation principle across physical, chemical, and biological dynamics is ambitious.

Method maturity: The mathematical framework of orthogonal predictive factorization is well-motivated. It provides a formal answer to the representation collapse and drift issues that have limited JEPAs in precision-critical domains.

Experimental integrity: The empirical evidence is exceptionally broad:

  • Coverage: 7 disparate domains (vision, molecular dynamics, clinical data, fluid/physical fields, robotics control, weather, cell biology).
  • Hard Benchmarks: 100-step molecular rollouts without divergence; 34.8% error reduction on Interventional Pong; over 1,000 clinical forecasting events.
  • Physical Grounding: Recovers Kepler’s third law scaling exponent (T∝a3/2  ⟹  slope −1.5T \propto a^{3/2} \implies \text{slope } -1.5) with empirical fit of -1.4991.
  • Wet-lab Validation: Biological intervention predictions confirmed in vitro (organoids, cell co-cultures) and in vivo (tumor fragments in mice).

Writing quality: Exemplary clarity. The paper avoids over-claiming, systematically contrasts against matched vanilla JEPA baselines, and open-sources reproducible code.

Verdict: strong accept — A landmark contribution that pushes self-supervised predictive representation learning into the territory of general scientific discovery.

Takeaways

  • Stop predicting entangled monolithic latent vectors for complex dynamical systems; decompose your target state into orthogonal factor pathways.
  • JEPAs can model far more than video: any system governed by time-series transitions (molecules, medical logs, planetary physics) can benefit from non-generative representation prediction.
  • Validate world models not merely on reconstruction loss, but on multi-step rollout stability, conservation law retention, and causal intervention accuracy.

论文: 2609.20800 作者: Taoyong Cui, Zhongyao Wang, Xinyue Xu, Weiyang Liu, Zhaochen Yu 分类: cs.CL, cs.AI, cs.LG

缺口

世界模型(World Models)被视为智能体实现高级认知与自主规划的精神沙盘:它能够推演未来状态、预演不同干预决策的连锁反应,从而让系统无需在现实世界中反复试错。

然而,长久以来世界模型的研究一直割裂在各自为政的垂直学科孤岛中。 视频生成模型(如 Sora、V-JEPA)死磕像素与时空 Patch;分子动力学依赖专门的几何图神经网络;气象预报仰赖球谐函数网络;而生物医学则困在基因调控网络里。

一个根本性的科学疑问始终悬而未决:是否存在一种普适通用的自监督预测学习原理,能够同时支配跨越宏观与微观的异构动力学系统?

杨立昆(Yann LeCun)提出的联合嵌入预测架构(JEPA)通过在抽象表征空间进行预测,彻底甩掉了像素级生成的算力浪费。 但在复杂的物理与科学动力学中,原生 JEPA 存在严重的**因子纠缠(Factor Entanglement)**问题:快变量与慢变量、守恒量与瞬时扰动、环境底色与干预动作全被揉进一个单一的隐空间向量里。 这种纠缠导致多步自回归推演迅速积累漂移误差,且根本无法精确预测因果干预的定向效应。

   核心挑战:割裂的跨领域动力学建模

   视觉世界         分子动力学           细胞生物学         天体运行
   (像素/图像块)    (原子与键能)         (基因表达谱)       (引力与开普勒轨道)
         \               |                    |                 /
          \              |                    |                /
           v             v                    v               v
   [ 领域专用模型 ] [ 专用物理模拟器 ]   [ 统计学轨迹分析 ] [ 经典常微分方程 ]
                          |
                          v
   根本疑问:能否用同一种预测学习范式统摄所有异构动力学系统?
                          |
                          v
   原生 JEPA 的局限:
     单一未解耦的隐表征空间 -> 快慢变量与守恒律混成一团
     -> 长程多步外推迅速漂移崩溃 -> 无法完成精准的定向因果干预
                          |
                          v
   本文解法:JEPA-Anything + 正交预测因子分解(OPF)
     将隐空间动力学解构为多条相互正交的专用演化通道

增量

一句话: JEPA-Anything 通过正交预测因子分解(OPF)将联合嵌入预测架构拓展为领域无关的通用世界模型底座,把复杂状态演化解耦为互补的正交潜因子通路,在 7 大异构科学领域全面刷新了推演精度,并精准复现了开普勒轨道指数。

核心机制

针对原生 JEPA 隐向量混杂的缺陷,**正交预测因子分解(Orthogonal Predictive Factorization, 简称 OPF)**将目标动力学流形剖解为多个互补的子空间:

  1. 正交通路分解(Orthogonal Pathway Decomposition):状态表征被显式解离为多组相互正交的潜因子空间: z=⨁k=1Kz(k),约束条件:⟨z(i),z(j)⟩≈0(∀i≠j)z = \bigoplus_{k=1}^K z^{(k)}, \quad \text{约束条件:} \langle z^{(i)}, z^{(j)} \rangle \approx 0 \quad (\forall i \neq j) 每个子空间专注于特定属性的动态:例如系统固有的几何守恒不变量、局部的快速高频振动、以及外部干预所触发的定向位移。
  2. 专属预测分支(Dedicated Predictors):各子因子沿专属的参数化通道进行演化预测,分别注入历史上下文与干预操作。
  3. 重组交叉注意力(Recombinant Co-Attention):各分支预测结果经由全局注意力层协同对齐,既确保了多变量联合演化的物理一致性,又杜绝了信息回流造成的表征塌缩。
   正交预测因子分解(OPF)内部拓扑

   当前上下文输入 (t) -----------------------------> 因子演化预测器
            |                                               |
            v                                               v
   +------------------+                            +------------------+
   |  通用主干编码器   |                            |   专属预测分支    |
   | (Transformer)    |                            |  (各行其道演化)   |
   +------------------+                            +------------------+
            |                                       |        |        |
            v                                       v        v        v
   正交潜因子: [ 守恒不变量 ]        [ 瞬态高频扰动 ]   [ 定向干预变量 ]
            \                     |                    /
             \                    |                   /
              v                   v                  v
         +-------------------------------------------------+
         |            正交几何对齐约束与协同注意力           |
         +-------------------------------------------------+
                                  |
                                  v
                    预测生成的下一时刻隐状态 z_{t+1}

这里的核喻是大型交响乐团的多轨专业录音调音台。 原生的预测模型就像在音乐厅中央放了一只劣质的全向麦克风。 当你试图预测两小节后的音乐走向时,震耳欲聋的大军鼓瞬间掩盖了双簧管细微的旋律,而且你绝不可能在不破坏鼓点节奏的前提下单独调整小提琴的音量(无法实施因果干预)。 JEPA-Anything 则为指挥家配备了专业的分轨拾音系统:小提琴、木管、铜管、打击乐各自拥有独立的指向性话筒与专属声道(正交专用通道)。 每个声道遵循自己的乐理逻辑独立推演(打击乐按节拍推演,长笛按旋律推进),最后在母带调音台上合成完美的交响乐。 由于轨道完全解耦,你可以随意修改小提琴的某个音符(定向干预),而绝对不会让鼓手的节拍发生丝毫错乱。

关键概念

  • 联合嵌入预测架构(JEPA):摒弃传统自回归或扩散模型生成底层像素/文本/坐标的沉重负担,直接在抽象表征空间预测遮蔽区域或未来时刻的状态,具有极致的计算效率与语义抽象能力。
  • 正交预测因子分解(OPF):在高维动力学演化中,通过施加代数正交惩罚,强制让表征拆解为彼此独立的动力学分量,从数学底层杜绝长期推演中的误差纠缠级联放大。
  • 干预泛化性(Intervention Generalization):世界模型不仅要能「被动看录像」,更要能在遭遇从未见过的外部干预指令时,准确预言系统将如何发生定向偏转。

框架转变

之前(领域垂直割裂与单体隐空间):
  特定领域 -> 研发专用架构 -> 单体隐向量直接外推
  -> 物理守恒量被瞬时噪声冲刷殆尽
  -> 多步自回归推演到 10~20 步即告彻底崩溃
  -> 实施局部干预时,不相关的物理属性产生灾难性联动漂移

之后(跨领域统一的正交解耦动力学):
  任意动力系统 -> 通用表征主干 -> 正交因子分轨演化预测
  -> 守恒量与高频扰动物理隔离
  -> 分子动力学稳定实现 100 步高保真滚动推演
  -> 严格复现开普勒天体轨道指数:拟合斜率达 -1.4991(理论值 -1.5)
  -> 生物干预预测通过小鼠活体与类器官湿实验证实

从「针对每个科学领域量身定制缝合怪模型」,核心转变在于:用正交因子解耦的联合嵌入自监督预测,首次实现了跨越宏观物理、微观分子与生物医学的通用世界模型。

专家评审

选题眼光: 极具雄心与前瞻性。 跳出单纯生成高清视频的窄门,将「世界模型」还原为其本质——对宇宙动力学演化法则的抽象表征与推演,展现了顶级的研究格局。

方法成熟度: 理论逻辑闭环且优雅。 针对 JEPA 在连续物理系统中因特征纠缠导致多步漂移的顽疾,给出了正交分解这一极具几何美感且可实施的解法,大幅提升了推演稳定性。

实验诚意: 广度与深度令人赞叹:

  • 跨度极广:横跨计算机视觉、分子动力学、临床多维时序、流体物理场、机器人控制、气象模拟与细胞生物学 7 大领域。
  • 硬核指标:在 10 项基准动力学任务中全面胜出;在干预版 Pong 游戏中将单次干预误差降低 34.8%;实现 4 种体系下最低的 100 步分子推演误差。
  • 物理定律验证:从无监督轨道数据中自发涌现天体运行规律,拟合出的开普勒第三定律指数为 -1.4991(理论值严格为 -1.5)。
  • 湿实验背书:所预测的生物靶点干预在类器官、肿瘤切片及活体小鼠体内实验中得到了明确证实。

写作功力: 架构清晰、实验交代详实客观,与原生 JEPA 的消融对照极其严密,并开源了全部代码。

判决: 强接收 (strong accept) — 自监督表征学习与科学 AI(AI for Science)交叉领域的里程碑之作,有力论证了跨系统通用世界模型的可能性。

要点总结

  • 当你在构建复杂物理或工程系统的时序预测模型时,停止直接使用单一黑盒隐向量做自回归推演;务必对状态进行正交因子解耦。
  • JEPA 绝不仅限于计算机视觉;任何具备状态转移特性的系统(分子模拟、医疗体检时序、工业机组运行)均可采用非生成式的抽象嵌入预测范式。
  • 评测世界模型优劣的硬标准不是看单步损失有多低,而是看长程多步推演是否发散、守恒物理量是否保持,以及因果干预响应是否真实准确。