Blog
1867 posts
- Agentic Societies Need a Social Harness 智能体社会需要一个社会性缰绳
- Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback 耦合的校准与学习:无目标域奖励反馈时缓解蒸馏中的教师偏差
- Decomposition Buys Integrity, Not Yield 分解买的是完整性,不是产出
- JustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management JustFit:用即时状态管理在 24 GiB 笔记本上跑 20 万 token
- When Should LLMs Abstain? Chain-of-Self-Questioning for Selective Risk Control 大模型何时该弃答?用自我追问链做选择性风险控制
- Attention Mean Fields Predict Average Representation Dynamics 注意力均值场能预测表征的平均演化
- Implementing a White-Box Undetectable Backdoor for Random Fourier Features 给随机傅里叶特征实现白盒不可检测后门
- Reasoning with Image Generation 用图像生成来推理
- Register Tokens for Bounded-State Reasoning in Diffusion Language Models 用寄存器 token 给扩散语言模型做定长状态推理
- Where Post-Training Quantization Breaks Text Embedders 训练后量化在文本嵌入模型上会从哪里断
- AGENTQ: A Checkpoint Can Pass Every Audit and Misbehave Once Quantized AGENTQ:一个检查点可以通过所有审计,却在量化后行为不端
- Measuring the Creativity of Frontier LLMs in Automated Research 度量前沿大模型在自动化科研中的「创造力」
- Performative Stability in Nearly the Minimum Number of Deployments 以「近乎最少的部署次数」达到表演性稳定
- TF-IDF and BM25 Are Exact KL Divergences TF-IDF 与 BM25 都是精确的 KL 散度
- VeriDx: A Correct Diagnosis Can Be Reached for the Wrong Reasons VeriDx:一个正确的诊断,可能是靠错误的理由得出的
- A Ranking Approach for Measuring Calibration 一种用「排序」来度量校准的方法
- Benign Landscapes and Worst-Case Hardness Can Coexist 「良性损失地形」与「最坏情形困难性」可以共存
- Embodied-BenchForge: Verify Each Artifact, or Defects Propagate Embodied-BenchForge:逐个产物做验证,否则缺陷会一路传播
- SAS: Stop Distilling Attention, Optimise the Ranking Instead SAS:别再蒸馏注意力分布,去优化「上下文排序」
- Type Diversity Explains Why Structural Generalisation Looks Harder 「类型多样性」解释了「结构泛化为何看起来更难」
- Artificial Id: When Drive Emerges From Persistence Alone 人工本我:当「驱动力」仅从「持续性」中涌现
- MoEs Overfit More to Repeated Data, and Sparsity Is What Drives It MoE 对「重复数据」过拟合更严重,而驱动力是「稀疏度」
- Distance Generalization: Changing the Gaps Without Changing the Length 距离泛化:在不改变长度的前提下改变「间隔」
- From Protocols to Evidence: Two Ways to Bound What an AI Claim May Assert 从协议到证据:给 AI 主张设界的两种方式
- GPU-CFR: Compiling the Game So the CPU Beats the GPU, Then the GPU Beats Everything GPU-CFR:先把博弈编译好,让 CPU 先赢过 GPU,再让 GPU 赢过所有
- The Gap-Entropy Conjecture: Entropy Is the Price of Not Knowing Which Arms Are Hard 缺口—熵猜想:熵,是「不知道哪些臂难」的代价
- BrainTaskonomy: Organise Pretraining and Transfer by Measured Learning Relations BrainTaskonomy:用「测出来的学习关系」来组织预训练与迁移
- IBIB: Enterprises Deploy Systems, Not Checkpoints IBIB:企业部署的是「系统」,而不是「检查点」
- IdeaAMBIG: Finding the Gap in a Method Spec Is the Hard Part IdeaAMBIG:难的是「找出方法规格里的缺口」
- Semigroup-JEPA: Better Features, Not Better Dynamics Semigroup-JEPA:赢在「更好的特征」,而不是「更好的动力学」
- State Credit, Not Gradient Decay, Is What Stops Recurrent Models Extrapolating 让循环模型无法长度外推的,是「状态信用」,而不是梯度衰减
- Nearly Tight Rademacher Bounds for Input-Dependent Sparsity 面向「输入相关稀疏性」的近乎紧致的 Rademacher 界
- Procedural Graphs: What-to-Do Knowledge as Triplets, Not a Growing History 过程图:把「该做什么」的知识做成三元组,而不是一段不断变长的历史
- ReCite: Semantic Similarity Finds Real Papers That Do Not Support Your Claim ReCite:语义相似会找到真实、但不支持你论断的论文
- SyncWorld: Actions Are Not a Universal Language in Pixel Space SyncWorld:在像素空间里,动作不是一门通用语言
- CUA-Universe: Why Computer-Use Agents Need Both the Mouse and the Terminal CUA-Universe:电脑智能体不能只靠点鼠标
- Does Your Agent's Memory Survive a Model Upgrade? 换掉大模型后,你的智能体其实已经失忆了
- From Interpretability Methods to Interpretable Models 可解释性研究走偏了十年:别比量角器,去量大脑
- Molecular Déjà Vu: Why Frontier LLMs Memorize Digits Instead of Learning Chemistry 分子既视感:大模型背出物理常数,不是学会了化学
- Necessary or Sufficient? Auditing LLM Explanations with Causal Interventions 大模型的理由是编的:因果审计揭穿可解释性幻象
- When Rewording the Prompt Flips the Robot's Grade 换句话问,机器人的满分就变零蛋
- Think-Verify-Revise: Merging VLMs with Dynamic Logic Tensor Networks 想、查、改:当大模型遇上动态逻辑张量网络
- UniMate: A Single Diffusion Model to Animate Any Skeleton UniMate:一个扩散大模型驱动万种骨骼动画
- When LLM Decompilers Recompile More and Preserve Less 反编译器变聪明了,但漏洞被它随手修没了
- WorldSculpt: Carving Interactive 3D Worlds from Video Streams WorldSculpt:把连体婴儿般的 3D 场景拆成可互动的真实世界
- Design Docs Are All You Need: A Repository Whose Main Branch Has Almost No Code Design Docs Are All You Need:一个主分支几乎没有代码的仓库
- KOPA-Bench: Synthesising Tool-Calling Data From Live API Execution KOPA-Bench:用「真实 API 执行」合成工具调用数据
- Task Graphs and Procedural Memory for Long-Horizon VLA Manipulation 用任务图与过程记忆支撑长时域 VLA 操作
- WearableQA: 4,084 Questions From 200 Real Users' Longitudinal Records WearableQA:来自 200 位真实用户长期记录的 4,084 道题
- What Matters, When? Distractors Break the Target Choice, Not the Skill 何时、什么才要紧?干扰物破坏的是「目标选择」,而不是操作技能
- Emergent Cheating and Whistleblowing in a Swarm of 100 Research Agents 100 个科研智能体群体中自发出现的作弊与举报
- Clean Engineering, Unstable Measurement: A Model Name Is Not a Frozen Instrument 干净的工程、不稳定的测量:模型名不是一个被冻结的仪器
- Auxiliary Views: Why Data Diversity Helps, and What Repetition Is Actually For 辅助视角:数据多样性为何有用,以及重复到底在做什么
- Legibility Is Not Interpretability: CoT Step Importance Is Only Partly in the Text 可读不等于可解释:思维链中「哪一步重要」只有一部分写在文本里
- Robust PAC Learning of Concurrent Stochastic Games: Certifying That No Equilibrium Exists 并发随机博弈的鲁棒 PAC 学习:为「均衡不存在」出具证明
- Cliff: Learn From the First Mistake, Ignore Everything After It Cliff:从「第一个错误」学习,忽略它之后的一切
- frb100-40: Settling a 20-Year Benchmark, and Reporting the Null Result Honestly frb100-40:终结一个 20 年的基准,并诚实地报告那个零结果
- SolarWM: Making Video World Models Reproducible Across Backbones SolarWM:让视频世界模型能跨主干复现
- Linguistic Illegibility: Why a Sandbox Cannot Trust What the Model Says About Itself 语言不可读性:为什么沙箱不能相信模型对自己的陈述
- User Feedback Is a Signal LLM Judges Cannot See 用户反馈是一种「大模型评审看不见」的信号
- Beyond Scores: How an LLM Judge Actually Decides Beyond Scores:大模型评审究竟是怎么打分的
- CordisBench: Can a Model Reason About the Harness It Is Rewriting? CordisBench:模型能推理它正在改写的脚手架吗?
- Mechanism Design for Alignment and Control: Incentivising Honesty and Obedience 对齐与控制中的机制设计:同时激励「诚实」与「服从」
- The Rise of Verbal Reinforcement Learning: Language as the Feedback Channel 言语强化学习的兴起:把自然语言当作反馈通道
- Quantization Damage Is Diffuse, So Spend the Next Bit Globally 量化损伤是弥散的,所以下一个比特应该全局花掉
- Aspire: Can Models Self-Evolve From a Vague Goal? Aspire:模型能从一句模糊的目标里自我演化吗?
- Auditing Anonymous AI Models: Identifying a Model That Does Not Want to Be Identified 审计匿名 AI 模型:识别一个不愿被识别的模型
- Constant Individual Regret in General Games 一般博弈中的常数个体遗憾
- LLM Post-Training as Brownfield Maintenance: Dataware, Not Recipes 把大模型后训练当作「棕地维护」:维护的是数据件,而不是配方
- Train Classical, Deploy Quantum: A Converged Loss Is Not Generalization 经典训练、量子部署:损失收敛不等于泛化
- A Formal Limit on Learning Meaning From Text Alone 从纯文本中学习「意义」的形式化极限
- Survey of Optimizers: There Is No Context-Free Replacement for AdamW 优化器综述:不存在能脱离语境替代 AdamW 的东西
- Learning Between the Peaks: How Anisotropy Reshapes Kernel Ridge Regression 在峰与峰之间学习:各向异性如何重塑核岭回归
- Agent Plugins Are Maintained Software: 8.8x Growth and a New Kind of Coupling 智能体插件是被维护的软件:8.8 倍增长,以及一种新型耦合
- When Robots Mishear Us: ASR Errors Are an Embodied Safety Problem 当机器人听错我们:语音识别错误是一个具身安全问题
- How Language Models Organize Moral Knowledge 语言模型如何组织道德知识
- Persona-Execution Separation: Let the Agent's Personality Drift, Freeze Its Actions 人格—执行分离:让智能体的「人格」自由漂移,把「行为」冻结下来
- Retrieval Heads Meet Vision: 1.7% of Attention Heads Do the Grounding 当检索头遇上视觉:1.7% 的注意力头承担了全部「定位」工作
- SWE-Prime: Training on 10% of the Trajectories Beats Training on All of Them SWE-Prime:只用 10% 的轨迹训练,效果反而更好
- Ellipsoid Fitting: A Sharp Threshold Set Only by the Fourth Moment 椭球拟合:一个只由四阶矩决定的锐利阈值
- Agentic Autoresearch: 81 Unattended Experiments and a Provable Structure 智能体自动科研:81 次无人值守实验,换来一个可证明的结构
- A Neutrino Model Knew More Than It Was Using 一个中微子模型知道的东西,比它在用的更多
- How Much Rank Does LoRA Need? A Task-Dependent Answer LoRA 到底需要多大的秩?一个依赖任务本身的答案
- Prefix Sliding: Most of a Reasoning Trace Stops Mattering Prefix Sliding:推理轨迹里的大部分内容会逐渐失去作用
- Trace Integrity: A Correct Answer Can Come from an Invalid Computation 轨迹完整性:正确答案也可能是「算错」得来的
- A Geometric Theory of Robust Fairness Audits: The Auditor Can Be the Fragile Part 稳健公平性审计的几何理论:脆弱的那一环可能正是审计本身
- BrowserForge: 203,238 Trajectories, One Per Website BrowserForge:203,238 条轨迹,每条来自一个不同网站
- Do Robotic World Models Really Follow Actions? Expert Demos Hide the Answer 机器人世界模型真的会跟随动作吗?专家演示掩盖了答案
- Reading Is Not Using: Long Contexts Break the Link Between Retrieval and Judgment 读到不等于用上:长上下文切断了检索与判断之间的联系
- What FID Hides: A Scalar Cannot Rank, Test and Diagnose at Once What FID Hides:一个标量无法同时承担排名、检验与诊断
- Correcting a Learned Invariant: World Models Know Physics They Still Violate 修正一个学到的守恒量:世界模型知道物理规律,却在想象时违反它
- How AI Assistance Affects Skill Development: Short-Term Gains, Longer-Term Losses AI 辅助如何影响技能成长:短期收益,长期代价
- Investigating Relational Reasoning in VLMs: How Much Is Seeing and How Much Is Language? 考察 VLM 的关系推理:多少来自「看见」,多少来自「语言」
- SWE Refactor Bench: Coding Agents Cannot Do a Whole-Repository Migration SWE Refactor Bench:编程智能体做不了整仓库迁移
- The Interaction Tax: Why Letting Agents Read Each Other Destroys Their Diversity 交互税:让智能体互读对方答案,为何会抹平它们的多样性
- Asymmetric Capacity Allocation: Your Critic Does Not Need to Be Smart 非对称容量分配:你的批评者不需要那么聪明
- CLEAR: Gating Safety On and Off Instead of Rewriting the Whole Model CLEAR:把安全能力「按需开关」,而不是重写整个模型
- Difficulty-Calibrated Paths: Let the Model Choose Its Own Interpolation Schedule 难度校准路径:让模型自己决定插值时刻表
- Move by Move: How LLMs Actually Conduct Therapy, and How to Steer It Move by Move:大模型到底怎么做心理咨询,以及如何引导它
- Primal Acceleration of Newton's Method: Cubic Rate from One Linear Solve 牛顿法的原空间加速:单次线性求解换来三次收敛率
- 4DAnyone: Turning One Casual Phone Video into a 4D Human 4DAnyone:一段随手拍的手机视频,如何变成 4D 数字人
- AI4AI-Bench: Can an Agent Actually Rewrite a Training Algorithm? AI4AI-Bench:智能体真的能改写训练算法吗?
- Catching the Rug: Spotting Solana Memecoin Fraud in the First Five Minutes Catching the Rug:用前五分钟的交易数据抓出 Solana 貔貅盘
- Inducing Task Models from Computer-Use Traces: Turning Raw Clicks into Auditable Workflows 从电脑操作轨迹中反推任务模型:把原始点击变成可审计的工作流
- Phantom Gains: Your Self-Improvement Numbers Are Probably Batching Artifacts Phantom Gains:你的自我改进指标,可能只是批推理的假象
- Chain-of-Experience: The Notebook Stays Open 经验链:让笔记本一直摊开着
- The Judge That Knows When to Google — And When to Shut Up 让 LLM 裁判在没把握时先去搜一下,再给你一张误判率证书
- Self-Improving Agents Are Mostly Measuring Noise 自我改进智能体:你测到的多半是噪声
- StagedWorkspace: The Agent Is Editing a File It Never Actually Read StagedWorkspace:智能体改的那个文件,它其实没读过
- Tokenization Is a Change of Coordinates, Not a Bigger Dictionary 分词是换坐标系,不是造更大的词典
- Reward Design as a Procedure: From Objectives to Causal Coverage to a Polytope of Weights 把奖励设计变成一道流程:从目标到因果覆盖再到权重多面体
- AVA-Encoder: Turning Films into Knowledge Graphs an Agent Can Actually Edit AVA-Encoder:把电影压成智能体能改的知识图谱
- CAM's Family Tree: A Method-Centered Map of Visual Explanation from CNNs to Foundation Models CAM 的家谱:从 CNN 到基础模型的视觉解释方法地图
- DARTree: Teaching a Chain-Trained Corrector to Think in Trees DARTree:让只会走直线的纠正头学会长成树
- DreamFly: Teaching a Drone to Remember, Rehearse, and Know When to Stop DreamFly:让无人机学会记忆、预演,以及知道何时停下
- Uncertainty as an Input, Not an Afterthought: LLM Sentiment Meets Small-Cap Portfolio Construction 把不确定性当作输入,而不是事后补丁:LLM 情绪信号与小盘股组合构建
- LittleLearner: What Happens When You Cap a Language Model at Fifth Grade LittleLearner:把语言模型的知识天花板锁在小学五年级
- Softmax Attention Was Always a Born-Rule Object (If You Live on the Simplex) 当数据活在概率单形上,Softmax 注意力本来就是玻恩规则的产物
- AdvFD: When Your Fréchet Loss Gets Gamed, Make the Metric Fight Back AdvFD:当 Fréchet 损失被刷分,就让度量自己反击
- Building the Jig Instead of Training the Hands: Test-Time Capability Transfer via Harnesses 不训练学生,只造夹具:用 Harness 在推理期完成能力迁移
- CausalSplat: Teaching Gaussian Splats to Answer 'Why', Not Just 'Where CausalSplat:让高斯泼溅从「指哪打哪」走向「听懂话里的意思」
- Conditional Independence Tests as the Weak Link in Causal Discovery 条件独立性检验:因果发现流水线上最脆弱的一环
- The Agent Did the Job — It Just Took the Scenic Route 任务完成了,但它绕了一大圈:技能型智能体的隐蔽成本攻击
- Feeling the Curvature Blindfolded: Curvature-Aware Zeroth-Order Test-Time Adaptation 蒙着眼摸出曲率:面向端侧测试时自适应的曲率感知零阶优化
- Diagram-MMU: Can MLLMs Actually Write the TikZ Behind a Figure? Diagram-MMU:多模态模型真能写出图背后的 TikZ 吗?
- Satellite Embeddings as Sub-Grid Descriptors: Teaching a Downscaler What the Ground Looks Like 把卫星嵌入当作次网格描述子:让降尺度模型知道地面长什么样
- Six Years of TrustNLP: Using a Workshop's Paper Stream as a Sensor for Where Trust Research Goes TrustNLP 六年志:把一个 workshop 的论文流当作可信研究的传感器
- Turning One Bit of 'Stop!' Into a Dense Cost Signal 把一声「停!」摊成一条稠密的代价信号
- StateFlow: Giving Generative Previsualization a Working Memory StateFlow:给生成式预演装上一块「工作台」
- Structural Silence: The Four Bottlenecks Below the Model 结构性沉默:模型之下的四重瓶颈
- VAKRA: When the Hard Part Isn't Calling the Tool VAKRA:真正难的不是调用工具
- XYZFlow: When More Context Beats Fewer Steps XYZFlow:与其减少步数,不如增加上下文
- Beyond a Bag of Features: When SAE Latent Sets Stop Behaving Like Sets 不只是特征袋:稀疏自编码器在集合层面的不稳定性
- DistMoE: Stitching Private Experts Into One MLLM Without Sharing Any Data DistMoE:不共享数据,把各家私有专家拼成一个多模态大模型
- Putting Physics Inside the Latent State: ELWM + Neural Time Fields 把物理放进隐空间:能量结构化世界模型与神经时间场
- Numbers as Words: FinATOM Turns Forecasting and Portfolio Allocation into Token Generation 把数字当词说出来:FinATOM 将收益预测与资产配置变成 token 生成
- Verifying That a Predictor's Probabilities Don't Contradict Each Other 如何验证一个预测器的概率不自相矛盾
- Building the Mine, Not the Miner: How AI Helped Tighten the Grothendieck Constant 人类挖矿脉,AI来钻探:一次关于 Grothendieck 常数的人机协作实录
- Code-Switching for Machines: Splicing Object Crops Into Sentences to Fix Referential Ambiguity 让机器「语码转换」:把物体图块塞进句子里,解决指代歧义
- Teaching Itself to Look: How CVPD Turns a Model's Own Blind Spots Into Dense Supervision 先看清,再学习:CVPD 如何把模型自己的视觉盲点变成密集监督信号
- Unlearning Without the Data: PRMU and the Corpus-Free Person Deletion Problem 没有数据也要遗忘:PRMU 与无语料的人物知识删除问题
- When RLVR, RLHF, and Agents Fight Over KV Cache: Admission Control for Mixed Rollouts 当 RLVR、RLHF 和 Agent 抢同一块 KV Cache:混合 Rollout 的准入控制
- Sterile Possession, Measured: Splitting Junk Ball-Hogging from Space-Creating Patience 把「无效控球」量化:区分死水式传导与创造空间的耐心
- Surgical WAM: Can Action-Free Endoscopic Video Teach a Robot to Operate? Surgical WAM:无动作标注的内窥镜视频,能教会手术机器人动手吗?
- Learning After Deployment: Teaching GUI Grounding Models to Reflect on Their Own Misses 部署之后还能学:让 GUI 定位模型反思自己点错的那一下
- The Illusion of Cross-Lingual Safety: Why Refusal Doesn't Travel to Low-Resource Languages 跨语言安全的幻觉:为什么拒答能力传不到低资源语言
- Making a Deepfake Detector Point at the Crime Scene 让视频鉴伪模型指出「案发时段」
- ArchAgent v2: When an Evolutionary Agent Out-Prefetches the Champion ArchAgent v2:当进化式智能体在预取器上击败人类冠军
- Thinking Without Words: BDH-CQ Buys ARC Reasoning for $0.0007 不说话的推理:BDH-CQ 用 0.0007 美元做一道 ARC 题
- Blast Radius: Predicting How Far a Prompt Reaches Before You Pay For It Blast Radius:在付钱之前,先算清一个 prompt 能炸多远
- COVER: Giving Video Temporal Grounders an Error Bar They Can Prove COVER:给视频时序定位装上可证明的误差条
- Confidence Has a Shape: Consilience for Verifier-Free Test-Time Scaling 置信度是一条曲线:用 Consilience 做无验证器的测试时扩展
- Decoding-Level Taboo: Kicking Models Off Their Comfortable Path 解码层禁忌:把模型踢出它的舒适通道
- DSLE: When the Benchmark Runs at 1x and Nobody Wins DSLE:当基准环境只能实时运行,而没人能赢
- When Two Minds Share One Head: The Training Dynamics of Thinking Mode Fusion 一个模型两种脑子:思考模式融合的训练动力学
- GENCO: One Neural Solver for Three Grid Problems GENCO:一个神经求解器,三类电网问题
- Integrate What You Know, Learn Only the Remainder: LDR and Extrapolative Video World Models 能算的就别学:LDR 与会外推的视频世界模型
- MirrorWorld: Teaching Video Diffusion Models What Belongs in a Mirror MirrorWorld:让视频扩散模型学会“镜子里该有什么
- MMDiff: Diffing SAEs to Find (and Flip) the Features Multimodal Training Added MMDiff:用 SAE 对比找出多模态训练新增的特征,并直接操控它
- SABRE: Turning VLM Stress Tests from a Collection Problem into a Compilation Problem SABRE:把 VLM 压力测试从「收集问题」变成「编译问题」
- SHE: Treating the Agent Harness as Something You Train, Not Ship SHE:把智能体的“安全外壳”当成可训练对象,而不是发布件
- SkillProx: Teaching Agent Skills to Shrink, Not Just Grow SkillProx:让智能体的技能库学会瘦身,而不只是长胖
- Plot Your Numbers: Feeding Time Series to VLMs as Images Cuts Tokens and Raises Accuracy 把数字画成图:用图像喂 VLM,token 更少、精度反而更高
- CoinRAG: Slicing the KV Cache Down to Nuggets CoinRAG:把 KV 缓存切到"信息颗粒"这一层
- Teaching a Post-Trained Model to Flip a Creativity Switch 让对齐后的模型学会自己按下"创造力开关
- Diffusion LLMs as Both Victim and Weapon: Safety Neurons Travel Across Architectures 扩散语言模型既是靶子也是武器:安全神经元会跨架构迁移
- Fisher-R1: Teaching Agents That a Correct p-Value Can Still Be Wrong Fisher-R1:算对了的 p 值,依然可能是错的
- When One AI Bosses Another: Driving Creates Behavior That Neither Agent Has Alone 当一个 AI 对另一个 AI 下命令:驱动造出了单独存在时不存在的行为
- When Grokking Runs Backwards: Muon, Modular Addition, and a Gain Knob Nobody Was Watching 顿悟之后的塌陷:Muon、模加法,以及没人盯着的那个音量旋钮
- SimWAM: Using Video Generation as a Training Signal, Then Throwing It Away SimWAM:把视频生成当成训练信号,然后扔掉它
- Strategy Before Steps: Can an LLM Agent Plan Total Syntheses Like a Chemist? 先有战略再有步骤:LLM 智能体能像化学家那样设计全合成吗?
- TEPA: When Memory Needs a Delete Key, Not Just a Save Button TEPA:智能体记忆缺的不是「存」,而是「撤销」
- Grading the Graders: Can LLMs Review National Standards? 给审稿人打分:大模型能审国标吗?
- Benchmarking the Benchmarks: A Reference-Free Audit for Conversational-Agent Evals 给评测集打分:对话智能体基准的无参考审计框架
- CalibForge: Stop Filtering Tasks, Start Tuning Them Against Solvers CalibForge:别再筛任务了,把任务对着求解器"调音
- When the Data Moves, So Must the Explanation: Rethinking XAI Evaluation 数据在变,解释也得跟着变:重新审视 XAI 的评估问题
- Does FLAIR Super-Resolution Erase or Hallucinate Small White-Matter Lesions? FLAIR 超分辨率会抹掉还是凭空造出小病灶?
- Auditing the Invisible: AI Transparency and Digital Sovereignty in Nigerian Shopping Apps 看不见的算法:尼日利亚购物 App 中的 AI 透明度与数字主权
- When Free Correct Labels Cost You a Logarithm: Optimal Rates Under Monotone Adversaries 免费的正确标签为何要收一个对数:单调对手下的最优速率
- Governing AI Agents With a Prepaid Meter: Resourced Authority as Mechanism Design 用预付电表管住 AI 智能体:把授权变成算力配额的机制设计
- RP-OPSD: Distilling Only Where the Reasoning Turns RP-OPSD:只在推理转弯处蒸馏
- Making VARMA Usable Again: Fixed-Size Statistics, Guaranteed Stability, No T in the Inner Loop 让 VARMA 重新可用:定长统计量、天生稳定的参数化、内层循环里没有 T
- The Low Frequency Trap: Why Video Models Can't Keep a Ledger 低频陷阱:视频模型为什么记不住一本账
- Features With Footnotes: Auditable Heart-Failure Feature Engineering by Multi-Agent Pipeline 带脚注的特征:用多智能体流水线做可审计的心衰特征工程
- Not the First Error, but the One That Killed It: TrajDebug and Error-Lifecycle Tracing 不是第一个错,而是致命的那个错:TrajDebug 与错误生命周期追踪
- Tytan: Letting the Database Describe Itself Tytan:让数据库自己说出它是什么
- UQ-Loc: Teaching LiDAR Localisation to Say How Sure It Is UQ-Loc:让 LiDAR 场景坐标回归学会说“我有多确定
- Closing the Last Log Factor: An Optimal Agnostic PAC Learner 抹掉最后一个 log 因子:最优的 agnostic PAC 学习器
- Stop When You Know: AV-AIVAT and the Real Price of Certified Early Stopping 证据够了就停:AV-AIVAT 与「可认证提前停止」的真实代价
- Beyond Sequence Order: Syntax-Informed Positional Embeddings 超越序列顺序:句法感知的位置编码
- EnvACE: Training Agents by Making the Policy Play the Environment EnvACE:让策略自己扮演环境来训练智能体
- Learning Globally Reusable Skills for Coding Agents 为编码智能体学习全局可复用的技能
- Selective Trust: Why 'Ignore the Context' Is Not Robustness 选择性信任:把上下文全当噪音,不叫鲁棒
- Reducing Belief in Conspiracy Theories As They Unfold 在阴谋论正在成形时降低对它的相信
- RRC: Turning LLM Judges Into Usable RL Rewards by Ranking, Not Scoring RRC:让生成式奖励模型在 RL 里真正好用——靠排序,而不是打分
- The Bitter Lesson of Tool Calling: When Code Beats JSON 工具调用的苦涩教训:写代码何时胜过填 JSON
- The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images 视觉工具调用的幻觉:对「看图思考」的因果审计
- Shadow Evaluations: Can AI Agents Actually Do Open-Ended AI Research? 影子评估:AI智能体真的能做开放式AI研究吗?
- KuTIE: Why LLM-Generated Kubernetes Security Patches Fail Without Live Cluster Topology KuTIE:没有实时集群拓扑,LLM生成的K8s安全补丁为什么总失败
- Are Tabular Foundation Models Robust in the Real World? 表格基础模型在真实世界中鲁棒吗?
- Instruction-Tuned LLMs Copy Human Syntax More Than Humans Do — But Less Precisely 指令微调的大模型比人类更爱抄语法——但抄得更粗糙
- MDTransformer: Squeezing Four Photonic Lanes Into One Waveguide MDTransformer:一根波导里塞进四条光子计算通道
- Mental World Modeling: When World Models Need to Simulate Minds, Not Just Matter 心理世界建模:世界模型需要模拟的不只是物理,还有人心
- OmegaUse-OfficeVal: Benchmarking LLM Agents on Office Tasks with Real-World Economics OmegaUse-OfficeVal:用真实经济标尺衡量 LLM 智能体的办公室任务能力
- VetClaw: Turning a Static Model into a Safety-Aware Veterinary Screening Agent VetClaw:将静态模型打造为安全感知的兽医疾病筛查智能体
- When Do Learned Proposals Actually Help Constraint Solvers? A Reality Check 学习提议真的帮到约束求解器了吗?一份实证检验
- GeoLens: Teaching Multimodal Models to Use a Full Toolkit on Ultra-High-Resolution Satellite Imagery GeoLens:教多模态模型在超高分辨率卫星图上用全套工具箱
- CHARM: Building Graph Foundation Models That Actually Transfer Across Multimodal Domains CHARM:真正实现多模态图跨域迁移的基础模型
- Desktop-Delta Bench: Teaching Computer-Use Agents to Verify Their Own Work Desktop-Delta Bench:教会计算机使用智能体验证自己的操作
- Falling Behind: How AI Race Dynamics Drive Unsafe Development 落后焦虑:AI竞赛动态如何驱动不安全的开发行为
- πR²: Making Robot Policies Reactive Without Losing Their Brains πR²:让机器人策略既反应灵敏又不丢掉大脑
- MemLens: Turning LLM Memory from a Junk Drawer into a Curated Library MemLens:将大模型记忆从杂物抽屉变为策展图书馆
- PDD: Skip Steps, Not Quality — Trajectory Distillation for Fast Video Generation PDD:跳步不降质——轨迹蒸馏加速视频生成
- Relay On-Policy Distillation: Passing the Baton When Students Go Astray 接力策略蒸馏:学生迷路时教师接棒
- Pictura: Self-Play Driving Without the God's-Eye View Pictura:无需上帝视角的自动驾驶自我博弈
- DITL: Letting the Dataset Tell You How Hard Each Sample Is DITL:让数据集自己告诉你每个样本有多难
- Reinformed Dreamer: Learning Better World Models with Latent Guidance 重构梦境者:通过潜在引导学习更优世界模型
- SAM Meets Muon: Why Spectral Geometry Unlocks Better Flat Minima SAM 遇上 Muon:谱范数几何如何通向更好的平坦极小值
- Spend Experts Where You Are Unsure: Adaptive MoE-LoRA Routing Based on Confidence 把专家花在刀刃上:基于置信度的自适应MoE-LoRA路由
- UniMem: Teaching LLM Agents When to Remember and When to Internalize UniMem:教 LLM 智能体何时记忆、何时内化
- Wonder: Building Playable Video Worlds with Camera Control and Long-Term Memory Wonder:构建支持相机控制与长期记忆的可交互视频世界
- ClinFusion: A Vision-Centric Multimodal LLM for Holistic Medical Understanding ClinFusion:面向全面医学理解的视觉中心多模态大语言模型
- DataOrchestra: Teaching LLMs to Treat Each Data Example Like a Solo Performance DataOrchestra:让每条训练数据都享受专属指挥
- Rethinking Classifier-Free Guidance for Robust Diffusion Distillation 重思分类器自由引导以实现鲁棒的扩散蒸馏
- A Tunable Quantum Ansatz: Balancing Trainability and Computational Power 可调节的量子电路:平衡可训练性与计算能力
- The Physics of Multi-Turn Planning: A Controlled Lab for Studying How LLMs Learn to Plan 多轮规划的物理学:一个可控实验室,研究大模型如何学会规划
- CARA: Teaching a Self-Driving Car to Explain Its Fears Using Stories CARA:用事故叙事教会自动驾驶车辆解释它的恐惧
- CausalForge: When AI Agents Do Causal Inference Research, Lean Proves It's Real CausalForge:当 AI 智能体做因果推断研究,用 Lean 证明它是真的
- Explainable RL for Air Traffic Control: A Preliminary Saliency Map Approach 可解释强化学习辅助空中交通管制:基于显著性图的初步探索
- kappa-LoRA: Smart Savings by Tuning Only the 'Crucial' LoRA Matrices κ-LoRA: 通过只调优'关键'LoRA矩阵实现智能节省
- Teaching Transformers to Design Quantum Chemistry Circuits 用 Transformer 学习设计量子化学电路
- MineValiCoder: When Tests Can't Be Trusted, Make Them Prove Themselves MineValiCoder:当测试不可信时,让测试先证明自己
- Robot-Factored World Models: Teaching AI to See Physics by Rendering the Robot Out 机器人因子化世界模型:通过渲染机器人教会AI理解物理
- Skill Self-Play: Co-Evolving Skills Push LLM Self-Improvement Forward 技能自博弈:用协同进化技能推动大模型能力边界
- SM4RT: Treating Motion as Physics, Not Pixels SM4RT:把运动当物理规律,而非像素飘移
- Twins: Unified Visual Tokens with Focal Loss for Balanced Multimodal Learning Twins:用焦点损失平衡统一视觉表征的多模态学习
- Barzilai-Borwein's Superlinear Dream Dies on an Open Set Barzilai-Borwein 方法的超线性收敛梦碎于开集之上
- Sycophancy Is Just the Symptom: How LLMs Update Moral Judgments Along Three Social Dimensions 谄媚只是表象:大语言模型如何沿三个社会维度更新道德判断
- EnsembleEGNN: Never Judge a Molecule by a Single Pose EnsembleEGNN:永远不要凭一张构象就给分子下定论
- MedGame: Turning Medical Cases into Interactive Story Games with LLMs MedGame:用大语言模型将医疗案例转化为互动叙事游戏
- Diagnosing Robot Language Understanding: Fixing Compositional Generalization with Bias-Aware Data Collection 诊断机器人语言理解:通过偏差感知数据收集修复组合泛化
- Seeing What Glare Sees: Gradient Saliency as a First-Class Output of Differentiable Rendering 看眩光所见:梯度显著性作为可微渲染的一流输出
- The Surprisal Tautology: Why a Foundational Theory in Psycholinguistics Makes No Testable Predictions 惊异悖论:为何心理语言学的一个基础理论无法做出可检验的预测
- Printing Quality Control Gets a Synthetic Data Lifeline 印刷质量控制迎来合成数据新解法
- When High Malaria Burden ≠ Weird Malaria: Consensus Anomaly Detection Uncovers Hidden Patterns in Ghana 高负担≠异常:加纳疟疾传播的共识异常检测揭示隐藏规律
- VCSD: Self-Distillation Without Any Tricks — Just Erase and Compare VCSD:不需要任何额外信号的自蒸馏——擦除即可对比
- VLM-IE3D: Teaching Vision-Language Models to See in 3D with Dual Geometry Tokens VLM-IE3D: 用双重几何Token教会视觉语言模型看懂三维世界
- Beyond Sufficiency: How TimePNS Reveals the Necessary Truths in Time Series 超越充分性:TimePNS 如何揭示时间序列中的必要真相
- Expanding Flow Maps: Learning How Big Your Output Should Be 扩展流映射:让模型学会输出该多大
- GraphVid: Controlling Video Generation with Scene Graphs, Not Scribbles GraphVid: 用场景图而非涂鸦来控制视频生成
- Progressive Seed Pruning: Spend Your Inference Budget Where It Matters 渐进式种子剪枝:把推理算力花在刀刃上
- OpenForgeRL: Training Harness-native Agents End-to-End with RL at Scale OpenForgeRL:在任意环境中端到端训练 Harness 原生 Agent
- SANA-Video 2.0: Bridging the Quality-Efficiency Gap in Video Generation with Hybrid Attention SANA-Video 2.0:用混合注意力弥合视频生成的质量-效率鸿沟
- Separating Camera and Object Motion in Self-Supervised Video Understanding 视频自监督学习中分离相机与物体运动
- Streaming World States for Consistent Multi-Agent Video Generation 流式世界状态实现一致的多智能体视频生成
- UniD: Teaching a Single Video Model Eight Scene Skills from Separate Textbooks UniD:用八本互不相干的教材教会一个视频模型八种场景理解能力
- Look Less, Think Faster: Coordinating Vision Tokens and LLM Compute in Multimodal Models 少看快想:多模态大模型中视觉Token与计算资源的联合调度
- From Expensive Giants to Accessible Minis: A VR+RL Control Stack for Small Humanoid Robots 从昂贵巨物到可及微缩:面向小型仿人机器人的VR+RL控制栈
- RECAP: When Your Model's Self-Report Can't Be Trusted, Train It to Be Independently Verified RECAP:当模型的自我报告不可信时,训练它接受独立验证
- Taming Variance in Domain Adaptation with Smart Pairing 用智能配对驯服域适应中的方差
- Causal Discovery Meets Irregular Time Series: Extending PCMCI+ Beyond Fixed Lags 因果发现遇上不规则时间序列:突破固定滞后的 PCMCI+ 扩展
- FlashRT: Letting AI Agents Optimize Your Multimodal Deployment Pipeline FlashRT:用 AI 智能体优化多模态应用部署流水线
- FlowMimic: Bootstrapping Video Editing Data from Image Edits via Pixel-Pair Warped Flow FlowMimic:通过像素对扭曲流场从图像编辑数据引导视频编辑
- GigaPath-Flash & GigaTIME-Flash: Making Pathology AI Actually Deployable GigaPath-Flash 与 GigaTIME-Flash:让病理 AI 真正可部署
- HOMIE: Unifying Human-Object Video Personalization with Multimodal Intelligence 人与物,和谐一体:HOMIE统一人-物视频个性化新范式
- It's Not What You Say, It's How You Say It: How Linguistic Framing Sneakily Persuades LLMs 你说什么不重要,你怎么说才重要:语言表达如何悄悄说服大语言模型
- Soft Prefixes Break LLM Logic — And What That Reveals About Model Stability 软前缀如何动摇大模型的逻辑判断
- Patch Policy: Bridging the Gap Between Lightweight Robots and Dense Vision Patch Policy:用轻量架构打通机器人与密集视觉之间的断层
- PPL-Factory: When Picking Training Data, Know Your Task and Count Your Pennies PPL-Factory: 挑选训练数据时,认清任务、算好账
- PIXAR-DG: Simple Domain Generalization for Pixel-Level Tampering Detection Across Modern VLMs PIXAR-DG:面向现代 VLM 的简单域泛化像素级图像篡改检测
- SWE-Pruner Pro: Let the Coding Agent Prune Its Own Context SWE-Pruner Pro:让编码智能体自己修剪上下文
- TPIPS: Teaching Vision Models to See Similarity Through Your Lens TPIPS:让视觉模型按你的视角理解图像相似度
- Three-Body Scattering: One-Step Generation via Energy-Based Motion 三体散射:基于能量驱动的一步生成模型
- Bridging RAG and Causal Inference: A Framework for Policy Learning 连接 RAG 与因果推断:基于检索增强的政策学习框架
- Thermodynamic Computing: Let Physics Do the Sampling 热力学计算:让物理本身完成采样
- ActiveVision: Can Multimodal LLMs Actually Observe, or Just Glance? ActiveVision:多模态大模型真的会「看」,还是只扫一眼?
- Cluster-Aware Matching via Laplacian Optimal Transport: Seeing the Forest and the Trees 通过拉普拉斯最优传输实现聚类感知匹配:既见树木,又见森林
- FVAttn: Fixing the Straggler Problem in Video Diffusion's Sparse Attention FVAttn:修复视频扩散稀疏注意力中的掉队者问题
- MotionForesight: Turning Video Priors into 3D Motion Predictors with a Lightweight Adapter MotionForesight:用轻量适配器将视频先验转化为三维运动预测器
- PagedWeight: Smart Memory Trading for MoE Models PagedWeight:为MoE大模型做智能内存交易
- PEARL: When Physics Meets RL for High-Dimensional Control PEARL:当物理遇上强化学习,攻克高维控制难题
- Searching Videos as Trees: How Self-Correcting Agents Crack Long Video QA 将视频搜索变成树导航:自校正智能体如何攻克长视频问答
- Teaching AI to Use Its Eyes: ToolSciVer's Tool-Augmented Approach to Scientific Claim Verification 教 AI 学会用眼睛:ToolSciVer 的工具增强科学声明验证方法
- Multi-Agent Systems Are an Information Bottleneck Problem: When Bounded Communication Helps or Hurts 多智能体系统是信息瓶颈问题:有限通信何时助益、何时拖累
- ARMOR++: Multi-Agent Orchestration Exposes Transferability Gaps in Deepfake Detectors ARMOR++:多智能体编排揭示深伪检测器的迁移性漏洞
- Decoding Market Emotion from Blockchain Activity: A Data-Driven Sentiment Classifier 从区块链活动解码市场情绪:数据驱动的情感分类器
- MCF-Net: When Heart Motion Meets Foundation Models for Infarction Localization MCF-Net:心脏运动与基础模型的多视角融合定位心肌梗死
- Decoupling Memory Updates from Inference: Real-Time Neural Novel View Synthesis 解耦记忆更新与推理:实时神经新视角合成
- TikStance: Teaching AI to Read Political Stance in TikTok's Multimodal Conversation Trees TikStance:让 AI 读懂 TikTok 多模态对话树中的政治立场
- Beyond Success Rate: How to Evaluate Security AI Agents by What They Actually Cost 超越成功率:如何按实际成本评估安全AI智能体
- Beyond the Leaderboard: Trustworthy Multimodal AI for Healthcare, Not Just Accurate Answers 超越排行榜:医疗多模态AI的可信之路,而非仅仅是答案准确
- Static Relevance Is a Lie: Bridge Documents in Agentic Search 静态相关性靠不住:智能体检索中的桥接文档
- Grounding Image Geo-localization: HoloGeo Mitigates Landmark Bias with Evidence-Driven Reasoning 立足图像地理定位:HoloGeo 以证据驱动推理缓解地标偏见
- MeanFlowNFT: Making Reward-Aligned Generation Fast with Average-Velocity Flows MeanFlowNFT:用平均速度流实现快速奖励对齐生成
- Mutable Low-Rank Sketches: Retrain-Free Recommendation That Actually Works 可变低秩草图:真正可行的免重训推荐系统
- SceneBind: Teaching Multimodal Models Where Things Are, Not Just What They Are SceneBind:让多模态模型不仅知道"是什么",还知道"在哪里
- SciDiagramEdit: Teaching AI to Edit Scientific Figures by Watching How Authors Revise SciDiagramEdit: 通过观察作者如何修改论文来教AI编辑科学图表
- SearchOS: Explicit State Management for Robust Multi-Agent Information-Seeking SearchOS:面向鲁棒多智能体信息检索的显式状态管理
- teLLMe Why: Turning Dashcam Data into Causal Questions with LLMs teLLMe Why:用大语言模型把行车记录仪数据变成因果问题
- AutoSynthesis: Can AI Agents Replace the Entire Meta-Analysis Pipeline? AutoSynthesis:AI智能体能否接管整条元分析流水线?
- Hierarchical Denoising: Teaching Video Models to Reason in Steps 层次化去噪:让视频模型学会分步推理
- LLMs Know the Parts But Fail at the Whole: Statistical Self-Consistency as a Blind Spot 大模型知道局部却算不对全局:统计自一致性作为盲点
- Poisoning the Well: How Public Forums Can Backdoor Language Models 毒源污染:公共论坛如何为语言模型埋入后门
- RoboTTT: Unlocking Long-Term Robot Memory with Test-Time Training RoboTTT:通过测试时训练解锁机器人长期记忆
- From Pixels to States: Why Video Generation Alone Won't Build Game Engines 从像素到状态:为什么单靠视频生成造不出游戏引擎
- Hindcast: Sealing the Evidence Room to Evaluate LLM Forecasters Hindcast:封存证据室来评估LLM预测能力
- MOJO: Teaching Neural Decoders to Learn Without Labels MOJO:让神经解码器无需标签也能学习
- Lighthouse RL: Strategic Resets for Sample-Efficient Circuit Design Lighthouse RL: 策略性重置实现样本高效的电路设计
- Reinforcement Learning Finds Stability Solutions Beyond Physics: Using Lyapunov Exponents as Universal Rewards 用李雅普诺夫指数做强化学习奖励:超越卡皮查摆的稳定控制
- M4World: Multi-View Multimodal Driving World Model with Object Manipulation and Minute-Long Streaming M4World:支持物体操控与分钟级长流生成的多视角多模态驾驶世界模型
- Why Deep Transformers Don't Collapse: Unifying Skip Connections and Norms Through the Lens of Rank 深度Transformer为何不崩溃:从秩的视角统一跳跃连接与归一化
- VideoRAE: Why Learn Video Latents from Scratch When Foundation Models Already Know So Much? VideoRAE:视频基础模型已经这么强了,为什么还要从零学潜空间?
- Parallel Transcripts: Running a Frozen 26B Diffusion LM for Speech Recognition 并行转录:用冻结的260亿参数扩散语言模型做语音识别
- Do AI Agents Know When a Task Is Simple? — Letting LLM Agents Right-Size Their Effort AI 智能体知道任务有多简单吗?——让 LLM 智能体量力而行
- From Templates to Formal Languages: A Scalable Pipeline for Generating Precise Analytic Geometry Problems 从模板到形式语言:一个可扩展的解析几何精确问题生成流水线
- PalmClaw: From Tapping Screens to Calling Tools — A Native On-Device Agent for Mobile Phones PalmClaw:从点击屏幕到调用工具——手机原生智能体框架
- Resist and Update: Certifying LLM Report Invariance Through Counterfactual Activation Clamping 抵抗与更新:通过反事实激活钳制验证LLM报告的因果不变性
- TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale TerraZero:无需示范的大规模程序化驾驶自博弈训练
- The Illusion of Robustness: Why Aggregate Accuracy Lies About LLM Reliability 稳健性幻觉:为什么聚合准确性掩盖了大模型的真实可靠性
- The Seriality Gap: Why Video Diffusion Models Can't Reason About Chain Reactions 序列性缺口:为什么视频扩散模型无法推理链式因果
- ViCo3D: Teaching Cars to Collaborate by Borrowing Vision Foundation Models ViCo3D:借力视觉基础模型,赋能车路协同三维感知
- Watermark Forensics: An Information-Theoretic Ladder Beyond Detection 水印取证:超越检测的信息论阶梯
- AdvancedMathBench: Testing the Frontier of LLM Mathematical Proof AdvancedMathBench:前沿大语言模型数学证明能力的试金石
- Teaching Video AI to Show Its Work: Evidence-Backed Video QA 让视频AI学会"交作业":证据支撑的视频问答
- HASTE: Rapid Building Damage Assessment with Post-Disaster Imagery Alone HASTE:仅用灾后影像快速评估建筑损毁情况
- Q-DIBA: When Quantum Backdoors Learn to Disguise Themselves Q-DIBA:当量子后门学会伪装
- Decoding the Unfair Judge: How LLM Biases Live in Hidden Geometric Space 解码不公正的法官:LLM偏见如何存在于隐藏几何空间中
- Why Transformers Learn to Reason in Low Dimensions: The Invariant Manifold of Inductive Circuits Transformer 推理学习的低维秘密:归纳电路的不变流形
- The First Map of Metacognition in Large Language Models 大语言模型元认知的首张全景图
- MM-ToolSandBox: When Smart Models Can't Read — A Benchmark That Exposes the Visual Perception Bottleneck in Tool-Calling Agents MM-ToolSandBox:当聪明的模型不会看图——一个暴露工具调用智能体视觉感知瓶颈的评测框架
- SpectraReward: Turning MLLMs into Zero-Shot Reward Models by Reading Prompts Back from Images SpectraReward:让多模态大模型通过"回读提示词"变成零样本奖励模型
- Requential Coding: Measuring What a Model Actually Learns, Not How Big It Is 重序编码:衡量模型实际学到了什么,而非它有多大
- ConceptSMILE: An Independent Auditor for Concept-Based AI Explanations ConceptSMILE:为概念级AI解释建立独立审计机制
- Deep Gaussian Processes Meet DAGs: Compositional Uncertainty on Graph Structures 深度高斯过程遇见DAG:图结构上的组合不确定性
- From Chatbot to Chip Architect: LLMs as the Next EDA Front-End Designers 从聊天机器人到芯片架构师:大语言模型作为下一代前端设计工具
- OpenLongTail: Generating Missing Camera Views to Scale Long-Tail Driving Data OpenLongTail:生成缺失视角以扩展长尾驾驶数据
- The Shape of Dreams: Topological EEG Analysis for Dream Content Classification and Synthesis 梦境的形状:基于拓扑的EEG梦境内容分类与信号合成
- Seeing is Learning: Visual Pretraining Beats Text for Language Models 看见即是学习:视觉预训练超越纯文本的语言模型训练
- Semantic Pareto-DQN: Multi-Objective RL Breaks the Fraud Collapse Trap 语义帕累托-DQN:用多目标强化学习破解金融异常检测中的"欺诈塌缩"陷阱
- Tokenizer Transplantation: Making English ASR Models Speak Bengali 词表移植:让英语 ASR 模型学会说孟加拉语
- Sign Language Translation Gets Real-Time: A Systems Paper Disguised as an SLT Paper 手语翻译实现实时化:一篇披着SLT外衣的系统论文
- Can LLM Agents Autonomously Hack IoT Devices? VEXAIoT Says Yes VEXAIoT:用 AI 智能体自动发现和利用 IoT 漏洞
- Canvas360: Teaching AI to Think in 360 Degrees Canvas360:让 AI 学会 360 度全景思维
- MulTTiPop: The Real-World Multitrack Benchmark That Reveals How Far Music Transcription Still Has to Go MulTTiPop:揭示自动音乐转录还有多远的多轨基准数据集
- Score Accuracy ≠ Stability: A Hidden Trap in Diffusion Sampling 分数精度≠稳定性:扩散采样中的隐性陷阱
- How 77,000 Students Actually Use AI Tutors — A Reality Check on EdChatbot Research 77,000 名学生如何实际使用 AI 助教——对教育聊天机器人研究的一次现实检验
- Agreement ≠ Validity: Why Your LLM Annotator Might Be Faking It 一致≠有效:你的LLM标注员可能在作弊
- ARDY: Real-Time Motion Generation That Actually Listens to You ARDY:真正能听懂你指令的实时动作生成
- UMAP's Hidden Graph: Treating the Discarded kNN Structure as First-Class Analysis UMAP 隐藏的图结构:把被丢弃的 kNN 图当作一等分析对象
- PanoLOG: Breaking the Parallax Curse in Panoramic 3D Reconstruction PanoLOG:打破全景3D重建中的视差魔咒
- IdeaGene: Benchmarking How AI Traces Scientific Thought Inheritance IdeaGene:基准测试AI追踪科学思想传承的能力
- LongE2V: Using Video Diffusion Priors to Reconstruct, Predict, and Interpolate from Sparse Event Streams LongE2V:用视频扩散先验从稀疏事件流中重建、预测和插值
- Video as Chain-of-Frame Reasoning 以视频为帧链推理
- SLORR: Making Neural Networks Compressible During Training with Minimal Overhead SLORR:以极低开销在训练中让神经网络变得可压缩
- Super Weights: Important But Untamable 超级权重:重要却不可驯服
- Quantized LLMs Aren't What They Seem: Why Accuracy Fails and What to Measure Instead 量化后的LLM并非表面那样:为何准确率会失效,以及应如何度量
- When Your Workflow Becomes a Fossil Record: Semantic Persistence for LLM Workflows 当工作流变成化石记录:LLM 工作流的语义持久化
- AutoPilot-VQA: Can VLMs Really Reason About Driving Incidents? AutoPilot-VQA:视觉语言模型真的能推理驾驶事故吗?
- Keeping the Teacher Honest: On-Policy Self-Distillation Fixes Long-Video Autoregressive Degradation 让老师保持诚实:在线策略自蒸馏修复长视频自回归退化
- UniClawBench: A Benchmark for Testing AI Agents in the Real World UniClawBench: 面向真实世界主动智能体的通用基准
- Wat3R: Teaching 3D Vision to See Through Water Wat3R:让3D视觉看透水域
- ZipDepth: Shrink Ray for AI Depth Perception, From Cloud to Pocket ZipDepth:为AI深度感知打造的缩放仪,从云端到掌心
- Agon: When Two AI Models Compete, Both Learn to Think Better Agon:当两个 AI 模型互相竞争,双方都学会了更好地思考
- Bridging the Size Gap: A Unified Sampling Theory for Variable-Size Learning 跨越尺寸鸿沟:面向可变尺寸学习的统一采样理论
- Jailbreak: LLM-Generated Storage Readers Bypass Database Engines for 27x Analytics Speedup 越狱:用大语言模型生成存储读取器,绕过数据库引擎实现27倍分析加速
- CO-LMLM: Teaching Language Models to Ask Questions in a Language Only Machines Understand CO-LMLM:让语言模型用机器才懂的语言提问
- DiaLLM: Understanding a Dialect Doesn't Mean You Can Speak One DiaLLM:听懂方言不等于能说方言
- Your LLM Knows What It Doesn't Know — But Won't Tell You 大模型知道自己不知道什么——但就是不说
- STRACE: Finding the Needle of Causality in the Haystack of Agent Traces STRACE:在智能体轨迹的干草堆中找到因果关系的针
- How Training Data Sculptures RoPE's Frequency Landscape 训练数据如何雕刻 RoPE 的频率地貌
- Red-Teaming the Rules: How Deployment Institutions Causally Shape Multi-Agent AI Safety 红队测试的目标是规则:部署制度如何因果性地决定多智能体AI安全
- AdaPrefix-GRPO: Turning the Hardest Problems into the Best Teachers AdaPrefix-GRPO:把最难的题变成最好的老师
- NOTES: Compressing the Design Space Before You Search It NOTES:先压缩设计空间,再搜索
- Scaling MoE Video Pretraining for Embodied Intelligence: LingBot-Video 面向具身智能的MoE视频预训练:LingBot-Video
- Sharpening Diffusion RLHF: Weighting Key Steps for 6x Sample Efficiency 为扩散模型 RLHF 提速:通过加权关键步骤实现 6 倍样本效率
- SkillCenter: Building a Verified Knowledge Base for AI Agents at Scale SkillCenter:为AI智能体构建大规模可验证知识库
- Why Delta Beats Gate: The Analysis Behind Linearizing Transformers Delta为何胜过Gate:Transformer线性化的分析基础
- Teaching VLMs to Think with Their Eyes: Grounding Physical Reasoning Through Visual-Action Alignment 教VLM用眼睛思考:通过视觉-动作对齐实现物理推理的落地
- CAIRN: Teaching AI to Navigate Homes by Learning the Floor Plan CAIRN:通过学习户型图让AI理解多房间家居场景
- DepthWeave-KV: Smarter KV Cache Compression by Sharing Across Layers and Adapting Per Token DepthWeave-KV:跨层共享 + 逐 Token 自适应的 KV 缓存压缩新方案
- ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation ELSA3D:用于统一3D理解与生成的弹性语义锚定
- ReChannel: Read Dense Fields Directly from Diffusion Tokens, Skip the VAE Decoder ReChannel:从扩散模型 Token 直接读取密集预测场,跳过 VAE 解码器
- Lift3D-VLA: Teaching Robots to See in 3D by Upgrading Their Flat Vision Lift3D-VLA:通过升级平面视觉让机器人学会三维操作
- Skeleton-Guided Autoregressive Video for Closed-Loop Driving Simulation 骨架引导的自回归视频生成:面向闭环自动驾驶仿真
- ProxyPose: Turning Pose Tracking into a Video Translation Problem ProxyPose:把六自由度位姿追踪变成视频翻译问题
- Rethinking Indic AI: When Language Technology Erases Cultural Worldviews 重新思考印度AI:语言技术如何无形中抹去文化世界观
- One Model to Rule All Vision Tasks: Unified Multimodal Generation for Computer Vision 一个模型统治所有视觉任务:面向计算机视觉的统一多模态生成
- From Context Overflow to Learned Compression: How CompactionRL Trains Long-Horizon LLM Agents 从上下文溢出到学习压缩:CompactionRL如何训练长时域LLM智能体
- Cortex: Bridging the Semantic-Kinematic Gap in Long-Horizon Robot Manipulation Cortex:弥合长时序机器人操作中的语义-运动学鸿沟
- Deform360: When Robots Learn to Feel and See Deformable Objects at Scale Deform360:让机器人在大规模下同时感知可变形物体的视觉与触觉
- Offline RL Without Bellman Completeness: FORE Drops the Hardest Assumption 离线强化学习无需贝尔曼完备性:FORE 去掉了最强假设
- CamVLA: Letting Robot Policies Figure Out Where the Camera Is CamVLA:让机器人策略自己弄清楚相机在哪
- Graph-as-Policy: Using Computation Graphs to Make Robots More Reliable for Variable Tasks 以图为策:利用计算图让机器人在可变任务中更可靠
- Graph Sparse Sampling: Break the Horizon Curse by Sharing Futures 图稀疏采样:用共享未来打破视野诅咒
- Verification as a Scaling Axis: LLM-as-a-Verifier's Probabilistic Approach to Solution Judging 验证即扩维:LLM-as-a-Verifier 如何用概率方法重构评判范式
- MV-Forcing: Bridging Long Video and Multi-View Generation with 4D Geometry MV-Forcing:用4D几何桥梁统一长视频与多视角生成
- PixWorld: Unifying 3D Reconstruction and Generation in Raw Pixel Space PixWorld:在原始像素空间统一3D重建与生成
- REDDIT: Fixing Timestamp Drift in ASR Without Forgetting — Replay-Based Distribution Editing REDDIT:基于重放分布编辑的 ASR 时间戳漂移修正与遗忘规避
- Search Beyond What Can Be Taught: Teaching Generators When to Look Things Up 超越已知边界:教会视觉生成器何时该求助外部搜索
- SynCity 3000: Scaling 3D Scene Generation with Convolutional Diffusion SynCity 3000: 用卷积扩散生成场景级3D内容
- Distilling the Change, Not the Champion: A Smarter Way to Transfer RL Improvements 蒸馏变化,而非冠军:一种更聪明的强化学习迁移方法
- One Object, Three Coordinates: A Unifying View of Discrete Diffusion 一个对象,三种坐标:离散扩散模型的统一视角
- When Does Neural Guidance Actually Speed Up Symbolic Solvers? 神经引导何时才能真正加速符号求解器?
- GeoMix: Bridging the Accuracy Gap in Descriptor-Free Visual Localization GeoMix:用全局上下文与多检测器训练弥合无描述子视觉定位的精度鸿沟
- Align4D: Turning Any Input into Animated 3D Objects Through Smart Alignment Align4D:通过对齐技术将任意输入转化为动态3D物体
- Beyond Adam: SOAP and Muon Optimizers for Faster MLIP Training 超越Adam:SOAP与Muon优化器加速机器学习原子间势能训练
- Controllable Traffic Agents with Behavior Latents: Steering Simulation Without Losing Realism 行为潜变量可控交通智能体:保持真实性的同时实现仿真可控
- DemoPSD: Teaching LLMs to Reason Without Leaking the Answer DemoPSD:教大模型推理,但别把答案藏进捷径
- The Long Con: How Persistent AI Codebases Enable Distributed Attacks 持久化代码仓库中的AI分布式攻击:藏在多次提交里的恶意
- Diffusion Self-Alignment: It Was Data Augmentation All Along 扩散模型自对齐:原来一直是数据增强在起作用
- LACUNA: When Unlearning Misses the Bullseye — A New Testbed for Parameter-Level Erasure Precision LACUNA:当机器遗忘偏离靶心——参数级擦除精度测试平台
- Real-Time Safety Alarms for LLMs: When to Pull the Emergency Brake 为大语言模型装上实时安全警报:何时该踩急刹车
- PointDiT: When Simplicity Beats Complexity in 3D Reconstruction PointDiT:简洁之道胜过复杂堆砌的3D重建
- Teaching LLMs to Play TV Detective: Reasoning-Based Speaker Recognition for Long Dramas 教大模型当电视剧侦探:基于推理的长剧说话人识别
- Seek to Segment: Teaching Agents to Actively Search 360° Scenes for Objects Seek to Segment:让智能体在360°全景场景中主动搜寻目标物体
- Training-Free Defense Against Typographic Attacks via Mechanistic Interpretability 无训练的排版攻击防御:基于机械可解释性的概念定位
- Teaching Vision-Language Models to Actually Look When They Think Twice 教会视觉语言模型在反思时真正看一眼
- When LLMs Say One Thing and Think Another: A Study of Emergent Social Compliance in Multi-Agent Debates 当大模型说一套想一套:多智能体辩论中涌现的社会性服从研究
- Embodied.cpp: A Portable C++ Runtime That Unifies Embodied AI Model Deployment Across Heterogeneous Robots Embodied.cpp:统一异构机器人上具身AI模型部署的可移植C++运行时
- From Function Calls to Function Factories: PAW Compiles Prompts into Tiny Adapters 从函数调用到函数工厂:PAW 将提示词编译成微型适配器
- ReContext: When LLMs Can See But Can't Use — Fixing the Last Mile of Long-Context Reasoning ReContext:大模型看得见却用不上——修补长上下文推理的最后一英里
- WorldDirector: When World Models Learn to Separate Directing from Filming WorldDirector:当世界模型学会把导演和摄影分开
- One Layer Is Enough: RL Post-Training Concentrates in the Middle of Transformers 一层足矣:RL 后训练的收益集中在 Transformer 中间层
- Neural Certificate Pricing: Learning Dual Prices to Solve Combinatorial Optimization 神经证书定价:学习对偶价格求解组合优化
- Perceive-to-Reason: Teaching Vision Models to See Before They Think 先看后想:感知与推理解耦的细粒度视觉推理框架
- Bridging RL and SFT: Training Language Models with Verifiable Rewards Plus Human Style 弥合强化学习与监督微调:用可验证奖励加人类风格训练语言模型
- Why Transformers Should Separate Memory from Prediction 为什么 Transformer 应该把记忆和预测分开
- Cross-Space Distillation: Bridging the Latent Gap Between Diffusion Teachers and Students 跨空间蒸馏:弥合扩散模型师生之间的潜在空间鸿沟
- FedLAB: Traceable Semantic Codebooks for Federated Multimodal Graph Foundation Learning FedLAB:面向联邦多模态图基础学习的可追溯语义码本
- Freeform Preference Learning: Teaching Robots with Multi-Axis Human Preferences 自由形式偏好学习:用多轴人类偏好教会机器人
- GEAR: Letting the Generator Teach Its Own Tokenizer GEAR:让生成器教会自己的分词器
- Introspective Coupling: Why Fixed Explanation Datasets Can Faithfully Track a Shifting Model 内省耦合:为何固定的解释数据集能忠实追踪变化中的模型
- QVal: A Training-Free Testbed for Evaluating Dense Supervision in LLM Agents QVal:无需训练的 LLM 智能体密集监督信号评估平台
- Teaching LLMs to Know What They Don't Know: Metacognitive Reinforcement Learning 教会大模型"知道自己不知道":元认知强化学习
- SpheRoPE: Making Diffusion Models Think in Spheres Without Retraining SpheRoPE:无需重训,让扩散模型在球面上思考
- Surrogate Fidelity: When Open Models Can (and Can't) Explain Closed Ones 替代保真度:开源模型何时能解释闭源模型?
- Mesa: Prioritizing Vulnerable Communication Channels for Securing Multi-Agent Systems 优先保护关键通道:多智能体系统安全新方法
- When Playing It Safe Backfires: Conservative Offline Training Makes Reward Hacking Worse 越保守越危险:离线训练的保守主义如何反噬在线对齐
- Agents-A1: Scaling the Agent Horizon Instead of Parameters Agents-A1:扩展智能体视野而非扩展参数量
- Self-Evolving Memory for LLM Agent World Models 让世界模型自我进化:LLM代理的智能驾驶教练
- Words Speak Louder Than Code: How Cognitive Heuristics Trick LLMs Into Missing Vulnerabilities 文字比代码更响亮:认知启发式如何欺骗大模型错过漏洞
- Agent-Native Immune System: Architecture, Taxonomy, and Engineering Agent原生免疫系统:架构、分类学与工程
- Democratic ICAI: Debating Our Way to Steering Principles from Preferences 民主逆宪法AI:通过辩论从偏好中推导指导原则
- DexCompose: Reusing Dexterous Policies for Multi-Task Manipulation with a Single Hand DexCompose: 用单只手复用灵巧策略完成多任务操作
- PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception PerceptionRubrics:校准多模态评估以对齐人类感知
- StructSplat: Generalizable 3D Gaussian Splatting from Uncalibrated Sparse Views StructSplat:从无标定稀疏视图进行可泛化3D高斯泼溅
- Surprises in Proper Positive-Only Learning: What You Can (and Can't) Learn from Positive Data Alone 正确正样本学习中的意外:仅从正数据中能(与不能)学到什么
- Towards Automating Scientific Review with Google's Paper Assistant Tool 迈向科学评审自动化:谷歌论文助手工具
- VGB for Masked Diffusion Model: Efficient Test-time Scaling for Reward Satisfaction and Sample Editing 掩码扩散模型的 VGB 方法:面向奖励满足与样本编辑的高效测试时缩放
- Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models 视觉默认,先验覆盖:视觉语言模型中感知-知识冲突的因果机制
- Which Nash Equilibrium? Solver-Dependent Selection on Zero-Sum Nash Polytopes 哪个纳什均衡?零和博弈纳什多面体上的求解器依赖性选择
- Autoregressive Boltzmann Generators: Breaking the Flow Bottleneck in Molecular Sampling 自回归玻尔兹曼生成器:打破分子采样的流模型瓶颈
- Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning 通过自主经验探索与事后经验利用赋能GUI代理的任务规划
- Error-Conditioned Neural Solvers: Learning to Correct Predictions by Reading Residual Fields 误差条件神经求解器:通过读取残差场学习纠正预测
- Language-Based Digital Twins for Elderly Cognitive Assistance 基于语言的数字孪生用于老年人认知辅助
- Mapping Political-Elite Networks in Europe with a Multilingual Joint Entity-Relation Extraction Pipeline 欧洲政治精英网络的多语言联合实体关系抽取管道
- Not All Actions Are Equal: Rethinking Conditioning for Dexterous World Models 并非所有动作都同等重要:重新思考灵巧世界模型的条件范式
- RoPEMover: Depth-Aware Object Relocation via Positional Embeddings 基于位置嵌入的深度感知物体重新定位
- SAM2Matting: Turning Video Trackers into High-Fidelity Matting Machines SAM2Matting:将视频追踪器变成高保真抠图机器
- Understanding Domain-Aware Distribution Alignment in Budgeted Entity Matching 理解预算有限实体匹配中的领域感知分布对齐
- When are likely answers right? On Sequence Probability and Correctness in LLMs 何时高概率答案是正确的?——关于LLM中序列概率与正确性的研究
- Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards 问答生成:通过自洽奖励实现自我进化的统一多模态理解与生成
- DanceOPD: On-Policy Generative Field Distillation DanceOPD:在线策略生成场蒸馏
- DnA: Denoising Attention for Visual Tasks DnA:去噪注意力,让视觉注意力更干净
- Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance 不要止步于模式!通过特征自引导减轻预训练流模型中的多样性崩溃
- Hallucination in World Models is Predictable and Preventable 世界模型中的幻觉:可预测且可预防
- VISE: Teaching Large Multimodal Models to Actually Look at Images VISE:教大模型多看图像少猜答案
- PhysiFormer: Forecasting 3D Object Motion via Diffusion in World Space PhysiFormer:在世界空间中通过扩散模拟力学
- RayPE: Ray-Space Positional Encoding for 3D-Aware Video Generation RayPE:面向3D感知视频生成的射线空间位置编码
- RL without Ground-Truth Solutions: RiVER Improves LLMs on Score-Based Tasks 无Ground-Truth的强化学习:RiVER在分数型任务上提升LLM
- World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays 世界动作模型助力持续模仿学习:基于循环生成回放
- Learning Action Priors for Cross-embodiment Robot Manipulation 面向跨实体机器人操作的动作先验学习
- Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment 模型取证:探究令人担忧的行为是否源于对齐问题
- MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation MVTrack4Gen:多视点跟踪作为4D视频生成的几何监督
- Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining 自然非领悟:预训练中规则存留的非对称控制
- Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents 后训练中的免费午餐:面向LLM智能体的进度优势
- On-Policy Self-Distillation with Sampled Demonstrations: The Hidden Cost of Reduced Output Diversity 带采样示范的在线自蒸馏:输出多样性降低的隐性代价
- Real-Time Voice AI Hears but Does Not Listen 实时语音AI:听见了,但没听进去
- RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments RevengeBench:从行为实验逆向工程代码空间策略
- Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal LLMs 同一证据,不同答案:多模态大模型的顺序敏感性审计
- TryOnCrafter: Camera-Controllable Video Virtual Try-on with a 4D Proxy TryOnCrafter:利用可渲染4D试穿代理实现摄像机轨迹可控的视频虚拟试穿
- BenchX: Benchmarking AI Models for Cancer Detection and Localization with Demographic and Protocol Biases BenchX:面向人口统计学与成像协议偏差的AI癌症检测模型基准评测
- Bridging the Manifold Gap: Riemannian Residual Line Search for One-Step Image Editing 弥合流形鸿沟:用于单步图像编辑的黎曼残差线搜索
- DiffusionBench: Time to Stop Evaluating Diffusion Transformers on ImageNet Alone DiffusionBench:别再只用ImageNet评估扩散Transformer了
- FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation FLAT:前馈潜在三角形泼溅用于几何精确的场景生成
- FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation FLUX3D:基于扩散对齐稀疏表示的高保真3D高斯生成
- GeoT2V-Bench: Diagnosing 3D Consistency of Camera-Prompted Text-to-Video via Reconstruction GeoT2V-Bench:基于3D重建诊断相机提示文本到视频的3D一致性
- Grading the Grader: How to Reliably Evaluate an Agentic Data Analysis System 给评分者评分:如何可靠评估智能数据分析系统
- InSight: Making VLAs Steerable for Self-Guided Skill Acquisition InSight:让视觉-语言-动作模型具备可控性,实现自主技能获取
- It's Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces 复杂之网:AI增强AAC界面的设计与评估
- IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation IV-CoT:隐式视觉思维链用于结构感知文本到图像生成
- New Bounds for the Last Iterate of the Stochastic subGradient Method 随机次梯度方法最后迭代的新界
- OpenThoughts-Agent: Data Recipes for Agentic Models OpenThoughts-Agent:智能体模型的数据配方
- Real vs. Complex Spectral Bases for Neural Operators: The Role of Green's Function Alignment 实谱与复谱基在神经算子中的对决:格林函数对齐的作用
- Spherical-to-ERP Epipolar Rectification for Single-Axis Disparity in 360 Stereo 球面到ERP极线校正:360度立体图像的单轴视差
- World Models in Pieces: Structural Certification for General Agents 分块的世界模型:通用智能体的结构性认证
- AIR: Adaptive Interleaved Reasoning with Code in MLLMs AIR:多模态大模型中自适应交错代码推理
- AutoDex: An Automated Real-World System for Dexterous Grasping Data Collection AutoDex:自动化真实世界灵巧抓取数据收集系统
- Can LLMs Reliably Self-Report Adversarial Prefills, and How? LLM 能可靠地自我报告对抗性前缀攻击吗?如何做到的?
- CoorDex: Coordinating Body and Hand Priors for Continuous Dexterous Humanoid Loco-Manipulation CoorDex:协调身体与手部先验实现连续灵巧人形机操作
- GeoFidelity-Bench: Do Text-to-Image Street-View Models Actually Know Which Road They’re Drawing? GeoFidelity-Bench:文生图街景模型到底知不知道自己在画哪条路?
- Keep The Essentials: Efficient Reference Conditioned Generation via Token Dropping 守住精华:基于Token丢弃的高效参考条件生成
- MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems? MAS-PromptBench:提示优化何时能改善多智能体大模型系统?
- Can AdamW Survive Heavy-Tailed Noise? An Open Problem with a Corridor Lower-Bound AdamW 能在重尾噪声下存活吗?一个带走廊下界设定的开放问题
- Randomized YaRN: Teaching LLMs to Reason Over Lengths They've Never Seen 随机YaRN:让大语言模型学会在未见过的超长文本中推理
- Tapered Language Models: A Free Lunch Hiding in Plain Sight 锥形语言模型:藏在眼皮底下的免费午餐
- Execution-State Capsules: Full-Stage Checkpoint on Device for Low-Latency AI 执行状态胶囊:面向低延迟设备端 AI 的完整阶段检查点机制
- How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech 语言指令如何影响语音生成?风格化TTS中的交叉注意力归因研究
- How Transparent is DiffusionGemma? DiffusionGemma 有多透明?
- LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents LedgerAgent:面向策略遵循的工具调用智能体的结构化状态管理
- Multi-Task Bayesian In-Context Learning 多任务贝叶斯上下文学习
- Predictability: A Fine-Grained Privacy Measure Beyond Differential Privacy 可预测性:超越差分隐私的细粒度隐私度量
- G2Rec: Unifying Global Graph Context and Semantic Tokenization for Generative Recommendation G2Rec:统一全局图结构与语义分词的生成式推荐
- Toward Calibrated Mixture-of-Experts Under Distribution Shift 分布偏移下的混合专家模型校准
- UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning UNIEGO:代理作为中介实现统一的第一人称视频表征学习
- VisDom: Sparse Novel View Synthesis with Visible Domain Constraint VisDom: 基于可见域约束的稀疏新视角合成
- CalTennis: Large Multi-View Tennis Video Dataset and Benchmark of Monocular-to-3D Pose Estimation CalTennis:大型多视角网球视频数据集及单目到3D姿态估计基准
- World Models Have No Persistence: When You Look Away, the Moon Stops Moving 世界模型缺乏持久状态核心:当你不看时,月亮就不走了
- JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising JanusMesh:跨空间去噪实现快速零样本3D视觉错觉生成
- Optimal Deterministic Multicalibration and Omniprediction 最优确定性多校准与全能预测
- SSD: Spatially Speculative Decoding Accelerates Autoregressive Image Generation SSD:空间投机解码加速自回归图像生成
- Only 15 Visual Cues Explain 80% of Social Biases in MLLMs 多模态模型中仅15种视觉线索驱动80%社会偏见
- The FID Lottery: Why Your Benchmark Number Is Luckier Than You Think FID彩票:你报告的那个数字,有多大成分是运气?
- The Token Is a Group Element: On Lie-Algebra Attention over Matrix Lie Groups 令牌即群元:论矩阵李群上的李代数注意力机制
- Thinking in Boxes: 3D Editing in Real Images Made Easy 用盒子思考:轻松编辑真实图像的3D变换
- TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning TimeProVe:先提案再验证——高效的长视频时序推理
- Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games 超越当前观察:受控非马尔可夫游戏中多模态大语言模型的评估
- Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents 数据智能体:通过自主编码体解释、建模和查询企业数据
- Do as I Do: Dexterous Manipulation Data from Everyday Human Videos 跟我做:从日常人类视频中提取灵巧操作数据
- Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play 通过多智能体虚拟博弈增强大语言模型决策能力
- Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States 用LOCUS解放法律:美国地方法规语料库
- Learning User Simulators with Turing Rewards 基于图灵奖励的用户模拟器学习
- OmniAgent: Native Active Perception as Reasoning for Long Video Understanding OmniAgent:将长视频理解化为主动感知推理的原生全模态智能体
- Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation 重新思考奖励监督:基于评分标准的自蒸馏
- The Chandra-Gaia Catalog of Counterparts: Resolving ambiguous Gaia matches to X-ray sources in the Chandra Source Catalog using Machine Learning 钱德拉-盖亚对应体目录:利用机器学习解决钱德拉源目录中X射线源与盖亚的模糊匹配
- UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning UBP2:面向高效偏好强化学习的不确定性平衡偏好规划
- Adaptive Volumetric Mechanical Property Fields Invariant to Resolution 自适应体积力学属性场:分辨率无关的材质预测
- EventDrive: Event Cameras for Vision-Language Driving Intelligence EventDrive:用事件相机驱动视觉-语言驾驶智能
- EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation EvolveNav: 零样本物体目标导航中的主动预思考与自演化记忆
- Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers 固定点推理器:稳定且自适应的深度循环Transformer
- FR3D: Future Dynamic 3D Reconstruction with Disentangled Ego-Motion FR3D:解耦自运动的未来动态三维重建
- Looped World Models: Iterative Latent Refinement as a New Scaling Axis for World Simulation 循环世界模型:迭代潜在状态精炼——世界模拟的新缩放维度
- Seeing Is Not Screening: Multimodal Hidden Instruction Attacks on Agent Skill Scanners 眼见不为真:面向智能体技能扫描器的多模态隐藏指令攻击
- Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification 统一多模态自回归建模:共享上下文-视觉分词器是关键
- Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement 视觉验证实现推理时引导与自主策略改进
- Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients 最近发展区策略优化:教师待在提示里,而非梯度里
- Benchmarking LLM Agents on Meta-Analysis: The Screening Bottleneck LLM智能体在元分析基准测试中的筛选瓶颈
- BRDFusion: Physics Meets Generation for Urban Scene Inverse Rendering BRDFusion:物理与生成模型融合,实现城市场景逆渲染
- Context-Aware RL: Teaching LLMs to Read the Fine Print 上下文感知强化学习:教大模型读懂字里行间
- DeepRubric: Evidence-Tree Rubric Supervision for Efficient RL of Deep Research Agents DeepRubric:基于证据树的评分标准监督,实现深度研究智能体的高效强化学习
- Exact Posterior Score Estimation for Solving Linear Inverse Problems 精确后验分数估计:求解线性逆问题的新范式
- ExpRL: Exploratory RL for LLM Mid-Training ExpRL:面向LLM中间训练的探索性强化学习
- Geometric Action Model for Robot Policy Learning 几何动作模型:机器人策略学习的3D飞跃
- HAMON: Passive Optical Sequence Mixing for Long-Horizon Forecasting HAMON:被动光学序列混合用于长程预测
- Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes 面向稀疏回合结果的VLA在线RL微调的分层优势加权
- KVEraser: Learning to Steer KV Cache for Efficient Localized Context Erasing KVEraser:学习引导KV缓存以实现高效局部上下文擦除
- Qwen-RobotWorld: Unifying Embodied World Modeling through Language-Conditioned Video Generation Qwen-RobotWorld:通过语言条件视频生成统一具身世界建模
- The Importance of Phase in Neural Representations: An Internal Oppenheim-Lim Test of Image Classifiers 相位在神经表示中的重要性:图像分类器的内部Oppenheim-Lim测试
- The Value Axis: Language Models Encode Whether They’re on the Right Track 价值轴:语言模型内部编码了‘当前方向是否正确’
- TokenPilot: Cache-Efficient Context Management for LLM Agents TokenPilot:面向LLM代理的缓存高效上下文管理
- Your Privacy My Cloak: How Differential Privacy Became a Shield for Backdoor Attacks 你的隐私,我的斗篷:差分隐私如何变成后门攻击的掩护
- How Hard Is It to Find the Worst Group? A Complexity Measure for Active Multi‑group Mean Estimation 最难的那组有多难?多组均值主动学习的复杂度度量
- AdaSR: Streaming Reasoning That Learns When to Think — A Hierarchical RL Approach AdaSR:自适应流式推理——分层强化学习教会模型何时思考
- AgentSpec: Why Your LLM Agent's Scaffold Matters More Than the Model AgentSpec:LLM智能体的脚手架比模型本身更重要
- ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning ClinHallu:诊断医学多模态大语言模型推理中阶段性幻觉的基准
- Compressed Computation is (probably) not Computation in Superposition 压缩计算(可能)不是叠加计算
- CORA: Bridging the Thinking-Answer Gap in Multimodal RLVR CORA:弥合多模态RLVR中的思考-答案鸿沟
- Gaze Heads: How VLMs Look at What They Describe 凝视头:视觉语言模型如何"看"它们所描述的东西
- Instruct-Particulate: Scaling Feed-Forward 3D Object Articulation with Kinematic Control 指令粒子:通过运动学控制实现前馈3D物体关节化的大规模扩展
- Memento: Reconstruct to Remember for Consistent Long Video Generation Memento:通过重建来记住——长视频生成中的主体一致性
- Optimal Hidden-Target Learning for Online Inventory Optimization on General Convex Sets 在线库存优化的最优隐目标学习:一般凸集上的最优遗憾界
- Persona-Pruner: Sculpting Lightweight Role-Playing Models Persona-Pruner:雕琢轻量角色扮演模型
- RATS! Patches Talk Through Registers: Emergent Parts in Register Attention Transformers RATS!补丁通过寄存器对话:寄存器注意力Transformer中涌现的部件
- RepFusion: Let Multimodal LLMs Denoise Visual Representations Directly RepFusion:让多模态大语言模型直接对视觉表示去噪
- Direct Cache-Based Synthesis for Parallel LLM Agent Branches 面向 LLM Agent 工作流中并行分支的隐空间直接合成
- When to Write and When to Suppress: Route-Specialized Dual Adapters for Memory-Assisted Knowledge Editing 何时写入,何时抑制:用于记忆辅助知识编辑的路由专用双适配器
- Before You Think: System 0, AI-Mediated Cognition and Cognitive Colonization 在你思考之前:系统0、AI中介认知与认知殖民
- EurekAgent: Environment Engineering is the Key to Autonomous Scientific Discovery EurekAgent:环境工程是自主科学发现的关键
- Operadic Consistency: Catching LLM Reasoning Failures Without Labels 操作一致性:无需标注的 LLM 推理失败检测信号
- SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation SkMTEB:斯洛伐克语大规模文本嵌入基准与模型适配
- Surflo: Consistent 3D Surface Flow with a Global State – Arbitrary Resolution from Unposed Views Surflo:带全局状态的3D表面流模型——从无位姿视图实现任意分辨率解码
- Agents-K1: From Flat Citations to Agent-Native Knowledge Graphs Agents-K1:从平铺引文到智能体原生的知识图谱
- Automated reproducibility assessments in the social and behavioral sciences using large language models 用大语言模型自动化社会与行为科学的可重复性评估
- Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation 密集监督,稀疏更新:论策略内蒸馏的稀疏性与几何结构
- EvoArena & EvoMem: Teaching LLM Agents to Remember How the World Changed EvoArena 与 EvoMem:教大语言模型代理记住世界是如何变化的
- Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction Flex4DHuman: 灵活的多视角视频扩散实现4D人体重建
- HyperTool: Beyond Step-Wise Tool Calls for Tool-Augmented Agents HyperTool:超越逐步工具调用的工具增强智能体
- Influcoder: Distilling Decoders' Gradient Influence Rankings into an Encoder for Data Attribution Influcoder:将解码器梯度影响排名蒸馏到编码器用于数据归因
- InterleaveThinker: Enabling Interleaved Text-Image Generation with Multi-Agent Reinforcement Learning InterleaveThinker: 用多智能体强化学习实现图文交叠生成
- Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning 通过检索增强强化微调学习类比推理
- Manas: Dexterous Manipulation of Articulated Tools Manas:铰接工具的灵巧操控框架
- Modality Forcing: Scalable Spatial Generation with Sparse Depth 模态强制:用稀疏深度实现可扩展的空间生成
- RepWAM: World Action Modeling with Representation Visual-Action Tokenizers RepWAM:用表示视觉-动作分词器进行世界行动建模
- SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning SpatialClaw:重新思考具身空间推理的动作接口
- Understanding Truncated Positional Encodings for Graph Neural Networks 理解图神经网络的截断位置编码
- World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible 世界追踪:超越可见的生成式像素对齐几何
- Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling 打破熵界:利用多令牌预测与拒绝采样加速强化学习训练
- Context-Driven Incremental Compression for Multi-Turn Dialogue Generation 面向多轮对话的上下文驱动增量压缩方法
- DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners? DIRECT:具身规划器中测试时计算应当何时何地分配?
- Doc-to-Atom: Learning to Compile and Compose Memory Atoms Doc-to-Atom:学习编译和组合记忆原子
- FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning FACTR 2:为通用机器人手臂学习外力感知以改进策略学习
- How Seemingly Inconsequential Design Choices Dictate LLM Performance in Pathology 看似无关紧要的输入配置如何决定病理学LLM的性能
- Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models 重路由,不删除:面向视觉语言模型的可恢复视觉令牌路由
- Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization 可验证环境是乐高积木:递归组合实现推理泛化
- VLGA: Adding Dense 3D Geometry to Vision-Language-Action Models for Autonomous Driving VLGA:为自动驾驶的视觉-语言-动作模型注入稠密3D几何
- Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs 我们的模型建立在哪些模型之上?审计现代LLM中不可见的依赖关系
- A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design 监督微调的统一视角:通过目标分布设计
- ABC-Bench: The First Benchmark for AI Agents Doing Real Biology Lab Work ABC-Bench:首个评估AI智能体真实生物实验能力的基准
- ARM: Unifying Image Understanding, Generation, and Editing with Discrete Tokens and RL ARM:用离散Token和强化学习统一图像理解、生成与编辑
- COGENT: Continuous Graph Emulators with Neural ODEs for Long-Term Physical Forecasting COGENT:基于连续图与神经常微分方程的长期物理预测仿真器
- Data Journalist Agent: From Raw Data to Verifiable Multimedia Stories 数据记者代理:从原始数据到可验证的多媒体故事
- EEVEE: Real-World Test-Time Prompt Learning for Self-Improving LLM Agents EEVEE:面向真实世界的测试时提示学习,让大模型智能体自我进化
- The LLM Automation Narrative Has Flaws: Variance and Error Magnitude Matter LLM自动化叙事存在缺陷:方差与错误幅度不容忽视
- Itô Maps for Any-Step SDEs: Single-Pass Stochastic Flow for Posterior Sampling 任意步SDE的伊藤映射:单次前向随机流用于后验采样
- Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization Lip Forcing:用于实时唇同步的少步自回归扩散
- Mean Flow Distillation: Robust and Stable Distillation for Flow Matching Models 均值流蒸馏:用于流匹配模型的鲁棒稳定蒸馏方法
- Next Forcing: Causal World Modeling with Multi-Chunk Prediction Next Forcing:基于多块预测的因果世界建模
- Piper: A Programmable Distributed Training System Piper:可编程的分布式训练系统
- Predicting Future Behaviors in Reasoning Models Enables Better Steering 预测推理模型未来行为以实现更优引导
- ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models ReasonAlloc:推理模型中解码阶段KV缓存预算的分层分配方法
- When to Align, When to Predict: A Phase Diagram for Multimodal Learning 何时对齐,何时预测:多模态学习的相位图
- Agency Transfer: How to Upgrade a Suboptimal Policy Without Starting from Scratch 代理人移交:如何不从头训练就升级一个次优策略
- Causal Evaluation of Formal Language Learnability: Why Correlation Fails 形式语言任务可学习性的因果评估:相关性为何失效
- Latent Spatial Memory for Video World Models 视频世界模型的潜在空间记忆
- MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models MemoryVLA++:视觉-语言-动作模型中的记忆与想象时间建模
- OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics OmniGameArena:面向VLM游戏智能体的统一UE5基准与改进动力学
- Benchmark Agent: Autonomous AI System for Continuous Benchmark Construction Benchmark Agent: 自主构建基准测试的智能体系统
- Complexity-Balanced Splitting: Allocating Neural Capacity Along the Diffusion Timeline 复杂度均衡分割:沿扩散时间线分配神经网络容量
- DNQ: Teaching Agents to Play Nash Equilibrium in Multi-Player Bidding Games DNQ: 教智能体在多人竞价博弈中打出纳什均衡
- HANDOFF: Teaching Humanoid Robots To Follow Simple Commands Through Expert Distillation HANDOFF:通过专家蒸馏让人形机器人听懂简单指令
- How Abundant Are Good Interpolators? A Large Deviation Analysis of Overparametrized Learning 好的插值器有多普遍?过参数化学习的大偏差分析
- Active Exploration Closes the Conjunctive Reasoning Gap Between Humans and LLMs 主动探索缩小人类与大语言模型在合取推理上的差距
- MLEvolve: Self-Evolving LLM Agents for Automated Machine Learning Discovery MLEvolve: 自演化大模型智能体的机器学习算法自动发现
- OpAI-Bench: Progressive Human-AI Co-Editing Reveals Non-Monotonic Detection Patterns OpAI-Bench:渐进式人机协同编辑揭示非单调检测模式
- PAR3D: Teaching 3D Vision Models to See Parts, Not Just Objects PAR3D:让3D视觉模型看到零件,而非只看到物体
- PC Layer: Stabilizing LLM Training Through Polynomial Weight Preconditioning PC层:通过多项式权重预处理稳定大语言模型训练
- Learning to Play Against Adaptive Opponents: Beyond External Regret in Repeated Games 对抗自适应对手的学习:重复博弈中超越外部遗憾
- RREDCoT: Credit Assignment for Chain-of-Thought Reasoning Without Extra Generation RREDCoT:无需额外生成的思维链功劳分配
- TailLoR: Steering Model Updates Away from Dominant Directions in Continual Learning TailLoR: 在持续学习中将模型更新导向长尾谱空间
- TempoVLA: Teaching Robots to Drive Fast and Brake Hard TempoVLA:让机器人学会快行与慢动
- You Only Index Once: Sharing Routing Across Layers for Efficient Long-Context LLMs 只索引一次:跨层共享路由实现高效长文本推理
- Code2LoRA: Zero-Token Repository Context via Hypernetwork-Generated Adapters Code2LoRA:通过超网络生成适配器实现零token代码库上下文
- Goedel-Architect: Blueprint-Driven Formal Theorem Proving Goedel-Architect: 蓝图驱动的形式化定理证明
- Pretraining RNNs Without Unrolling: Supervised Memory Training 无需展开的 RNN 预训练:监督式记忆训练
- Self-Augmenting Retrieval for Diffusion Language Models 扩散语言模型的自增强检索
- Astra: Teaching Vision Models to Think by Imagining What They Haven't Seen Astra:让视觉模型通过想象未见之物来思考
- Audio-Interaction: From Offline Audio Models to Always-On Streaming Perception Audio-Interaction:从离线音频模型到永不停机的流式感知
- Failed Reasoning Traces Tell You What Is Fixable (But Not by Reading Them) 失败推理轨迹告诉你什么可修复(但不靠阅读它们)
- DistIL: Learning from Rich Feedback via Distributional Imitation DistIL:通过分布式模仿从丰富反馈中学习
- Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data 自我评估能力已经存在:用最少数据唤醒基础大模型中的潜在评判校准
- StreamMA: Pipelining Multi-Agent Reasoning by Streaming Partial Results StreamMA:通过流式传输部分结果实现多智能体推理流水线化
- Agentic Chain-of-Thought Steering: Budget-Aware Control for Efficient LLM Reasoning 智能体思维链引导:预算感知的高效语言模型推理控制
- Formalizing the Binding Problem: Measuring How Vision Models Link Features to Objects 形式化绑定问题:测量视觉模型如何将特征关联到对象
- Humanoid-GPT: Scaling Transformer Architecture to Billion-Frame Motion Control Humanoid-GPT:将 Transformer 架构扩展到十亿帧级动作控制
- Imaginative Perception Tokens: Teaching Vision Models to Hallucinate Correctly 想象感知标记:教视觉模型正确地幻想
- Language Models Need Sleep: Continual Learning Through Memory Consolidation and Dreaming 语言模型需要睡眠:通过记忆巩固和做梦实现持续学习
- NewtPhys: Testing Whether Vision Models Actually Understand Physics NewtPhys:视觉模型真的懂物理吗?
- When Reasoning Models Bluff: Quantifying Faithful Confidence in LLMs 当推理模型虚张声势:量化大语言模型的真实置信度
- SimuScene: Physics-in-the-Loop 3D Scene Reconstruction from Single Images SimuScene:单图物理闭环三维场景重建
- Skill-RM: Unifying Reward Models Through Agent-Based Skill Execution Skill-RM:用智能体技能统一奖励模型的评估标准
- AdaCodec: Predictive Visual Coding for Efficient Video MLLMs AdaCodec: 视频多模态大模型的预测式视觉编码
- AFUN: Teaching Robots Where and How to Touch AFUN:教机器人在哪里、怎样触碰
- ClinEnv: Testing AI Doctors Through Real Hospital Admissions, Not Multiple Choice ClinEnv:用真实住院病例而非选择题测试 AI 医生
- SubFit: Breaking the Layer Boundary in LLM Compression SubFit:打破大模型压缩的层级边界
- SPAWN: Training-Free Concept Insertion in Autoregressive World Models SPAWN:无需训练的世界模型概念注入技术
- HERO'S JOURNEY: Testing Rule Induction Through Text-Based Goal Execution HERO'S JOURNEY: 通过文本游戏测试复杂规则归纳能力
- HumanNOVA: Fast 3D Human Avatar Generation from Single Images HumanNOVA:从单张图片快速生成3D人体化身
- LongLive-RAG: Treating Video History as Searchable Memory to Fix Autoregressive Drift LongLive-RAG:将视频历史当作可搜索记忆来修复自回归漂移
- Perceptual Judgment Bias: Teaching Multimodal LLM Judges to Trust Their Eyes Over Plausible Text 感知判断偏差:教多模态大模型评判员相信眼睛而非文字
- Modeling Depth Ambiguity: A Mixture-Density Representation for Flying-Point-Free Depth Estimation 建模深度歧义:消除飞点的混合密度深度估计
- Verifiable Neural Safety Filters for Human-Robot Interaction Through Conformal Prediction 基于保形预测的可验证神经安全滤波器:让机器人在人机交互中安全学习
- RoboDream: Scaling Robot Learning Through Embodiment-Aware World Models RoboDream:通过具身感知世界模型扩展机器人学习
- Thinking in Blender: From Pixels to Editable 3D Code with Vision-Language Models 用 Blender 思考:视觉-语言模型从像素到可编辑 3D 代码
- Evidence-Augmented Machine Learning for Self-Harm Detection in Emergency Department Triage Notes 证据增强型机器学习:从急诊分诊记录中检测自残行为
- VLMs as Teachers: Adaptive Test-Time Optimization for Video Reasoning 视觉语言模型作为教师:视频推理的自适应测试时优化
- CoFiDA-M: Baking Clinical Reasoning into Image-Only Skin Cancer Screening CoFiDA-M: 将临床推理烘焙进纯图像皮肤癌筛查模型
- CHARM: Teaching Time-Series Models to Speak Through Channel Descriptions CHARM:让时序模型通过通道描述开口说话
- Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions 语言模型学会构式语义,更不用说句法:配对焦点构式理解的实证研究
- C4G: Compact Gaussian Queries for Globally Coherent 4D Scene Reconstruction C4G:用紧凑高斯查询实现全局一致的4D场景重建
- StateKV: Linear-Time Video Understanding Through Recurrent State Compression StateKV:通过循环状态压缩实现线性时间视频理解
- LongTraceRL: Teaching LLMs Long-Context Reasoning Through Search Trajectories and Fine-Grained Rewards LongTraceRL:通过搜索轨迹和细粒度奖励教会大模型长文本推理
- Lumos-Nexus: Training Cheap, Inferring Expensive for High-Fidelity Video Generation Lumos-Nexus:训练用轻量模型,推理换高清生成器
- nuReasoning: Teaching Self-Driving Cars to Think Through Edge Cases nuReasoning:教自动驾驶汽车思考边缘场景
- GRW: Teaching Machines to Read Gestures That Actually Mean Something GRW:教机器识别真正有意义的手势
- Representation Forcing: Eliminating VAE Bottlenecks in Unified Multimodal Models 表征强制:消除统一多模态模型中的 VAE 瓶颈
- SOCO: Testing Whether Vision Models Really Understand Object Parts SOCO:测试视觉模型是否真正理解物体部件
- Stateful Monitoring Catches Distributed Agent Attacks Across User Accounts 有状态监控:捕获跨账户分布式智能体攻击
- SurGe: Fixing the Bumpy Surfaces in Single-Image 3D Reconstruction SurGe:修复单图3D重建中的粗糙表面
- TunerDiT: Training-Free Progressive Steering for Multi-Event Video Generation TunerDiT:无需训练的多事件视频生成渐进式引导
- What Gets Unmasked First? Trajectory Analysis of Diffusion Models for Graph-to-Text Generation 先解码什么?图到文本生成中扩散模型的轨迹分析
- AdaState: Self-Evolving Anchors for Streaming Video Generation AdaState:流式视频生成中的自演化锚点
- FlatSounds: Testing Whether Video-to-Audio Models Understand Physics or Just Fake It FlatSounds:测试视频生成音频模型是真懂物理还是装懂
- Demystifying Data Organization for Enhanced LLM Training 揭秘数据组织:提升大语言模型训练效率的新视角
- FedTSV: Fair Federated Learning Through Trajectory Shapley Value FedTSV:基于轨迹 Shapley 值的公平联邦学习
- GMOS: Grounding Moving Object Segmentation in 3D Space and Time GMOS:在三维空间和时间中定位运动物体分割
- Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents 局部连贯,全局失调:多组件LLM智能体的组合不一致性边界
- NeuROK: Learning Physics Through Neural Object Kinematics NeuROK:通过神经物体运动学学习物理
- REST3D: From Floating Objects to Stable Scenes — Physics-Aware 3D Reconstruction REST3D:从悬浮物体到稳定场景——物理感知的3D重建
- Tiny but Trusted: Teaching Small Vision Models to Explain Time-Series Anomalies 小而可信:教小型视觉模型解释时序异常
- GAVIS: Uncertainty-Aware 3D Gaussian Splatting with Anisotropic Visibility for Active Mapping GAVIS:基于各向异性可见性场的不确定性感知3D高斯溅射主动建图
- COMPOSE: Predicting Future Math by Combining Citation Networks and Formal Proofs COMPOSE: 结合引用网络与形式化证明预测未来数学定理
- DynaFLIP: Teaching Robots to See Motion, Not Just Scenes DynaFLIP:教机器人看懂运动,而非仅识别场景
- HullFT: Geometric Test-Time Finetuning via Convex Hulls and Gradient Reuse HullFT:基于凸包和梯度复用的几何测试时微调
- SchGen: Teaching LLMs to Design Circuit Boards Through Semantic Code SchGen:用语义代码教大模型设计电路板
- VideoMLA: Compressing Video Diffusion Memory by 92.7% with Low-Rank Latent Attention VideoMLA:用低秩潜在注意力将视频扩散内存压缩92.7%
- GPIC: A Giant Permissive Image Corpus for Visual Generation GPIC:面向视觉生成的巨型开放图像语料库
- LLMSurgeon: Reverse-Engineering the Training Diet of Black-Box Language Models LLMSurgeon:逆向工程黑盒语言模型的训练配方
- When AI Agents Fake Understanding: A Physicist's 12-Day Reality Check 当 AI 智能体假装理解:一位物理学家的 12 天现实检验
- Reasoning in Memory: Teaching LLMs to Think Without Speaking 记忆推理:教大模型不说话也能思考
- YoCausal: Testing Whether Video Models Understand Causality or Just Memorize Time YoCausal:视频模型是理解因果还是只会记忆时间?
- Beyond Binary: Physics-Grounded Contact Representation for Sim-to-Real Dexterous Manipulation 超越二值:基于物理原理的接触表征实现仿真到真实的灵巧操作
- Gradient Probes: Finding Hidden Biases in Vision Models Without Labels 梯度探针:无需标注即可发现视觉模型中的隐藏偏见
- Calibrating Conservatism for Scalable Oversight 校准保守性:可扩展监督的新方法
- CaMBRAIN: Streaming EEG Analysis with Causal State Space Models CaMBRAIN:基于因果状态空间模型的流式脑电信号分析
- NEO-ov: Native One-Vision Models Without Modular Stitching NEO-ov:无需模块拼接的原生统一视觉模型
- Gamma-World: Scaling World Models to Multi-Agent Interactive Simulation Gamma-World:将世界模型扩展到多智能体交互模拟
- OmniVerifier-M1: Teaching Multimodal Models to Show Their Work with Symbolic Verification OmniVerifier-M1:用符号验证教多模态模型「展示推理过程」
- PEFT-Arena: Rethinking Parameter-Efficient Finetuning Through the Stability-Plasticity Lens PEFT-Arena:从稳定性-可塑性视角重新审视参数高效微调
- Personal Visual Memory from Explicit and Implicit Evidence 从显性与隐性证据构建个人视觉记忆
- Bidirectional Evolutionary Search: Breaking the Autoregressive Bottleneck in Language Model Self-Improvement 双向进化搜索:打破语言模型自我改进中的自回归瓶颈
- Alignment Tampering: When AI Systems Game Their Own Training Process 对齐篡改:AI 系统如何操纵自己的训练过程
- FinHarness: Inline Safety Monitoring for Finance LLM Agents FinHarness:金融 LLM 智能体的内联安全监控
- LocateAnything: Parallel Box Decoding for Fast Vision-Language Grounding LocateAnything:用并行框解码实现快速视觉-语言定位
- MATCHA: Teaching Metrics to Spot Contradictions, Not Just Overlap MATCHA:教会评估指标识别矛盾,而非仅看重叠
- MobileMoE: Bringing Mixture-of-Experts to Your Smartphone MobileMoE:把混合专家模型装进手机
- Channel-wise Vector Quantization: Tokenizing Images by Detail Levels Instead of Spatial Patches 通道级向量量化:按细节层次而非空间块对图像编码
- From Model Scaling to System Scaling: Why Agent Architecture Matters More Than Model Size 从模型扩展到系统扩展:为什么智能体架构比模型规模更重要
- Helix4D: Inheriting 3D Quality for Complex Dynamic Mesh Generation Helix4D:继承3D质量生成复杂动态网格
- Language Models Need Sleep: Consolidating Context Through Offline Recurrence 语言模型需要睡眠:通过离线循环巩固上下文
- LoopMDM: Selective Layer Looping Accelerates Masked Diffusion Language Models LoopMDM:选择性层循环加速掩码扩散语言模型
- Complete-μE: Tune Dense Once, Transfer to All MoE Configurations Complete-μE:调一次稠密模型,迁移到所有 MoE 配置
- ETCHR: Teaching Image Editors to Think Like Reasoners ETCHR:让图像编辑器学会推理式思考
- From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills 从原始经验到技能消费:模型生成智能体技能的系统性研究
- GenRecon: Scaling Generative 3D Priors to Multi-View Scene Reconstruction GenRecon:将生成式 3D 先验扩展到多视角场景重建
- HorizonStream: Taming Temporal Chaos in Streaming 3D Reconstruction HorizonStream:驯服流式3D重建中的时间混乱
- LaMo: Learning Motion Priors from Unlabeled Video for Physically Realistic Generation LaMo:从无标注视频中学习运动先验以生成物理真实的视频
- LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws 大语言模型作为噪声信道:从香农定理看模型容量与扩展定律
- PiD: Rethinking Latent Decoding as Pixel Diffusion for Fast High-Resolution Generation PiD:将潜在解码重构为像素扩散,实现快速高分辨率生成
- SkillOpt: Training Agent Skills Like Neural Network Weights SkillOpt:像训练神经网络权重一样训练智能体技能
- Training-Free Looped Transformers: Retrofitting Recurrence at Test Time 免训练循环Transformer:测试时改装递归结构
- AwareVLN: Teaching Navigation Agents to Know Where They Are and What They've Done AwareVLN: 让导航智能体知道自己在哪、做了什么
- Finite-Particle Convergence Rates for Conservative and Non-Conservative Drifting Models 保守与非保守漂移模型的有限粒子收敛速率
- Integrable Elasticity via Neural Demand Potentials 通过神经需求势函数实现可积弹性
- Remember to be Curious: Persistent 3D Memory Solves Exploration Amnesia 记得保持好奇:持久3D记忆解决探索失忆症
- SDPM: Diffusion Models Meet Survival Analysis Without Time Discretization SDPM:扩散模型遇上生存分析,无需时间离散化
- Cambrian-P: Teaching Video Models to See Through the Camera's Eyes Cambrian-P: 让视频模型学会用相机的眼睛看世界
- AI Chatbots as News Intermediaries: High Accuracy Masks Regional Inequity and Retrieval Fragility AI 聊天机器人作为新闻中介:高准确率掩盖的地区不平等与检索脆弱性
- FAME: Pinpointing Which Log Line Failed Using Mixture-of-Experts FAME:用混合专家模型精准定位故障日志行
- GesVLA: Teaching Robots to Understand Pointing Gestures GesVLA: 让机器人理解手势指向
- LCGuard: Protecting Privacy in Multi-Agent LLM Systems Through Adversarial KV Cache Transformation LCGuard:通过对抗性 KV 缓存变换保护多智能体 LLM 系统的隐私
- MotiMotion: Teaching Video Models to Reason About Physics Before They Animate MotiMotion: 让视频生成模型先推理物理再动画
- Sensor2Sensor: Converting Dashcam Videos into Multi-Modal AV Sensor Data Sensor2Sensor:将行车记录仪视频转换为多模态自动驾驶传感器数据
- The Matching Principle: Unifying Robustness Through Deployment Nuisance Geometry 匹配原理:通过部署干扰几何统一鲁棒性
- ConvexTok: Optimal Tokenization Through Linear Programming ConvexTok:通过线性规划实现最优分词
- Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs 它往哪边动了?诊断并克服视频大语言模型的方向运动盲区
- DecQ: Bridging Reconstruction and Generation in Vision Autoencoders with Detail Queries DecQ:用细节查询弥合视觉自编码器的重建与生成鸿沟
- DeltaBox: Millisecond-Level Checkpoint/Rollback for AI Agent State Exploration DeltaBox:为 AI 智能体状态探索提供毫秒级检查点/回滚
- Gated DeltaNet-2: Decoupling Memory Erasure from Content Writing in Linear Attention Gated DeltaNet-2:线性注意力中解耦记忆擦除与内容写入
- MOSS: Teaching Agents to Rewrite Their Own Source Code MOSS:让智能体改写自己的源代码
- Vector Policy Optimization: Training LLMs for Diversity to Unlock Test-Time Search 向量策略优化:训练多样性以解锁测试时搜索
- Agent JIT Compilation: Compiling Web Automation Tasks into Parallel Code 智能体即时编译:将网页自动化任务编译为并行代码
- DelTA: Teaching Language Models Which Tokens Actually Matter for Reasoning DelTA:教会语言模型哪些词元真正影响推理
- Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning 平衡推理器:学习吸引子实现可扩展推理
- Fixed-Point Distillation: One-Step Discrete Diffusion via Iterative Correction 不动点蒸馏:通过迭代修正实现单步离散扩散
- RELEX: Extrapolating RLVR Training via Rank-1 Trajectories RELEX:通过秩-1 轨迹外推强化学习训练
- The Stochastic-Deterministic Boundary: A Design Framework for Production LLM Agents 随机-确定性边界:生产级 LLM 智能体的设计框架
- CaMo: Teaching Vision-Language Models to Understand Camera Movement CaMo:教会视觉语言模型理解相机运动
- ClinSeekAgent: From Passive Evidence Consumption to Active Clinical Evidence Hunting ClinSeekAgent:从被动接收证据到主动搜寻临床证据
- Staged Training for Vision-Language Models: Separating Perception from Reasoning 视觉语言模型的分阶段训练:将感知与推理解耦
- KoRe: Compressing Knowledge Graphs into Discrete Tokens for LLMs KoRe:将知识图谱压缩为离散token注入大语言模型
- Policy-Aware Rubric Rewards: Teaching AI What Matters Now, Not Just What Matters Most 策略感知的评分奖励:教 AI 当下有用的,而非仅仅重要的
- PixVerve: Scaling Text-to-Image Generation to 100 Megapixels PixVerve:将文生图扩展到 1 亿像素
- TIDE: Lossless MoE Diffusion LLM Inference via I/O-Aware Expert Offloading TIDE:基于 I/O 感知专家卸载的无损 MoE 扩散 LLM 推理
- TideGS: Training Billion-Scale 3D Gaussian Splatting via Out-of-Core Optimization TideGS:通过外存优化训练十亿级3D高斯溅射
- When Does Model Collapse Occur in Structured Interactive Learning? 结构化交互学习中的模型坍塌何时发生?
- WorldString: Learning Actionable Object Representations for Physical World Models WorldString:为物理世界模型学习可操作的物体表征
- IAMFlow: Training-Free Identity-Aware Memory for Consistent Long Video Generation IAMFlow:无需训练的身份感知记忆框架,实现一致性长视频生成
- Aurora: Bridging User Intent and Video Editing Models with Tool-Using Agents Aurora:用工具增强智能体连接用户意图与视频编辑模型
- Code as Agent Harness: When Programs Become the Infrastructure for AI Agents 代码作为智能体线束:当程序成为 AI 智能体的基础设施
- DashAttention: Making Sparse Attention Fully Differentiable with Adaptive Block Selection DashAttention:通过自适应块选择实现全可微稀疏注意力
- ESI-Bench: Teaching AI to Explore Before It Reasons ESI-Bench:让AI学会先探索再推理
- General Preference Reinforcement Learning: Multi-Dimensional Quality for Open-Ended LLM Alignment 通用偏好强化学习:开放式大模型对齐的多维质量建模
- Predictable Confabulations: How Model Size and Data Frequency Govern LLM Factual Recall 可预测的幻觉:模型规模与数据频率如何支配大语言模型的事实召回
- Spectral Progressive Diffusion: Growing Resolution Along the Denoising Path 频谱渐进扩散:沿去噪路径逐步增长分辨率
- Vision-OPD: Teaching Multimodal LLMs to See Fine Details by Learning from Their Own Zoomed-In Views Vision-OPD:让多模态大模型通过学习自己的放大视角来看清细节
- Argus: Assembling Evidence Jigsaws for Deep Research with Cooperative Agents Argus:用协作智能体拼装证据拼图的深度研究系统
- When More Reasoning Hurts: Why Programmatic State Beats LLM Deliberation in Adversarial Environments 当推理越多越糟:为什么在对抗环境中程序化状态优于LLM深思
- Fully Open Meditron: Building Auditable Clinical AI from the Ground Up 全开放 Meditron:从零构建可审计的临床 AI
- Autonomous AI Systems Generate Expert-Level Disease Forecasts 自主AI系统生成专家级疾病预测
- Predicting Magnetic Structures from Atoms Alone: A Neural Network Trained on Real Experiments 仅从原子坐标预测磁结构:在真实实验数据上训练的神经网络
- Grep vs. Vector Search: How Agent Harnesses Shape Retrieval Performance Grep 对比向量检索:智能体框架如何重塑检索性能
- RefDecoder: Fixing Video Generation's Decoder Blindness Problem RefDecoder:修复视频生成中解码器的「失明」问题
- Text Knows What, Tables Know When: Bridging Clinical Narratives and EHR Data for Precise Timeline Reconstruction 文本知道什么,表格知道何时:融合临床叙事与电子病历数据的精确时间线重建
- VGGT-Ω: Scaling Feed-Forward 3D Reconstruction with Efficient Architecture and Self-Supervised Learning VGGT-Ω:用高效架构和自监督学习扩展前馈式三维重建
- Warp-as-History: Zero-Shot Camera Control for Video Generation via Warped Pseudo-History Warp-as-History:通过扭曲伪历史实现视频生成的零样本相机控制
- Spherical Flow Matching: Aligning Latent Geometry for Better Image Generation 球面流匹配:对齐潜在几何以改进图像生成
- Articraft: Programming Articulated 3D Assets at Scale with LLM Agents Articraft:用大语言模型智能体规模化生成可动3D资产
- ATLAS: Bridging Agentic and Latent Visual Reasoning with Functional Tokens ATLAS:用功能性词元连接智能体推理与隐式视觉推理
- EntityBench: A Benchmark for Character-Consistent Long Video Generation EntityBench:长视频生成中的实体一致性基准
- Shodh-MoE: Solving Multi-Physics Interference with Sparse Expert Routing Shodh-MoE:用稀疏专家路由解决多物理干扰
- EviScreen: Evidence-Based Disease Screening with Historical Case Retrieval EviScreen: 基于历史病例检索的证据推理疾病筛查
- From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing 从规划到像素:学习规划与编排以实现开放式图像编辑
- FutureSim: Testing AI Agents on Real-World Event Prediction FutureSim:用真实世界事件测试AI智能体
- MetaBackdoor: Hijacking LLMs Through Position, Not Content MetaBackdoor:通过位置而非内容劫持大语言模型
- OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation OpenDeepThink:通过 Bradley-Terry 聚合实现并行推理
- PDI-Bench: Quantifying Geometric Coherence in Video Generation Models PDI-Bench:量化视频生成模型的几何一致性
- RAVEN: Bridging Training-Inference Distribution Gap in Autoregressive Video Generation RAVEN:弥合自回归视频生成中的训练-推理分布鸿沟
- SANA-WM: Minute-Scale World Modeling with 36× Faster Inference SANA-WM:分钟级世界模型,推理速度提升36倍
- VGGT-Edit: Native 3D Scene Editing via Feed-Forward Residual Field Prediction VGGT-Edit:通过前馈残差场预测实现原生3D场景编辑
- Tensor Similarity: A Weight-Based Metric for Comparing Neural Network Mechanisms 张量相似度:基于权重的神经网络机制比较度量
- EVA-Bench: Testing Voice Agents End-to-End with Simulated Conversations EVA-Bench:用模拟对话端到端测试语音代理
- TFlow: Multi-Agent LLMs Communicate Through Weight Perturbations Instead of Text TFlow:多智能体大模型通过权重扰动而非文本通信
- AEvo: Meta-Editing the Evolution Process Itself AEvo:元编辑进化过程本身
- History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions 历史锚点:先前行为如何将大语言模型决策引向不安全行动
- Negation Neglect: Why Fine-tuning Makes Models Believe What They're Told Is False 否定忽视:为什么微调让模型相信它被告知是假的东西
- VERIMED: Catching Ambiguity in Safety Requirements Before Code Ships VERIMED:在代码发布前捕获安全需求中的歧义
- OmniLiDAR: Unified Multi-Domain LiDAR Generation via Text-Conditioned Diffusion OmniLiDAR:基于文本条件扩散的统一多域激光雷达生成
- Parallel Scan Recurrent Neural Quantum States: Making RNNs Scalable for Quantum Many-Body Problems 并行扫描循环神经量子态:让RNN在量子多体问题上可扩展
- Fast Vector Quantization with Provable Guarantees via Randomized Hadamard Transform 随机 Hadamard 变换实现快速向量量化的可证明保证
- QLAM: Using Quantum Superposition to Remember Long Sequences QLAM:用量子叠加态记忆长序列
- R-DMesh: Solving the Pose Misalignment Problem in Video-Guided 3D Animation R-DMesh:解决视频引导3D动画中的姿态错位问题
- Cross-Sample Prediction Churn: The Hidden Instability in Scientific ML 跨样本预测波动:科学机器学习中的隐藏不稳定性
- MMProLong: Training Vision-Language Models for 128K+ Context with Balanced Data MMProLong:用平衡数据训练 128K+ 上下文的视觉语言模型
- Unlocking Patch-Level Features for CLIP-Based Class-Incremental Learning 解锁 CLIP 类增量学习中的局部特征
- WARDEN: Transcribing Endangered Languages with Just 6 Hours of Data WARDEN:用 6 小时数据转录濒危语言
- AlphaGRPO: Teaching Multimodal Models to Think Before They Generate AlphaGRPO:让多模态模型学会「先思考再生成」
- Sparse-to-Dense Reward Allocation: Why Your Best Data Should Train Your Best Model First 稀疏到密集的奖励分配原则:为什么最好的数据应该先训练最强的模型
- CausalCine: Real-Time Multi-Shot Video Generation Through Interactive Directing CausalCine:通过交互式导演实现实时多镜头视频生成
- Covering the Long Tail: Synthetic Data for Complex Computer-Use Actions 覆盖长尾:为复杂计算机操作合成训练数据
- EgoForce: Absolute 3D Hand Pose from Egocentric Monocular Video via Forearm Guidance EgoForce: 通过前臂引导从单目第一视角视频重建绝对3D手部姿态
- Elastic Attention Cores for Scalable Vision Transformers 弹性注意力核心:可扩展视觉Transformer的线性复杂度方案
- WebEye: When Visual Perception Needs to Google First WebEye:视觉感知需要先搜索的时代
- KV-Fold: Treating KV Cache as a Fold Accumulator for Long-Context Inference KV-Fold: 将 KV 缓存视为折叠累加器的长文本推理
- Fast-Slow Learning: Teaching LLMs to Adapt Without Forgetting 快慢学习:让大语言模型在适应中保持记忆
- LongMemEval-V2: Teaching Agents to Remember Like Experienced Colleagues LongMemEval-V2:让智能体像资深同事一样记忆
- MEME: When LLM Agents Forget What They Just Learned MEME:当 LLM 智能体忘记刚学到的东西
- OmniNFT: Teaching Diffusion Models to Generate Synchronized Audio-Video with Modality-Aware Reinforcement Learning OmniNFT: 用模态感知强化学习教扩散模型生成同步音视频
- Pion: Training Neural Networks by Rotating Weight Matrices Instead of Adding to Them Pion:通过旋转权重矩阵而非累加来训练神经网络
- AmbiSuR: Resolving Photometric Ambiguity in Gaussian Splatting for Robust Surface Reconstruction AmbiSuR:解决高斯溅射中的光度歧义以实现鲁棒表面重建
- Reward Hacking in Rubric-Based Reinforcement Learning 基于评分标准的强化学习中的奖励欺骗
- Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts 路由器学习专家的几何结构:稀疏混合专家模型中的几何耦合
- SenseNova-U1: Breaking the Understanding-Generation Divide in Vision-Language Models SenseNova-U1:打破视觉语言模型的理解-生成鸿沟
- Attractor Models: Making Iterative Refinement Scalable Through Fixed-Point Learning 吸引子模型:通过不动点学习让迭代优化变得可扩展
- Task-Adaptive Embedding Refinement via Test-time LLM Guidance 测试时 LLM 引导的任务自适应嵌入优化
- ToolCUA: Teaching Agents When to Click and When to Call APIs ToolCUA: 教会智能体何时点击、何时调用工具
- StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction StraTA:通过战略轨迹抽象激励智能体强化学习
- AI Co-Mathematician: Accelerating Mathematicians with Agentic AI AI协同数学家:用智能体AI加速数学家工作
- Verifier-Backed Hard Problem Generation for Mathematical Reasoning 基于验证器的数学推理难题生成
- EMO: Pretraining Mixture of Experts for Emergent Modularity EMO:预训练混合专家模型以实现涌现模块化
- UniPool: A Globally Shared Expert Pool for Mixture-of-Experts UniPool:混合专家模型的全局共享专家池
- AI Co-Mathematician: From Proof Assistants to Research Partners AI 协同数学家:从证明助手到研究伙伴
- EMO: Training Modular Language Models Through Document-Level Expert Pooling EMO:通过文档级专家池训练模块化语言模型
- Recursive Agent Optimization: Teaching AI to Divide and Conquer Through Self-Delegation 递归智能体优化:教会 AI 通过自我委派实现分而治之
- StraTA: Teaching AI Agents to Think Before They Act with Strategic Planning StraTA:让 AI 智能体先想后做的战略规划框架
- UniPool: Breaking the Per-Layer Expert Ownership Rule in Mixture-of-Experts UniPool:打破混合专家模型的逐层专家所有权规则
- Design Conductor 2.0: AI Agent Builds Hardware Accelerator in 80 Hours Design Conductor 2.0:AI 智能体 80 小时造出硬件加速器
- Beyond Sampling: Analytical Estimation of Neural Network Outputs at Initialization 超越采样:神经网络初始化输出的解析估计
- Grokability in Five Inequalities: When AI Discovers New Math 五个不等式中的可理解性:当 AI 发现新数学
- Implicit Representations of Grammaticality in Language Models 语言模型中的隐式语法表征
- LongSeeker: Teaching Agents to Forget What Doesn't Matter LongSeeker:教会智能体遗忘无关信息
- OpenSearch-VL: Opening the Black Box of Multimodal Search Agents OpenSearch-VL:打开多模态搜索智能体的黑箱
- PhysForge: From Static Meshes to Interactive Physics-Grounded 3D Assets PhysForge:从静态网格到可交互的物理基础3D资产
- Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval 线性联想记忆的锐容量阈值:从赢者通吃到列表检索
- First-Token Confidence: A Zero-Cost Hallucination Detector 首词置信度:零成本的幻觉检测器
- Attention as Featurizer: How Transformers Learn Nonlinear Regression In-Context 注意力作为特征器:Transformer 如何在上下文中学习非线性回归
- Experience-RAG Skill: Teaching Agents to Choose the Right Retrieval Strategy Experience-RAG Skill:让智能体学会选择正确的检索策略
- Audio-Visual Intelligence: A Survey of Foundation Models Bridging Sound and Sight 视听智能:连接声音与视觉的基础模型综述
- EQUITRIAGE: When AI Triage Systems Inherit Gender Bias from Clinical Practice EQUITRIAGE:当AI分诊系统继承临床实践中的性别偏见
- UniReasoner: Closing the Understanding-Generation Gap in Text-to-Image Models UniReasoner:弥合文生图模型的理解-生成鸿沟
- OpenSeeker-v2: How an Academic Team Beat Industry Giants at Search Agents Using Just 10k Examples OpenSeeker-v2:学术团队如何用1万条数据在搜索智能体上击败工业巨头
- Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours 智能体时代的 AI 红队重构:从数周到数小时
- Rethinking Reasoning-Intensive Retrieval: From Topical Matching to Evidence Portfolio Construction 重新思考推理密集型检索:从主题匹配到证据组合构建
- Clinical LLM Safety Doesn't Scale Like Accuracy: Evidence Quality Matters More Than Model Size 临床大模型的安全性不随准确率扩展:证据质量比模型规模更重要
- SymptomAI: Why Structured Interviews Beat Chat for Medical Diagnosis SymptomAI:为什么结构化问诊比闲聊更适合医疗诊断
- UniCorrn: One Model for All Visual Correspondence Tasks UniCorrn:统一所有视觉对应任务的单一模型
- ActDiff-VC: Compressing Video to Almost Nothing by Teaching Diffusion Models to Hallucinate Wisely ActDiff-VC:通过教扩散模型「明智幻觉」将视频压缩到极致
- JACTUS: Unifying Model Compression and Adaptation in a Single Subspace JACTUS:在统一子空间中同时完成模型压缩与适配
- Edge-Efficient Image Restoration: Distilling Transformers into State-Space Models 边缘高效图像修复:将 Transformer 蒸馏到状态空间模型
- FlexSQL: Breaking the Fixed Pipeline in Text-to-SQL with On-Demand Database Interaction FlexSQL:用按需数据库交互打破文本转SQL的固定流水线
- HAAS: Governance as a Design Variable in Human-AI Task Allocation HAAS:将治理作为人机任务分配的设计变量
- Reinforcement Learning for Multi-Agent LLM Systems: The Orchestration Trace Framework 多智能体大语言模型系统的强化学习:编排轨迹框架
- SpecKV: Adaptive Speculation Length for Compressed LLMs SpecKV:压缩模型的自适应推测长度
- Teaching Compact Models to Think: Knowledge Distillation for Cross-Language Clone Detection 教小模型学会推理:跨语言代码克隆检测的知识蒸馏
- Trust, but Verify: Peeling Low-Bit Transformer Networks for Training Monitoring 信任但验证:剥离低比特Transformer网络以监控训练
- VideoNet: Reviving Action Recognition with Domain-Specific Challenges VideoNet:用领域专属动作重振动作识别
- Can AI Coding Agents Actually Do Science? Testing LLMs on Materials Science Reproducibility AI 编码智能体真的会做科学吗?在材料科学可重复性上测试大语言模型
- Local Attention as a Temporal Logic Operator: Why Windowed Transformers See More 局部注意力作为时序逻辑算子:为何窗口化 Transformer 反而看得更多
- Directed Social Regard: Measuring Who Gets Praised and Who Gets Blamed in the Same Message 定向社会态度:在同一消息中测量谁被赞扬、谁被指责
- Validation-Driven Chart Generation: Teaching LLMs to See What They Draw 验证驱动的图表生成:教会 LLM 看见自己画的东西
- GeoContra: Turning Fluent GIS Code into Verifiable Spatial Analysis GeoContra:让流畅的 GIS 代码变成可验证的空间分析
- HyCOP: Learning PDEs as Interpretable Programs Instead of Black-Box Operators HyCOP:将偏微分方程学习为可解释程序而非黑盒算子
- GenLIP: Teaching Vision Transformers to Speak the Language of LLMs GenLIP:让视觉Transformer说大语言模型的语言
- LightKV: Prompt-Guided Vision Token Compression for Efficient LVLM Inference LightKV:提示引导的视觉token压缩,让多模态大模型推理更轻量
- Map2World: User-Controlled 3D World Generation from Arbitrary Segment Maps Map2World:从任意分割地图生成用户可控的3D世界
- Persistent Visual Memory: Keeping Vision Alive Through Long Text Generation 持久视觉记忆:在长文本生成中保持视觉感知
- Posterior Augmented Flow Matching: Training Generative Models with Multiple Target Trajectories 后验增强流匹配:用多目标轨迹训练生成模型
- RunAgent: Bridging Natural Language Plans and Deterministic Execution RunAgent:连接自然语言计划与确定性执行
- SAVGO: Steering Policy Updates Through Value Geometry in Continuous Control SAVGO:用价值几何引导连续控制中的策略更新
- When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models 当大模型不再遵循步骤:语言模型程序执行能力的诊断研究
- When Patient Chatbots Leak Everything: A Security Autopsy of Medical RAG Systems 当患者聊天机器人泄露一切:医疗 RAG 系统的安全解剖
- Adaptive Wavelets Tame Extreme Loss Imbalance in Physics-Informed Neural Networks 自适应小波驯服物理信息神经网络中的极端损失失衡
- Beyond Nash: Computing Equilibria That Minimize Coalition Deviation Incentives 超越纳什:计算最小化联盟偏离激励的均衡
- Continuous-Tone Simple Points: Making Topology Differentiable for Deep Learning Segmentation 连续色调简单点:让拓扑结构可微分以用于深度学习分割
- Quantum Autoencoders as Adversarial Purifiers: Training-Free Defense for Quantum Classifiers 量子自编码器作为对抗净化器:量子分类器的免训练防御
- Machine Learning Maps the Vicsek Flocking Model's Phase Space 用机器学习绘制 Vicsek 群集模型的相图
- Action Motifs: Learning Reusable Movement Patterns from Human Pose 动作基元:从人体姿态中学习可复用的运动模式
- AEGIS: When AI-Generated Academic Images Fool Even the Best Detectors AEGIS:当AI生成的学术图像连最好的检测器都能骗过
- Exploration Hacking: When AI Models Learn to Sabotage Their Own Training 探索劫持:AI 模型学会破坏自己的训练
- GenWildSplat: Feed-Forward 3D Reconstruction from Wild Internet Photos GenWildSplat:从互联网照片直接重建3D场景
- HERMES++: Bridging Scene Understanding and Future Prediction in Autonomous Driving HERMES++:自动驾驶中场景理解与未来预测的统一
- LaST-R1: Teaching Robots to Think Before They Act with Reinforcement Learning LaST-R1:让机器人在行动前先思考的强化学习方法
- LLMs as Graph Surgeons: Cleaning Noisy Brain Signal Networks for Seizure Detection 大模型做图外科医生:为癫痫检测清理脑电信号网络
- OmniRobotHome: Making Multi-Human Multi-Robot Collaboration Actually Work OmniRobotHome:让多人多机器人协作真正可行
- PhyCo: Teaching Video Models to Respect Physics Through Controllable Material Properties PhyCo:通过可控材料属性让视频模型学会尊重物理规律
- Representation Fréchet Loss for Visual Generation 表示空间中的 Fréchet 损失用于视觉生成
- Sequential Inference for Gaussian Processes: Bridging Signal Processing and Modern ML 高斯过程的序列推断:信号处理与现代机器学习的桥梁
- Stop Holding Your Breath: CT-Informed Gaussian Splatting for Dynamic Bronchoscopy 别再屏住呼吸:用CT引导的高斯散射重建动态支气管镜
- Strait: Priority-Aware ML Inference Serving Through Contention Modeling Strait:通过竞争建模实现优先级感知的机器学习推理服务
- Synthetic Computers at Scale: Training AI Agents in Simulated User Worlds 规模化合成计算机:在模拟用户世界中训练 AI 智能体
- From Pixel Painters to World Simulators: The Five Levels of Visual Generation 从像素画师到世界模拟器:视觉生成的五层进化
- AnimateAnyMesh++: Scaling 4D Mesh Animation with 300K Identities and Variable-Length Generation AnimateAnyMesh++:用30万身份和变长生成扩展4D网格动画
- Neural Assemblies Learn Causal Direction Without Backpropagation 神经集群无需反向传播即可学习因果方向
- ClassEval-Pro: Exposing the Class-Level Code Generation Gap ClassEval-Pro:揭示类级代码生成的能力断层
- ClawGym: A Complete Training Pipeline for File-Manipulating AI Agents ClawGym:文件操作型 AI 智能体的完整训练管线
- Color-Encoded Illumination for High-Speed Volumetric Scene Reconstruction 彩色编码照明:用普通相机捕捉高速3D场景
- FaaSMoE: Serverless Mixture-of-Experts for Multi-Tenant AI Serving FaaSMoE:多租户混合专家模型的无服务器架构
- GSCNet: Aligning Mismatched Thermal-RGB UAV Images with Semantic Graphs GSCNet:用语义图对齐无人机热红外-可见光错位图像
- HyCNNs: Exponentially More Efficient Convex Neural Networks HyCNNs:指数级更高效的凸神经网络
- Learning to Tune ADMM: Adaptive Relaxation with Convergence Guarantees 学习调优 ADMM:带收敛保证的自适应松弛策略
- SEAL: Teaching Diffusion Models to Personalize Stickers Without Memorizing Backgrounds SEAL:教会扩散模型个性化贴纸而不记住背景
- Select to Think: Teaching Small Models to Choose Wisely Instead of Guessing Blindly 选择性思考:教小模型明智选择而非盲目猜测
- When Transformers Meet Particle Physics: Proving Deep Learning Converges to Stochastic Flows 当 Transformer 遇见粒子物理:证明深度学习收敛到随机流
- Three-Step Nav: Teaching Vision-Language Robots to Plan Like Humans Navigate Three-Step Nav:让视觉语言机器人像人类一样规划导航
- TIDE: Cross-Architecture Distillation for Diffusion Language Models TIDE:扩散语言模型的跨架构蒸馏
- World2VLM: Teaching Vision Models to Imagine Motion by Distilling World Models World2VLM:通过蒸馏世界模型教会视觉模型想象运动
- The Paradox of AI Fluency: Why Expert Users Fail More (But Succeed Better) AI 熟练度悖论:为什么专家用户失败更多(但成功得更好)
- Carbon-Taxed Transformers: Compressing LLMs by Penalizing Computational Waste 碳税变压器:通过惩罚计算浪费来压缩大语言模型
- Conditional Misalignment: How Safety Interventions Create Context-Triggered Backdoors 条件性失配:安全干预如何制造上下文触发的后门
- DV-World: Testing Data Visualization Agents in Real Enterprise Workflows DV-World:在真实企业工作流中测试数据可视化智能体
- From Syntax to Emotion: How LLMs Internally Recognize Feelings 从句法到情感:大语言模型如何在内部识别情绪
- Escaping Cold Start: Training Reasoning Models with Tsallis Loss 逃离冷启动:用 Tsallis 损失训练推理模型
- Luminol-AIDetect: Detecting AI Text by How It Breaks Under Shuffling Luminol-AIDetect:通过打乱文本看AI如何露馅
- No Pedestrian Left Behind: Adaptive Traffic Signals That Watch and Wait 不让行人掉队:会观察、会等待的自适应交通信号
- QCalEval: Teaching Vision Models to Read Quantum Calibration Plots QCalEval:教视觉模型读懂量子校准图
- Recursive Multi-Agent Systems: Scaling Collaboration Through Latent-Space Loops 递归多智能体系统:通过潜空间循环扩展协作
- RESTestBench: When LLMs Generate API Tests from Requirements, How Do We Know They Actually Work? RESTestBench:当大模型从需求生成 API 测试时,我们如何知道它们真的有效?
- Robust Deepfake Detection via Calibrated Multi-Stream Ensembles 通过校准多流集成实现鲁棒深度伪造检测
- Three Models of RLHF Annotation: Extension, Evidence, and Authority RLHF 标注的三种模型:延伸、证据与权威
- Geometric Algebra for Language: Beyond Vectors to Multivectors 语言的几何代数:从向量到多重向量
- When Wrong Rewards Help: Why Imperfect Feedback Can Accelerate RL Training 错误奖励何时有益:不完美反馈如何加速强化学习训练
- GFlowState: Visualizing GFlowNet Training Dynamics GFlowState:可视化GFlowNet训练动态
- Temporal Taskification: Making Time a First-Class Citizen in RL 时间任务化:让时间成为RL中的第一等公民
- TraceScope: Interactive Phishing URL Triage TraceScope:交互式钓鱼URL分流
- Causality-Encoded Diffusion Models 因果编码扩散模型
- Transient Turn Injection Attack 瞬时轮次注入攻击
- EVENT5Ws: Large Open-Domain Event Extraction Benchmark EVENT5Ws:大规模开放域事件抽取基准
- Bounding the Black Box: AI Risk Regulation 为黑盒划界:AI风险监管
- Seeing Fast and Slow: Learning Temporal Flow in Videos 快慢兼察:学习视频中的时间流
- Replay-Buffer Engineering for Quantum Circuit Optimization 用于量子电路优化的重放缓冲区工程
- Machine Behavior in Moral Dilemmas 道德困境中的机器行为
- Addressing Image Authenticity with GenAI Cameras 通过生成式AI相机解决图像真实性问题
- RedirectQA: Non-Verbatim Memorization in LLMs RedirectQA:LLM中的非字面记忆
- MODEE: Multimodal Open-Domain Event Extraction MODEE:多模态开放域事件抽取
- Fine-Tuning Regimes: Different Fine-Tuning Methods Define Distinct CL Problems 微调机制:不同的微调方法定义不同的持续学习问题
- ASR Evaluation Using Generative LLMs: A New Paradigm 使用生成式LLM进行ASR评估:新的范式
- LayerTracer: Joint Task-Particle and Vulnerable-Layer Analysis for LLMs LayerTracer:LLM联合任务粒子和脆弱层分析框架
- SuperIgor: Self-Guided Plan Extraction for Instruction Following SuperIgor:用于指令跟随的自我引导计划提取
- LLMs Outperform Humans in Fraud Detection 大语言模型在欺诈检测中超越人类
- GRPO-VPS: Verifiable Process Supervision for Effective Reasoning GRPO-VPS:用于有效推理的可验证过程监督
- The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning 视觉的代价:获得可信赖的多模态推理
- ORPHEAS: Greek-English Embedding Model for RAG ORPHEAS:用于RAG的希腊语-英语嵌入模型
- TACO: Terminal Agent Compression via Self-Evolving Observation Rules TACO:通过自进化观察规则压缩终端智能体
- LLMs Impact on Peer Review: A Fine-Grained Analysis LLM对同行评审的影响:细粒度分析
- Cross-Model Consistency of AI-Generated Exercise Prescriptions AI生成运动处方跨模型一致性研究
- SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal LLMs SafetyALFRED:评估多模态LLM的安全意识规划
- Micro Language Models Enable Instant Responses 微型语言模型实现即时响应
- GRIL: Grounded Reasoning via Interactive Reinforcement Learning GRIL:通过交互式强化学习实现扎根推理
- Discovering Shared Logical Subspace for LLM Logical Reasoning 发现共享逻辑子空间以提升LLM逻辑推理
- Benign Overfitting in Adversarial Training for Vision Transformers 视觉Transformer对抗训练中的良性过拟合
- VLA Foundry: Unified Framework for Training Vision-Language-Action Models VLA Foundry:训练视觉-语言-动作模型的统一框架
- FB-NLL: Feature-Based Noisy Labels in Personalized Federated Learning FB-NLL:个性化联邦学习中的基于特征的噪声标签处理
- FASTER: Value-Guided Sampling for Fast Reinforcement Learning FASTER:用于快速强化学习的价值引导采样
- Safe Continual Reinforcement Learning in Non-stationary Environments 非平稳环境中的安全持续强化学习
- Phase Transitions in Fluctuations of Random Neural Networks 随机神经网络泛函波动的相变
- Generalization at the Edge of Stability 稳定性边缘的泛化
- Bi-CMPStereo: Bridging Event and Frame Cameras Through Bidirectional Cross-Modal Prompting Bi-CMPStereo:通过双向跨模态提示连接事件相机与帧相机
- Cloning is as Hard as Learning for Stabilizer States 稳定子态的克隆与学习同样困难
- Quantum-Inspired Node Embeddings: When Structure Beats Features in Graph Neural Networks 量子启发的节点嵌入:图神经网络中结构何时胜过特征
- SegWithU: Single-Pass Uncertainty Estimation for Medical Image Segmentation via Perturbation Energy SegWithU:通过扰动能量实现医学图像分割的单次前向不确定性估计
- ORCA: Making SVM Decisions Transparent Through Orthogonal Polynomial Kernels ORCA:用正交多项式核让支持向量机决策透明化
- AD4AD: Teaching Self-Driving Cars to Say 'I Don't Know' AD4AD:让自动驾驶汽车学会说「我不知道」
- AnimationBench: The First Benchmark That Actually Gets Animation AnimationBench:首个真正理解动画的评测基准
- Muon Over AdamW: A Systematic Benchmark of Optimizers for Tabular Deep Learning Muon 优于 AdamW:表格深度学习优化器的系统性基准测试
- When AI Judges Disagree With Themselves: Diagnosing LLM Evaluation Reliability 当AI评委自相矛盾:诊断大语言模型评估的可靠性
- Why Language Models Fail at Longer Problems: A Shortest Path Study 语言模型为何在更长问题上失效:最短路径研究
- GlobalSplat: Compact 3D Reconstruction via Global Scene Tokens GlobalSplat:通过全局场景令牌实现紧凑3D重建
- Can Language Models Navigate Space Without Seeing? Probing Viewpoint Rotation Understanding 语言模型能否在无视觉条件下理解空间?视角旋转理解的探针研究
- LeapAlign: Teaching Flow Models to Jump Through Time for Better Image Generation LeapAlign: 让流匹配模型学会时间跳跃以生成更好的图像
- MM-WebAgent: Hierarchical Planning Meets Multimodal Webpage Generation MM-WebAgent:层级规划驱动的多模态网页生成
- Prism: Symbolic Superoptimization of Tensor Programs Prism:张量程序的符号超优化器
- R3D: Fixing 3D Robot Learning with Better Architecture and Training R3D:用更好的架构和训练修复3D机器人学习
- RAD-2: Teaching Self-Driving Cars to Plan Like a Chess Player with a Coach RAD-2: 让自动驾驶像有教练的棋手一样规划路径
- Think in Latent Thoughts: Reasoning-Driven Sign Language Translation 隐式思维链:推理驱动的手语翻译新范式
- TokenLight: Controlling Image Lighting Like Adjusting Sliders on a Mixing Board TokenLight:像调音台推子一样精确控制图像光照
- Why Vision-Language Models Can't Read the Room: The Emotion Recognition Paradox 视觉语言模型为何读不懂表情:情感识别的悖论
- CRAFT: Building Better Reasoning Chains Through Consensus Knowledge Graphs CRAFT:通过共识知识图谱构建更优推理链
- From Vibes to Metrics: Formalizing How Users Really Evaluate LLMs 从感觉到指标:形式化用户真实评估大模型的方式
- PreRL: Teaching Language Models to Reason by Pruning Wrong Paths Before Fine-Tuning PreRL:在微调前通过剪枝错误路径教会语言模型推理
- Steering as Adaptation: Rethinking How We Change Language Model Behavior 转向即适配:重新思考如何改变语言模型行为
- HiVLA: Decoupling Robot Reasoning from Motor Control HiVLA: 解耦机器人推理与运动控制
- LongCoT: When AI Models Hit the Wall of Extended Reasoning LongCoT:当AI模型撞上长程推理的墙
- One Token per Frame: Extreme Video Compression for Long-Form Understanding 每帧一个令牌:长视频理解的极限压缩
- SpatialEvo: Teaching AI Spatial Reasoning Through Self-Play with Physics SpatialEvo:通过物理自对弈教会 AI 空间推理
- TREX: Automating the Full LLM Training Pipeline with Multi-Agent Tree Search TREX:用多智能体树搜索自动化大语言模型训练全流程
- UI-Zoomer: Smart Zoom-In for Finding Tiny Buttons in Screenshots UI-Zoomer:用不确定性驱动的智能放大找到界面中的小按钮
- DDTree: Turning Block Diffusion Drafters into Tree-Based Speculative Decoders DDTree:将块扩散草稿器转化为树状推测解码器
- HypoExplore: Neural Architecture Discovery as Scientific Inquiry HypoExplore:将神经架构发现变成科学探索
- Generative Refinement Networks: Painting Images Like an Artist, Not a Printer 生成式精化网络:像画家而非打印机那样生成图像
- Lightning OPD: Making Teacher-Student Distillation 4x Faster by Going Offline Lightning OPD:离线蒸馏让大模型训练提速4倍
- LogicEval: Testing Whether AI Can Actually Fix Logic Bugs in Real Software LogicEval: 测试 AI 能否真正修复真实软件中的逻辑漏洞
- Lyra 2.0: Building Persistent 3D Worlds from Video Generation Lyra 2.0:从视频生成构建持久3D世界
- One Token Away from Collapse: How Simple Word Bans Break Instruction-Tuned LLMs 一词之差即崩溃:指令微调模型的脆弱性
- When Teachers Fail Students: The Hidden Dynamics of On-Policy Distillation 当教师模型误导学生:在线蒸馏的隐藏动力学
- See, Point, Refine: Teaching AI Agents to Click Precisely Through Trial and Error 看、指、修正:通过试错让 AI 智能体学会精准点击
- AiScientist: Autonomous Long-Horizon ML Research Through File-as-Bus Architecture AiScientist:基于文件总线架构的自主长周期机器学习研究
- Looped Reasoning Models: How Cyclic Layers Learn to Think in Stages 循环推理模型:循环层如何学会分阶段思考
- Triadic Suffix Tokenization: Teaching LLMs to Count by Giving Numbers a Backbone 三元后缀分词法:给数字装上脊梁骨,教会大模型数数
- ClawGuard: Deterministic Defense Against Prompt Injection in LLM Agents ClawGuard:LLM智能体的确定性提示注入防御
- ClawGUI: The Missing Infrastructure for GUI Agents That Actually Ship ClawGUI:让 GUI 智能体真正落地的缺失基础设施
- Decomposing Hidden Measurement Error in LLM Evaluation Pipelines 拆解大模型评估管道中的隐藏测量误差
- Meerkat: Finding Rare Safety Violations Across Agent Traces with Clustering and Adaptive Search Meerkat:用聚类和自适应搜索在智能体轨迹中发现罕见安全违规
- General365: When LLMs Excel at Math but Stumble on Everyday Reasoning General365:当大模型擅长数学却在日常推理上栽跟头
- GenTac: Teaching Machines to Dream Up Soccer Tactics Like a Coach GenTac:让机器像教练一样构想足球战术
- Object-Centric Vision Meets LMMs: From Scene Understanding to Precise Object Control 以物体为中心的视觉遇见大型多模态模型:从场景理解到精确物体控制
- LottieGPT: Teaching AI to Generate Editable Vector Animations LottieGPT:让AI生成可编辑的矢量动画
- MLLM-as-a-Judge Exhibits Model Preference Bias 多模态大模型评判者的模型偏好偏差
- Relax: Asynchronous RL Training for Omni-Modal AI at Scale Relax:面向全模态AI的异步强化学习训练引擎
- Physics Simulators as Infinite Training Data: Teaching LLMs to Reason Through Synthetic Worlds 物理模拟器作为无限训练数据:通过合成世界教会大模型推理
- SyncFix: Multi-View Consistency for 3D Reconstruction Refinement SyncFix:通过多视图同步修复三维重建
- UniToolCall: Standardizing How AI Agents Use Tools UniToolCall:统一 AI 智能体的工具使用范式
- BERT-as-a-Judge: Lightweight Semantic Evaluation for LLM Outputs BERT-as-a-Judge:轻量级语义评估方法
- Do Vision Language Models Need to Process Image Tokens? 视觉语言模型需要处理图像令牌吗?
- E3-TIR: Balancing Expert Guidance and Self-Exploration in Tool-Using AI Agents E3-TIR:工具使用AI智能体中专家指导与自主探索的平衡
- From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models 从推理到智能体:大语言模型强化学习中的信用分配
- Optical Hardware Acceleration for Transformer Attention: Replacing Softmax with Lithium Niobate Modulators 用光学硬件加速 Transformer 注意力:用铌酸锂调制器替代 Softmax
- The Hidden Architecture of Harm: How LLMs Store Dangerous Knowledge in Compressed Neural Circuits 危害的隐藏架构:大语言模型如何在压缩神经回路中存储危险知识
- MANYFAKE: Exposing How Modern Fake News Hides in Plain Sight MANYFAKE:揭示现代假新闻如何藏身于真实叙事之中
- RecaLLM: Teaching Language Models to Think and Retrieve in Tandem RecaLLM: 让语言模型学会边思考边检索
- When Agents Speak Different Languages: The Communication Cost of Mismatched Semantic Spaces 当智能体说不同的语言:语义空间错配的通信代价
- Tango: Smarter Token Pruning for Efficient Video LLMs Tango: 为高效视频大语言模型设计的智能令牌剪枝
- UIPress: Optical Token Compression for UI-to-Code Generation UIPress:UI转代码生成中的光学令牌压缩
- VisionFoundry: Teaching Vision Models to See with Synthetic Data VisionFoundry: 用合成数据教视觉模型「看」
- VISOR: Teaching AI Agents to Search Visual Documents Without Getting Lost VISOR:让 AI 智能体在视觉文档中搜索而不迷失方向
- VL-Calibration: Teaching Vision-Language Models to Separate 'I Can't See' from 'I Don't Know' VL-Calibration:教会视觉-语言模型区分「看不清」和「想不通」
- Fail2Drive: Why Self-Driving Models Memorize Routes Instead of Learning to Drive Fail2Drive:为什么自动驾驶模型只会背路线而不会真正开车
- RewardFlow: Steering Diffusion Models with Multi-Objective Optimization RewardFlow: 用多目标优化引导扩散模型
- Scal3R: Test-Time Training Meets Large-Scale 3D Reconstruction Scal3R: 测试时训练遇上大规模三维重建
- SelfEvo: Teaching 4D Vision Models to Improve Themselves Without Labels SelfEvo:让4D视觉模型在无标注下自我进化
- SIM1: Grounding Simulation in Reality to Scale Deformable Object Manipulation SIM1:将仿真锚定于现实以扩展可变形物体操作
- Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models Act Wisely:在多模态智能体中培养元认知工具使用能力
- Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest AI 聊天机器人中的广告?大语言模型如何应对利益冲突
- ClawBench: Can AI Agents Complete Everyday Online Tasks? ClawBench:AI 智能体能完成日常在线任务吗?
- Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models 揭秘 OPD:大语言模型的长度膨胀与稳定化策略
- Faithful GRPO: Improving Visual Spatial Reasoning via Constrained Policy Optimization Faithful GRPO:通过约束策略优化改进视觉空间推理
- Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding 元学习上下文推理:无需训练即可跨被试解码大脑活动
- OpenVLThinkerV2: Gaussian GRPO for Balanced Multimodal Reasoning OpenVLThinkerV2:用于平衡多模态推理的高斯GRPO
- PIArena: A Platform for Prompt Injection Evaluation PIArena:提示注入攻击评估统一平台
- PSI: Shared State as the Missing Layer for Coherent AI-Generated Instruments in Personal AI Agents PSI:共享状态作为个人 AI 智能体中连贯 AI 生成工具的缺失层
- Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts 看见但不思考:多模态 MoE 中的路由分心现象
- SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions SUPERNOVA:通过自然指令强化学习激发大语言模型的通用推理能力
- What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal 表示引导的驱动机制:拒绝引导的机制化案例研究
- Claw-Eval: Making Agent Evaluation Actually Trustworthy Claw-Eval:让智能体评估真正可信
- Exclusive Unlearning: Forgetting Everything Except What Matters 排他性遗忘:只记住重要的,忘掉其他一切
- Gym-Anything: Scaling Computer-Use Agents to Real Economic Tasks Gym-Anything:将计算机使用智能体扩展到真实经济任务
- In-Place Test-Time Training: Making LLMs Learn While They Work 原地测试时训练:让大语言模型边工作边学习
- PoM: Trading Quadratic Attention for Linear Polynomial Mixing PoM:用线性多项式混合替代二次方注意力
- Target Policy Optimization: Decoupling What to Learn from How to Learn in RL 目标策略优化:在强化学习中解耦「学什么」与「怎么学」
- Beyond Next-Token Prediction: Building Coherent World Models Through Multi-Token Prediction and Latent Semantic Enhancement 超越下一词预测:通过多词预测和潜在语义增强构建连贯世界模型
- Who Governs the Machine? A Taxonomy for AI Identity Governance Across Borders 谁来管理机器?跨境AI身份治理分类体系
- Are Latent Reasoning Models Easily Interpretable? 隐式推理模型真的可解释吗?
- CoDE-Stop: When to Stop Thinking by Watching Confidence Dance CoDE-Stop:通过观察置信度动态决定何时停止推理
- SA-HGNN: Predicting Power Outages by Teaching Neural Networks About Geography SA-HGNN:让神经网络理解地理,预测极端天气停电
- TriAttention: Trigonometric KV Compression for Efficient Long Reasoning TriAttention:基于三角函数的 KV 压缩实现高效长推理
- Vero: Cracking Open the Black Box of Visual Reasoning with Full-Stack RL Vero: 用全栈强化学习撬开视觉推理的黑箱
- BAS: A Decision-Theoretic Approach to Evaluating LLM Confidence BAS:从决策论视角重新定义大语言模型置信度评估
- Second-Guessing Success: Gradient Boosting within a Single Attention Layer 学会‘回过头看’:在单个注意力层内实现梯度提升
- Scaling's Dead End: The Discrete Action Bottleneck in VLA Models 扩展的死胡同:VLA 模型中动作离散化的信息瓶颈
- Batched Contextual Reinforcement: A Simple Way to Make Reasoning Models Think Shorter Batched Contextual Reinforcement:让推理模型更短更省的一种简单训练法
- Generative World Renderer: Using AAA Game Worlds to Train Better Inverse and Forward Rendering Generative World Renderer:把 AAA 游戏世界变成可训练的真实渲染数据源
- Omni123: A Single Autoregressive Model for Both 2D Images and 3D Generation Omni123:把2D图像生成和3D生成塞进同一个自回归模型
- Steerable Visual Representations: Letting Vision Features Listen to Language Without Becoming CLIP 可转向的视觉表征:让视觉特征听懂语言,但不变成语言主导
- Replacing Softmax with an Int8-Friendly Surrogate for Edge Transformers 用适合 Int8 的 Softmax 替身加速边缘 Transformer
- CliffSearch: Co-Evolving Theory and Code for Scientific Algorithm Discovery CliffSearch:理论与代码共同进化的科学算法发现框架
- Self-Distillation: Teaching LLMs to Code Better Using Only Their Own Outputs 自蒸馏:让大模型仅用自己的输出学会更好地编程
- ORCA: Making LLM Reasoning Cheaper Through Smart Calibration ORCA:通过智能校准让大模型推理更省钱
- YC-Bench: Testing AI Agents on Year-Long Startup Simulation YC-Bench:用一年期创业模拟测试 AI 智能体
- Universal YOCO: Recursive Depth Scaling with Constant Memory Universal YOCO:常量内存下的递归深度扩展
- When Training Breaks Transparency: A Framework for Safe Chain-of-Thought Optimization 当训练破坏透明性:链式思维安全优化框架
- Architecting Secure AI Agents: System-Level Defenses Against Indirect Prompt Injection 构建安全AI智能体:系统级防御间接提示注入攻击
- Transformer Models for Automatic Loop Parallelization: When DistilBERT Meets Compiler Analysis 用 Transformer 自动识别可并行循环:当 DistilBERT 遇上编译器分析
- GeoCodeBench: Why AI Still Can't Write PhD-Level 3D Vision Code GeoCodeBench:为什么AI还写不出博士级3D视觉代码
- YARN: Teaching Machines to Reason by Analogy Through Structured Abstraction YARN:通过结构化抽象教机器进行类比推理
- MONA Camera Dropbox: From Oracle Approval to Learned Overseers in Reward-Hacking Mitigation MONA 相机投递箱:从预言机审批到学习型监督者的奖励黑客缓解
- Hybrid RL-LLM Framework for Robotic Manipulation: Bridging Low-Level Control and High-Level Reasoning 混合强化学习-大语言模型机器人操作框架:连接底层控制与高层推理
- OmniRoam: Panoramic Video Generation for Infinite Scene Exploration OmniRoam:全景视频生成实现无限场景漫游
- Refined Detection for Gumbel Watermarking Gumbel 水印的精细化检测
- Cost-Aware LLM Routing with NeuralUCB: Balancing Quality and Efficiency 用 NeuralUCB 做成本感知的大模型路由:在质量和效率间找平衡
- Triadic Cognitive Architecture: Teaching AI Agents When to Stop Thinking 三元认知架构:教会AI智能体何时停止思考
- Tracking Equivalent Mechanistic Interpretations Across Neural Networks 追踪神经网络间的等价机制解释
- Tucker Attention: Unifying Efficient Attention Through Tensor Decomposition Tucker 注意力:通过张量分解统一高效注意力机制
- Video Models Reason Early: Exploiting Plan Commitment for Maze Solving 视频模型的早期推理:利用计划承诺解决迷宫问题
- IF4: Adaptive Block-Scaled Data Types for Smarter 4-bit Quantization IF4:自适应块缩放数据类型——更聪明的4比特量化
- Gen-Searcher: Teaching Image Generators to Search Before They Draw Gen-Searcher:让图像生成模型先搜索再作画
- HandX: Scaling Bimanual Motion and Interaction Generation HandX:双手运动与交互生成的规模化突破
- On-the-fly Repulsion in the Contextual Space for Rich Diversity in Diffusion Transformers 扩散Transformer中的上下文空间即时排斥机制:实现丰富多样性
- HyperP: Transferable Hyperparameter Scaling for Stable Language Model Training HyperP: 可迁移的超参数缩放实现稳定语言模型训练
- VeoPlace: Teaching Vision-Language Models to Place Chips Like Human Designers VeoPlace:让视觉语言模型像人类设计师一样布局芯片
- SOLE-R1: Teaching Robots Through Video Reasoning Alone SOLE-R1:仅用视频推理教会机器人
- SonoWorld: Turning a Single Image into an Immersive 3D Audio-Visual World SonoWorld:从一张图片生成沉浸式3D视听场景
- Stop Probing, Start Coding: Why Sparse Autoencoders Fail at Compositional Generalization 别探测了,写代码吧:稀疏自编码器为何在组合泛化上失败
- Temporal Credit Is Free: Why Recurrent Networks Don't Need Full Backpropagation 时间信用是免费的:为什么循环网络不需要完整反向传播
- GaussianGPT: Autoregressive 3D Scene Generation via Token Prediction GaussianGPT:通过token预测实现自回归3D场景生成
- AnyHand: Scaling Hand Pose Estimation with 6.6M Synthetic RGB-D Images AnyHand:用660万张合成RGB-D图像突破手部姿态估计瓶颈
- Back to Basics: Revisiting ASR in the Age of Voice Agents 回归本质:语音代理时代的ASR重审
- BizGenEval: Testing If AI Can Actually Design Business Documents BizGenEval: 测试 AI 能否真正设计商业文档
- Drive My Way: Teaching Self-Driving Cars Your Personal Driving Style Drive My Way: 让自动驾驶学会你的个人驾驶风格
- LGTM: Breaking the 4K Barrier in Feed-Forward 3D Gaussian Splatting LGTM:突破前馈式3D高斯泼溅的4K分辨率瓶颈
- MuRF: Multi-Resolution Fusion for Vision Foundation Models MuRF: 视觉基础模型的多分辨率融合
- Natural-Language Agent Harnesses: Making Agent Control Logic Portable and Inspectable 自然语言智能体线束:让智能体控制逻辑可移植、可检视
- PSDesigner: Mimicking Human Creative Workflow for Automated Graphic Design PSDesigner: 模仿人类创意工作流的自动化平面设计系统
- SlotVTG: Teaching Video Models to See Objects, Not Shortcuts SlotVTG: 让视频模型看见物体而非捷径
- WriteBack-RAG: Training Knowledge Bases Through Evidence Distillation WriteBack-RAG:通过证据蒸馏训练知识库
- MegaFlow: Zero-Shot Large Displacement Optical Flow MegaFlow:零样本大位移光流估计
- PackForcing: Training on 5-Second Clips to Generate 2-Minute Videos PackForcing: 用5秒短视频训练生成2分钟长视频
- R-C2: Teaching AI to Think Consistently Across Vision and Language R-C2: 让 AI 在视觉与语言间保持一致性思考
- ShotStream: Real-Time Interactive Video Storytelling via Causal Multi-Shot Generation ShotStream: 基于因果多镜头生成的实时交互式视频叙事
- Vega: Teaching Self-Driving Cars to Follow Natural Language Instructions Vega: 让自动驾驶汽车听懂人话
- Anti-I2V: Protecting Photos from Malicious AI Video Generation Anti-I2V: 保护照片免受恶意AI视频生成
- Chameleon: Geometry-Grounded Episodic Memory for Long-Horizon Robot Manipulation Chameleon: 面向长时域机器人操作的几何基础情景记忆
- When AI Judges Code: The Hidden Biases Between LLMs and Human Developers 当AI评判代码:大模型与人类开发者之间的隐藏偏见
- DreamerAD: 80x Faster World Model RL for Autonomous Driving via One-Step Latent Diffusion DreamerAD:通过单步潜在扩散实现自动驾驶世界模型强化学习的80倍加速
- EndoVGGT: Graph Neural Networks Learn Tissue Geometry Beyond Occlusions EndoVGGT: 图神经网络跨越遮挡学习组织几何
- Latent-WAM: Compact World Models for Autonomous Driving via Spatial-Aware Compression Latent-WAM: 通过空间感知压缩实现自动驾驶的紧凑世界模型
- LensWalk: Teaching AI Agents to Control Their Own Eyes When Watching Videos LensWalk:让AI智能体主动控制自己如何观看视频
- Polynomial Speedup in Diffusion Models with the Multilevel Euler-Maruyama Method 多层次欧拉-丸山方法:扩散模型的多项式加速
- RAVEN: Teaching Foundation Models to Predict Hospital Visits by Learning What Comes Back RAVEN:通过学习复发模式让基础模型预测就诊事件
- TAG: Fixing Robot Vision's 'Close Enough' Problem with Target-Agnostic Guidance TAG:用目标无关引导解决机器人视觉的「差不多」问题
- Free-Market Algorithm: When Economics Meets Optimization 自由市场算法:当经济学遇上优化问题
- TextFlow: Training-Free Scene Text Editing via Flow Manifold Steering TextFlow: 基于流形引导的免训练场景文字编辑
- Trust Region Bayesian Optimization: Taming High-Dimensional Constrained Black-Box Problems with Penalties 信赖域贝叶斯优化:用惩罚法驯服高维约束黑盒问题
- VFIG: Teaching Vision Models to Redraw Technical Figures as Editable Vectors VFIG: 教视觉模型把技术图表重绘成可编辑矢量
- Chameleon: Episodic Memory for Long-Horizon Robotic Manipulation Chameleon: 面向长时程机器人操作的情景记忆
- Evidence of an Emergent 'Self' in Continual Robot Learning 持续机器人学习中涌现“自我”的证据
- Improving Lean4 Autoformalization via Cycle Consistency Fine-tuning 通过循环一致性微调改进Lean4自动形式化
- MARCH: Three Agents for Grounded Self-Checking in RAG MARCH:用三智能体自检降低RAG幻觉
- One View Is Enough: Training Novel View Synthesis on Unpaired Internet Images 单视图足矣:用互联网无配对图像训练新视角生成
- RealMaster: Bridging the Sim-to-Real Gap in Video Generation RealMaster:跨越渲染视频到真实视频的鸿沟
- Retrieval Improvements Do Not Guarantee Better Answers: A Study of RAG for AI Policy QA 检索改进并不保证更好的答案:AI政策问答中的RAG研究
- SpecEyes: Speculative Acceleration for Agentic Multimodal LLMs SpecEyes: 智能体多模态大模型的推测加速
- The Stochastic Gap: A Markovian Framework for Pre-Deployment Reliability and Oversight-Cost Auditing in Agentic Artificial Intelligence Stochastic Gap: 面向代理式人工智能的部署前可靠性与监督成本审计的马尔可夫框架
- The Stochastic Gap: A Markovian Framework for Pre-Deployment Reliability and Oversight-Cost Auditing in Agentic Artificial Intelligence Stochastic Gap: 面向代理式人工智能的部署前可靠性与监督成本审计的马尔可夫框架
- UniGRPO: Unified Reinforcement Learning for Reasoning-Driven Image Generation UniGRPO: 推理驱动的视觉生成统一策略优化
- VISOR: Sparse Attention for Efficient Vision-Language Models Without Visual Token Loss VISOR:通过稀疏注意力实现高效视觉语言模型而不丢失视觉信息
- Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions 更弱假设下的拜占庭鲁棒与差分隐私联邦优化
- DAK-UCB: Diversity-Aware Prompt Routing for LLMs and Generative Models DAK-UCB:面向大语言模型和生成模型的多样性感知提示路由
- Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe 揭秘长周期工具使用智能体的强化学习:一份完整配方
- Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models 先描述再行动:通过蒸馏语言-动作世界模型实现主动智能体引导
- DualCoT-VLA: Teaching Robots to Think Visually and Linguistically Before Acting DualCoT-VLA: 让机器人在行动前同时进行视觉和语言思考
- UNITE: Single-Stage Joint Training of Image Tokenization and Diffusion Generation UNITE:图像分词与扩散生成的单阶段联合训练
- Failure of Contextual Invariance in Gender Inference with Large Language Models 大语言模型性别推断中的上下文不变性失效
- ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM Alignment ImplicitRM:从隐式偏好数据中进行无偏奖励建模以实现LLM对齐
- Knowledge Access Beats Model Size: Memory Augmented Routing for Persistent AI Agents 知识获取胜过模型规模:持久AI智能体的记忆增强路由
- MemCollab: Cross-Agent Memory Collaboration via Contrastive Trajectory Distillation MemCollab:通过对比轨迹蒸馏实现跨智能体记忆协作
- SPLADE-Code: Bringing Learned Sparse Retrieval to Code Search SPLADE-Code:将学习型稀疏检索引入代码搜索
- Polaris: A Gödel Agent Framework for Small Language Models through Experience-Abstracted Policy Repair Polaris:通过经验抽象策略修复实现小语言模型的哥德尔智能体框架
- Geometric Latent Diffusion: Repurposing Foundation Models for Multi-view Synthesis 几何隐空间扩散:重用基础模型实现多视角合成
- ROM: Catching AI Models Before They Overthink ROM:在AI模型过度思考前及时刹车
- Set-Valued Prediction for Large Language Models with Feasibility-Aware Coverage Guarantees 面向大语言模型的集值预测及可行性感知覆盖保证
- Sparser, Faster, Lighter Transformer Language Models 更稀疏、更快速、更轻量的Transformer语言模型
- daVinci-MagiHuman: Single-Stream Architecture Unifies Audio-Video Generation daVinci-MagiHuman:单流架构统一音视频生成
- ThinkJEPA: Bridging Dense Prediction and Semantic Reasoning in Latent World Models ThinkJEPA:在潜在世界模型中连接密集预测与语义推理
- Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low Retention 视频大语言模型的统一时空token压缩:极低保留率下的突破
- UniMotion: Treating Motion as a First-Class Continuous Modality UniMotion: 将动作视为一等连续模态
- The End of Coding: Andrej Karpathy on AI Agents, AutoResearch, and the Loopy Era 编程的终结:Andrej Karpathy 谈 AI 智能体、自动研究与循环时代
- Confidence-Based Decoding is Provably Efficient for Diffusion Language Models 基于置信度的解码对扩散语言模型具有可证明的高效性
- Decoupling Exploration and Policy Optimization: Uncertainty Guided Tree Search for Hard Exploration 解耦探索与策略优化:面向困难探索的不确定性引导树搜索
- Gumbel Distillation for Parallel Text Generation Gumbel蒸馏实现并行文本生成
- Scaling DoRA: High-Rank Adaptation via Factored Norms and Fused Kernels 扩展DoRA:通过分解范数和融合核实现高秩适配
- TiCo: Time-Controllable Training for Spoken Dialogue Models TiCo:面向口语对话模型的时间可控训练
- Jensen Huang on NVIDIA, Rack-Scale AI, and Extreme Co-Design Jensen Huang 谈 NVIDIA、机架级 AI 与极致协同设计
- CubiD: Breaking the Dimensionality Ceiling in Discrete Visual Generation CubiD:突破离散视觉生成的维度天花板
- Pitfalls in Evaluating Interpretability Agents Pitfalls in Evaluating Interpretability Agents
- Spectral Alignment in Forward-Backward Representations via Temporal Abstraction Spectral Alignment in Forward-Backward Representations via Temporal Abstraction
- The Y-Combinator for LLMs: Solving Long-Context Rot with λ-Calculus LLM的Y组合子:用λ演算解决长上下文退化问题
- Var-JEPA: A Variational Formulation of Joint-Embedding Predictive Architecture Var-JEPA: A Variational Formulation of Joint-Embedding Predictive Architecture
- Adapt4Me: Uncertainty-Aware Authoring for Personalizing ASR to Non-normative Speech Adapt4Me: Uncertainty-Aware Authoring for Personalizing ASR to Non-normative Speech
- Chain-of-Adaptation: Surgical Vision-Language Adaptation with Reinforcement Learning Chain-of-Adaptation: Surgical Vision-Language Adaptation with Reinforcement Learning
- Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on LLMs Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on LLMs
- An Agentic Multi-Agent Architecture for Cybersecurity Risk Management An Agentic Multi-Agent Architecture for Cybersecurity Risk Management
- Enhancing HAL Representations via Attention-Based Pooling for Text Classification Enhancing HAL Representations via Attention-Based Pooling for Text Classification
- Design-OS: A Specification-Driven Framework for Engineering System Design Design-OS: A Specification-Driven Framework for Engineering System Design
- Semantic Token Clustering for Efficient Uncertainty Quantification in LLMs Semantic Token Clustering for Efficient Uncertainty Quantification in LLMs
- The Robot's Inner Critic: Self-Refinement of Social Behaviors through VLM-based Replanning The Robot's Inner Critic: Self-Refinement of Social Behaviors through VLM-based Replanning
- Learning Dynamic Belief Graphs for Theory-of-Mind Reasoning Learning Dynamic Belief Graphs for Theory-of-Mind Reasoning
- Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM CoT Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM CoT
- AI Agents Can Autonomously Perform Experimental High Energy Physics AI Agents Can Autonomously Perform Experimental High Energy Physics
- Adaptive Greedy Frame Selection for Long Video Understanding Adaptive Greedy Frame Selection for Long Video Understanding
- Improving Generalization on Cybersecurity Tasks with Multi-Modal Contrastive Learning Improving Generalization on Cybersecurity Tasks with Multi-Modal Contrastive Learning
- VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
- LumosX: Personalized Video Generation with Identity-Attribute Relations LumosX: Personalized Video Generation with Identity-Attribute Relations
- From Masks to Pixels and Meaning: VLM Image Tampering From Masks to Pixels and Meaning: VLM Image Tampering
- AI Agents Can Already Autonomously Perform Experimental High Energy Physics AI 智能体已经可以自主执行实验高能物理学
- Do VLMs Actually Need Vision Transformers? A Case for SSM Encoders 视觉语言模型真的需要视觉Transformer吗?状态空间模型编码器的潜力
- Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on LLMs 演化越狱:LLM 的自动化多目标长尾攻击
- Experience is the Best Teacher: Motivating Effective Exploration in RL for LLMs 经验是最好的老师:激励 LLM 强化学习中的有效探索
- Learning Dynamic Belief Graphs for Theory-of-mind Reasoning 学习动态信念图进行心智理论推理
- Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM CoT Evaluation 测量忠实度取决于如何测量:LLM 思维链评估中的分类器敏感性
- Nemotron-Cascade 2: Teaching a Small Model to Think Big with CascadeRL and On-Policy Distillation Nemotron-Cascade 2:用级联强化学习和在线蒸馏让小模型学会深度推理
- Pitfalls in Evaluating Interpretability Agents 评估可解释性智能体的陷阱
- R-Equivalence on Cubic Surfaces: Closing a50-Year Gap with AI-Assisted Profs 三次曲面上的R-等价:用AI辅助证明填补五十年空白
- Semantic Token Clustering for Efficient Uncertainty Quantification in LLMs 语义令牌聚类:LLM 中的高效不确定性量化
- Spectrally-Guided Diffusion Noise Schedules: Tailoring Noise to Image Frequency Content 频谱引导的扩散噪声调度:根据图像频率内容定制噪声
- The Y-Combinator for LLMs: Solving Long-Context Rot with λ-Calculus LLM 的 Y 组合子:用 λ 演算解决长上下文腐烂
- Terence Tao on How the World's Top Mathematician Uses AI Terence Tao 谈世界顶级数学家如何使用 AI
- AgentFactory: Growing a Library of Executable Agents Instead of Piling Up Prompts AgentFactory:用可执行子智能体积累经验,而不是堆砌提示词
- MUD: Faster Transformer Training via Triangular Momentum Whitening MUD:用三角动量去相关加速 Transformer 训练
- Teaching Robots to Follow Rules Without Retraining: STL-Constrained Action Distribution Shaping 无需重训练让机器人遵守规则:基于STL的动作分布整形
- Stop Breaking Things: How Graph-Based Impact Analysis Cuts AI Coding Regressions by 70% 别再搞坏代码了:图结构影响分析如何将 AI 编码回归率降低 70%
- Steering Image Generation with Text Embedding Interpolation — No Training Required 无需训练,用文本嵌入插值精准操控图像生成
- Demystifying Video Reasoning: How Diffusion Models Think Through Denoising, Not Frames 揭秘视频推理:扩散模型通过去噪步骤而非帧序列进行思考
- Making LLM Reasoning Work on Your Phone: Budget-Aware Adapters for Edge Deployment 让大模型推理在手机上跑起来:面向边缘部署的预算感知适配器
- M^3: Plugging a Matching Head into Multi-View Foundation Models for Monocular SLAM M^3:为多视图基础模型插上匹配头,让单目 SLAM 真正好用
- Online Experiential Learning: Teaching LLMs From Their Own Deployment History 在线经验学习:让大模型从自身部署经历中持续成长
- SegviGen: Turning 3D Generators into Part Segmentation Tools SegviGen:将3D生成模型改造为零件分割工具
- Stochastic Resetting as a Learning Accelerator in Reinforcement Learning 随机重置:强化学习中的收敛加速器
- Fixing Vision Transformers' Hidden Bias: Why Your Model Cares Where Things Are (Even When It Shouldn't) 修复视觉Transformer的隐藏偏见:为什么你的模型在意东西在哪儿(即使它不该在意)
- WorldCam: Grounding Interactive Game Worlds in Camera Geometry WorldCam:用相机几何锚定交互式游戏世界
- WorldDrive: Teaching World Models to Speak the Planner's Language WorldDrive:让世界模型说规划器的语言
- CyCLeGen: Unifying Vision Understanding and Generation Through Cycle-Consistent Loops CyCLeGen: 通过循环一致性回路统一视觉理解与生成
- Directional Routing in Transformers: When the Conductor Matters More Than the Orchestra Transformer 中的方向路由:指挥比乐团更重要
- FAR-Drive: Building Interactive Driving Simulators That Don't Drift Off Course FAR-Drive:构建不会偏离轨道的交互式驾驶模拟器
- GlyphPrinter: Teaching AI to Render Text by Showing It What's Wrong, Not Just What's Right GlyphPrinter: 通过展示错误而非仅展示正确来教AI渲染文字
- Seoul World Model: Grounding Video Generation in Real Cities 首尔世界模型:在真实城市中锚定视频生成
- HorizonMath: Testing AI's Ability to Discover New Mathematics HorizonMath:测试AI发现新数学的能力
- LLM as Graph Kernel: Rethinking Message Passing on Text-Rich Graphs 将大语言模型作为图核:重新思考文本丰富图上的消息传递
- DeepVision-VLA: Keeping Visual Context Alive Through the Depths of Robot Action Models DeepVision-VLA:在机器人动作模型的深层保持视觉上下文活跃
- Mixture-of-Depths Attention: Letting Deep Layers Remember Shallow Insights 深度混合注意力:让深层记住浅层的洞见
- Neural Networks Meet Lithography: Accelerating EUV Mask Simulation with Physics-Informed Learning 神经网络遇上光刻:用物理信息学习加速极紫外掩模仿真
- DOMINO: Teaching Robots to Catch Moving Objects DOMINO:教机器人抓取运动物体
- Machine Learning Meets Celestial Mechanics: Clustering 22,000 Saturn Moon Orbits 机器学习遇上天体力学:对22,000条土星卫星轨道聚类
- Defending Brain Tumor Classifiers: Diffusion Denoising Meets Matrix Factorization 保卫脑肿瘤分类器:扩散去噪遇上矩阵分解
- Teaching AI Agents to Remember: From Execution to Expertise in Computational Science 让AI智能体学会记忆:从执行任务到积累专业知识
- Privacy Lives in a Few Critical Weights: Surgical Fine-Tuning for Membership Privacy 隐私藏在少数关键权重中:手术式微调保护成员隐私
- Constitutional Governance for LLM-Mediated Multi-Agent Systems: When Cooperation Needs Guardrails 大语言模型多智能体系统的宪法治理:当合作需要护栏
- MXNorm: Hijacking Block Scales from Low-Precision Matmuls to Speed Up Normalization MXNorm:劫持低精度矩阵乘法的块尺度来加速归一化
- Selecting Training Data by Reading the Model's Mind: Neuron Activation Patterns for Instruction Tuning 通过读取模型的"心思"选择训练数据:用神经元激活模式指导指令微调
- Out of Sight, Out of Mind? Testing Whether Video World Models Understand Hidden State Evolution 眼不见心不烦?测试视频世界模型能否理解隐藏状态演化
- Perceive What Matters: Smart Scheduling for Real-Time Robot Perception 感知关键信息:实时机器人感知的智能调度
- PhysMoDPO: Teaching Motion Models to Respect Physics Through Preference Learning PhysMoDPO:通过偏好学习让运动模型尊重物理规律
- Stop Predicting Pixels: Why Latent Space Learning Beats Next-Frame Prediction for Physical Systems 别再预测像素了:为什么潜空间学习在物理系统中胜过下一帧预测
- Steve-Evolving: Self-Improving Minecraft Agents Through Experience Diagnosis and Knowledge Distillation Steve-Evolving:通过经验诊断和知识蒸馏实现自我进化的 Minecraft 智能体
- World Scene Graphs: Tracking Objects Through Occlusion in 3D Video Understanding 世界场景图:在3D视频理解中追踪被遮挡的物体
- Visual-ERM: Teaching AI to Judge Code by What It Renders, Not What It Says Visual-ERM:让AI通过渲染结果而非代码文本来评判代码质量
- BiGain: Making Diffusion Models Fast Without Choosing Between Quality and Accuracy BiGain:让扩散模型加速时不必在质量和准确率之间二选一
- DreamVideo-Omni: Teaching Video Models to Remember Faces While Following Motion Scripts DreamVideo-Omni:让视频模型同时记住人脸、听懂动作指令
- EndoCoT: Teaching Diffusion Models to Think Before They Draw EndoCoT:让扩散模型先思考再生成
- Reasoning Judges in RL Alignment: A Double-Edged Sword 推理型裁判在强化学习对齐中的双刃剑效应
- GRADE: Can AI Models Actually Edit Images With Domain Knowledge? GRADE:多模态模型真的懂学科知识吗?
- SciMDR: Teaching Models to Reason Across Full Scientific Papers SciMDR:让模型真正读懂完整科学论文
- AutoGaze: Teaching Video Models to Look Before They See AutoGaze:让视频模型先选帧再看帧
- DVD: Turning Video Diffusion Models into Deterministic Depth Estimators DVD:让视频扩散模型变成确定性深度估计器
- EVATok: Stop Wasting Tokens on Boring Video Frames EVATok:别再把 Token 浪费在无聊的视频帧上了
- Energy-Based Fine-Tuning: Teaching LLMs to Match Meaning, Not Just Tokens 基于能量的微调:让语言模型匹配语义而非逐词预测
- MM-CondChain: A Benchmark That Finally Tests Whether Vision-Language Models Can Actually Follow Conditional Logic MM-CondChain:终于有人认真测试多模态模型的条件推理能力了
- OmniStream: One Backbone to See, Reconstruct, and Act in Real Time OmniStream:一个骨干网络,实时感知、重建与行动
- ELIT: One Diffusion Model, Any Compute Budget ELIT:一个扩散模型,任意算力预算
- SceneAssistant: Closing the Loop on Text-to-3D Scene Generation with Visual Feedback SceneAssistant:用视觉反馈闭环驱动开放词汇3D场景生成
- One Architecture to Rule Them All: Separable Neural Networks as a Unified Primitive 一种架构统一预测与生成:可分离神经网络作为通用基元
- Spatial-TTT: Teaching Video Models to Remember Where Things Are Spatial-TTT:让视频模型记住空间在哪里
- STAMP: Smarter Text Privacy by Treating Tokens Differently STAMP:按词分配隐私预算的文本保护新框架
- Color by Coordinates: Finding HSL Structure in FLUX's Latent Space 坐标即色彩:在 FLUX 潜空间中发现 HSL 结构
- FIRM: Building Reward Models That Actually Know What They're Scoring FIRM:让奖励模型真正理解它在评什么
- Watch and Think at the Same Time: VST Brings Reasoning Into Live Video Streams 边看边想:VST 让视频大模型在直播流中同步推理
- Agentar-Fin-OCR: Document Parsing for Financial PDFs with Cross-Page Continuity Agentar-Fin-OCR:具备跨页连续性的金融文档解析系统
- The Illusion of Agreement: Why LLM Judges Fake Consensus and How to Fix It 共识的幻觉:LLM 评判者为何制造虚假一致,以及如何破解
- COMIC: Teaching AI Agents to Write Sketch Comedy Through Studio Role-Play COMIC:通过工作室角色扮演教AI智能体创作喜剧小品
- Does AI See like Art Historians? Interpreting How VLMs Recognize Artistic Style AI 看画的方式和艺术史家一样吗?解读视觉语言模型识别艺术风格的机制
- DynVLA: Predicting World Dynamics Before Acting in Autonomous Driving DynVLA:自动驾驶中先预测世界动态再决策
- IsalGraph: Encoding Graphs as Strings a Language Model Can Actually Read IsalGraph:让语言模型能读懂的图结构字符串编码
- Leech Lattice Vector Quantization: Making Optimal 24D Sphere Packing Practical for LLM Compression Leech格向量量化:让24维最优球堆积在大语言模型压缩中落地
- LiTo: Teaching Neural Networks to See Shine LiTo:让神经网络看懂光泽
- Neural Field Thermal Tomography: Hard Physics Constraints Meet Continuous 3D Reconstruction 神经场热层析成像:硬物理约束遇上连续3D重建
- Binary Routing in Transformer MLPs: How Language Models Make Discrete Decisions with Continuous Signals Transformer MLP 中的二值路由:语言模型如何用连续信号做离散决策
- Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity 太鲜艳了不真实?生成式色彩保真度的基准测试与校准
- BEACON: Teaching Robots to Navigate by Looking Around Corners BEACON:教机器人通过「看穿遮挡」来导航
- CREATE: A Benchmark for Testing Associative Creativity in Language Models CREATE:测试语言模型联想创造力的基准
- How Emotions Tip the Scales in Collective Decisions 情绪如何在集体决策中打破平衡
- When Neural Networks Learn to Use Interference: How Feature Correlations Create Semantic Structure 当神经网络学会利用干扰:特征相关性如何创造语义结构
- Generative Drifting Unmasked: It's Score Matching with a Frequency Problem 生成漂移的真面目:带频率问题的分数匹配
- Model Merging in LLMs: A Comprehensive Survey Through the FUSE Framework 大语言模型的模型融合:FUSE框架下的全景综述
- ReCoSplat: Teaching 3D Reconstruction to Self-Correct Through Visual Feedback ReCoSplat:通过视觉反馈让3D重建学会自我纠错
- Physics-Guided Representation Learning for Global Carbon Flux Estimation 物理引导的表征学习用于全球碳通量估算
- Think Before You Lie: How Reasoning Improves Honesty in LLMs 三思而后诚:推理如何提升大语言模型的诚实度
- When AI Guides Get Social: How Blind Users Treat VR Assistants Differently Around Others 当AI向导变得社交化:盲人用户在他人面前如何不同地对待VR助手
- EcoAI-Resilience: Optimizing AI Deployment for Sustainability and Economic Resilience EcoAI-Resilience:面向可持续性与经济韧性的AI部署优化
- AI Finds Bigger Efficiency Gap in Classic Trading Mechanism AI发现经典交易机制存在更大效率缺口
- Teaching AI Agents to Judge, Not Just Imitate: Agentic Critical Training 教AI智能体学会判断而非模仿:智能体批判性训练
- AutoAdapt: Turning LLM Domain Adaptation from Art into Engineering AutoAdapt:把大模型领域适配从手艺活变成工程活
- Can Language Models Beat FLAC? Benchmarking Neural Compression for Full-Fidelity Audio 语言模型能否击败 FLAC?全保真音频的神经压缩基准测试
- Self-Conditioned GANs Learn Trajectory Patterns Without Labels or Maps 自条件 GAN 无需标签或地图即可学习轨迹模式
- ER-Pose: Breaking Free from Bounding Boxes in Real-Time Pose Estimation ER-Pose: 摆脱边界框束缚的实时姿态估计
- AFIB: A Multi-Dimensional Benchmark for Financial Intelligence in Large Language Models AFIB:大语言模型金融智能的多维评测基准
- FVG-PT: Keeping Vision Models Focused During Prompt Tuning FVG-PT:让视觉模型在提示调优时保持专注
- HiAR: Reversing the Order of Video Generation to Fix Error Accumulation HiAR:反转视频生成顺序以解决误差累积问题
- Impermanent: Testing Time Series Models on Data That Won't Sit Still Impermanent:在不断变化的数据上测试时序模型
- Momentum SVGD-EM: Adding Flywheels to Particle-Based EM 动量 SVGD-EM:为粒子化 EM 算法装上飞轮
- OfficeQA Pro: Why Frontier AI Still Fails at Enterprise Document Reasoning OfficeQA Pro:为什么前沿AI在企业文档推理上仍然失败
- Your Voice Assistant Is Leaking Your Identity: Privacy Attacks on Full-Duplex Speech Models 你的语音助手正在泄露身份:全双工对话模型的隐私攻击
- Scale Space Diffusion: Why Process Noise at Full Resolution? 尺度空间扩散:为什么要全分辨率处理噪声?
- SERQ: Single-Matrix Error Correction for 4-bit LLM Quantization SERQ: 用单矩阵修正4比特大模型量化误差
- Split Federated Learning: Optimizing Where to Cut the Model for Accuracy and Speed 分割联邦学习:优化模型切分位置以提升精度与速度
- Structural Causal Bottleneck Models: Dimension Reduction That Respects Causality 结构因果瓶颈模型:尊重因果关系的降维方法
- TildeOpen LLM: Teaching a 30B Model to Speak 34 European Languages Fairly TildeOpen LLM: 用课程学习让300亿参数模型公平掌握34种欧洲语言
- Video2LoRA: One Hypernetwork to Rule All Video Control Conditions Video2LoRA: 用超网络统一视频生成的语义控制
- When Physics Priors Meet Scale: AllScAIP's Data-Driven Path to Long-Range Molecular Interactions 当物理先验遇上规模:AllScAIP用数据驱动捕获长程分子相互作用
- AV-Unified: Teaching One Model to Handle All Audio-Visual Tasks AV-Unified: 让单一模型处理所有视听任务
- BEVLM: Teaching Self-Driving Cars to Think in Bird's-Eye View with Language Model Wisdom BEVLM:用语言模型的智慧教会自动驾驶从鸟瞰视角思考
- CODEC: Tracing Causal Paths Through Neural Networks by Decomposing Contributions, Not Just Activations CODEC:通过分解贡献而非激活值追踪神经网络的因果路径
- LiveSense: Turning Your Laptop's Wi-Fi Into a Centimeter-Precision Radar LiveSense: 把笔记本 Wi-Fi 变成厘米级雷达
- Why MLLMs Look Bad at Classification: It's the Benchmark, Not the Model 多模态大模型分类表现差?问题出在评测上
- Teaching Diffusion Models to Say 'No': A Geometric Approach to Linguistic Negation 教扩散模型说「不」:语言否定的几何方法
- Omni-Diffusion: Ditching Autoregression for Masked Diffusion in Multimodal Models Omni-Diffusion:用掩码扩散替代自回归的多模态模型
- Penguin-VL: Why Vision-Language Models Don't Need CLIP Anymore Penguin-VL:视觉语言模型为何不再需要CLIP
- Teaching Surgical AI to Think Like a Surgeon: Mining Reasoning from Video Lectures 让手术AI像外科医生一样思考:从教学视频中挖掘推理能力
- DEBISS: Building a Multi-Layered Spoken Debate Corpus for Real Argumentation DEBISS:为真实论辩构建多层标注的口语辩论语料库
- Teaching Neural Networks to Respect Quantum Physics: Kraus-Constrained Sequence Learning 让神经网络懂物理:量子轨迹的Kraus约束序列学习
- Teaching Bangla AI to Say 'I Don't Know': A Dataset for Unanswerable Questions 教孟加拉语AI说「我不知道」:无答案问题数据集
- Thermodynamic Response Functions in Singular Bayesian Models 奇异贝叶斯模型中的热力学响应函数
- Fixing Holes in Real-Time 3D Streams: A Transformer Approach to Multi-View Inpainting 实时3D流的补洞术:基于Transformer的多视角修复
- CalibAtt: Making Video Generation 1.6x Faster by Learning Which Attention to Skip CalibAtt:通过学习跳过哪些注意力让视频生成快1.6倍
- Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation 审查模型:挖掘隐藏知识的天然实验场
- Cheap Thrills: Training Neural Optimizers with Imperfect Labels 廉价标签训练神经优化器:不完美也能用
- EdgeDAM: Bringing Distractor-Aware Memory to Mobile Object Tracking EdgeDAM: 在移动设备上实现抗干扰的目标跟踪
- FaceCam: Teaching Video Models to Move the Camera Around Faces Without Breaking Them FaceCam: 让视频模型学会围着人脸转镜头而不把脸拍坏
- HALP: Catching Hallucinations Before They Happen HALP:在幻觉发生之前就将其捕获
- Fact-Checking Without Retrieval: Tapping Into What LLMs Already Know 无需检索的事实核查:挖掘大模型的内在知识
- POET-X: Training Billion-Parameter LLMs on a Single GPU Through Efficient Orthogonal Transformations POET-X: 通过高效正交变换在单卡上训练十亿参数大模型
- Reasoning Theater: When AI Models Pretend to Think 推理剧场:当AI模型假装思考
- RoboPocket: Teaching Robots with Your Phone by Showing Them Their Future Mistakes RoboPocket: 用手机让机器人看见自己未来会犯的错
- SurvHTE-Bench: Finally, a Fair Fight for Survival Treatment Effect Methods SurvHTE-Bench: 生存分析中异质性治疗效应估计的标准擂台
- Why Transformers Have Weird Spikes and Attention Sinks: It's the Architecture, Not the Data Transformer 为什么会出现激活尖峰和注意力汇聚:架构使然,非数据所致
- Seeing Invisible Gas in 3D: Neural Radiance Fields Meet Hyperspectral Infrared Imaging 让看不见的气体现形:神经辐射场遇上高光谱红外成像
- MM-Lifelong: Teaching AI to Remember Across Days, Weeks, and Months MM-Lifelong:教AI跨越天、周、月记忆的数据集与智能体基线
- Bias-Bounded Evaluation: Making LLM Judges Provably Fair 偏差有界评估:让大模型评委可证明公平
- τ-Knowledge: Testing AI Agents Where Knowledge Meets Action τ-Knowledge: 在知识与行动交汇处测试AI智能体
- AgentIR: Teaching Retrievers to Read AI Agents' Minds AgentIR: 让检索器读懂AI智能体的心思
- Why Quantization Breaks: It's Not Just Outliers, It's Misalignment 量化为何失效:不只是离群值,更是方向错位
- Teaching Web Agents to Ignore Lies: A Co-Evolutionary Defense Against Cross-Modal Attacks 教会网页智能体识破谎言:对抗跨模态攻击的协同进化防御
- Breaking Safety Alignment by Reshaping Activation Distributions, Not Just Directions 重塑激活分布而非仅移除方向:突破大模型安全对齐的新视角
- EgoPoseFormer v2: Teaching AR/VR Headsets to See Your Whole Body from Just Your Head EgoPoseFormer v2: 让 AR/VR 头显从头部视角推断全身姿态
- Helios: Making 14B Video Models Run at 19.5 FPS Without the Usual Tricks Helios: 让140亿参数视频模型跑到19.5帧每秒,不靠常规加速术
- Teaching Robots to Peel: Learning Fine-Grained Manipulation Through Human Preference 教机器人削皮:通过人类偏好学习精细操作
- A Safety Knob for Protein Generators: Steering Away from Toxicity Without Retraining 蛋白质生成器的安全旋钮:无需重训练即可规避毒性
- Catching AI Misbehavior in Real-Time: Monitoring Reward Hacking Through Internal Activations 实时捕捉AI作弊行为:通过内部激活监测奖励黑客
- RoboCasa365: Building a Standardized Testing Ground for Generalist Home Robots RoboCasa365:为通用家务机器人构建标准化测试场
- Making AI Agents Robust Without Killing Their Intelligence: Adversarially-Aligned Jacobian Regularization 让AI智能体既鲁棒又聪明:对抗对齐的雅可比正则化
- SaFeR: Teaching AI to Generate Dangerous-But-Possible Traffic Scenarios SaFeR: 让AI生成危险但可行的交通场景
- SELDON: Fast Supernova Forecasting with Neural ODEs for the Era of 10 Million Alerts Per Night SELDON: 用神经常微分方程预测超新星——应对每晚千万级警报的时代
- TumorFlow: Synthesizing Brain Tumor Growth with Physics-Guided Generative Models TumorFlow:用物理引导的生成模型合成脑肿瘤生长过程
- ZipMap: Compressing 3D Reconstruction from Quadratic to Linear Time ZipMap:将3D重建从二次时间压缩到线性时间
- Beyond Language Modeling: How Vision and Language Scale Differently in Multimodal Pretraining 超越语言建模:视觉与语言在多模态预训练中的不对称缩放规律
- Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation 超越任务完成:通过过程感知评估揭示LLM智能体的虚假成功
- When LLMs Switch Mid-Conversation: The Hidden Cost of Model Handoffs 大模型对话中途换人:模型切换的隐性代价
- Inherited Goal Drift: When Strong AI Agents Copy Bad Habits from Weak Ones 继承性目标漂移:强大的AI智能体会从弱智能体那里学坏
- Odin: Teaching Knowledge Graphs to Explore Themselves Odin:让知识图谱学会自主探索
- Speculative Speculative Decoding: Parallelizing the Parallelizer 投机式投机解码:给并行加速器再加速
- TAO-Attack: Two-Stage Optimization for Breaking LLM Safety Guardrails TAO-Attack: 双阶段优化突破大模型安全防线
- ULTRA: Teaching Humanoids to Act on Intent, Not Just Mimic Moves ULTRA: 让人形机器人理解意图而非模仿动作
- Utonia: One Encoder to Rule All Point Clouds Utonia:一个编码器统治所有点云
- Why Adam Beats SGD: The Math Behind Second-Moment Normalization Adam 为何胜过 SGD:二阶矩归一化的数学原理
- Teaching Multimodal Models to Know When They're Wrong 教会多模态模型识别自己的错误
- CHIMERA: Teaching Small Models to Reason with 9K Synthetic Examples CHIMERA: 用9千条合成数据让小模型学会推理
- Conformal Policy Control: Safe Exploration with Probabilistic Guardrails 保形策略控制:用概率护栏实现安全探索
- Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization 曲率加权容量分配:基于最小描述长度的层自适应大语言模型优化框架
- The Leaderboard Illusion: Why 93% of Top AV Perception Models Aren't Production-Ready 排行榜幻觉:为何93%的顶级自动驾驶感知模型无法投产
- Frontier Models Can Take Actions at Low Probabilities 前沿模型能以极低概率执行动作
- Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact 有知识无智慧:大语言模型与预期影响的错位测量
- KVSlimmer: Why Keys Cluster but Values Don't—and How to Compress Accordingly KVSlimmer: 为什么键会聚类而值不会——以及如何据此压缩
- Multi-Head Low-Rank Attention: Making Compressed Attention Actually Parallelizable 多头低秩注意力:让压缩注意力真正可并行
- AgentSkillOS: Managing AI Agent Skills Like an Operating System AgentSkillOS:像操作系统一样管理AI智能体技能
- pySpatial: Teaching Language Models to Think in 3D Through Code pySpatial: 让语言模型通过代码学会三维思考
- Reasoning Core: Teaching Language Models to Think Through Procedural Symbolic Data Reasoning Core:用程序化符号数据教会语言模型推理
- SageBwd: Making INT8 Attention Work for Pre-training, Not Just Fine-tuning SageBwd: 让INT8注意力机制在预训练中也能用,不只是微调
- Teaching Neural Networks That Red=Blue=Green: Symbol-Equivariant Reasoning 教神经网络懂得红=蓝=绿:符号等变推理模型
- Tool Verification for Test-Time Reinforcement Learning: Stopping AI Models from Confidently Learning the Wrong Thing 测试时工具验证强化学习:阻止AI模型自信地学错
- Notes: Head of Claude Code on What Happens After Coding Is Solved 笔记:Claude Code 负责人谈编程被解决之后的世界
- Notes: AI Is Critical for Humanity's Survival - Cisco President Jeetu Patel 笔记:AI对人类生存至关重要 - Cisco总裁Jeetu Patel
- Notes: The Design Process Is Dead - What's Replacing It 笔记:设计流程已死——取而代之的是什么
- A Minimal Agent for Automated Theorem Proving 自动定理证明的极简智能体
- Why Neural Networks Need Linear, Orthogonal Representations to Generalize Compositionally 神经网络为何需要线性正交表征才能组合泛化
- Controllable Reasoning Models Are Private Thinkers 可控推理模型是隐私保护的思考者
- CUDA Agent: Teaching AI to Write High-Performance GPU Code Through Reinforcement Learning CUDA Agent: 用强化学习教会AI编写高性能GPU代码
- DARE-bench: A Ground-Truth Benchmark for Data Science Instruction Following DARE-bench:数据科学指令遵循的真值基准
- Do LLMs Benefit From Their Own Words? 大模型需要记住自己说过的话吗?
- Efficient Discovery of Approximate Causal Abstractions via Neural Mechanism Sparsification 通过神经机制稀疏化高效发现近似因果抽象
- Enhancing Spatial Understanding in Image Generation via Reward Modeling 通过奖励建模增强图像生成中的空间理解能力
- Memory Caching: Giving RNNs a Growing Memory Without the Quadratic Cost 记忆缓存:让循环网络拥有可增长记忆而不付出二次代价
- Mode Seeking meets Mean Seeking for Fast Long Video Generation 模式寻求遇上均值寻求:快速长视频生成的双轨策略
- SafeGen-LLM: Enhancing Safety Generalization in Task Planning for Robotic Systems SafeGen-LLM: 增强机器人系统任务规划中的安全泛化能力
- SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Caching SenCache: 基于敏感度感知缓存加速扩散模型推理
- Taming Momentum: Rethinking Optimizer States Through Low-Rank Approximation 驯服动量:用低秩近似重新思考优化器状态
- UFO-4D: Reconstructing Dynamic 3D Scenes from Two Unposed Images UFO-4D: 从两张无位姿图像重建动态三维场景
- UMPIRE: Uncertainty Quantification for Multimodal Large Language Models with Incoherence-adjusted Semantic Volume UMPIRE: 基于不一致性调整语义体积的多模态大语言模型不确定性量化
- Runtime-Reconfigurable Bitwise Systolic Arrays for Multi-Precision Neural Network Acceleration 面向多精度神经网络加速的运行时可重构位级脉动阵列架构
- Deep Ensemble Graph Neural Networks for Probabilistic Cosmic-Ray Direction and Energy Reconstruction in Autonomous Radio Arrays 基于深度集成图神经网络的自主射电阵列宇宙线方向与能量概率重建
- Memory-Efficient MCTS: Extending GRAVE for Resource-Constrained Game Playing 内存高效的蒙特卡洛树搜索:面向资源受限环境的GRAVE算法扩展
- Mean Estimation from Coarse Data: Characterizations and Efficient Algorithms 粗粒度数据的均值估计:可识别性刻画与高效算法
- Model Agreement via Anchoring: A Unified Framework for Reducing Prediction Disagreement 通过锚定实现模型一致性:减少预测分歧的统一框架
- Retrieve and Segment: Bridging the Supervision Gap in Open-Vocabulary Segmentation with Few-Shot Learning 检索与分割:少样本学习能否弥合开放词汇分割中的监督差距?
- Sensor Generalization for Adaptive Sensing in Event-based Object Detection via Joint Distribution Training 通过联合分布训练实现事件相机目标检测的传感器泛化与自适应感知
- Toward Expert Investment Teams: A Multi-Agent LLM System with Fine-Grained Trading Tasks 构建专家投资团队:基于细粒度交易任务的多智能体大语言模型系统
- Understanding Usage and Engagement in AI-Powered Scientific Research Tools: The Asta Interaction Dataset 理解AI驱动科研工具的使用与参与模式:Asta交互数据集
- Utilizing LLMs for Industrial Process Automation 利用大语言模型实现工业过程自动化
- A Dataset is Worth 1 MB: Efficient Dataset Distribution via Pseudo-Labels 数据集仅需1MB:通过伪标签实现高效数据集分发
- AIQI: The First Model-Free Universal AI Agent AIQI:首个无模型通用AI智能体
- Agency and Architectural Limits: Why Optimization-Based Systems Cannot Be Norm-Responsive 主体性与架构限制:为何基于优化的系统无法响应规范
- AgentDropoutV2: Optimizing Information Flow in Multi-Agent Systems via Test-Time Rectify-or-Reject Pruning AgentDropoutV2: 通过测试时修正-拒绝剪枝优化多智能体系统信息流
- Differentiable Zero-One Loss via Hypersimplex Projections 通过超单纯形投影实现可微分的零一损失
- DDTSR: Enabling Human-Like Responsiveness in Spoken Dialogue Through Dual-Track Streaming DDTSR:通过双轨流式响应实现类人对话反应速度
- FlashOptim: Cutting Neural Network Training Memory by Over 50% FlashOptim: 将神经网络训练内存减少超过50%
- LLM Novice Uplift on Dual-Use, In Silico Biology Tasks 大语言模型对新手在双重用途生物学任务中的能力提升研究
- MediX-R1: Open-Ended Reinforcement Learning for Medical Multimodal Models MediX-R1: 面向医疗多模态模型的开放式强化学习
- ParamMem: Augmenting Language Agents with Parametric Reflective Memory ParamMem: 通过参数化反思记忆增强语言智能体
- Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning 规模无法克服语用学问题:报告偏差对视觉-语言推理的影响
- SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation SeeThrough3D: 文本到图像生成中的遮挡感知3D控制
- SOTAlign: Semi-Supervised Alignment of Vision and Language Models via Optimal Transport SOTAlign: 通过最优传输实现视觉与语言模型的半监督对齐
- VGG-T³: Scaling 3D Reconstruction to Thousands of Images with Linear Complexity VGG-T³: 线性复杂度下的大规模三维重建
- Why Diffusion Language Models Struggle with Truly Parallel (Non-Autoregressive) Decoding? 扩散语言模型为何难以实现真正的并行(非自回归)解码?
- CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation CoLoGen: 统一图像生成中概念-定位二元性的渐进式学习
- DySCO: Dynamic Attention-Scaling Decoding for Long-Context Language Models DySCO: 面向长文本语言模型的动态注意力缩放解码
- Enhancing LLM-Based Test Generation by Eliminating Covered Code 通过消除已覆盖代码增强基于大语言模型的测试生成
- GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL GUI-Libra: 通过动作感知监督和部分可验证强化学习训练原生GUI智能体进行推理与行动
- Improving Parametric Knowledge Access in Reasoning Language Models 提升推理语言模型的参数化知识访问能力
- NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors NoLan: 通过动态抑制语言先验缓解大型视觉-语言模型中的物体幻觉
- Off-The-Shelf Image-to-Image Models Are All You Need To Defeat Image Protection Schemes 现成的图像转换模型足以破解图像保护方案
- PanoEnv: Teaching Vision-Language Models 3D Spatial Reasoning in Panoramic Environments PanoEnv: 通过强化学习在全景环境中探索3D空间智能
- Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets 翻译中的恢复:基准测试和数据集自动化翻译的高效流程
- RobustVisRAG: Causality-Aware Vision-Based Retrieval-Augmented Generation under Visual Degradations RobustVisRAG: 视觉退化条件下的因果感知视觉检索增强生成
- Solaris: Building a Multiplayer Video World Model in Minecraft Solaris:在Minecraft中构建多人视频世界模型
- When AI Writes, Whose Voice Remains? Quantifying Cultural Marker Erasure in LLMs 当AI代笔时,谁的声音还在?量化大语言模型中的文化标记消除现象
- When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign Adapters 当LoRA背叛信任:通过伪装成良性适配器对文本到图像模型植入后门
- WHOLE: Holistic Hand-Object Reconstruction from Egocentric Videos WHOLE: 从第一人称视频中整体重建手部与物体运动
- World Guidance: World Modeling in Condition Space for Action Generation 世界引导:条件空间中的世界建模用于动作生成
- Aletheia Tackles FirstProof Autonomously: AI Agent Solves 6 of 10 Mathematical Research Problems Aletheia自主挑战FirstProof:AI智能体解决10道数学研究问题中的6道
- Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs 从试错中学习:具身大语言模型的反思式测试时规划
- LogicGraph: Benchmarking Multi-Path Logical Reasoning via Neuro-Symbolic Generation and Verification LogicGraph: 通过神经符号生成与验证的多路径逻辑推理基准
- Motivation is Something You Need: A Neuroscience-Inspired Dual-Model Training Framework 动机驱动训练:受神经科学启发的双模型训练框架
- NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning NoRD: 无需推理的数据高效视觉-语言-动作自动驾驶模型
- CLIPGlasses: Teaching CLIP to Understand 'Not' Without Retraining CLIPGlasses: 让CLIP理解否定语义的即插即用框架
- On Data Engineering for Scaling LLM Terminal Capabilities 大语言模型终端能力扩展的数据工程研究
- PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A PaperTrail: 基于声明-证据映射的学术问答系统溯源界面
- Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning Spa3R: 用于3D视觉推理的预测性空间场建模
- Squint: Fast Visual Reinforcement Learning for Sim-to-Real Robotics Squint: 面向仿真到真实机器人的快速视觉强化学习
- Test-Time Training with KV Binding Is Secretly Linear Attention 基于键值绑定的测试时训练本质上是线性注意力机制
- Tool Building as a Path to Superintelligence 工具构建:通向超级智能的路径
- UPipe: Breaking the Memory Barrier in Long-Context Transformer Training with Headwise Chunking UPipe: 通过注意力头级分块突破长上下文Transformer训练的内存瓶颈
- Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training 为什么Pass@k优化会降低Pass@1性能:大语言模型后训练中的提示干扰现象
- VBVR: A Million-Scale Dataset for Video Reasoning at Scale VBVR:百万级视频推理数据集与基准测试
- Anatomy of Agentic Memory: Taxonomy and Empirical Analysis of Evaluation and System Limitations 智能体记忆系统解剖:评估与系统局限性的分类与实证分析
- City Editing: Hierarchical Agentic Execution for Dependency-Aware Urban Geospatial Modification 城市编辑:面向依赖感知的城市地理空间修改的层次化智能体执行框架
- DefenseSplat: Enhancing the Robustness of 3D Gaussian Splatting via Frequency-Aware Filtering DefenseSplat: 通过频率感知滤波增强3D高斯溅射的鲁棒性
- LAD: Learning Advantage Distribution for Reasoning LAD: 学习优势分布以增强推理能力
- Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device Mobile-O: 移动设备上的统一多模态理解与生成
- NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning NovaPlan: 通过闭环视频语言规划实现零样本长时域机器人操作
- Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks Skill-Inject: 测量智能体对技能文件攻击的脆弱性
- TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics TOPReward: 将Token概率作为机器人零样本奖励的隐藏信号
- Training-Free Cross-Architecture Merging for Graph Neural Networks 图神经网络的免训练跨架构融合
- Benchmarking Graph Neural Networks in Solving Hard Constraint Satisfaction Problems 图神经网络求解困难约束满足问题的基准测试研究
- CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation CapNav: 基于能力条件的室内导航视觉语言模型基准测试
- CoPeDiT: Self-Perceptive Diffusion Transformer for Unified 3D MRI Synthesis CoPeDiT: 基于自感知扩散Transformer的统一3D MRI合成方法
- Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control 生成现实:基于手部和相机控制的以人为中心的交互式视频生成世界模拟
- MemStream: Scaling Token Budgets for Enhanced Video Stream Understanding with Dynamic KV-Cache MemStream: 通过动态键值缓存扩展令牌预算以增强视频流理解
- Latent Equivariant Operators for Robust Object Recognition: Promise and Challenges 用于鲁棒目标识别的潜在等变算子:前景与挑战
- PRISM-FCP: Byzantine-Resilient Federated Conformal Prediction via Partial Sharing PRISM-FCP: 通过部分共享实现拜占庭容错的联邦保形预测
- RVR: Retrieve-Verify-Retrieve for Comprehensive Question Answering RVR:用于全面问答的检索-验证-检索框架
- SARAH: Spatially Aware Real-time Agentic Humans SARAH: 空间感知的实时智能体人类
- Self-Aware Object Detection via Degradation Manifolds 基于退化流形的自感知目标检测
- SPQ: An Ensemble Technique for Large Language Model Compression SPQ: 大语言模型压缩的集成技术
- Subgroups of U(d) Induce Natural RNN and Transformer Architectures U(d)的子群诱导自然的RNN和Transformer架构
- The Geometry of Noise: Why Diffusion Models Don't Need Noise Conditioning 噪声的几何学:扩散模型为何不需要噪声条件
- Unifying Approach to Uniform Expressivity of Graph Neural Networks 图神经网络统一表达能力的统一方法
- VIRAASAT: Advancing Cultural Reasoning in LLMs Through Multi-Hop Knowledge Graphs VIRAASAT: 通过多跳知识图谱推进大语言模型的文化推理能力
- CLEF HIPE-2026: Evaluating Person-Place Relation Extraction from Multilingual Historical Texts CLEF HIPE-2026: 多语言历史文本中人物-地点关系抽取评测
- CORAL: Correspondence Alignment for Improved Virtual Try-On CORAL:通过对应关系对齐改进虚拟试衣效果
- Language Models Show Selective Typological Alignment in Differential Argument Marking 语言模型在差异论元标记中展现选择性类型学对齐
- IntRec: Intent-based Retrieval with Contrastive Refinement IntRec: 基于意图的对比细化检索框架
- Unmasking the Factual-Conceptual Gap in Persian Language Models 揭示波斯语语言模型中的事实-概念鸿沟
- A.R.I.S.: Deep Learning-Powered E-Waste Classification for Automated Recycling A.R.I.S.: 基于深度学习的电子废物自动分类回收系统
- OSI-FL: One-Shot Incremental Federated Learning with Catastrophic Forgetting Resilience OSI-FL: 具有灾难性遗忘弹性的单轮增量联邦学习
- FAMOSE: A ReAct Approach to Automated Feature Discovery FAMOSE: 基于ReAct范式的自动化特征发现方法
- Human-level 3D Shape Perception Emerges from Multi-view Learning 通过多视角学习实现人类水平的三维形状感知
- MARS: Margin-Aware Reward-Modeling with Self-Refinement MARS:基于边际感知的自我精炼奖励建模
- Mine and Refine: Optimizing Graded Relevance in E-commerce Search Retrieval 挖掘与精炼:电商搜索检索中的分级相关性优化
- Multi-Round Human-AI Collaboration with User-Specified Requirements 基于用户指定需求的多轮人机协作
- OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents OpenEarthAgent: 工具增强型地理空间智能体统一框架
- Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting 通过细粒度细节定向推进黑盒大型视觉语言模型攻击前沿
- Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting Reverso: 用于零样本预测的高效时间序列基础模型
- Sink-Aware Pruning for Diffusion Language Models 扩散语言模型的注意力汇聚感知剪枝
- SMAC: Score-Matched Actor-Critics for Robust Offline-to-Online Transfer SMAC: 基于分数匹配的演员-评论家算法实现稳健的离线到在线迁移
- UniLID: Tokenizer-Based Language Identification for Low-Resource and Dialect Settings UniLID: 基于分词器的低资源和方言语言识别方法
- When to Trust the Cheap Check: Weak and Strong Verification for Reasoning 何时信任廉价检查:推理中的弱验证与强验证
- When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs 当视觉凌驾于语言之上:评估和缓解视觉-语言-动作模型中的反事实失效
- Gemini 3.1 Pro: Google's Greatest Model Ever — Full Review & Testing Gemini 3.1 Pro:谷歌有史以来最强模型——完整评测
- Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents 先校准再行动:大语言模型智能体中的成本感知探索
- Causality is Key for Interpretability Claims to Generalise 因果性是可解释性主张泛化的关键
- HERO: Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation HERO: 面向开放词汇视觉移动操作的人形机器人末端执行器控制学习
- SAW-Bench: Evaluating Egocentric Situated Awareness in Multimodal Foundation Models SAW-Bench:评估多模态基础模型的第一人称情境感知能力
- Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology 测量2025年中期大语言模型对生物学新手实验表现的辅助效果
- PCAS: Policy Compiler for Secure Agentic Systems PCAS:面向安全智能体系统的策略编译器
- Protecting the Undeleted in Machine Unlearning 机器遗忘中对未删除数据的保护
- REFINE: Reinforcement Learning Framework for Fast Weight Models with Next-Sequence Prediction REFINE:基于下一序列预测的快速权重模型强化学习框架
- Saliency-Aware Multi-Route Thinking: Revisiting Vision-Language Reasoning 显著性感知多路径思考:重新审视视觉-语言推理
- SODA: Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens SODA: 通过交错语义、声学和文本标记扩展开放离散音频基础模型
- Frisch & Tinbergen (1969): Founding Econometrics and Dynamic Economic Models 弗里希与丁伯根(1969):创立计量经济学,建立经济过程的动态模型
- Samuelson (1970): Raising Economics to a Science with Mathematics 萨缪尔森(1970):用数学将经济学提升为科学
- Kuznets (1971): Economic Growth, National Income, and the Inverted U 库兹涅茨(1971):经济增长、国民收入与倒U型曲线
- Hicks & Arrow (1972): General Equilibrium and the Limits of Social Choice 希克斯与阿罗(1972):一般均衡与社会选择的边界
- Leontief (1973): Input-Output Analysis and the X-Ray of the Economy 列昂惕夫(1973):投入产出分析——给经济拍X光
- Myrdal & Hayek (1974): Money, Fluctuations, and the Clash of Two Worldviews 缪尔达尔与哈耶克(1974):货币、波动与两种世界观的碰撞
- Kantorovich & Koopmans (1975): Linear Programming and the Optimal Allocation of Resources 康托罗维奇与库普曼斯(1975):线性规划与资源最优分配
- Friedman (1976): Monetarism, Consumption Analysis, and the Complexity of Stabilization 弗里德曼(1976):货币主义、消费分析与稳定政策的复杂性
- Ohlin & Meade (1977): Why Nations Trade and How Capital Flows 俄林与米德(1977):国家为何贸易,资本如何流动
- Simon (1978): Bounded Rationality and Decision-Making in Organizations 西蒙(1978):有限理性与组织中的决策过程
- Schultz & Lewis (1979): Development Economics, Human Capital, and the Dual Economy 舒尔茨与刘易斯(1979):发展经济学、人力资本与二元经济
- Klein (1980): Macroeconometric Models and the Art of Economic Forecasting 克莱因(1980):宏观计量经济模型与经济预测的艺术
- Tobin (1981): Financial Markets, Tobin's Q, and the Real Economy 托宾(1981):金融市场、托宾Q理论与实体经济
- Stigler (1982): Industrial Structure, Information, and the Capture of Regulation 斯蒂格勒(1982):产业结构、信息与监管俘获
- Debreu (1983): Mathematical Proof That Markets Can Work 德布鲁(1983):市场可以运作的数学证明
- Stone (1984): The Architect of National Accounts 斯通(1984):国民核算体系的建筑师
- Modigliani (1985): The Life Cycle of Saving and the Irrelevance of Capital Structure 莫迪利安尼(1985):储蓄的生命周期与资本结构无关性
- Buchanan (1986): Public Choice Theory — Economics Meets Politics 布坎南(1986):公共选择理论——当经济学遇见政治学
- Solow (1987): Why Economies Grow — And Why Capital Alone Isn't Enough 索洛(1987):经济为何增长——为什么仅靠资本远远不够
- Allais (1988): Market Efficiency and the Paradox That Broke Expected Utility 阿莱(1988):市场效率与击破期望效用的悖论
- Haavelmo (1989): Giving Econometrics a Rigorous Foundation 哈维尔莫(1989):为计量经济学奠定严谨基础
- Markowitz, Miller & Sharpe (1990): The Birth of Modern Finance 马科维茨、米勒、夏普(1990):现代金融学的诞生
- Coase (1991): Transaction Costs, Property Rights, and Why Firms Exist 科斯(1991):交易成本、产权与企业为何存在
- Becker (1992): The Economic Approach to Everything Human 贝克尔(1992):用经济学解释一切人类行为
- Fogel & North (1993): When Economics Met History 福格尔、诺斯(1993):当经济学遇见历史学
- Harsanyi, Nash & Selten (1994): The Mathematics of Strategic Thinking 海萨尼、纳什、泽尔腾(1994):战略思维的数学
- Lucas (1995): Rational Expectations and the Death of Old Macroeconomics 卢卡斯(1995):理性预期与旧宏观经济学的终结
- Mirrlees & Vickrey (1996): Designing Incentives When Information Is Hidden 莫里斯、维克里(1996):信息隐藏时如何设计激励
- Merton & Scholes (1997): The Formula That Priced the Future 默顿、斯科尔斯(1997):为未来定价的公式
- Sen (1998): Economics with a Human Face 森(1998):有人性面孔的经济学
- Mundell (1999): The Impossible Trinity and the Intellectual Father of the Euro 蒙代尔(1999):不可能三角与欧元的思想之父
- Heckman & McFadden (2000): The Science of Individual Choices and Hidden Biases 赫克曼、麦克法登(2000):个体选择与隐藏偏差的科学
- Akerlof, Spence & Stiglitz (2001): When One Side Knows More Than the Other 阿克洛夫、斯彭斯、斯蒂格利茨(2001):当一方比另一方知道更多
- Kahneman & Smith (2002): The Mind vs. The Market 卡尼曼、史密斯(2002):心智对决市场
- Engle & Granger (2003): Taming Time Series — Volatility and Cointegration 恩格尔、格兰杰(2003):驯服时间序列——波动率与协整
- Kydland & Prescott (2004): Why Good Intentions Make Bad Policy, and Why Recessions Might Be Efficient 基德兰德、普雷斯科特(2004):为什么好意图造成坏政策,为什么衰退可能是有效率的
- Aumann & Schelling (2005): The Game Theory of War and Peace 奥曼、谢林(2005):战争与和平的博弈论
- Phelps (2006): Why You Can't Buy Low Unemployment with High Inflation Forever 菲尔普斯(2006):为什么不能永远用高通胀换取低失业
- Hurwicz, Maskin & Myerson (2007): Engineering Institutions That Make People Tell the Truth 赫维茨、马斯金、迈尔森(2007):设计让人说真话的制度
- Krugman (2008): Why Countries That Make the Same Things Trade With Each Other 克鲁格曼(2008):为什么生产相同东西的国家互相贸易
- Ostrom & Williamson (2009): Beyond Markets and Governments — How Humans Actually Organize 奥斯特罗姆、威廉姆森(2009):超越市场与政府——人类实际如何组织
- Diamond, Mortensen & Pissarides (2010): Why Jobs Go Unfilled While People Go Unemployed 戴蒙德、莫滕森、皮萨里德斯(2010):为什么职位空缺与失业并存
- Sargent & Sims (2011): Untangling Cause and Effect in the Macroeconomy 萨金特、西姆斯(2011):解开宏观经济中的因果关系
- Roth & Shapley (2012): The Economics of Matchmaking — From Marriage to Kidneys 罗斯、沙普利(2012):配对经济学——从婚姻到肾脏
- Fama, Hansen & Shiller (2013): Can You Beat the Market? It Depends on Your Time Horizon 法玛、汉森、席勒(2013):你能战胜市场吗?取决于你的时间跨度
- Tirole (2014): How to Tame Powerful Firms Without Breaking the Market 梯若尔(2014):如何驯服强大的企业而不破坏市场
- Deaton (2015): What People Buy Tells You Everything About How They Live 迪顿(2015):人们买什么告诉你他们如何生活的一切
- Hart & Holmström (2016): The Science of Writing Better Contracts 哈特、霍尔姆斯特伦(2016):编写更好合同的科学
- Thaler (2017): Nudging Humans Toward Better Choices Without Taking Away Their Freedom 塞勒(2017):在不剥夺自由的情况下助推人类做出更好的选择
- Nordhaus & Romer (2018): Growth's Greatest Gift and Greatest Threat 诺德豪斯、罗默(2018):增长最大的礼物与最大的威胁
- Banerjee, Duflo & Kremer (2019): Fighting Poverty One Experiment at a Time 班纳吉、迪弗洛、克雷默(2019):一次一个实验地对抗贫困
- Milgrom & Wilson (2020): The Science of Selling Things Nobody Knows the Value Of 米尔格罗姆、威尔逊(2020):拍卖无人知晓价值之物的科学
- Card, Angrist & Imbens (2021): When Life Runs the Experiment for You 卡德、安格里斯特、因本斯(2021):当生活替你做实验
- Bernanke, Diamond & Dybvig (2022): Why Banks Fail and Why It Matters 伯南克、戴蒙德、迪布维格(2022):银行为什么会倒闭以及为什么重要
- Goldin (2023): Two Centuries of Data on Why Women Still Earn Less 戈尔丁(2023):两个世纪的数据揭示女性为何仍然收入较低
- Acemoglu, Johnson & Robinson (2024): Why Some Nations Are Rich and Others Are Poor 阿西莫格鲁、约翰逊、罗宾逊(2024):为什么有些国家富裕而另一些贫穷
- Mokyr, Aghion & Howitt (2025): Why Growth Keeps Going — and Why It Might Stop 莫基尔、阿吉翁、霍伊特(2025):增长为何持续——以及为何可能停止
- Avey-B: An Attention-Free Bidirectional Encoder for Efficient NLP Avey-B: 面向高效自然语言处理的无注意力双向编码器
- CrispEdit: Low-Curvature Projections for Scalable Non-Destructive LLM Editing CrispEdit: 基于低曲率投影的可扩展无损大语言模型编辑
- Developing AI Agents with Simulated Data: Why, what, and how? 使用模拟数据开发AI智能体:为什么、是什么、怎么做?
- Dex4D: Task-Agnostic Point Track Policy for Sim-to-Real Dexterous Manipulation Dex4D: 面向仿真到真实迁移的任务无关点轨迹策略用于灵巧操作
- Neural Scaling Laws for Boosted Jet Tagging 增强喷注标记的神经网络缩放定律
- Operationalising the Superficial Alignment Hypothesis via Task Complexity 通过任务复杂度量化表层对齐假设
- Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching 感知型人形机器人跑酷:通过动作匹配链接动态人类技能
- The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety 对齐崩溃的几何学:微调如何破坏安全性
- Validating LLM Simulations as Behavioral Evidence: When Can AI Replace Human Subjects? 验证大语言模型模拟作为行为证据:何时AI可以替代人类受试者?
- VideoSketcher: Leveraging Video Diffusion Models for Sequential Sketch Generation VideoSketcher: 利用视频扩散模型实现序列化草图生成
- Dario Amodei: The Highest-Stakes Financial Model in History Dario Amodei:史上最高风险的财务模型
- OpenClaw: The Viral AI Agent that Broke the Internet OpenClaw:席卷互联网的病毒式AI智能体
- Claude Sonnet 4.6: Near Opus-Level Intelligence at Half the Price Claude Sonnet 4.6:以一半价格获得接近Opus级别的智能
- Boundary Point Jailbreaking: Breaking Black-Box LLM Safeguards with Binary Feedback 边界点越狱:仅用二进制反馈突破黑盒大语言模型防护
- BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames BPP: 通过关注关键历史帧实现长上下文机器人模仿学习
- Deep Dense Exploration for LLM Reinforcement Learning via Pivot-Driven Resampling 基于枢轴驱动重采样的大语言模型强化学习深度密集探索
- EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing EditCtrl: 解耦局部与全局控制的实时生成式视频编辑
- ForesightSafety Bench: A Comprehensive Framework for Evaluating Frontier AI Risks ForesightSafety Bench: 面向安全AI的前沿风险评估与治理框架
- Hunt Globally: Deep Research AI Agents for Drug Asset Scouting in Investing, Business Development, and Search & Evaluation 全球搜寻:用于投资、商业拓展和搜索评估的药物资产侦察深度研究AI智能体
- Sphere Encoder: Single-Pass Image Generation via Spherical Latent Space 球面编码器:通过球形潜在空间实现单次前向传播图像生成
- Knowing When Not to Answer: Abstention-Aware Scientific Reasoning 知道何时不回答:具有弃权意识的科学推理
- Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization 长上下文,弱聚焦:大语言模型在隐私与个性化中暴露的扩展性缺陷
- Rethinking Diffusion Models with Symmetries through Canonicalization for Molecular Graph Generation 通过规范化重新思考对称性扩散模型:分子图生成的新范式
- Scaling Beyond Masked Diffusion Language Models 超越掩码扩散语言模型的规模化研究
- Symmetry in Language Statistics Shapes the Geometry of Model Representations 语言统计中的对称性塑造模型表征的几何结构
- Synergistic Intra- and Cross-Layer Regularization Losses for MoE Expert Specialization 协同层内与跨层正则化损失促进MoE专家特化
- ThermEval: A Structured Benchmark for Evaluating Vision-Language Models on Thermal Imagery ThermEval: 热成像视觉语言模型结构化评估基准
- When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift 当基准测试说谎时:真实分布偏移下的恶意提示词分类器评估
- Krites: Asynchronous Verified Semantic Caching for Tiered LLM Architectures Krites:面向分层LLM架构的异步验证语义缓存
- Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision 对话式图像分割:通过可扩展监督实现抽象概念的像素级定位
- CoPE-VideoLM: Codec Primitives For Efficient Video Language Models CoPE-VideoLM: 基于编解码器原语的高效视频语言模型
- FlashSchNet: Fast and Accurate Coarse-Grained Neural Network Molecular Dynamics FlashSchNet: 快速精确的粗粒化神经网络分子动力学
- FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control FlexAM: 灵活的外观-运动解耦实现多功能视频生成控制
- PSI: Learning Robot Manipulation from Human Videos with Simulation-Filtered Grasping PSI: 通过仿真过滤抓取从人类视频学习机器人操作
- In-Context Autonomous Network Incident Response: An End-to-End Large Language Model Agent Approach 上下文自主网络事件响应:端到端大语言模型智能体方法
- LongStream: Enabling Kilometer-Scale 3D Reconstruction with Gauge-Decoupled Streaming LongStream: 基于规范解耦流式处理的公里级3D重建
- Quantization-Robust LLM Unlearning via Low-Rank Adaptation 基于低秩适配的量化鲁棒大语言模型遗忘
- Semantic Chunking and the Entropy of Natural Language 语义分块与自然语言的熵
- Teaching AI Fundamentals to Translation Professionals: A Technical Curriculum for Language-Oriented AI 面向翻译专业人员的AI基础教学:语言导向人工智能技术课程
- Is Online Linear Optimization Sufficient for Strategic Robustness? 在线线性优化是否足以实现策略鲁棒性?
- iUzawa-Net: Real-Time Neural Network Solver for Nonsmooth PDE Optimal Control iUzawa-Net: 非光滑偏微分方程最优控制的实时神经网络求解器
- Implicit Regularization in Langevin Dynamics with Projected Noise 投影噪声朗之万动力学中的隐式正则化
- Flow-Guided Neural Operators: Dynamic Noise Levels for Self-Supervised Time-Series Learning 流引导神经算子:时间序列自监督学习中的动态噪声水平
- AttentionRetriever: Attention Layers are Secretly Long Document Retrievers AttentionRetriever: 注意力层是隐藏的长文档检索器
- Community Concealment from Unsupervised Graph Learning-Based Clustering 基于无监督图学习聚类的社区隐藏技术
- Creative Ownership in the Age of AI: Rethinking Copyright Infringement AI时代的创作所有权:重新思考版权侵权
- ExtractBench: Exposing the Limits of LLMs in Complex Structured Data Extraction ExtractBench: 揭示大语言模型在复杂结构化数据提取中的局限性
- Function-Space Decoupled Diffusion for Forward and Inverse Modeling in Carbon Capture and Storage 函数空间解耦扩散模型用于碳捕获与封存的正向和反向建模
- On-Policy Context Distillation for Language Models 语言模型的在线策略上下文蒸馏
- Stroke of Surprise: Progressive Semantic Illusions in Vector Sketching 惊喜之笔:矢量素描中的渐进式语义错觉
- T3D: Accelerating Diffusion Language Models with Trajectory Self-Distillation T3D: 通过轨迹自蒸馏加速扩散语言模型
- KeplerAgent: Physics-Guided LLM Agent for Scientific Equation Discovery KeplerAgent: 物理引导的大语言模型智能体用于科学方程发现
- UniT: Unified Multimodal Chain-of-Thought Test-time Scaling UniT: 统一多模态思维链测试时扩展
- CATTS: Confidence-Aware Test-Time Scaling for Multi-Step Web Agents CATTS: 面向多步骤网络智能体的置信度感知测试时扩展
- Capability-Oriented Training Induced Alignment Risk 能力导向训练引发的对齐风险
- CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use CM2: 基于检查清单奖励的多轮多步骤智能体工具使用强化学习
- DeepSight: An All-in-One LM Safety Toolkit DeepSight: 大模型安全一体化工具包
- GigaBrain-0.5M*: Advancing Vision-Language-Action Models Through World Model-Based Reinforcement Learning GigaBrain-0.5M*: 通过世界模型强化学习推进视觉-语言-动作模型
- MonarchRT: Efficient Attention for Real-Time Video Generation MonarchRT: 实时视频生成的高效注意力机制
- Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment 扩展验证比扩展策略学习更有效:视觉-语言-动作对齐的新范式
- When Speech Recognition Fails Where It Matters: The Hidden Cost of Transcription Errors 语音识别在关键场景下的失效:转录错误的隐藏代价
- Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penalty 停止不必要的反思:通过自适应反思和长度协调惩罚训练高效推理的大型推理模型
- The Pensieve Paradigm: Stateful Language Models Mastering Their Own Context 冥想盆范式:具有状态管理能力的语言模型
- Asymmetric Prompt Weighting for Reinforcement Learning with Verifiable Rewards 基于可验证奖励的强化学习中的非对称提示加权
- Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning 数据重复优于数据扩展:长链思维监督微调的新发现
- Diffusion-Pretrained Dense and Contextual Embeddings for Web-Scale Retrieval 基于扩散预训练的密集与上下文嵌入模型用于网络规模检索
- FormalJudge: A Neuro-Symbolic Paradigm for Agentic Oversight FormalJudge: 智能体监督的神经符号范式
- From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design 从缓冲区到寄存器:通过混合键合3D NPU协同设计解锁细粒度FlashAttention
- From Circuits to Dynamics: Understanding and Stabilizing Failure in 3D Diffusion Transformers 从电路到动力学:理解并稳定3D扩散Transformer的失效模式
- MKNA: Language-Driven Agent for Autonomous Materials Discovery MKNA: 语言驱动的自主材料发现智能体
- GENIUS: Evaluating Generative Fluid Intelligence in Multimodal Models GENIUS: 生成式流体智能评估基准
- SVDA: Making Vision Transformers Interpretable for Monocular Depth Estimation SVDA:让视觉Transformer在单目深度估计中变得可解释
- LaSSM: Efficient 3D Instance Segmentation with Local Aggregation and State Space Models LaSSM: 基于局部聚合和状态空间模型的高效3D实例分割
- ROCKET: Training-Free Model Compression via Knapsack Optimization and Sparse Factorization ROCKET: 基于背包优化和稀疏分解的免训练模型压缩方法
- TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model TabICLv2: 更优、更快、可扩展的开源表格基础模型
- Weight Decay Improves Language Model Plasticity 权重衰减提升语言模型可塑性
- YOR: Building Affordable Mobile Manipulators for Generalizable Robotics Research YOR:面向通用机器人研究的低成本移动操作平台
- Pretraining with Token-Level Adaptive Latent Chain-of-Thought 使用Token级自适应潜在思维链进行预训练
- Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement Learning 通过端到端强化学习实现压缩记忆上的动态长上下文推理
- Discovering Interpretable Algorithms by Decompiling Transformers to RASP 通过将Transformer反编译为RASP发现可解释算法
- DirMoE: Dirichlet-routed Mixture of Experts DirMoE:基于Dirichlet路由的混合专家模型
- Document Reconstruction Unlocks Scalable Long-Context RLVR 文档重建解锁可扩展的长上下文RLVR
- Understanding Dynamic Compute Allocation in Recurrent Transformers 理解循环Transformer中的动态计算分配
- Emergent Search and Backtracking in Latent Reasoning Models 潜在推理模型中涌现的搜索与回溯
- FlexMoRE: A Flexible Mixture of Rank-heterogeneous Experts for Efficient LLMs FlexMoRE:用于高效LLM的灵活异构秩混合专家
- GSS: Gated Subspace Steering for Selective Memorization Mitigation in LLMs GSS:门控子空间引导实现LLM选择性记忆缓解
- Latent Reasoning with Supervised Thinking States 基于监督思维状态的潜在推理
- Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models 离散潜在空间中的下一概念预测产生更强的语言模型
- Prism: Spectral-Aware Block-Sparse Attention Prism:频谱感知的块稀疏注意力
- QUOKA: Query-Oriented KV Selection For Efficient LLM Prefill QUOKA:面向查询的KV选择实现高效LLM预填充
- New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR 新技能还是更锐利的原语?RLVR中推理涌现的概率视角
- DyTopo: Dynamic Topology Routing for Multi-Agent Reasoning via Semantic Matching DyTopo:基于语义匹配的多智能体推理动态拓扑路由
- Shared LoRA Subspaces for Almost Strict Continual Learning 共享LoRA子空间实现近乎严格的持续学习
- AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions AgenticPay:用于买卖交易的多智能体LLM谈判系统
- CoPE: Clipped RoPE as a Free Lunch for Long Context LLMs CoPE:裁剪RoPE——长上下文LLM的免费午餐
- CORAL: Correctness-Optimized Residual Activation Lens for Inference-Time Steering CORAL:面向正确性优化的残差激活透镜用于推理时引导
- InftyThink+: Infinite-Horizon Reasoning via Reinforcement Learning InftyThink+:通过强化学习实现无限视野推理
- Mining Generalizable Activation Functions via Evolutionary Search 通过进化搜索挖掘可泛化的激活函数
- Multi-Task GRPO: Reliable LLM Reasoning Across Tasks 多任务GRPO:跨任务的可靠LLM推理
- OmniMoE: Efficient Mixture-of-Experts via Atomic Expert Orchestration OmniMoE:通过原子专家编排实现高效混合专家模型
- Pseudo-Invertible Neural Networks: Generalizing Moore-Penrose to Deep Learning 伪可逆神经网络:将Moore-Penrose推广到深度学习
- RWML: Reinforcement World Model Learning for LLM-based Agents RWML:基于LLM智能体的强化世界模型学习
- Revisiting the Shape Convention of Transformer Language Models 重新审视Transformer语言模型的形状约定
- Elon Musk: The Entire Tech Tree Is Converging Elon Musk:整个科技树正在汇聚
- Harvard CS50 (2026) – Full Computer Science University Course 哈佛CS50 (2026) – 完整计算机科学大学课程
- The Key to State Reduction in Linear Attention: A Rank-based Perspective 线性注意力状态压缩的关键:基于秩的视角
- Reinforced Attention Learning: Optimizing Where to Look, Not What to Say 强化注意力学习:优化关注位置而非输出内容
- Rethinking the Trust Region in LLM Reinforcement Learning 重新思考LLM强化学习中的信任域
- TinyLoRA: Learning to Reason in 13 Parameters TinyLoRA:用13个参数学会推理
- Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-Tuning 学习率很重要:原版LoRA可能就够了
- Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models 大语言模型强化微调中的熵动力学
- State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI 2026年AI现状:LLM、编程、缩放定律、中国、智能体、GPU、AGI
- Pipeline Parallelism from Scratch – Building Distributed AI Training 从零构建流水线并行 - 分布式AI训练
- RAG & MCP Fundamentals – A Hands-On Crash Course RAG与MCP基础 - 实战速成课程
- How to Benchmark Embedding Models On Your Own Data 如何在自己的数据上评测嵌入模型
- Building Agentic AI Workloads – Crash Course 构建智能体AI工作负载 - 速成课程
- From RNNs to Transformers: The Complete Neural Machine Translation Journey 从RNN到Transformer:神经机器翻译完整之旅
- Become an AI Researcher – LLM, Math, PyTorch, Neural Networks, Transformers 成为AI研究员 - LLM、数学、PyTorch、神经网络、Transformer
- How AI Will Change Software Engineering – Martin Fowler on The Pragmatic Engineer AI将如何改变软件工程 - Martin Fowler访谈
- From Swift to Mojo: High-Performance AI Engineering with Chris Lattner 从Swift到Mojo:Chris Lattner谈高性能AI工程
- Notes: Transformers Explained - The Discovery That Changed AI Forever 笔记:Transformer 详解 - 永远改变 AI 的发现
- Nick Lane: The Universe Favors Life Nick Lane:宇宙偏爱生命
- Notes: The 7 Most Powerful Moats For AI Startups 笔记:AI 创业公司的 7 大护城河
- Anthropic Head of Pretraining on Scaling Laws, Compute, and the Future of AI Anthropic预训练负责人谈扩展定律、算力与AI的未来
- Notes: The Future of Software Creation with Replit CEO Amjad Masad 笔记:Replit CEO Amjad Masad 谈软件创作的未来
- Notes: Building Cursor At 23, Taking On GitHub Copilot 笔记:23岁创建Cursor,挑战GitHub Copilot
- Notes: OpenAI vs. DeepSeek vs. Qwen - Comparing Open Source LLM Architectures 笔记:OpenAI vs. DeepSeek vs. Qwen - 开源 LLM 架构对比
- Anthropic Co-founder: Building Claude Code, Lessons From GPT-3 & LLM System Design Anthropic联合创始人:构建Claude Code、GPT-3的经验教训与LLM系统设计
- Scaling and the Road to Human-Level AI | Anthropic Co-founder Jared Kaplan 扩展与通往人类级AI之路 | Anthropic联合创始人Jared Kaplan
- Demis Hassabis: Future of AI, Simulating Reality, Physics and Video Games Demis Hassabis:AI的未来、模拟现实、物理与电子游戏
- Notes: How Replit Went From $10M to $100M ARR In Just 9 Months 笔记:Replit 如何在 9 个月内从 1000 万美元 ARR 增长到 1 亿美元
- Notes: AI is Revolutionizing Scientific Discovery 笔记:AI 正在革新科学发现
- DHH: Future of Programming, AI, Ruby on Rails, Productivity & Parenting DHH:编程的未来、AI、Ruby on Rails、生产力与育儿
- Notes: Perplexity's Race to Build Agentic Search 笔记:Perplexity 构建智能搜索的竞赛
- Andrew Ng: Building Faster with AI 吴恩达:用AI更快地构建
- François Chollet: How We Get To AGI François Chollet:我们如何实现AGI
- Software Engineering with LLMs in 2025: Reality Check 2025年LLM软件工程:现实检验
- Fei-Fei Li: Spatial Intelligence is the Next Frontier in AI 李飞飞:空间智能是AI的下一个前沿
- George Church — A Billion Years of Evolution in an Afternoon George Church — 一个下午完成十亿年的进化
- Satya Nadella: Microsoft's AI Bets, Hyperscaling, Quantum Computing Breakthroughs Satya Nadella:微软的AI赌注、超大规模扩展与量子计算突破
- Sam Altman: The Future of OpenAI, ChatGPT's Origins, and Building AI Hardware Sam Altman:OpenAI的未来、ChatGPT的起源与AI硬件建设
- Notes: Software Is Changing (Again) 笔记:软件正在(再次)改变
- Terence Tao: Hardest Problems in Mathematics, Physics & the Future of AI 陶哲轩:数学、物理中最难的问题与AI的未来
- TDD, AI Agents and Coding with Kent Beck – The Pragmatic Engineer TDD、AI代理与编程 - Kent Beck访谈
- Sundar Pichai: CEO of Google and Alphabet Sundar Pichai:Google和Alphabet CEO
- Video-Mirai: Teaching Causal Video Models to Think Ahead Video-Mirai: 让因果视频模型学会前瞻
- From Software Engineer to AI Engineer – Janvi Kalra on The Pragmatic Engineer 从软件工程师到AI工程师 - Janvi Kalra访谈
- Hung-yi Lee ML 2025 Lecture 12: How Language Models Learn to Speak 李宏毅机器学习2025 第十二讲:语言模型如何学会说话
- Hung-yi Lee ML 2025 Lecture 11: Model Merging - Equipping Your Foundation Model with Task Vectors 李宏毅机器学习2025 第十一讲:模型合并 - 为基础模型装备任务向量
- Hung-yi Lee ML 2025 Lecture 10: Model Editing - AI Microsurgery 李宏毅机器学习2025 第十讲:模型编辑 - 人工智能的微创手术
- Hung-yi Lee ML 2025 Lecture 9: Why Do You Care So Much About Benchmarks? 李宏毅机器学习2025 第九讲:你这么认这个评分系统干什么啊?
- Hung-yi Lee ML 2025 Lecture 8: When LLMs Think Too Much - Controlling Reasoning Length 李宏毅机器学习2025 第八讲:当大模型想太多 - 控制推理长度
- Hung-yi Lee ML 2025 Lecture 7: How DeepSeek-R1 and Reasoning Models Think Deeply 李宏毅机器学习2025 第七讲:DeepSeek-R1等推理模型如何进行深度思考
- Hung-yi Lee ML 2025 Lecture 6: Post-Training and the Forgetting Problem 李宏毅机器学习2025 第六讲:后训练与遗忘问题
- Hung-yi Lee ML 2025 Lecture 5: The Power and Limits of Pretrain-Alignment 李宏毅机器学习2025 第五讲:预训练-对齐范式的强大与极限
- Hung-yi Lee ML 2025 TA Session: Multi-GPU Training for LLMs 李宏毅机器学习2025 助教课:利用多GPU训练大型语言模型
- Hung-yi Lee ML 2025 Lecture 4: Beyond Transformers - Linear Attention, Mamba, and the RNN Renaissance 李宏毅机器学习2025 第四讲:超越Transformer - 线性注意力、Mamba与RNN的复兴
- ThePrimeagen: Programming, AI, ADHD, Productivity, Addiction, and God ThePrimeagen:编程、AI、ADHD、生产力、成瘾与上帝
- Hung-yi Lee ML 2025 Lecture 3: The Neuroscience of AI - Inside Language Model Mechanisms 李宏毅机器学习2025 第三讲:AI的脑科学 - 语言模型内部运作机制剖析
- Hung-yi Lee ML 2025 Lecture 2: AI Agents - From AlphaGo to LLM-Powered Autonomy 李宏毅机器学习2025 第二讲:AI Agent - 从AlphaGo到大型语言模型驱动的自主系统
- Notes: How I Use LLMs 笔记:我如何使用大语言模型
- Hung-yi Lee ML 2025 Lecture 1: Understanding Generative AI Breakthroughs 李宏毅机器学习2025 第一讲:一堂课搞懂生成式AI的技术突破
- S1: Simple Test-Time Scaling S1:简单的测试时扩展
- Notes: Deep Dive into LLMs like ChatGPT 笔记:深入理解 ChatGPT 等大语言模型
- AI Engineering with Chip Huyen – The Pragmatic Engineer AI工程与Chip Huyen - Pragmatic Engineer访谈
- DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters DeepSeek、中国、OpenAI、NVIDIA、xAI、台积电、星门与AI超级集群
- Marc Andreessen: Trump, Power, Tech, AI, Immigration & Future of America Marc Andreessen:特朗普、权力、科技、AI、移民与美国的未来
- TouchDesigner Tutorial: Working with Time in Interactive Systems TouchDesigner教程:交互系统中的时间处理
- TouchDesigner Tutorial: Basics of Instancing TouchDesigner教程:实例化基础
- TouchDesigner Tutorial: Channel Shifting RGB Effects TouchDesigner教程:RGB通道偏移效果
- TouchDesigner Tutorial: Image Displacement Techniques TouchDesigner教程:图像位移技术
- TouchDesigner Tutorial: Instancing with CHOPs TouchDesigner教程:使用CHOPs进行实例化
- TouchDesigner Tutorial: Instancing with TOPs TouchDesigner教程:使用TOPs进行实例化
- TouchDesigner Tutorial: Working with the Noise TOP TouchDesigner教程:使用Noise TOP
- TouchDesigner Tutorial: Slit Scan using the Time Machine TOP TouchDesigner教程:使用Time Machine TOP实现狭缝扫描
- TouchDesigner Tutorial: Basic Render Setup TouchDesigner教程:基础渲染设置
- TouchDesigner Tutorial: Lights & Shadows TouchDesigner教程:灯光与阴影
- TouchDesigner Tutorial: Texturing Geometry with MATs TouchDesigner教程:使用MATs为几何体贴图
- TouchDesigner Tutorial: Working with Cameras TouchDesigner教程:使用相机
- TouchDesigner Tutorial: Essential Python Expressions TouchDesigner教程:Python表达式基础
- TouchDesigner Tutorial: Scripting Events with Callbacks TouchDesigner教程:使用回调脚本处理事件
- TouchDesigner Tutorial: The User Interface TouchDesigner教程:用户界面
- TouchDesigner Tutorial: Working with CHOPs TouchDesigner教程:使用CHOPs
- TouchDesigner Tutorial: Working with COMPs TouchDesigner教程:使用COMPs
- TouchDesigner Tutorial: Working with DATs TouchDesigner教程:使用DATs
- TouchDesigner Tutorial: Working with SOPs TouchDesigner教程:使用SOPs
- TouchDesigner Tutorial: Working with TOPs TouchDesigner教程:使用TOPs
- ReAct: Synergizing Reasoning and Acting in Language Models ReAct: Synergizing Reasoning and Acting in Language Models
- Chinchilla: Training Compute-Optimal Large Language Models Chinchilla: Training Compute-Optimal Large Language Models
- Chapter 24: Managing the IoT Box 第24章 管理IoT盒子
- Chapter 20: Remote Procedure Calls in Odoo 第20章 Odoo中的远程过程调用(RPC)
- Chapter 22: Point of Sale 第22章 POS(销售点)
- LoRA: Low-Rank Adaptation of Large Language Models LoRA:大型语言模型的低秩适应
- Chapter 23: Managing Emails in Odoo 第23章 在Odoo中管理Email
- Chapter 19: Managing, Deploying, and Testing with Odoo.sh 第19章 使用Odoo.sh管理、部署和测试
- Chapter 18: Automated Test Cases 第18章 自动化测试用例
- Chapter 17: In-App Purchasing with Odoo 第17章 Odoo的应用内购买
- Chapter 15: Web Client Development 第15章 网页客户端开发
- Chapter 14: CMS Website Development 第14章 CMS网站开发
- Chapter 13: Web Server Development 第13章 Web服务端开发
- Chapter 12: Automation, Workflows, Email, and Printouts 第12章 自动化、工作流、Email和打印件
- Chapter 10: Access Security 第10章 权限安全
- Chapter 21: Performance Optimization 第21章 性能优化
- Chapter 8: Advanced Server-Side Development Techniques 第8章 高级服务端开发技巧
- Chapter 4: Application Models 第4章 应用模型
- Chapter 9: Backend Views 第9章 后端视图
- Chapter 5: Basic Server-Side Development 第5章 基本服务端开发
- Chapter 3: Creating Odoo Add-on Modules 第3章 创建Odoo插件模块
- Chapter 7: Debugging Modules 第7章 调试模块
- Chapter 1: Installing the Odoo Development Environment 第1章 安装Odoo开发环境
- Chapter 11: Internationalization 第11章 国际化
- Chapter 2: Managing Odoo Server Instances 第2章 管理Odoo服务端实例
- Chapter 6: Managing Module Data 第6章 管理模块数据
- Chapter 16: Odoo Web Library (OWL) 第16章 Odoo Web Library (OWL)
- Denoising Diffusion Probabilistic Models (DDPM) Denoising Diffusion Probabilistic Models (DDPM)
- GPT-3: Language Models are Few-Shot Learners GPT-3:语言模型是少样本学习者
- Scaling Laws for Neural Language Models Scaling Laws for Neural Language Models
- Chapter 23: Managing Emails in Odoo 第23章 在Odoo中管理email
- Chapter 21: Performance Optimization 第21章 性能优化
- Chapter 22: Point of Sale 第22章 POS(销售点)
- Chapter 19: Managing, Deploying, and Testing with Odoo.sh 第19章 使用Odoo.sh管理、部署和测试
- Chapter 18: Automated Test Cases 第18章 自动化测试用例
- Chapter 17: In-App Purchasing with Odoo 第17章 Odoo的应用内购买
- Chapter 15: CMS Website Development 第15章 CMS网站开发
- Chapter 20: Remote Procedure Calls in Odoo 第20章 Odoo中的远程过程调用(RPC)
- Chapter 14: Web Server Development 第14章 网页服务端开发
- Chapter 13: Automation, Workflows, and Printouts 第13章 自动化、工作流和打印件
- Chapter 16: Web Client Development 第16章 网页客户端开发
- Chapter 12: Internationalization 第12章 国际化
- Chapter 11: Access Security 第11章 权限安全
- Chapter 24: IoT Box 第24章 IoT盒子
- Chapter 10: Backend Views 第10章 后端视图
- Chapter 9: Advanced Server-Side Development Techniques 第9章 高级服务端开发技巧
- Chapter 8: Debugging 第8章 调试
- Chapter 7: Module Data 第7章 模块数据
- Chapter 6: Basic Server-Side Development 第6章 基本服务端开发
- Chapter 5: Application Models 第5章 应用模型
- Chapter 4: Creating Odoo Add-on Modules 第4章 创建Odoo插件模块
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
- Chapter 3: Server Deployment 第3章 服务器部署
- Chapter 2: Managing Odoo Server Instances 第2章 管理Odoo服务器实例
- Chapter 1: Installing the Odoo Development Environment 第1章 安装Odoo开发环境
- Chapter 4: Automating Regular Administrative Activities 第4章 自动化常规运维活动
- Chapter 10: Basic Networking - Socket Programming 第10章 网络基础 - Socket编程
- Chapter 13: Building Graphical User Interfaces 第13章 创建图形化用户界面
- Chapter 2: Debugging and Profiling Python Scripts 第2章 Python脚本调试和性能测试
- Chapter 8: Documentation and Reporting 第8章 文档和报告
- Chapter 6: File Archiving, Encrypting, and Decrypting 第6章 文件存档、加密和解密
- Chapter 11: Handling Emails with Python Scripting 第11章 使用Python脚本处理邮件
- Chapter 5: Handling Files, Directories, and Data 第5章 文件、目录和数据处理
- Chapter 18: MySQL and SQLite Database Administration 第18章 MySQL和SQLite数据库管理
- Chapter 1: Python Scripting Overview 第1章 Python脚本概述
- Chapter 12: Remote Monitoring of Hosts over Telnet and SSH 第12章 使用Telnet和SSH远程监控主机
- Chapter 15: SOAP and REST API Communication 第15章 SOAP和REST API通讯
- Chapter 17: Statistics Gathering and Reporting 第17章 数据收集及报表
- Chapter 7: Text Processing and Regular Expressions 第7章 文本处理和正则表达式
- Chapter 3: Unit Testing - Introduction to the Unit Testing Framework 第3章 单元测试-单元测试框架的介绍
- Chapter 16: Web Scraping - Extracting Data from Websites 第16章 网络抓取 - 从网站上提取有用的信息
- Chapter 14: Working with Apache and Other Log Files 第14章 处理Apache和其它的日志文件
- Chapter 9: Working with Various File Types 第9章 操作各类文件
- Chapter 14: Deploying and Maintaining Production Instances 第14章 Odoo 12开发之部署和维护生产实例
- Chapter 13: Creating Website Frontend Features 第13章 Odoo 12开发之创建网站前端功能
- Chapter 12: Reports and Server-Side QWeb 第12章 Odoo 12开发之报表和服务端 QWeb
- Chapter 11: Kanban Views and Client-Side QWeb 第11章 Odoo 12开发之看板视图和用户端 QWeb
- Chapter 10: Backend Views - Designing the User Interface 第10章 Odoo 12开发之后台视图 - 设计用户界面
- Chapter 9: External API - Integrating Third-Party Systems 第9章 Odoo 12开发之外部 API - 集成第三方系统
- Chapter 8: Business Logic - Supporting Business Processes 第8章 Odoo 12开发之业务逻辑 - 业务流程的支持
- Chapter 7: Recordsets - Working with Model Data 第7章 Odoo 12开发之记录集 - 使用模型数据
- Chapter 6: Models - Structuring the Application Data 第6章 Odoo 12开发之模型 - 结构化应用数据
- Chapter 5: Import, Export, and Module Data 第5章 Odoo 12开发之导入、导出以及模块数据
- Chapter 4: Extending Modules 第4章 Odoo 12 开发之模块继承
- Chapter 3: Creating Your First Odoo Application 第3章 Odoo 12 开发之创建第一个 Odoo 应用
- Chapter 2: Preparing the Development Environment 第2章 Odoo 12开发之开发环境准备
- Chapter 1: Quick Start Using the Developer Mode 第1章 使用开发者模式快速入门 Odoo 12
- BERT: Pre-training of Deep Bidirectional Transformers BERT: Pre-training of Deep Bidirectional Transformers
- GPT-1: Improving Language Understanding by Generative Pre-Training GPT-1: Improving Language Understanding by Generative Pre-Training
- Attention Is All You Need: The Paper That Started the AI Revolution Attention Is All You Need:开启AI革命的论文
- Mixture-of-Experts (MoE): Outrageously Large Neural Networks Mixture-of-Experts (MoE): Outrageously Large Neural Networks
- Google's Neural Machine Translation System Google's Neural Machine Translation System
- Deep Residual Learning for Image Recognition (ResNet) Deep Residual Learning for Image Recognition (ResNet)
- Distilling the Knowledge in a Neural Network Distilling the Knowledge in a Neural Network
- Sequence to Sequence Learning with Neural Networks Sequence to Sequence Learning with Neural Networks
- Neural Machine Translation by Jointly Learning to Align and Translate Neural Machine Translation by Jointly Learning to Align and Translate
- Generative Adversarial Networks: The Revolutionary Framework for AI Generation Generative Adversarial Networks: The Revolutionary Framework for AI Generation
- Word2Vec: Efficient Estimation of Word Representations Word2Vec: Efficient Estimation of Word Representations
- AlexNet: ImageNet Classification with Deep Convolutional Neural Networks AlexNet:使用深度卷积神经网络进行ImageNet分类
- Brook for GPUs: Stream Computing on Graphics Hardware Brook for GPUs: Stream Computing on Graphics Hardware