Paper: 2609.05381 Authors: Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler Categories: cs.AI
The Gap
Large language models (LLMs) are increasingly evaluated on scientific regression benchmarks, such as predicting molecular solvation free energy, lipophilicity, or binding affinity directly from SMILES strings. When an LLM posts an exceptionally low Mean Absolute Error (MAE) that rivals specialized graph neural networks or density functional theory (DFT) simulations, community excitement follows: the model appears to have discovered quantum chemical principles from natural language pre-training.
Yet a regression score cannot distinguish prediction from retrieval. If a molecule’s exact experimental property value is already recorded in the pre-training corpus, a model does not need to learn chemical mechanics; it only needs to look up the number. In the physical sciences, this leads to an absurd paradox: on the popular FreeSolv benchmark, the published experimental uncertainty is around 0.6 kcal/mol. If an LLM achieves an error of 0.025 kcal/mol, it has surpassed the precision of physical reality. Such precision cannot be prediction — it is photographic recall of a text table.
Prior work acknowledged general data contamination, but lacked systematic diagnostic tools to isolate verbatim retrieval from legitimate inductive generalization in continuous molecular regression.
[SCIENTIFIC BENCHMARK] SMILES String -> LLM -> Continuous Physical Value
|
v
[CORE DILEMMA] Does a low error indicate chemical understanding
or verbatim recall of published literature?
|
+-------------------------+-------------------------+
v v v
[PHYSICAL UNCERTAINTY] [AUDIT SUITE] [REASONING ABLATION]
FreeSolv noise ~ 0.6 22 frontier LLMs across Zero-shot vs Heavy CoT:
kcal/mol; LLM reports 12 regression datasets Does thinking suppress
MAE of 0.025 kcal/mol! audited for digit matches or amplify lookup?
| | |
+-------------------------+-------------------------+
v
[EVIDENCE] Over 50% of frontier LLMs exhibit verbatim retrieval
on 5 major benchmarks; high reasoning increases retrieval by 89%;
suppressing retrieval collapses performance differences across models.
The Increment
One sentence: Before this paper, frontier LLMs were credited with emergent chemical intuition on molecular benchmarks; after it, their benchmark dominance is revealed to be digit-level retrieval of literature tables, where test-time reasoning acts as an index lookup rather than physical inference.
Core Mechanism
The authors conducted an exhaustive audit of 22 frontier models across 12 molecular property regression benchmarks (including FreeSolv, ESOL, Lipophilicity, and BACE).
The diagnostic framework operates across three key stages:
- Verbatim Digit Alignment: Testing whether predictions match published literature values to exact decimal places (e.g. reproducing 3 to 5 significant digits matching legacy papers).
- Reasoning-Level Auditing: Comparing the same models under low-reasoning (direct output) versus high-reasoning (extended Chain-of-Thought). Strikingly, extended reasoning triggered 89% more verbatim retrieval instances, indicating that internal chain-of-thought allows models to traverse associative memory paths to locate memorized citations.
- Retrieval Interruption (Canonical vs Non-Canonical SMILES): By transforming the input representation (e.g. generating randomized non-canonical SMILES, atom-reordered representations, or synonymous IUPAC names), the authors severed the direct token-lookup trigger. Once retrieval was interrupted, the performance gap between top frontier models largely vanished, converging to a baseline level of actual generalizable prediction.
RETRIEVAL VS PREDICTION FORENSIC PIPELINE
[Input: Molecule SMILES]
|
+-------+-------+
| |
v v
[Standard SMILES] [Transformed/Scrambled SMILES]
| |
v v
[Frontier LLM] [Frontier LLM]
(CoT Memory) (Forced Inductive Path)
| |
v v
Exact Match: Degraded to Baseline Error:
0.025 kcal/mol ~ 1.15 kcal/mol
(Memorized Table) (True Inductive Generalization)
To visualize this dynamic, consider a structural metaphor of a trivia contestant who memorized the encyclopedia volume on world capitals. When asked the exact elevation of the capital of Burundi to four decimal places, a geographer would estimate based on topography and altitude maps, accepting a margin of error. The trivia contestant pauses, browses their mental index card of page 412, and recites the exact elevation down to the inch. If you disguise the country by giving only its coordinates, the contestant’s superpower vanishes instantly, revealing they know nothing about mountainous geography.
Key Concepts
- Digit-Level Verbatim Retrieval: The reproduction of empirical benchmark target values to exact decimal places stored from training papers, exceeding the ground-truth experimental measurement precision.
- Reasoning-Induced Retrieval Amplification: The counterintuitive phenomenon where increasing test-time reasoning tokens enables the model to locate and extract obscure memorized tabular data rather than calculating the answer.
- Retrieval Interruption: A methodology of perturbing syntactic surface forms (e.g. SMILES randomization) to break literal token associative memory without altering the underlying physical graph.
Framework Shift
Before (Naive Benchmark Faith):
SMILES ---> [Frontier LLM] ---> [Low Prediction Error] ---> Claim: "Model understands chemistry"
(Ignores the fact that error is lower than real-world instrument noise)
After (Forensic Contamination Audit):
SMILES ---> [Frontier LLM] ---> [Digit Match Audit] ---> Flag: Literature Retrieval
SMILES* --> [Interrupted] ---> [Physical Error] ---> Reality: True generalization ceiling
From trusting leaderboard MAE as an indicator of scientific insight to auditing whether predictions violate physical measurement uncertainty boundaries, the paradigm shifts from praise of emergent intelligence to forensic detection of literature memorization.
Expert Assessment
Problem choice: Outstanding. As LLMs are pitched for drug discovery and materials science, relying on models that merely quote old papers risks costly dead ends in wet labs.
Method maturity: The methodology is meticulous. Testing 22 different frontier models across 12 datasets with controlled reasoning budgets prevents the conclusions from being dismissed as quirks of a single architecture.
Experimental integrity: The authors cite the original paper errors and instrumentation tolerances (such as FreeSolv’s 0.6 kcal/mol limit), demonstrating that the LLM was scoring well below the noise floor of the physical experiment.
Writing quality: Clear, evidence-driven, and devoid of sensationalism. The tables showing exact digit matches across multiple models are unassailable.
Verdict: strong accept — A definitive study that changes how scientific machine learning benchmarks must be constructed and audited.
Takeaways
- If an AI model’s benchmark error is lower than the experimental uncertainty of the original apparatus, suspect memorization immediately.
- Extended reasoning (CoT) is not just a problem-solver; it is also a search engine into the pre-training corpus.
- For scientific benchmarks, always evaluate on perturbed, non-canonical representations (e.g. randomized SMILES or newly synthesized compounds) to separate memory from reasoning.
论文: 2609.05381 作者: Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler 分类: cs.AI
缺口
近年来,前沿大语言模型(LLM)被频繁应用于分子化学与材料科学,直接根据分子的 SMILES 字符串预测其水合自由能、脂水分配系数或结合亲和力。 当顶级大模型在公开基准上刷出极低的平均绝对误差(MAE),甚至超越专业的图神经网络(GNN)与密度泛函理论(DFT)模拟时,整个社区往往为之振奋,认为大模型仅凭自然语言预训练就“涌现”出了量子化学直觉。
然而,在连续数值回归任务中,基准得分根本无法区分真正的预测(Prediction)与文献检索(Retrieval)。 如果一个分子及其属性数值早已白纸黑字印在预训练文献中,模型根本不需要理解化学,只需要把数字“背出来”。 这在物理科学中引出了一个极度荒谬的悖论:在著名的 FreeSolv 水合能基准中,由于原始实验测量的局限,官方设定的默认实验不确定度约为 0.6 kcal/mol。 但实测中,某顶尖大模型在 FreeSolv 上的预测误差居然只有 0.025 kcal/mol! 一个误差远小于物理世界测量精度的模型,根本不可能是在进行物理预测——它只是以照相机的精度,复述了论文表格里的数字。
此前的研究虽然隐约意识到数据污染的存在,但缺乏系统性的量化手段来刺破连续回归评测中的字面记忆假象。
[科学评测基准] 分子 SMILES 字符串 -> 大模型 -> 预测物理连续值
|
v
[核心拷问] 超低预测误差究竟代表学会了化学法则,
还是仅仅背诵了训练集里的论文附表?
|
+-------------------------+-------------------------+
v v v
[物理测量不确定度极限] [基准污染审计体系] [推理强度消融]
FreeSolv 噪声 ~ 0.6 覆盖 22 个前沿大模型与 对比零样本直出与长思考 CoT:
kcal/mol;大模型误差竟 12 个权威回归基准, “深入思考”究竟是在推理,
低至 0.025 kcal/mol! 逐位比对浮点数有效数字 还是在扩大文献检索范围?
| | |
+-------------------------+-------------------------+
v
[实测事实] 5 个主流基准中超过 50% 的模型存在字面背诵;
开启高强度思考后,字面检索行为暴增 89%;
打断记忆后,所有前沿模型的真实预测能力瞬间收敛回平庸基线。
增量
一句话: 在这篇论文之前,前沿大模型在分子基准上的统治级表现被赞誉为通晓微观物理法则;在这篇论文之后,学界看清了其本质大多是浮点数级别的文献死记硬背,而“增加思考时间”往往只是在增强记忆检索的命中率。
核心机制
研究团队对全球 22 个前沿大模型在 12 个主流分子物性回归基准(涵盖 FreeSolv、ESOL、Lipophilicity、BACE 等)上展开了深度法证式审计。
其核心审计机制由三部分组成:
- 数字级字面匹配(Digit-Level Alignment):比对模型生成的预测数值与数十年前原始文献中的浮点数字符串(精确到小数点后 3 至 5 位有效数字)。
- 推理层级消融实验(Reasoning-Level Auditing):对比模型在“直接回答”与“延长思维链(CoT)”下的表现。 震撼的发现是:在高推理强度下,模型被判定为字面背诵的频次反而飙升了 89%——这证明内省式的思维链并不总是在进行严密的逻辑推导,而是在模型的千亿参数记忆库中充当了语义联想检索探针。
- 检索阻断机制(Retrieval Interruption):通过打乱 SMILES 的原子遍历顺序(生成非标准非规范 SMILES),或者替换为同义但冷门的分子描述。 一旦切断了字面字符串与文献标题的精确关联,顶级模型的断层领先优势彻底灰飞烟灭,预测误差全部回退到真正的泛化水平。
分子记忆与预测阻断实验流程
[输入: 分子 SMILES 描述]
|
+-------+-------+
| |
v v
[标准规范 SMILES] [扰动/非标准 SMILES]
| |
v v
[前沿大模型] [前沿大模型]
(思维链检索模式) (强制归纳推理模式)
| |
v v
精确数字复述: 误差大幅回升至物理基线:
0.025 kcal/mol ~ 1.15 kcal/mol
(命中训练集表格) (真正的冷启动泛化能力)
可以用一个百科全书死记硬背者的核喻来理解这个现象: 一位考生在参加地理知识抢答。 当被问到某冷门岛国的海拔时,真正的地理学者会根据气候与地貌断层估算出一个区间; 而死记硬背者闭上眼睛思索五秒,准确报出“海拔 1284.37 米”——与 1982 年出版的百科全书上的排印印刷一字不差。 只要你把题目换成经纬度坐标,死记硬背者就彻底哑口无言。 这证明他心里根本没有群山与海洋,只有油墨印出的铅字。
关键概念
- 字面数字检索(Digit-Level Verbatim Retrieval):模型未做物理建模,而是直接从预训练语料中提取出小数点后数位完全一致的原始实验数据。
- 推理放大检索(Reasoning-Induced Retrieval Amplification):增加推理 token 数量反而给模型提供了更充裕的上下文寻址空间,帮助其精准锚定冷门文献片段的反直觉现象。
- 检索阻断(Retrieval Interruption):利用图等价但表层文本异构的表示法(如随机化 SMILES),在不改变化学结构的前提下破坏文本层面的字面关联。
框架转变
之前(对公开榜单的盲目信任):
分子文本 ---> [前沿大模型] ---> [超低预测误差] ---> 结论:“大模型自学掌握了深层化学知识”
(完全忽视了预测误差已经违反物理测量噪声底线的事实)
之后(严谨的法证式污染审计):
分子文本 ---> [前沿大模型] ---> [有效数字法证审计] ---> 诊断:文献字面背诵
异构文本 ---> [阻断记忆后] ---> [真实物理误差重测] ---> 真相:真实归纳泛化能力依然有限
从迷信基准榜单上的 MAE 数字,转向依据物理测量极限来审判模型输出的合法性,核心转变在于将“真正的物理常识推理”与“海量文献照相记忆”严格区分开来。
专家评审
选题眼光: 极具洞察力与批判精神。 在 AI for Science 烈火烹油的当下,如果药企或实验室依据大模型在公开榜单上的“神级表现”投入上千万元的湿实验研发,极可能因为文献记忆的虚假繁荣而遭遇重大挫败。
方法成熟度: 评测逻辑严丝合缝。 不仅测试了 22 款主流基座模型,而且设计了严密的非规范 SMILES 阻断对照组,使“模型并非真正理解物理规律”的结论坚不可摧。
实验诚意: 翻阅对比了半个世纪以来的原始化学实验文献与测定仪器公差,用“误差低于实验不确定度”这一铁证彻底拆穿了神话。
写作功力: 逻辑清晰,层层剥茧,是近年来大模型基准审计领域的典范之作。
判决: 强接收(Strong Accept) — 重新定义了科学大模型评估的行业标准。
要点总结
- 凡是大模型在物理/化学任务中的预测精度超过了实验仪器本身的公差底线,几乎百分之百是文献污染背诵。
- 测试时推理(CoT)并不总是逻辑演绎,它同样是大模型在浩瀚记忆中翻找文献索引的有效工具。
- 评测科学 AI 模型时,必须采用异构表示扰动或全新合成的分子,彻底屏蔽字面检索通道。