Paper: 2609.14057 Authors: Yiheng Zhao, Mengzhuo Chen, Chengming Hu, Pengyi Liao, Yiran Pang Categories: cs.CL
The Gap
Frontier LLMs are increasingly capable of conducting automated research, yet their creativity in this setting has not been systematically evaluated. That is a gap with a particular shape: capability is measured by outputs — did the experiment run, did the paper get written — while creativity is the property that determines whether the research was worth doing. A system that produces competent, unoriginal work passes every capability check.
The measurement problem is what makes this hard. “Creativity” resists a single score, and the obvious composite would average away the differences that matter. So the paper’s contribution is a set of metrics along dimensions chosen deliberately, rather than one number.
CAPABILITY IS MEASURED; THE PROPERTY THAT MAKES RESEARCH WORTH DOING IS NOT
FRONTIER LLMs are increasingly capable of CONDUCTING AUTOMATED RESEARCH
YET THEIR CREATIVITY IN THIS SETTING HAS NOT BEEN SYSTEMATICALLY
EVALUATED
|
v
[THE GAP HAS A PARTICULAR SHAPE]
CAPABILITY is measured by OUTPUTS
did the experiment RUN
did the paper GET WRITTEN
CREATIVITY is the property that determines WHETHER THE RESEARCH WAS
WORTH DOING
-> a system producing COMPETENT, UNORIGINAL WORK PASSES EVERY
CAPABILITY CHECK
[AND THE MEASUREMENT PROBLEM IS WHAT MAKES THIS HARD]
"CREATIVITY" RESISTS A SINGLE SCORE
-> the OBVIOUS COMPOSITE would AVERAGE AWAY THE DIFFERENCES THAT
MATTER
-> so the contribution is A SET OF METRICS ALONG DELIBERATELY CHOSEN
DIMENSIONS, rather than ONE NUMBER
The Increment
One sentence: Before this paper, automated-research creativity was unmeasured; after it, sixteen metrics split valueness from three kinds of novelty show models differ sharply on the novelty dimension that most predicts research performance.
Core Mechanism
The first split is valueness versus novelty, and it is the right first cut: valueness assesses whether each proposed idea is useful, while novelty is evaluated from three perspectives. Keeping them separate matters because they can dissociate — an idea can be genuinely novel and useless, or useful and entirely derivative, and a combined score would blur the two failure modes.
The three novelty perspectives are progressively harder to satisfy, which is what makes the decomposition informative:
- Exact-Match P-Novelty — whether the same idea has appeared before. The weakest test: it detects literal repetition and nothing subtler.
- Variable-level P-Novelty — whether a previously unexplored variable or variable combination is explored. A stronger test, because the same idea can be restated in new terms and pass the exact-match test while exploring nothing new. This one asks whether the research space was extended.
- H-Novelty — whether the idea directly follows retrieved external knowledge or departs from it. The most interesting of the three, because it measures the relationship to prior work rather than overlap with it. An idea that follows directly from what was retrieved is closer to competent summarisation than to research.
The result is a profile rather than a ranking, and the contrast within it is the finding: models achieve relatively similar scores on most creativity metrics, but differ substantially in Variable-level P-Novelty, which reflects the breadth of research-space exploration. So most of the metrics fail to separate the models — which is itself informative, since it means the usual dimensions of comparison are not where the difference lies.
And the paper connects the metric to outcomes: further correlation and idea-level performance analyses show that Variable-level P-Novelty is the creativity dimension most strongly associated with research performance. That closes the loop in the way that matters for a metric — not “this measures something interesting” but “this measures the thing that predicts the outcome you care about.” It also justifies the paper’s structure after the fact: the dimension on which models differ is the dimension that predicts performance, so the disagreement between models is substantively important rather than incidental.
[THE FIRST SPLIT: VALUENESS VERSUS NOVELTY]
VALUENESS assesses whether each proposed idea IS USEFUL
NOVELTY is evaluated FROM THREE PERSPECTIVES
<- keeping them SEPARATE matters because they can DISSOCIATE
<- an idea can be GENUINELY NOVEL AND USELESS, or USEFUL AND
ENTIRELY DERIVATIVE
-> a COMBINED score would BLUR THE TWO FAILURE MODES
[THE THREE NOVELTY PERSPECTIVES ARE PROGRESSIVELY HARDER TO SATISFY]
-- which is what makes the decomposition INFORMATIVE
[1] EXACT-MATCH P-NOVELTY
whether THE SAME IDEA HAS APPEARED BEFORE
<- the WEAKEST TEST: detects LITERAL REPETITION and NOTHING
SUBTLER
[2] VARIABLE-LEVEL P-NOVELTY
whether A PREVIOUSLY UNEXPLORED VARIABLE OR VARIABLE
COMBINATION IS EXPLORED
<- STRONGER: the SAME IDEA can be RESTATED IN NEW TERMS and
PASS the exact-match test while EXPLORING NOTHING NEW
<- this one asks whether THE RESEARCH SPACE WAS EXTENDED
[3] H-NOVELTY
whether the idea DIRECTLY FOLLOWS RETRIEVED EXTERNAL KNOWLEDGE
or DEPARTS FROM IT
<- the MOST INTERESTING: it measures THE RELATIONSHIP TO PRIOR
WORK rather than OVERLAP WITH IT
<- an idea that FOLLOWS DIRECTLY from what was retrieved is
CLOSER TO COMPETENT SUMMARISATION THAN TO RESEARCH
[THE RESULT IS A PROFILE RATHER THAN A RANKING -- AND THE CONTRAST
WITHIN IT IS THE FINDING]
models achieve RELATIVELY SIMILAR SCORES ON MOST METRICS
BUT DIFFER SUBSTANTIALLY IN VARIABLE-LEVEL P-NOVELTY
<- which reflects THE BREADTH OF RESEARCH-SPACE EXPLORATION
<- MOST METRICS FAIL TO SEPARATE THE MODELS
-> itself informative: the USUAL DIMENSIONS OF COMPARISON are
NOT WHERE THE DIFFERENCE LIES
[AND THE METRIC IS CONNECTED TO OUTCOMES]
CORRELATION AND IDEA-LEVEL PERFORMANCE ANALYSES show VARIABLE-LEVEL
P-NOVELTY IS THE CREATIVITY DIMENSION MOST STRONGLY ASSOCIATED WITH
RESEARCH PERFORMANCE
<- closes the loop in the way that matters for a METRIC:
not "this measures SOMETHING INTERESTING"
but "this measures THE THING THAT PREDICTS THE OUTCOME YOU
CARE ABOUT"
<- and it JUSTIFIES THE STRUCTURE AFTER THE FACT: the dimension on
which models DIFFER is the dimension that PREDICTS PERFORMANCE
-> so the disagreement between models is SUBSTANTIVELY
IMPORTANT rather than INCIDENTAL
Think of it as assessing chefs by their menus rather than by whether dinner arrived. Dinner arriving is capability, and everything in the kitchen can pass that test. To judge creativity you would look at the menu along separable axes: is this dish any good (valueness); has this exact dish been served here before (exact match); does it use an ingredient combination nobody has tried (variable-level); and is it a variation on what the last restaurant down the street is doing, or something that stands on its own (H-novelty). The paper’s finding is that most chefs look alike on most of these — and that the one axis where they genuinely differ is the one that predicts whether the restaurant does well. Which is also why collapsing the axes into one number would have been a mistake: the difference lives on a single axis and an average would have hidden it.
Key Concepts
- Valueness and novelty as separable axes: useful ideas and novel ideas dissociate, so a combined score blurs two failure modes.
- Three novelty tests of increasing strength: exact match, variable-level exploration, and departure from retrieved knowledge. The progression is what makes the decomposition informative rather than redundant.
- Variable-level P-Novelty as research-space breadth: whether an unexplored variable or combination appears. It catches restatement, which the exact-match test cannot.
- H-Novelty as relationship to prior work: following retrieved knowledge directly is closer to summarisation than research. It measures relation rather than overlap.
- A profile that separates on one axis: similar scores on most metrics, substantial differences on one. It shows the usual dimensions of comparison are not where the difference is.
- Correlation with performance as the validation: the differing dimension is the one most associated with research performance, which makes the disagreement substantive.
Framework Shift
Before (capability as the measure):
assess automated research by outputs: did it run, was it written
-> competent and unoriginal work passes
-> "creativity" has no systematic measurement
-> a composite score would average the dimensions together
After (a profile across separated dimensions):
valueness separated from three novelty perspectives
-> exact-match, variable-level, and departure from retrieved knowledge
-> models score similarly on most metrics
-> they differ substantially on variable-level P-Novelty
-> and that is the dimension most associated with performance
From measuring whether automated research produces output, to measuring whether the ideas extend the research space, the core shift is that the creativity dimension separating models is a single axis that a composite score would have averaged away.
Expert Assessment
Problem choice: Excellent, and the gap is framed in the way that makes measurement the right response. Capability and creativity can dissociate, and a system producing competent unoriginal work is exactly the case a capability benchmark cannot detect — so building metrics is not a refinement of existing evaluation but a different question.
Method maturity: The decomposition is well designed, and the progression across the three novelty tests is what gives it content: exact match catches literal repetition, variable-level catches restatement, and H-Novelty measures the relationship to prior work rather than its overlap. Keeping valueness separate from novelty is the right first cut, since the two dissociate. And connecting the metric to research performance is what makes it a validated dimension rather than an interesting one — a claim about prediction rather than description.
Experimental integrity: The most valuable finding is that most metrics do not separate the models, because reporting that is less flattering than a clean ranking and it reframes where comparison should focus. The idea-level performance analysis is what supports the association with performance, and the paper presents the correlation as an association rather than as a causal claim. The limitation is that performance is itself measured by some criterion, so the correlation validates the metric against that criterion rather than against research value in an absolute sense — a caveat the framing acknowledges implicitly by speaking of research performance.
Writing quality: The three novelty perspectives are named and justified individually, which makes the decomposition usable rather than a list. Because the practical consequence is “compare models on variable-level P-Novelty”, a short passage on how that is computed — what counts as a variable, and what counts as unexplored — would let a reader apply the metric to other settings.
Verdict: strong accept — it separates the creativity dimensions that a capability benchmark conflates, shows the models differ on one axis while agreeing on the rest, and validates that axis against research performance.
Takeaways
- Separate usefulness from novelty. An idea can be novel and useless or useful and derivative, and one score cannot express both.
- Test novelty at increasing strength. Exact-match, variable-level and relationship tests catch different things, and the strongest is the most informative.
- Report where a metric fails to separate. Similar scores across most dimensions is a finding about where the differences are not.
- Validate the metric against an outcome. A dimension that predicts performance is worth comparing on, where an interesting one may not be.
论文: 2609.14057 作者: Yiheng Zhao, Mengzhuo Chen, Chengming Hu, Pengyi Liao, Yiran Pang 分类: cs.CL
缺口
前沿 LLM 越来越有能力进行自动化科研,然而它们在这件事上的「创造力」从未被系统地评估过。 这个缺口形态特殊:能力由输出来衡量——实验跑起来了吗、论文写出来了吗——而创造力才是决定这项研究值不值得做的性质。一个产出称职但毫无原创工作的系统,能通过所有能力检查。
而正是测量上的困难让这件事棘手。“创造力”抗拒单一分数,而显而易见的合成分会把真正要紧的那些差异平均掉。所以论文的贡献是一组沿着刻意选定的维度构造的指标,而不是一个数。
能力被测量;而"让研究值得做"的那个性质没有被测量
前沿 LLM 越来越有能力「进行自动化科研」
然而它们在这件事上的「创造力」从未被系统地评估过
|
v
[这个缺口形态特殊]
「能力」由「输出」来衡量
实验跑起来了吗
论文写出来了吗
「创造力」才是决定"这项研究值不值得做"的那个性质
-> 一个产出「称职但毫无原创」工作的系统,
能通过所有能力检查
[而正是测量上的困难让这件事棘手]
"创造力"「抗拒单一分数」
-> 显而易见的「合成分」会把「真正要紧的那些差异平均掉」
-> 所以贡献是一「组」沿着刻意选定维度构造的指标,
而不是一个数
增量
一句话: 在这篇论文之前,自动化科研的创造力未被度量;在这篇论文之后,十六项指标把”有用性”与三种新颖性分开,显示模型在”最能预测科研表现”的那个新颖性维度上差异显著。
核心机制
第一刀切在「有用性 vs 新颖性」,而这是正确的第一刀:有用性评估每个被提出的想法是否有用,而新颖性从三个视角评估。 把两者分开很重要,因为它们可以解离——一个想法可以真正新颖却毫无用处,也可以有用却完全派生,而一个合并分数会把这两种失效模式模糊在一起。
三种新颖性视角的满足难度是递进的,这正是让这个拆解有信息量的原因:
- 精确匹配 P-Novelty(Exact-Match)——同一个想法此前是否出现过。 最弱的检验:它能检出字面重复,除此之外什么也检不出。
- 变量层面 P-Novelty(Variable-level)——是否探索了一个此前未被探索的变量或变量组合。 更强的检验,因为同一个想法可以用新的说法重述、通过精确匹配检验,却什么新东西都没探索。这一个问的是研究空间是否被扩展了。
- H-Novelty——这个想法是直接跟随检索到的外部知识、还是偏离了它。 三者中最有意思的一个,因为它度量的是与已有工作的「关系」,而不是与它的重叠。一个直接跟随检索结果的想法,更接近称职的综述,而不是科研。
结果是「一组画像」而不是一个排名,而其中的对比才是发现:模型在多数创造力指标上分数相对接近,却在「变量层面 P-Novelty」上差异显著,而这一维反映的正是研究空间探索的广度。所以多数指标没能把模型区分开来——这本身就有信息量,因为它意味着通常用来比较的那些维度并不是差异所在。
而论文把指标与结果连了起来:进一步的相关性与「想法层面」的表现分析表明,变量层面 P-Novelty 是与科研表现关联最强的那个创造力维度。 这以一种对指标而言要紧的方式闭合了回路——不是”它度量了某个有意思的东西”,而是”它度量的正是能预测你所关心结果的那个东西”。它也在事后为论文的结构提供了正当性:模型之间产生差异的那个维度,正是预测表现的那个维度,所以模型之间的分歧是实质性的、而不是偶然的。
[第一刀:有用性 vs 新颖性]
「有用性」评估每个被提出的想法「是否有用」
「新颖性」从「三个视角」评估
<- 把两者分开很重要,因为它们「可以解离」
<- 一个想法可以「真正新颖却毫无用处」,也可以
「有用却完全派生」
-> 一个「合并分数」会把这两种失效模式「模糊在一起」
[三种新颖性视角的满足难度是递进的]
——这正是让这个拆解「有信息量」的原因
[1] 「精确匹配 P-NOVELTY」
同一个想法此前是否出现过
<- 最弱的检验:能检出「字面重复」,除此之外什么也检不出
[2] 「变量层面 P-NOVELTY」
是否探索了一个此前未被探索的变量或变量组合
<- 更强:同一个想法可以「用新的说法重述」、
通过精确匹配检验,却「什么新东西都没探索」
<- 这一个问的是「研究空间是否被扩展了」
[3] 「H-NOVELTY」
这个想法是「直接跟随检索到的外部知识」、
还是「偏离了它」
<- 三者中最有意思:它度量的是「与已有工作的关系」,
而不是与它的「重叠」
<- 一个直接跟随检索结果的想法,
「更接近称职的综述,而不是科研」
[结果是「一组画像」而不是一个排名——而其中的对比才是发现]
模型在「多数指标」上分数「相对接近」
「却在「变量层面 P-NOVELTY」上差异显著」
<- 而这一维反映的正是「研究空间探索的广度」
<- 「多数指标没能把模型区分开来」
-> 这本身就有信息量:「通常用来比较的那些维度
并不是差异所在」
[而指标与结果被连了起来]
相关性与「想法层面」的表现分析表明:「变量层面 P-NOVELTY
是与科研表现关联最强的那个创造力维度」
<- 以一种对指标而言要紧的方式闭合了回路:
不是"它度量了某个有意思的东西"
而是"它度量的正是「能预测你所关心结果的那个东西」"
<- 也在事后为论文结构提供了正当性:
「模型之间产生差异的那个维度,正是预测表现的那个维度」
-> 所以模型之间的分歧是「实质性的」、而不是「偶然的」
可以用**“靠菜单而不是靠”晚饭有没有端上来”来评判厨师”来理解这件事: 晚饭端上来了是能力**,后厨里的一切都能通过这项检验。要评判创造力,你会沿着彼此可分的轴去看菜单:这道菜好不好(有用性);这道菜此前有没有在本店上过(精确匹配);它用的食材组合有没有人试过(变量层面);而它是对街那家店最近做法的一个变体、还是自成一路(H-Novelty)。 论文的发现是:大多数厨师在这些轴上的大多数项看起来都差不多——而他们真正有差异的那条轴,正是能预测这家店开得好不好的那条。 这也正是为什么把各条轴压成一个数会是个错误:差异只住在一条轴上,而平均值会把它藏起来。
关键概念
- 以「有用性」与「新颖性」作为可分的轴: 有用的想法与新颖的想法会解离,所以合并分数会模糊两种失效模式。
- 三种强度递进的新颖性检验: 精确匹配、变量层面探索、以及与检索知识的偏离。这个递进才让拆解有信息量而不是冗余。
- 以变量层面 P-Novelty 作为研究空间的广度: 是否出现了一个未被探索的变量或组合。它能抓住重述,而精确匹配做不到。
- 以 H-Novelty 作为与已有工作的关系: 直接跟随检索到的知识,更接近综述而非科研。它度量的是关系而不是重叠。
- 只在一条轴上分开的画像: 多数指标分数相近、一条轴上差异显著。它说明通常的比较维度不是差异所在。
- 以与表现的相关性作为验证: 产生差异的那个维度正是与科研表现关联最强的——这才让分歧具有实质性。
框架转变
之前(以能力作为度量):
用输出评估自动化科研:跑起来了吗、写出来了吗
-> 称职而无原创的工作可以通过
-> "创造力"没有系统的测量
-> 合成分会把各维度平均掉
之后(一个在可分维度上的画像):
有用性与三种新颖性视角分开
-> 精确匹配、变量层面、以及与检索知识的偏离
-> 模型在多数指标上分数相近
-> 在变量层面 P-Novelty 上差异显著
-> 而那正是与表现关联最强的维度
从”测量自动化科研是否产出了输出”,转变为”测量这些想法是否扩展了研究空间”,核心转变在于:区分模型的那个创造力维度是单独一条轴,而一个合成分本会把它平均掉。
专家评审
选题眼光: 极好,而缺口的框定方式使”测量”成为正确的应对。 能力与创造力可以解离,而”产出称职却无原创工作的系统”恰恰是能力基准检不出的那类情形——所以构建指标不是对既有评测的改良,而是另一个问题。
方法成熟度: 拆解设计得好,而三种新颖性检验的递进才给了它内容:精确匹配抓字面重复、变量层面抓重述、H-Novelty 度量与已有工作的关系而非重叠。把有用性与新颖性分开是正确的第一刀,因为两者会解离。而把指标与科研表现连起来,才让它成为一个被验证的维度、而不是一个有趣的维度——那是一个关于预测的主张,而不是关于描述的主张。
实验诚意: 最有价值的发现是”多数指标没能区分模型”——因为报告这一点比给一个干净的排名更不好看,而它重新指出了比较应当聚焦在哪里。“想法层面”的表现分析支撑了与表现的关联,而论文把相关性呈现为关联而非因果主张。 局限是:表现本身也是由某个标准衡量的,所以这个相关性是把指标对着那个标准验证,而不是在绝对意义上对着”研究价值”验证——这个保留条件通过使用”科研表现”这一措辞被隐含地承认。
写作功力: 三种新颖性视角各自被命名并给出理由,这使拆解可用而不只是一份清单。 由于实际后果是”在变量层面 P-Novelty 上比较模型”,若能补一小段讲清它如何计算——什么算一个变量、什么算未被探索——会让读者能把该指标应用到别的场景。
判决: 强接收(Strong Accept) — 它把能力基准所混淆的创造力维度分开,表明模型在一条轴上不同、在其余轴上相近,并对着科研表现验证了那条轴。
要点总结
- 把有用性与新颖性分开。一个想法可以新颖而无用、或有用而派生,一个分数无法同时表达两者。
- 用递进的强度检验新颖性。精确匹配、变量层面与关系检验各自抓住不同的东西,而最强的那一个信息量最大。
- 报告指标在哪里没能区分。多数维度上分数相近,是一条关于”差异不在哪里”的发现。
- 对着结果验证指标。一个能预测表现的维度值得用来比较,而一个”有趣”的维度未必。