Paper: 2609.14016 Author: Ivan Silajev Categories: cs.IR, cs.CL
The Gap
TF-IDF and BM25 are two of the most widely used methods for scoring query-document relevance, decades old and deployed everywhere. And neither has a standard probabilistic derivation that justifies it as a statistical method within a unified framework.
That absence has consequences that are easy to underrate. A heuristic that works can be tuned, combined and used, but it cannot be reasoned about. You cannot say what it estimates, so you cannot say when it should fail, what a sensible variant would be, or how it relates to a method derived from different principles. And comparison is stuck at the experimental level: two scorers can be benchmarked against each other indefinitely without any account of whether they are measuring the same quantity.
TWO UBIQUITOUS SCORERS WITHOUT A DERIVATION
TF-IDF and BM25: TWO OF THE MOST WIDELY USED METHODS for scoring
query-document relevance
decades old, deployed everywhere
|
v
AND NEITHER HAS A STANDARD PROBABILISTIC DERIVATION that justifies it
as a STATISTICAL METHOD WITHIN A UNIFIED FRAMEWORK
|
v
[WHY THAT ABSENCE MATTERS, AND IS EASY TO UNDERRATE]
a heuristic that works can be TUNED, COMBINED and USED
but it CANNOT BE **REASONED ABOUT**:
YOU CANNOT SAY WHAT IT ESTIMATES
-> so you cannot say WHEN IT SHOULD FAIL
-> nor what a SENSIBLE VARIANT would be
-> nor HOW IT RELATES to a method derived from DIFFERENT PRINCIPLES
|
v
AND COMPARISON IS STUCK AT THE EXPERIMENTAL LEVEL
-> two scorers can be BENCHMARKED AGAINST EACH OTHER INDEFINITELY
without any account of WHETHER THEY ARE MEASURING THE SAME
QUANTITY
The Increment
One sentence: Before this paper, TF-IDF and BM25 were heuristics with no probabilistic interpretation; after it, both are exact KL divergences between two probability models, giving them a shared basis and making theoretical comparison possible.
Core Mechanism
The claim is that both scoring methods admit an exact interpretation as Kullback-Leibler divergences between two probability models. Two properties of that result are worth isolating, because they determine what it is good for.
It is exact, not approximate. An asymptotic or approximate derivation would explain why the scorers behave sensibly in some limit; an exact identification says the score is the divergence, so statements about the divergence are statements about the scorer. That is the difference between a heuristic with an explanation and a statistic.
And it is a shared basis, which is the part that unlocks comparison. The paper’s stated payoff is a common theoretical basis for TF-IDF and BM25 that clarifies what they measure and allows them to be compared theoretically with other information retrieval methods rather than only experimentally. So the contribution is not two separate derivations that happen to exist; it is that both land in the same framework, which is what makes them commensurable.
One detail deserves emphasis because it is the kind of thing that decides whether a derivation applies to reality: the paper treats the BM25 variant that includes the plus-one correction in the IDF term, which is the one used in practice, and also discusses the original formulation without that correction. That is a deliberate choice to derive the deployed method rather than the published idealisation. Derivations of a formula nobody runs are of theoretical interest; deriving the variant in production is what lets the theory bear on actual systems — and the fact that the paper covers both means the difference between them becomes visible rather than assumed away.
THE CLAIM
BOTH scoring methods admit an EXACT INTERPRETATION AS
KULLBACK-LEIBLER DIVERGENCES BETWEEN TWO PROBABILITY MODELS
TWO PROPERTIES DETERMINE WHAT THE RESULT IS GOOD FOR
[1] IT IS **EXACT**, NOT APPROXIMATE
<- an ASYMPTOTIC or APPROXIMATE derivation would explain why
the scorers behave sensibly IN SOME LIMIT
<- an EXACT IDENTIFICATION says THE SCORE **IS** THE
DIVERGENCE
-> statements about the DIVERGENCE are statements about THE
SCORER
<- the difference between a HEURISTIC WITH AN EXPLANATION and
a STATISTIC
[2] IT IS A **SHARED BASIS** -- the part that UNLOCKS COMPARISON
the stated payoff: a COMMON THEORETICAL BASIS FOR TF-IDF AND
BM25 that
CLARIFIES WHAT THEY MEASURE
ALLOWS THEM TO BE COMPARED THEORETICALLY with other
information retrieval methods RATHER THAN ONLY EXPERIMENTALLY
<- the contribution is NOT two separate derivations that happen
to exist
<- it is that BOTH LAND IN THE SAME FRAMEWORK
-> which is what makes them COMMENSURABLE
ONE DETAIL DESERVES EMPHASIS -- it decides whether a derivation APPLIES
TO REALITY
the paper treats THE BM25 VARIANT THAT INCLUDES THE PLUS-ONE
CORRECTION IN THE IDF TERM, WHICH IS THE ONE USED IN PRACTICE
and ALSO discusses the ORIGINAL FORMULATION WITHOUT THAT CORRECTION
<- a DELIBERATE CHOICE to derive THE DEPLOYED METHOD rather than
THE PUBLISHED IDEALISATION
<- derivations of a formula NOBODY RUNS are of THEORETICAL interest
<- deriving THE VARIANT IN PRODUCTION is what lets the theory BEAR
ON ACTUAL SYSTEMS
-> and covering BOTH means the DIFFERENCE BETWEEN THEM becomes
VISIBLE rather than ASSUMED AWAY
Think of it as discovering that two folk remedies are, in fact, dosing the same active ingredient. Two traditions each arrived at a working preparation by trial and error, and practitioners can compare which works better on which ailment — but only experimentally, and only for the ailments tried. Identifying the common active ingredient converts the comparison: now you can reason about dose-response, predict where each preparation should fail, and relate both to a compound derived from first principles. The detail about the plus-one variant is the equivalent of checking that the ingredient is present in the preparation people actually take, not only in the historical recipe.
Key Concepts
- A missing derivation as a real limitation: without knowing what a scorer estimates, you cannot say when it fails or what a sensible variant would be. It is what makes the gap more than a matter of tidiness.
- Exactness rather than approximation: the score is the divergence, so theory about the divergence transfers to the scorer without an intervening limit.
- A shared framework as the enabler of comparison: both scorers landing in one basis is what makes them commensurable with each other and with other methods.
- Deriving the deployed variant: the plus-one IDF correction is what runs in practice. Theory aimed at the published idealisation can miss the systems it is meant to explain.
- Covering both formulations: it makes the difference between them a result rather than an assumption.
Framework Shift
Before (heuristics with no stated quantity):
score relevance with TF-IDF or BM25
-> no probabilistic derivation
-> cannot say what either estimates
-> cannot predict failures or justify variants
-> comparison is only experimental
After (exact divergences in a shared framework):
both are exact KL divergences between two probability models
-> the plus-one BM25 variant in practice is the one derived, with the
uncorrected original discussed alongside
-> a common basis clarifies what they measure
-> theoretical comparison with other IR methods becomes possible
From two effective scorers whose properties can only be observed, to two exact divergences in one framework whose properties can be reasoned about, the core shift is that a widely used heuristic can turn out to have been measuring something all along.
Expert Assessment
Problem choice: Excellent, and the framing of the gap is more useful than “this has no theory”. The paper identifies what is lost specifically — the ability to state what is estimated, and therefore to predict failure and compare theoretically — which is what makes the derivation worth having rather than merely satisfying.
Method maturity: Two choices raise this above a formal exercise. Deriving the deployed variant rather than the published formula is the decision that makes the theory applicable, and it is the kind of detail that is easy to skip and fatal to usefulness. And treating the two scorers as landing in one framework, rather than producing two independent derivations, is what delivers the comparison payoff. Whether the KL interpretation is the most natural one is a separate question, but an exact identification in a standard framework is a strong result regardless.
Experimental integrity: This is a theory contribution, so the claims are mathematical and the reading is of definitions. The identification is exact, which means it can be checked directly rather than tested statistically, and the inclusion of both BM25 variants means the result does not depend on silently choosing a convenient formulation. The limitation is scope of interpretation: an exact KL divergence says what the score computes, not that ranking by it is optimal for retrieval — usefulness to a user is a separate question the derivation does not answer, and the paper’s claim of clarifying “what they measure” is appropriately limited to that.
Writing quality: The abstract states the result, names the pragmatic variant choice, and states the payoff in one sequence, which is the right order for a derivation with a practical detail. Because the value depends on the specific form, writing out the two probability models explicitly in the paper’s own presentation would let a reader check the identification at a glance.
Verdict: strong accept — it supplies the missing probabilistic basis for two ubiquitous scorers, derives the version actually deployed rather than the idealised formula, and places both in one framework so theoretical comparison becomes possible.
Takeaways
- Distinguish useful heuristics from understood ones. A method can be tuned and deployed for decades without anyone being able to say what it estimates.
- Prefer exact identifications to asymptotic ones when available. If the score is the quantity, theory about the quantity transfers without a limit.
- Derive the variant that runs. A derivation of the published formula can leave the deployed system unexplained.
- Ask whether two methods measure the same thing. Placing them in a common framework turns an endless benchmark competition into a comparison.
论文: 2609.14016 作者: Ivan Silajev 分类: cs.IR, cs.CL
缺口
TF-IDF 与 BM25 是使用最广的两个「查询—文档相关性」打分方法,已有数十年历史、部署得到处都是。而两者都没有一个标准的概率论推导,来把它们论证为「统一框架内的一个统计方法」。
这种缺席有后果,而且容易被低估。一个有效的启发式可以被调参、被组合、被使用,但它无法被「推理」。你说不出它在估计什么,于是你说不出它应该在何时失效、一个合理的变体该是什么样、以及它与一个由不同原理推导出来的方法是什么关系。而比较被卡在实验层面:两个打分器可以被无休止地互相 benchmark,却没有任何说法能回答”它们是否在测量同一个量”。
两个无处不在、却没有推导的打分器
TF-IDF 与 BM25:使用最广的「查询—文档相关性」打分方法
数十年历史,部署得到处都是
|
v
而两者都「没有标准的概率论推导」,来把它们论证为
「统一框架内的一个统计方法」
|
v
[这种缺席为何要紧,且容易被低估]
有效的启发式可以被「调参、组合、使用」
但它「无法被推理」:
「你说不出它在估计什么」
-> 于是你说不出「它应该在何时失效」
-> 也说不出一个「合理的变体」该是什么样
-> 更说不出它与"一个由不同原理推导出来的方法""是什么关系"
|
v
而比较被卡在「实验层面」
-> 两个打分器可以被「无休止地互相 BENCHMARK」,
却没有任何说法能回答"它们「是否在测量同一个量」"
增量
一句话: 在这篇论文之前,TF-IDF 与 BM25 是没有概率解释的启发式;在这篇论文之后,两者都是两个概率模型之间的精确 KL 散度,从而获得共同基础,并使理论层面的比较成为可能。
核心机制
主张是:这两种打分方法都容许”被精确地解释为「两个概率模型之间的 Kullback-Leibler 散度」”。 这个结果有两条性质值得单独拎出来,因为它们决定了它有什么用。
它是精确的,不是近似的。 一个渐近的或近似的推导,解释的是”这些打分器在某个极限下为何表现合理”;而精确的同一性说的是”这个分数就是那个散度”,因此关于散度的陈述就是关于打分器的陈述。这就是”一个有解释的启发式”与”一个统计量”之间的差别。
而它是一个共同基础,这一点才解锁了比较。 论文所陈述的回报是一个TF-IDF 与 BM25 的共同理论基础,它澄清了它们在测量什么,并使它们可以与其他信息检索方法在理论上比较、而不只是实验上比较。所以贡献不是两个恰好都存在的独立推导;而是两者落在同一个框架里——这才让它们可以相互度量。
有一个细节值得强调,因为正是这类细节决定了推导是否适用于现实:论文处理的是「IDF 项中包含加一修正」的那个 BM25 变体,也就是实践中真正在用的那个;同时也讨论了不含该修正的原始形式。 这是一个刻意的选择:去推导被部署的方法,而不是已发表的理想化。推导一个没人跑的公式只有理论趣味;推导生产中在跑的那个变体,才让理论作用于真实系统——而论文把两者都覆盖,意味着它们之间的差别变成了一个结果,而不是被假设掉。
主张
「两种打分方法都容许被精确地解释为
「两个概率模型之间的 KULLBACK-LEIBLER 散度」」
两条性质决定了这个结果「有什么用」
[1] 它是「精确的」,不是近似的
<- 渐近的或近似的推导,解释的是"这些打分器
「在某个极限下」为何表现合理"
<- 「精确的同一性」说的是"这个分数「就是」那个散度"
-> 关于散度的陈述,就是关于打分器的陈述
<- "一个有解释的启发式"与"一个统计量"之间的差别
[2] 它是一个「共同基础」——正是它「解锁了比较」
所陈述的回报:一个 TF-IDF 与 BM25 的「共同理论基础」,它
「澄清了它们在测量什么」
「使它们可以与其他信息检索方法在「理论上」比较、
而不只是「实验上」比较」
<- 贡献「不是」两个恰好都存在的独立推导
<- 而是「两者落在同一个框架里」
-> 这才让它们「可以相互度量」
「一个细节值得强调」——它决定推导是否「适用于现实」
论文处理的是「IDF 项中包含加一修正」的 BM25 变体,
「也就是实践中真正在用的那个」
同时「也讨论」不含该修正的原始形式
<- 一个「刻意」的选择:去推导「被部署的方法」,
而不是「已发表的理想化」
<- 推导一个「没人跑的公式」只有理论趣味
<- 推导「生产中在跑的那个变体」,才让理论「作用于真实系统」
-> 而把两者都覆盖,意味着它们之间的差别变成了
「一个结果」,而不是被「假设掉」
可以用**“发现两种民间偏方其实在服同一种有效成分”来理解这件事: 两套传统各自靠试错摸出了一个能用的配方,而实践者可以在”哪个对哪种症状更管用”上做比较——但只能在实验上**、而且只在试过的那些症状上。识别出共同的有效成分就把比较变了:现在你可以就剂量—反应做推理、预测两种配方各自应当在何处失效、并把两者与一个从第一原理推导出来的化合物联系起来。 而”加一修正”那个细节,相当于去核实这种成分出现在人们实际服用的那个配方里,而不只是出现在历史文献的原始方子里。
关键概念
- 以”缺少推导”作为一项真实局限: 不知道一个打分器在估计什么,你就说不出它何时失效、也说不出一个合理变体该是什么样。这才让这个缺口不止于”整洁问题”。
- 精确而非近似: 分数就是那个散度,因此关于散度的理论无需经过一个极限就能迁移到打分器上。
- 以共享框架作为比较的使能条件: 两个打分器落在同一个基础里,才让它们彼此之间、以及与其他方法之间可通约。
- 推导被部署的那个变体: 加一的 IDF 修正才是实践中在跑的。瞄准已发表理想化的理论,可能恰恰错过它本要解释的系统。
- 覆盖两种形式: 它让两者之间的差别成为一个结果,而不是一个假设。
框架转变
之前(没有可陈述量的启发式):
用 TF-IDF 或 BM25 给相关性打分
-> 没有概率论推导
-> 说不出任何一者在估计什么
-> 无法预测失效、也无法为变体正名
-> 比较只在实验层面
之后(共享框架内的精确散度):
两者都是两个概率模型之间的精确 KL 散度
-> 被推导的是实践中那个加一的 BM25 变体,
并与未修正的原始形式一并讨论
-> 一个共同基础澄清了它们在测量什么
-> 与其他 IR 方法的理论比较成为可能
从”两个只能被观察其性质的有效打分器”,转变为”同一框架内两个可被推理的精确散度”,核心转变在于:一个被广泛使用的启发式,可能一直就在测量某个东西。
专家评审
选题眼光: 极好,而对缺口的框定比”这没有理论”更有用。 论文具体指出了失去的是什么——无法陈述它在估计什么,因而无法预测失效、无法做理论比较——这才让一个推导值得拥有,而不只是令人满足。
方法成熟度: 有两个选择把它提升到”形式练习”之上。 推导被部署的变体、而不是已发表的公式,是让理论可用的决定;这是那种很容易被跳过、却对有用性致命的细节。 而把两个打分器视为落在一个框架里、而不是产出两个独立推导,才交付了”比较”这份回报。KL 解释是否是最自然的解释是另一个问题,但在一个标准框架里给出精确同一性本身就是强结果。
实验诚意: 这是一项理论贡献,因此主张是数学性的,该读的是定义。 同一性是精确的,这意味着它可以被直接验算,而不必被统计检验;而把两个 BM25 变体都包含进来,意味着结论不依赖于”悄悄选一个方便的公式”。 局限在解释范围:一个精确的 KL 散度说的是”这个分数计算了什么”,而不是”按它排序对检索最优”——“对使用者的有用性”是一个推导没有回答的独立问题;论文”澄清它们在测量什么”这一主张也被恰当地限定在这里。
写作功力: 摘要按一个顺序给出了结果、点明了务实的变体选择、并陈述了回报——对一个带实践细节的推导来说,这个顺序是对的。 由于价值取决于具体形式,若能在论文自身的呈现里把那两个概率模型显式写出,读者就能一眼验算这个同一性。
判决: 强接收(Strong Accept) — 它为两个无处不在的打分器补上了缺失的概率基础,推导的是实际被部署的版本而不是理想化公式,并把两者放进同一个框架,使理论比较成为可能。
要点总结
- 区分有效的启发式与被理解的启发式。一个方法可以被调参、被部署数十年,而没有人能说出它在估计什么。
- 在可能时,优先精确同一性而不是渐近结论。如果分数就是那个量,关于该量的理论无需经过极限就能迁移。
- 推导在生产中跑的那个变体。对已发表公式的推导,可能让被部署的系统仍然无从解释。
- 问一句两个方法是否在测量同一个东西。把它们放进一个共同框架,能把一场无休止的 benchmark 竞赛变成一次比较。