Paper: 2609.14018 Authors: Zhong Cao, Shuying Chen Categories: cs.AI, cs.CL
The Gap
The paper opens with a sentence that states the whole problem: a correct diagnosis can still be reached for the wrong reasons. Getting the right answer and reasoning correctly are different achievements, and in clinical settings the difference is not academic — a right answer reached by luck carries no information about what happens on the next patient.
And the structure of the failure is specific to the domain. Every disease hypothesis creates obligations: key evidence must be checked, alternatives must be ruled out, contradictions must be resolved, useful tests must be considered, and closure must be justified. These are not stylistic preferences; they are what clinical reasoning consists of. A hypothesis that is entertained without its obligations being discharged has not been reasoned about — it has been asserted.
And existing evaluations miss exactly this. Current evaluations of medical LLMs mostly focus on final answers, local steps, or isolated facts. Each of the three looks past the obligations in a different way: final answers ignore the process, local steps do not track obligations accumulated across a trajectory, and isolated facts do not connect to any hypothesis at all. So the quantity that matters — whether commitments were honoured — is not measured in any of them.
A CORRECT DIAGNOSIS CAN BE REACHED FOR THE WRONG REASONS
GETTING THE RIGHT ANSWER and REASONING CORRECTLY are DIFFERENT
ACHIEVEMENTS
<- in CLINICAL settings the difference is NOT ACADEMIC: a right
answer reached by LUCK carries NO INFORMATION about what
happens ON THE NEXT PATIENT
|
v
[AND THE STRUCTURE OF THE FAILURE IS SPECIFIC TO THE DOMAIN]
EVERY DISEASE HYPOTHESIS CREATES OBLIGATIONS:
key evidence MUST BE CHECKED
alternatives MUST BE RULED OUT
contradictions MUST BE RESOLVED
useful tests MUST BE CONSIDERED
closure MUST BE JUSTIFIED
<- NOT stylistic preferences: this is WHAT CLINICAL REASONING
**CONSISTS OF**
-> a hypothesis entertained WITHOUT its obligations discharged has
NOT BEEN REASONED ABOUT -- it has been ASSERTED
[AND EXISTING EVALUATIONS MISS EXACTLY THIS]
current evaluations of medical LLMs mostly focus on
FINAL ANSWERS | LOCAL STEPS | ISOLATED FACTS
<- each looks past the obligations in a DIFFERENT way:
FINAL ANSWERS ignore the PROCESS
LOCAL STEPS do NOT TRACK OBLIGATIONS ACCUMULATED ACROSS A
TRAJECTORY
ISOLATED FACTS do NOT CONNECT TO ANY HYPOTHESIS AT ALL
-> the quantity that matters -- WHETHER COMMITMENTS WERE HONOURED --
is NOT MEASURED in ANY of them
The Increment
One sentence: Before this paper, medical-LLM evaluation scored answers, steps or facts; after it, a disease-centric framework links free-form reasoning to structured disease profiles and shows many diagnostic errors are broken commitments made earlier, not isolated mistakes.
Core Mechanism
The design move is to give the reasoning something structured to be checked against: VeriDx links free-form diagnostic reasoning to structured disease profiles. The pairing is what makes obligations checkable. Free-form reasoning supplies the trajectory; the structured profile supplies what a hypothesis of that disease owes. Neither alone yields a verification: unstructured reasoning has nothing to compare to, and a profile alone says nothing about what the model actually did.
And the output is a three-way status per hypothesis, which is the right granularity: it tracks whether each hypothesis is satisfied, unresolved, or violated its clinical obligations. The distinction between the last two is the important one. Unresolved means an obligation was left open — the model never got to it. Violated means it was contradicted — the model did something inconsistent with its own hypothesis. Those have different remedies and different severities, and collapsing them into “wrong” would lose the distinction.
The taxonomy of what gets exposed is concrete, and each item is a different failure of commitment: missing critical tests, unresolved differentials, ignored contradictions, unsupported claims, and premature closure. Reading them together, they are all cases of a hypothesis-with-obligations where the obligations were not discharged — which is the paper’s thesis in list form. Premature closure is worth singling out because it is a recognised clinical failure mode with real consequences: settling on a diagnosis before the alternatives have been excluded.
The instantiation is appropriately narrow and demanding: complex respiratory diagnosis, using guideline-derived disease profiles and expert-annotated longitudinal cases. Three properties of that choice matter. Guideline-derived profiles give the obligations external authority rather than the authors’ opinion. Expert-annotated cases give ground truth for the annotation. And longitudinal cases are where commitments accumulated across a trajectory become visible — a single time point would not show a differential left unresolved over time.
And the headline finding is a reinterpretation, which is what makes the framework worth building: many diagnostic errors are not isolated mistakes, but broken commitments made earlier in the reasoning process. So the error is not located where the wrong answer appears. It was created earlier, when an obligation went undischarged, and the final wrong answer is a downstream consequence. That reframing changes what a fix would target — not the final decision, but the point in the trajectory where the commitment was dropped.
[THE DESIGN MOVE: GIVE THE REASONING SOMETHING STRUCTURED TO BE
CHECKED AGAINST]
VERIDX LINKS FREE-FORM DIAGNOSTIC REASONING TO STRUCTURED DISEASE
PROFILES
<- the PAIRING is what makes OBLIGATIONS CHECKABLE:
FREE-FORM REASONING supplies THE TRAJECTORY
THE STRUCTURED PROFILE supplies WHAT A HYPOTHESIS OF THAT
DISEASE **OWES**
<- NEITHER ALONE YIELDS A VERIFICATION:
UNSTRUCTURED REASONING has NOTHING TO COMPARE TO
A PROFILE ALONE says NOTHING ABOUT WHAT THE MODEL ACTUALLY DID
[AND THE OUTPUT IS A THREE-WAY STATUS PER HYPOTHESIS -- THE RIGHT
GRANULARITY]
it tracks whether each hypothesis is SATISFIED, UNRESOLVED, or
VIOLATED ITS CLINICAL OBLIGATIONS
<- the distinction between the LAST TWO is the IMPORTANT ONE:
UNRESOLVED = an obligation LEFT OPEN -- the model NEVER GOT TO
IT
VIOLATED = it was CONTRADICTED -- the model did something
INCONSISTENT WITH ITS OWN HYPOTHESIS
<- DIFFERENT REMEDIES and DIFFERENT SEVERITIES
-> collapsing them into "WRONG" would LOSE THE DISTINCTION
[THE TAXONOMY OF WHAT GETS EXPOSED IS CONCRETE -- each item a
DIFFERENT FAILURE OF COMMITMENT]
MISSING CRITICAL TESTS
UNRESOLVED DIFFERENTIALS
IGNORED CONTRADICTIONS
UNSUPPORTED CLAIMS
PREMATURE CLOSURE
<- read together, ALL are cases of a hypothesis-with-obligations
where the OBLIGATIONS WERE NOT DISCHARGED
-> the paper's thesis IN LIST FORM
<- PREMATURE CLOSURE deserves singling out: a RECOGNISED CLINICAL
FAILURE MODE with REAL CONSEQUENCES -- settling on a diagnosis
BEFORE the alternatives have been excluded
[THE INSTANTIATION IS APPROPRIATELY NARROW AND DEMANDING]
COMPLEX RESPIRATORY DIAGNOSIS
GUIDELINE-DERIVED disease profiles
EXPERT-ANNOTATED longitudinal cases
<- GUIDELINE-DERIVED profiles give the obligations EXTERNAL
AUTHORITY rather than THE AUTHORS' OPINION
<- EXPERT-ANNOTATED cases give GROUND TRUTH for the ANNOTATION
<- LONGITUDINAL cases are where COMMITMENTS ACCUMULATED ACROSS A
TRAJECTORY become VISIBLE
-> a SINGLE TIME POINT would NOT SHOW A DIFFERENTIAL LEFT
UNRESOLVED OVER TIME
[AND THE HEADLINE FINDING IS A REINTERPRETATION -- which is what makes
the framework WORTH BUILDING]
MANY DIAGNOSTIC ERRORS ARE NOT ISOLATED MISTAKES, BUT BROKEN
COMMITMENTS MADE EARLIER IN THE REASONING PROCESS
-> the ERROR IS NOT LOCATED WHERE THE WRONG ANSWER APPEARS
-> it was CREATED EARLIER, when an OBLIGATION WENT UNDISCHARGED
-> the FINAL WRONG ANSWER is a DOWNSTREAM CONSEQUENCE
-> that REFRAMING CHANGES WHAT A FIX WOULD TARGET: NOT the final
decision, but THE POINT IN THE TRAJECTORY WHERE THE COMMITMENT
WAS DROPPED
Think of it as auditing a diagnosis the way you would audit a differential checklist rather than the discharge note. The note says the right thing, and reading it tells you nothing about how the decision was made. The checklist tells you what was owed: was the test ordered, was the alternative excluded, was the contradiction addressed, was the closure argued. Note that the checklist distinguishes “never got to it” from “contradicted yourself”, which matters because they are different problems — an omission to be filled versus an inconsistency to be resolved. And the reframing is the point: if most errors are items left unchecked earlier, then improving the final decision is treating a symptom, and the intervention belongs where the checklist was abandoned.
Key Concepts
- Obligations created by a hypothesis: evidence to check, alternatives to exclude, contradictions to resolve, tests to consider, closure to justify. They are what clinical reasoning consists of, so a hypothesis without them has been asserted rather than reasoned about.
- Why existing evaluations miss it: answers ignore process, steps miss accumulated obligations, facts connect to no hypothesis. Each looks past the commitment structure differently.
- Linking free-form reasoning to structured profiles: the pairing that makes obligations checkable. The trajectory supplies behaviour, the profile supplies what is owed.
- Satisfied, unresolved, violated: the three-way status, where the last two must stay separate. An omission and an inconsistency need different remedies.
- Broken commitments rather than isolated mistakes: the reinterpretation. The error is created earlier than the wrong answer appears, which changes where a fix belongs.
- Guideline-derived profiles and longitudinal cases: external authority for the obligations, and a time span long enough for accumulated commitments to be visible.
Framework Shift
Before (score answers, steps or facts):
evaluate medical reasoning by final answer, local step, or fact recall
-> a correct diagnosis reached by luck scores the same as one reasoned to
-> obligations accumulated across a trajectory are not tracked
-> errors appear where the wrong answer does
After (verify hypothesis obligations):
free-form reasoning linked to structured disease profiles
each hypothesis marked satisfied, unresolved, or violated
-> missing tests, unresolved differentials, ignored contradictions,
unsupported claims, premature closure
-> on guideline-derived profiles and expert-annotated longitudinal cases
-> many errors are broken commitments made earlier, not local mistakes
From asking whether the diagnosis came out right, to asking whether each hypothesis discharged the obligations it created, the core shift is that clinical correctness is a property of a trajectory’s commitments rather than of its final line.
Expert Assessment
Problem choice: Excellent, and the opening observation is the one that justifies the framework. “A correct diagnosis can still be reached for the wrong reasons” is exactly the failure a final-answer evaluation cannot see, and in a clinical setting the practical stakes of that blindness are high — a lucky right answer says nothing about the next patient.
Method maturity: The design’s logic is sound: obligations become checkable only when free-form reasoning is paired with something structured describing what is owed, which is why neither half suffices alone. The three-way status is the right granularity, and insisting on the unresolved-versus-violated distinction is what keeps the output diagnostic rather than merely evaluative. Deriving profiles from guidelines rather than from the authors’ judgement gives the obligations external authority, and using longitudinal cases is necessary for obligations that accumulate over time to be observable at all.
Experimental integrity: The headline finding is a reinterpretation rather than a score, and it is the more valuable form: locating errors earlier than they surface changes where an intervention belongs. The scope is narrow — one clinical domain, respiratory diagnosis — which is appropriate for a framework paper that needs guideline-derived profiles to exist. The limitation is that the verification depends on the quality of the structured profiles and the annotation, so disagreements about clinical obligations would propagate into the metric, and the paper’s use of guideline-derived profiles is the mitigation rather than a demonstration that obligation extraction is uncontroversial.
Writing quality: The obligation list is stated up front and every subsequent component maps onto it, which makes the framework easy to follow. Because the finding’s practical value is a change of target, a short example of one trajectory — the hypothesis, the obligation left unresolved, and the downstream error it produced — would make the reframing immediately concrete for a clinical reader.
Verdict: strong accept — it identifies the failure mode that answer-level evaluation cannot see, makes hypothesis obligations checkable by pairing free-form reasoning with structured profiles, and reports that errors are created earlier than they appear.
Takeaways
- Ask whether a right answer was reasoned to. A correct conclusion reached by luck carries no information about the next case.
- Track obligations, not just steps. Local step scoring misses commitments accumulated and abandoned across a trajectory.
- Separate omission from contradiction. “Never got to it” and “contradicted yourself” need different remedies, and one label loses the difference.
- Locate errors where they were created. If most failures are commitments broken earlier, fixing the final decision treats a symptom.
论文: 2609.14018 作者: Zhong Cao, Shuying Chen 分类: cs.AI, cs.CL
缺口
论文开篇的一句话就把整个问题说清了:一个正确的诊断,仍然可能是靠错误的理由得出的。 得到正确答案与正确推理是两种不同的成就;而在临床场景里,这个差别不是学术性的——靠运气得到的正确答案,对”下一个病人会怎样”不携带任何信息。
而失效的结构是这个领域特有的。每一个疾病假设都会产生「义务」:必须核查关键证据、必须排除替代诊断、必须解决矛盾、必须考虑有用的检查、并且必须为收束给出理由。 这些不是风格偏好;它们就是临床推理所「包含」的内容。一个被提出、却未履行其义务的假设,并未被推理——它只是被断言。
而既有评测恰恰漏掉了这一点。 当前对医学 LLM 的评测主要聚焦于最终答案、局部步骤或孤立事实。 这三者各自以不同方式越过了那些义务:最终答案忽略过程;局部步骤不追踪跨轨迹累积的义务;孤立事实根本不与任何假设相连。所以真正要紧的那个量——承诺是否被履行——在它们之中一个都没有被测量。
一个正确的诊断,可能是靠错误的理由得出的
「得到正确答案」与「正确推理」是「两种不同的成就」
<- 在临床场景里,这个差别「不是学术性的」:
靠运气得到的正确答案,
对"下一个病人会怎样"「不携带任何信息」
|
v
[而失效的结构是这个领域特有的]
「每一个疾病假设都会产生义务」:
必须「核查」关键证据
必须「排除」替代诊断
必须「解决」矛盾
必须「考虑」有用的检查
并且必须为「收束」给出理由
<- 这些「不是风格偏好」:它们「就是临床推理所包含的内容」
-> 一个被提出、却未履行其义务的假设,
「并未被推理——它只是被断言」
[而既有评测恰恰漏掉了这一点]
当前对医学 LLM 的评测主要聚焦于
「最终答案 | 局部步骤 | 孤立事实」
<- 三者各自以「不同方式」越过了那些义务:
「最终答案」忽略「过程」
「局部步骤」不追踪「跨轨迹累积的义务」
「孤立事实」根本不与任何假设相连
-> 真正要紧的那个量——「承诺是否被履行」——
在它们之中「一个都没有被测量」
增量
一句话: 在这篇论文之前,医学 LLM 的评测打分的是答案、步骤或事实;在这篇论文之后,一个以疾病为中心的框架把自由形式的推理与结构化疾病画像连起来,并表明许多诊断错误是更早时被打破的承诺,而不是孤立的失误。
核心机制
设计上的动作,是给推理一个「结构化的东西」去对照:VeriDx 把自由形式的诊断推理与结构化的疾病画像连起来。 正是这种配对让义务可被核查。自由形式的推理提供轨迹;结构化画像提供”关于该疾病的假设欠着什么”。任何一半单独都产不出验证:无结构的推理没有可对照的东西,而只有画像则完全说不出模型实际做了什么。
而输出是每个假设的「三态」——粒度是对的:它追踪每个假设是「履行了」「悬置未决」、还是「违反了」它的临床义务。后两者之间的区分才是要紧的那个。“悬置”意味着某项义务被留在那里——模型从未走到它。“违反”意味着它被矛盾了——模型做出了与它自己的假设不一致的事。这两者需要不同的补救、也有不同的严重程度,而把它们压成”错了”就会丢掉这个区别。
被暴露出来的清单很具体,而每一项都是一种不同的承诺失效**:缺失的关键检查、未解决的鉴别诊断、被忽视的矛盾、缺乏支撑的主张、以及过早收束。 把它们放在一起读,全部都是”一个有义务的假设、其义务未被履行”的情形——这就是论文的论点以清单形式呈现。过早收束值得单独点出,因为它是一个被承认的临床失效模式、后果真实:在替代诊断尚未被排除之前就定下诊断。
所实例化的场景范围适当狭窄、要求也高:复杂呼吸系统诊断,使用由指南推导出的疾病画像与专家标注的纵向病例。 这个选择有三条性质要紧。由指南推导的画像给义务外部权威,而不是作者的意见。专家标注的病例为标注提供真值。而纵向病例才是跨轨迹累积的承诺变得可见的地方——单一时间点不可能显示一个鉴别诊断被搁置了很长时间。
而头条发现是一次「重新解释」,也正是让这个框架值得构建的原因:许多诊断错误不是孤立的失误,而是推理过程中更早时被打破的承诺。 所以错误并不位于错误答案出现的地方。它是在更早被造出来的——当某项义务未被履行时——而最终的错误答案是下游的后果。这个重构改变了修复该瞄准什么:不是最终决定,而是轨迹中那个承诺被丢下的位置。
[设计上的动作:给推理一个结构化的东西去对照]
VERIDX 把自由形式的诊断推理与结构化的疾病画像连起来
<- 正是这种「配对」让义务「可被核查」:
自由形式的推理提供「轨迹」
结构化画像提供"关于该疾病的假设「欠着什么」"
<- 任何一半单独都产不出验证:
无结构的推理「没有可对照的东西」
只有画像则「完全说不出模型实际做了什么」
[输出是每个假设的「三态」——粒度是对的]
追踪每个假设是「履行了」「悬置未决」、还是「违反了」
它的临床义务
<- 后两者之间的区分才是「要紧的那个」:
「悬置」= 某项义务「被留在那里」——模型「从未走到它」
「违反」= 它「被矛盾了」——模型做出了
「与它自己的假设不一致」的事
<- 「需要不同的补救、也有不同的严重程度」
-> 把它们压成"错了"就会「丢掉这个区别」
[被暴露出来的清单很具体——每一项都是一种「不同的承诺失效」]
缺失的关键检查
未解决的鉴别诊断
被忽视的矛盾
缺乏支撑的主张
过早收束
<- 放在一起读,全部都是"一个有义务的假设、
其义务未被履行"的情形
-> 论文的论点「以清单形式」呈现
<- 「过早收束」值得单独点出:一个「被承认的临床失效模式、
后果真实」——在替代诊断尚未被排除之前就定下诊断
[所实例化的场景范围适当狭窄、要求也高]
「复杂呼吸系统诊断」
由「指南推导出」的疾病画像
「专家标注」的纵向病例
<- 由指南推导的画像给义务「外部权威」,而不是作者的意见
<- 专家标注的病例为标注提供「真值」
<- 「纵向」病例才是「跨轨迹累积的承诺」变得可见的地方
-> 「单一时间点」不可能显示一个鉴别诊断被搁置了很长时间
[而头条发现是一次「重新解释」——也正是让这个框架值得构建的原因]
「许多诊断错误不是孤立的失误,
而是推理过程中更早时被打破的承诺」
-> 错误「并不位于错误答案出现的地方」
-> 它是在「更早」被造出来的——当某项义务未被履行时
-> 最终的错误答案是「下游的后果」
-> 这个重构改变了「修复该瞄准什么」:不是最终决定,
而是「轨迹中那个承诺被丢下的位置」
可以用**“像审计一份鉴别诊断清单、而不是审计出院小结那样去审计一次诊断”来理解这件事: 小结上写的是对的,而读它完全告诉不了你这个决定是怎么做出的。 清单告诉你欠着什么**:检查开了没有、替代诊断排除了没有、矛盾处理了没有、收束论证过没有。 注意那张清单把”从未走到它”与”自相矛盾”区分开来——这很重要,因为它们是不同的问题:一个是待补的遗漏,一个是待解的不一致。 而重构才是要点:如果大多数错误是更早时未被勾选的条目,那么”改进最终决定”就是在治标,而干预应当落在清单被放弃的那个位置。
关键概念
- 由假设产生的义务: 要核查的证据、要排除的替代、要解决的矛盾、要考虑的检查、要为收束给出的理由。它们就是临床推理的内容,所以一个没有这些的假设是被断言的,而不是被推理的。
- 既有评测为何漏掉它: 答案忽略过程、步骤漏掉累积的义务、事实不与任何假设相连。三者各自以不同方式越过了那套承诺结构。
- 把自由形式推理与结构化画像连起来: 这种配对让义务可被核查。轨迹提供行为,画像提供欠着什么。
- 履行、悬置、违反: 三态,其中后两者必须保持分开。一个遗漏与一个不一致需要不同的补救。
- 是被打破的承诺,而不是孤立的失误: 这次重新解释。错误在被造出来的位置上早于错误答案的出现——这改变了修复应当落在哪里。
- 由指南推导的画像与纵向病例: 给义务以外部权威,并给出足够长的时间跨度,让累积的承诺得以可见。
框架转变
之前(给答案、步骤或事实打分):
用最终答案、局部步骤或事实回忆来评估医学推理
-> 靠运气得出的正确诊断与推理得出的得分相同
-> 跨轨迹累积的义务未被追踪
-> 错误出现在错误答案出现的地方
之后(核查假设的义务):
自由形式推理与结构化疾病画像相连
每个假设被标记为已履行、悬置或违反
-> 缺失检查、未解决的鉴别、被忽视的矛盾、
缺乏支撑的主张、过早收束
-> 在指南推导的画像与专家标注的纵向病例上
-> 许多错误是更早时被打破的承诺,而不是局部失误
从”问诊断结果对不对”,转变为”问每一个假设是否履行了它所产生的义务”,核心转变在于:临床正确性是「轨迹的承诺」的属性,而不是「它最后一行」的属性。
专家评审
选题眼光: 极好,而开篇那个观察正是让框架获得正当性的那个。 “一个正确的诊断仍可能靠错误的理由得出”恰恰是”最终答案式评测”看不见的失效;而在临床场景里,这种盲区的实际代价很高——一个靠运气得到的正确答案,对下一个病人什么也没说。
方法成熟度: 设计的逻辑是站得住的:只有当自由形式的推理与”描述欠着什么”的结构化东西配对时,义务才变得可核查——这就是任何一半单独都不够的原因。三态是正确粒度;而坚持区分”悬置”与”违反”,才让输出是诊断性的、而不只是评价性的。把画像从指南而非作者判断中导出,给义务以外部权威;而使用纵向病例,对”跨时间累积的义务”能否被观察到来说是必要的。
实验诚意: 头条发现是一次重新解释而不是一个分数,而这是更有价值的形态:把错误定位到它们浮出水面之前,改变了干预应当落在哪里。 范围很窄——一个临床领域、呼吸系统诊断——这对一篇需要”指南推导画像”存在的框架论文来说是恰当的。 局限是:验证依赖结构化画像与标注的质量,因此关于临床义务的分歧会传导进这个指标;而论文使用”指南推导画像”是缓解措施,而不是”义务抽取没有争议”的演示。
写作功力: 义务清单被放在最前面,而后续每个组件都对应到它,这让框架容易跟读。 由于这个发现的实际价值是一次目标的改变,若能给一条轨迹的简短示例——那个假设、被悬置的那项义务、以及它所造成的下游错误——会让这次重构对临床读者立刻具体起来。
判决: 强接收(Strong Accept) — 它识别出”答案层评测看不见”的失效模式,通过把自由形式推理与结构化画像配对使假设义务可被核查,并报告出”错误在被看到之前就被造出来了”。
要点总结
- 问一句这个正确答案是否”推理出来”的。靠运气得到的正确结论,对下一个病例不携带任何信息。
- 追踪义务,而不只是步骤。局部步骤打分漏掉了跨轨迹累积并被放弃的承诺。
- 把遗漏与矛盾分开。“从未走到它”与”自相矛盾”需要不同的补救,而一个标签会丢掉这个区别。
- 把错误定位到它们被造出来的地方。如果大多数失效是更早时被打破的承诺,那么修最终决定就是在治标。