Press ESC to close Press ⌘K or Ctrl+K to open

Recent Posts

AGENTQ: A Checkpoint Can Pass Every Audit and Misbehave Once Quantized

AGENTQ:一个检查点可以通过所有审计,却在量化后行为不端

Measuring the Creativity of Frontier LLMs in Automated Research

度量前沿大模型在自动化科研中的「创造力」

Performative Stability in Nearly the Minimum Number of Deployments

以「近乎最少的部署次数」达到表演性稳定

TF-IDF and BM25 Are Exact KL Divergences

TF-IDF 与 BM25 都是精确的 KL 散度

VeriDx: A Correct Diagnosis Can Be Reached for the Wrong Reasons

VeriDx:一个正确的诊断,可能是靠错误的理由得出的

A Ranking Approach for Measuring Calibration

一种用「排序」来度量校准的方法

Benign Landscapes and Worst-Case Hardness Can Coexist

「良性损失地形」与「最坏情形困难性」可以共存

Embodied-BenchForge: Verify Each Artifact, or Defects Propagate

Embodied-BenchForge:逐个产物做验证,否则缺陷会一路传播

SAS: Stop Distilling Attention, Optimise the Ranking Instead

SAS:别再蒸馏注意力分布,去优化「上下文排序」

Type Diversity Explains Why Structural Generalisation Looks Harder

「类型多样性」解释了「结构泛化为何看起来更难」

Artificial Id: When Drive Emerges From Persistence Alone

人工本我:当「驱动力」仅从「持续性」中涌现

MoEs Overfit More to Repeated Data, and Sparsity Is What Drives It

MoE 对「重复数据」过拟合更严重,而驱动力是「稀疏度」

View all posts →

Series