The paper's evidence chain is unusually rigorous: each control was forced by an anomaly, not planned a priori. (1) Gene-identity leakage: under random splits, 98.3% of test variants share genes with training data; under LOGO-CV this drops to zero, collapsing the tokenization gap (Fig. 3a). (2) Memorization canary: a one-hot (position, codon) lookup achieves AUC 0.891 under random splits and 0.739 under LOGO-CV, exceeding every neural model's 0.507β0.564 . (3) Probe implementation artifact: the surviving +4.7 pp residual decomposes into six defensible defaults; seed 42 exceeds all 20 alternative seeds. (4) Best-epoch selection bias: regression gains of +6.4 to +23.7 pp reverse to -37.0 to -10.8 pp under validation-based early stopping.
The asymmetric collapse confirms that the thinner the true signal, the larger the share of apparent advantage that leakage suppliesβconsistent with I(Ο;Y|A) β 0.04 bits vs I(A;Y) β 0.13 bits.
Strengths: Causal matched-pair tokenization ablation with synthetic-CDS negative control; exploratory analyses labeled as such; direction of effect (not magnitude) is the primary claimβrobust to multiple testing. Limitations acknowledged: restricted to frozen-embedding probing on human ClinVar SynPath (n=2,840); LOGO-CV does not control gene-family or GC-strata leakage; two-point scale comparison insufficient. Unaddressed blindspot: whether the FungalExpr residual +2.8β6.5 pp gain is genuine codon-usage signal or species-level leakage remains open. A codon model clearing the 0.739 canary under leakage-controlled evaluation would falsify the central claim.
Know what changed, what holds up, and what remains uncertain. Every Friday. No ads.