The paper reports that most methods are poorly calibrated under null conditions (spots permuted to remove spatial variability), producing p-values that are either too conservative or too liberal, while SPARK and SPARK-X are among the best-calibrated in their evaluation.
Skeptical implication: If p-values are miscalibrated, then βsignificant SVG setsβ chosen by thresholds like p<0.05 may yield biased false-positive/false-negative rates, which can contaminate downstream spatial domain clustering and biological downstream interpretation. The paperβs own recommendation is to rely on ranking thresholds (e.g., top-K genes) rather than p-value significance.
SpatialDE is a classic Gaussian-process-based spatial variability approach that underlies much subsequent work.
The benchmarking paperβs calibration warning is especially relevant here: even when GP-based models can detect spatial variability, correct p-value calibration can still fail depending on test construction and data preprocessingβso ranking-only conclusions should not be over-extended to inferential claims.
Know what changed, what holds up, and what remains uncertain. Every Friday. No ads.