Yan, Harper, and Li review 34 SVG methods and organize them by biological target and statistical machinery. The key distinction is between overall SVGs—non-random spatial expression without external labels—cell-type-specific SVGs—spatial variation within a cell type—and spatial-domain-marker SVGs—higher expression in a predefined domain. Because these hypotheses differ, comparing methods by the number of genes detected can compare different scientific tasks rather than competing solutions.
The review’s warnings are supported by an independent benchmark of 31 real datasets spanning nine technologies plus nine scDesign3 simulations. That study found limited agreement in significant calls, expression-level bias in several statistics, sparsity sensitivity in neighborhood-based methods, and imperfect FDR calibration for some Giotto and Moran’s-I analyses. It also found that approximately 900–1100 selected SVGs gave the best clustering performance in one E9.5 mouse-embryo analysis, illustrating that downstream utility is not equivalent to statistical significance.
This is a strong conceptual review, but its conclusions are not themselves a systematic effect-size synthesis: the supplied paper text does not report a preregistered search strategy, formal risk-of-bias assessment, quantitative aggregation of method performance, or a reproducible scoring framework for the 34 methods. The “34” count is consequently time-dependent and may omit unpublished, newly released, or difficult-to-classify tools. The authors do, however, explicitly address several apparent weaknesses—technology and tissue differences, multi-sample analysis, double-dipping, negative controls, and benchmark design—so these should be viewed as proposed research priorities rather than overlooked limitations. Their criticism of double-dipping is especially important where domains are learned and then reused for marker testing, because feature selection and inference can be statistically dependent. Confidence: high for the taxonomy’s conceptual usefulness; moderate for universal claims about method superiority or validity.
Bottom line: use the taxonomy to define the biological question before choosing a detector; treat rankings, p-values, and downstream clustering as different evidence types; and require category-matched, platform-aware validation rather than assuming one universally best SVG method.
Know what changed, what holds up, and what remains uncertain. Every Friday. No ads.