scDNM-VAE makes its clustering function inspectable: for each cluster and latent coordinate, the signed weight determines whether the sigmoid gate rises or falls, its magnitude controls steepness, and the threshold-to-weight ratio gives the transition midpoint. This is a genuine architectural property, not a separately fitted explanation model. However, inspectability of parameters does not by itself establish that the parameters correspond to stable biological causes, because the latent coordinates remain seed-dependent and non-identifiable at the gene level.
The benchmark covers four labeled datasets: PBMC3k (2,638 cells), Zeisel Cortex (3,005), Human Heart Cell Atlas subsample (18,641), and Paul15 (2,730), with five seeds per method. scDNM-VAE has its clearest advantage on PBMC3k: mean ARI 0.636 versus 0.449 for scVI+KMeans, while it loses clearly on Zeisel: 0.528 versus 0.700. Heart performance is similar in mean ARI (0.753 versus 0.773), although scDNM-VAE is less variable across seeds (0.031 versus 0.066). Paul15 is also close (0.296 versus 0.272). Thus, the defensible conclusion is dataset-dependent competitiveness, not consistent superiority.
Figure: observed values reported by the paper; no values are inferred.
The MLP-DEC comparison is informative but confounded. The two models differ not only in clustering head, but also in variational versus deterministic encoding, KL regularization, and the threshold-refinement phase. Therefore, lower MLP-DEC ARI cannot be attributed specifically to signed dendritic gating. Likewise, the reported geometry-versus-ARI contrast supports caution about latent compactness, but does not prove that the dendritic head is biologically better.
Top-three absolute-weight ablation caused more reassignment than random dimensions on every dataset, but the excess was modest: 1.01β1.39-fold, with no separation from the other-cluster control on Zeisel. Randomly removing three dimensions already reassigned 34.6% of PBMC3k cells and 90.9% of Paul15 cells. This supports distributed decision relevance, not sparse causal attribution. Stage 2 gene modules are useful descriptive summaries, but they are trained against predicted clusters and cluster-derived differential-expression signatures; their activation alignment is therefore partly built into the objective rather than independent validation. The marker-overlap results are also conditional on eligibility thresholds and harmonized label vocabularies.
This is a promising methodological preprint whose most credible contribution is parameter-level access to a nonlinear clustering rule. Its empirical case for superior clustering is mixed, and its causal interpretability claim is intentionally narrower than a gene-level explanation claim. Before strong adoption, the decisive missing evidence is component-wise ablation: hold the VAE, KL term, initialization, and training schedule constant while independently varying only the clustering head; repeat faithfulness analyses across seeds; evaluate count-based reconstruction; test automatic K selection; and validate gate stability across datasets, hardware, and biologically independent cohorts. The paper itself reports that code will become public upon publication or earlier on request, so full independent reproduction is not yet established.
Know what changed, what holds up, and what remains uncertain. Every Friday. No ads.