Marek et al. (Nature, 2022) pooled three large neuroimaging datasetsβABCD (n = 3,928 with strict motion denoising; 9β10 y, 21 sites), HCP (n β 900; 22β35 y, single scanner, ~60 min RSFC), and UK Biobank (subsampled to n = 900 from 32,572; 40β69 y, 6 min RSFC)βtotaling ~50,000 individuals to quantify how BWAS effect sizes and reproducibility scale with sample size. Across 11 million univariate associations, the median effect size was |r| β 0.01, and only the top 1% reached |r| > 0.06; the largest association that replicated out-of-sample was |r| = 0.16 .
At n = 25, the 99% confidence interval for univariate associations was r Β± 0.52βeffects can be strongly inflated by chance. Even in split halves (n = 1,964), top 1% effects remained inflated by r = 0.07 (78%) on average. At n = 1,000, false-negative rates were 75β100% and half of significant associations were inflated by β₯100%; maximum power was 0.68 at n = 3,928. Replication rates were ~5% at typical sample sizes (n < 500) and only 25% at n = 1,964 (P < 0.05). Bonferroni-corrected detection of the top 1% effects required n = 9,500 for 80% power vs n = 2,200 uncorrected .
Note: replication points are reported values (25%, 5%); intermediate points are illustrative interpolation for visualization onlyβsee original Fig. 3e for exact curves.
RSFC-based multivariate prediction (SVR) achieved stronger out-of-sample associations (maximum RSFCβcrystallized intelligence r_pred = 0.39 vs univariate r = 0.16), yet in-sample estimates remained inflated by Ξr = β0.29 even at n β 2,000. Out-of-sample multivariate associations correlated robustly with univariate effect sizes (r = 0.79) .
Size-matched subsamples (n = 900) yielded comparable RSFCβcognitive ability distributions: HCP |r| > 0.12, ABCD |r| > 0.11, UKB |r| > 0.09 (top 1%), suggesting effect-size limitations are not scanner- or site-specific but universal to BWAS with current methods .
The paper is transparent about its own limits: sociodemographic covariate adjustment decreased the strongest effects (top 1% Ξr = β0.014), RSFC reliability varied widely across datasets (UKB r = 0.39 vs HCP r = 0.79), and theoretical maximum effect sizes are capped by biology and MRI physics. The observational design precludes causal inference, and the authors appropriately note that small-sample, within-person, lesion, and interventional designs remain essential for other neuroimaging questions. A notable limitation is that conclusions rest on RSFC and cortical thicknessβother modalities (e.g., task fMRI activations showed matched distributions here) and demographic confounds may behave differently. Recent work building directly on this paper shows that study-design features (longitudinal designs, increased between-subject age variability) can boost standardized effect sizes substantially (~380% increase for GMV-age longitudinal vs cross-sectional RESI), providing a partial counterpoint to the 'thousands required' framing for specific covariate associations . The authors also have declared financial interests in NOUS Imaging Inc. (FIRMM motion-monitoring software), managed by their institutions.
Confidence in the core conclusion is high (strong evidence, ~50,000 participants, rigorous motion denoising, cross-dataset replication, publicly available code at gitlab.com/DosenbachGreene/bwas). The claim would be challenged if a large-scale multi-site BWAS demonstrated robust, stable, non-inflated associations at typical sample sizes, or if improved measurement reliability (e.g., longer RSFC, multi-echo, better phenotyping) substantially increased effect sizes in thousands-of-participant data.
Simulated illustration based on the paper's reported n=25 CI width (r Β± 0.52) scaled as 1/βn and analytic power for |r| = 0.06 at P < 0.05. Assumes bivariate normal correlations, no attrition, and approximates the paper's bootstrap behavior; not observed data.
Generated scientific data; not direct experimental measurements.Know what changed, what holds up, and what remains uncertain. Every Friday. No ads.