Why BGPT?
logo

Review Claim by Claim

Check what supports each statement: experiments, reported results, scope, and limitations.Know what the science actually supports before you trust the answer.

Press Enter ↡ to review paper


     Quick Explanation



    Marek et al. (2022) show that brain-wide associations (BWAS) between MRI features and behavior are tiny (median |r| β‰ˆ 0.01; top 1% |r| > 0.06) in ~50,000 participants across ABCD, HCP, and UK Biobank, making typical small-sample neuroimaging studies severely underpowered and prone to inflated, irreproducible effects; robust univariate replication requires thousands of participants, with stronger reproducibility for functional (RSFC) versus structural, cognitive versus mental health measures, and multivariate versus univariate approaches.


     Long Explanation



    Core Claim and Evidence

    Marek et al. (Nature, 2022) pooled three large neuroimaging datasetsβ€”ABCD (n = 3,928 with strict motion denoising; 9–10 y, 21 sites), HCP (n β‰ˆ 900; 22–35 y, single scanner, ~60 min RSFC), and UK Biobank (subsampled to n = 900 from 32,572; 40–69 y, 6 min RSFC)β€”totaling ~50,000 individuals to quantify how BWAS effect sizes and reproducibility scale with sample size. Across 11 million univariate associations, the median effect size was |r| β‰ˆ 0.01, and only the top 1% reached |r| > 0.06; the largest association that replicated out-of-sample was |r| = 0.16 .

    At n = 25, the 99% confidence interval for univariate associations was r Β± 0.52β€”effects can be strongly inflated by chance. Even in split halves (n = 1,964), top 1% effects remained inflated by r = 0.07 (78%) on average. At n = 1,000, false-negative rates were 75–100% and half of significant associations were inflated by β‰₯100%; maximum power was 0.68 at n = 3,928. Replication rates were ~5% at typical sample sizes (n < 500) and only 25% at n = 1,964 (P < 0.05). Bonferroni-corrected detection of the top 1% effects required n = 9,500 for 80% power vs n = 2,200 uncorrected .

    Note: replication points are reported values (25%, 5%); intermediate points are illustrative interpolation for visualization onlyβ€”see original Fig. 3e for exact curves.

    Multivariate Findings

    RSFC-based multivariate prediction (SVR) achieved stronger out-of-sample associations (maximum RSFC–crystallized intelligence r_pred = 0.39 vs univariate r = 0.16), yet in-sample estimates remained inflated by Ξ”r = βˆ’0.29 even at n β‰ˆ 2,000. Out-of-sample multivariate associations correlated robustly with univariate effect sizes (r = 0.79) .

    Cross-Dataset Generality

    Size-matched subsamples (n = 900) yielded comparable RSFC–cognitive ability distributions: HCP |r| > 0.12, ABCD |r| > 0.11, UKB |r| > 0.09 (top 1%), suggesting effect-size limitations are not scanner- or site-specific but universal to BWAS with current methods .

    Critical Assessment and Blindspots

    The paper is transparent about its own limits: sociodemographic covariate adjustment decreased the strongest effects (top 1% Ξ”r = βˆ’0.014), RSFC reliability varied widely across datasets (UKB r = 0.39 vs HCP r = 0.79), and theoretical maximum effect sizes are capped by biology and MRI physics. The observational design precludes causal inference, and the authors appropriately note that small-sample, within-person, lesion, and interventional designs remain essential for other neuroimaging questions. A notable limitation is that conclusions rest on RSFC and cortical thicknessβ€”other modalities (e.g., task fMRI activations showed matched distributions here) and demographic confounds may behave differently. Recent work building directly on this paper shows that study-design features (longitudinal designs, increased between-subject age variability) can boost standardized effect sizes substantially (~380% increase for GMV-age longitudinal vs cross-sectional RESI), providing a partial counterpoint to the 'thousands required' framing for specific covariate associations . The authors also have declared financial interests in NOUS Imaging Inc. (FIRMM motion-monitoring software), managed by their institutions.

    Confidence and What Would Change It

    Confidence in the core conclusion is high (strong evidence, ~50,000 participants, rigorous motion denoising, cross-dataset replication, publicly available code at gitlab.com/DosenbachGreene/bwas). The claim would be challenged if a large-scale multi-site BWAS demonstrated robust, stable, non-inflated associations at typical sample sizes, or if improved measurement reliability (e.g., longer RSFC, multi-echo, better phenotyping) substantially increased effect sizes in thousands-of-participant data.

    Author Reviews



    Feedback:    

    Updated: September 25, 2026

     BGPT Paper Review



    Study Novelty

    90%

    First systematic quantification of BWAS effect-size distributions and replication rates as a function of sample size across three of the largest human neuroimaging datasets, directly challenging prevailing practice and setting a new standard for the field.



    Scientific Quality

    90%

    Rigorous bootstrapping, cross-dataset replication (ABCD/HCP/UKB), strict motion denoising sensitivity analyses, matched subsamples, and transparent code/data sharing. Minor concerns: declared financial interests (NOUS/FIRMM) and reliance on observational associations; some conclusions generalized from RSFC and cortical thickness primarily.



    Study Generality

    80%

    Findings apply broadly to brain-behavior association research across ages (9–69 y), modalities, and cohorts; conclusions about power, inflation, and replication generalize to other correlation-based biomarker fields (e.g., genomics, psychology).



    Study Usefulness

    90%

    Directly reshapes study design, funding, and publication standards for neuroimaging; provides concrete sample-size targets (thousands for univariate, ~2,000+ for multivariate) and advocates data sharing, out-of-sample reporting, and within-person/interventional designs.



    Study Reproducibility

    90%

    All three datasets are openly accessible under consortium rules; analysis code (gitlab.com/DosenbachGreene/bwas), pipelines, and ARMS participant lists are publicly released; methods are exhaustively documented.



    Explanatory Depth

    90%

    Provides a mechanistic statistical account (sampling variability β†’ inflation β†’ significance-selection paradox β†’ regression-to-mean replication failure) that explains why small-sample BWAS fail, not merely that they do.


    🎁 Authors: Collect 500 Free Science Tokens (β‰ˆ $50.0 USD)

    Claim My Author Tokens

    Use for 125 days of free BGPT access (4 tokens = 1 day) or trade/sell (β‰ˆ $50.0 USD)

     Top Data Sources ExportMCP



     DataGen



    Simulated illustration based on the paper's reported n=25 CI width (r ± 0.52) scaled as 1/√n and analytic power for |r| = 0.06 at P < 0.05. Assumes bivariate normal correlations, no attrition, and approximates the paper's bootstrap behavior; not observed data.

    Generated scientific data; not direct experimental measurements.

     Hypothesis Graveyard



    Small-sample BWAS fail because of methodological variability or pipeline differences: falsified by Marek et al.'s demonstration that effect-size distributions match across ABCB/HCP/UKB pipelines and that sampling variability alone reproduces the observed inflation and replication failures.


    Task-fMRI activations yield larger brain-behavior effects than RSFC and thus need smaller samples: largely falsified; task-activation and RSFC effect-size distributions were closely matched in HCP, and combined task+rest connectivity produced the same top-1% threshold (|r| > 0.06) as RSFC alone.

     Science Art


    Paper Review: Reproducible brain-wide association studies require thousands of individuals Science Art

     Science Movie



    Make a narrated HD Science movie for this answer ($32 per minute)




     Discussion


    Stay current without chasing every paper.

    Know what changed, what holds up, and what remains uncertain. Every Friday. No ads.


    My BGPT