Ask questions across claims linked to experiments, exact results, limitations, and sources.Know what the science actually supports before you trust the answer.
Press Enter β΅ to search evidence
Explore by Goal
"The most incomprehensible thing about the world is that it is comprehensible."
- Albert Einstein
Quick Analysis Plan
Copied
The claim that brain-wide association study (BWAS) replication requires thousands of participants is strongly supported by resampling analyses across ABCD (nβ3,928 split discovery/replication), HCP (n=900), and UK Biobank (n=900) datasets, where univariate effect sizes were tiny (median |r|β0.01, largest replicated |r|=0.16) and small-sample results were inflated with sign errors and false negatives, though multivariate RSFC approaches showed stronger out-of-sample prediction (SVR r_pred up to 0.39) and remain undercharacterized in their own sample-size dependencies.
Long Analysis Plan
Evidence Battle-Test: Do BWAS Replications Require Thousands of Participants?
Yes β and the support is unusually rigorous. The core evidence comes from a Nature paper (2022) that resampled three large-scale datasets (ABCD, HCP, UK Biobank) to model how sample size affects brain-wide association study (BWAS) effect sizes and replication.
Reported observations across ABCD (3,928 youth split evenly into discovery/replication), HCP (900 adults), and UK Biobank (900 subsamples from 32,572 middle-aged adults):
Univariate BWAS effect sizes were tiny β median |r| β 0.01; top 1% exceeded |r| > 0.06; largest replicated out-of-sample correlation was |r| = 0.16.
Simulation showed substantial inflation, false negatives, and sign errors in smaller samples; replication probability was low even with rigorous denoising.
Multivariate SVR on RSFC predicted cognitive phenotypes with r_pred up to 0.39 out-of-sample β stronger than univariate, but still with in-sample inflation.
Effect size distributions were similar across ABCD, HCP, and UKB when size-matched, supporting the generality of the finding across age ranges (9β69 years).
BGPT inference: The authors' interpretation β that thousands are needed β is well-supported for univariate RSFC/cortical-thickness BWAS targeting cognitive and psychopathology phenotypes. However, the claim's strongest form rests on the assumption that typical BWAS effect sizes reflect biological reality rather than measurement noise. If behavioral measurement reliability improves substantially (e.g., reliable task-based phenotypes), effect sizes might rise, reducing the required n below the 'thousands' threshold β this remains untested in these data.
Blind spots and limitations:
Observational, not causal β the resampling logic assumes that in the very large ABCD sample the observed effect sizes converge on the true population values, which is an assumption, not a proof.
Focus limited to RSFC and cortical thickness; task fMRI, DTI, or multimodal measures are not addressed, and may have different effect size distributions.
Site/scanner harmonization reduces but does not eliminate residual confounds; demographic covariate adjustment choices affect results.
Potential for pipeline selection bias (a specific denoising pipeline was used); alternative pipelines could yield different effect size distributions.
The conclusion is about current BWAS practices; it does not preclude methodological innovations that shift the sample-size requirement downward.
Falsification test: The authors specify what would overturn this finding: a new large-scale BWAS where small-sample results reliably replicate across independent datasets with stable effect sizes matching in-sample estimates, or where reliable replication is achieved with far fewer than thousands of participants.
Confidence level: High for the descriptive claim that current BWAS practices yield tiny effects requiring large samples; moderate for the prescriptive claim that 'thousands' is a fixed universal requirement, since it depends on measurement reliability and pipeline choices that could evolve.
We'll email you the results when your analysis is finished.
Hypothesis Graveyard
'BWAS failure is due to poor preprocessing/denoising pipelines' β rejected: the Nature analysis used a rigorously denoised pipeline yet effect sizes remained tiny, indicating the ceiling is biological/phenotypic, not purely methodological.