Using banded ridge encoding models on NSD fMRI (two subjects, 10,000 images each) with fixed DINOv2 visual features and varying MPNet language predictors, the authors show that the caption-narrative contrast in the additional predictive contribution of vision, ΞV(L) = r(V+L) β r(L), flips sign when only the text predictor's construction changes. In Places, one caption vs full Localized Narrative gave +0.0118 [+0.0066, +0.0166] (subj01) and +0.0147 [+0.0092, +0.0196] (subj02); after matching four concatenated captions (41.7 words) to the narrative (41.3 words), the contrast reversed to β0.0307 [β0.0351, β0.0262] and β0.0232 [β0.0278, β0.0192]. Narrative truncations to caption length yielded the same reversal (β0.121 to β0.141), and even two truncation strategies differing only in text position differed from each other (0.020/0.019) β predictor construction matters at fixed word count.
Length matching shifted every ROI toward the caption predictor by β0.030 to β0.050 with all bootstrap intervals excluding zero, indicating a globally consistent effect rather than a Places-specific anomaly (reported, Figure 1B).
The decomposition (Eq. 2) shows the contrast is driven mostly by language-only prediction: in Places the original joint terms were +0.0008/β0.0003 versus language-only terms β0.0110/β0.0149; after matching, language-only terms (+0.0488/+0.0378) exceeded joint terms (+0.0181/+0.0146), and language-only dominated in every ROI and comparison. Random k-word subsamples showed monotonic dose-response, and averaging eight independent 10-word views recovered ~76-78% of the full-narrative gap β consistent with reduced estimator variance, greater content coverage, or both, which the authors correctly cannot disentangle. Scrambled-text contrasts stayed negative in Faces/Bodies/EBA (β0.038 to β0.084) but reversed sign in Places (+0.017/+0.022), qualifying the lexical-semantic interpretation; scrambling also reduced language-model accuracy, so the authors treat it as an identifiability control. Shared-image analysis (1,000 images) was weaker in Places (subj02 interval included zero) than in other category-selective ROIs β BGPT notes the headline Places reversal is the least robust finding under this check.
Strengths: unusually transparent reporting (Appendix tables, aggregator and threshold sensitivity, text-selection redraws, regularization grid diagnostics), honest scope statements, and practical falsification criteria. Blindspots: only two subjects with no population inference; the matched four-caption condition still confounds annotator count (4 vs 1) and redundancy with length; word count and type-token ratio are collinear, so the within-narrative dose-response does not isolate length; results are specific to DINOv2/MPNet and may not generalize to other feature spaces. The core methodological claim β nested contrasts do not identify represented content by themselves β is well supported within this scope and actionable for the encoding-model community, but small effect sizes (|contrast| β² 0.06 correlation units) and preprint status warrant replication before broad adoption.
Know what changed, what holds up, and what remains uncertain. Every Friday. No ads.