Chen, Thomas, Tenenbaum, and Saxe propose that relationships act like latent features of a sociological environment, analogous to physical obstacles: a dyad's intimacy lowers the discomfort of vulnerable actions (shared utensils, bodily contact, disclosure), which observers invert alongside desire and physical effort to explain behavior. They operationalize this with a hybrid architecture β Llama-3.3-70B generates context-specific alternative actions and scores goal-satisfaction, effort, and risk; a softmax inverse-planning model then performs exact enumeration over these elicited features .
The empirical support is unusually disciplined: six preregistered experiments (N=1,554 retained of 1,564), 16 food and 16 non-food vignettes spanning substance, space, and privacy domains, leave-one-scenario-out cross-validation, and public code/data on GitHub and Zenodo . Key finding: only the full model reproduces relationship-dependent inference β high-risk sharing is stronger evidence of desire in formal relationships, and weaker evidence of intimacy when desire is high or low-risk alternatives are effortful.
Simulated note: the chart reproduces the reported correlations and CIs from the paper's Results text; no novel data are plotted .
Strengths. The novelty is real: prior computational relationship work covered social connectedness and vicarious utility, but the formalityβintimacy axis as an environmental constraint in inverse planning had not been formalized. The cross-domain generalization is compelling β food-fitted utility weights predicted non-food judgments at r=0.977 vs 0.988 for domain-fitted weights, arguing the structure is not a saliva-specific contamination response .
Weaknesses and blind spots. The authors transparently report five preregistration deviations; critically, under the fully preregistered specification the full model's held-out log-likelihood advantage over the vanilla model was not reliable in Studies 1b and 3a (95% CIs crossing zero) β the primary-metric reversal and post-hoc comparison-set reweighting (fitted Ξ· up to 8.85 in Study 2a) are the load-bearing changes . Other limits: US-only participants and contexts (LM-elicited features inherit Western norms); only the intimacy dimension tested; the 'we-agent' abstraction cannot capture asymmetric, reluctant, or nonconsensual action; and the very high correlations near the noise ceiling partly reflect vignette design with strong experimental manipulations rather than free-form everyday observation. BGPT inference (not authors' claim): the observed~90% explainable variance at the cell level (Table S7) is impressive but the paradigm measures third-party reading of explicit relationship labels, not implicit relational knowledge in live interaction.
What would change the conclusion. A replication where the full model's advantage disappears under the preregistered specification, or where intimacy no longer modulates desire inferences after controlling desire and effort, would falsify the framework. Cross-cultural tests (where sharing norms differ) are the sharpest future test the authors themselves propose.
Know what changed, what holds up, and what remains uncertain. Every Friday. No ads.