Gilestro (Imperial College London, preprint 16 Sep 2026) maps multigenerational AI model populations onto population genetics: training on model output is drift (Wright-Fisher), verified real data is immigration, weight averaging is blending inheritance, merging complementary specialists is recombination (Fisher-Muller), and merge failure is reproductive isolation (Bateson-Dobzhansky-Muller incompatibilities). Claims are tested at three tiers: an exact NumPy/SciPy Wright-Fisher simulator, trained networks (RNN, MLP, convolutional VAE on MNIST), and LoRA specialists on Qwen2.5 bases (0.5B/1.5B/7B) plus SmolLM2-1.7B .
Grounding = immigration, a count not a fraction. In the inheritance model (K=1000 Zipf items, n=200 samples/gen, 100 lineages), the closed-form equilibrium H_eq β H*Β·2m/(2m+1) was matched to within 0.5%: one real sample/generation keeps ~67% of source diversity, ten keep ~95% regardless of training-set size. A convolutional VAE self-trained on MNIST collapsed from 30 digit modes to 1 in 15 generations ungrounded, but ~10% real data held all 30 (4 replicates) .
Merging = recombination. Three LoRA specialists merged (weight-averaged or TIES) beat every parent in every seed: 0.65 vs 0.59 (5 seeds, 0.5B) and 0.87 vs 0.81 (3 seeds, 7B) overall. On hard 7B tasks, plain averaging only matched the best specialist (0.41) while routing scored 0.50 in every seed β the gap governed by "headroom," not model size alone .
Six-generation LLM society. Obligate contemporary merging collapsed from 0.65 to 0.27 once partners stopped being complementary; ancestor merging (3 generations back) beat contemporary merging in every seed (0.66 vs 0.27). Notably, the author's own modifier-theory prediction failed: merge refusals tracked generation, not complementarity (partial Spearman Ο = β0.07, CI β0.21 to 0.09), and a fixed stop matched the declinable arm .
Speciation = conflicting conventions. After alignment (Git Re-Basin permutation + per-unit rescaling), same-task MLP pairs merged at parent level (residual β0.001) while conflicting-label pairs kept their barrier (0.502β0.497, 3 replicates). Divergence without conflict (up to 6.4Γ training) produced no isolation β merges improved instead (0.50 specialists β 0.955 merged). Pre-merge functional disagreement predicted merge damage (Ο β +0.45, condition-clustered bootstrap) whereas LoRA-update cosine did not β and the cosine's apparent Ο=+0.60 collapsed to +0.03 when shared-data pairs were added, a confound control the merge-prediction literature reportedly lacked .
Strengths: pre-registered falsifiers with two honest failures reported (confidence-weighting, complementarity-tracking refusals); exact closed forms verified in CI with 151 correctness tests; bitwise-reproducible artifacts and a committed repository; three tiers of evidence rather than simulation alone. Weaknesses and blind spots: single-author work co-produced heavily with AI models (code, analysis, first draft), which raises accountability and verification questions beyond the author's declared responsibility; LLM evidence confined to two model families and LoRA adapters; the six-generation society had no differential reproduction and one base model; epistasis, outbreeding depression, and mate-pool breadth rest on inheritance-model simulation only; several sample sizes are small (3 seeds in key comparisons, where some per-seed wins are not statistically significant); LLM-scale grounding claims are cited from prior work, not re-run here. The analogy's boundaries are stated but the mappings (e.g., weights as alleles, conventions as loci) remain metaphorical in places, and the blending-inheritance proposition for weight averaging applies strictly only by analogy since networks are nonlinear in weights.
What would disprove it: a multigenerational loop where diversity retention tracks the real-data fraction rather than count would falsify the central grounding law; merges failing after alignment without conflicting conventions would falsify the speciation claim; broader model families failing to reproduce the disagreement-predicts-damage result would undercut the diagnostic.
Know what changed, what holds up, and what remains uncertain. Every Friday. No ads.