The paper (Machine Learning journal, published 15 October 2025, open access) extends the authors' earlier Discovery Science conference work into a scalable, two-step benchmarking methodology: (1) select a judge LLM via Cohen's ΞΊ agreement with a manually curated 800-pair control set; (2) evaluate base-prompt safety across 10 bias categories, then stress-test only categories scoring β₯ Ο = 0.5 with 7 jailbreak techniques (3 variants each), using misunderstanding filtering with a knee-derived threshold Ο = 0.33 .
DeepSeek V3 671B was the best judge (ΞΊ 0.82, macro-F1 0.861, accuracy 0.869), followed by Gemini 2.0 Flash (ΞΊ 0.74) β notably, commercial giants GPT-4o (ΞΊ 0.66), Claude 3.5 Sonnet (0.65), and Llama 3.1 405B (0.64) trailed substantially .
Initial base-prompt safety was uneven: religion (0.70) and sexual orientation (0.65) safest; age (0.24), disability (0.25), socioeconomic status (0.31) worst; all intersectional categories scored below their isolated counterparts . Phi-4 (0.64) and Gemma2 27B (0.635) were the safest models despite small scale, while DeepSeek V3 (0.405) and GPT-4o mini (0.205) scored lowest β evidence that training/architecture matters more than scale for baseline safety .
Under jailbreaks, machine translation was the most effective attack (mean effectiveness 0.34), then refusal suppression (0.30) and prompt injection (0.29); Scottish Gaelic was the hardest variant; reward incentive (0.05) and role-playing (0.04) least effective. No model remained above the safety threshold under at least one attack, with DeepSeek V3 (expected safety reduction 0.45), Gemma2 27B (0.37), and Gemini 2.0 Flash (0.34) most vulnerable and Llama 3.1 8B most resilient .
Later generations showed higher base safety (GPT-4o 0.455 vs GPT-3.5 Turbo 0.245; Phi-4 0.640 vs Phi-3 0.495; Gemma2 27B 0.635 vs Gemma 7B 0.440), yet newer, more capable models were more susceptible to contextual reframing and obfuscation attacks β a genuine safety-capability trade-off . Four Llama-derived medical LLMs (Bio-Medical, JSL-Med, Med42-v2, UltraMedical) scored lower on safety than general-purpose Llama, with numbers only shown graphically .
What would change the conclusions: replication with an independent judge or human labels on the full response set; jailbreak testing of initially unsafe categories; and multiple-dataset statistical testing (no CIs or significance tests for the safety comparisons are reported). Confidence in the central findings is moderate-to-high for the judge comparison and direction of attack effectiveness, moderate for absolute safety scores.
Author reviews:
Know what changed, what holds up, and what remains uncertain. Every Friday. No ads.