Fluent critiques, weak facts.
Critique skill is 7.04/10, but finding recall is 66% and flaw spotting is 68%.
BGPT
REFUTE v3 · Scientific honesty leaderboard
Can AI read science honestly?
A fluent AI answer can quietly turn weak evidence into a strong claim. REFUTE tests whether models know the difference.
Point estimates, not definitive ranks. The top three are separated by only 1.3 points.
Why this matters
Scientific-sounding is not the same as scientifically honest.
People use AI to understand research, from lab findings to medical news. The dangerous failure is often subtle: turning “may” into “does,” hiding a flaw, or sounding certain when the data are weak.
A careful reader asks: What did the study show? What would prove it wrong? How sure should we be?
The main result
Truth Score, every complete model
Truth Score combines six parts of scientific judgment. Grok-4.2 leads by point estimate, but only 1.3 points separate the top three.
Scores use the v3 formula. Blank or unparsed answers count as wrong. Small differences are not a reliable ordering.
The best model scores 74.5 out of 100. Leading models are close overall, but their weaknesses differ. A rank alone does not tell you which model to trust for a specific scientific task.
The useful tradeoff
The best critic also knows when to be unsure.
Some models write strong critiques but express too much confidence. The frontier shows who best combines critical reasoning with calibrated uncertainty.
Gemma-4-31B, gpt-oss-120b, and Llama-3.3-70B are omitted from this scatter because they stretch the axes and obscure the central comparison. They remain in the full ranking and table.
Why one score is not enough
The same score can hide different risks.
One model may remember findings but overstate confidence. Another may critique elegantly while missing basic facts. That difference matters when choosing an AI reader for research, medicine, journalism, or policy.
Knows the paper, cannot calibrate.
Finding recall is 80%, but critique skill is 4.47/10 and Brier skill is below baseline.
Weak across the board.
Finding recall is 61%, flaw spotting is 54%, and falsifier choice is 60%.
Optional: how model families differ
Family sizes vary, so these means are descriptive. Staying within scope is close to solved, while honest-summary judgment has more headroom.
Benchmark anatomy
What REFUTE tests
REFUTE turns careful scientific reading into four concrete questions. Calibration and open-ended critique are measured separately.
Know the finding
Recall the actual result, not a plausible near-miss from the same study.
Name what overturns it
Choose the concrete observation that would weaken the claim.
Stay within evidence
Choose the conclusion the experiment supports without scope creep.
Spot the honest summary
Distinguish a faithful report from a quietly distorted one.
The Replicate route returns a successful status but an empty answer on REFUTE biology and medical prompts, usually after only 2 to 7 output tokens. Ordinary non-science prompts work. This appears to be route-level safety filtering. We do not convert refusals into a score.
Supplementary detail
Full results and methods
Use Truth Score for the headline. Use the components to understand why a model got there.
Open the complete 15-model table
| # | Model | Truth | Know | Flaws | Falsify | Scope | Skill | Brier↓ |
|---|---|---|---|---|---|---|---|---|
| 1 | Grok-4.2 | 74.5 | 85% | 81% | 98.8% | 98.3% | 7.59 | 0.173 |
| 2 | Claude-Opus-4.7 | 73.4 | 78.8% | 84% | 97.5% | 98.3% | 6.95 | 0.166 |
| 3 | Claude-Opus-4.6 | 73.2 | 78.8% | 83% | 93.8% | 95% | 6.61 | 0.149 |
| 4 | Grok-3-Mini | 72.2 | 80% | 83% | 96.2% | 98.3% | 7.46 | 0.189 |
| 5 | Grok-4.3 | 71.5 | 80% | 81% | 97.5% | 100% | 7.61 | 0.198 |
| 6 | Gemini-3.1-Pro | 71.2 | 88.8% | 85% | 98.8% | 100% | 6.42 | 0.216 |
| 7 | Kimi-K2.6 | 71.0 | 77.5% | 80% | 92.5% | 100% | 6.42 | 0.163 |
| 8 | GPT-5.2 | 69.2 | 82.5% | 79% | 85% | 98.3% | 7.04 | 0.191 |
| 9 | GPT-5.6 | 67.8 | 82.5% | 86% | 95% | 100% | 7.03 | 0.287 |
| 10 | Gemma-4-31B | 67.2 | 82.5% | 77% | 97.5% | 96.7% | 5.62 | 0.205 |
| 11 | GPT-5.4 | 64.3 | 76.2% | 75% | 95% | 98.3% | 6.96 | 0.242 |
| 12 | DeepSeek-V4-Pro | 62.9 | 77.5% | 80% | 83.8% | 100% | 6.32 | 0.246 |
| 13 | Grok-4.1-Fast | 59.4 | 66.2% | 68% | 80% | 95% | 7.04 | 0.228 |
| 14 | gpt-oss-120b | 57.5 | 80% | 74% | 76.2% | 96.7% | 4.47 | 0.494 |
| 15 | Llama-3.3-70B | 46.3 | 61.3% | 54% | 60% | 76.7% | 4.47 | 0.238 |
Five models have partial results but no Truth Score: GLM-5.1, Qwen3-235B, Qwen3.5-397B, GLM-5, and Cogito-v2.1. GPT-5.6 critique skill uses two judges because the third original judge was unavailable for the rerun.
Run REFUTE
Load a public split and grade the final answer letter.
from datasets import load_dataset
items = load_dataset(
"BGPT-OFFICIAL/refute",
"refute_discrimination_hard",
split="train",
)
# Grade the final line: ANSWER=A/B/C/D
What to know before citing
- Truth Scores are point estimates.
- Items were fixed before frontier evaluation.
- Blank and unparsed answers count as wrong.
- Scope is recent English-language empirical research.
- MCQ axes are judge-free; critique skill is 15% of Truth Score.
Open benchmark
Cite REFUTE
@misc{bgpt_refute_2026,
title = {REFUTE: Reasoning Over Evidence Benchmark},
author = {{BGPT Team}},
year = {2026},
url = {https://huggingface.co/datasets/BGPT-OFFICIAL/refute}
}
Built by BGPT from structured evidence extracted from full-text scientific papers.