A benchmark you can vote in
Which AI writes the best summary? You be the judge.
100 of OpenRouter's most-used models each summarized the same paper in one paragraph. AI judges ranked them. Now compare a few pairs yourself and see which judge thinks like you.
Ground truth: none. Nobody can say which summary is truly best. The ranking is an AI jury's opinion, and your votes are the human check. How we test the judges against real ground truth →
Your turn
You don't need to know the paper: these summaries were written for exactly you. Pick the better one for the question asked. Keys: ← A · → B · ↓ skip. Models stay anonymous until the end.
Who wrote the summaries you saw?
Add your picks to the study
Opens a prefilled GitHub issue; nothing is sent until you press submit there.
Leaderboard
The AI jury's ranking (4 judges pooled). Strength = how often the jury would prefer that summary over a typical one (50% = average). vs next = how often it's preferred over the model ranked just below. Tap a model to read its summary.
Is the judging any good?
Here there's no right answer, so we test the same judges on eight tasks where there is: exact error counts, next-day stock returns, sandboxed code timings, tomorrow's weather and more.
Humans vs judges
Ground truth: the picks of site visitors who submitted ballots (not experts).
Loading…
Made with pairsort
This whole leaderboard came from one open-source Python library that ranks anything with AI judges: summaries, ideas, papers, candidates, prompts.
pip install pairsort
pairsort sort ideas.txt "Which idea has more impact?"