pairsort · summary showdown

A benchmark you can vote in

Which AI writes the best summary? You be the judge.

100 of OpenRouter's most-used models each summarized the same paper in one paragraph. AI judges ranked them. Now compare a few pairs yourself and see which judge thinks like you.

100models, one paper
8%of all pairs judged (400 of 4,950)
$0.11for Jev's 4,800 judgments

Ground truth: none. Nobody can say which summary is truly best. The ranking is an AI jury's opinion, and your votes are the human check. How we test the judges against real ground truth →

Your turn

You don't need to know the paper: these summaries were written for exactly you. Pick the better one for the question asked. Keys: A · B · skip. Models stay anonymous until the end.

A

B

the paper (PDF)

Leaderboard

The AI jury's ranking (4 judges pooled). Strength = how often the jury would prefer that summary over a typical one (50% = average). vs next = how often it's preferred over the model ranked just below. Tap a model to read its summary.

Full results & method →

Is the judging any good?

Here there's no right answer, so we test the same judges on eight tasks where there is: exact error counts, next-day stock returns, sandboxed code timings, tomorrow's weather and more.

See the evals → Showdown results & method

Humans vs judges

Ground truth: the picks of site visitors who submitted ballots (not experts).

Loading…

Made with pairsort

This whole leaderboard came from one open-source Python library that ranks anything with AI judges: summaries, ideas, papers, candidates, prompts.

pip install pairsort
pairsort sort ideas.txt "Which idea has more impact?"

About pairsort → Quickstart GitHub