pairsort

Evals

Every eval states its ground truth first: something a script can recount, or something that only became known after the judges had answered. The same pairs go to Jev and to three general LLM judges (DeepSeek V4.1 Flash, Gemma 4 31B, Nemotron 3.5 Lightning), so you can compare quality and cost side by side.

Verifiable Exact error counts

Truth: planted false statements and facts in fictional documents, recountable by a script.

Jev 96% of pairs right on accuracy, 94% on completeness, for 2¢.

Results →
Market Next-day stock moves

Truth: realized returns after judges read pre-open SEC filings. Two sessions so far.

Monday: Jev τ +0.14 (p = 0.035), best judge. Tuesday: nobody beat guessing.

Results →
Code Which code runs faster?

Truth: sandboxed wall-clock timing of 25 correct implementations.

Jev picks the faster one 97% of the time, for 1.2¢.

Results →
Ladder Damage, one step at a time

Truth: the rung number; each rung adds one logged error, deletion or swap.

Jev 94% of pairs right, calibration error 0.033.

Results →
4 languages Same truth, any language

Truth: identical content in English, Spanish, German and Japanese.

Jev gives the same answer in all four 96% of the time, best of the judges.

Results →
Weather Tomorrow, today

Truth: airport observations of the next day's high and rain; rankings committed first.

First round resolves 2026-09-25.

Details →
Synthetic + papers Known orderings

Truth: a known latent order, and 16 fictional abstracts with levels set by their author.

Papers: Jev AUC 0.994, τ 0.86. Synthetic: calibration error 0.158 → 0.013.

Results →
Humans Humans vs judges

Truth: none. Reader ballots from the Summary Showdown, compared with each judge.

Collecting ballots: vote on the showdown to add yours.

Method →

The 100-model Summary Showdown has no ground truth at all. It shows what a jury of judges thinks, cheaply (results & method), and readers can vote.


Home · How it works · All docs