Evals
Every eval states its ground truth first: something a script can recount, or something that only became known after the judges had answered. The same pairs go to Jev and to three general LLM judges (DeepSeek V4.1 Flash, Gemma 4 31B, Nemotron 3.5 Lightning), so you can compare quality and cost side by side.
Truth: planted false statements and facts in fictional documents, recountable by a script.
Jev 96% of pairs right on accuracy, 94% on completeness, for 2¢.
Results → Market Next-day stock movesTruth: realized returns after judges read pre-open SEC filings. Two sessions so far.
Monday: Jev τ +0.14 (p = 0.035), best judge. Tuesday: nobody beat guessing.
Results → Code Which code runs faster?Truth: sandboxed wall-clock timing of 25 correct implementations.
Jev picks the faster one 97% of the time, for 1.2¢.
Results → Ladder Damage, one step at a timeTruth: the rung number; each rung adds one logged error, deletion or swap.
Jev 94% of pairs right, calibration error 0.033.
Results → 4 languages Same truth, any languageTruth: identical content in English, Spanish, German and Japanese.
Jev gives the same answer in all four 96% of the time, best of the judges.
Results → Weather Tomorrow, todayTruth: airport observations of the next day's high and rain; rankings committed first.
First round resolves 2026-09-25.
Details → Synthetic + papers Known orderingsTruth: a known latent order, and 16 fictional abstracts with levels set by their author.
Papers: Jev AUC 0.994, τ 0.86. Synthetic: calibration error 0.158 → 0.013.
Results → Humans Humans vs judgesTruth: none. Reader ballots from the Summary Showdown, compared with each judge.
Collecting ballots: vote on the showdown to add yours.
Method →The 100-model Summary Showdown has no ground truth at all. It shows what a jury of judges thinks, cheaply (results & method), and readers can vote.
← Home · How it works · All docs