pairsort

pairsort docs

page what’s inside
Quickstart rank anything in one line; the 60-second tour
Evals all eight evals on one page, each with its ground truth and headline result
Summary Showdown results 100 popular models summarize the PKPD paper: leaderboard, findings, judges, method
Usage install, CLI, input formats, Python API, calibration, the 16-paper worked example
How it works pairwise questions, PKPD Eq. 7, Bradley–Terry, the guards, blending, meta-judge, referee, budgets
Judges & backends Jev via OpenRouter, the LLM fallback, TypeSafe, open Jev models, the /v1/systemone shim
Verifiable eval judges vs exact, recountable counts in freshly generated fictional documents
Market eval judges rank 80 stocks from pre-open SEC filings; truth = realized next-day return
Degradation ladder summaries damaged one logged step at a time: does confidence track the damage?
Code runtime judges pick the faster of two implementations; truth = sandboxed timing
Cross-lingual same truth in 4 languages: does the judge agree with itself?
Weather rank cities by tomorrow’s high and rain before it happens; resolves daily
Evaluation synthetic judge + 16-paper demo: ROC/AUC, calibration, cost vs quality, guard ablations
Humans vs judges the voting site, ballots, and the human-agreement graph

Or just play the Summary Showdown.