pairsort docs
| page | what’s inside |
|---|---|
| Quickstart | rank anything in one line; the 60-second tour |
| Evals | all eight evals on one page, each with its ground truth and headline result |
| Summary Showdown results | 100 popular models summarize the PKPD paper: leaderboard, findings, judges, method |
| Usage | install, CLI, input formats, Python API, calibration, the 16-paper worked example |
| How it works | pairwise questions, PKPD Eq. 7, Bradley–Terry, the guards, blending, meta-judge, referee, budgets |
| Judges & backends | Jev via OpenRouter, the LLM fallback, TypeSafe, open Jev models, the /v1/systemone shim |
| Verifiable eval | judges vs exact, recountable counts in freshly generated fictional documents |
| Market eval | judges rank 80 stocks from pre-open SEC filings; truth = realized next-day return |
| Degradation ladder | summaries damaged one logged step at a time: does confidence track the damage? |
| Code runtime | judges pick the faster of two implementations; truth = sandboxed timing |
| Cross-lingual | same truth in 4 languages: does the judge agree with itself? |
| Weather | rank cities by tomorrow’s high and rain before it happens; resolves daily |
| Evaluation | synthetic judge + 16-paper demo: ROC/AUC, calibration, cost vs quality, guard ablations |
| Humans vs judges | the voting site, ballots, and the human-agreement graph |
Or just play the Summary Showdown.