Research log Small Model Experimentation
GitHub

Universal-Line Medium-Tier Measurement

Two tie-flips from the goal: eight family wins, zero losses, menders and rites at zero

The one idea you need

The receipt forensics showed the program's all-families goal has been judged on the wrong benchmark tier: the quick tier's coarse scoring manufactures unbeatable ties, while the bigger medium tier has seen the goal pass nine times historically. But this program's own clean-trained models — the ones forbidden from ever seeing benchmark-like data — have never been measured on medium at all. This experiment runs that measurement: four already-built models, one fresh sealed seed, no training.

The question

Where do the line's best clean models actually stand on the medium benchmark tier, and which families still block the all-families goal there?

What we found

The clean models' first medium-tier outing put all three at eight-of-ten family wins over the base — matching the best historical arms — and the top two lost NOTHING: they only tied on two families where both they and the base scored zero. The install-carrier model leads the aggregate and the quick-tier ranking inverted (the old quick champion came last of the trained three). The all-families goal now needs exactly two zeros flipped: menders, a genuine capability gap, and rites, which this very lineage has already scored on elsewhere.

Why it matters

The program goal — one trained model beating base on every family at once — needs the right courtroom. This is the first time the clean line gets judged in it, and whatever happens, the next training dose inherits exact per-family targets instead of two fake walls.

Strict family wins vs base8/10all three trained models
Losses (top two arms)0only ties at menders and rites
Best medium aggregate0.338hygiene_explore vs base 0.057
Ties left to flip2menders and rites, both at zero
On this page
  1. Results at a glance
  2. Overview
  3. Report
    1. Summary
    2. Research Program Fit
    3. Method
    4. Results
    5. Controls
    6. Oracle Versus Deployable Evidence
    7. Next Stage
    8. Artifact Manifest
  4. Experiment log
  5. Data files
  6. Reproduce
  7. Related

Results at a glance 1

Each model's score on the ten benchmark families at medium tier

How to read

Grouped bars per family: the untouched base beside the three trained models.

00.20.40.60.8chroniclechroniclelockpicklockpickmendersmendersmiragemirageritesritessiftstacksiftstacksirenssirensstockadestockadetoolsmithtoolsmithwarrenwarren

Takeaway → The trained models beat the base almost everywhere; the two zero-zero ties at menders and rites are all that separates the best model from the all-families goal.

Data table
public benchmark familybasedesigned_freshreplay_repeathygiene_explore
chronicle00.50.40.5
lockpick00.10.40.2
menders0000
mirage00.50.40.6
rites00.100
siftstack00.50.40.5
sirens0.40.60.60.6
stockade00.1350.1590.113
toolsmith0.10.7130.5220.7
warren0.0670.050.10.167

Numbers from experiments/qwen35_4b_universal_medium_tier_measurement/runs/benchmark/medium_tb1024_seed78150_measurement/measurement_readout.json

Technical framing

Per-family scores at medium tier, seed 78150 (tb 1024) — The universal line's first medium-tier paired event: aggregates base 0.0567 / designed_fresh 0.3197 / replay_repeat 0.2981 / hygiene_explore 0.3379 (quick ordering inverted). All three treated arms took 8/10 strict family wins vs base; hygiene_explore and replay_repeat lost nothing and tied only at menders and rites (both 0.0) — the recorded goal gate is two tie-flips wide. Sirens resolved to a strict win at medium granularity exactly as the tier forensics predicted; base sat inside the historical envelope on every family.

In the author’s words from the Overview · “Results”

The single event ran clean (all four arms authenticated, within budget, base inside the historical envelope on every family) (table on the experiment page). Per-family (base → hygiene_explore): chronicle 0→0.5, lockpick 0→0.2, menders 0→0, mirage 0→0.6, rites 0→0, siftstack 0→0.5, sirens 0.4→0.6, stockade 0→0.113, toolsmith 0.1→0.7, warren 0.067→0.167. The quick aggregate ordering INVERTED at medium: replay_repeat (0.5081 quick, best-ever) ranks last of the treated at 0.2981, while the install carrier hygiene_explore leads — the non-convex tier-Pareto frontier reappears inside the universal line. Sirens resolved exactly as the forensics predicted: base 0.4, every treated arm 0.6 — a strict win at medium granularity where quick manufactured ties. … Read the full result →

Overview

The contamination-free universal line's first medium-tier paired benchmark event: four published composites, one fresh sealed seed, the goal gate recorded at the granularity where history says it is winnable.

Research Program

  • Program: agentic_breadth_installation.
  • Program question: can synthetic curricula install general capability that lifts the held-out aggregate without a negative family?
  • Prior anchors: the tier forensics (goal gate 9/94 at medium vs 1/84 at quick; base never at a family ceiling at medium; menders/sirens constants are quick artifacts); the goal-gap pilot (7 families up, 0 down at quick, gate failed on the two artifacts); replay_repeat 0.5081 best-ever quick aggregate.

Question

Where does the line's Pareto set actually stand at medium: does the quick ordering hold, how close is each arm to the recorded all-ten-families goal gate, and which families block the next dose?

Setup

  • Model: only Qwen/Qwen3.5-4B, revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
  • Arms (explicit merges, tree-hash bound at event time): base (tree 26d8ee48…, weights b654e033…), designed_fresh (93433aa2…), replay_repeat (4c4f3561…), hygiene_explore (9eb653d7…).
  • Event: tier medium, think budget 1,024, sealed fresh seed 78,150, trusted gateway only, one-seed ledger, sequential same-seed runs in frozen order.
  • Readings (no promotion bars): medium aggregates + ordering vs quick; recorded goal gate (strict wins vs base per family); base sanity envelope vs the forensics' historical distribution; blocking families per arm.

Run

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_universal_medium_tier_measurement/scripts/run.py --smoke
.venv/bin/python -B experiments/qwen35_4b_universal_medium_tier_measurement/scripts/run.py --stage benchmark

Results

The single event ran clean (all four arms authenticated, within budget, base inside the historical envelope on every family):

armmedium aggregatestrict wins vs basetieslosses
base0.0567
hygiene_explore0.33798/10menders, ritesnone
designed_fresh0.31978/10menderswarren (0.050 vs 0.067)
replay_repeat0.29818/10menders, ritesnone

Per-family (base → hygiene_explore): chronicle 0→0.5, lockpick 0→0.2, menders 0→0, mirage 0→0.6, rites 0→0, siftstack 0→0.5, sirens 0.4→0.6, stockade 0→0.113, toolsmith 0.1→0.7, warren 0.067→0.167.

  • The quick aggregate ordering INVERTED at medium: replay_repeat (0.5081 quick, best-ever) ranks last of the treated at 0.2981, while the install carrier hygiene_explore leads — the non-convex tier-Pareto frontier reappears inside the universal line.
  • Sirens resolved exactly as the forensics predicted: base 0.4, every treated arm 0.6 — a strict win at medium granularity where quick manufactured ties.
  • Menders stayed at 0 for ALL four arms — for the contamination-free line it is a genuine capability gap, not purely an instrument artifact (though replay_refresh once scored 0.125 at quick, so it is a marginal-capability-plus-item-draw gap, not an absolute wall).

Interpretation

The program has never been closer: hygiene_explore beats base strictly on eight families, loses none, and needs exactly two ties flipped — menders > 0 and rites > 0 — to pass the recorded goal gate. rites is demonstrably elicitable in this lineage (designed_fresh scored 0.1 in this same event; replay flipped it at quick). menders is the binding constraint: two designed same-shape attempts already failed (the closed tracefix axis), gym-trained arms historically reached 0.3–0.4 there, and the clean line's only nonzero was one quick item. Successor design must attack menders with a genuinely new mechanism argument, with rites carried alongside, from the hygiene_explore parent, under the pooled_k3 retention protocol.

Knowledgebase Update

  • Program evidence updated: the medium map, the ordering inversion, and the two-tie goal-gate position recorded.
  • Program backlog updated: forensics successor closed; the two-family install (menders + rites) from the hygiene_explore parent is the funded next branch.
  • Claim ledger updated: no new claim; the tier-frontier law gains a universal-line replication.

Artifacts

  • data/design_receipt.json: seed/tier/budget/model/gateway/forensics pins.
  • reports/preregistration.md, reports/benchmark_design_review.md: contract and authorization.

Report

Rendered from reports/report.md

Summary

The universal line's first medium-tier paired event is closed: four published composites on sealed seed 78,150 at think budget 1,024. hygiene_explore leads at 0.3379 (the quick ordering inverted — replay_repeat, best-ever at quick, ranks last of the treated at medium), all three treated arms hit 8/10 strict family wins versus base (the historical mode from the tier forensics), and hygiene_explore and replay_repeat carry zero losses: the recorded goal gate is exactly two tie-flips wide (menders and rites, both at 0). Sirens resolved to a strict win at medium granularity as the forensics predicted; menders stayed at 0 for every arm including base and is the line's binding constraint; base sat inside the historical envelope on all ten families.

Research Program Fit

The tier forensics moved the goal gate's venue to medium; this cell supplies the line's first medium-tier map — either the recorded milestone or the exact blocking families for the next dose.

Method

See the preregistration.

Results

runs/benchmark/medium_tb1024_seed78150_measurement/measurement_readout.json: aggregates 0.0567 / 0.3197 / 0.2981 / 0.3379 (base / designed_fresh / replay_repeat / hygiene_explore); goal gate 8/10-8/10-8/10 with blockers menders+warren (designed_fresh, the only strict loss of the event at warren 0.050 vs 0.067) and menders+rites (the other two, ties at 0); envelope all-inside; every arm within budget (wall 136–230 s).

Controls

All arms published with committed merge receipts; tree hashes recomputed pre-event; one-seed ledger; identical benchmark source inventory across arms; base sanity envelope pinned from the forensics analysis.

Oracle Versus Deployable Evidence

Gateway aggregate and public family scores only; benchmarks/ never read.

Next Stage

Closed. Funded successor: a menders+rites install from the hygiene_explore parent under the pooled_k3 retention protocol — menders requires a genuinely new mechanism argument (same-shape trace-repair doses are closed by kill rule); rites is elicitable in the lineage.

Artifact Manifest

Four composite pins external with committed receipts; everything else in-repo.

Experiment log 5

Show the running log (5 entries, 2026-07-15)

2026-07-15 — Model-free design freeze

  • Opened as the tier forensics' funded successor: the goal gate's venue moves to medium, where the universal line has never been measured.
  • Frozen: four published composites (base / designed_fresh / replay_repeat / hygiene_explore, tree-hash bound), tier medium, think budget 1,024, sealed fresh seed 78,150, one-seed ledger, four preregistered readings (aggregate ordering, recorded goal gate, base sanity envelope, blocking families), no promotion logic anywhere.
  • No model, GPU, or benchmark event has run; nothing trains in this cell.

2026-07-15 — Measurement pipeline built (still model-free)

  • Implemented the four-script pipeline: gen_design_receipt.py (frozen pins + --check byte-identity + refuse-overwrite), run_benchmark.py (event-time tree recompute per arm, one-seed ledger that refuses ANY prior entry, trusted-gateway invocation in the frozen order, safe failure receipts), check_benchmark.py (the four preregistered readings from the gateway receipts only), and the run.py harness (--smoke and the single --stage benchmark behind PASS_BENCHMARK_EVENT).
  • Correction while pinning: the scaffold docs labeled b654e033…16db as base's TREE hash; it is base's reserialized WEIGHTS hash (per the goal-gap pilot's frozen external-weights block). Both identities are now pinned and enforced: base tree 26d8ee48…b677 (recomputed from disk), base weights b654e033…16db.
  • Design receipt generated after full on-disk verification (all four composite trees recomputed and matched) and re-verified byte-identically twice; sha256 e3dfc87434b9eea3173db3d7f5a2b0c2fb501154d94784b5a601adc51de51422.
  • Seed-freshness audit inside the receipt: no seed-context use of 78150 anywhere under experiments/, knowledge/, or research_programs/ outside this experiment's own declarations.
  • Removed the scaffold's src/vllm_runner.py and its test: no engine code outside the trusted gateway exists in this cell (see src/README.md).
  • 37 unit tests green (goal-gate counting, envelope, ordering, ledger refusal, receipt authentication, cross-module frozen constants, the forensics FAMILIES tuple byte-for-byte); smoke green; all three entry points verified to fail closed pre-commit.

2026-07-15 — Adversarial review fixes (two MAJOR, three minor)

  • MAJOR 1 closed: run_benchmark.py itself now enforces the PASS_BENCHMARK_EVENT verdict, adds the review and preregistration to its committed-at-HEAD list, and re-runs gen_design_receipt.py --check (code pins, seed audit, quick pins) at the seed-consuming boundary — a direct invocation can no longer consume seed 78150 with unreviewed or drifted code. Verified live: direct invocation refuses at the review gate; a one-byte drift of check_benchmark.py fails the --check.
  • MAJOR 2 closed: write-ahead one-seed ledger. An opened record is appended before the first gateway call and a closed record after the summary; any closed/legacy record refuses forever, and a crashed event's opened record forces recovery through explicit --resume (matching record required) — deleting the event directory can no longer silently re-consume the seed.
  • Minors: run.py forwards --resume only on explicit operator request (never auto); both receipt loaders reject non-finite or out-of-[0,1] scores (a NaN can no longer silently drop a family from the strict-win partition); smoke now FAILS a published event whose terminal measurement_readout.json is missing.
  • Design receipt regenerated after the code-pin change (full deep verification re-run) and re-verified byte-identically twice; new sha256 26034d8383a146cc3ddb3d8c67e564a53ee2262c6f99ac1269ce1f8482536cad. Seed-freshness audit still clean.
  • Tests now 46 (opened/closed ledger semantics, resume matching, finiteness guards in both loaders, resume-forwarding and seed-boundary-gate contracts); smoke green.

2026-07-15 — Adversarial review: seed-boundary hardened pre-freeze

  • Three-lens review with adversarial verification confirmed two MAJOR findings, both fixed and drill-verified before commit: the seed-consuming runner now enforces the review verdict and the design receipt's code-pin re-check inside its own pre-event block (a one-byte drift of the readings evaluator demonstrably trips it), and the one-seed ledger writes an opened record before the first gateway call so a crashed event can never be silently re-consumed by deleting scratch artifacts.
  • Minors fixed with it: operator-explicit --resume, finite-[0,1] score validation in both loaders, smoke requiring the terminal readout of any published event.
  • Receipt regenerated (code pins changed) with full deep tree verification; 46 tests green; smoke green. Verdict: PASS_BENCHMARK_EVENT.

2026-07-15 — The event (single model event) and closure

  • CI green on the freeze commit; run.py --stage benchmark ran the four gateway events in frozen order on sealed seed 78,150; every arm authenticated (deep tree recompute), within budget, ledger opened before the first run and closed after the summary; the readout re-derived byte-identically.
  • Readings: hygiene_explore 0.3379 > designed_fresh 0.3197 > replay_repeat 0.2981 > base 0.0567 (quick ordering inverted); all three treated arms 8/10 strict wins vs base; hygiene_explore and replay_repeat zero losses with ties only at menders and rites (both 0.0); designed_fresh loses warren 0.050 vs 0.067; base inside the historical envelope on all ten families; sirens resolved to a strict win (0.4 vs 0.6) exactly as the forensics predicted.
  • Closure: the goal gate is two tie-flips wide. rites is elicitable in the lineage (0.1 in this event for designed_fresh); menders is the binding constraint and needs a new mechanism argument per the standing kill rule on same-shape trace-repair doses.

Data files 1

Result tables and metrics copied from the experiment folder — preview inline or open the raw file.

Reproduce

Smoke test

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_universal_medium_tier_measurement/scripts/run.py --smoke

Full run

checkpointed scripts/run.py --stage benchmark only

Run steps are documented inside the experiment folder (README and scripts).

Browse the experiment folder on GitHub ↗