Medium Intermediate Budget Probe
The one idea you need
The previous probe asked whether eight times the thinking allowance would move the two stuck benchmark families — and never got an answer, because the benchmark's referee enforces a wall-clock limit per model and refused the very first model at the door. This successor halves the ask: four times the thinking allowance, one fresh seed, same four models. Either the event fits and finally answers the question, or a second refusal proves the thinking-time lever simply does not fit this benchmark's rules and the program accepts a nine-of-ten ceiling for its next training move.
The question
Does four-times thinking room fit the benchmark's wall-clock rules — and if it fits, do the stuck families finally move?
What we found
Same outcome as the first probe, one setting lower: the benchmark's wall-clock referee refused the base model at four-times thinking allowance before any trained model ran. Per the plan written before the event, a second refusal closes the thinking-time lever entirely — no more budget probes at any setting. The complete answer cost two sealed seeds and zero exposed scores. The program's honest ceiling with every currently-believable training path is nine of ten families; the stuck debugging-style family now needs a different class of idea, not another variation.
Why it matters
This is the last cheap test standing between the program and an honest strategic fork: a reopened ten-family goal, or a committed nine-family ceiling with the proven state-tracking dose as the next move.
On this page
Results at a glance 1
How to read
Each bar is a preregistered stage; one means complete.
Takeaway → Design, review, and freeze completed; the event stopped at the first arm and the pre-written consequence closed the lever.
Data table
| preregistered experiment stage | completed checkpoint |
|---|---|
| intake | 1 |
| design receipt frozen | 1 |
| delta review | 1 |
| 4-model tb4096 event | 0 |
Technical framing
Preregistered stage completion — BUDGET_GATE_STOP, second and final: the gateway refused base at medium/tb4096 exactly as at tb8192 (budget_gate_failed, no score emitted, seed 78153 spent). Per the frozen consequence the thinking-budget lever is closed entirely for paired medium events. The program's reachable ceiling with believable training paths is 9/10 families; menders requires a different mechanism class (dose scale or on-policy episode training).
In the author’s words from the Overview · “Results”
The second preregistered stop fired identically to the first: the gateway refused base at medium/tb4096 with budget_gate_failed (exit 2, no score emitted, nothing exposed); zero treated arms ran; the opened ledger record marks seed 78,153 spent; the failure receipt is preserved. Base fits medium at tb1024 (157 s) but not at 4,096 or 8,192 — the gateway's per-arm wall budget binds between 1× and 4× thinking room for the slowest common denominator.
Overview
The thinking-budget lever's last test: the same four published composites, medium tier, think budget 4,096, fresh sealed seed — either the event fits the gateway's wall budget and answers whether serial-compute room moves menders/rites, or a second BUDGET_GATE_STOP closes the lever entirely.
Research Program
- Program:
agentic_breadth_installation. - Program question: can synthetic curricula install general capability that lifts the held-out aggregate without a negative family?
- Prior anchors: the tb8192 probe (BUDGET_GATE_STOP at base, first arm, minimal cost); the medium measurement (hygiene_explore at 8/10, ties menders/rites); the two-tie install (repair kill rule extended — training paths cap at 9/10); C44 serial-compute; the truncation-cascade history.
Question
Does medium/tb4096 fit the gateway's per-arm wall budget for all four arms — and if so, does menders (zero everywhere at tb1024) or rites (zero for three of four arms) move with 4× thinking room?
Setup
- Arms (identical pins to both predecessor events):
base,designed_fresh,replay_repeat,hygiene_explore; base first; the stop applies whichever arm trips. - Event: medium, think budget 4,096, sealed seed 78,153, hardened seed-boundary runner, write-ahead one-seed ledger.
- Readings (identical to the reviewed predecessor): scoped movement booleans; implementation-verified budget contrast vs tb1024/78,150; recorded goal gate; budget integrity.
- Frozen consequence: a second BUDGET_GATE_STOP closes the thinking-budget lever entirely for paired medium events.
Run
PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_medium_intermediate_budget_probe/scripts/run.py --smoke
.venv/bin/python -B experiments/qwen35_4b_medium_intermediate_budget_probe/scripts/run.py --stage benchmarkResults
The second preregistered stop fired identically to the first: the gateway refused base at medium/tb4096 with budget_gate_failed (exit 2, no score emitted, nothing exposed); zero treated arms ran; the opened ledger record marks seed 78,153 spent; the failure receipt is preserved. Base fits medium at tb1024 (157 s) but not at 4,096 or 8,192 — the gateway's per-arm wall budget binds between 1× and 4× thinking room for the slowest common denominator.
Interpretation
- Per the frozen consequence, the thinking-budget lever is closed ENTIRELY for paired medium events: no further budget probes at any setting without a new mechanism argument. Two seeds bought a complete and final answer about the one non-training lever on the binding menders constraint.
- The program's honest position: menders has now defeated three SFT pedagogies at small dose AND the deployment-budget lever. The reachable ceiling for the funded statechain-only successor is 9/10 families. Paths that remain believable for menders are different mechanism CLASSES — dose scale (C43 precedent: partial installs were data-limited), or on-policy episode training — each requiring its own intake and calibration.
Knowledgebase Update
- Program evidence updated: the lever's closure and the two-stop cost recorded.
- Program backlog updated: budget probes retired; statechain-only dose is the funded branch; a dose-scale intake for menders is the queued divergent bet.
- Claim ledger updated: no.
Artifacts
data/design_receipt.json: seed/tier/budget/model/gateway/contrast-source pins.reports/preregistration.md,reports/benchmark_design_review.md: contract and authorization.
Report
Rendered from reports/report.md
Summary
The thinking-budget lever's last test ended in its second and final preregistered stop: the gateway's per-arm wall budget refused base at medium/tb4096 before any treated arm ran, exactly as at tb8192. Per the frozen consequence the lever is closed entirely for paired medium events. Total cost of the complete answer: two sealed seeds, two single-arm gateway refusals, zero exposed scores. The program's reachable ceiling with currently-believable training paths is 9/10 families; menders now requires a different mechanism class (dose scale or on-policy episode training) with its own intake.
Research Program Fit
The tb8192 probe stopped at the gateway's wall-budget gate at minimal cost; this intermediate setting is the only remaining believable configuration for the one non-training lever on the binding menders constraint.
Method
See the preregistration.
Results
runs/benchmark/medium_tb4096_seed78153_measurement/base.failure.json: gateway exit 2, budget_gate_failed, score_emitted: false; single opened ledger record; no treated arm consumed compute; no readout exists by construction.
Controls
Identical model pins to both predecessor events; within-event pairing on one seed; implementation-signature equality verified fail-closed before any cross-event contrast; the stop contract prices failure at one arm's compute.
Oracle Versus Deployable Evidence
Gateway aggregates and public family scores only; benchmarks/ never read.
Next Stage
Closed; the lever is retired. The funded branch is the statechain-only dose (9/10 ceiling accepted); the queued divergent bet is a menders dose-scale intake.
Artifact Manifest
Four composite pins external with committed receipts; everything else in-repo.
Experiment log 2
Show the running log (2 entries, 2026-07-15)
2026-07-15 — Model-free design freeze
- Opened as the tb8192 stop's preregistered successor: think budget 4,096 on fresh sealed seed 78,153, same four published composites, same post-review readings (scoped movement booleans; fail-closed implementation-signature contrast vs the pinned tb1024/78,150 summary), same base-first stop contract — strengthened: a second BUDGET_GATE_STOP closes the thinking-budget lever entirely for paired medium events.
- No model event has run; nothing trains in this cell.
2026-07-15 — The event: second stop; the lever closes
- CI green on the freeze; the event opened the ledger and ran base first; the gateway refused it with
budget_gate_failedat tb4096, exactly as at tb8192. Zero treated arms ran; seed 78,153 spent. - Per the frozen consequence the thinking-budget lever is closed ENTIRELY for paired medium events. The complete answer cost two seeds and two single-arm refusals with nothing exposed.
Reproduce
Smoke test
PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_medium_intermediate_budget_probe/scripts/run.py --smokeFull run
checkpointed scripts/run.py --stage benchmark onlyRun steps are documented inside the experiment folder (README and scripts).