Research log Small Model Experimentation
GitHub

Medium Intermediate Budget Probe

Refused at the door again: the thinking-time lever is closed for good

The one idea you need

The previous probe asked whether eight times the thinking allowance would move the two stuck benchmark families — and never got an answer, because the benchmark's referee enforces a wall-clock limit per model and refused the very first model at the door. This successor halves the ask: four times the thinking allowance, one fresh seed, same four models. Either the event fits and finally answers the question, or a second refusal proves the thinking-time lever simply does not fit this benchmark's rules and the program accepts a nine-of-ten ceiling for its next training move.

The question

Does four-times thinking room fit the benchmark's wall-clock rules — and if it fits, do the stuck families finally move?

What we found

Same outcome as the first probe, one setting lower: the benchmark's wall-clock referee refused the base model at four-times thinking allowance before any trained model ran. Per the plan written before the event, a second refusal closes the thinking-time lever entirely — no more budget probes at any setting. The complete answer cost two sealed seeds and zero exposed scores. The program's honest ceiling with every currently-believable training path is nine of ten families; the stuck debugging-style family now needs a different class of idea, not another variation.

Why it matters

This is the last cheap test standing between the program and an honest strategic fork: a reopened ten-family goal, or a committed nine-family ceiling with the proven state-tracking dose as the next move.

Arms run before the stop1base first, by design — again
Budget settings refused28192 and 4096; 1024 fits
Total cost of the full answer2 seedszero scores exposed
Honest ceiling now9/10families, with believable training paths
On this page
  1. Results at a glance
  2. Overview
  3. Report
    1. Summary
    2. Research Program Fit
    3. Method
    4. Results
    5. Controls
    6. Oracle Versus Deployable Evidence
    7. Next Stage
    8. Artifact Manifest
  4. Experiment log
  5. Reproduce
  6. Related

Results at a glance 1

Where the experiment stopped

How to read

Each bar is a preregistered stage; one means complete.

00.250.50.751intakeintake1design receipt frozendesign receipt frozen1delta reviewdelta review14-model tb4096 event4-model tb4096 event0

Takeaway → Design, review, and freeze completed; the event stopped at the first arm and the pre-written consequence closed the lever.

Data table
preregistered experiment stagecompleted checkpoint
intake1
design receipt frozen1
delta review1
4-model tb4096 event0

Numbers from experiments/qwen35_4b_medium_intermediate_budget_probe/runs/benchmark/medium_tb4096_seed78153_measurement/base.failure.json

Technical framing

Preregistered stage completion — BUDGET_GATE_STOP, second and final: the gateway refused base at medium/tb4096 exactly as at tb8192 (budget_gate_failed, no score emitted, seed 78153 spent). Per the frozen consequence the thinking-budget lever is closed entirely for paired medium events. The program's reachable ceiling with believable training paths is 9/10 families; menders requires a different mechanism class (dose scale or on-policy episode training).

In the author’s words from the Overview · “Results”

The second preregistered stop fired identically to the first: the gateway refused base at medium/tb4096 with budget_gate_failed (exit 2, no score emitted, nothing exposed); zero treated arms ran; the opened ledger record marks seed 78,153 spent; the failure receipt is preserved. Base fits medium at tb1024 (157 s) but not at 4,096 or 8,192 — the gateway's per-arm wall budget binds between 1× and 4× thinking room for the slowest common denominator.

Overview

The thinking-budget lever's last test: the same four published composites, medium tier, think budget 4,096, fresh sealed seed — either the event fits the gateway's wall budget and answers whether serial-compute room moves menders/rites, or a second BUDGET_GATE_STOP closes the lever entirely.

Research Program

  • Program: agentic_breadth_installation.
  • Program question: can synthetic curricula install general capability that lifts the held-out aggregate without a negative family?
  • Prior anchors: the tb8192 probe (BUDGET_GATE_STOP at base, first arm, minimal cost); the medium measurement (hygiene_explore at 8/10, ties menders/rites); the two-tie install (repair kill rule extended — training paths cap at 9/10); C44 serial-compute; the truncation-cascade history.

Question

Does medium/tb4096 fit the gateway's per-arm wall budget for all four arms — and if so, does menders (zero everywhere at tb1024) or rites (zero for three of four arms) move with 4× thinking room?

Setup

  • Arms (identical pins to both predecessor events): base, designed_fresh, replay_repeat, hygiene_explore; base first; the stop applies whichever arm trips.
  • Event: medium, think budget 4,096, sealed seed 78,153, hardened seed-boundary runner, write-ahead one-seed ledger.
  • Readings (identical to the reviewed predecessor): scoped movement booleans; implementation-verified budget contrast vs tb1024/78,150; recorded goal gate; budget integrity.
  • Frozen consequence: a second BUDGET_GATE_STOP closes the thinking-budget lever entirely for paired medium events.

Run

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_medium_intermediate_budget_probe/scripts/run.py --smoke
.venv/bin/python -B experiments/qwen35_4b_medium_intermediate_budget_probe/scripts/run.py --stage benchmark

Results

The second preregistered stop fired identically to the first: the gateway refused base at medium/tb4096 with budget_gate_failed (exit 2, no score emitted, nothing exposed); zero treated arms ran; the opened ledger record marks seed 78,153 spent; the failure receipt is preserved. Base fits medium at tb1024 (157 s) but not at 4,096 or 8,192 — the gateway's per-arm wall budget binds between 1× and 4× thinking room for the slowest common denominator.

Interpretation

  • Per the frozen consequence, the thinking-budget lever is closed ENTIRELY for paired medium events: no further budget probes at any setting without a new mechanism argument. Two seeds bought a complete and final answer about the one non-training lever on the binding menders constraint.
  • The program's honest position: menders has now defeated three SFT pedagogies at small dose AND the deployment-budget lever. The reachable ceiling for the funded statechain-only successor is 9/10 families. Paths that remain believable for menders are different mechanism CLASSES — dose scale (C43 precedent: partial installs were data-limited), or on-policy episode training — each requiring its own intake and calibration.

Knowledgebase Update

  • Program evidence updated: the lever's closure and the two-stop cost recorded.
  • Program backlog updated: budget probes retired; statechain-only dose is the funded branch; a dose-scale intake for menders is the queued divergent bet.
  • Claim ledger updated: no.

Artifacts

  • data/design_receipt.json: seed/tier/budget/model/gateway/contrast-source pins.
  • reports/preregistration.md, reports/benchmark_design_review.md: contract and authorization.

Report

Rendered from reports/report.md

Summary

The thinking-budget lever's last test ended in its second and final preregistered stop: the gateway's per-arm wall budget refused base at medium/tb4096 before any treated arm ran, exactly as at tb8192. Per the frozen consequence the lever is closed entirely for paired medium events. Total cost of the complete answer: two sealed seeds, two single-arm gateway refusals, zero exposed scores. The program's reachable ceiling with currently-believable training paths is 9/10 families; menders now requires a different mechanism class (dose scale or on-policy episode training) with its own intake.

Research Program Fit

The tb8192 probe stopped at the gateway's wall-budget gate at minimal cost; this intermediate setting is the only remaining believable configuration for the one non-training lever on the binding menders constraint.

Method

See the preregistration.

Results

runs/benchmark/medium_tb4096_seed78153_measurement/base.failure.json: gateway exit 2, budget_gate_failed, score_emitted: false; single opened ledger record; no treated arm consumed compute; no readout exists by construction.

Controls

Identical model pins to both predecessor events; within-event pairing on one seed; implementation-signature equality verified fail-closed before any cross-event contrast; the stop contract prices failure at one arm's compute.

Oracle Versus Deployable Evidence

Gateway aggregates and public family scores only; benchmarks/ never read.

Next Stage

Closed; the lever is retired. The funded branch is the statechain-only dose (9/10 ceiling accepted); the queued divergent bet is a menders dose-scale intake.

Artifact Manifest

Four composite pins external with committed receipts; everything else in-repo.

Experiment log 2

Show the running log (2 entries, 2026-07-15)

2026-07-15 — Model-free design freeze

  • Opened as the tb8192 stop's preregistered successor: think budget 4,096 on fresh sealed seed 78,153, same four published composites, same post-review readings (scoped movement booleans; fail-closed implementation-signature contrast vs the pinned tb1024/78,150 summary), same base-first stop contract — strengthened: a second BUDGET_GATE_STOP closes the thinking-budget lever entirely for paired medium events.
  • No model event has run; nothing trains in this cell.

2026-07-15 — The event: second stop; the lever closes

  • CI green on the freeze; the event opened the ledger and ran base first; the gateway refused it with budget_gate_failed at tb4096, exactly as at tb8192. Zero treated arms ran; seed 78,153 spent.
  • Per the frozen consequence the thinking-budget lever is closed ENTIRELY for paired medium events. The complete answer cost two seeds and two single-arm refusals with nothing exposed.

Reproduce

Smoke test

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_medium_intermediate_budget_probe/scripts/run.py --smoke

Full run

checkpointed scripts/run.py --stage benchmark only

Run steps are documented inside the experiment folder (README and scripts).

Browse the experiment folder on GitHub ↗