Research log Small Model Experimentation
GitHub

Qwen35 4b Menders Dose Scale

Ten times the lessons, zero learning: the last teaching route closes

The one idea you need

Nine of the ten benchmark families now beat the untouched base model on every sealed test seed; exactly one family — the one shaped like fix-it-and-check-again work — still ties at zero margin. An earlier experiment tried to teach that skill with 80 practice episodes on invented toy machines and installed nothing (0 of 20 fresh problems solved). Every other cheap route is closed by pre-registered kill rules; the one move still allowed is giving the SAME lesson ten times more data. This cell trains on 800 two-round repair episodes across eight invented machines (the four originals with brand-new instances plus four newly designed ones, every knob explicitly bounded in the written rules), against a matched control that sees exactly the same token exposure, from the same parent model.

The question

At ten times the dose that scored zero, does ANY fresh-instance transfer of the repair-loop skill appear — and if the trained model passes the calibrated local gate, does the last blocking benchmark family finally move?

What we found

["The scale bet returned the cleanest possible no. After ten times the training data — eight hundred feedback-repair lessons across eight invented machines — the trained model solved exactly one of forty fresh test episodes, the same single lucky guess as both untrained comparison models. More lessons did not overcome the zero; the skill simply does not install this way. Forgetting also crept past the allowed margin against the parent. With this, every supervised-teaching route to the one family blocking the program's ten-family goal is closed by pre-written rules: three teaching styles, doses from eighty to eight hundred lessons, and the thinking-time levers. What remains: a fundamentally different training class (learning from the model's own attempts with live feedback), or standing on the program's honest position — the full ten-family sweep demonstrated on two of four sealed seeds."]

Why it matters

This is the program's last permitted mechanism at its last blocking family. A nonzero result prices the next dose rung; a zero kills the whole scale class and redirects the budget to genuinely different mechanisms. Either outcome converts an open question into a measured fact.

Fresh-instance transfer1/40= both untrained controls exactly
Dose curve0/20 → 1/4080 rows → 800 rows: flat at floor
Retention vs parent-5.33outside the calibrated ±5 band
SFT routes now closedall3 pedagogies × 80-800 rows + budget levers
On this page
  1. Results at a glance
  2. Overview
  3. Report
    1. Summary
    2. Research Program Fit
    3. Method
    4. Results
    5. Controls
    6. Oracle Versus Deployable Evidence
    7. Interpretation
    8. Next Experiments
    9. Artifact Manifest
  4. Experiment log
  5. Reproduce
  6. Related

Results at a glance 1

Fresh-episode test scores after ten times the training

How to read

Three bars: the model trained on 800 repair lessons beside its untrained parent and a matched control.

020406080holdout correct (40)holdout correct (40)111retention pooled mean (104)retention pooled mean (104)68.765.763.3

Takeaway → One in forty for all three — the training installed nothing the untrained models did not already have.

Data table
gate instrumenthygiene_explore_parentreplay_ctl3feedloop_scale (800-row dose)
holdout correct (40)111
retention pooled mean (104)68.765.763.3

Numbers from experiments/qwen35_4b_menders_dose_scale/runs/local/seed88037_promotion.json

Technical framing

Eight-formalism fresh-instance holdout (of 40) and pooled retention, seed 88037 gate — NOT_PROMOTED + DOSE_SCALE_NULL: at 10x the failed dose the candidate's fresh-instance transfer is exactly the untrained controls' (1/40 each — the guess floor; the mechanical 'nonzero' branch fired but the strict control comparison adjudicates the null). Dose curve flat at floor: 0/20 at 80 rows, 1/40-at-control at 800. Retention failed the parent band (-5.33). The dose-scale mechanism class and the added-diversity variant close together; menders is closed to every tested SFT configuration; seed 78158 permanently sealed.

In the author’s words from the Overview · “Results”

All staged runs completed clean (control loss 0.3899, candidate 0.5298 at 2,280 rows/285 updates each; merges pinned; 12-run authenticated gate). The gate: candidate 1/40 on the eight-formalism fresh-instance holdout — exactly equal to the parent (1/40) and the replay control (1/40); every strict axis bar failed; pooled retention 63.33 vs parent 68.67 (−5.33, outside the ±5 band) and vs replay 65.67 (−2.33, inside); caps and parsed bands failed against both. NOT promoted; sealed seed 78,158 never opened. … Read the full result →

Overview

Lifecycle 20 — the dose-SCALE cell aimed at the last blocking family. Nine benchmark families now hold vs base on every sealed seed; menders alone gates the all-families goal (0-margin ties). Three small-dose pedagogies failed at it; the scale hypothesis (C43: partial installs were data-limited) is the one permitted mechanism class. The dose: 800 episode-feedback rows — 10x the reference cell's failed 80-row dose (u_feedloop scored 0/20 on fresh instances there).

Research Program

  • Program: agentic_breadth_installation
  • Program question: can synthetic curricula install general capability that lifts the held-out aggregate without a negative family?
  • Prior anchors: the reference cell qwen35_4b_feedback_loop_state_chain_install (lifecycle 15: u_feedloop dead at 0/20 inside an 80-row mixed dose); qwen35_4b_statechain_only_dose (lifecycle 18: the proven sibling skill promoted locally and converted rites; its event recorded the parent's first 10/10); qwen35_4b_goal_gate_confirmation (the parent's 10/10 confirmed on three sealed seeds — menders margins single-item thin); menders closed for three small-dose pedagogies + the budget lever.

Question

Does the feedback-loop episode lesson install at 800 rows (10x the failed dose) — and even short of promotion, does ANY fresh-instance u_feedloop transfer appear at 10x? That dose-response reading is preregistered and non-gating, and it carries a stated DOSE x DIVERSITY confound (formalism diversity doubled 4→8 alongside the 10x dose): nonzero at 10x is evidence that scale-plus-diversity reopens the family (the 10x dose is the dominant delta, not a pure dose-response isolate); a 0 at 10x closes the dose-scale mechanism class AND the added-diversity variant together for this skill.

Hypothesis

C43 (meta-induction bankability): partial installs were data-limited — the shift-induction skill moved 0.087→0.40 with dose and plateaued only below the execute ceiling. The feedback-loop lesson at 80 rows inside a mixed dose taught nothing measurable; if the C43 mechanism transfers, 10x the rows on eight formalisms (double the surface diversity) should move fresh-instance transfer off zero, and an installed act→observe→revise loop is the program's best-calibrated bet at the menders family.

Setup

  • Parent and adapter base: the hygiene_explore composite (tree 9eb653d7…), fresh rank-32/alpha-64 adapters, no warm start, training seed 71.
  • Corpus: data/sft_feedloop_scale.jsonl (080c3603…), 800 rows, construction seed 77,150, eight formalisms x 100 — troughline/trinketcord/crankwheel/sigilslate reused from the reference cell as FRESH instances (zero row-overlap receipts) plus four NEW legality-bounded formalisms (barrowyoke, balesled, millround, skeinreel). Every reviewed invariant kept: >=2 legal fix candidates after round-1 evidence (the wrong attempt among them), exactly 1 after rounds 1+2, extended-grammar exclusion audit with a per-formalism probe scope recorded row-by-row (numeric parameters to 12; item parameters over the full pools; the two named-container machines — troughline and barrowyoke — additionally probe every op's CONTAINER dimension over the full pool via a tolerant probe apply where phantom containers start empty; out-of-bound alternatives excluded only by the rendered legality clause), think targets quantifying over legal steps, easy repairs (the lesson is the loop). Banned vocabulary extended with the statechain cells' surface pools; fresh-surface grep audit + zero row-overlap receipts vs 36 pinned predecessor sources.
  • Arms: replay_ctl3 (control, trains FIRST) and feedloop_scale (candidate).
  • Exposure: exact zero-delta 3-axis MILP at namespace seed 55,140 — 1,280 shared replay core + a 1,000-row variable block per arm (control: 1,000 replay slots; candidate: 800 treatment + 200 replay fillers); 2,280 rows/arm, 285 optimizer updates at accum 8 (1,878,709 forward tokens, 771,405 targets, 867,281 mass x5 per arm; zero encoder skips). POOL BIND (documented in the stream manifest): the 2,240-row replay pool cannot fill 1,280 core + 1,000 distinct control rows, and the treatment's long answer spans are unreachable from the 960 non-core rows alone, so the control block draws from the full pool under an ARM-LEVEL multiplicity cap of 2 (575 solver-minimized repeats; no replay row is seen more than twice in the control arm's epoch; the candidate arm is duplicate-free). RESIDUAL BIAS DIRECTION (stated in the manifest and preregistration): repetition plausibly deflates the replay control slightly, making candidate-vs-replay comparisons marginally easier; the parent-anchored bars bind independently and are unaffected, and the retention band vs replay is conservative in the direction that costs the candidate nothing.
  • Local gate: axis holdout 88,037 (40 u_feedloop, 5 per formalism; strict TOTAL over both controls, no per-kind split — single-kind dose) + retention pooled over screens 88,038/88,039/88,040 under pooled_k3 bands on pooled sums (correct >= -15, caps <= +9, parsed >= -9 vs BOTH controls; i.e. +-5/3/3 on means). PLUS the preregistered NON-GATING dose-response reading vs the reference cell's frozen 0/20 baseline, rendered per formalism, both consequence statements recorded either way.
  • Conditional benchmark (only on promotion): medium, tb1024, ONE sealed fresh seed 78,158, four models (base, parent, replay_ctl3, feedloop_scale), hardened seed-boundary runner with the receipt-pinned closed-ledger pattern (the closed record sha-pins the summary AND all four gateway receipts). Pilot gate = candidate aggregate strictly > base AND > replay_ctl3 AND > parent; goal gate recorded either way; frozen power statement: menders > 0 for the candidate on this seed is the reading of consequence; any 10/10 feeds a fresh confirmation cell before any claim.
  • Standalone lineage: the confirmation cell's six-stage hygiene_explore package copied byte-identically (datasets, manifest, three trainers, merger, vendored root adapter ad2ef4fa…/cd764ae8…) and EXTENDED with this cell's stage 7 (the candidate's own training; produced pins are post-training TODO-PINs); scripts/rebuild_lineage.py --verify-inputs is wired into smoke.

Run

Smoke:

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_menders_dose_scale/scripts/run.py --smoke

Full (one stage per pushed checkpoint, each behind its review verdict):

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_menders_dose_scale/scripts/run.py --stage train-control
# then: train-candidate, merge-arms, local, benchmark

Results

All staged runs completed clean (control loss 0.3899, candidate 0.5298 at 2,280 rows/285 updates each; merges pinned; 12-run authenticated gate). The gate: candidate 1/40 on the eight-formalism fresh-instance holdout — exactly equal to the parent (1/40) and the replay control (1/40); every strict axis bar failed; pooled retention 63.33 vs parent 68.67 (−5.33, outside the ±5 band) and vs replay 65.67 (−2.33, inside); caps and parsed bands failed against both. NOT promoted; sealed seed 78,158 never opened.

The dose-response reading, honestly adjudicated: 80 rows → 0/20 (0%); 800 rows → 1/40 (2.5%) with BOTH untrained controls also at 1/40 (2.5%) — the mechanical 'nonzero' trigger fired but the preregistered strict control comparison reads it as the guess floor, not transfer. Per the frozen zero-consequence's substance, the dose-scale mechanism class AND the added-diversity variant close together. Separate deployable evidence from oracle/hidden evaluation. The dose-response reading (candidate u_feedloop axis total vs the frozen 0/20 baseline, per formalism) is recorded either way and never feeds the promotion verdict.

Interpretation

The map hardens decisively. The menders-shaped skill (two-round eliminative repair) is not installable in this model by ANY tested SFT configuration: three pedagogies (asserted, demonstrated-search, episode-feedback), doses from 80 to 800 rows, four to eight surface formalisms, and the deployment budget levers are ALL closed by preregistered kill rules and null results. This extends C38/C48: the eliminative-inference wall is dose-independent in the tested range — more data does not overcome a zero, unlike C43's data-limited partial install (which amplified a 0.087, not a 0). What remains believable for menders: on-policy episode training (a genuinely different mechanism class, its own charter) — or accepting the program's demonstrated-not-confirmed position, which stands at two full 10/10 sweeps across four sealed seeds with nine families holding everywhere. The zero-root lineage rebuild remains queued as the provenance question.

Knowledgebase Update

  • Program evidence updated:
  • Program backlog updated:
  • Claim ledger updated:

Artifacts

  • scripts/ — generators, builders, exact-exposure solver/validator, gate, trainers' wrappers, merger copy, benchmark runner, lineage rebuilder
  • data/ — frozen corpus + manifest, exact-exposure streams + receipts, four gate input files + local design receipt, design receipt, lineage package
  • runs/ — staged GPU outputs (training/merges/local/benchmark; none yet)
  • reports/artifact_manifest.yaml

Report

Rendered from reports/report.md

Summary

The dose-scale bet closed NOT_PROMOTED with the cleanest possible null: at 10× the failed dose (800 episode-feedback rows across eight legality-bounded formalisms, exact zero-delta exposure vs a replay control), the candidate scored 1/40 on fresh instances — exactly equal to both untrained controls — while pooled retention fell outside the parent band (−5.33). The frozen "nonzero" branch fired mechanically (1 > 0) but the strict control comparison adjudicates it as the guess floor. Per the frozen consequences, the dose-scale mechanism class and its added-diversity variant close together: the menders-shaped eliminative-repair skill is not installable in this model by any tested SFT configuration (three pedagogies, 80–800 rows, 4–8 formalisms, plus the closed budget levers). Sealed seed 78,158 was never opened.

Research Program Fit

Method

Results

Controls

Oracle Versus Deployable Evidence

Interpretation

Next Experiments

Artifact Manifest

Four frozen gate inputs, the corpus, streams, and the full lineage package are in-repo; the vendored root and (post-training) adapters and merges live in this cell's own artifact storage with hashes pinned in receipts and reports/artifact_manifest.yaml.

Experiment log 5

Show the running log (5 entries, 2026-07-16)

Scaffold

Created as a new experiment scaffold.

2026-07-16 — Model-free design freeze (lifecycle 20)

  • Opened as the dose-SCALE cell at the last blocking family: nine families hold vs base on every sealed seed and menders alone gates the goal (0-margin ties). Three small-dose pedagogies failed there; dose scale (C43: partial installs were data-limited) is the one permitted mechanism class. Dose: 800 u_feedloop episode rows — 10x the reference cell's failed 80-row dose (0/20 on fresh instances).
  • Frozen: 800-row corpus (080c3603…, construction seed 77,150) on EIGHT legality-bounded formalisms — troughline/trinketcord/crankwheel/ sigilslate reused as fresh instances (zero row-overlap receipts vs 36 pinned predecessor sources) + four new ones (barrowyoke, balesled, millround, skeinreel; fresh-surface grep audit, 47 claimed tokens, zero hits). All reviewed episode invariants kept (>=2 round-1 candidates with the wrong attempt among them, unique after rounds 1+2, extended-grammar exclusion audit; banned vocabulary extended with the statechain cells' surface pools, retaining only the reused feedloop formalisms' nouns).
  • Exact zero-delta 3-axis MILP at namespace seed 55,140: 1,280 shared core

    • 1,000-row variable blocks; 2,280 rows/arm, 285 updates (1,878,709 /

    771,405 / 867,281 per arm; zero skips; 1,280 position-aligned rows). POOL BIND recorded honestly: the predecessor formulation (pairwise disjoint blocks over the 960 non-core rows) is arithmetically infeasible at this dose — the pool is 40 rows short AND the treatment's answer-token mass (mass-minus-nonzero = 4x answer, an encoder identity) is unreachable from the non-core rows alone. The frozen geometry is kept by letting the control block draw from the full pool under an ARM-LEVEL multiplicity cap of 2; the MILP minimizes repeats (575, solver-proven optimal at gap 0) and the candidate arm stays duplicate-free. Documented in the stream manifest, independently re-validated, and unit-tested.

  • Local gate frozen: axis holdout 88,037 (40 u_feedloop, 5 per formalism)

    • three 104-row pooled_k3 retention screens 88,038/88,039/88,040

    (canonical generator 70fc722a…); overlap receipts vs all prior gates 88,013-88,036, all corpora and streams (this cell's and predecessors'), and each other; gen_local_gate --check green twice. Promotion: axis total strictly > parent AND > replay_ctl3, pooled bands (-15/+9/-9 on sums) vs BOTH controls. Preregistered NON-GATING dose-response reading vs the frozen 0/20 baseline (reference promotion receipt sha d232a1be…), rendered per formalism, both consequence statements recorded either way.

  • Conditional benchmark frozen: medium/tb1024, ONE sealed fresh seed 78,158, four arms (base b654e033…, parent 9eb653d7…, replay_ctl3, feedloop_scale), receipt-pinned closed-ledger runner (the closed record sha-pins the summary AND all four gateway receipts; unopened events demand a clean slate; crashed summaries reconcile byte-identically — the confirmation cell's fix class, now standard). Power statement: menders > 0 for the candidate is the reading of consequence; any 10/10 feeds a fresh confirmation cell before any claim.
  • Standalone lineage package: the confirmation cell's six-stage package copied byte-identically (6 datasets + 3 trainers + merger + vendored root ad2ef4fa…/cd764ae8… under this cell's own large_artifacts) and EXTENDED with stage 7 (the candidate's training; dataset = feedloop_scale.jsonl 3aee5f5e…; produced pins are post-training TODO-PINs with explicit pending markers; the GPU rebuild refuses while pending). rebuild_lineage.py --verify-inputs green (7 datasets, 4 trainers, merger, 6 root files) and wired into smoke.
  • Seeds: 77150/55140/88037/88038/88039/88040/78158 verified grep-fresh in seed contexts (zero hits). Training seed 71 verified fresh in the qwen35_4b lineage's training-seed contexts; its only grep hits are run artifacts of the retired sparse_support_memory_executor track and one unrelated scorer test's run_seed — the same artifact class the seed-67 audit accepted — recorded in the design receipt rather than substituted. No substitution required anywhere.
  • 152 tests green; run.py --smoke green; boundary drills refuse without verdicts (train/merge/local/benchmark all fail closed on missing reviews, unpinned TODO-PINs, and dirty git). No GPU stage has run; no model has been loaded.

2026-07-16 — Four-lens review: multiplicity deviation ACCEPTABLE; four minors fixed

  • The review adjudicated the control-arm multiplicity deviation ACCEPTABLE (all four lenses independently; solver-proven minimal; parent-anchored gates prevent false promotion). Zero majors. Four minors fixed:

    1. BIAS DIRECTION stated explicitly where the deviation is documented

    (stream manifest row_duplication.residual_bias_direction, preregistration, README): repetition plausibly DEFLATES the replay control slightly, making candidate-vs-replay comparisons marginally easier; the parent-anchored bars bind independently and are unaffected; the retention band vs replay is conservative in the direction that costs the candidate nothing. 2. DOSE x DIVERSITY CONFOUND written into both frozen consequence statements (check_local.DOSE_RESPONSE_CONSEQUENCES, preregistration, README, site brief): a nonzero reading is evidence that SCALE-PLUS-DIVERSITY reopens the family (10x dose the dominant delta; formalisms doubled 4->8 with it — not a pure dose-response isolate); a zero at 10x closes the scale class AND the added-diversity variant together. Reading stays non-gating. 3. EXTENDED-GRAMMAR AUDIT: chose the preferred option — the audit now PROBES the container dimension for troughline and barrowyoke over the full module pools via a tolerant probe apply (phantom containers start empty; an op touching one can never reproduce the wanted state; probe apply verified equal to the bounded apply on every bounded op, per row). The probe scope is now stated exactly per formalism (EXTENDED_PROBE_SCOPE, recorded row-by-row in the audit and in the corpus manifest; sigilslate's slot indices are declared structural and NOT probed past 4 — the old "items over the full pools" wording is gone). Corpus bytes UNCHANGED (080c3603… — regeneration verified byte-identical; the probe adds enumeration, not rng draws); gate runner inputs and retention sources unchanged; only the axis SOURCE file changed (it embeds the per-row audit). 4. DEAD CODE: run.py local_stage's unreachable check_local --out recovery branch REMOVED; the stage now requires both receipts after the eval and re-adjudicates verify-only; the real post-crash recovery path (manual check_local.py <local_receipt> --out <promotion>) is documented in the stage docstring.

  • Receipt regeneration cascade (--check green twice where applicable): corpus manifest 5617c2c5… (was f354f2f9…; corpus 080c3603… unchanged), stream manifest 5738f0b8… (was 37080454…; both streams byte-identical: 02275b95…/3aee5f5e…), stream token receipt 46ae4cf6… (was 6d1377c3…; exposure numbers unchanged), axis source local_tasks_seed88037 5ec590bf… (was 2547dec1…; all runner inputs + retention sources byte-identical), local design receipt 7ac40653… (was a1a8d969…), design receipt cce84ff4… (was a1f90e93…). train_trial/run.py/check_design pins updated accordingly.
  • 154 tests green (two new container-probe/scope suites); run.py --smoke green; boundary drills re-verified. Still no GPU stage, no commit.

2026-07-16 — Adversarial review: deviation adjudicated, minors fixed, freeze

  • Four lenses, zero blockers/majors; all four independently accepted the pool-bind deviation (control multiplicity 2, solver-proven minimal, parent-anchored gates unaffected). Six minors fixed across two rounds: bias direction stated, dose×diversity confound acknowledged in both frozen consequences, container-dimension probe extended with per-row scope records (corpus bytes unchanged), dead recovery branch removed.
  • 154 tests green; smoke green; receipts regenerated (--check twice); PASS_EXPENSIVE_RUN and PASS_CONTROL_TRAINING granted.

2026-07-16 — Training, the gate, and the null that closes the class

  • Both arms trained clean behind green checkpoints at the enlarged stream; merges published; the 12-run gate executed with full authentication.
  • Verdict NOT_PROMOTED: candidate 1/40 on the fresh-instance holdout, exactly equal to parent (1/40) and replay (1/40) — the guess floor, not transfer; retention 63.33 failed the parent band (−5.33) with caps and parsed bands failing against both controls. Seed 78,158 permanently sealed.
  • The dose-response reading adjudicated honestly: the mechanical 'nonzero' branch fired (1 > 0) but both untrained controls scored the same 1/40, so the substantive reading is NULL AT CONTROL LEVEL — 0/20 at 80 rows, control-floor at 800. The dose-scale mechanism class and the added-diversity variant close together, extending the repair kill rule to every tested SFT dose size and shape.

Reproduce

Smoke test

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_menders_dose_scale/scripts/run.py --smoke

Full run

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_menders_dose_scale/scripts/run.py --stage train-control (then train-candidate, merge-arms, local, benchmark; one stage per pushed checkpoint)

Run steps are documented inside the experiment folder (README and scripts).

Browse the experiment folder on GitHub ↗