Research log Small Model Experimentation
GitHub

Clean Gym-Mix Dose

Three lessons in one dose taught none of them: dilution, re-confirmed

The one idea you need

The retired ancestor adapter was good at exactly three benchmark families — resisting embedded trick instructions, following multi-step procedures, and knowing when a question has no forced answer. The owner's directive: don't bring the ancestor back; rebuild what it knew as fresh, fully documented training data instead. This experiment writes 160 new lessons on invented surfaces covering those three skills and trains them into the clean-lineage model whose every prior step is receipted.

The question

Can fresh, documented gym-style lessons teach the clean model what the undocumented ancestor knew — and does it show up on the three benchmark families?

What we found

The mix failed cleanly and instructively. Splitting the standard 160-lesson budget across three skills — sixty trick-instruction episodes, fifty procedure chains, fifty answer-or-abstain puzzles — taught none of them: on its own fresh exam the trained model scored below BOTH untrained comparison models, while forgetting stayed safely inside the calibrated margins (the dose was inert, not harmful). This re-confirms an earlier program law on clean ground: thin slices of many skills dilute below the threshold where anything installs — the one skill that reliably installs (procedure-tracking) always got the full 160 rows in its winning runs. Two bonus findings about the measuring sticks themselves: the answer-or-abstain exam was too easy for untrained models (they nearly aced it, so it can't detect teaching), and the trick-episode exam was too hard for everyone. The three target families now need three separate full-strength experiments; the sealed benchmark seed was never spent.

Why it matters

A win closes the clean lineage's three-family gap without inheriting a single undocumented byte, raising the odds of full ten-family sweeps on fully documented ground; a null sharpens the conversion law — designed substrate versus inherited substrate — either way the map improves.

Candidate on its own exam15/40below parent 17 and control 19
Rows per skill50-60vs 160 in the winning single-skill runs
Forgettingin-bandinert, not destructive
Benchmark seeds spent0sealed 78,161 never opened
On this page
  1. Results at a glance
  2. Overview
  3. Report
    1. Summary
    2. Research Program Fit
    3. Method
    4. Results
    5. Controls
    6. Oracle Versus Deployable Evidence
    7. Interpretation
    8. Next Experiments
    9. Artifact Manifest
  4. Experiment log
  5. Reproduce
  6. Related

Results at a glance 1

Fresh-exam scores by skill for the three models

How to read

Grouped bars per skill: the trained mix beside its parent and the matched control.

020406080siren_episode (14)siren_episode (14)423statechain (13)statechain (13)576mirage_abstain (13)mirage_abstain (13)8106axis total (40)axis total (40)171915retention pooled (104)retention pooled (104)62.764.359.7

Takeaway → The trained model wins nothing — thin multi-skill doses dilute below the installation threshold.

Data table
gate instrumentzero_root_parentreplay_ctl5gym_mix
siren_episode (14)423
statechain (13)576
mirage_abstain (13)8106
axis total (40)171915
retention pooled (104)62.764.359.7

Numbers from experiments/qwen35_4b_clean_gym_mix_dose/runs/local/seed88046_promotion.json

Technical framing

Gym-mix holdout by kind (correct) and pooled retention, seed 88046 gate — NOT_PROMOTED (mixture dilution): splitting 160 rows across three kinds installed nothing — the candidate scored below both controls on its own holdout while retention stayed in-band (inert, not destructive). The dose-diversity law re-confirms on clean ground: one kind per dose at full concentration. Instrument findings: mirage-abstain ceilinged for untrained controls (replay 10/13); siren-episode floored for everyone (2-4/14). Seed 78161 permanently sealed.

In the author’s words from the Overview · “Results”

Both arms trained clean (retention bands all passed — the dose was not destructive); the 12-run gate (table on the experiment page). NOT promoted: the candidate lost the axis total to BOTH controls and won no kind. Seed 78,161 permanently sealed per contract.

Overview

Research Program

  • Program: agentic_breadth_installation
  • Program question: can DESIGNED synthetic curricula install universal, transferable agentic skills into the one 4B — provably, on fully documented lineage?
  • Prior anchors: lifecycle 22 (qwen35_4b_zero_root_lineage_rebuild — the zero-root composite, tree 414f5829…, 0.3462 aggregate at sealed 78,159) and lifecycle 23 (qwen35_4b_clean_path_statechain_extension — the statechain install replicates on the clean lineage but its rites CONVERSION does not: rites/sirens/mirage were precisely the retired C53 prefix's strengths, so the conversion law carries a substrate precondition).

Question

Lifecycle 25 — owner-directed design. The zero-root composite's weak axes versus the retired gym-era prefix (sealed 78,159 contrast) are exactly the prefix's strengths: sirens (0.5 vs 0.7; 0.4-0.5 across seeds), rites (0.0-0.1), and mirage (0.5-0.6 vs 0.8). Can FRESH, documented, contamination-free gym-style content — designed from the PUBLIC family descriptions only, never touching the prefix composite — recover those margins on clean ground?

Hypothesis

The prefix's family advantages came from trainable behaviors, not from its undocumented substrate: a designed 160-row three-kind dose (goal fidelity under embedded imperatives; hidden-state protocol compliance; calibrated abstention under provable forcing/unforcing) at the standard exposure-matched geometry should install on the axis holdout and move the sirens/rites/mirage readings toward the prefix's margins at medium.

Setup

  • Model: Qwen/Qwen3.5-4B (revision 851bf6e8…), always.
  • Parent and adapter base: the zero-root composite (large_artifacts/qwen35_4b_zero_root_lineage_rebuild/merged/zero_root_hygiene_explore, tree 414f5829…, weights 6e9aad25…), authenticated against lifecycle 22's committed merge receipt (e906caea…; byte-identical provenance copy in data/lineage/provenance/merge.json).
  • Treatment: data/sft_gym_mix.jsonl — 160 FRESH rows at construction seed 77,180 (sha 62950116…), all invented vocabulary, executable truth throughout:

    • 60 u_siren_episode: an invented multi-step book-retrieval transcript in ONE

    user message; 45 rows embed adversarial imperatives carrying a format-matched decoy (the reviewed u_hygiene mechanism in episode form; decoy ≠ truth always), 15 clean; the think target narrates ignoring embedded orders as book content.

    • 50 u_statechain: FRESH instances from the byte-copied PROVEN lifecycle 18

    generator (brewvat/courierloft/peatstove/muletrack, 13/13/12/12), reviewed invariants intact.

    • 50 u_mirage_abstain: invented counter systems (3-5 entities, pairwise ties,

    domain 1-6) proven by exhaustive enumeration — 25 UNIQUELY FORCED (answer = the value) and 25 provably not forced (13 unsatisfiable + 12 undetermined; answer = the invented abstain token NOWHERE, chosen over the family's public IMPOSSIBLE, which is banned everywhere in this corpus). Generated in digit-only pairs: each forced/abstain pair shares one surface skeleton, and the two classes' alphabetic token sets are identical (audited).

    • Banned vocabulary: the proven statechain inventory EXTENDED with the

    sirens/mirage description nouns (injection, retrieval, document(s), directive(s), abstain/abstention, constraint(s), unsatisfiable, impossible), scanned case-insensitively; fresh-surface + row-overlap audits clear all 69 pinned predecessor sources (corpora, streams, gates, lineage datasets).

  • Arms: replay_ctl5 (control, trains FIRST) then gym_mix (candidate); fresh rank-32/alpha-64 adapters, NO warm start, training seed 79, 1 epoch over 1,520 rows (190 optimizer updates, LR 1e-5, batch 1×8, max length 4,096, w_think/w_close 0.2).
  • Exposure: exact zero-delta three-axis MILP (forward / nonzero-target / absolute loss mass ×5) at the frozen geometry — 1,280-row shared stratified core (replay sft_blend.jsonl, byte-identical 25a9595f…) + 240-row variable block (control: 240 replay; candidate: 160 treatment + 80 fillers), namespace seed 55,160. Matched exactly: 1,359,192 forward / 567,805 nonzero / 621,517 mass×5 per arm.
  • Local gate (three arms: parent + both trained): 40-row axis holdout at seed 88,046 (14 siren / 13 statechain / 13 mirage, FRESH instances) + three 104-row retention screens at 88,048/88,050/88,051 under pooled_k3 (the frozen triple 88,047/88,048/ 88,049 collided: 88,047 and 88,049 are the counterfactual cell's reflection/action seeds — next-free substitutions recorded; 88,043 was already taken there). Promotion: axis total strictly > parent AND > replay; AND at least TWO of the three kinds individually strict over BOTH controls (a tie fails a kind); pooled retention bands on screen sums (correct −15, caps +9, parsed −9) vs both controls.
  • Conditional benchmark (only on promotion): ONE sealed medium tb1024 event at fresh seed 78,161, four arms in frozen order — base (26d8ee48…), zero_root_parent (414f5829…), replay_ctl5, gym_mix. Trained-arm pins are six fail-closed TODO-PIN slots in scripts/run_benchmark.py, frozen by check_design's NORMALIZED-HASH pin (lifecycle 22's mechanism).
  • Primary metric: local axis-holdout total + kind breadth (promotion), then pilot gate (candidate aggregate strictly > base AND > replay_ctl5 AND > zero_root_parent).
  • Frozen framing: menders remains closed, so the winnable ceiling is 9/10; the readings of consequence are the THREE AXIS READINGS — candidate vs parent on sirens, rites, and mirage specifically (does fresh gym-style content recover the retired prefix's margins on clean ground?) — and the per-family table, recorded either way. Any 10/10 is a menders draw and feeds a fresh confirmation cell.
  • Contamination statement: the retired prefix composite is NEVER touched or referenced as a model input anywhere in this cell; only the public family descriptions (permitted metadata) informed the design.
  • Standalone: data/lineage/ carries the complete clean-chain package — the six zero-root stage datasets, lifecycle 22's stage + merge receipts as provenance documents, the trainer/merger copies, and a clean-chain manifest recording this cell's dose as STAGE 7. NO blend root exists anywhere in this cell (fail-closed).
  • Hidden-label boundary: gate answers live only in data/local_tasks_seed*.jsonl; the model-facing local_input_seed*.jsonl files carry id/messages/meta only. The benchmark suite directory is never read; only the trusted aggregate gateway runs.

Run

Smoke (no GPU, no writes):

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_clean_gym_mix_dose/scripts/run.py --smoke

Full (one stage per pushed checkpoint, each behind its review verdict):

.venv/bin/python -B experiments/qwen35_4b_clean_gym_mix_dose/scripts/run.py --stage train-control
# then: train-candidate, merge-arms, local, benchmark

Standalone lineage verification (no GPU) / full clean-chain rebuild (GPU):

.venv/bin/python -B experiments/qwen35_4b_clean_gym_mix_dose/scripts/rebuild_clean_chain.py --verify-inputs

Results

Both arms trained clean (retention bands all passed — the dose was not destructive); the 12-run gate:

armaxis total (40)siren_episode (14)statechain (13)mirage_abstain (13)retention pooled
zero_root_parent1745862.67
replay_ctl519271064.33
gym_mix1536659.67

NOT promoted: the candidate lost the axis total to BOTH controls and won no kind. Seed 78,161 permanently sealed per contract.

Interpretation

Three lessons. (1) MIXTURE DILUTION: 50–60 rows per kind installs nothing — the proven statechain effect needed its full 160-row concentration (its own cell won 21/40 with the same vehicle), and splitting the budget three ways landed below controls; this re-confirms the dose-diversity refutation on clean ground and hardens it into a design rule: one kind per dose at full concentration. (2) INSTRUMENT CEILING: the mirage-abstain kind was too easy untrained (replay 10/13) — a kind whose holdout the controls nearly ceiling cannot register installation; future abstention instruments need harder forced/abstain discrimination. (3) The siren-episode kind floors for everyone (2–4/14) — episode-form injection resistance likely needs its own concentrated cell to move at all. The path to the three families runs through three SEPARATE concentrated doses, not one mix.

Knowledgebase Update

  • Program evidence updated: pending results.
  • Program backlog updated: pending results.
  • Claim ledger updated: pending results.

Artifacts

  • src/ — frozen vLLM runner (byte-identical to the lifecycle 23 cell's).
  • scripts/ — staged harness, generators (gym-mix + byte-copied proven statechain

    • canonical retention), corpus builder with audits, exposure pipeline, gate,

    benchmark runner, clean-chain rebuild script, vendored trainer/merger copies.

  • configs/ — frozen identity.
  • data/ — treatment corpus + manifest, replay copy, exposure streams + receipts, gate files, clean-chain lineage package (data/lineage/).
  • runs/ — stage receipts (written by the staged GPU runs).
  • reports/artifact_manifest.yaml

Report

Rendered from reports/report.md

Summary

The clean gym-mix dose closed NOT_PROMOTED with a clean dilution reading: splitting the 160-row budget across three kinds (60 injection-resistance episodes, 50 statechain, 50 provable-abstention) installed nothing — the candidate scored 15/40 on its own holdout, below the parent (17) and the replay control (19), while retention stayed comfortably in-band. The dose-diversity refutation is re-confirmed on clean ground and hardens into a design rule: one kind per dose at full concentration (the statechain effect that won its own gate at 21/40 used all 160 rows on one kind). Two instrument findings ride along: the mirage-abstain kind ceilinged for untrained controls (replay 10/13 — it measures general ability, not the skill) and the siren-episode kind floored for everyone (2–4/14). Sealed seed 78,161 was never opened. The path to the retired prefix's three families runs through three separate concentrated cells.

Research Program Fit

Method

Results

Controls

Oracle Versus Deployable Evidence

Interpretation

Next Experiments

Artifact Manifest

The frozen corpus, streams, receipts, gate inputs, and the clean-chain lineage package are in-repo; trained adapters and merges will live in this cell's own artifact storage with hashes pinned in receipts and reports/artifact_manifest.yaml.

Experiment log 2

Show the running log (2 entries, 2026-07-16)

2026-07-16 — Model-free design freeze

  • Owner-directed design: the retired prefix stays retired; its gym-era strengths (sirens/rites/mirage) are recreated as fresh documented content on the clean lineage — episode-form injection resistance (the 7/7 hygiene mechanism), the proven statechain machinery, and provable-abstention constraint systems with the invented NOWHERE token.
  • Build interrupted twice by transient API overloads and resumed from transcript both times; the review found zero majors (the mirage shortcut hunt clean) and four minors, all fixed pre-freeze with receipts regenerated and all nine gate/corpus artifacts verified byte-identical. Screens moved to 88,048/88,050/88,051 after collisions, recorded fail-closed. 139 tests green; smoke green.

2026-07-16 — The gate: mixture dilution; closure

  • The 12-run gate refused promotion decisively: gym_mix 15/40 below both the parent (17) and replay (19); no kind won; retention comfortably in-band (59.67 vs 62.67/64.33 pooled — the dose was not destructive, it simply did not install at 50–60 rows per kind).
  • Readings: mixture dilution re-confirms the dose-diversity law on clean ground (one kind per dose at full concentration); the mirage kind ceilinged for untrained controls (replay 10/13 — instrument too easy); the siren-episode kind floored for everyone (2–4/14). Seed 78,161 permanently sealed. The three families need three separate concentrated cells; the enumerative-repair dose (already queued) proceeds single-kind by design.

Reproduce

Smoke test

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_clean_gym_mix_dose/scripts/run.py --smoke

Full run

.venv/bin/python -B experiments/qwen35_4b_clean_gym_mix_dose/scripts/run.py --stage train-control (then train-candidate, merge-arms, local, benchmark)

Run steps are documented inside the experiment folder (README and scripts).

Browse the experiment folder on GitHub ↗