Research log Small Model Experimentation
GitHub

Hygiene-Explore De-stacked Dose with Medium Pilot

Skills came back; forgetting blocked promotion

The one idea you need

Strip the practice set down to the only two lessons that have stuck in every measurement — ignoring smuggled instructions, and route-finding under a step budget — and teach them to the cleanest earlier model, before the pile-up of later practice rounds that eventually stopped helping.

The question

Do the two reliable lessons come back cleanly on a fresh start — proving the recent stall was about over-stacking practice, not the lessons — and does the bigger benchmark tier then credit them?

What we found

On the fresh start, both reliable lessons came back decisively — the best skill-test result of the whole session (15 vs 11 and 8 of 20) — proving the earlier stall was about over-stacking practice on one model, not about the lessons. But this direct dose made the model forget ten retained answers, and the screen correctly refused it. Comparing receipts across trials isolates the cause: the one dose that never forgot had a full review round between practice doses; this one skipped it.

Why it matters

Two questions became one law each: stalling = over-stacking (confirmed by recovery), and forgetting = skipping the review round at the dose boundary. The next recipe — review round, then dose — already has its review-round model built and receipted, making it the highest-probability trial yet.

Recovery flagsboth TRUEexplore 7 vs 4/6; hygiene 8 vs 4/5
Skill test15 / 20session's best — vs 11 (review) and 8 (parent)
Retention cost-1058 vs 68 — band was -5; screen refused
Isolated causeno review roundthe never-forgot dose had one between doses
On this page
  1. Results at a glance
  2. Overview
  3. Report
    1. Summary
    2. Research Program Fit
    3. Method
    4. Results
    5. Controls
    6. Oracle Versus Deployable Evidence
    7. Next Stage
    8. Artifact Manifest
  4. Experiment log
  5. Reproduce
  6. Related

Results at a glance 1

Skills recovered on the fresh start; forgetting cost the promotion

How to read

Left: unseen skill-test results per model. Right: retained-skill results per model.

020406080clean_parentclean_parent8687replay_cleanreplay_clean116619hygiene_explorehygiene_explore155811

Takeaway → The trimmed dose wins the skill test decisively but forgets too much without a review round between doses — the screen refused it, and the fix is already queued.

Data table
explicit merged compositeaxis holdout correct (of 20)retention correct (of 104)cap contacts
clean_parent8687
replay_clean116619
hygiene_explore155811

Numbers from experiments/qwen35_4b_hygiene_explore_destack_medium/reports/report.md

Technical framing

Seed-88018 gate: axis holdout and retention by arm — Both recovery flags TRUE (interference confirmed; content decay refuted) with the session's best axis result, but the direct dose broke the retention bands (58 vs 68/66) and the gate refused promotion; seed 78,148 sealed. Cross-receipt isolation: replay interleaving at the dose boundary protects retention.

In the author’s words from the Overview · “Results”

Both arms trained cleanly (control 0.3831, candidate 0.4613 train loss; 0 skips) and merged. The frozen 124-task gate at seed 88,018 (normalized grading, both kinds detectable, both required to win): axis holdout of 20 — candidate 15, replay control 11, parent 8; per-kind candidate/parent/replay: explore 7/4/6 (win), hygiene 8/4/5 (win). RECOVERY FLAGS: explore_win: true, hygiene_win: true — the preregistered de-stacking reading is positive. Retention of 104: candidate 58/93/11 versus parent 68/98/7 and replay 66/86/19 — the candidate broke the correct band against both controls (−10/−8 vs the −5 band), the parent cap band (+4 vs +3), and the parent parse band (−5 vs −3). No promotion; seed 78,148 permanently sealed; the medium pilot never ran.

Overview

The minimal proven-install dose: only the two lessons that installed in every prior measurement (injection hygiene, budgeted route search), at their replicated per-kind dose, trained on the CLEAN surface-general parent — testing whether v2's failure was lineage saturation rather than content decay — with a gate that requires both installs to recover and the conditional pilot at the medium tier.

Research Program

  • Program: agentic_breadth_installation.
  • Program question: can synthetic curricula install general capability that lifts the held-out aggregate without a negative family?
  • Prior anchors: the four-event installed-lesson map (hygiene 4/4 kind wins; explore most); the v2 kill-rule closure and third-dose interference law; the dose-two precedent (axis_on_replay installed cleanly with retention byte-equal).

Question

Do the two replicated installs recover cleanly when de-stacked onto the clean lineage at dose two — adjudicating interference versus content decay — and does the medium tier then convert them?

Hypothesis

Interference, not content decay, explains v2: the same lessons at the same per-kind dose on a twice-dosed-maximum lineage recover their installs. Falsified if either kind fails to strictly win its fresh holdout against both the parent and matched replay.

Setup

  • Model: only Qwen/Qwen3.5-4B, revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
  • Parent: designed_fresh composite (tree 93433aa2...255, weights 0a3b89cd...979); warm start from its adapter (36f41095...442).
  • Corpus (construction seed 77,119): 80 rows — u_hygiene 40 (co-location-hardened v2 lesson), u_explore 40 (unchanged); executable truth; banned vocabulary audited.
  • Arms: replay_clean (control) and hygiene_explore (candidate); 1,280-core + 240-block exact three-axis MILP (candidate block = 80 treatment + 160 fillers; slot seed 55,121); training seed 55; 190 updates; zero skips.
  • Gate: fresh seed 88,018; 20-task two-kind holdout + 104-task retention; corrected detectability bar (with two kinds, BOTH must strictly win; ties fail; fail-closed if undetectable); retention bands unchanged; documented answer normalization; unconditional recovery flags in the receipt.
  • Conditional pilot: sealed seed 78,148, MEDIUM tier, think budget 1,024; candidate aggregate strictly above base, replay control, and parent; every-family-versus-base recorded as the goal gate (8-of-92 historical medium passes).
  • Hidden boundary: benchmarks/ unread.

Run

Smoke:

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --smoke

Checkpointed stages:

.venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --stage train-control
.venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --stage train-candidate
.venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --stage merge-arms
.venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --stage local
.venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --stage benchmark

Results

Both arms trained cleanly (control 0.3831, candidate 0.4613 train loss; 0 skips) and merged. The frozen 124-task gate at seed 88,018 (normalized grading, both kinds detectable, both required to win): axis holdout of 20 — candidate 15, replay control 11, parent 8; per-kind candidate/parent/replay: explore 7/4/6 (win), hygiene 8/4/5 (win). RECOVERY FLAGS: explore_win: true, hygiene_win: true — the preregistered de-stacking reading is positive. Retention of 104: candidate 58/93/11 versus parent 68/98/7 and replay 66/86/19 — the candidate broke the correct band against both controls (−10/−8 vs the −5 band), the parent cap band (+4 vs +3), and the parent parse band (−5 vs −3). No promotion; seed 78,148 permanently sealed; the medium pilot never ran.

Interpretation

Three frozen readings. (1) RECOVERY CONFIRMED: both replicated installs returned decisively on the clean lineage at matched exposure — v2's failure was lineage interference, not content decay; the escalation rule does not fire. (2) The retention cost isolates a mechanism: the one prior dose-two event that was retention-byte-equal (axis_on_replay) had a dedicated full replay round between doses; this trial dosed directly and paid ten retention points. Replay interleaving between designed doses protects retention — consistent with every replay-refresh observation in the line, and now localized to the dose boundary. (3) The gate did its job precisely: the axis instrument certified the installs while the retention bands correctly refused a candidate that forgets — this is the strongest axis result of the session (15/20, +4 over the best control) on a model that must not be deployed.

Terminal Disposition

No later event is authorized here. Seed 78,148 is spent-by-sealing. The published composites and receipts are preserved; the replay_clean composite (retention 66/86/19 at the gate) and its adapter provide the interleaved-replay parent that the retention law points to for any successor dose, which requires its own intake and lifecycle.

Knowledgebase Update

  • Program evidence updated: recovery confirmation and the replay-interleaving retention law recorded.
  • Program backlog updated: the interleaved-replay dose successor queued with calibration notes.
  • Claim ledger updated: no.

Artifacts

  • data/sft_hygiene_explore.jsonl, data/corpus_manifest.json: frozen corpus.
  • data/stream_manifest.json, data/stream_token_receipt.json: exposure receipts.
  • data/local_tasks_seed88018.jsonl, data/local_input_seed88018.jsonl, data/local_design_receipt.json: frozen gate.
  • reports/preregistration.md, reports/design_review.md: contract and authorization.
  • reports/artifact_manifest.yaml: external parent and conditional model-artifact plan.

Report

Rendered from reports/report.md

Summary

Model-free construction under way: the minimal proven-install dose (hygiene + explore, 80 rows) on the clean surface-general parent at dose two, exact-exposure replay control, a two-kind gate where both installs must recover, unconditional recovery flags, and the conditional medium pilot. No model event has run; seed 78,148 sealed.

Research Program Fit

The de-stacking test the third-dose-interference law calls for: adjudicates interference versus content decay with the highest calibrated gate probability of any trial this session, and funds the line's first medium-tier paired event on promotion.

Method

See the preregistration.

Results

  • Training (1,520 rows, 0 skips, 190 updates each): control 0.3831, candidate 0.4613 train loss.
  • Gate (seed 88,018, normalized grading, both kinds detectable and required): axis holdout of 20 — candidate 15, replay 11, parent 8; explore 7/4/6 (win), hygiene 8/4/5 (win); RECOVERY both true.
  • Retention of 104: candidate 58/93/11 vs parent 68/98/7 and replay 66/86/19 — correct band failed vs both, caps and parse vs parent. Not promoted; seed 78,148 permanently sealed.

Controls

clean_parent baseline; replay_clean exact-exposure control; fresh two-kind holdout; retention screen; normalization identical to v2.

Oracle Versus Deployable Evidence

Executable truth grades outputs only; benchmarks/ remains unread.

Next Stage

None. Closed per contract with the recovery confirmation and the replay-interleaving retention law recorded; the interleaved-replay dose successor is queued in the program backlog.

Artifact Manifest

Model-free artifacts in-repo; parent artifacts external with pinned checksums.

Experiment log 6

Show the running log (6 entries, 2026-07-15)

2026-07-15 — Model-free design freeze

  • Opened after the v2 kill-rule closure recorded third-dose interference on the axis adapter lineage. This trial de-stacks: only the two lessons with replicated installs (hygiene 4/4 kind wins; explore most), at their proven per-kind dose, on the CLEAN designed_fresh lineage at dose two — the interference law's safe region.
  • Corpus: 80 rows from the byte-identical v2 generator at seed 77,119.
  • Gate: two-kind holdout where BOTH installs must strictly win, plus the standard retention screen; unconditional recovery flags preregistered; the escalation rule (no further dose permutations if recovery fails) is frozen.
  • Conditional pilot at the MEDIUM tier, sealed seed 78,148.
  • No model, GPU, training, local, or benchmark event has run.

2026-07-15 — Model-free pipeline run (freeze → measure → materialize → validate → design → gate)

  • Adapted the full pipeline from the v2 staged-repair predecessor (build/measure/materialize/validate/train/merge/gate/eval/benchmark/harness), retargeted to the designed_fresh clean parent and the two de-stacked kinds; every fail-closed convention kept (hash pins, --check byte-identity, TODO-PIN fail-closed, encoder binding, merge self-pin in the gate receipt, shared finalize_promotion writer, full benchmark CLI, weight recomputation). gen_axis_v2.py, gen_curriculum.py, train_think.py, src/vllm_runner.py copied byte-identical from the predecessor.
  • Froze the de-stacked corpus at construction seed 77,119: data/sft_hygiene_explore.jsonl sha256 8b3e97919c62cbb0893add281dc1d3ae881aa0138d0d1721043fec26b0c22cf1 (80 rows; hygiene 40 / explore 40; balance 27/40 injected, 17 co-located), manifest cbc9ae6d132d09b9bac2eb43010c2f3eb051993493d6ea659950ac07f0a1e903; replay blend copied byte-identically (25a9595f…abf0c2) from the parent experiment.
  • Measured exact spans (source_token_lengths.json f67b916688cf8cdcd182963e43883433157445285d3abab7b9c1991b290a50ea; treatment vector forward 19,582 / nonzero 5,793 / mass×5 9,665); the three-axis MILP solved optimally in 0.65 s: both 240-row variable blocks at forward 139,986 / nonzero 58,961 / mass×5 67,773; arm totals 1,367,212 / 574,619 / 629,207; 1,280 position-aligned shared rows; zero skips. Streams: replay_clean.jsonl 2189d160…ce0017, hygiene_explore.jsonl 82aa1a78…4112f; manifest 16b5c1a8…5baac; independent validation receipt stream_token_receipt.json f74988c3647f206b1b379ea482bcf9803cb1916d0b45a3a446fa3c80f199da12.
  • Froze the design receipt (data/design_receipt.json 5319f208cd63c3482d3b81b2e291619418a3bde6dfbc311cdce9c5a084113982) binding the parent identity (tracked merge receipt ab3f20cc…6acc2, tree 93433aa2…255, weights 0a3b89cd…979, adapter 36f41095…442 / 5966461b…055) and the lifecycle substring contracts.
  • Froze the 124-task gate at seed 88,018: local_tasks_seed88018.jsonl 597f10a44674cc12e5f499be8de6804bb040985019b18aacc5527339a26857eb, runner input 2d58da21…12565, design receipt c4952ca9…9b442; zero canonical-message overlap proven against both frozen corpora, both materialized streams, regenerated construction rows, prior local seeds 88,000–88,017, and all five predecessor frozen gates (88,013–88,017).
  • Filled the stream pins (materialize/validate/train_trial exposure constants); left fail-closed as TODO-PIN: PUBLISHED_ARM_HASHES (both arms, filled after each training stage publishes) and EXPECTED_TREE_SHA256 (both arms, filled after merge-arms publishes); verified each aborts with the pin message.
  • run.py --smoke green end to end (every --check byte-identical, 50 unit tests, py_compile); training remains sealed behind the pushed-checkpoint gates and the PASS_CONTROL_TRAINING / PASS_CONTROL_MERGE verdicts.
  • No model, GPU, training, local, or benchmark event has run.

2026-07-15 — Authenticated control training

  • A first launch attempt failed pre-GPU on a shell working-directory slip (no model event, no artifact); the stage relaunched cleanly from the same green checkpoint.
  • train-control ran from freeze commit e773ed5e (clean synced green main): replay_clean trained 1,520/1,520 rows with 0 skipped over 190 updates; receipt/log published and pinned fail-closed.

2026-07-15 — Authenticated candidate training

  • train-candidate ran only after control checkpoint b063c9c6 matched origin/main with both workflows green and a clean worktree.
  • hygiene_explore trained 1,520/1,520 rows with 0 skipped over 190 updates; receipt/log published and pinned fail-closed. Merges are next.

2026-07-15 — Authenticated explicit composites

  • merge-arms ran only after candidate checkpoint a002d5ae matched origin/main with both workflows green; both arms merged (128/128 modules, fingerprint-verified); tree pins filled fail-closed. The one frozen 124-task gate event at seed 88,018 is the only next stage.

2026-07-15 — Recovery gate: installs recovered; retention bands failed; closed

  • The gate ran from merge checkpoint 866f23ce: three authenticated engine runs over the 124-row input at seed 88,018 with normalized grading.
  • Axis holdout of 20: candidate 15, replay 11, parent 8; explore 7/4/6 (win), hygiene 8/4/5 (win); RECOVERY both true — the de-stacking reading is positive (interference confirmed; content decay refuted; the escalation rule does not fire).
  • Retention: 58/93/11 vs parent 68/98/7 and replay 66/86/19 — the correct band failed against both controls, the cap and parse bands against the parent. No promotion; seed 78,148 permanently sealed.
  • Cross-receipt isolation: the retention-safe dose-two precedent had a full replay round between doses; this direct dose did not — replay interleaving protects retention.

Reproduce

Smoke test

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --smoke

Full run

checkpointed scripts/run.py stages only

Run steps are documented inside the experiment folder (README and scripts).

Browse the experiment folder on GitHub ↗