Hygiene-Explore De-stacked Dose with Medium Pilot
The one idea you need
Strip the practice set down to the only two lessons that have stuck in every measurement — ignoring smuggled instructions, and route-finding under a step budget — and teach them to the cleanest earlier model, before the pile-up of later practice rounds that eventually stopped helping.
The question
Do the two reliable lessons come back cleanly on a fresh start — proving the recent stall was about over-stacking practice, not the lessons — and does the bigger benchmark tier then credit them?
What we found
On the fresh start, both reliable lessons came back decisively — the best skill-test result of the whole session (15 vs 11 and 8 of 20) — proving the earlier stall was about over-stacking practice on one model, not about the lessons. But this direct dose made the model forget ten retained answers, and the screen correctly refused it. Comparing receipts across trials isolates the cause: the one dose that never forgot had a full review round between practice doses; this one skipped it.
Why it matters
Two questions became one law each: stalling = over-stacking (confirmed by recovery), and forgetting = skipping the review round at the dose boundary. The next recipe — review round, then dose — already has its review-round model built and receipted, making it the highest-probability trial yet.
On this page
Results at a glance 1
How to read
Left: unseen skill-test results per model. Right: retained-skill results per model.
Takeaway → The trimmed dose wins the skill test decisively but forgets too much without a review round between doses — the screen refused it, and the fix is already queued.
Data table
| explicit merged composite | axis holdout correct (of 20) | retention correct (of 104) | cap contacts |
|---|---|---|---|
| clean_parent | 8 | 68 | 7 |
| replay_clean | 11 | 66 | 19 |
| hygiene_explore | 15 | 58 | 11 |
Numbers from experiments/qwen35_4b_hygiene_explore_destack_medium/reports/report.md
Technical framing
Seed-88018 gate: axis holdout and retention by arm — Both recovery flags TRUE (interference confirmed; content decay refuted) with the session's best axis result, but the direct dose broke the retention bands (58 vs 68/66) and the gate refused promotion; seed 78,148 sealed. Cross-receipt isolation: replay interleaving at the dose boundary protects retention.
In the author’s words from the Overview · “Results”
Both arms trained cleanly (control 0.3831, candidate 0.4613 train loss; 0 skips) and merged. The frozen 124-task gate at seed 88,018 (normalized grading, both kinds detectable, both required to win): axis holdout of 20 — candidate 15, replay control 11, parent 8; per-kind candidate/parent/replay: explore 7/4/6 (win), hygiene 8/4/5 (win). RECOVERY FLAGS: explore_win: true, hygiene_win: true — the preregistered de-stacking reading is positive. Retention of 104: candidate 58/93/11 versus parent 68/98/7 and replay 66/86/19 — the candidate broke the correct band against both controls (−10/−8 vs the −5 band), the parent cap band (+4 vs +3), and the parent parse band (−5 vs −3). No promotion; seed 78,148 permanently sealed; the medium pilot never ran.
Overview
The minimal proven-install dose: only the two lessons that installed in every prior measurement (injection hygiene, budgeted route search), at their replicated per-kind dose, trained on the CLEAN surface-general parent — testing whether v2's failure was lineage saturation rather than content decay — with a gate that requires both installs to recover and the conditional pilot at the medium tier.
Research Program
- Program:
agentic_breadth_installation. - Program question: can synthetic curricula install general capability that lifts the held-out aggregate without a negative family?
- Prior anchors: the four-event installed-lesson map (hygiene 4/4 kind wins; explore most); the v2 kill-rule closure and third-dose interference law; the dose-two precedent (axis_on_replay installed cleanly with retention byte-equal).
Question
Do the two replicated installs recover cleanly when de-stacked onto the clean lineage at dose two — adjudicating interference versus content decay — and does the medium tier then convert them?
Hypothesis
Interference, not content decay, explains v2: the same lessons at the same per-kind dose on a twice-dosed-maximum lineage recover their installs. Falsified if either kind fails to strictly win its fresh holdout against both the parent and matched replay.
Setup
- Model: only
Qwen/Qwen3.5-4B, revision851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a. - Parent:
designed_freshcomposite (tree93433aa2...255, weights0a3b89cd...979); warm start from its adapter (36f41095...442). - Corpus (construction seed 77,119): 80 rows —
u_hygiene40 (co-location-hardened v2 lesson),u_explore40 (unchanged); executable truth; banned vocabulary audited. - Arms:
replay_clean(control) andhygiene_explore(candidate); 1,280-core + 240-block exact three-axis MILP (candidate block = 80 treatment + 160 fillers; slot seed 55,121); training seed 55; 190 updates; zero skips. - Gate: fresh seed 88,018; 20-task two-kind holdout + 104-task retention; corrected detectability bar (with two kinds, BOTH must strictly win; ties fail; fail-closed if undetectable); retention bands unchanged; documented answer normalization; unconditional recovery flags in the receipt.
- Conditional pilot: sealed seed 78,148, MEDIUM tier, think budget 1,024; candidate aggregate strictly above base, replay control, and parent; every-family-versus-base recorded as the goal gate (8-of-92 historical medium passes).
- Hidden boundary:
benchmarks/unread.
Run
Smoke:
PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --smokeCheckpointed stages:
.venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --stage train-control
.venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --stage train-candidate
.venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --stage merge-arms
.venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --stage local
.venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --stage benchmarkResults
Both arms trained cleanly (control 0.3831, candidate 0.4613 train loss; 0 skips) and merged. The frozen 124-task gate at seed 88,018 (normalized grading, both kinds detectable, both required to win): axis holdout of 20 — candidate 15, replay control 11, parent 8; per-kind candidate/parent/replay: explore 7/4/6 (win), hygiene 8/4/5 (win). RECOVERY FLAGS: explore_win: true, hygiene_win: true — the preregistered de-stacking reading is positive. Retention of 104: candidate 58/93/11 versus parent 68/98/7 and replay 66/86/19 — the candidate broke the correct band against both controls (−10/−8 vs the −5 band), the parent cap band (+4 vs +3), and the parent parse band (−5 vs −3). No promotion; seed 78,148 permanently sealed; the medium pilot never ran.
Interpretation
Three frozen readings. (1) RECOVERY CONFIRMED: both replicated installs returned decisively on the clean lineage at matched exposure — v2's failure was lineage interference, not content decay; the escalation rule does not fire. (2) The retention cost isolates a mechanism: the one prior dose-two event that was retention-byte-equal (axis_on_replay) had a dedicated full replay round between doses; this trial dosed directly and paid ten retention points. Replay interleaving between designed doses protects retention — consistent with every replay-refresh observation in the line, and now localized to the dose boundary. (3) The gate did its job precisely: the axis instrument certified the installs while the retention bands correctly refused a candidate that forgets — this is the strongest axis result of the session (15/20, +4 over the best control) on a model that must not be deployed.
Terminal Disposition
No later event is authorized here. Seed 78,148 is spent-by-sealing. The published composites and receipts are preserved; the replay_clean composite (retention 66/86/19 at the gate) and its adapter provide the interleaved-replay parent that the retention law points to for any successor dose, which requires its own intake and lifecycle.
Knowledgebase Update
- Program evidence updated: recovery confirmation and the replay-interleaving retention law recorded.
- Program backlog updated: the interleaved-replay dose successor queued with calibration notes.
- Claim ledger updated: no.
Artifacts
data/sft_hygiene_explore.jsonl,data/corpus_manifest.json: frozen corpus.data/stream_manifest.json,data/stream_token_receipt.json: exposure receipts.data/local_tasks_seed88018.jsonl,data/local_input_seed88018.jsonl,data/local_design_receipt.json: frozen gate.reports/preregistration.md,reports/design_review.md: contract and authorization.reports/artifact_manifest.yaml: external parent and conditional model-artifact plan.
Report
Rendered from reports/report.md
Summary
Model-free construction under way: the minimal proven-install dose (hygiene + explore, 80 rows) on the clean surface-general parent at dose two, exact-exposure replay control, a two-kind gate where both installs must recover, unconditional recovery flags, and the conditional medium pilot. No model event has run; seed 78,148 sealed.
Research Program Fit
The de-stacking test the third-dose-interference law calls for: adjudicates interference versus content decay with the highest calibrated gate probability of any trial this session, and funds the line's first medium-tier paired event on promotion.
Method
See the preregistration.
Results
- Training (1,520 rows, 0 skips, 190 updates each): control 0.3831, candidate 0.4613 train loss.
- Gate (seed 88,018, normalized grading, both kinds detectable and required): axis holdout of 20 — candidate 15, replay 11, parent 8; explore 7/4/6 (win), hygiene 8/4/5 (win); RECOVERY both true.
- Retention of 104: candidate 58/93/11 vs parent 68/98/7 and replay 66/86/19 — correct band failed vs both, caps and parse vs parent. Not promoted; seed 78,148 permanently sealed.
Controls
clean_parent baseline; replay_clean exact-exposure control; fresh two-kind holdout; retention screen; normalization identical to v2.
Oracle Versus Deployable Evidence
Executable truth grades outputs only; benchmarks/ remains unread.
Next Stage
None. Closed per contract with the recovery confirmation and the replay-interleaving retention law recorded; the interleaved-replay dose successor is queued in the program backlog.
Artifact Manifest
Model-free artifacts in-repo; parent artifacts external with pinned checksums.
Experiment log 6
Show the running log (6 entries, 2026-07-15)
2026-07-15 — Model-free design freeze
- Opened after the v2 kill-rule closure recorded third-dose interference on the axis adapter lineage. This trial de-stacks: only the two lessons with replicated installs (hygiene 4/4 kind wins; explore most), at their proven per-kind dose, on the CLEAN designed_fresh lineage at dose two — the interference law's safe region.
- Corpus: 80 rows from the byte-identical v2 generator at seed 77,119.
- Gate: two-kind holdout where BOTH installs must strictly win, plus the standard retention screen; unconditional recovery flags preregistered; the escalation rule (no further dose permutations if recovery fails) is frozen.
- Conditional pilot at the MEDIUM tier, sealed seed 78,148.
- No model, GPU, training, local, or benchmark event has run.
2026-07-15 — Model-free pipeline run (freeze → measure → materialize → validate → design → gate)
- Adapted the full pipeline from the v2 staged-repair predecessor (build/measure/materialize/validate/train/merge/gate/eval/benchmark/harness), retargeted to the
designed_freshclean parent and the two de-stacked kinds; every fail-closed convention kept (hash pins,--checkbyte-identity, TODO-PIN fail-closed, encoder binding, merge self-pin in the gate receipt, sharedfinalize_promotionwriter, full benchmark CLI, weight recomputation).gen_axis_v2.py,gen_curriculum.py,train_think.py,src/vllm_runner.pycopied byte-identical from the predecessor. - Froze the de-stacked corpus at construction seed 77,119:
data/sft_hygiene_explore.jsonlsha2568b3e97919c62cbb0893add281dc1d3ae881aa0138d0d1721043fec26b0c22cf1(80 rows; hygiene 40 / explore 40; balance 27/40 injected, 17 co-located), manifestcbc9ae6d132d09b9bac2eb43010c2f3eb051993493d6ea659950ac07f0a1e903; replay blend copied byte-identically (25a9595f…abf0c2) from the parent experiment. - Measured exact spans (
source_token_lengths.jsonf67b916688cf8cdcd182963e43883433157445285d3abab7b9c1991b290a50ea; treatment vector forward 19,582 / nonzero 5,793 / mass×5 9,665); the three-axis MILP solved optimally in 0.65 s: both 240-row variable blocks at forward 139,986 / nonzero 58,961 / mass×5 67,773; arm totals 1,367,212 / 574,619 / 629,207; 1,280 position-aligned shared rows; zero skips. Streams:replay_clean.jsonl2189d160…ce0017,hygiene_explore.jsonl82aa1a78…4112f; manifest16b5c1a8…5baac; independent validation receiptstream_token_receipt.jsonf74988c3647f206b1b379ea482bcf9803cb1916d0b45a3a446fa3c80f199da12. - Froze the design receipt (
data/design_receipt.json5319f208cd63c3482d3b81b2e291619418a3bde6dfbc311cdce9c5a084113982) binding the parent identity (tracked merge receiptab3f20cc…6acc2, tree93433aa2…255, weights0a3b89cd…979, adapter36f41095…442/5966461b…055) and the lifecycle substring contracts. - Froze the 124-task gate at seed 88,018:
local_tasks_seed88018.jsonl597f10a44674cc12e5f499be8de6804bb040985019b18aacc5527339a26857eb, runner input2d58da21…12565, design receiptc4952ca9…9b442; zero canonical-message overlap proven against both frozen corpora, both materialized streams, regenerated construction rows, prior local seeds 88,000–88,017, and all five predecessor frozen gates (88,013–88,017). - Filled the stream pins (materialize/validate/train_trial exposure constants); left fail-closed as TODO-PIN:
PUBLISHED_ARM_HASHES(both arms, filled after each training stage publishes) andEXPECTED_TREE_SHA256(both arms, filled after merge-arms publishes); verified each aborts with the pin message. run.py --smokegreen end to end (every--checkbyte-identical, 50 unit tests, py_compile); training remains sealed behind the pushed-checkpoint gates and the PASS_CONTROL_TRAINING / PASS_CONTROL_MERGE verdicts.- No model, GPU, training, local, or benchmark event has run.
2026-07-15 — Authenticated control training
- A first launch attempt failed pre-GPU on a shell working-directory slip (no model event, no artifact); the stage relaunched cleanly from the same green checkpoint.
train-controlran from freeze commite773ed5e(clean synced green main):replay_cleantrained 1,520/1,520 rows with 0 skipped over 190 updates; receipt/log published and pinned fail-closed.
2026-07-15 — Authenticated candidate training
train-candidateran only after control checkpointb063c9c6matchedorigin/mainwith both workflows green and a clean worktree.hygiene_exploretrained 1,520/1,520 rows with 0 skipped over 190 updates; receipt/log published and pinned fail-closed. Merges are next.
2026-07-15 — Authenticated explicit composites
merge-armsran only after candidate checkpointa002d5aematchedorigin/mainwith both workflows green; both arms merged (128/128 modules, fingerprint-verified); tree pins filled fail-closed. The one frozen 124-task gate event at seed 88,018 is the only next stage.
2026-07-15 — Recovery gate: installs recovered; retention bands failed; closed
- The gate ran from merge checkpoint
866f23ce: three authenticated engine runs over the 124-row input at seed 88,018 with normalized grading. - Axis holdout of 20: candidate 15, replay 11, parent 8; explore 7/4/6 (win), hygiene 8/4/5 (win); RECOVERY both true — the de-stacking reading is positive (interference confirmed; content decay refuted; the escalation rule does not fire).
- Retention: 58/93/11 vs parent 68/98/7 and replay 66/86/19 — the correct band failed against both controls, the cap and parse bands against the parent. No promotion; seed 78,148 permanently sealed.
- Cross-receipt isolation: the retention-safe dose-two precedent had a full replay round between doses; this direct dose did not — replay interleaving protects retention.
Reproduce
Smoke test
PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_hygiene_explore_destack_medium/scripts/run.py --smokeFull run
checkpointed scripts/run.py stages onlyRun steps are documented inside the experiment folder (README and scripts).