Clean-Path Statechain Extension
The one idea you need
Earlier work proved two things separately: a small 160-example training dose reliably teaches the model to track hidden running state (and that skill carries over to a benchmark family it had always failed), and the program's best model can be rebuilt from scratch using only documented, contamination-free training steps. This experiment combines them: apply the exact proven dose, byte-for-byte, to the clean rebuilt model.
The question
Does the proven state-tracking dose install just as well on the fully documented clean-lineage model, and does it again unlock the benchmark family it unlocked before?
What we found
["Split verdict with a clean lesson. The state-tracking dose installed for the THIRD time on its third different parent — the program's most reliable trained effect — and the clean-lineage model beat the untouched base by 2.7 times while keeping its memory inside the calibrated margins. But the headline hope failed: on the original lineage this dose had TRIPLED the protocol-compliance benchmark family, and on the fully clean lineage that conversion vanished (zero, versus the original's 0.30). The pattern reads clearly: that family was one of exactly three the undocumented ancestor adapter was good at, so the taught skill seems to convert into benchmark scores only where the ancestor's training already tilled the soil. One consolation footnote: the clean model scored a strict win on the otherwise-impossible debugging family through a lucky draw. The fully-documented model — every training step receipted from the official base, zero contamination anywhere — stands as the mission's reference artifact."]
Why it matters
If it works, the program has ONE headline model whose every training step is documented, receipted, and reproducible from a blank start — no mystery ingredients anywhere — while carrying the strongest install the program has proven.
On this page
Results at a glance 1
How to read
Grouped bars per family; note rites, where the clean candidate scored zero.
Takeaway → The clean lineage wins broadly over base but the taught skill's family conversion did not survive the ancestor's removal.
Data table
| public benchmark family | base | zero_root_parent | replay_ctl4 | statechain_clean |
|---|---|---|---|---|
| chronicle | 0.1 | 0.7 | 0.6 | 0.6 |
| lockpick | 0.1 | 0 | 0.1 | 0 |
| menders | 0 | 0 | 0 | 0.017 |
| mirage | 0 | 0.6 | 0.3 | 0.7 |
| rites | 0.1 | 0.1 | 0 | 0 |
| siftstack | 0 | 0.5 | 0.6 | 0.5 |
| sirens | 0.4 | 0.5 | 0.4 | 0.4 |
| stockade | 0.1 | 0.1 | 0.132 | 0.273 |
| toolsmith | 0.3 | 0.8 | 0.72 | 0.71 |
| warren | 0.133 | 0.217 | 0.267 | 0.133 |
Numbers from experiments/qwen35_4b_clean_path_statechain_extension/runs/benchmark/medium_tb1024_seed78160_pilot/summary.json
Technical framing
Per-family scores at medium tier, sealed seed 78160 (tb 1024) — PILOT_NOT_PROMOTED + CONVERSION_NOT_REPLICATED: the statechain install held on its third parent (local 21/40 strict; retention in-band — 3-for-3) and the clean candidate beat base 2.7x and its replay control, but lost to its parent by 0.018 and the rites conversion came back FALSE (candidate 0.0 vs the original lineage's 0.300) — the data-to-family converter is lineage-dependent: rites/sirens/mirage were exactly the C53 prefix's strengths. Footnote: a strict menders draw-win (0.017) landed while rites collapsed; this seed drew the strongest base ever (0.1234), squeezing all treated arms to 6/10.
In the author’s words from the Overview · “Results”
Local gate: PROMOTED on all eight checks — the statechain install's third replication, on its third distinct parent (axis 21/40 strictly over parent 19 and replay 16; pooled retention 61.33 vs 62.33/63.0, deep inside the calibrated bands). Training-loss property recorded: the clean chain fits the replay surface at ~1.3 versus the original lineage's ~0.43 while performing within ~10% at the benchmark (loss-level ≠ capability, dramatically). Sealed event at 78,160 (all arms authenticated; the six-slot normalized pin held through the fill) (table on the experiment page). Pilot: candidate > base ✓, > replay ✓, > parent ✗ (−0.018) — NOT promoted, the same shape as the original statechain cell. … Read the full result →
Overview
Research Program
- Program:
agentic_breadth_installation - Program question: can DESIGNED synthetic curricula install universal, transferable agentic skills into the one 4B — provably, on fully documented lineage?
- Prior anchors: lifecycle 18 (
qwen35_4b_statechain_only_dose— the statechain dose installs, 21/40 axis strict over both controls, and CONVERTS to the rites family: 0.300 vs 0.100/0.100 paired at sealed 78,154) and lifecycle 22 (qwen35_4b_zero_root_lineage_rebuild— the six documented stages replayed from a fresh zero-initialized adapter produce the zero-root composite, tree414f5829…, weights6e9aad25…, 0.3462 aggregate with 7/10 strict wins and zero losses at sealed 78,159).
Question
Lifecycle 23 — the mission's cleanest artifact. Does the PROVEN statechain converter dose, applied byte-identically to the ZERO-ROOT composite, produce a single installed model whose ENTIRE lineage is documented and contamination-free end-to-end — and does the rites conversion replicate ON THE CLEAN LINEAGE?
Hypothesis
The statechain install is a property of the dose, not of the blend-rooted parent it was first proven on: the same 160 frozen rows at the same exposure-matched geometry should clear the same calibrated gate from the zero-root composite, and the local install should again convert to the rites family at medium.
Setup
- Model: Qwen/Qwen3.5-4B (revision
851bf6e8…), always. - Parent and adapter base: the zero-root composite (
large_artifacts/qwen35_4b_zero_root_lineage_rebuild/merged/zero_root_hygiene_explore, tree414f5829…, weights6e9aad25…), authenticated against lifecycle 22's committed merge receipt (e906caea…; byte-identical provenance copy indata/lineage/provenance/merge.json). - Treatment:
data/sft_statechain_only.jsonl— the source cell's frozen 160-row corpus copied BYTE-IDENTICALLY (ab6c7845…); fresh instances would change the treatment, so the byte-copy is the controlled choice. Replay poolsft_blend.jsonl(25a9595f…) byte-identical to every predecessor copy. - Arms:
replay_ctl4(control, trains FIRST) thenstatechain_clean(candidate); fresh rank-32/alpha-64 adapters, NO warm start, training seed 73, 1 epoch over 1,520 rows (190 optimizer updates, LR 1e-5, batch 1×8, max length 4,096, w_think/w_close 0.2). - Exposure: exact zero-delta three-axis MILP (forward / nonzero-target / absolute loss mass ×5) at the frozen geometry — 1,280-row shared stratified core + 240-row variable block (control: 240 replay; candidate: 160 treatment + 80 fillers), namespace seed 55,150.
- Local gate (three arms: parent + both trained): 40-row statechain axis holdout at seed 88,041 (10 per formalism, FRESH instances from the copied generator) + three 104-row retention screens at 88,042/88,044/88,045 under pooled_k3 (88,043 is taken by
qwen35_4b_counterfactual_plan_reflection_transfer— documented skip). Promotion: axis total strictly > parent AND > replay_ctl4; pooled retention bands on screen sums (correct −15, caps +9, parsed −9) vs BOTH controls. - Conditional benchmark (only on promotion): ONE sealed medium tb1024 event at fresh seed 78,160, four arms in frozen order — base (
26d8ee48…), zero_root_parent (414f5829…), replay_ctl4, statechain_clean. Trained-arm pins are six fail-closed TODO-PIN slots inscripts/run_benchmark.py, frozen by check_design's NORMALIZED-HASH pin (lifecycle 22's mechanism). - Primary metric: local axis-holdout total (promotion), then pilot gate (candidate aggregate strictly > base AND > replay_ctl4 AND > zero_root_parent).
- Frozen framing: menders is closed, so the winnable ceiling is 9/10; the readings of consequence are (a) the rites conversion ON THE CLEAN LINEAGE (candidate rites vs parent/replay rites, paired) and (b) the fully documented model's per-family profile. Any 10/10 is a menders draw and feeds a fresh confirmation cell before any claim.
- Standalone:
data/lineage/carries the complete clean-chain package — the six zero-root stage datasets, lifecycle 22's stage + merge receipts as provenance documents, the trainer/merger copies, and a clean-chain manifest recording this cell's dose as STAGE 7. NO blend root exists anywhere in this cell (fail-closed). - Hidden-label boundary: gate answers live only in
data/local_tasks_seed*.jsonl; the model-facinglocal_input_seed*.jsonlfiles carry id/messages/meta only. The benchmark suite directory is never read; only the trusted aggregate gateway runs.
Run
Smoke (no GPU, no writes):
PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_clean_path_statechain_extension/scripts/run.py --smokeFull (one stage per pushed checkpoint, each behind its review verdict):
.venv/bin/python -B experiments/qwen35_4b_clean_path_statechain_extension/scripts/run.py --stage train-control
# then: train-candidate, merge-arms, local, benchmarkStandalone lineage verification (no GPU) / full clean-chain rebuild (GPU):
.venv/bin/python -B experiments/qwen35_4b_clean_path_statechain_extension/scripts/rebuild_clean_chain.py --verify-inputsResults
Local gate: PROMOTED on all eight checks — the statechain install's third replication, on its third distinct parent (axis 21/40 strictly over parent 19 and replay 16; pooled retention 61.33 vs 62.33/63.0, deep inside the calibrated bands). Training-loss property recorded: the clean chain fits the replay surface at ~1.3 versus the original lineage's ~0.43 while performing within ~10% at the benchmark (loss-level ≠ capability, dramatically).
Sealed event at 78,160 (all arms authenticated; the six-slot normalized pin held through the fill):
| arm | aggregate | goal gate vs base | rites |
|---|---|---|---|
| base | 0.1234 (strongest base draw yet) | — | 0.100 |
| zero_root_parent | 0.3517 | 6/10 | 0.100 |
| statechain_clean | 0.3333 | 6/10 (incl. a strict MENDERS win, 0.017) | 0.000 |
| replay_ctl4 | 0.3119 | 6/10 | 0.000 |
Pilot: candidate > base ✓, > replay ✓, > parent ✗ (−0.018) — NOT promoted, the same shape as the original statechain cell. The frozen conversion reading: converts_on_clean_lineage: false — candidate rites 0.0 against the original-lineage conversion's 0.300.
Interpretation
Three durable readings. (1) The statechain INSTALL is robust — three parents, three promotions, retention held each time under calibrated bands. (2) The CONVERSION is lineage-dependent: 1-for-2, expressed on the prefix lineage and absent on the clean one, and the pattern is legible — rites/sirens/mirage were precisely the C53 prefix's strengths, so the designed dose appears to convert only where the prefix's substrate already leans toward the family. The program's one proven data→family mechanism thus carries a substrate precondition, which scopes the conversion law honestly. (3) Per-seed goal gates swing on base's own draws: this seed's base took rites 0.1/warren 0.133/lockpick 0.1 and squeezed every treated arm to 6/10 — more evidence that per-seed sweep readings are rate measurements, never single-event claims. The clean lineage remains the mission's best-documented artifact: 2.7× base aggregate, fully receipted stages 1–7, zero contamination anywhere in its history.
Knowledgebase Update
- Program evidence updated: pending results.
- Program backlog updated: pending results.
- Claim ledger updated: pending results.
Artifacts
src/— frozen vLLM runner (byte-identical to the source cell's).scripts/— staged harness, exposure pipeline, gate, benchmark runner, clean-chain rebuild script, vendored trainer/merger copies.configs/— frozen identity.data/— byte-copied corpora, exposure streams + receipts, gate files, clean-chain lineage package (data/lineage/).runs/— stage receipts (written by the staged GPU runs).reports/artifact_manifest.yaml
Report
Rendered from reports/report.md
Summary
The clean-path extension closed with a robust install and a scoping null. The statechain dose promoted on its THIRD distinct parent (21/40 axis strict over both controls; pooled retention deep in-band) — the install is now the program's most replicated designed effect. At the sealed event the clean candidate beat base 2.7× (0.3333 vs 0.1234, the strongest base draw yet) and its replay control, lost to its parent by 0.018 (pilot not promoted), and the frozen conversion reading came back FALSE: candidate rites 0.000 against the original lineage's 0.300 conversion. The data→family conversion is 1-for-2 and lineage-dependent — rites/sirens/mirage were exactly the C53 prefix's strengths, so the converter appears to require substrate the clean lineage lacks. Footnote: a strict menders WIN (0.017 draw) landed while rites collapsed. The clean lineage remains the mission's best-documented artifact: stages 1–7 fully receipted from the official base, zero contamination anywhere.
Research Program Fit
Method
Results
Controls
Oracle Versus Deployable Evidence
Interpretation
Next Experiments
Artifact Manifest
The frozen corpus, streams, receipts, gate inputs, and the complete clean-chain lineage package are in-repo; trained adapters and merges will live in this cell's own artifact storage with hashes pinned in receipts and reports/artifact_manifest.yaml.
Experiment log 3
Show the running log (3 entries, 2026-07-16)
2026-07-16 — Model-free design freeze
- Opened as the mission's cleanest artifact: the proven statechain converter on the zero-root parent, entire lineage documented end-to-end, blend-root absence enforced fail-closed.
- Treatment byte-copied from the proven cell (regenerates byte-identically); exposure exact zero-delta at the standard geometry; gates at 88,041 + 88,042/88,044/88,045 (88,043 taken, skipped); six-slot normalized-hash pin on the benchmark runner with fill-state-agnostic mutation fixtures (the lifecycle-22 lesson applied at build time). 127 tests green; smoke green; zero seed substitutions.
2026-07-16 — Local gate: PROMOTED (third install replication); benchmark authorized
- The 12-run pooled_k3 gate promoted statechain_clean on all eight checks: axis 21/40 strictly over the parent (19) and replay (16); pooled retention 61.33 vs 62.33/63.0 — the converter installs on its third distinct parent with retention held. Training-loss note recorded (clean chain ~1.3 vs original ~0.43 on identical data; loss-level ≠ capability).
- Six pin slots filled from committed receipts; the normalized-hash pin verified post-fill; PASS_BENCHMARK_EVENT granted; sealed 78,160 opens at the next green checkpoint.
2026-07-16 — The sealed event and closure
- All four arms ran clean at 78,160. Aggregates 0.1234 / 0.3517 / 0.3119 / 0.3333 (base — the strongest base draw yet / parent / replay / candidate). Pilot NOT promoted (beat base and replay; lost to the parent by 0.018). The frozen conversion reading: converts_on_clean_lineage FALSE — candidate rites 0.0 vs the original lineage's 0.300; the conversion is 1-for-2 and lineage-dependent, with the legible pattern that rites/sirens/mirage were the C53 prefix's strengths. Footnote: the candidate took a strict menders WIN (0.017, a draw) while rites collapsed — family movements remain draw-coupled.
- The clean lineage stands as the mission's best-documented artifact: 2.7× base, stages 1–7 fully receipted, zero contamination end-to-end.
Data files 1
Result tables and metrics copied from the experiment folder — preview inline or open the raw file.
Reproduce
Smoke test
PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_clean_path_statechain_extension/scripts/run.py --smokeFull run
PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_clean_path_statechain_extension/scripts/run.py --stage train-control (then train-candidate, merge-arms, local, benchmark; one stage per pushed checkpoint)Run steps are documented inside the experiment folder (README and scripts).