Research log Small Model Experimentation
GitHub

Clean-Path Statechain Extension

The skill installs anywhere; the family conversion needed the old ancestor's soil

The one idea you need

Earlier work proved two things separately: a small 160-example training dose reliably teaches the model to track hidden running state (and that skill carries over to a benchmark family it had always failed), and the program's best model can be rebuilt from scratch using only documented, contamination-free training steps. This experiment combines them: apply the exact proven dose, byte-for-byte, to the clean rebuilt model.

The question

Does the proven state-tracking dose install just as well on the fully documented clean-lineage model, and does it again unlock the benchmark family it unlocked before?

What we found

["Split verdict with a clean lesson. The state-tracking dose installed for the THIRD time on its third different parent — the program's most reliable trained effect — and the clean-lineage model beat the untouched base by 2.7 times while keeping its memory inside the calibrated margins. But the headline hope failed: on the original lineage this dose had TRIPLED the protocol-compliance benchmark family, and on the fully clean lineage that conversion vanished (zero, versus the original's 0.30). The pattern reads clearly: that family was one of exactly three the undocumented ancestor adapter was good at, so the taught skill seems to convert into benchmark scores only where the ancestor's training already tilled the soil. One consolation footnote: the clean model scored a strict win on the otherwise-impossible debugging family through a lucky draw. The fully-documented model — every training step receipted from the official base, zero contamination anywhere — stands as the mission's reference artifact."]

Why it matters

If it works, the program has ONE headline model whose every training step is documented, receipted, and reproducible from a blank start — no mystery ingredients anywhere — while carrying the strongest install the program has proven.

Install replications3/3three parents, three promotions
Rites conversion on clean ground0.0vs 0.30 on the prefix lineage
Clean model vs base2.7×0.3333 vs 0.1234 (strongest base draw)
Documented stages7fully receipted, zero contamination
On this page
  1. Results at a glance
  2. Overview
  3. Report
    1. Summary
    2. Research Program Fit
    3. Method
    4. Results
    5. Controls
    6. Oracle Versus Deployable Evidence
    7. Interpretation
    8. Next Experiments
    9. Artifact Manifest
  4. Experiment log
  5. Data files
  6. Reproduce
  7. Related

Results at a glance 1

The four models across the ten families at the sealed event

How to read

Grouped bars per family; note rites, where the clean candidate scored zero.

00.20.40.60.8chroniclechroniclelockpicklockpickmendersmendersmiragemirageritesritessiftstacksiftstacksirenssirensstockadestockadetoolsmithtoolsmithwarrenwarren

Takeaway → The clean lineage wins broadly over base but the taught skill's family conversion did not survive the ancestor's removal.

Data table
public benchmark familybasezero_root_parentreplay_ctl4statechain_clean
chronicle0.10.70.60.6
lockpick0.100.10
menders0000.017
mirage00.60.30.7
rites0.10.100
siftstack00.50.60.5
sirens0.40.50.40.4
stockade0.10.10.1320.273
toolsmith0.30.80.720.71
warren0.1330.2170.2670.133

Numbers from experiments/qwen35_4b_clean_path_statechain_extension/runs/benchmark/medium_tb1024_seed78160_pilot/summary.json

Technical framing

Per-family scores at medium tier, sealed seed 78160 (tb 1024) — PILOT_NOT_PROMOTED + CONVERSION_NOT_REPLICATED: the statechain install held on its third parent (local 21/40 strict; retention in-band — 3-for-3) and the clean candidate beat base 2.7x and its replay control, but lost to its parent by 0.018 and the rites conversion came back FALSE (candidate 0.0 vs the original lineage's 0.300) — the data-to-family converter is lineage-dependent: rites/sirens/mirage were exactly the C53 prefix's strengths. Footnote: a strict menders draw-win (0.017) landed while rites collapsed; this seed drew the strongest base ever (0.1234), squeezing all treated arms to 6/10.

In the author’s words from the Overview · “Results”

Local gate: PROMOTED on all eight checks — the statechain install's third replication, on its third distinct parent (axis 21/40 strictly over parent 19 and replay 16; pooled retention 61.33 vs 62.33/63.0, deep inside the calibrated bands). Training-loss property recorded: the clean chain fits the replay surface at ~1.3 versus the original lineage's ~0.43 while performing within ~10% at the benchmark (loss-level ≠ capability, dramatically). Sealed event at 78,160 (all arms authenticated; the six-slot normalized pin held through the fill) (table on the experiment page). Pilot: candidate > base ✓, > replay ✓, > parent ✗ (−0.018) — NOT promoted, the same shape as the original statechain cell. … Read the full result →

Overview

Research Program

  • Program: agentic_breadth_installation
  • Program question: can DESIGNED synthetic curricula install universal, transferable agentic skills into the one 4B — provably, on fully documented lineage?
  • Prior anchors: lifecycle 18 (qwen35_4b_statechain_only_dose — the statechain dose installs, 21/40 axis strict over both controls, and CONVERTS to the rites family: 0.300 vs 0.100/0.100 paired at sealed 78,154) and lifecycle 22 (qwen35_4b_zero_root_lineage_rebuild — the six documented stages replayed from a fresh zero-initialized adapter produce the zero-root composite, tree 414f5829…, weights 6e9aad25…, 0.3462 aggregate with 7/10 strict wins and zero losses at sealed 78,159).

Question

Lifecycle 23 — the mission's cleanest artifact. Does the PROVEN statechain converter dose, applied byte-identically to the ZERO-ROOT composite, produce a single installed model whose ENTIRE lineage is documented and contamination-free end-to-end — and does the rites conversion replicate ON THE CLEAN LINEAGE?

Hypothesis

The statechain install is a property of the dose, not of the blend-rooted parent it was first proven on: the same 160 frozen rows at the same exposure-matched geometry should clear the same calibrated gate from the zero-root composite, and the local install should again convert to the rites family at medium.

Setup

  • Model: Qwen/Qwen3.5-4B (revision 851bf6e8…), always.
  • Parent and adapter base: the zero-root composite (large_artifacts/qwen35_4b_zero_root_lineage_rebuild/merged/zero_root_hygiene_explore, tree 414f5829…, weights 6e9aad25…), authenticated against lifecycle 22's committed merge receipt (e906caea…; byte-identical provenance copy in data/lineage/provenance/merge.json).
  • Treatment: data/sft_statechain_only.jsonl — the source cell's frozen 160-row corpus copied BYTE-IDENTICALLY (ab6c7845…); fresh instances would change the treatment, so the byte-copy is the controlled choice. Replay pool sft_blend.jsonl (25a9595f…) byte-identical to every predecessor copy.
  • Arms: replay_ctl4 (control, trains FIRST) then statechain_clean (candidate); fresh rank-32/alpha-64 adapters, NO warm start, training seed 73, 1 epoch over 1,520 rows (190 optimizer updates, LR 1e-5, batch 1×8, max length 4,096, w_think/w_close 0.2).
  • Exposure: exact zero-delta three-axis MILP (forward / nonzero-target / absolute loss mass ×5) at the frozen geometry — 1,280-row shared stratified core + 240-row variable block (control: 240 replay; candidate: 160 treatment + 80 fillers), namespace seed 55,150.
  • Local gate (three arms: parent + both trained): 40-row statechain axis holdout at seed 88,041 (10 per formalism, FRESH instances from the copied generator) + three 104-row retention screens at 88,042/88,044/88,045 under pooled_k3 (88,043 is taken by qwen35_4b_counterfactual_plan_reflection_transfer — documented skip). Promotion: axis total strictly > parent AND > replay_ctl4; pooled retention bands on screen sums (correct −15, caps +9, parsed −9) vs BOTH controls.
  • Conditional benchmark (only on promotion): ONE sealed medium tb1024 event at fresh seed 78,160, four arms in frozen order — base (26d8ee48…), zero_root_parent (414f5829…), replay_ctl4, statechain_clean. Trained-arm pins are six fail-closed TODO-PIN slots in scripts/run_benchmark.py, frozen by check_design's NORMALIZED-HASH pin (lifecycle 22's mechanism).
  • Primary metric: local axis-holdout total (promotion), then pilot gate (candidate aggregate strictly > base AND > replay_ctl4 AND > zero_root_parent).
  • Frozen framing: menders is closed, so the winnable ceiling is 9/10; the readings of consequence are (a) the rites conversion ON THE CLEAN LINEAGE (candidate rites vs parent/replay rites, paired) and (b) the fully documented model's per-family profile. Any 10/10 is a menders draw and feeds a fresh confirmation cell before any claim.
  • Standalone: data/lineage/ carries the complete clean-chain package — the six zero-root stage datasets, lifecycle 22's stage + merge receipts as provenance documents, the trainer/merger copies, and a clean-chain manifest recording this cell's dose as STAGE 7. NO blend root exists anywhere in this cell (fail-closed).
  • Hidden-label boundary: gate answers live only in data/local_tasks_seed*.jsonl; the model-facing local_input_seed*.jsonl files carry id/messages/meta only. The benchmark suite directory is never read; only the trusted aggregate gateway runs.

Run

Smoke (no GPU, no writes):

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_clean_path_statechain_extension/scripts/run.py --smoke

Full (one stage per pushed checkpoint, each behind its review verdict):

.venv/bin/python -B experiments/qwen35_4b_clean_path_statechain_extension/scripts/run.py --stage train-control
# then: train-candidate, merge-arms, local, benchmark

Standalone lineage verification (no GPU) / full clean-chain rebuild (GPU):

.venv/bin/python -B experiments/qwen35_4b_clean_path_statechain_extension/scripts/rebuild_clean_chain.py --verify-inputs

Results

Local gate: PROMOTED on all eight checks — the statechain install's third replication, on its third distinct parent (axis 21/40 strictly over parent 19 and replay 16; pooled retention 61.33 vs 62.33/63.0, deep inside the calibrated bands). Training-loss property recorded: the clean chain fits the replay surface at ~1.3 versus the original lineage's ~0.43 while performing within ~10% at the benchmark (loss-level ≠ capability, dramatically).

Sealed event at 78,160 (all arms authenticated; the six-slot normalized pin held through the fill):

armaggregategoal gate vs baserites
base0.1234 (strongest base draw yet)0.100
zero_root_parent0.35176/100.100
statechain_clean0.33336/10 (incl. a strict MENDERS win, 0.017)0.000
replay_ctl40.31196/100.000

Pilot: candidate > base ✓, > replay ✓, > parent ✗ (−0.018) — NOT promoted, the same shape as the original statechain cell. The frozen conversion reading: converts_on_clean_lineage: false — candidate rites 0.0 against the original-lineage conversion's 0.300.

Interpretation

Three durable readings. (1) The statechain INSTALL is robust — three parents, three promotions, retention held each time under calibrated bands. (2) The CONVERSION is lineage-dependent: 1-for-2, expressed on the prefix lineage and absent on the clean one, and the pattern is legible — rites/sirens/mirage were precisely the C53 prefix's strengths, so the designed dose appears to convert only where the prefix's substrate already leans toward the family. The program's one proven data→family mechanism thus carries a substrate precondition, which scopes the conversion law honestly. (3) Per-seed goal gates swing on base's own draws: this seed's base took rites 0.1/warren 0.133/lockpick 0.1 and squeezed every treated arm to 6/10 — more evidence that per-seed sweep readings are rate measurements, never single-event claims. The clean lineage remains the mission's best-documented artifact: 2.7× base aggregate, fully receipted stages 1–7, zero contamination anywhere in its history.

Knowledgebase Update

  • Program evidence updated: pending results.
  • Program backlog updated: pending results.
  • Claim ledger updated: pending results.

Artifacts

  • src/ — frozen vLLM runner (byte-identical to the source cell's).
  • scripts/ — staged harness, exposure pipeline, gate, benchmark runner, clean-chain rebuild script, vendored trainer/merger copies.
  • configs/ — frozen identity.
  • data/ — byte-copied corpora, exposure streams + receipts, gate files, clean-chain lineage package (data/lineage/).
  • runs/ — stage receipts (written by the staged GPU runs).
  • reports/artifact_manifest.yaml

Report

Rendered from reports/report.md

Summary

The clean-path extension closed with a robust install and a scoping null. The statechain dose promoted on its THIRD distinct parent (21/40 axis strict over both controls; pooled retention deep in-band) — the install is now the program's most replicated designed effect. At the sealed event the clean candidate beat base 2.7× (0.3333 vs 0.1234, the strongest base draw yet) and its replay control, lost to its parent by 0.018 (pilot not promoted), and the frozen conversion reading came back FALSE: candidate rites 0.000 against the original lineage's 0.300 conversion. The data→family conversion is 1-for-2 and lineage-dependent — rites/sirens/mirage were exactly the C53 prefix's strengths, so the converter appears to require substrate the clean lineage lacks. Footnote: a strict menders WIN (0.017 draw) landed while rites collapsed. The clean lineage remains the mission's best-documented artifact: stages 1–7 fully receipted from the official base, zero contamination anywhere.

Research Program Fit

Method

Results

Controls

Oracle Versus Deployable Evidence

Interpretation

Next Experiments

Artifact Manifest

The frozen corpus, streams, receipts, gate inputs, and the complete clean-chain lineage package are in-repo; trained adapters and merges will live in this cell's own artifact storage with hashes pinned in receipts and reports/artifact_manifest.yaml.

Experiment log 3

Show the running log (3 entries, 2026-07-16)

2026-07-16 — Model-free design freeze

  • Opened as the mission's cleanest artifact: the proven statechain converter on the zero-root parent, entire lineage documented end-to-end, blend-root absence enforced fail-closed.
  • Treatment byte-copied from the proven cell (regenerates byte-identically); exposure exact zero-delta at the standard geometry; gates at 88,041 + 88,042/88,044/88,045 (88,043 taken, skipped); six-slot normalized-hash pin on the benchmark runner with fill-state-agnostic mutation fixtures (the lifecycle-22 lesson applied at build time). 127 tests green; smoke green; zero seed substitutions.

2026-07-16 — Local gate: PROMOTED (third install replication); benchmark authorized

  • The 12-run pooled_k3 gate promoted statechain_clean on all eight checks: axis 21/40 strictly over the parent (19) and replay (16); pooled retention 61.33 vs 62.33/63.0 — the converter installs on its third distinct parent with retention held. Training-loss note recorded (clean chain ~1.3 vs original ~0.43 on identical data; loss-level ≠ capability).
  • Six pin slots filled from committed receipts; the normalized-hash pin verified post-fill; PASS_BENCHMARK_EVENT granted; sealed 78,160 opens at the next green checkpoint.

2026-07-16 — The sealed event and closure

  • All four arms ran clean at 78,160. Aggregates 0.1234 / 0.3517 / 0.3119 / 0.3333 (base — the strongest base draw yet / parent / replay / candidate). Pilot NOT promoted (beat base and replay; lost to the parent by 0.018). The frozen conversion reading: converts_on_clean_lineage FALSE — candidate rites 0.0 vs the original lineage's 0.300; the conversion is 1-for-2 and lineage-dependent, with the legible pattern that rites/sirens/mirage were the C53 prefix's strengths. Footnote: the candidate took a strict menders WIN (0.017, a draw) while rites collapsed — family movements remain draw-coupled.
  • The clean lineage stands as the mission's best-documented artifact: 2.7× base, stages 1–7 fully receipted, zero contamination end-to-end.

Data files 1

Result tables and metrics copied from the experiment folder — preview inline or open the raw file.

Reproduce

Smoke test

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_clean_path_statechain_extension/scripts/run.py --smoke

Full run

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_clean_path_statechain_extension/scripts/run.py --stage train-control (then train-candidate, merge-arms, local, benchmark; one stage per pushed checkpoint)

Run steps are documented inside the experiment folder (README and scripts).

Browse the experiment folder on GitHub ↗