Research log Small Model Experimentation
GitHub

Dose-Diversity Mechanism Cell

Forgetting is intrinsic; the trade is priced

The one idea you need

Every practice dose has cost the model about ten retained answers — except one, which used a larger, more varied practice set. After the schedule theory died in a direct test, one measurement remains: teach that larger varied set directly, with no review round, and see whether variety itself is what protects memory.

The question

Does a bigger, more varied practice set protect retained skills where the small focused set does not — or is forgetting simply the price of practice on this model?

What we found

Variety does not protect memory: the larger varied practice set also cost nine retained answers, the known ten-point case reproduced exactly, and even pure review cost five on this fresh screen — so the one past case with zero forgetting was measurement luck. Meanwhile the skills themselves keep landing: the hygiene lesson went ten-for-ten, its seventh straight category win, and the varied-dose model posted the best overall skill score with the cleanest finishing behavior.

Why it matters

Four trials of questions collapse into one priced law: on this training setup, teaching new skills costs five to ten retained answers, full stop. The choice ahead is engineering — change the training vehicle, or accept and price the trade — and both paths are preregistered.

Verdictintrinsicvariety -9; known case -10 reproduced; review -5
Hygiene lesson10 / 10seventh straight category win — perfect record
Best skill score26 / 40the varied dose, with the fewest run-ons (5)
Lucky precedentresolvedthe zero-forgetting case was screen fortune
On this page
  1. Results at a glance
  2. Overview
  3. Report
    1. Summary
    2. Research Program Fit
    3. Method
    4. Results
    5. Controls
    6. Oracle Versus Deployable Evidence
    7. Next Stage
    8. Artifact Manifest
  4. Experiment log
  5. Reproduce
  6. Related

Results at a glance 1

Every dose pays the memory tax; the skills land anyway

How to read

Bars per model: retained-skill total (of 104) and unseen skill total (of 40).

020406080clean_parentclean_parent7024replay_cleanreplay_clean6523axis160_directaxis160_direct6126hygiene_explore_directhygiene_explore_direct6024

Takeaway → The parent keeps the most memory; every trained arm pays five to ten points for its gains — the trade is real, two-sided, and now priced.

Data table
explicit merged compositeretention correct (of 104)axis holdout correct (of 40)
clean_parent7024
replay_clean6523
axis160_direct6126
hygiene_explore_direct6024

Numbers from experiments/qwen35_4b_dose_diversity_mechanism_cell/reports/report.md

Technical framing

Seed-88020 verdict gate: retention and axis by arm — Verdict REFUTED_INTRINSIC: the diverse dose also broke the retention band (-9), the known -10 reproduced, and replay itself measured -5 — the retention cost is intrinsic to this vehicle and the sole safe precedent was screen fortune. Hygiene hit 10/10 (seven consecutive wins).

In the author’s words from the Overview · “Results”

The single arm trained cleanly (0.5068 train loss, 0 skips) and merged. The four-arm gate at seed 88,020 (fresh screen, normalized grading): retention correct of 104 — clean_parent 70, replay_clean 65 (−5), axis160_direct 61 (−9), hygiene_explore_direct 60 (−10, reproducing its known cost exactly). Preregistered verdict: REFUTED_INTRINSIC — the diverse dose broke the band too. Axis holdout of 40: axis160_direct 26 (best; hygiene 10/10 — the seventh consecutive hygiene win, now perfect; caps 5, best), hygiene_explore_direct 24, clean_parent 24, replay_clean 23.

Overview

The single missing measurement that adjudicates why designed doses cost retention: the verified 160-row/4-kind corpus dosed DIRECTLY from the clean parent (no replay round), judged on a fresh screen alongside the re-measured 80-row/2-kind dose (known −10), the replay round, and the parent — with a preregistered three-way verdict.

Research Program

  • Program: agentic_breadth_installation.
  • Program question: can synthetic curricula install general capability that lifts the held-out aggregate without a negative family?
  • Prior anchors: the interleaving refutation and its escalation rule; the sole retention-safe dose event (160-row corpus) versus two ~−10 events (80-row corpus); hygiene six-for-six as the install probe.

Question

Does corpus size/diversity at matched per-kind dose protect retention (SUPPORTED), is the retention cost intrinsic to dosing this vehicle (REFUTED_INTRINSIC), or does the known −10 fail to reproduce on a fresh screen (SCREEN_FORTUNE_SUSPECT)?

Hypothesis

Diversity: the 160/4-kind block dilutes per-kind gradient pressure on the shared representation that retained skills live on. Each alternative outcome selects a different, preregistered successor.

Setup

  • Model: only Qwen/Qwen3.5-4B, revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
  • Parent: designed_fresh composite (tree 93433aa2...255); warm start from its adapter (36f41095...442).
  • New arm: axis160_direct — the twice-verified 160-row v1 axis corpus (sha e7a95d73...79e), inherited byte-identically, in the standard exact-exposure stream (slot seed 55,123; training seed 57; 190 updates).
  • Gate (seed 88,020): 40-row v1-kind axis holdout + 104-row retention screen; four arms — the new merge plus three published composites (hygiene_explore_direct with its known −10, replay_clean, clean_parent); normalization unchanged; NO promotion — the receipt carries the per-arm table and the three-way diversity_mechanism verdict (bands: SUPPORTED if axis160 retention ≥ parent−5 while hygiene_explore ≤ parent−6; REFUTED_INTRINSIC if axis160 ≤ parent−6; SCREEN_FORTUNE_SUSPECT if the −10 fails to reproduce at ≥ parent−5).
  • No benchmark stage and no aggregate seed: a mechanism cell mints no claims.

Run

Smoke:

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_dose_diversity_mechanism_cell/scripts/run.py --smoke

Checkpointed stages:

.venv/bin/python -B experiments/qwen35_4b_dose_diversity_mechanism_cell/scripts/run.py --stage train-candidate
.venv/bin/python -B experiments/qwen35_4b_dose_diversity_mechanism_cell/scripts/run.py --stage merge-candidate
.venv/bin/python -B experiments/qwen35_4b_dose_diversity_mechanism_cell/scripts/run.py --stage local

Results

The single arm trained cleanly (0.5068 train loss, 0 skips) and merged. The four-arm gate at seed 88,020 (fresh screen, normalized grading): retention correct of 104 — clean_parent 70, replay_clean 65 (−5), axis160_direct 61 (−9), hygiene_explore_direct 60 (−10, reproducing its known cost exactly). Preregistered verdict: REFUTED_INTRINSIC — the diverse dose broke the band too. Axis holdout of 40: axis160_direct 26 (best; hygiene 10/10 — the seventh consecutive hygiene win, now perfect; caps 5, best), hygiene_explore_direct 24, clean_parent 24, replay_clean 23.

Interpretation

The mechanism question is answered: at this vehicle (rank-32 LoRA continued in place, 190 updates, LR 1e-5), designed doses cost roughly five to ten retention points intrinsically — corpus diversity does not protect it, replay interleaving does not protect it (prior refutation), and even a pure replay round costs about five on a fresh screen. The single retention-byte-equal precedent was screen fortune, exactly as the SCREEN_FORTUNE alternative anticipated for the OTHER arm. Meanwhile the installs themselves are unambiguous: hygiene is now seven-for-seven across every parent, dose size, and recipe, and the diverse dose posted the best axis total and termination in this event. The program-level law: install-versus-retention is a real, priced trade at this vehicle. Successors must either change the vehicle (rank, loss weighting, update count) or preregister gates that price the trade rather than demand its absence.

Terminal Disposition

No later event is authorized here. No benchmark seed existed. All four composites and the verdict receipt are preserved. Per the preregistered branch, the funded successor is a dose-vehicle study (rank / loss weights / update count as single variables against this same gate design), with its own intake.

Knowledgebase Update

  • Program evidence updated: the intrinsic-cost verdict, the screen-fortune resolution, and hygiene's seven-for-seven recorded.
  • Program backlog updated: the vehicle study is the funded successor; the recipe search stays closed.
  • Claim ledger updated: no.

Artifacts

  • data/sft_axis160.jsonl, data/corpus_manifest.json: inherited twice-verified corpus.
  • data/stream_manifest.json, data/stream_token_receipt.json: exposure receipts.
  • data/local_tasks_seed88020.jsonl, data/local_input_seed88020.jsonl, data/local_design_receipt.json: frozen gate.
  • reports/preregistration.md, reports/design_review.md: contract and authorization.
  • reports/artifact_manifest.yaml: external composite pins.

Report

Rendered from reports/report.md

Summary

Model-free construction under way for the escalation rule's funded successor: one new training arm (the twice-verified 160-row corpus, direct from the clean parent) gated on fresh instruments alongside three published composites, with a preregistered three-way mechanism verdict and no benchmark stage.

Research Program Fit

The single-variable cell that adjudicates the retention-cost mechanism after the interleaving refutation.

Method

See the preregistration.

Results

  • Training (one arm): 1,520 rows, 0 skips, 190 updates, 0.5068 train loss.
  • Gate (seed 88,020, four weight-authenticated arms): retention of 104 — clean_parent 70, replay_clean 65 (−5), axis160_direct 61 (−9), hygiene_explore_direct 60 (−10, reproducing its known cost). Axis of 40 — axis160_direct 26 (hygiene 10/10, caps 5), hygiene_explore_direct 24, clean_parent 24, replay_clean 23.
  • Preregistered verdict: REFUTED_INTRINSIC.

Controls

Three published, weight-authenticated composites re-measured on the same fresh screen, including the known −10 arm; hygiene rides as the six-for-six install probe.

Oracle Versus Deployable Evidence

Executable truth grades outputs only; benchmarks/ remains unread.

Next Stage

None. The verdict selects the dose-vehicle study (or the priced-trade gate path) as the funded successor; the recipe search stays closed.

Artifact Manifest

Inherited corpus and model-free artifacts in-repo; comparison composites external with pinned receipts.

Experiment log 4

Show the running log (4 entries, 2026-07-15)

2026-07-15 — Model-free design freeze

  • Opened as the escalation rule's funded successor after the interleaving refutation. Single variable: corpus size/diversity at matched per-kind dose.
  • The twice-verified 160-row v1 axis corpus inherited byte-identically; one training arm from the clean parent; the gate re-measures the known −10 arm on the same fresh screen; the three-way verdict is preregistered with its successor selection.
  • Seeds 55123/57/88020; no aggregate seed exists for this cell.
  • No model, GPU, training, local, or benchmark event has run.

2026-07-15 — Authenticated candidate training

  • train-candidate ran only after the amended freeze checkpoint matched origin/main with both workflows green and a clean worktree.
  • axis160_direct trained 1,520/1,520 rows with 0 skipped over 190 updates; receipt/log published and pinned fail-closed. The single merge is next.

2026-07-15 — Authenticated candidate composite

  • merge-candidate ran only after the training checkpoint matched origin/main with both workflows green; the composite merged (128/128 modules, fingerprint-verified); the tree pin filled fail-closed. The one frozen four-arm gate event at seed 88,020 is the only next stage.

2026-07-15 — Verdict: REFUTED_INTRINSIC; cell closed

  • The four-arm gate ran from the merge checkpoint: retention 70 / 65 / 61 / 60 (parent / replay / axis160 / hygiene-explore) — the diverse dose broke the band (−9), the known −10 reproduced, and replay itself measured −5.
  • Verdict REFUTED_INTRINSIC per the frozen partition: the retention cost is intrinsic to this dose vehicle; the sole retention-safe precedent was screen fortune. Hygiene reached 10/10 (seventh consecutive win); axis160_direct posted the best axis total (26/40) and best caps (5).
  • The preregistered successor is the dose-vehicle study. No benchmark seed existed; the cell mints no claim.

Reproduce

Smoke test

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_dose_diversity_mechanism_cell/scripts/run.py --smoke

Full run

checkpointed scripts/run.py stages only

Run steps are documented inside the experiment folder (README and scripts).

Browse the experiment folder on GitHub ↗