Rank-Capacity Vehicle Cell
The one idea you need
The forgetting tax is now a law: every practice dose costs five to ten retained answers on this training setup. The leading suspect is capacity — the small add-on module carries the whole training history, so new lessons must evict something. This trial doubles the module's size and gives the dose its own fresh module on a frozen base model.
The question
Does a double-size, fresh add-on module absorb the practice dose without evicting retained skills — or is the forgetting tax deeper than capacity?
What we found
The trial could not answer the capacity question — by design. Its built-in guard required the known nine-point forgetting case to reproduce before trusting any comparison, and on this fresh screen that case measured only five. Across the last four screens the same models' forgetting scores wobble by three to four points, which rivals the five-point rule itself. The double-size module measured seven points of forgetting and a weaker skill result, but none of that is judgeable until the measuring stick is calibrated.
Why it matters
A measurement rule that can be tripped by luck must be caught by its own guard — and it was. The next step costs no training at all: re-measure the built models on several fresh screens, size the wobble, and set rules that only real effects can pass. Every future forgetting claim depends on it.
On this page
Results at a glance 1
How to read
Bars per model: retained-skill totals on this screen.
Takeaway → The known forgetting case failed to reproduce, so no capacity conclusion was drawn — the measuring stick gets calibrated next, at zero training cost.
Data table
| explicit merged composite | retention correct (of 104) | axis holdout correct (of 40) |
|---|---|---|
| clean_parent | 69 | 17 |
| axis160_direct (r32) | 64 | 21 |
| axis160_r64 | 62 | 19 |
Numbers from experiments/qwen35_4b_rank_capacity_vehicle_cell/reports/report.md
Technical framing
Seed-88021 gate: retention and axis by arm (unadjudicated) — Verdict SCREEN_INSTABILITY: the r32 arm's known -9 re-measured at -5, tripping the guard; pooled across four gates the screen's seed noise (±3-4) rivals the ±5 band. The eval-only calibration study is the preregistered successor.
In the author’s words from the Overview · “Results”
The rank-64 arm trained cleanly (0.5554 train loss, 0 skips) and merged onto the composite. The three-arm gate at seed 88,021: retention correct of 104 — clean_parent 69, axis160_direct 64 (−5), axis160_r64 62 (−7). The rank-32 arm's known −9 failed to reproduce (−5, inside the band), so the ordered partition returned SCREEN_INSTABILITY and no capacity inference is made. Axis holdout of 40: axis160_direct 21, axis160_r64 19 (install_preserved: false — the fresh rank-64 adapter also under-delivered the install on this screen), clean_parent 17.
Overview
The vehicle study's sharpest single variable: a FRESH rank-64 adapter trained on the clean parent composite with the same twice-verified corpus and exposure, judged on a fresh screen beside the published rank-32 arm (known −9 retention) and the parent — asking whether capacity removes the intrinsic retention tax while preserving the install.
Research Program
- Program:
agentic_breadth_installation. - Program question: can synthetic curricula install general capability that lifts the held-out aggregate without a negative family?
- Prior anchors: the priced intrinsic-tax law (diversity and interleaving both refuted; screen fortune resolved); the interference arc's independent capacity implication; hygiene seven-for-seven as the install probe.
Question
Does doubling adapter capacity (rank 64, fresh adapter on the composite) absorb the dose without evicting retained skills (CAPACITY_SUPPORTED), fail to (CAPACITY_REFUTED), or does the known −9 fail to reproduce (SCREEN_INSTABILITY)?
Hypothesis
The tax is subspace competition: the rank-32 adapter carries the whole lineage, and a new dose must evict something. A fresh rank-64 adapter on the frozen composite separates the dose's parameters from the lineage's entirely — the cleanest capacity test available.
Setup
- Model: only
Qwen/Qwen3.5-4B, revision851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a. - Base for training: the
designed_freshmerged composite (weights0a3b89cd...979) via the per-experiment trainer's new--model-pathargument (encode_row byte-unchanged); FRESH rank-64/alpha-128 adapter, no warm start; otherwise identical geometry (1,520 rows, 190 updates, LR 1e-5, seed 58; inherited corpuse7a95d73...79e; slot seed 55,124). - Merge: the rank-64 adapter applied onto the composite weights (scale 2.0, fingerprint and module checks unchanged).
- Gate (seed 88,021): 40-task v1-kind axis holdout + 104-task retention; three weight-authenticated arms (
axis160_r64, publishedaxis160_direct,clean_parent); ordered verdict partition — SCREEN_INSTABILITY (the −9 fails to reproduce at ≥ −5) takes precedence, then CAPACITY_SUPPORTED (r64 ≥ −5 while r32 ≤ −6), else CAPACITY_REFUTED; plus aninstall_preservedflag (r64 axis total ≥ r32's on this screen). - No benchmark stage and no aggregate seed.
Run
Smoke:
PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_rank_capacity_vehicle_cell/scripts/run.py --smokeCheckpointed stages:
.venv/bin/python -B experiments/qwen35_4b_rank_capacity_vehicle_cell/scripts/run.py --stage train-candidate
.venv/bin/python -B experiments/qwen35_4b_rank_capacity_vehicle_cell/scripts/run.py --stage merge-candidate
.venv/bin/python -B experiments/qwen35_4b_rank_capacity_vehicle_cell/scripts/run.py --stage localResults
The rank-64 arm trained cleanly (0.5554 train loss, 0 skips) and merged onto the composite. The three-arm gate at seed 88,021: retention correct of 104 — clean_parent 69, axis160_direct 64 (−5), axis160_r64 62 (−7). The rank-32 arm's known −9 failed to reproduce (−5, inside the band), so the ordered partition returned SCREEN_INSTABILITY and no capacity inference is made. Axis holdout of 40: axis160_direct 21, axis160_r64 19 (install_preserved: false — the fresh rank-64 adapter also under-delivered the install on this screen), clean_parent 17.
Interpretation
The instability guard did exactly its job. Pooling the last four gates, the same composites' retention deltas scatter by three to four points across fresh screens (−9 then −5 for the rank-32 arm; −10 twice for the two-lesson arm; −5 for replay; −7 for rank-64): the 104-task retention screen's seed-to-seed noise is comparable to the ±5 band, which means single-screen delta comparisons near the band edge — including parts of the intrinsic-tax evidence — carry more draw noise than the bands assumed. The honest consolidated statement: doses cost retention on the order of five to ten points with screen noise of roughly ±3, and no single-screen reading should adjudicate a five-point band. Per the frozen branch, the successor is a retention-screen calibration study: the published composites re-measured across several fresh screens, eval-only, to size seed variance and set bands that separate real effects from draws.
Terminal Disposition
No later event is authorized here. No benchmark seed existed. The rank-64 composite and all receipts are preserved. Per the preregistered branch, the funded successor is the screen-calibration study; vehicle inference stays open until bands are recalibrated.
Knowledgebase Update
- Program evidence updated: the instability verdict and the pooled noise estimate recorded.
- Program backlog updated: the calibration study is the funded successor; capacity remains untested pending calibrated bands.
- Claim ledger updated: no.
Artifacts
data/sft_axis160.jsonl,data/corpus_manifest.json: inherited twice-verified corpus.data/stream_manifest.json,data/stream_token_receipt.json: exposure receipts.data/local_tasks_seed88021.jsonl,data/local_input_seed88021.jsonl,data/local_design_receipt.json: frozen gate.reports/preregistration.md,reports/design_review.md: contract and authorization.reports/artifact_manifest.yaml: external composite pins.
Report
Rendered from reports/report.md
Summary
Model-free construction under way for the vehicle study's first variable: a fresh rank-64 adapter on the clean-parent composite, same corpus and exposure, judged beside the re-measured rank-32 arm and the parent under an ordered three-way capacity verdict. No benchmark stage.
Research Program Fit
The direct capacity test the intrinsic-tax law and the interference arc both point to.
Method
See the preregistration.
Results
- Training (one arm, fresh rank-64 on the composite): 1,520 rows, 0 skips, 190 updates, 0.5554 train loss.
- Gate (seed 88,021): retention of 104 — clean_parent 69, axis160_direct 64 (−5), axis160_r64 62 (−7); axis of 40 — axis160_direct 21, axis160_r64 19 (install_preserved false), clean_parent 17.
- Verdict: SCREEN_INSTABILITY (the known −9 re-measured at −5); no capacity inference.
Controls
The published rank-32 composite re-measured on the same fresh screen; the parent baseline; hygiene as the seven-for-seven install probe; screen-instability branch guards the comparison.
Oracle Versus Deployable Evidence
Executable truth grades outputs only; benchmarks/ remains unread.
Next Stage
None. The preregistered branch funds the eval-only retention-screen calibration study; vehicle inference stays open until bands are recalibrated.
Artifact Manifest
Inherited corpus and model-free artifacts in-repo; comparison composites external with pinned receipts.
Experiment log 4
Show the running log (4 entries, 2026-07-15)
2026-07-15 — Model-free design freeze
- Opened as the vehicle study's first single-variable cell after the intrinsic-tax verdict. Rank capacity is the sharpest candidate mechanism (independently implicated by the interference arc).
- One trainer delta (
--model-path) lets a FRESH rank-64 adapter train on the frozen clean-parent composite; the corpus, exposure geometry, and gate design are inherited unchanged; the published rank-32 arm is re-measured on the same screen; the ordered three-way verdict is frozen with successor selection. - Seeds
55124/58/88021; no aggregate seed exists. - No model, GPU, training, local, or benchmark event has run.
2026-07-15 — Authenticated rank-64 training
train-candidateran only after the freeze checkpoint matchedorigin/mainwith both workflows green; the fresh rank-64/alpha-128 adapter trained on the pinned clean-parent composite (full-weights preflight) with 1,520/1,520 rows, 0 skipped, 190 updates; receipt/log published and pinned fail-closed.
2026-07-15 — Authenticated rank-64 composite
merge-candidateran only after the training checkpoint matchedorigin/mainwith both workflows green; the rank-64 adapter merged onto the clean-parent composite (scale 2.0, 128/128 modules, fingerprint-verified); the tree pin filled fail-closed. The one frozen three-arm verdict gate at seed 88,021 is the only next stage.
2026-07-15 — Verdict: SCREEN_INSTABILITY; cell closed
- The three-arm gate ran from the merge checkpoint: retention 69 / 64 / 62 (parent / r32 / r64). The r32 arm's known −9 re-measured at −5, tripping the instability guard; no capacity inference was made. Axis: r32 21, r64 19 (install_preserved false), parent 17.
- Pooled across the last four gates, same-composite retention deltas scatter ±3-4 points — the screen's seed noise is comparable to the ±5 band. The preregistered successor is an eval-only retention-screen calibration study.
Reproduce
Smoke test
PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_rank_capacity_vehicle_cell/scripts/run.py --smokeFull run
checkpointed scripts/run.py stages onlyRun steps are documented inside the experiment folder (README and scripts).