Research log Small Model Experimentation
GitHub

Rank-Capacity Vehicle Cell

Screen too noisy to judge; guard fired

The one idea you need

The forgetting tax is now a law: every practice dose costs five to ten retained answers on this training setup. The leading suspect is capacity — the small add-on module carries the whole training history, so new lessons must evict something. This trial doubles the module's size and gives the dose its own fresh module on a frozen base model.

The question

Does a double-size, fresh add-on module absorb the practice dose without evicting retained skills — or is the forgetting tax deeper than capacity?

What we found

The trial could not answer the capacity question — by design. Its built-in guard required the known nine-point forgetting case to reproduce before trusting any comparison, and on this fresh screen that case measured only five. Across the last four screens the same models' forgetting scores wobble by three to four points, which rivals the five-point rule itself. The double-size module measured seven points of forgetting and a weaker skill result, but none of that is judgeable until the measuring stick is calibrated.

Why it matters

A measurement rule that can be tripped by luck must be caught by its own guard — and it was. The next step costs no training at all: re-measure the built models on several fresh screens, size the wobble, and set rules that only real effects can pass. Every future forgetting claim depends on it.

Guardfiredthe known -9 re-measured at -5 — comparison refused
Screen wobble±3-4same models across four fresh screens
Rank-64 readings-7 / 19retention / skill — recorded, not adjudicated
Next step costeval onlyno training; calibrate the screen itself
On this page
  1. Results at a glance
  2. Overview
  3. Report
    1. Summary
    2. Research Program Fit
    3. Method
    4. Results
    5. Controls
    6. Oracle Versus Deployable Evidence
    7. Next Stage
    8. Artifact Manifest
  4. Experiment log
  5. Reproduce
  6. Related

Results at a glance 1

The guard refused the comparison: the screen wobbles as much as the rule

How to read

Bars per model: retained-skill totals on this screen.

020406080clean_parentclean_parent6917axis160_direct (r32)axis160_direct (r32)6421axis160_r64axis160_r646219

Takeaway → The known forgetting case failed to reproduce, so no capacity conclusion was drawn — the measuring stick gets calibrated next, at zero training cost.

Data table
explicit merged compositeretention correct (of 104)axis holdout correct (of 40)
clean_parent6917
axis160_direct (r32)6421
axis160_r646219

Numbers from experiments/qwen35_4b_rank_capacity_vehicle_cell/reports/report.md

Technical framing

Seed-88021 gate: retention and axis by arm (unadjudicated) — Verdict SCREEN_INSTABILITY: the r32 arm's known -9 re-measured at -5, tripping the guard; pooled across four gates the screen's seed noise (±3-4) rivals the ±5 band. The eval-only calibration study is the preregistered successor.

In the author’s words from the Overview · “Results”

The rank-64 arm trained cleanly (0.5554 train loss, 0 skips) and merged onto the composite. The three-arm gate at seed 88,021: retention correct of 104 — clean_parent 69, axis160_direct 64 (−5), axis160_r64 62 (−7). The rank-32 arm's known −9 failed to reproduce (−5, inside the band), so the ordered partition returned SCREEN_INSTABILITY and no capacity inference is made. Axis holdout of 40: axis160_direct 21, axis160_r64 19 (install_preserved: false — the fresh rank-64 adapter also under-delivered the install on this screen), clean_parent 17.

Overview

The vehicle study's sharpest single variable: a FRESH rank-64 adapter trained on the clean parent composite with the same twice-verified corpus and exposure, judged on a fresh screen beside the published rank-32 arm (known −9 retention) and the parent — asking whether capacity removes the intrinsic retention tax while preserving the install.

Research Program

  • Program: agentic_breadth_installation.
  • Program question: can synthetic curricula install general capability that lifts the held-out aggregate without a negative family?
  • Prior anchors: the priced intrinsic-tax law (diversity and interleaving both refuted; screen fortune resolved); the interference arc's independent capacity implication; hygiene seven-for-seven as the install probe.

Question

Does doubling adapter capacity (rank 64, fresh adapter on the composite) absorb the dose without evicting retained skills (CAPACITY_SUPPORTED), fail to (CAPACITY_REFUTED), or does the known −9 fail to reproduce (SCREEN_INSTABILITY)?

Hypothesis

The tax is subspace competition: the rank-32 adapter carries the whole lineage, and a new dose must evict something. A fresh rank-64 adapter on the frozen composite separates the dose's parameters from the lineage's entirely — the cleanest capacity test available.

Setup

  • Model: only Qwen/Qwen3.5-4B, revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
  • Base for training: the designed_fresh merged composite (weights 0a3b89cd...979) via the per-experiment trainer's new --model-path argument (encode_row byte-unchanged); FRESH rank-64/alpha-128 adapter, no warm start; otherwise identical geometry (1,520 rows, 190 updates, LR 1e-5, seed 58; inherited corpus e7a95d73...79e; slot seed 55,124).
  • Merge: the rank-64 adapter applied onto the composite weights (scale 2.0, fingerprint and module checks unchanged).
  • Gate (seed 88,021): 40-task v1-kind axis holdout + 104-task retention; three weight-authenticated arms (axis160_r64, published axis160_direct, clean_parent); ordered verdict partition — SCREEN_INSTABILITY (the −9 fails to reproduce at ≥ −5) takes precedence, then CAPACITY_SUPPORTED (r64 ≥ −5 while r32 ≤ −6), else CAPACITY_REFUTED; plus an install_preserved flag (r64 axis total ≥ r32's on this screen).
  • No benchmark stage and no aggregate seed.

Run

Smoke:

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_rank_capacity_vehicle_cell/scripts/run.py --smoke

Checkpointed stages:

.venv/bin/python -B experiments/qwen35_4b_rank_capacity_vehicle_cell/scripts/run.py --stage train-candidate
.venv/bin/python -B experiments/qwen35_4b_rank_capacity_vehicle_cell/scripts/run.py --stage merge-candidate
.venv/bin/python -B experiments/qwen35_4b_rank_capacity_vehicle_cell/scripts/run.py --stage local

Results

The rank-64 arm trained cleanly (0.5554 train loss, 0 skips) and merged onto the composite. The three-arm gate at seed 88,021: retention correct of 104 — clean_parent 69, axis160_direct 64 (−5), axis160_r64 62 (−7). The rank-32 arm's known −9 failed to reproduce (−5, inside the band), so the ordered partition returned SCREEN_INSTABILITY and no capacity inference is made. Axis holdout of 40: axis160_direct 21, axis160_r64 19 (install_preserved: false — the fresh rank-64 adapter also under-delivered the install on this screen), clean_parent 17.

Interpretation

The instability guard did exactly its job. Pooling the last four gates, the same composites' retention deltas scatter by three to four points across fresh screens (−9 then −5 for the rank-32 arm; −10 twice for the two-lesson arm; −5 for replay; −7 for rank-64): the 104-task retention screen's seed-to-seed noise is comparable to the ±5 band, which means single-screen delta comparisons near the band edge — including parts of the intrinsic-tax evidence — carry more draw noise than the bands assumed. The honest consolidated statement: doses cost retention on the order of five to ten points with screen noise of roughly ±3, and no single-screen reading should adjudicate a five-point band. Per the frozen branch, the successor is a retention-screen calibration study: the published composites re-measured across several fresh screens, eval-only, to size seed variance and set bands that separate real effects from draws.

Terminal Disposition

No later event is authorized here. No benchmark seed existed. The rank-64 composite and all receipts are preserved. Per the preregistered branch, the funded successor is the screen-calibration study; vehicle inference stays open until bands are recalibrated.

Knowledgebase Update

  • Program evidence updated: the instability verdict and the pooled noise estimate recorded.
  • Program backlog updated: the calibration study is the funded successor; capacity remains untested pending calibrated bands.
  • Claim ledger updated: no.

Artifacts

  • data/sft_axis160.jsonl, data/corpus_manifest.json: inherited twice-verified corpus.
  • data/stream_manifest.json, data/stream_token_receipt.json: exposure receipts.
  • data/local_tasks_seed88021.jsonl, data/local_input_seed88021.jsonl, data/local_design_receipt.json: frozen gate.
  • reports/preregistration.md, reports/design_review.md: contract and authorization.
  • reports/artifact_manifest.yaml: external composite pins.

Report

Rendered from reports/report.md

Summary

Model-free construction under way for the vehicle study's first variable: a fresh rank-64 adapter on the clean-parent composite, same corpus and exposure, judged beside the re-measured rank-32 arm and the parent under an ordered three-way capacity verdict. No benchmark stage.

Research Program Fit

The direct capacity test the intrinsic-tax law and the interference arc both point to.

Method

See the preregistration.

Results

  • Training (one arm, fresh rank-64 on the composite): 1,520 rows, 0 skips, 190 updates, 0.5554 train loss.
  • Gate (seed 88,021): retention of 104 — clean_parent 69, axis160_direct 64 (−5), axis160_r64 62 (−7); axis of 40 — axis160_direct 21, axis160_r64 19 (install_preserved false), clean_parent 17.
  • Verdict: SCREEN_INSTABILITY (the known −9 re-measured at −5); no capacity inference.

Controls

The published rank-32 composite re-measured on the same fresh screen; the parent baseline; hygiene as the seven-for-seven install probe; screen-instability branch guards the comparison.

Oracle Versus Deployable Evidence

Executable truth grades outputs only; benchmarks/ remains unread.

Next Stage

None. The preregistered branch funds the eval-only retention-screen calibration study; vehicle inference stays open until bands are recalibrated.

Artifact Manifest

Inherited corpus and model-free artifacts in-repo; comparison composites external with pinned receipts.

Experiment log 4

Show the running log (4 entries, 2026-07-15)

2026-07-15 — Model-free design freeze

  • Opened as the vehicle study's first single-variable cell after the intrinsic-tax verdict. Rank capacity is the sharpest candidate mechanism (independently implicated by the interference arc).
  • One trainer delta (--model-path) lets a FRESH rank-64 adapter train on the frozen clean-parent composite; the corpus, exposure geometry, and gate design are inherited unchanged; the published rank-32 arm is re-measured on the same screen; the ordered three-way verdict is frozen with successor selection.
  • Seeds 55124/58/88021; no aggregate seed exists.
  • No model, GPU, training, local, or benchmark event has run.

2026-07-15 — Authenticated rank-64 training

  • train-candidate ran only after the freeze checkpoint matched origin/main with both workflows green; the fresh rank-64/alpha-128 adapter trained on the pinned clean-parent composite (full-weights preflight) with 1,520/1,520 rows, 0 skipped, 190 updates; receipt/log published and pinned fail-closed.

2026-07-15 — Authenticated rank-64 composite

  • merge-candidate ran only after the training checkpoint matched origin/main with both workflows green; the rank-64 adapter merged onto the clean-parent composite (scale 2.0, 128/128 modules, fingerprint-verified); the tree pin filled fail-closed. The one frozen three-arm verdict gate at seed 88,021 is the only next stage.

2026-07-15 — Verdict: SCREEN_INSTABILITY; cell closed

  • The three-arm gate ran from the merge checkpoint: retention 69 / 64 / 62 (parent / r32 / r64). The r32 arm's known −9 re-measured at −5, tripping the instability guard; no capacity inference was made. Axis: r32 21, r64 19 (install_preserved false), parent 17.
  • Pooled across the last four gates, same-composite retention deltas scatter ±3-4 points — the screen's seed noise is comparable to the ±5 band. The preregistered successor is an eval-only retention-screen calibration study.

Reproduce

Smoke test

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_rank_capacity_vehicle_cell/scripts/run.py --smoke

Full run

checkpointed scripts/run.py stages only

Run steps are documented inside the experiment folder (README and scripts).

Browse the experiment folder on GitHub ↗