Research log Small Model Experimentation
GitHub

Goal-Gap Axis Curriculum

Skills installed; benchmark chose review

The one idea you need

Instead of teaching everything at once, teach exactly the four abilities where the held-out benchmark has been stuck for months: fixing a buggy program from its wrong output, finding a route under a strict step budget, ignoring instructions smuggled inside documents, and following a freshly written procedure with interacting switches.

The question

Can four purpose-built practice sets — each varied across many invented styles — install the four abilities every previous model failed on, without damaging anything else?

What we found

The targeted practice worked on its own terms — the first screen pass in this program's history (28 vs 22 and 18 of 40 on unseen tasks, with zero forgetting). On the held-out benchmark it beat the untrained model by a wide margin (0.42 vs 0.11) with seven families up, none down, and one stuck family flipped. But plain review of old material scored even higher (0.51), so the frozen rule closed the experiment. Two families never moved for anyone: program repair stayed at zero and injection resistance at exactly one-half, for every model.

Why it matters

Targeted practice installs skills without forgetting, but installed skills and benchmark family scores are separated by more than vocabulary. And the cheapest interventionre-practicing old material — keeps compounding the total score. The goal is now blocked by exactly two frozen families, which is a much sharper question than it was yesterday.

Screen result28 vs 22 / 18first promotion in this line; zero forgetting
Benchmark vs base+0.317 families up, 0 down, warren flipped
But plain review0.51beat the candidate's 0.42 — rule closed the trial
Frozen families2program repair at 0; injection resistance at exactly 0.5
On this page
  1. Results at a glance
  2. Overview
  3. Report
    1. Summary
    2. Research Program Fit
    3. Method
    4. Results
    5. Controls
    6. Oracle Versus Deployable Evidence
    7. Next Stage
    8. Artifact Manifest
  4. Experiment log
  5. Data files
  6. Reproduce
  7. Related

Results at a glance 1

Benchmark totals: practice beat base widely, review beat practice

How to read

One bar per model: the untrained base, the axis-practice model, its parent, and the plain-review model.

00.20.40.6basebase0.108axis_curriculumaxis_curriculum0.422designed_fresh_parentdesigned_fresh_parent0.464replay_repeatreplay_repeat0.508

Takeaway → Every trained model crushes base; plain review of old material tops the table, so the targeted-practice model did not advance.

Data table
explicit merged compositemenagerie quick aggregate
base0.108
axis_curriculum0.422
designed_fresh_parent0.464
replay_repeat0.508

Numbers from experiments/qwen35_4b_goal_gap_axis_curriculum_target_match/reports/report.md

Technical framing

Seed-78144 quick aggregate by arm (tb 1024) — The axis candidate beat base +0.3138 with 7 positive / 0 negative families (warren flipped) but lost the aggregate to its parent and to the replay control, which flipped rites and posted the line's best recorded aggregate. menders (0) and sirens (0.500) were identical for every arm.

In the author’s words from the Overview · “Results”

Both arms trained cleanly (replay 0.3776, axis 0.4884 train loss; 0 skips each) and merged; the frozen 144-task gate event PROMOTED axis_curriculum — the first promotion in this program's universal line: axis holdout 28/40 versus parent 22 and replay 18 (kind wins on hygiene 9-vs-5/5, explore 7-vs-6/3, tracefix 4-vs-3/2; protocol tied at the preregistered control ceiling), with retention byte-equal to the parent (71/95/9) while replay drifted (65/89/15). The conditional aggregate pilot at seed 78,144 (quick, think budget 1,024) then ran all four composites: base 0.1085, axis_curriculum 0.4223, parent 0.4644, replay_repeat 0.5081. … Read the full result →

Overview

Attack the benchmark's four empirically stuck families directly: designed, contamination-free, single-turn atom curricula built only from public axis descriptions — multi-formalism program repair, budgeted unique-shortest route search, instruction hygiene with parseable-wrong decoys, and branch-balanced protocol execution — trained on the surface-general designed_fresh parent against a three-axis exact-exposure replay control, with a prospectively achievable two-instrument gate.

Research Program

  • Program: agentic_breadth_installation.
  • Program question: can synthetic curricula install general capability that lifts the held-out aggregate without a negative family?
  • Prior anchors: goal-gap forensics (the all-families goal passed once in 65 quick events; failures concentrate on four families whose gym analogues near-mastered internally without transferring); the fresh-surface trial's positive surface-generality reading and its structurally unpassable induct floor.

Question

Does axis-targeted designed-atom content — varied formalisms within each stuck axis — install held-out-surface capability on those axes without regressing the thirteen retained skills, where single-surface harvested episodes did not transfer?

Hypothesis

Coverage-without-transfer is a surface-binding failure: each gym axis lived on one surface, so the installed skill bound to the skin. Designed atoms that vary formalism and vocabulary within each axis should bind to structure (as the fresh-surface trial proved for the generic dose), and the hygiene block's decoy construction makes injection-compliance a parseable wrong answer, scoring resistance directly. Exact three-axis exposure matching against replay makes any win attributable to content.

Setup

  • Model: only Qwen/Qwen3.5-4B, revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
  • Parent: authenticated designed_fresh composite (tree 93433aa2...0255, weights 0a3b89cd...7979); warm start from its adapter (36f41095...b442). Runtime LoRA forbidden.
  • Treatment corpus (construction seed 77,117): sft_axis160.jsonl — 40 rows each of u_tracefix (four invented executable formalisms, unique repair enforced by exhaustive enumeration), u_explore (unique shortest route under an exact move budget, compact frontier-notation thinking), u_hygiene (embedded directives carry format-matched decoys; 30 injected / 10 clean), u_protocol (documented flags plus tally; all three closing branches trained). Banned-vocabulary audit covers benchmark family names, gym families and their flavor nouns, and every predecessor surface pool and attribute set.
  • Streams: 1,280 shared position-aligned replay rows + one 240-row variable block per arm (candidate = 160 treatment + 80 fillers; control = 240 replay), EXACT equality on forward tokens, loss-bearing targets, and absolute loss mass (MILP, namespace seed 55,118); 1,520 rows per arm.
  • Training: one event per arm, 190 updates, LR 1e-5, rank 32 alpha 64, think/close weights 0.2/0.2, seed 52, zero skips required.
  • Local gate (one event, fresh seed 88,014, three merged composites): two instruments — a 40-task axis holdout (10 per axis, unseen seed) and a 104-task retention screen (8 per original skill). Promotion: candidate strictly beats parent AND replay on axis-holdout total and on at least 3 of 4 axis kinds; retention within non-inferiority bands (correct ≥ each control − 5; caps ≤ each control + 3; parsed ≥ each control − 3); route abstentions ≤ 4. No absolute per-kind floors.
  • Conditional aggregate: sealed fresh seed 78,144, quick tier, think budget 1,024, four composites (base, parent, replay control, candidate). Gates: candidate aggregate strictly above base, replay, and parent; the every-family-strict-versus-base record is reported as the goal gate with the frozen power statement (a quick-tier pilot failure of that gate is expected even under the hypothesis; family-level confirmation belongs to the medium tier).
  • Hidden boundary: benchmarks/ unread; higher-tier confirmation and matched-compute sample-more remain mandatory before any universal claim.

Run

Smoke:

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_goal_gap_axis_curriculum_target_match/scripts/run.py --smoke

Checkpointed stages (each requires its prerequisite committed at a clean, pushed, green main):

.venv/bin/python -B experiments/qwen35_4b_goal_gap_axis_curriculum_target_match/scripts/run.py --stage train-control
.venv/bin/python -B experiments/qwen35_4b_goal_gap_axis_curriculum_target_match/scripts/run.py --stage train-candidate
.venv/bin/python -B experiments/qwen35_4b_goal_gap_axis_curriculum_target_match/scripts/run.py --stage merge-arms
.venv/bin/python -B experiments/qwen35_4b_goal_gap_axis_curriculum_target_match/scripts/run.py --stage local
.venv/bin/python -B experiments/qwen35_4b_goal_gap_axis_curriculum_target_match/scripts/run.py --stage benchmark

Results

Both arms trained cleanly (replay 0.3776, axis 0.4884 train loss; 0 skips each) and merged; the frozen 144-task gate event PROMOTED axis_curriculum — the first promotion in this program's universal line: axis holdout 28/40 versus parent 22 and replay 18 (kind wins on hygiene 9-vs-5/5, explore 7-vs-6/3, tracefix 4-vs-3/2; protocol tied at the preregistered control ceiling), with retention byte-equal to the parent (71/95/9) while replay drifted (65/89/15).

The conditional aggregate pilot at seed 78,144 (quick, think budget 1,024) then ran all four composites: base 0.1085, axis_curriculum 0.4223, parent 0.4644, replay_repeat 0.5081. The candidate beat base by +0.3138 with SEVEN strictly positive families, three ties, and zero negatives — including flipping warren (0 → 0.125), one of the historically stuck families — but lost the aggregate to both its parent (−0.042) and the replay control (−0.086), so the pilot gate fails. Notably the replay control flipped rites (0 → 0.125) and posted the highest aggregate this line has recorded at any seed, and menders (0 for every arm) and sirens (exactly 0.500 for every arm) did not move for anyone.

Interpretation

Three separable findings. (1) The axis-atom mechanism installs: first-ever local promotion, +6/+10 on held-out axis tasks with zero retention cost. (2) Local axis wins under-convert at quick tier: hygiene's near-doubling locally left sirens pinned at 0.500, and tracefix's win left menders at zero — though warren did flip, consistent with the explore lesson (or 1/8-granularity noise; single seed). (3) Replay continuation compounds again: one more replay round beat everything on aggregate while flipping rites, replicating the line's replay-is-active-intervention law. Per the frozen drift-versus-content reading, the candidate's aggregate loss to replay is a content-opportunity-cost result at matched exposure, not retention damage (retention was byte-equal to the parent). The goal's wall is now precisely two families wide: menders and sirens are frozen for every arm at every seed at this tier configuration.

Terminal Disposition

No later event is authorized here. Seed 78,144 is consumed and recorded; do not re-run any stage or benchmark the composites from this directory. The published replay_repeat composite (aggregate 0.5081, receipts under runs/) is the strongest known artifact and the presumptive parent for successors; axis_curriculum remains the only artifact with a positive axis-installability reading.

Knowledgebase Update

  • Program evidence updated: local promotion, pilot negative, per-family record, and the replay-compounding observation recorded.
  • Program backlog updated: the menders/sirens wall and the replay-compounding line are queued with calibration notes.
  • Claim ledger updated: no; single-seed pilot events do not mint claims.

Artifacts

  • data/sft_axis160.jsonl, data/corpus_manifest.json: frozen axis corpus.
  • data/sft_blend.jsonl: frozen replay pool (byte-identical to predecessors).
  • data/stream_manifest.json, data/stream_token_receipt.json: exposure receipts.
  • data/local_tasks_seed88014.jsonl, data/local_input_seed88014.jsonl, data/local_design_receipt.json: frozen two-instrument gate.
  • reports/preregistration.md, reports/design_review.md: prospective contract and adversarial authorization.
  • reports/artifact_manifest.yaml: external parent and conditional model-artifact plan.

Report

Rendered from reports/report.md

Summary

Model-free construction is under way: a 160-row designed axis corpus targeting the benchmark's four empirically stuck families (built only from public axis descriptions on fresh vocabulary), a three-axis exact-exposure replay control from the surface-general designed_fresh parent, and a two-instrument local gate (40-task unseen-seed axis holdout plus 104-task retention screen) whose promotion conditions are prospectively achievable — relative wins and non-inferiority bands, no absolute per-kind floors. No model event has run; aggregate seed 78,144 is sealed.

Research Program Fit

Goal-gap forensics over 65 quick-tier and 92 medium-tier historical paired events show the all-families goal passing exactly once, with failures concentrated in four families whose gym analogues near-mastered internally without transferring. This trial attacks those axes directly with designed atoms whose formalisms vary within each axis — the surface-diversity lever the fresh-surface trial just validated for the generic dose.

Method

  • u_tracefix: program repair across four invented executable formalisms; the corrected instruction is provably the unique single-step repair.
  • u_explore: unique-shortest route under an exact move budget with compact frontier-notation thinking.
  • u_hygiene: records embedding adversarial directives whose decoy value is the only parseable wrong answer — injection compliance is scored by construction.
  • u_protocol: documented flags plus tally with all three closing branches trained in balance.
  • Control: exact-exposure replay continuation (MILP; forward tokens, loss-bearing targets, loss mass all delta-zero).
  • Gate: candidate must strictly beat parent AND replay on the 40-task axis holdout total and on at least 3 of 4 axis kinds, while staying within retention non-inferiority bands on the 104-task original-skill screen. Only the tracefix kind varies formalism within its axis, so its win reads as structure-binding; the other three kinds' wins read as installation on the axis surface.
  • Conditional aggregate at sealed seed 78,144 (quick, tb 1,024): candidate strictly above base, replay, and parent on aggregate; the every-family-versus-base record is reported under the frozen quick-tier power statement.

Results

  • Training (1,520 rows, 0 skips, 190 updates each): replay 0.3776, axis 0.4884 train loss.
  • Local gate (seed 88,014): axis holdout of 40 — candidate 28, parent 22, replay 18; per-kind candidate/parent/replay: explore 7/6/3, hygiene 9/5/5, protocol 8/8/8 (control-ceiling tie), tracefix 4/3/2. Retention of 104: candidate 71/95/9 (correct/parsed/caps) = parent exactly; replay 65/89/15. All ten checks passed; PROMOTED.
  • Aggregate pilot (seed 78,144, quick, tb 1,024): base 0.1085, axis_curriculum 0.4223, parent 0.4644, replay_repeat 0.5081. Candidate vs base +0.3138: 7 families strictly positive, 3 ties (menders 0=0, rites 0=0, sirens 0.5=0.5), 0 negative; warren flipped 0→0.125. Replay vs base: 7 positive with rites flipped and ties at menders/sirens/warren. Pilot gate failed on both aggregate comparisons.
  • Corpus hash e7a95d73...686e; balance: tracefix formalisms 13/13/9/5, protocol branches 15/14/11, hygiene 30 injected / 10 clean.

Controls

The parent composite is the baseline; the exactly-matched replay continuation is the mechanism-falsifying control; the axis holdout at an unseen seed separates installation from memorization; the retention screen guards the thirteen retained skills.

Oracle Versus Deployable Evidence

Executable truth grades outputs only and is stripped from every model-facing byte. benchmarks/ remains unread.

Next Stage

None. Closed per the frozen contract after the pilot negative; seed 78,144 is consumed and recorded. The published replay_repeat composite is the presumptive successor parent; the queued next steps (replay-compounding measurement; menders/sirens instrument forensics) carry calibration notes in the program backlog.

Artifact Manifest

Model-free artifacts tracked in-repo; parent artifacts external with pinned checksums; future adapters and composites under large_artifacts/ with tracked receipts.

Experiment log 6

Show the running log (6 entries, 2026-07-14 → 15)

2026-07-14 — Model-free design freeze

  • Opened after the fresh-surface budget-commit trial closed green with a positive surface-generality reading and a structurally unpassable induct floor; this trial claims the program's queued goal-gap successor slot.
  • Goal-gap forensics over 65 quick-tier and 92 medium-tier historical paired events located the benchmark's reproducible blockers in four public families whose gym analogues near-mastered internally without transferring; the four axis-lesson kinds here target exactly those failure modes with designed single-turn atoms on fresh vocabulary (construction seed 77,117).
  • Froze data/sft_axis160.jsonl (160 rows, sha e7a95d73...686e): unique-repair program fixing across four invented formalisms, unique-shortest budgeted route search with frontier-notation thinking, injection-hygiene records with parseable-wrong decoys (30 injected / 10 clean), and branch-balanced documented protocols (15/14/11 across the three closing branches).
  • Reserved fresh construction/slot-match/training/gate/aggregate seeds 77117/55118/52/88014/78144; aggregate sealed.
  • Preregistered the full contract: three-axis exact-exposure replay control from the designed_fresh parent, a two-instrument gate (40-task axis holdout

    • 104-task retention screen) with relative wins and non-inferiority bands and

    NO absolute per-kind floors, and a four-model conditional aggregate pilot at seed 78,144 with the frozen quick-tier power statement.

  • No model, GPU, training, local, or benchmark event has run.

2026-07-14 — Authenticated control training

  • train-control ran only after design-freeze commit e1064249 matched origin/main with both workflows green and a clean worktree.
  • replay_repeat trained 1,520/1,520 rows with 0 skipped over 190 updates (train loss 0.3776, 1,320 wrapper seconds); receipt/log published and pinned fail-closed. The candidate arm remains untrained until this checkpoint publishes green.

2026-07-14 — Authenticated candidate training

  • train-candidate ran only after control checkpoint 00ddd91e matched origin/main with both workflows green and a clean worktree.
  • axis_curriculum trained 1,520/1,520 rows with 0 skipped over 190 updates (train loss 0.4884, 1,366 wrapper seconds); receipt/log published and pinned fail-closed. Merges are the only next stage.

2026-07-14 — Authenticated explicit composites

  • merge-arms ran only after the candidate checkpoint aedc1770 matched origin/main with both workflows green; PASS_CONTROL_MERGE and the merge self-pin were required and verified.
  • Both arms merged through the pinned external merger (scale 2.0, 128/128 nonzero modules, fingerprint-verified); receipts and logs under runs/merges/; both merged-tree pins now filled fail-closed in the local evaluator. The one frozen 144-task gate event is the only next stage.

2026-07-14 — Local gate event: axis_curriculum PROMOTED

  • The one frozen 144-task gate event ran from merge checkpoint 68916e45 (clean synced green main): three sequential authenticated engine runs.
  • Axis holdout (of 40): axis_curriculum 28, parent 22, replay 18. Per-kind: explore 7 vs 6/3 (win), hygiene 9 vs 5/5 (win), protocol 8 vs 8/8 (tie — recorded as the preregistered control-ceiling case), tracefix 4 vs 3/2 (win) — 3 of 4 kind wins.
  • Retention (of 104): axis_curriculum 71 correct / 95 parsed / 9 caps — byte-equal to the parent's totals; replay drifted to 65/89/15.
  • All ten preregistered checks pass; axis_curriculum is promoted. The conditional aggregate at sealed seed 78,144 is the only next stage.

2026-07-15 — Conditional aggregate pilot: negative; experiment closed

  • The benchmark stage ran from promotion checkpoint 4adf6b7a (clean synced green main): weight-recomputation binding for all four composites, then four gateway events on never-touched seed 78,144 (quick, think budget 1,024, identical source inventory across arms; seed recorded in the ledger).
  • Aggregates: base 0.1085; axis_curriculum 0.4223; designed_fresh_parent 0.4644; replay_repeat 0.5081.
  • Candidate versus base: +0.3138 aggregate; per family 7 strictly positive, 3 ties (menders 0=0, rites 0=0, sirens 0.5=0.5), 0 negative; warren flipped (0 → 0.125). Replay versus base: 7 positive, ties menders/sirens/warren; rites flipped (0 → 0.125).
  • Pilot gate FAILS: the candidate lost the aggregate to parent (−0.042) and replay (−0.086). Per the frozen contract the experiment closes; per the frozen readings, the loss is content-opportunity-cost at matched exposure (retention was byte-equal to parent locally), and the every-family-vs-base record (7/3/0) is reported under the quick-tier power statement.
  • menders scored 0 and sirens exactly 0.500 for every arm — the goal's wall is now two families wide and precisely localized.

Data files 1

Result tables and metrics copied from the experiment folder — preview inline or open the raw file.

Reproduce

Smoke test

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_goal_gap_axis_curriculum_target_match/scripts/run.py --smoke

Full run

checkpointed scripts/run.py stages only

Run steps are documented inside the experiment folder (README and scripts).

Browse the experiment folder on GitHub ↗