Research log Small Model Experimentation
GitHub

Qwen3.5-4B Tool-Seeded Banking: does tool-search + banking cross the depth-3 wall?

Search finds the harder skill

The one idea you need

This small model writes little programs by chaining list operations. It normally improves by restudying solutions it found itself — but it almost never discovers three-operation chains, so none enter its study pile. Here a plain brute-force search supplies those missing three-chain answers for it to copy.

The question

If a small model never discovers a three-step program on its own, can a mechanical search find those solutions and retraining teach it to produce them?

What we found

Partly. A brute-force search found working three-step programs the model never produces itself, and retraining on them lifted three-step success from a hard zero to 5 of 40 fresh, never-seen tasks when it can reason across sixteen tries — a real, significant crossing. But in one shot, the production setting, it still solved none, versus 15% on two-step tasks. The knowledge crossed the wall, yet barely stuck in the weights.

Why it matters

To push a small model one composition-step deeper, use an external search to supply examples it cannot generate itself — but expect the new skill only with test-time reasoning, not one-shot production calls. Seed each depth separately; deeper installs weaker.

Three-step success, reasoning across sixteen tries0 of 40 → 5 of 40novel unseen tasks unlocked, up from a hard zero
Three-step success, single best shotstill 0%the skill never installed for one-shot deployment
Two-step success, reasoning across sixteen tries23% → 42%shallower chains install far more strongly
Hard programs the search found130 of 130three-step tasks brute-forced on CPU, no model involved
On this page
  1. Results at a glance
  2. Overview
  3. Report
    1. Summary
    2. Research Program Fit
    3. Method
    4. Pre-registered verdicts
    5. Interpretation
    6. Honesty notes / limits
    7. Next Experiments
    8. Artifact Manifest
  4. Experiment log
  5. Figures
  6. Data files
  7. Reproduce
  8. Related

Results at a glance 2

Where the tool-seeded training helps and where it doesn't

How to read

Bars grouped by program depth — two, three, and four chained operations. Within each group: the untrained model's success across sixteen tries, the retrained model's success across sixteen tries, and the retrained model's single-shot production success. Taller is better.

0%20%40%60%depth 2depth 222.5%42.5%15%depth 3 (the wall)depth 3 (the wall)0%12.5%0%depth 4depth 40%0%0%

Takeaway → At three operations the retrained multi-try bar jumps off zero while the untrained bar stays flat — but its single-shot bar is still zero, and four operations never moves. Knowledge crossed, not deployable.

Data table
conditionbase think cov@16banked_tool think cov@16banked_tool no-think greedy@1 (deployable)
depth 222.5%42.5%15%
depth 3 (the wall)0%12.5%0%
depth 40%0%0%

Numbers from

Technical framing

Tool-seeded banking crosses depth-3 (0->0.125) where self-banking gave 0 -- but weakly — C21: self-banking depth-1+2 gave depth-3 coverage exactly 0.00 (base samples ~0 depth-3 to bank). Here we harvest depth-3 via an interpreter-backed explorer (CPU, no external model, 130/130 solved), add to the SAME depth-1+2 pairs, and bank. Depth-3 think coverage crosses 0.00->0.125 (5/40 distinct novel held-out tasks, significant vs the 0/40 floor) -- validating the recipe: tools explore, banking installs. BUT the install is weak and test-time-dominated: depth-3 no-think greedy@1 stays 0.00, vs depth-2 which installs deployably (greedy@1 0.15). The installer's efficacy decays with depth (echoing C19: depth-3 representation is a thread).

Three-step success climbs with more tries but stays low

How to read

Share of three-step tasks solved (vertical) as allowed tries rise from one to sixteen (horizontal). Flat bottom line: the untrained model, pinned at zero. Middle line: the retrained model. Top line: that same retrained model on easier two-step tasks. Higher is better.

0%20%40%60%51015depth 2 banked_tool (contrast)depth 2 banked_to…depth 3 banked_tooldepth 3 banked_to…depth 3 base

Takeaway → The retrained three-step line climbs steadily off zero to 5 of 40 while the untrained line stays flat at zero — real learning, but far below the two-step line's 42%.

Data table
k (samples)depth 3 basedepth 3 banked_tooldepth 2 banked_tool (contrast)
10%2.3%15.3%
20%4.3%21.9%
30%5.8%26%
40%7.1%28.9%
50%8.2%31%
60%9%32.7%
70%9.7%34.1%
80%10.3%35.3%
90%10.8%36.4%
100%11.1%37.4%
110%11.5%38.3%
120%11.7%39.2%
130%12%40.1%
140%12.2%40.9%
150%12.3%41.7%
160%12.5%42.5%

Numbers from

Technical framing

Depth-3 coverage@k (think): banked_tool lifts off zero; base stays pinned — The banked model's depth-3 think-coverage climbs to 0.125 with more samples while base is flat at 0.00. Depth-2 (contrast) climbs to 0.42. The crossing is real but modest, and it is test-time think-search coverage; deployable no-think single-shot barely moves at depth 3.

In the author’s words from the Overview · “Results”

CROSSED-BUT-WEAK. Depth-3 think coverage 0.00 (0/40) -> 0.125 (5/40 distinct novel tasks) — significant vs the 0/40 floor where C21 gave exactly 0 (tools explore, banking installs). But weak/test-time-dominated: no-think depth-3 0.025, greedy@1 0.00, vs depth-2 deployable greedy@1 0.15. Depth-4 stayed 0. See reports/report.md, analysis/tool_seeded_banking.png.

Overview

Research Program

  • Program: structured_execution_and_compilers
  • Program question: the C21 positive control — does seeding banking with tool-found depth-3 solutions cross the depth-3 wall self-banking couldn't?
  • Prior anchors: C21 (self-banking can't climb), C12 (tool-search cracks depth-3), C19 (depth-3 representation is a thread).

Question

C21: self-banking gives depth-3 coverage 0.00 (base samples ~0 depth-3 to bank). Does harvesting depth-3 via an interpreter-backed explorer, then banking, install depth-3?

Hypothesis

Pre-registered (reports/prereg.md, hardened by reports/design_review.md): P1 explorer solves >=90% depth-3; P2 banked_tool depth-3 think cov >= 0.15 & >= base+0.10 & >=5 distinct; P3 generalization; P4 deployable + next rung.

Setup

  • Model: Qwen3.5-4B (only permitted model). No external model — explorer is interpreter brute-search over families' own 16-op DSL (CPU).
  • Harvest depth-3 (130/130 solved) + C21's exact depth-1+2 pairs = 260 pairs {1:47,2:83,3:130}. Bank -> banked_tool.
  • Eval: frozen PAIRED held-out set (behavioral function-signature dedup), n=40/depth 2/3/4; think (primary) + no-think (deployable).

Run

Smoke: python scripts/tool_harvest.py --smoke Full: python scripts/tool_harvest.py --n-depth3 130 && bash runs/launch.sh && python scripts/analyze.py

Results

CROSSED-BUT-WEAK. Depth-3 think coverage 0.00 (0/40) -> 0.125 (5/40 distinct novel tasks) — significant vs the 0/40 floor where C21 gave exactly 0 (tools explore, banking installs). But weak/test-time-dominated: no-think depth-3 0.025, greedy@1 0.00, vs depth-2 deployable greedy@1 0.15. Depth-4 stayed 0. See reports/report.md, analysis/tool_seeded_banking.png.

Interpretation

Validates the C21 recipe (explorer + installer, both required) but reveals the installer's efficacy decays with depth (echoes C19/C20). Each rung must be seeded; extendable with diminishing efficiency.

Knowledgebase Update

  • Program evidence updated: research_programs/structured_execution_and_compilers/evidence.md (C22)
  • Claim ledger updated: C22 added

Artifacts

  • scripts/tool_harvest.py (CPU explorer + combine), scripts/train_lora.py, scripts/eval_ladder.py (frozen paired + behavioral dedup), scripts/analyze.py, scripts/common.py
  • data/{train,tool_depth3,eval_frozen,train_tasks}.jsonl, runs/eval_{base,banked}_{think,nt}.json, runs/verdict.json
  • analysis/tool_seeded_banking.png, reports/{prereg,report,design_review}.md
  • runs/banked_tool_adapter/ — adapter (~180MB, moved out of repo; regenerable)

Report

Rendered from reports/report.md

Summary

C21 showed self-banking is coverage-seed-bounded: banking depth-1+2 self-solutions unlocked ZERO depth-3 coverage (base 0.00 → banked 0.00), because the base samples ≈0 depth-3 solutions to bank. The predicted fix: seed banking with depth-3 solutions found by an explorer the base lacks. This tests it — harvest depth-3 via interpreter-backed search (a tool, no external model), bank, and measure held-out depth-3 coverage.

Answer: CROSSED-BUT-WEAK. Tool-seeded banking does cross the wall self-banking couldn't — but installing a depth-3 composition is far harder than a depth-2 one, and the crossing is modest and mostly test-time.

depth 2depth 3 (the wall)depth 4
think cov@16 base → banked0.23 → 0.420.00 → 0.125 (5/40 unlocked)0.00 → 0.00
no-think cov@16 base → banked0.03 → 0.170.00 → 0.025 (1/40)0.00 → 0.00
no-think greedy@1 base → banked0.03 → 0.150.00 → 0.000.00 → 0.00
  • The recipe is validated (the explorer was the missing ingredient). Depth-3 think coverage rose from a hard 0/40 (0.00 — identical to C21's self-banking) to 5/40 (0.125) on a frozen, behaviorally-deduped held-out set. That clears base's 95% upper CI (~0.075), unlocks ≥5 distinct novel depth-3 rules, and is highly significant vs the 0/40 floor (p<0.01). The ONLY change from C21's banked1 was adding tool-found depth-3 pairs — so tools explore, banking installs, exactly as C21 predicted.
  • But installing gets harder with depth. The same banking recipe that installs depth-2 strongly (think 0.42, no-think 0.17, deployable greedy@1 0.15) installs depth-3 weakly (think 0.125, no-think 0.025, greedy@1 0.00). The depth-3 gain is almost entirely test-time think-search; it barely banks into single-shot weights. This echoes C19 (the depth-3 representation is a thread): even with perfect training data, the model absorbs deep compositions far less than shallow ones.
  • No free next rung. Depth-4 stayed 0.00 → 0.00 — banking depth-3 does not leap to depth-4 (P4 held); each rung must be seeded, consistent with the rung-by-rung recipe.

Research Program Fit

The positive control the C13C21 arc predicts. Closes the loop: C21 (self-banking can't climb) → C12 (tools crack depth-3) → this (tools + banking cross depth-3, weakly). Design hardened by an adversarial multi-agent review (reports/design_review.md).

Method

Substrate list, families.LIST_PRIMS (the SAME 16-op DSL as eval/C21 — search written on families.py, not C12's 23-op decompose_lib, so no vocabulary mismatch). No external model anywhere.

  • Explorer (CPU-only, no model): for each depth-3 TRAIN task, brute-enumerate the interpreter over the substrate's own primitives (BFS, global state-dedup, max_depth 3) to DISCOVER an op-sequence correct on all task examples; render via families.reference_code. Solved 130/130 depth-3 tasks (mean found-depth 3.00, all sandbox-verified) — what monolithic sampling gets ≈0 of. (C12 found guided≈brute at depth-3, so the crack is composition-structure + interpreter, not the model's planning; framed as "interpreter-backed search," not "the model's planning.")
  • Seed set: C21's EXACT 130 depth-1+2 self-harvest pairs + the 130 tool-found depth-3 pairs = 260 pairs {1:47, 2:83, 3:130}. The ONLY delta from C21's banked1 is the depth-3 seeds; LoRA recipe, prompt (8-visible ident_prompt), everything else held identical.
  • Bank: QLoRA-SFT single-shot prompt→code → banked_tool.
  • Eval: frozen PAIRED held-out set (generated once, reused by every arm), n=40/depth at depths 2/3/4, with behavioral function-signature dedup (a task excluded if its function on a fixed probe set matches ANY training rule — catches alternate decompositions). Primary: think coverage@16 (matched to C21). Secondary: no-think greedy@1 + coverage@16 (deployable). base vs banked_tool.

Pre-registered verdicts

  • P1 (explorer works): HELD strongly — interpreter search solved 130/130 depth-3 tasks; built 130 pairs.
  • P2 (unlock): CROSSED-BUT-WEAK. banked_tool depth-3 think cov = 0.125 (5/40 distinct), ≥ base + 0.10 (✓) and ≥ 5 distinct (✓) and above base's 95% upper CI (✓, significant), but just under the pre-registered strong 0.15 point-estimate (0.125). A real, significant unlock of modest magnitude.
  • P3 (generalization): HELD — the 5 unlocked tasks are on the frozen behaviorally-deduped held-out set (novel depth-3 rules disjoint from training).
  • P4 (deployable + next rung): the depth-3 gain does NOT survive into deployable no-think single-shot (greedy@1 0.00, no-think cov 0.025) — it is test-time-dominated. Depth-4 did not rise (no free next rung).

Interpretation

  • The C21 recipe is confirmed in direction: tools are the explorer, banking is the installer, and both are required. The identical banking that gave exactly 0.00 at depth-3 from self-samples (C21) gives a significant 0.125 (5 novel tasks) once seeded with tool-found depth-3 solutions. So the wall is crossable by self-training — provided an external search seeds each rung the base cannot reach.
  • New nuance the arc did not have: the installer's efficacy decays with depth. Banking installs depth-2 robustly and deployably (greedy@1 0.15) but depth-3 only weakly and mostly at test-time (greedy@1 0.00). Given C19 (the depth-3 inverse is barely represented) and C20 (it is not steerable), this suggests the deep wall resists installation too: even perfect training data lands only a thread of depth-3 capability in one QLoRA round. Extending the frontier a full rung likely needs more depth-3 data / more training, or is capped by a representational bottleneck.
  • The full recipe, precisely: to extend the frontier one depth — (1) an explorer the base lacks (tool-search / enumeration) reaches the next rung; (2) banking installs it, but weakly, and mostly as test-time-accessible rather than single-shot; (3) each rung must be seeded (no free leap). The frontier is extendable, but with diminishing installation efficiency the deeper you go.

Honesty notes / limits

  • The crossing landed at 0.125 — just under the pre-registered 0.15 strong bar, though significant vs the base 0/40 floor and clearing the ≥5-distinct and ≥+0.10 criteria. Reported as CROSSED-BUT-WEAK, not a clean pass.
  • Single QLoRA round, 130 depth-3 pairs. A dose–response (more pairs / epochs) and whether depth-3 ever banks into deployable single-shot are untested — the weak deployable install may be data/training-limited or a representational cap (C19).
  • The depth-2 coverage also rose (0.23→0.42), partly from the larger combined training set; the depth-3 isolation is clean (base 0.00 floor, only depth-3 seeds added).

Next Experiments

  • Dose–response: vary depth-3 tool-pairs (40/130/400) — does depth-3 install strengthen toward deployable greedy@1, or plateau (representational cap)?
  • Iterate the rung: after banking depth-3, does tool-search on the banked model harvest depth-4 more cheaply (does the installed depth-3 make depth-4 search/coverage easier)?

Artifact Manifest

See reports/artifact_manifest.yaml. Key: scripts/tool_harvest.py (CPU explorer), scripts/train_lora.py, scripts/eval_ladder.py (frozen paired set + behavioral dedup), scripts/analyze.py, data/train.jsonl, data/tool_depth3.jsonl, data/eval_frozen.jsonl, runs/eval_{base,banked}_{think,nt}.json, runs/verdict.json, analysis/tool_seeded_banking.png, reports/design_review.md. Adapter (runs/banked_tool_adapter, ~180MB) omitted from git.

Experiment log 1

Show the running log (1 entry)

Scaffold

Created as a new experiment scaffold.

Figures 1

tool seeded banking
tool seeded banking · analysis/

Data files 4

Result tables and metrics copied from the experiment folder — preview inline or open the raw file.

Reproduce

Smoke test

python scripts/tool_harvest.py --smoke

Full run

python scripts/tool_harvest.py --n-depth3 130 && bash runs/launch.sh && python scripts/analyze.py

Run steps are documented inside the experiment folder (README and scripts).

Browse the experiment folder on GitHub ↗