Agentic Breadth Installation
Install general agentic capability in Qwen3.5-4B via breadth-first expert iteration on a firewall-clean multi-family gym, arbitrated by the blackbox menagerie instrument.
What we have learned
Seed Experiments
- Experiment:
qwen35_4b_gauntlet_breadth_round1— gym built (12 families, 10 trained + 2 held out), two fast training rounds run, first menagerie-arbitrated install in the corpus.
Confirmed Claims
- Claim: C49 (Confirmed) — vLLM 0.24 runtime LoRA is a silent no-op for Qwen3.5-4B PEFT adapters; on-vs-off behavioral gate required; deploy via merged composite checkpoints or the HF backend.
- Claim: C50 (Promising) — breadth-first expert iteration with emission-seam-weighted loss installs substrate-general agentic competence: menagerie quick +0.223/+0.294 on two fresh paired seeds (HF backend, deterministic), gym +0.518 including never-trained held-out families.
Negative Findings
- Finding: full-weight SFT on the model's own naturally-closed verified chains (round 1) installs nothing measurable — near-self-distillation; the deployment-critical post-force-close state must be in-distribution and the gradient concentrated on the answer/action emission seam.
Finding: stallwright (bounded optimization) is unharvestable at round 1 — the base model never concludes its optimization deliberation even at a 4096-token think budget; the axis moved only by transfer (+0.395 gym) and its menagerie analogue (stockade) did not move.
Claim: C53 (Promising) — the second wall: nine paired events across six escalation levers land the treated model at 0.375-0.447 aggregate; the emission-policy install is a one-time step and no self-verified-output recipe variant moves the band; a single rich family installs nearly the full effect (breadth is an axis-aligned increment, not the mechanism).
- Claim: C54 (Promising) — novel serial-compute mechanisms (compression advantage + skin-shuffle, co-trained from base = the apex arm) DECISIVELY clear the +0.32 MEDIUM bar for the first time (+0.345, all events), breaking the wall gold-procedure supervision (C53) could not; but no single 4B adapter clears quick AND medium — a non-convex tier-Pareto frontier.
Current Read
Breadth + strict verifiers + emission-seam supervision is the first recipe in the corpus to move the blackbox instrument, and the locality laws (C43/C45/C48) do not extend to it. The binding deployed constraint at current difficulty is the truncation cascade (consume budget → force-close → verbose restart → no parseable answer); repairing commit-from-partial- reasoning transfers across substrates the model never trained on. Trust only paired same-backend menagerie comparisons; never a vLLM adapter arm (C49). Next: the residual is a capability core — scaffold-distillation of tool-found solutions, on-policy RL at residual failures, and failure-forensics curricula are the queued beyond-recipe mechanisms (see C53 next tests). Same-recipe scaling is closed.
qwen35_4b_counterfactual_evidence_acquisition_curriculum (2026-07-13 — Lineage-locality infeasible)
The exact transaction-replay start checkpoint was compared directly with the C54 apex anchor on the frozen 48-context block before interface behavior or training. Median centered non-target logit drift was 0.110735 against the registered 0.10 ceiling; all 48 row-level drift estimates exceeded 0.10. The entropy guard passed (+0.013636 versus a -0.05 floor), varentropy was essentially flat (+0.000297, diagnostic only), and rendered prompts matched. The formal verdict is LINEAGE_LOCALITY_INFEASIBLE.
Per preregistration, interface selection, acquisition qualification, all three training arms, behavior, transfer, retention, uncertainty analysis, and Menagerie remained sealed. This result does not test counterfactual evidence acquisition, transition-balanced replay, or capability installation. It establishes only that this exact parent-anchor pair was ineligible under the new direct 0.10 attribution contract.
Read: eligibility under a direct locality ceiling is not inherited from lineage or from a prior looser or different-context gate. A successor must use a new experiment and fresh locality contexts, begin from apex itself or prospectively select a fixed apex-compatible parent using outcome-free locality blocks, then independently re-establish complete-loop behavior and acquisition headroom before training. Do not raise the observed ceiling, swap lineage after the result, or rescue this directory.
Deep-Advantage MOPD (2026-07-15 — Terminal capability negative)
qwen35_4b_deep_advantage_mopd repeated same-prefix qualification on two new 192-state blocks from the immutable 40/60 joint soup. Deep was selected on 28/26 states and beat soup on disjoint audit branches by +0.1650/+0.1220 (pooled +0.1421, one-sided 95% lower bound +0.1230); it beat quick by +0.2000/+0.1420 (lower bound +0.1534). Every support, block-sign, and uncertainty gate passed, authorizing exact-logit locality.
The locality round needed three fixed candidate batches and found 90 deep routes, from which it froze 60 deep units, 20 soup anchors, and 60 disjoint matched non-advantage controls. The five-update 15-deep/5-soup pilot passed: centered non-target drift was 0.02760, entropy drop was 3.11%, and exact target loss improved 0.01293→0.01170. The exact measurement is one midpoint token per consumed unit. Seeds 42, 43, and 44 subsequently completed all four registered integration rounds, establishing replicated route supply and optimizer-gate completion.
All matched controls are now built. Four-round full-prefix non-advantage MOPD, wrong-teacher quick MOPD, and off-policy continuation SFT each completed every consume-once update and passed their frozen training gates; the 25%/50%/75% deep parameter soups have exhaustive model-byte receipts. Independent canonical replay passes for aggregate controls receipt 103ef4cc0b24d7c10666b6f0adfcd4dfae4720415c7fbbc76b681ab79162640b. The sealed two-block comparison is terminal negative. Primary seed 42's pooled joint deltas are −0.006845 versus deep (one-sided 95% LCB −0.012839), −0.001300 versus the soup initialization, −0.001872 versus off-policy SFT, −0.003706 versus soup75, and −0.169239 versus soup best-of-eight (LCB −0.175468). Both block means versus deep are negative, every better-source stratum cell fails, and seeds 43/44 also trail deep by −0.003450/−0.005660 pooled joint. Retention passes, while untouched brinework and spindle transfer improve +0.015625/+0.010590.
There is a small real mechanism signal: primary seed 42 beats matched non-advantage MOPD by +0.005619 pooled joint (LCB +0.000582) and wrong-teacher by +0.005312 (LCB +0.000099). Advantage routing and teacher identity therefore affect the update direction, but the operator does not cross the source frontier and loses to simple interpolation and sample-more. The terminal analyzer emitted stop_before_benchmark_cli; no benchmark was authorized or opened.
Quick also passed diagnostically on 29/18 routes in this fresh replication, after failing one soup-relative block in the predecessor. This is not license to add it to the locked deep-only treatment. It makes the later two-teacher design more worthwhile, but those estimator improvements are not sufficient. First require a direct-bf16 deployment-parity microtrial whose causal update survives merge and beats the source, interpolation, and sample-more. A later two-teacher return still requires cross-fitted direct advantage prediction, adaptive branch allocation (including zero quick allocation), and a third untouched block.
Specialist Integration Attempt (2026-07-12; stopped before training)
qwen35_4b_specialist_policy_integration attempted the accepted next mechanism: four live-state DAgger/execution-RL capability producers followed conditionally by same-origin MOPD integration. Its new compound environments, runtime, regenerated incumbent, explicit merge, and 7/7 behavioral-installation gates passed. The disjoint compound macro was 0.135 versus the <0.60 headroom bar.
The full baseline exposed a terminal design error before training: the sole tools family ferrier scored 0.994, making its frozen S0 + 0.10 target 1.094 under a score ceiling of 1.0. Because all four specialists were mandatory, the run stopped before best-of-8, DAgger, GRPO, teacher audit, integration, or benchmark exposure. This does not test MOPD efficacy. It establishes a reusable portfolio rule: endpoint headroom is insufficient; every mandatory teacher's theoretical improvement bar must be feasible on a disjoint baseline before capability-production spend.
Pareto Policy Integration Qualification (2026-07-12 — Negative prerequisite)
qwen35_4b_pareto_policy_integration corrected the earlier gate rather than lowering it. Independently regenerated and behavior-gated C54 blend and apex policies were compared on two fresh contamination-safe procedural blocks. blend - apex on quick capability was negative in both blocks (-0.00693, -0.03789), pooling to -0.02241 with a one-sided 95% lower bound of -0.04897. apex - blend on deep capability was replicated (+0.04563, lower bound +0.03401) but six retention cells regressed beyond 0.02. All protocol checks passed.
The external menagerie tier ranking therefore did not supply a clean quick/deep teacher crossover on the distribution where distillation would run. The experiment stopped before teacher audit or any MOPD update, so MOPD remains untested. The strategic correction is to estimate teacher choice on same-prefix states: a future run should freeze a disjoint verified continuation-advantage router, not assume that an aggregate instrument label is a local teacher label.
qwen35_4b_interactive_policy_curriculum (2026-07-11/12 — Negative)
The full-sequence live-state DAgger warm start failed its preregistered mechanism gate before RL or Menagerie. Against the regenerated C53 incumbent, train-family episode macro fell 0.6048→0.3517 (−0.2531; paired-bootstrap 95% CI [−0.2954, −0.2103]) and three untouched families fell 0.6850→0.3519 (−0.3331; CI [−0.3804, −0.2869]). Atom retention stayed inside its −0.03 guard (−0.0215), parsing stayed perfect, and natural closure improved by 10–13pp, ruling out generic collapse.
The failure was semantic-operator capture: only 55/2,270 training targets were VERIFY; after training, loomfix produced 600 PATCH and zero RUN actions, while the untouched analogous patchwheel produced 599 RULE and one RUN. Thus DAgger taught a fluent shared observe/revise trace but erased the scarce verify/commit pivots that make the loop effective. Live-state correctness is not enough: a shared update must preserve the incumbent operator distribution and behavior outside corrected states. Entropy/outcome variance can route state acquisition, but must not become token pressure. RL, matched controls, and Menagerie were correctly cancelled; zero benchmark seeds were consumed.
qwen35_4b_think_ftpo_round1 (2026-07-11, C52 — Negative)
The first different-mechanism recipe after C50's re-saturation: single-position preference training (FTPO) on outcome-conditioned think-block pivot points (prefix-tree divergence of n=16 verifier-scored rollouts). Preregistered mechanism gate FAILED (−0.039/−0.076 vs a +0.05 bar on held-out band tasks); the shuffled-label control degraded identically, so the harm is the training regime, not the steering signal. Guards localize the channel: no C29-style collapse (the two-tier logit tether works), no-think channel clean — the damage is think-flow convergence (natural close halves at think@2048). Read: FTPO's safety/efficacy requires the rejected token to be a CONFIDENT OUTLIER (loop initiators, lexical attractors); near-parity pivot tokens violate the precondition and the ε-margin objective's collateral dominates. Menagerie was correctly never exposed (mechanism-gate rule). Census bonus: repetition loops are ~0.1% at deployed budgets — the loop-FTPO variant belongs to 16k+ only. That result queued confident-wrong-turn filtering (failing branch's token also locally dominant); round 2 below has now tested it.
qwen35_4b_think_ftpo_round2 (2026-07-11, C52 — Low-dose null)
The registered confident-wrong-turn rescue also failed, but isolated the next bottleneck. A frozen entropy/varentropy selector retained 155 failed-argmax pivots. Conventional demotion, bounded positive-only uplift, and shuffled uplift all failed exact-logit locality (mean per-row median non-target drift 0.229/0.145/0.120 logits vs a 0.10 ceiling). Pull-up was materially safer than push-down and true labels beat shuffled labels on the gym (+6.25pp) and fresh repository agent (+13.89pp, paired-bootstrap CI touching zero), so the steering directions contain some signal. They did not elicit breadth: repository hidden- test pass was base 43/72, uplift 39/72, demote 34/72, shuffled 29/72; fresh whitebox uplift was +0.26pp at think@1024 and −3.06pp at think@2048. Every coarse C49/collapse/no-think/gym guard passed, but P1/P2/P3 did not, so menagerie remained sealed (zero seeds consumed).
Read: confident-outlier geometry is necessary but not sufficient; the active constraint is context locality of the parameter update. Entropy/varentropy can route and diagnose pivots, but higher varentropy was not safer (lowest-V quartile had the cleanest uplift drift). Do not scale this LoRA recipe. Require a lower-dose or context-gated mechanism to clear P1 before another harvest or agentic transfer run.
qwen35_4b_repo_search_compress_bank (2026-07-12 — Negative)
Executable search and replay compression produced a superficially excellent install on the six trained repository families: apex 40/48 versus compact 48/48, with every compact trajectory reduced to exactly INSPECT→PATCH→VERIFY→COMMIT. That behavior did not transfer. On four wholly held-out algorithm families, compact fell 49/72→25/72 (−33.3pp; paired 95% CI [−44.4,−22.2]) and lost to matched-compute sampling by 18.1pp. Locality also failed (0.386 centered-logit drift versus 0.15), invalid actions rose 9.3%→26.0%, and verification retention fell 1.00→0.88. Menagerie remained sealed.
The transition audit resolves why exact operator balance was insufficient. After a failed visible test, apex chose another patch on 24/26 next actions; compact did so on 0/48. It still committed after every passed test. All 18 recursive-overlay trajectories repeated the same rejected patch. Success-only minimization installed a family-specific happy path while deleting the verifier-conditioned recovery policy. Marginal operator counts do not preserve conditional transition structure. Because the necessary gate cancelled the action-only arm, plan-gradient attribution remains open; the supported negative is the complete compact plan-plus-action recipe.
qwen35_4b_verifier_conditioned_recovery_bank (2026-07-12 — Locality-gated negative)
The direct C54 successor repaired the missing intervention unit. Fifty-seven model-found repository repairs produced 399 rows/arm balanced at seven public state→action transitions, including rejected-patch→changed-patch, failed-test→diagnose/revise, and passed-test→commit. Every bank/replay/firewall gate passed. On 60 fresh trained-family recovery cases, frozen apex scored 0.483, matched happy-action training 0.817, recovery action-only 0.850, and recovery-plus-plan 0.917. The plan arm used 480 mean sampled tokens versus 2,340 for base and was selected by the frozen rule.
The headline recipe nevertheless stopped at exact-logit locality: selected drift was 0.303 versus a 0.15 ceiling and unrelated entropy fell 0.106 nats, so no transfer family or Menagerie seed was exposed. Exploratory controls localize the damage. Happy action and recovery action-only passed locality at 0.083 and 0.098; action-only unrelated entropy was flat (+0.006). Nominal 5% plan-token mass produced a 29.5% larger merge-delta norm than action-only and a step-10 pre-clip gradient of 42.1 versus 1.8. Seam entropy/varentropy explains why: every JSON action-start token was already rank 1, while imposed ordinary-state plan starts were ranks ~8,404 (inspect→patch), ~1,163 (patch→verify), ~135 (start→inspect), and 3 (pass→commit). Plan SFT drove all to rank 1 and near-zero entropy.
Read: conditional action banking contains a strong, parameter-local recovery signal; broad lexical plan imitation adds trained-family efficiency but violates locality because token-mass weighting ignores realized surprisal/gradient. The next experiment should interpolate the already-trained reason delta under a locality-first gate, with full-dose action-only as the safe anchor. Do not expose the untouched four-family transfer blocks until a scaled checkpoint passes.
qwen35_4b_recovery_reason_locality_interpolation (2026-07-12 — Local but policy-gated)
The frozen action→reason path disproved the simplest non-separability reading. All four action-anchored mixtures passed the parent's exact locality instrument: drift was 0.100/0.104/0.111/0.121 for λ=.10/.18/.24/.30, with bounded entropy and varentropy change, while full reason reproduced 0.303. The two independently trained deltas strongly cancel in weight space: summed mixed-delta norm falls from 29.17 at action to 23.47 at λ=.30. This is not a monotone dose path.
Behavior also has a sharp safe optimum. On the 60-case selection block λ=.18 scored 0.967, versus 0.483 base, 0.817 happy, 0.850 action, and 0.917 full reason. It reached 0.933 failed-test and 1.000 rejected-patch success. However, no candidate cleared the registered policy gates: λ=.18 had 0.104 invalid actions/turn versus a 0.077 ceiling, and only 0.333 immediate rejected change versus the 0.60 bar. Confirmation, transfer, and Menagerie stayed sealed.
Post-stop forensics distinguish the failures. Every one of λ=.18's 24 invalid steps closed thinking and exhausted the exact 256-answer-token cap inside a long exact-replacement JSON payload; this is a real harness payload bottleneck, not repetition or free-form slop. Conversely, all 30 rejected cases changed the patch within two turns and solved: 20 used sensible INSPECT→PATCH and 10 used PATCH→VERIFY. The immediate-only proxy rejected retained conditional recovery.
Read: a locality-safe recovery policy exists on this weight path, but the registered harness cannot deploy it cleanly. The next experiment should freeze λ=.18, enlarge tool-answer capacity for every arm under matched compute, and measure rejected→changed-patch within two turns. This result does not support a transfer or benchmark claim.
qwen35_4b_same_prefix_advantage_routing (2026-07-12 — Route-gated negative)
The clean successor to the two earlier specialist stops finally measured both same-origin policies on exact soup states. Across 384 fresh states and 9,216 teacher/student continuations, deep independently passed both contrasts in both blocks; the combined router also passed. Quick did not: selected quick beat the soup by +0.2009 in block 0 but lost by -0.0253 in block 1. The frozen rule stopped before locality, MOPD, controls, or Menagerie.
The mechanism audit rejects the tempting threshold repair. Quick block 1 had an apparent +0.319 selection margin, yet only 6/26 states remained strict quick winners on independent audit. Requiring observed margins of +0.10 or +0.25 left the soup-relative audit mean negative. Absolute policy scores were reliable (r=0.79--0.86); statewise winner conditioning was not. Routing was also atom-heavy (101/288 atoms, 10/96 episodes), so composable episode evidence is especially thin.
Read: deep is the first source policy to clear the intended local-teacher prerequisite, but the required two-teacher composition does not exist under this labeler. A new deep-only routed-MOPD experiment is the shortest test of the update kernel. Reintroducing quick requires cross-fitted direct advantage prediction and a third untouched block; otherwise retire it rather than tune a margin. This result is not evidence against MOPD.
qwen35_4b_recovery_payload_budget_harness (2026-07-12 — Confirm-gated negative)
The matched interface repair validated the predecessor's post-stop diagnosis. The fixed locality-safe λ=.18 checkpoint passed a third disjoint locality block (0.114 centered-logit drift; entropy Δ −0.0059), while a 512-token answer slot reduced candidate cap hits to 0.5%/7.8%/7.9% of turns across calibration/dev/confirm. Valid rejected-patch and failed-test change-within-two reached 100% on both transfer blocks. The old invalid-action and immediate-transition stops were therefore harness/proxy failures.
The candidate passed every development gate at 0.7125 recovery, +0.125 versus base and +0.2125 versus equal-reservation sample-more, with exact normal-task retention. Independent confirmation stopped the claim: candidate and action-only both scored 0.6875, missing the frozen candidate ≥ action +0.03 bar. All other confirm checks passed, including +0.0875 versus base, +0.225 versus sample-more, nonnegative deltas on all four held-out families, perfect two-turn recovery, locality, and exact normal retention. Menagerie remained sealed.
Paired forensics locate a new deployable opportunity. Candidate and action-only had a 0.7875 hidden-success union on both dev and confirm, with 10/6 and 8/8 exclusive wins. Confirmation action-only wins clustered in pattern_router rejected-patch states while candidate wins clustered in rate_buckets; neither global policy dominates. Hidden union is oracle-only. The warranted successor is bounded public-verifier branching between the two local policies, compared to equal-compute sampling, followed only conditionally by transition-balanced winner banking. Do not tune another scalar reason dose or route on family identity.
qwen35_4b_recovery_verifier_branch_tournament (2026-07-12 — Prospective infeasibility)
The predecessor's replicated 0.7875 action/reason union did not transfer to four new procedural repository families. On 80 prospective-dev recovery cases, C54 base scored 0.6125, λ=.18 and action-only each scored 0.7375, and their deterministic hidden union reached only 0.7500. Equal-reservation pass-if-either sample-more scored 0.7375 for λ=.18 and 0.7500 for action. The union therefore failed every frozen +0.03 feasibility contrast before the public selector was scored. Confirmation, winner banking, and Menagerie stayed sealed.
The paired failure anatomy is more useful than the aggregate null: 58 cases were solved by both sources, one by candidate only, one by action only, and 20 by neither. Every shared failure was the new atomic_reservations family. Both source policies retained 1.00 changed-patch-within-two on both controlled states, so this is not another recovery-loop deletion. Traces repeatedly fixed whole-request validation or input immutability separately and then regressed the other; action sample-more assembled the full conjunction once in 20 cases.
Read: public selection cannot manufacture proposal coverage from globally local policies with the same semantic core. Retire branch arbitration for this recovery line. The next intervention should install the missing transactional validate-copy-commit invariant from executable tool-found solutions across diverse training families, mix the existing conditional recovery bank as replay, and require transfer to structurally different transactional families plus broad recovery retention before Menagerie.
qwen35_4b_transaction_invariant_recovery_curriculum (2026-07-13 — Transfer-dev negative)
The fixed action-seam curriculum passed exact locality against C54 apex (0.119 centered non-target drift; entropy +0.011, varentropy −0.0002) and strongly installed six trained transaction families: 0.817 versus recovery parent 0.517 and matched replay-only 0.383. Both changed-patch-within-two transitions were 1.00; invalid actions and answer-cap contacts improved. Thus executable programmatic supervision can locally move semantic coding proposals without deleting the loop or broad neighboring logits.
It did not meet unseen transfer. On four transaction families, primary was 0.719, parent 0.703, replay-only 0.641, and equal-reservation parent sample-more 0.703. Candidate-parent paired CI was [−0.031,+0.078], below the registered +0.10 bar. All family and interface guards passed, but atomic reservations—the only headroom family—remained 0/16 for every arm. Confirmation, broad retention, and Menagerie stayed sealed.
The proposal audit advances the mechanism despite the score null. Every one of the candidate's 16 first atomic patches contained copied state, all-resource validation, and atomic per-request commit; none contained the separate negative amount exception. After the visible test reported the omission, every trajectory overcorrected by raising on all unavailable/insufficient requests, destroying required False decisions. The missing unit is now verifier-faithful validation-policy discrimination, not transaction structure or recovery syntax. Next: near-correct counterexample states that isolate one policy distinction at a time, with complete recovery replay and matched extra- transaction controls.
qwen35_4b_validation_policy_counterexample_curriculum (2026-07-13 — Calibration infeasible)
The one-transition residual update itself was local: candidate-to-C54 drift was 0.109, entropy changed +0.021, and varentropy −0.011. Candidate and matched extra-transaction control each received 36 steps, 336 rows, zero think loss, and identical transition/operator action mass from the same learned transaction parent. The candidate changed only 24 post-diagnosis revision rows.
Controls-first calibration made the causal comparison impossible. The parent and matched control each solved all 48 fresh train-skin recovery cases, with perfect failed-test and rejected-patch changed-within-two and zero invalid actions. All 48 parent first changed patches already included negative handling, copied state, and ordinary False rejection. The theoretical candidate ceiling could not clear +15/+10, so candidate scientific behavior, transfer, retention, and Menagerie stayed sealed.
Read: making the rule explicit and the partial state otherwise correct removed the predecessor's residual. The original atomic miss was conditional on its more implicit contract and proposal dynamics, not evidence of a generic inability to write the distinction. Add a substrate-headroom gate before capability production: qualify multiple semantic conflicts under the exact prompt/verifier distribution, require replicated non-saturation, and only then bank/train on disjoint skins. Do not lower the current bars or expose the trained candidate post-stop.
qwen35_4b_semantic_policy_headroom_tournament (2026-07-13 — Instrument failure)
The no-training qualification ran the exact learned transaction parent on two content-disjoint 72-case blocks. Its formal verdict is INSTRUMENT_FAIL: answer-cap contacts were 43/356 turns (0.1208) and 46/363 (0.1267), above the frozen 0.05 ceiling. Explicit controls passed at 9/9 and 8/9; invalid actions were only 0.0169/0.0138. Most cap contacts contained a valid tool call followed by run-on, but all 12 endpoint failures contacted the cap, so the registered stop cannot be waived. No training or Menagerie event ran.
No semantic axis qualified even descriptively under the frozen rule. Negative and non-integer failed-test recovery were 9/9 across all three representations in both blocks. Blank-resource recovery was 8/9 and 7/9, but only record was inside the 0.15–0.80 band in A and only tuple in B; replicated two-shape support was absent.
The trajectory-state contrast changes strategy. Every one of 72 failed-test cases reached a fully correct patch, and the four terminal misses were later regressions. Without test output, inferred-contract rejected cases produced 0/54 fully correct first patches; none of all 72 rejected trajectories read the visible tests before first patching, although 64/72 eventually reached correct workspaces after later evidence. Read: the missing unit is active specification acquisition before proposal, not post-failure semantic revision. The warranted successor should counterfactually pair nearly identical issue/source states with flipped public evidence, teach evidence inspection then evidence-faithful first patches, replay the full conditional loop, and beat matched replay and sample-more on held-out evidence channels before Menagerie.
qwen35_4b_universal_curriculum (2026-07-13 — First pilot specialization negative)
The inherited designed-curriculum scaffold was not admissible: 16/600 induction traces contradicted their answers, at least 33/600 nominal two-step rules collapsed to one primitive, its smoke command was dead, its shell failed open, and the advertised run had never started. A replacement deterministic 13-skill, six-surface generator now enforces executable truth, induction query-identifiability, genuine dead ends, behavioral depth, byte determinism, and zero tokenizer skips.
The first frozen arm continued C53 blend for one epoch on 800 designed rows. It installed locally: fresh synthetic accuracy rose 0.500->0.692, parse rate 0.615->0.962, and cap contacts fell 10/26->1/26. Firewall-clean paired quick@1024 (seed 78131, merged qwen_vllm) rejected universality. Candidate was 0.3073 versus base 0.1667 (+0.1406) and blend 0.4458 (-0.1385); chronicle and siftstack each gained +0.75 over base, but rites, stockade, and warren each fell -0.125. Six families were positive, one zero, three negative.
The from-base factorial arm then co-trained all 800 designed and 2,240 replay rows. Its effective-batch-8 recovery completed 3,040/3,040 rows with zero skips and loss 1.366. On prospectively frozen local seed 88002 it reached 0.692 accuracy and recovered routing to 2/2, but parse rate was 0.846 (<0.90) and cap contacts were 4/26 (>2); induction and execution were each 0/2. It failed two local gates, so benchmark seed 78132 remained sealed.
Read: human-designed executable supervision has real held-out signal, but a designed-only continuation specializes and displaces the mature broad policy, while from-base replay co-training does not reliably install concise execution. Local accuracy is not a retention or emission proxy. The result-separated qwen35_4b_universal_replay_anchor successor now tests the more direct integration repair: low-rate mature warm start with replay in every optimizer window versus a matched replay-only refresh. No shared claim changes from this negative parent factorial.
qwen35_4b_universal_replay_anchor (2026-07-13 — Designed arm negative; replay anchor advances)
Both matched 1,520-row continuations started from C53 blend, used 190 effective-batch-8 steps at 1e-5, and were explicitly merged. The candidate substituted 400 truth-audited designed rows for 400 replay rows; the mechanism control used replay throughout and received 17.3% more forward-token compute. The candidate passed its frozen synthetic gate at 0.731 accuracy, 0.962 parse, one cap contact, and zero feasible-route abstentions.
Firewall-clean paired quick@1024 seed 78133 rejected the designed candidate. warm_union scored 0.4238: +0.2488 versus base, -0.0172 versus blend, and -0.0613 versus replay refresh. It regressed rites -0.125 below base and strictly improved only five families. Every universality rule except positive aggregate failed.
The replay-only control is the strategic result. replay_refresh scored 0.4851, +0.3101 versus base and +0.0441 versus blend. All ten family deltas were nonnegative and eight were positive; rites and sirens tied base. This is not a universal-feature claim, because strict all-family lift and replication are absent. It does show that C53's broad replay policy was not saturated and that replay is an active capability intervention, not a neutral retention control.
Read: the arm with a 26% designed substitution still failed broad retention even under a low-rate replay-anchored warm start. Because replay refresh had 17.3% more forward-token exposure despite matched steps, the failure rejects this arm but does not isolate designed content as the cause of the whole gap. The next result-separated test should start from the authenticated replay-refreshed policy, reduce designed density by an order of magnitude, and match both optimizer steps and forward-token exposure against replay continuation. Require retention of all eight observed gains and strict lift of every family; do not tune on seed 78133 or reuse this directory.
qwen35_4b_universal_low_density_token_match (2026-07-13 — Exact-token local negative)
Three 1,520-row continuations started from the authenticated replay-refresh adapter and received exactly 190 effective-batch-8 updates and 1,429,053 forward tokens. A common 1,440-row replay core occupied identical slots. The 0-, 40-, and 80-row designed doses replaced two, one, or zero independently token-matched 40-row replay blocks. All three trained with zero tokenizer skips.
Fresh local seed 88004 rejected every arm before merge or benchmark. Replay repeat scored 0.500 accuracy, 0.538 parse, and 13 cap contacts; designed40 scored 0.500/0.538/12; designed80 scored 0.538/0.615/10. The inherited replay-refresh anchor was 0.538/0.577/11. Every candidate passed the feasible-route abstention check but failed the frozen accuracy ≥0.65, parse ≥0.90, and cap-contact ≤2 requirements. The promotion receipt contained no eligible arm, and benchmark seed 78134 remained unconsumed.
Read: exact forward-token parity removes the prior compute-dose ambiguity at these low densities. Forty or 80 designed rows are insufficient to install concise local execution from the strong replay anchor. The 80-row arm directionally improves parseability and cap behavior over replay repeat, but only ties the inherited anchor's accuracy and remains far outside the gate. This does not measure broad retention and does not reject intermediate doses or termination-focused mechanisms. A successor must use a new directory and fresh seeds, preserve exact-token replay controls, and pass a prospectively frozen local gate before any benchmark event.
qwen35_4b_universal_mid_density_token_match (2026-07-13 — 160-row near miss; 240-row reversal)
Three 1,520-row continuations started from the authenticated replay-refresh adapter and each received 190 effective-batch-8 updates and exactly 1,405,510 forward tokens with zero skips. A common 1,280-row replay core occupied identical slots. Replay repeat retained three token-matched replay blocks; the designed arms replaced two or three 80-row blocks, each covering all 13 truth-audited skills.
Fresh local seed 88005 rejected every arm before merge or benchmark. Anchor and replay repeat each scored 17/26 accuracy, 18/26 parse, and 9 cap contacts. designed160 improved to 19/26 accuracy, 23/26 parse, and 3 cap contacts; designed240 fell back to 17/26, 22/26, and 5. Every candidate passed accuracy ≥0.65 and had zero feasible-route abstentions, but none met parse ≥0.90 and cap contacts ≤2. The 160-row arm missed each remaining bar by one case. Promotion was empty, and aggregate seed 78135 remained unconsumed.
Read: representative designed density has a nonmonotonic local optimum near 160 rows. Relative to exact-token replay, that arm adds two correct cases, five parsed answers, removes six cap contacts, and shortens mean output by about 218 tokens. Another 80 generic designed rows reverses the accuracy gain and worsens parse/cap behavior, so further dose interpolation is not the warranted next move. A new experiment should hold the 160-row capability mix fixed and isolate concise answer commitment or termination with an exact-token active control and fresh local seed. This result contains no broad-retention evidence and does not license a lower gate.
qwen35_4b_universal_close_weight_token_match (2026-07-14 — Close-weight mechanism negative)
Three short exact-token continuations started from the authenticated designed160 adapter. All received 320 rows, 286,814 forward tokens, 40 effective-batch-8 updates, and zero skips. The active control replayed only incumbent data. The two target arms shared byte-identical fresh execute/induct rows; their sole assigned-loss contrast was weight 0.2 versus 1.0 on the natural </think> span.
Fresh local seed 88006 rejected both treatments before merge or benchmark. The immediate parent scored 16/26 accuracy, 20/26 parse, and six cap contacts; replay repeat scored 14/26, 18/26, and eight. Ordinary target training scored 15/26, 23/26, and three, while close-weighted training scored 16/26, 23/26, and three. All arms passed the route-abstention check. Close weighting missed each frozen numeric bar by one case/contact, promotion was empty, and aggregate seed 78136 remained sealed.
Read: fresh target data improves emission relative to parent and replay, but the byte-identical contrast rejects higher close-span loss as the cause. Ordinary and close-weighted arms have identical parse and cap metrics, both remain 0/4 on the targeted execute/induct cases, and close only adds one non-target abstention win. Its parent-relative accuracy tie is a three-kind-for-three-kind redistribution, not a generalized install. Retire close-span dose tuning. A successor needs a different bounded-computation/canonical-answer commitment mechanism, fresh seeds, an active replay control, and the unchanged local gate before any broad evaluation.
qwen35_4b_universal_search_scaffold_token_match (2026-07-14 — Staged-search mechanism negative)
Two short continuations started independently from the authenticated close_xi near-miss. The candidate replaced 80 of 120 variable replay rows with 16 each of executable apply, fit, reject, execute, and two-branch search lessons. The active control used replay only. Both arms had 320 rows, exactly 286,814 forward tokens, 40 effective-batch-8 updates, ordinary thought/close weights, and zero skips; 200 replay slots were byte-identical.
Fresh local seed 88007 rejected the scaffold before merge or benchmark. Parent, replay, and scaffold scored 18/26, 16/26, and 16/26 correct; every arm parsed 23/26 and contacted the cap three times. Scaffold was 0/2 execute, 0/2 induct, and 0/2 probe, versus parent 1/2, 1/2, and 2/2. It failed accuracy, parse, cap, execute, and induct checks; promotion was empty and aggregate seed 78137 remained sealed.
Read: separately supervising canonical two-operation search substates does not make them reusable at the deployment interface. Scaffold gains two cases and loses four versus parent; mean output lengthens to 520.5 tokens from 434.2. On both execute misses it computes the correct final state in visible thought but over-explains to the cap, while both probe regressions show damaged independent simulation/scoring. Do not add more canonical two-op/two-branch lessons. A successor must be result-separated and prospectively test variable-depth natural-language state tables, hypothesis scoring, and verified answer commitment under fresh seeds and exact-token replay control.
qwen35_4b_universal_state_table_compiler_token_match (2026-07-14 — State-table mechanism negative)
Two short continuations again started independently from authenticated close_xi. The candidate replaced 80 of 120 variable replay rows with 20 each of variable-depth natural-language execution tables, independently recomputed hypothesis scores, first-error repair, and verified commit. Candidate and replay each used 320 rows, exactly 286,814 forward tokens, 40 updates, zero skips, and 200 position-aligned identical replay rows.
Fresh local seed 88008 rejected the candidate before merge or benchmark. Parent/replay/candidate scored 19/16/16 correct, parsed 23/21/22, and contacted the cap 3/5/5 times. Their execute+induct+probe subtotals were 4/6, 2/6, and 1/6. Candidate was 0/2 execute, 0/2 induction, and 1/2 probe; it failed five absolute gates and every strict relative check. Promotion was empty and aggregate seed 78138 remains sealed.
Read: truth-audited natural-language tables can improve isolated computation without installing the deployed procedure. The candidate gained one trace and one optimize case versus both controls and computed one state semantically correctly before losing only on whitespace. But it treated a cycle declaration as an extra operation, repeated both induction cases to cap, miscounted probe distinctness, and reached one correct execute result without committing. The idealized traces remained off-policy relative to actual failure prefixes. Retire another hand-authored trace surface. A successor should use fresh parent rollouts and executable-oracle corrections at the first observable failure prefix, while retaining exact-token replay, fresh seeds, and the unchanged strict local gate.
qwen35_4b_universal_on_policy_prefix_repair_token_match (2026-07-14 — On-policy prefix mechanism negative)
The successor collected 288 fresh parent rollouts and found 230 reachable failures, then selected exactly ten from each of six failure classes. Its candidate masked the realized parent prefix and supervised executable-oracle correction from the first machine-observable failure. Candidate and replay independently trained 320 rows, 40 updates, zero skips, and exactly 304,313 forward tokens; 200 replay positions were identical. The candidate carried 33,421 fewer supervised target tokens, a registered intervention caveat. Both adapters were explicitly merged and authenticated before one same-vLLM local event.
Fresh seed 88009 rejected the candidate. Parent/replay/candidate scored 16/18/15 correct, 24/23/23 parsed, 2/3/3 cap contacts, and 2/1/0 of six on execute+induct+probe. Candidate was 0/2 on every target kind, failed six absolute checks and all four strict relative checks, and had only one paired win against four losses versus replay. It improved no per-kind count; order, probe, and trace each lost one. Local/promotion hashes are b4b333ca...b8c8 / 1e048e75...f5c; all raw hashes revalidated, no benchmark data was read, and aggregate seed 78139 remains sealed.
Read: collecting failures on-policy fixes substrate mismatch but conditioning loss on long realized failure prefixes does not teach the earlier decisions needed to avoid or repair analogous fresh trajectories. Cap-boundary selection and reduced target exposure remain coupled, so this rejects the complete matched-forward-compute recipe, not every on-policy objective. Retire long masked failure-prefix continuation. A successor should supervise short pre-failure decision boundaries and match nonzero target exposure (or include an exact target-token control) before another local gate.
qwen35_4b_universal_failure_selected_restart_target_match (2026-07-14 — Clean-restart mechanism negative)
This successor removed both registered predecessor confounds. It selected four fresh parent failures per each of 13 skills but discarded every failed trajectory, teaching 52 truth-audited solutions from the original prompt. Candidate and replay each used 320 rows, 297,731 forward tokens, 126,796 loss-bearing targets, absolute loss mass 27,632.8, 40 updates, zero skips, and 200 aligned byte-identical replay rows. Both arms independently started from the same authenticated replay parent, then deployed as complete authenticated composites through one same-vLLM local event.
Fresh seed 88010 rejected the candidate. Parent/replay/candidate scored 17/16/15 correct, 21/22/25 parsed, 5/4/1 cap contacts, and 2/2/0 of six on execute+induct+probe. Candidate was 0/2 separately on all three target kinds, missed the 17/26 accuracy floor, and failed all four strict total/target comparisons. Local and empty-promotion hashes are 39fe68b9...de9e / 4c381fbd...6759; all 78 raw requests and model-tree boundaries authenticated, no benchmark data was read, and aggregate seed 78140 remains sealed.
Read: clean restarts reliably changed bounded emission without installing semantic competence. Relative to parent, the candidate produced four more parses and four fewer cap contacts with 34 fewer mean sampled tokens, yet lost two correct tasks and erased both probe successes. Removing the wrong prefix and matching target exposure therefore do not make hand-authored oracle traces policy-compatible. Retire this balanced oracle-restart package. The warranted next test is policy-supported successful-sibling distillation: on fresh procedural tasks, train only where greedy fails but a prospectively sampled same-model sibling is short and verifier-correct, then compare against exact-exposure replay and matched-compute sample-more.
qwen35_4b_universal_successful_sibling_target_match (2026-07-14 — Prerequisite stop)
The registered same-parent trial materialized 624 fresh tasks, 48 per each of 13 skills, and collected one authenticated greedy event before opening sibling sampling. The event completed 624/624 rows and 296,259 sampled tokens at 859.6 tok/s with no recovery or rerun. Model-free grading found 227 hard failures overall.
The prospective four-failures-per-skill prerequisite was impossible: count and route had zero hard failures and select had two. The experiment therefore stopped STOP_INSUFFICIENT_GREEDY_FAILURES; no sibling input, sibling event, training arm, local result, or benchmark result exists, and aggregate seed 78141 remains sealed.
Read: this is not evidence against successful-sibling distillation. It falsifies the design assumption that every universal skill needs and can supply failure-only repair data from the current parent. A successor should treat the ten skills with at least four failures as the residual intervention set and preserve saturated skills with active replay and an unchanged all-skill retention gate. Reuse of the published immutable collection is legitimate only in a new result directory with a new sampling seed and prospective residual policy.
qwen35_4b_universal_residual_successful_sibling_target_match (2026-07-14 — Terminal availability stop)
The residual successor inherited the immutable 624-task source and 227-failure inventory, prospectively treated the ten skills with at least four hard failures (225 rows), and completed its single authenticated same-parent n=16 event at seed 66117 from published-green commit fc5a333b: 3,600/3,600 outputs, 2,337,087 sampled tokens at 739.2 tok/s, no recovery or rerun.
The frozen model-free selection ran from green checkpoint 915a7c62 and qualified 855/3,600 siblings (natural stop, closed canonical thinking, exact answer, ≤768 thinking tokens; dominant rejections: over the short budget 1,527, wrong answer 1,359). Per-task availability was execute 29, optimize/probe/repair 21, state/trace/verify 12, order 11, abstain 6 — and induct 2, below the mandatory four. The outcome is STOP_INSUFFICIENT_SUCCESSFUL_SIBLINGS with zero selected rows; inventory/receipt hashes are 60c95b7a...083e / d3926daf...ad01. No training corpus, adapter, local result, or benchmark result exists; seeds 50/88012/78142 remain unconsumed and benchmark data was never read.
Read: this closes same-parent successful-sibling mining as a universal-curriculum source. Nine of ten residual skills supplied quota easily, so the residual/retention separation worked; the design failed only at induct, where 46 failure tasks and 736 samples yielded two supported tasks. The parent's policy support is empty exactly at the program's wall skill (C38/C39): what greedy decoding cannot do, temperature-0.6 re-sampling within a short-thinking budget cannot supply either. Curriculum signal for the wall must come from designed synthetic data that does not depend on parent policy support — per the queued bounded-computation plus canonical-answer-commitment successor spec — not from harvesting the parent's own successes.
qwen35_4b_universal_fresh_surface_budget_commit_target_match (2026-07-15 — Terminal local negative; positive surface-generality reading)
The bounded-computation successor trained three exactly-matched arms from the authenticated replay_after_close parent (three-axis MILP: forward 1,356,964, nonzero targets 576,718, loss mass ×5 631,326 per arm; zero deltas, zero skips) and evaluated them in one frozen 104-task original-surface gate at seed 88,013 — training rendered only six fresh surfaces, so the gate doubled as a surface- transfer test.
Totals (correct/parsed/caps of 104): parent 63/87/18; replay 62/91/13; designed_fresh 69/97/7; budget_commit 62/88/16. Mean generated tokens 515.8/534.0/357.2/396.2. designed_fresh passed the correct/parse/caps/ abstention bars and won ALL FOUR preregistered strict comparisons (total and the 24-row execute+induct+probe subtotal versus both parent and replay) — the preregistered surface-generality reading is POSITIVE: the designed dose binds to structure, not surface vocabulary, and simultaneously repairs termination (−11 caps, +10 parses, −31% generation length vs parent). But induct was 0/8 for EVERY arm including the parent, so the induct ≥ 4/8 floor was structurally unpassable; no candidate promoted, and aggregate seed 78,143 is permanently sealed. budget_commit was at-or-below replay on every headline number: the bounded-scan lesson did not generalize to termination and its 40-row substitution cost semantics.
Read: (1) the first positive mechanism reading in the universal line — the designed dose is surface-general and think-economical; (2) the budget-commit content lever is retired; (3) the per-kind induct floor is now known to be unpassable for this lineage at n=8 (the C38/C39 wall made exact), so the line's local gate can never promote a continuation regardless of treatment quality. Successors must either attack induction with a fundamentally different mechanism or preregister floors that are achievable given the known wall — and the program's goal-gap forensics (menders/warren/sirens/rites, recorded in the successor intake) point the next attack at the benchmark's actual bottleneck rather than the local gate's hardest kind.
qwen35_4b_goal_gap_axis_curriculum_target_match (2026-07-15 — First local promotion; aggregate pilot negative)
The goal-gap successor trained a 160-row designed axis corpus (four stuck-family axes, public descriptions only, fresh vocabulary, executable truth) against a three-axis exact-exposure replay control from the designed_fresh parent, and became the first universal-line experiment to PASS its local gate: axis holdout 28/40 versus parent 22 and replay 18 (hygiene 9 vs 5/5, explore 7 vs 6/3, tracefix 4 vs 3/2, protocol tied at the control ceiling), with retention byte-equal to the parent (71/95/9 of 104) while replay drifted (65/89/15).
The conditional aggregate pilot then consumed seed 78,144 (quick, tb 1,024, four weight-authenticated composites): base 0.1085, axis_curriculum 0.4223, parent 0.4644, replay_repeat 0.5081. The candidate beat base +0.3138 with 7 strictly positive families, 3 ties, 0 negatives — flipping warren — but lost the aggregate to parent and replay, so the pilot gate fails and the experiment closed per contract. The replay control flipped rites and posted the line's highest recorded aggregate at any seed.
Read: (1) designed axis atoms INSTALL (first promotion ever; zero retention cost) but under-convert to the corresponding quick-tier families — task-level capability and family-level scoring are separated by more than surface (sirens stayed exactly 0.500 despite hygiene nearly doubling locally; menders stayed 0 despite the tracefix win). (2) Replay continuation compounds aggregate a third consecutive time (0.4410→0.4851 at 78133; 0.4644→0.5081 here); it is the strongest single intervention this line has measured and the presumptive parent for successors. (3) The all-families goal is now blocked by exactly two families frozen for every arm at every seed at this tier configuration: menders (0 everywhere) and sirens (0.500 everywhere). Successors must either explain those two constants (instrument-level forensics from public metadata and score behavior only) or find a mechanism that moves them; another same-shape axis dose is not a believable next step for menders after two failed transfer attempts (loomfix, tracefix).
qwen35_4b_axis_replay_stack_medium_target_match (2026-07-15 — Local negative on the breadth bar; stack survival and replay-drift readings)
The stack trial retrained the inherited axis corpus from the 0.5081 replay-compounded parent against a replay-squared exact-exposure control. At the frozen 144-task gate (seed 88,015): axis holdout candidate 24/40 vs parent 18 vs replay_squared 15 (hygiene 9/5/5, tracefix 2/1/0, protocol 8/8/3 — tied at the parent ceiling for the second consecutive experiment — explore 5/4/7); retention candidate 64/98/6 vs 65/92/12 vs 64/86/18. Nine of ten checks passed; the 3-of-4 kind-breadth bar alone failed, so seed 78,145 sealed and the medium-tier pilot never ran.
Read: (1) STACK SURVIVAL — the axis install transfers across parents (+6 axis total twice, hygiene 9/10 twice, best-in-event termination both times); (2) REPLAY ROUND-TWO DRIFT — the second replay round degraded every local quality number (parse 86, caps 18, axis 15/40 with wild kind variance), so the aggregate compounding at seed 78,144 is aggregate-specific or seed-fortunate, not a general quality gain; (3) INSTRUMENT FLAW — the protocol holdout ties at the parent ceiling in two independent experiments, silently tightening 3-of-4 into 3-of-3; successors must handle undetectable kinds prospectively. Queued next (calibrated): a training-free fresh-instrument re-adjudication of the published composites with a detectability-corrected breadth bar and a conditional medium pilot — the mechanism evidence is replicated, the blocker is gate noise, and the cost is merge/eval only.
qwen35_4b_axis_stack_readjudication_medium_pilot (2026-07-15 — Corrected-bar negative; the three-replication mechanism map)
The training-free re-adjudication judged the published stack composites on a third fresh instrument (seed 88,016) with the detectability-corrected bar. All four kinds were detectable; required wins 3. Candidate 22/40 vs parent 15 and replay_squared 18 — the axis-total win's third replication — with kind wins on explore (7/3/6) and hygiene (7/5/5), a third consecutive protocol tie with the parent (7/7/5), and a tracefix loss (1/0/2). Retention 65/98/5 vs 61/92/12 and 66/91/13. Two wins < 3: NOT_PROMOTED; seed 78,146 sealed; the medium pilot never ran.
Read — the mechanism map across three preregistered fresh instruments: INSTALLED: hygiene (won all three events), explore (two of three), and think-economy/termination (caps halved in every event); the axis TOTAL won all three events (+7, +6, +6). NOT INSTALLED: tracefix (4/10 → 2/10 → 1/10, trending to chance — multi-formalism program repair does not take at this dose from these parents) and protocol (tied the parent every time — the parent already carries the skill, so the lesson is redundant dose). The corrected bar worked as designed and the remaining deficit is CONTENT, not measurement. Queued successor (calibrated): axis corpus v2 that keeps hygiene/explore, replaces protocol with a lesson targeting capability the parent lacks, and redesigns trace-repair from this line's own 432-completion-per-arm raw failure outputs (own-experiment data, no benchmark exposure). Do not re-measure the existing composites again; do not reuse sealed seeds.
qwen35_4b_axis_corpus_v2_staged_repair (2026-07-15 — Kill rule fired; third-dose interference)
The forensics-driven v2 (staged repair lessons with demonstrated bounded search; co-location-hardened hygiene; unchanged explore) trained cleanly and met the frozen 154-task gate at seed 88,017 with normalized grading. Axis holdout of 50: candidate 19, parent 19, replay_repeat3 25; per-kind candidate/parent/replay: bugfind 3/0/3, bugmend 3/4/2, retrace 1/2/5, explore 5/7/9, hygiene 7/6/6. Retention 66/98/4 vs 71/98/3 vs 69/95/8. The kill rule fired (u_bugfind_win and u_bugmend_win both false); seed 78,147 sealed.
Two program laws:
- TRACE-REPAIR AXIS CLOSED. Two content designs (asserted search; demonstrated staged search built from quantified forensics) across four fresh-instrument events produced zero robust repair installs. The skill this axis needs is not installable in this model by ~30-55-row designed doses at rank-32, regardless of pedagogy. Any future attack requires a different mechanism argument (per the frozen rule, not a v3).
- THIRD-DOSE INTERFERENCE. The third consecutive designed dose continued in place on one adapter lineage tied its parent on the axis total, lost the previously-installed explore edge, and dropped retention by five — while the third replay round won the entire axis holdout (25/50, explore 9/10, retrace 5/10). Combined with the stack trial's replay-drift reading, the adapter lineage is saturated as a vehicle for further designed doses: future doses need a fresh adapter from a clean parent, and replay continuation remains the strongest single broad-instrument move.
qwen35_4b_hygiene_explore_destack_medium (2026-07-15 — Recovery confirmed; retention bands failed; the replay-interleaving law)
The de-stacking test trained the two replicated installs (hygiene 40 co-location- hardened, explore 40) directly from the clean designed_fresh adapter against matched replay. At the frozen 124-task gate (seed 88,018): axis holdout 15/20 vs replay 11 and parent 8, with BOTH preregistered recovery flags true (explore 7/4/6, hygiene 8/4/5) — the strongest axis result of the session. Retention: 58/93/11 vs 68/98/7 and 66/86/19 — the correct band failed against both controls (−10/−8 vs −5), caps and parse against the parent. Not promoted; seed 78,148 sealed.
Three laws sharpen:
- RECOVERY CONFIRMED. v2's stall was lineage interference, not content decay: the same lessons at the same dose on a clean lineage reinstall decisively at matched exposure. The escalation rule does not fire.
- REPLAY INTERLEAVING PROTECTS RETENTION. The only retention-safe dose-two event (axis_on_replay: byte-equal retention) had a dedicated full replay round between doses; this direct dose paid ten retention points. Combined with every replay-refresh observation, the dose boundary is where replay belongs.
- The gate architecture works as designed: it certified installs and refused a forgetting candidate in one event.
Queued successor (calibrated): the interleaved-replay dose — hygiene+explore warm-started from this experiment's OWN replay_clean adapter (the replay round already exists with receipts), same gate design at fresh seeds. That recipe exactly reproduces the retention-safe precedent with the proven-install content; honest gate probability is the session's highest yet.
qwen35_4b_interleaved_replay_dose_medium (2026-07-15 — Interleaving refuted; escalation fired; dose-recipe search closed)
The direct test of the replay-interleaving retention law trained the verified hygiene+explore corpus from the receipted interleaving replay round. At the frozen 124-task gate (seed 88,019): axis candidate 11/20 vs parent 7 and replay 6 — hygiene won its SIXTH consecutive event (7/2/3) — but explore lost (4/5/3) and retention broke against both controls (59 vs 68/69; −9/−10 against a −5 band), reproducing the direct dose's cost almost exactly DESPITE the interleaved parent. Not promoted; seed 78,149 sealed.
Laws updated, one by refutation:
- REPLAY-INTERLEAVING LAW REFUTED. Replay at the dose boundary does not protect retention. The single retention-safe dose event (axis_on_replay) owes its safety to something else — corpus composition (160 rows/4 kinds), lineage depth, or screen-seed fortune. Cross-receipt inference proposed the law; the preregistered direct test killed it. Record both.
- HYGIENE IS UNCONDITIONAL. Six consecutive kind wins across every parent, dose size, and recipe — the single most robust installed lesson the program has produced.
- ESCALATION FIRED. Per the frozen rule, the dose-recipe search is closed: the ~10-point retention cost of this two-lesson dose is not a scheduling artifact. The only funded successor in this line is a dose-vehicle mechanism study (adapter rank/capacity, loss weighting, dose size, or optimizer dynamics), with its own intake. The 160-row/4-kind composition difference from the retention-safe precedent is that study's first variable.
qwen35_4b_dose_diversity_mechanism_cell (2026-07-15 — REFUTED_INTRINSIC: the retention trade is priced)
The escalation rule's funded mechanism cell trained the twice-verified 160-row corpus directly from the clean parent and gated it at fresh seed 88,020 beside three published composites. Retention correct of 104: clean_parent 70, replay_clean 65 (−5), axis160_direct 61 (−9), hygiene_explore_direct 60 (−10 — its known cost reproduced exactly). Preregistered verdict: REFUTED_INTRINSIC. Axis holdout: axis160_direct best at 26/40 with hygiene 10/10 (SEVEN consecutive hygiene wins, now perfect) and best caps (5).
The four-lifecycle retention arc closes with a priced law: at this vehicle (rank-32 LoRA continued in place, 190 updates, LR 1e-5), designed doses cost ~5–10 retention points intrinsically. Diversity does not protect it (this cell); interleaving does not protect it (prior refutation); a pure replay round itself costs ~5 on a fresh screen; the sole byte-equal precedent was screen fortune. Installs remain unambiguous throughout — the trade is real on both sides. Successors must change the vehicle (rank, loss weighting, update count — single variables against this same gate design) or preregister gates that price the trade. The recipe search remains closed.
qwen35_4b_rank_capacity_vehicle_cell (2026-07-15 — SCREEN_INSTABILITY: the guard fired; bands need calibration)
The vehicle study's first cell trained a fresh rank-64/alpha-128 adapter on the clean-parent composite (trainer's one-argument --model-path delta; encode_row byte-identity enforced by an AST test) and gated it at seed 88,021 beside the published rank-32 arm and the parent. Retention: parent 69, r32 64 (−5), r64 62 (−7). The r32 arm's known −9 failed to reproduce, tripping the preregistered SCREEN_INSTABILITY guard: no capacity inference. Axis: r32 21, r64 19 (install_preserved false), parent 17.
Consolidated across the four most recent gates, same-composite retention deltas scatter ±3–4 points between fresh screens (r32: −9 then −5; two-lesson: −10, −10; replay: −5; r64: −7). The 104-task retention screen's seed noise is comparable to the ±5 band, so single-screen band adjudications near the edge — including parts of the intrinsic-tax chain — carry real draw noise. Standing summary: doses cost retention ~5–10 points with screen noise ~±3; no five-point band should be adjudicated by one screen.
Funded successor (preregistered branch): an eval-only retention-screen calibration study — the published composites re-measured across several fresh screens to size seed variance directly and set bands (or pooled-screen protocols) that separate real effects from draws. Vehicle inference (rank, weighting, updates) stays open until then.
qwen35_4b_retention_screen_calibration (2026-07-15 — CALIBRATION_READ_COMPLETE: the band was ~1.2 SD wide; the tax law revises downward)
The instability guard's funded successor measured the measuring stick: the five published composites re-run on four fresh 104-row retention screens (seeds 88,022–88,025; 20 authenticated engine runs; zero training). The adversarial design review corrected the estimand pre-freeze — bands govern same-screen DELTAS versus the parent, so the calibration pools the delta-vs-parent SD (common screen difficulty cancels; the draft's level SD was wrong in both directions).
Readings: delta SD pooled 4.27 (per-arm 5.68/4.27/3.59/3.10; level SD 4.81 descriptive) → recommended band 9 and frozen protocol pooled_k3. The ±5 single-screen band every prior gate used was ~1.2 SD wide — but ±5 applied to the MEAN of three pooled fresh screens is almost exactly 2 SD (2 × 4.27/√3 = 4.9), so the historical band size survives as a pooled-k3 rule. All five historical single-screen tax readings (−9 axis160_direct, −10/−10 hygiene_explore_direct, −7 axis160_r64, −5 replay_clean) fall INSIDE their arms' pooled ± 2·SD intervals; the pooled deltas are −3.75, −2.25, −0.75, −0.75. Same-composite single readings swing −10 to +4 across screens; screen 88,025 ran commonly hard (parent 64 vs 67–69).
The standing law revises: designed doses cost ~1–4 retention points pooled (not 5–10; the old figure was single-screen draws from a ±4.3-SD process). Installs remain unambiguous; the trade is real but several times cheaper than priced. Vehicle, descriptive only: rank-64 pooled −0.75 versus rank-32's −3.75 (+3.0 favoring capacity, within noise) — the capacity question stays open and is now cheaply adjudicable under pooled_k3 with both arms already published.
qwen35_4b_menders_sirens_tier_forensics (2026-07-15 — CONSTANTS_ARE_INSTRUMENT_ARTIFACTS: the goal gate's venue moves to medium)
The backlog's queued prerequisite ran as pure receipt analysis (2,278 committed gateway files, 356 cleaned family-score rows, zero GPU, zero seeds, benchmarks/ never read). The goal-gap pilot's standing claim — menders = 0 and sirens = 0.500 for every arm at every seed at quick/tb1024 — has three committed counterexamples at the line's own instrument (base sirens 0.375 at 78,131; candidate menders 0.021 at 78,131; replay_refresh menders 0.125 at 78,133): the constants are item-draw artifacts of the quick tier's 1/8-step granularity, not model walls.
The decisive tier read, from paired within-event strict-win adjudication: the goal gate (all ten families strictly above base) passed 9 of 94 historical medium arm-events versus 1 of 84 at quick — the medium mode is 8/10 strict wins with 20 events at 9/10. Base never sits at a family ceiling at medium (0/95 events; quick 2/82), sirens leaves its 0.500 sticking point (base exactly-0.5 in 14/95 medium events vs 49/82 quick, spanning 0.2–0.6), and menders stays beatable (base zero in 54/95, max 0.3; treated arms reached 0.4). Near-miss blockers: menders/sirens/warren at quick; menders/rites/warren at medium.
Honest limit carried on every reading: the nine medium passers were gym-trained arms from the old line (trained ON menagerie-family data); instrument feasibility is established, line transfer is not — the contamination-free universal arms (best 0.5081 quick aggregate) have never been measured at medium. Funded successor: that measurement — base plus the line's best published composites, one fresh sealed medium seed, tb1024, paired same-backend, the goal gate recorded from the same event.
qwen35_4b_universal_medium_tier_measurement (2026-07-15 — MEASUREMENT_READ_COMPLETE: eight wins, zero losses, two ties from the goal)
The forensics' funded successor ran the universal line's first medium-tier paired event: four published composites (trees deep-verified) on sealed fresh seed 78,150 at tb 1,024, one-seed write-ahead ledger, base inside the historical envelope on every family, all arms within budget. The seed-consuming runner was hardened pre-freeze by adversarial review (the review-verdict and code-pin checks now live at the boundary itself; a one-byte drift of the readings evaluator trips them).
Readings: hygiene_explore 0.3379 > designed_fresh 0.3197 > replay_repeat 0.2981 > base 0.0567 — the quick ordering INVERTED (replay_repeat, 0.5081 best-ever at quick, ranks last of the treated at medium): the non-convex tier-Pareto frontier (C54) replicates inside the universal line, and the install carrier leads where it matters. All three treated arms took 8/10 strict family wins versus base — the historical mode — and hygiene_explore/replay_repeat lost NOTHING: ties only at menders and rites (both 0.0). designed_fresh's sole strict loss was warren (0.050 vs 0.067). Sirens resolved to a strict win (base 0.4, every arm 0.6) exactly as the forensics predicted.
Program position: the recorded goal gate is two tie-flips wide for hygiene_explore. rites is elicitable in this lineage (designed_fresh 0.1 in this same event; replay flipped it at quick). menders is the binding constraint — 0 for every clean arm at both tiers on every seed except one quick item (replay_refresh 0.125 at 78,133), while gym-trained arms historically reached 0.3–0.4 there; the same-shape trace-repair dose is closed by kill rule, so the successor must bring a genuinely new mechanism argument for menders, carry rites alongside, start from the hygiene_explore parent, and gate retention under pooled_k3.
qwen35_4b_feedback_loop_state_chain_install (2026-07-15 — NOT_PROMOTED, split install: state-chains teach, feedback-repair fails a third time)
The two-tie install ran the full ladder cleanly (paired fresh rank-32 adapters from the hygiene_explore parent, exact zero-delta exposure, control first, merges pinned, 12-run authenticated gate). The adversarial review had corrected one MAJOR pre-freeze (unbounded op grammars versus the finite uniqueness enumeration — 15 rows admitted out-of-grammar valid fixes; every parameterized op now carries a documented legality clause and an extended-grammar exclusion audit).
Verdict NOT_PROMOTED on three frozen bars, and the split IS the reading: u_statechain INSTALLED — 11/20 on fresh instances, strict over parent (7) and the strong replay control (10); narrated hidden-state tracking is teachable at an 80-row dose (C14's state-chain law reaching the episode protocol). u_feedloop FAILED COMPLETELY — 0/20 on fresh instances of the four formalisms it trained 80 rows on, below both untrained controls (1/20): repair-with-feedback is the THIRD failed pedagogy at the menders-shaped skill (after asserted single-turn repair and demonstrated bounded search, both closed by kill rule). Retention: candidate −3.0 pooled vs the parent (inside the revised 1–4-point tax law) but −5.67 vs the replay control — 0.67 outside the calibrated ±5 pooled band — because replay itself GAINED +2.67 over the parent, repeating the replay-compounding law at the retention instrument. The pooled_k3 protocol's first live use measured event delta SD 4.08 versus the calibration's 4.27: the instrument performs as designed. Axis total tied replay 11–11 (ties fail). Sealed seed 78,151 was never opened and is permanently sealed.
Program consequences: (1) extend the repair kill rule — no small designed dose of ANY tested pedagogy (asserted, demonstrated-search, episode-feedback) installs the repair-shaped skill; menders is open only to genuinely different mechanism classes (scale, scaffolding, or non-SFT levers). (2) The statechain lesson is a proven install and the rites-relevant successor is a statechain-only dose (drop the dead feedloop rows), which also relieves the retention pressure that came from competing against replay's own gain.
qwen35_4b_medium_budget_probe_measurement (2026-07-15 — BUDGET_GATE_STOP: the 8x thinking lever is infeasible at medium)
The budget probe asked whether serial-compute room alone moves the menders/rites floors (the 9-versus-10 goal-ceiling question) and closed on its preregistered stop at the minimum possible cost: the trusted gateway's hard per-arm wall budget refused base at medium/tb8192 (safe diagnostic budget_gate_failed, exit 2, no score emitted, nothing exposed) before any treated arm ran — the frozen order ran base first precisely because the line's quick-tier power statements flagged this risk. Seed 78,152 is spent by the write-ahead ledger's opened record; no retry and no lower-budget re-run are permitted inside the directory.
The delta review had corrected two MAJORs pre-freeze (movement booleans scoped to arm/family pairs at zero in the pinned tb1024 event — designed_fresh's rites was already 0.1 there, falsifying the original premise; and fail-closed benchmark-implementation-signature equality before any cross-budget contrast). Both amendments carry to any successor.
Standing read: the budget lever survives only at intermediate settings (tb2048–4096); base ran medium/tb1024 in 157 s but hygiene_explore was the slowest arm (230 s), so either could bind at the gate. One further preregistered intermediate-budget probe is the lever's last believable test; a second stop closes the lever entirely and fixes the statechain successor's 9/10 ceiling as the program's honest position.
qwen35_4b_medium_intermediate_budget_probe (2026-07-15 — second BUDGET_GATE_STOP: the thinking-budget lever is closed)
The lever's preregistered last test refused identically to the first: the gateway's per-arm wall budget rejected base at medium/tb4096 (budget_gate_failed, exit 2, nothing exposed) before any treated arm ran. Base fits medium at tb1024 (157 s) but not at 4× or 8× thinking room — the wall budget binds between 1× and 4× for the slowest common denominator, and per the frozen consequence the thinking-budget lever is CLOSED ENTIRELY for paired medium events; no further budget probes at any setting without a new mechanism argument. Total cost of the complete answer: two sealed seeds (78,152 / 78,153), two single-arm refusals, zero exposed scores.
Program position after the lever's closure: menders has defeated three small-dose SFT pedagogies AND the deployment-budget lever. The reachable ceiling for currently-believable training paths is 9/10 families (hygiene_explore + a rites flip). The mechanism classes that remain believable for menders — dose SCALE (C43: partial installs were data-limited; all three failed menders attempts were 80–160 rows) and on-policy episode training — each need their own intake, calibration, and kill rules. The funded branch is the statechain-only dose toward the 9/10 position; the queued divergent bet is the dose-scale cell.
qwen35_4b_statechain_only_dose (2026-07-15 — PROMOTED locally; rites CONVERTED; the parent recorded the first 10/10)
Lifecycle 15's funded successor ran the full ladder clean and produced the program's densest single event. Local gate (the first pooled_k3 promotion): axis 21/40 strictly over replay_ctl2 (19) and the parent (17); pooled retention 64.67 vs 66.67/67.33 — inside the calibrated ±5 bands (−2.0/−2.67; the revised tax law priced it right); per-formalism counts recorded (peatstove lost to both controls; muletrack floor-hard).
The medium event at sealed seed 78,154 returned three readings:
- PILOT NOT PROMOTED: the candidate strictly beat base (0.3494 vs 0.0800) and its exposure-matched replay control (0.3157) but lost to its parent by 0.017 (lockpick/siftstack/sirens gave back what rites gained). The parent remains the portfolio's best single model.
- THE CONVERSION: candidate rites 0.300 versus 0.100 for BOTH matched controls on the same seed — the first demonstrated local-install→family transfer in program history. The axis→family under-conversion law has its first counterexample, with an end-to-end causal chain: designed state-chain episodes → local holdout install → benchmark family movement, all paired and exposure-matched.
- THE RECORDED PASS: hygiene_explore_parent goal_gate_pass TRUE — 10/10 strict family wins vs base (menders 0.017; warren 0.150 vs 0.100), zero ties, zero losses, aggregate 0.3663 vs 0.0800. The first all-families pass by a contamination-free arm. The frozen "9/10 ceiling" was a draw-dependent floor-tie, exactly as the tier forensics predicted: menders was marginal capability plus item luck, never an absolute wall.
Honest scope, frozen into every document: single-item margins at menders and warren on ONE seed. The confirmation law — independent fresh sealed seeds plus a same-backend matched-compute sample-more baseline — governs before any claim. The confirmation cell is the immediate funded successor; nothing else outranks it.
qwen35_4b_goal_gate_confirmation (2026-07-15 — AGGREGATE_ONLY: the sweep repeated once; the goal narrows to one family)
The mandatory replication of the recorded 10/10 ran clean: three independent sealed medium seeds, both arms authenticated, every closed ledger record carrying receipt pins, the readout refusing anything not provenance-anchored (the review had caught and fixed exactly that gap pre-freeze). It is also the program's first standalone-compliant cell under the owner directive: the full six-stage lineage package (copied datasets, fixed-seed manifest, three vendored trainer variants + merger, the C53-era root adapter vendored with its provenance boundary stated, rebuild_lineage.py verified in smoke).
Verdict AGGREGATE_ONLY under the frozen ordered partition. The aggregate transfer is unconditional: 0.3287/0.3737/0.3837 versus base 0.0586/0.1122/0.0982 — strict wins on all three seeds, 4/4 all-time with the discovery, never close. The all-families sweep replicated once: seed 78,157 passed 10/10 (two full sweeps across four independent sealed seeds). The frozen 2/3 bar failed because 78,155 (9/10) and 78,156 (8/10) were blocked entirely by TIES — menders at a 0.0 margin on both, warren once (warren WON +0.267 on 78,155) — with zero strict losses anywhere in the event.
Program position, stated exactly: the goal's primary condition is DEMONSTRATED on two of four independent sealed seeds and NOT CONFIRMED at the preregistered majority bar. The gate is localized to a single family: menders, where both arms sit at zero on most draws and every tested small-dose pedagogy plus the budget lever are closed. The funded successor is the menders dose-scale intake (C43 precedent: partial installs were data-limited; all failed menders attempts were 80–160 rows) with a precisely-known success criterion — any reliable nonzero menders yield completes the gate; the zero-root lineage rebuild stays queued as the provenance question.
qwen35_4b_menders_dose_scale (2026-07-16 — NOT_PROMOTED + DOSE_SCALE_NULL: the last SFT route to menders closes)
The scale bet ran the full ladder clean (fresh rank-32 adapters from the hygiene_explore parent at an enlarged 2,280-row stream, 285 updates, exact zero-delta exposure with the control's solver-proven-minimal multiplicity-2 repetition disclosed and direction-of-bias stated; the four new formalisms hand-simulated clean in review) and returned the cleanest possible null: at 10× the failed dose the candidate scored 1/40 on the eight-formalism fresh-instance holdout — EXACTLY equal to both untrained controls (1/40 each). The frozen "nonzero" branch fired mechanically (1 > 0) but the strict control comparison adjudicates it as the guess floor. The dose curve is flat at floor: 0/20 at 80 rows, control-level at 800. Retention fell outside the parent band (63.33 vs 68.67, −5.33) with caps/parsed bands failing against both controls — the larger dose costs more and buys nothing. Sealed seed 78,158 was never opened.
Standing law, now closed at every tested point: the menders-shaped eliminative-repair skill is NOT installable in this model by supervised fine-tuning — three pedagogies (asserted repair, demonstrated bounded search, episode feedback), doses 80→800 rows, surface diversity 4→8 formalisms, and the deployment thinking-budget levers are all closed by preregistered kill rules and adjudicated nulls. Unlike C43's data-limited partial install (which amplified a 0.087), scale does not overcome a zero: this hardens C38/C48 into a dose-independent wall for the eliminative-inference class. Remaining believable classes for menders: on-policy episode training (its own charter, genuinely different mechanism), or standing on the program's demonstrated-not-confirmed position (two 10/10 sweeps across four sealed seeds; nine families hold everywhere; the aggregate transfer unconditional at 4/4).
qwen35_4b_repair_verifier_signal_probe (2026-07-16 — SIGNAL_ABSENT: the map completes)
The on-policy charter's feasibility gate ran clean after its review caught and removed a narration confound (the first build textually marked the wrong option as the failed attempt; the frozen design presents pure failure evidence with symmetric unmarked candidates, a 33-token marker audit at zero hits, and a test-pinned 53.25% shortcut ceiling). Verdict: SIGNAL_ABSENT — think 103/200 (51.5%), nothink 98/200 (49.0%), both at the coin-flip floor and below the shortcut ceiling, with cap contacts at 7.5% (no budget scoping). The model finishes its reasoning and still cannot tell which of two handed repairs works.
The scientific content: the C29 read-only-verifier dissociation is SKILL-SCOPED. Verifying a fix against two trials is execution the model nominally has, but the multi-constraint bookkeeping (two candidates × two trials × outcome comparison) sits behind the same wall as proposal — extending C32/C36/C38/C48 to the verification seat. An on-policy loop would have no reward signal to climb: the class closes for menders by the frozen rule.
THE PROGRAM MAP IS COMPLETE. Every mechanism class for the one goal-gating family is closed by a preregistered rule: three SFT pedagogies (asserted, demonstrated-search, episode-feedback) across 80–800 rows and 4–8 formalisms; the deployment thinking budget (two gateway stops); and on-policy selection (no verification signal). The goal position stands exactly: aggregate transfer unconditional (4/4 sealed seeds, 3–6× margins), the all-families sweep demonstrated on two of four sealed seeds, not confirmed at the frozen 2/3 bar, and the blocking skill precisely characterized — the model can neither produce, nor buy with serial compute, nor recognize multi-constraint eliminative repairs. The statechain→rites conversion remains the program's proven install mechanism; the zero-root rebuild remains queued as provenance.
qwen35_4b_zero_root_lineage_rebuild (2026-07-16 — ZERO_ROOT_DEGRADED, mildly: the prefix priced at ~0.036; the documented recipe carries ~90%)
The provenance question closed with numbers. Six stage replays from a fresh zero-initialized adapter (same datasets, fixed seeds 42/43/44/47/51/55, recorded trainer variants; two honest bumps logged — a mid-stage CUDA-fragmentation OOM resumed under the expandable-segments allocator, and one masked-then-corrected test-fixture failure) produced a composite measured beside the original at sealed 78,159: aggregates 0.0713 / 0.3824 / 0.3462 (base / original / zero-root). The zero-root arm took 7/10 strict wins with ZERO losses; the original read 9/10 (menders tie — the fifth data point on the ~50% sweep rate). Verdict ZERO_ROOT_DEGRADED under the frozen ≥ original−1 rule.
The contrast is the durable finding: the undocumented C53-era prefix is worth ~0.036 aggregate, concentrated exactly in sirens/rites/mirage (coherent with its gym-era training profile) — and it SUPPRESSED chronicle/siftstack/stockade, where the clean rebuild wins outright. So the documented contamination-free recipe carries ~90% of the transfer on its own; the headline model's recorded sweeps lean on the prefix margin (the clean-provenance upgrade is not available); and every prior reading involving this lineage now carries a quantified scope. The mapped clean path, if funded: a zero-root lineage extended with the statechain converter — fully documented end-to-end. The cell also validated the review-installed normalized-hash pin through a real post-merge fill.
qwen35_4b_clean_path_statechain_extension (2026-07-16 — install 3-for-3; the conversion is lineage-dependent)
The clean-path cell ran the full ladder (fresh adapters on the zero-root parent; the treatment byte-copied from the proven converter cell; the six-slot normalized pin held through a real fill; the training-loss property recorded — the clean chain fits the replay surface at ~1.3 versus the original's ~0.43 while performing within ~10% at the benchmark: loss-level ≠ capability, dramatically).
Local: PROMOTED on all eight checks — the statechain install's THIRD replication on its third distinct parent (21/40 strictly over parent 19 and replay 16; pooled retention −1.0/−1.67). The install is the program's most replicated designed effect.
Sealed 78,160: base 0.1234 (the strongest base draw yet — it took rites 0.1, warren 0.133, lockpick 0.1 and squeezed every treated arm to 6/10), parent 0.3517, candidate 0.3333 (beat base 2.7× and the replay control; lost to the parent by 0.018 — pilot not promoted, the same shape as the original cell). THE SCOPING NULL: converts_on_clean_lineage FALSE — candidate rites 0.000 versus the original lineage's 0.300 conversion. The data→family converter is 1-for-2 and lineage-dependent, and the pattern is legible: rites/sirens/mirage were exactly the C53 prefix's strengths, so conversion appears to require substrate the prefix supplied. Footnote: the candidate took a strict menders WIN (0.017, a draw) while rites collapsed — family movements remain draw-coupled.
Standing consequences: (1) the install law strengthens to 3-for-3 with retention held under calibrated bands every time; (2) the conversion law is SCOPED — demonstrated on the prefix lineage, absent on the clean one at its one seed; any future conversion claim must state its lineage; (3) the clean lineage (stages 1–7 fully receipted from the official base, zero contamination) is the mission's reference artifact at 2.7× base aggregate.
qwen35_4b_sweep_rate_consolidation (2026-07-16 — CONSOLIDATED: the terminal claim on all the evidence)
Terminal bookkeeping with a correction. The campaign's closures cited an informal "~50% sweep rate" computed over the 78,154–78,157 window (2/4), later held as "a fifth data point" at 2/5 — both omitting the earlier 78,150 reading. This cell collected ALL SIX recorded goal-gate readings of the reference composite from sha-pinned byte-copied summaries, recomputed every verdict from per-family scores (forensics-identical counting, cross-checked against any recorded gate block), and corrected the figure everywhere with visible errata.
The consolidated terminal statement: strict wins 8/10, 10/10, 9/10, 8/10, 10/10, 9/10 across seeds 78,150–78,159 — sweep rate 2/6 = 0.333 (exact 95% CI [0.043, 0.777]; Beta posterior mean 0.375), ZERO strict losses in all sixty family comparisons, aggregate strict win 6/6, and menders blocking every miss at a 0-margin draw (rites and warren once each; warren WON +0.267 at 78,155). Base draw note: base rites was 0.0 on all six seeds; chronicle drew >0 for base on exactly the two passing seeds. Errata landed in the confirmation README/report, the zero-root README, three synthesis paragraphs, and both carrier briefs; no per-seed fact changed anywhere. The program's flagship number now rests on the complete record with its honest uncertainty.
qwen35_4b_clean_gym_mix_dose (2026-07-16 — NOT_PROMOTED: mixture dilution re-confirmed on clean ground)
The owner-directed cell (recreate the retired prefix's strengths as fresh documented content) ran the full ladder clean and the gate refused decisively: the three-kind mix scored 15/40 on its own holdout — BELOW the parent (17) and the replay control (19) — winning no kind, with retention comfortably in-band (59.67 vs 62.67/64.33 pooled; inert, not destructive). Per-kind: siren_episode 3 vs 4/2 (floored for everyone at 2–4/14), statechain 6 vs 5/7 (the proven skill did NOT re-install at 50 rows — its winning cells used the full 160), mirage_abstain 6 vs 8/10 (the instrument ceilinged for untrained controls — replay 10/13 — so it measures general ability, not the skill). Sealed seed 78,161 was never opened.
Standing law hardened: MIXTURE DILUTION — thin per-kind doses (50–60 rows) install nothing; the dose-diversity refutation now has a clean- lineage replication and becomes a design rule: one kind per dose at full concentration. Instrument rules recorded: an abstention instrument must be hard enough that untrained controls do not ceiling; episode-form injection instruments need difficulty calibration before they can register installs. The retired prefix's three families remain open and now require three separate concentrated cells; the enumerative-repair dose (queued, single-kind by design) is unaffected and proceeds.
qwen35_4b_enumerative_repair_protocol (2026-07-16 — installed at the starkest contrast; failed on its own terms at the family)
The anti-cleverness bet ran the full ladder clean. Local gate: PROMOTED with the program's starkest mechanism contrast — 9/40 canonical-next on fresh instances versus BOTH controls at exactly 0/40 (nobody enumerates systematically untrained), retention deep in-band. The fidelity cascade localizes the leak: parseable 19/40 (long-prompt answer formatting is the bottleneck), legal 18, untried 16, canonical-next 9 (56% ordering discipline once a legal untried candidate emerges).
Sealed 78,162: pilot NOT promoted (0.3252 vs both controls at 0.3502; the dose cost ~0.025 aggregate). The frozen ordered menders rule — installed pre-event after the review caught one-sided consequences — fired FAILED_ON_ITS_OWN_TERMS: candidate menders 0.0 at 22.5% fidelity (below the frozen 0.50 precondition), and the untrained replay control drew a menders item (0.1), denying even the clean zero-contrast. The pure-enumeration SFT route closes at this dose on its own preregistered terms; a formatting-targeted variant requires new evidence per calibrate-and-diverge.
Standing laws sharpened: (1) protocols install 5-for-5 — even this partially-installed discipline cleared strict bars against absolute-zero controls; (2) INSTALL≠CONVERSION hardens further — the wall between a locally-installed skill and a benchmark family stands even for the mechanism the corpus's own laws favored; (3) menders draws continue to land on arbitrary arms (the replay control this time), re-confirming draw-domination at this granularity.
qwen35_4b_count_dont_walk_enumeration (2026-07-16 — MECHANISM_ANSWER: first menders move over all controls; the taught mechanism refuted en route)
The truncation-forensics successor ran the full ladder clean and produced the program's first preregistered positive at the family. Local gate: count_walk PROMOTED (axis totals strictly over BOTH controls, retention in-band) — but the cell's own hypothesis died there: the new expression_cost reading showed the trained candidate still thinking to the 1,024-token cap (median 1,024; truncations 25/40 vs replay 27, parent 32) despite 160 rows of verified ≤105-token constant-cost targets, and canonical-next fidelity reached only 7/40 = 0.175 (replay drew 5/40 by itself off the rendered-ranges prompt delta shown to all arms). Two laws land: SHORT THINKING DOES NOT INSTALL at a 160-row dose — the verbose enumeration is the model's own preference, not a walk-pedagogy artifact — and PROMPT-RENDERED STRUCTURE LIFTS EVERY ARM (the strictly-above-both-controls clause is what kept the claim honest).
Sealed 78,163: menders 0.1 with base, parent, AND replay all at exactly 0.0 → the frozen positive-precedence branch (MECHANISM_ANSWER) fired, the first candidate-vs-all-controls menders movement in program history; the candidate also topped the aggregate outright (0.3312 vs 0.3298/0.2950/0.0753) with a 9-win/1-tie/0-loss goal gate vs base. Honest scope, frozen at closure: one episode in one sealed seed; the reference cell's untreated replay control drew 0.1 on a different seed; and whatever converted is NOT the taught compact computation (refuted above) — candidate-specific deltas vs replay (menders +0.1, mirage +0.4, stockade +0.21) are all single-seed. The confirmation doctrine funds an EVAL-ONLY multi-seed confirmation on the same two committed composites before any claim.
qwen35_4b_count_walk_menders_confirmation (2026-07-17 — AMBIGUOUS: the menders reading does not replicate; no claim)
The eval-only four-seed confirmation (78164-78167, 16 sealed runs, all budget-clean, implementation signature identical across every receipt and the prior event) read out mechanically: candidate one full episode (78164) + one partial (78167, recorded-never-counted), replay control ALSO one full episode (78164) — hits 1 of the required 2, episode totals tied 1-1 with a control, verdict AMBIGUOUS. Frozen claim applied: no claim; further spending on this contrast requires a mechanism-differentiated NEW design, not more seeds of the same. The pre-GPU review earned its keep twice: banker's-rounding episode conversion would have manufactured phantom episodes from partial draws (fixed to floor semantics before any seed was spent — the 78167 partial is exactly the input class that could have corrupted the verdict), and the noise pricing (background full-episode rate ~3/29 per arm-event) is what makes the honest reading unambiguous: the untreated control's hit is the priced coincidence, and 78163's MECHANISM_ANSWER was most likely the same coincidence landing on the candidate. MENDERS REMAINS WITHOUT A CONFIRMED MOVER; single-episode draws at the family are noise until a design clears a multi-event bar. Descriptive, never gating: count_walk topped the aggregate at 2 of 4 seeds (0.398, 0.392 — the program's best aggregate readings to date).
qwen35_4b_count_walk_replay_compound (2026-07-17 — BOUNDED: replay compounding hits diminishing returns at stage 8)
The chain's stage-8 replay-refresh (fresh rank-32 adapter on the count_walk composite, seed 86, mirroring the recipe that added aggregate at stages 1/4/7) ran the full ladder clean and returned the first NEGATIVE compounding result in the chain's history. Sealed 78168: candidate 0.3420 vs parent 0.3626 (a real -0.0206 loss, tie guard inactive) with warren -0.15 past the family slack -> frozen verdict BOUNDED: "the replay-compounding law hits diminishing returns at stage 8 on this parent; the count_walk composite remains the reference; further aggregate pushes need a different move class." LAW: replay compounding is not unbounded — on a replay-saturated parent it redistributes across families (2 up, 3 down, net negative) rather than accumulating. This retires the strongest reflexive broad move as the default next step and forces move-class diversification. The pre-GPU adversarial workflow earned its keep: it caught an aggregate tie-guard gap (ulp-level rendering of true rational ties could have flipped BOUNDED->COMPOUNDED) and the chain-wide standalone-reproduction violation, both fixed pre-freeze; and a mid-ladder orchestration slip (pinning the receipt's internal merge_receipt_sha256 field instead of the committed-file sha) was caught by direct pin-vs-runtime verification before the seed was spent. count_walk remains the program's best aggregate artifact (0.3626 here, records 0.398/0.392 at 78164/78167).
qwen35_4b_state_track_install (2026-07-17 — INSTALLED_TRANSFER: a divergent skill adds where replay bounded; single seed)
After replay compounding BOUNDED at stage 8, the divergent-move-class bet paid off at stage 9. A single-kind state-tracking dose (running ledger through declarative updates, 160 rows, seed 87, made to look nothing like any benchmark family) on the count_walk parent scored 0.3260 at sealed 78169 vs parent 0.3004 and base 0.1675 — beats both, zero families below the one-episode slack -> frozen INSTALLED_TRANSFER. LAW (candidate): the divergent-skill move class is NOT bounded where replay is — a non-overlapping designed curriculum can still add aggregate on a replay-saturated parent, exactly as the install-universal-features doctrine predicts. Gains vs parent land on agentic families reachable by state-tracking transfer (siftstack +0.2, lockpick +0.1, mirage +0.1). HONEST SCOPE: single seed; the parent's aggregate swings 0.30-0.36 across seeds and 0.3260 sits in that band, so per the confirmation doctrine (which killed the menders reading in lifecycle 28) the lift needs a fresh eval-only multi-seed confirmation before state_track is the durable reference. The ultimate every-family-beats-base bar is still unmet (warren 0.200 < base 0.367, count_walk-inherited). The three-lens adversarial workflow found the cell clean (zero MAJORs) and the file-sha merge-receipt lesson from lifecycle 29 was applied at pinning. Funded successor: multi-seed confirmation of the aggregate lift.
qwen35_4b_state_track_confirmation (2026-07-17 — CONFIRMED, directional and soft)
The six-seed eval-only paired confirmation of lifecycle 30's INSTALLED_TRANSFER returned CONFIRMED but honestly soft. Paired deltas (state_track - count_walk, same seed) mean +0.0207, wins 4/6 -> frozen CONFIRMED; but SD 0.0453, paired t=1.12 (5 df), not strictly significant (p~0.16), exactly as the preregistered LIBERAL rule (alpha~0.31, NOT_CONFIRMED the decisive outcome) anticipated. The mean matches the single-seed 78169 lift (+0.0256); all 7 seeds give mean +0.0214, 5/7 positive. LAW (updated): a divergent designed skill produced a REAL but SMALL and NOISY transferable aggregate gain (~+0.02) that replicates directionally - the install-universal-features doctrine is supported, not proven at strict significance at this N. state_track adopted as the program reference composite. The paired design (same seed both arms) was essential: it cancels the parent's 0.30-0.36 marginal seed swing that would otherwise swamp a +0.02 effect. NEXT PHASE (owner-approved): measure base/count_walk/state_track on a real agentic coding harness (duet-eval driving Pi against the local composites), measurement-only, no contamination, to test whether the menagerie proxy gain translates to real coding.
A measurable instrument, and the raw-base control the program never had (2026-07-25)
qwen35_4b_realrepo_agentic_instrument. C63/C64 left this program unable to measure its own next step: execution-selected best-of-N reaches 0.818 on the 11-task synthetic holdout and 0.909 on toolz, with 11/11 holdout tasks solvable, so no +0.05 effect is resolvable; and no raw-base pi baseline exists anywhere, meaning "the warm-start improved pi deployment" is inferred from a TRL-env engagement contrast, never measured.
Built: 200 execution-verified stub-a-function tasks over 15 real OSS libraries -- 138 train / 62 held-out -- firewalled by REPO with each repo's side fixed before any task was generated, so a held-out task lives in a codebase never harvested. Scoring is per-test (SWE-bench style): each task carries the fail_to_pass set its stub breaks, and the repo's baseline passing set minus those is pass_to_pass, which makes "edit the test instead of the code" unrewardable.
Three design traps were caught, each of which would have produced confident nonsense:
- Requiring a globally GREEN suite discarded 11 of 24 libraries over one or two environment-dependent failures (a package-metadata check, a test wanting >6 GB). Those tests fail identically before and after an agent's edit, so per-test sets exclude them by construction.
- The editable-install trap:
uv pip install -e .+ copy-per-episode makes pytest in the copy import the ORIGINAL source forsrc/layouts, so the stub has no effect and every episode scores a free 1.0. Guarded three ways, the strongest being that a task is admitted only if stubbing it actually breaks tests -- so an import leak yields ZERO tasks instead of free reward. -vcannot read a suite whose own addopts enable pytest-xdist;wcwidthbaselined at "0 passed" until-rAwas added. Adding it recovered 8 libraries.
Operational: the pi trajectory for one real-repo episode is ~4 MB of JSON events. The prior runners captured that with subprocess.run(stdout=PIPE), which buffers a child's entire output in the PARENT -- the mechanism behind eight WSL VM deaths (docs/wsl_stability.md). Output is now spooled to disk and the trajectories are retained deliberately as the harvest substrate for the Line-1 successor.
Scorecard
- Program: charter
- Current read: breadth-first expert iteration remains the first blackbox-arbitrated install (+0.223/+0.294 menagerie quick; C49/C50), but C53 closes same-recipe scaling at a robust second wall. Transaction action supervision locally installed proposal structure without transfer. Direct failed-test semantic revision is now descriptively near-saturated: all 72 headroom cases reached a correct patch, while inferred rejected-state first-patch correctness was 0/54 and test inspection before patching 0/72. Structure, evidence acquisition before proposal, and explicit policy editing are separate units. The headroom result is formally instrument-failed because answer-cap contacts exceeded its frozen ceiling. The counterfactual evidence-acquisition strategy remains open: its first curriculum stopped at
LINEAGE_LOCALITY_INFEASIBLE(0.110735 direct start-to-apex drift versus 0.10, entropy retained) before behavior or training, so it is not a capability negative. The universal-curriculum line now has three consecutive procedural-interface negatives: canonical staged search scored 16/26, natural-language state tables scored 16/26, and on-policy failure-prefix correction scored 15/26 against replay's 18/26. The last arm was 0/6 on execute+induct+probe despite 230 reachable parent failures, so on-policy substrate did not overcome late-prefix teacher forcing. - Active experiments: none on this exact MOPD line; the deep-advantage experiment is finished negative after sealed confirmation and its frozen stop prevented benchmark exposure.
qwen35_4b_counterfactual_evidence_acquisition_curriculumis finished at the preregistered lineage prerequisite; an apex-rooted or prospectively parent-qualified successor has not opened. The latest universal on-policy prefix experiment is terminal negative; no aggregate seed was consumed. - Strong anchors:
qwen35_4b_gauntlet_breadth_round1,qwen35_4b_gauntlet_frontier(C53/C54),qwen35_4b_specialist_policy_integration,qwen35_4b_pareto_policy_integration,qwen35_4b_same_prefix_advantage_routing,qwen35_4b_deep_advantage_mopd,qwen35_4b_verifier_conditioned_recovery_bank,qwen35_4b_recovery_reason_locality_interpolation,qwen35_4b_recovery_payload_budget_harness,qwen35_4b_recovery_verifier_branch_tournament,qwen35_4b_transaction_invariant_recovery_curriculum,qwen35_4b_validation_policy_counterexample_curriculum,qwen35_4b_counterfactual_evidence_acquisition_curriculum,qwen35_4b_universal_search_scaffold_token_match,qwen35_4b_universal_state_table_compiler_token_match,qwen35_4b_universal_on_policy_prefix_repair_token_match,qwen35_4b_think_ftpo_round2(C52),qwen35_4b_interactive_policy_curriculum,qwen35_4b_repo_search_compress_bank. - Avoid repeating: un-gated vLLM runtime LoRA; near-self-distillation; filtering deployment-critical force-closes; backend-mixed scores; infeasible specialists; external tier labels; four-branch statewise argmax or posthoc route margins; scaling NF4 MOPD before direct-bf16 merge survival is proven; non-local thought LoRA; unbalanced DAgger; success-only banks without recovery transitions; more generic transaction families; canonical two-operation/two-branch search scaffolds; idealized state-table traces; long masked failure-prefix continuation with reduced target exposure; capability-production training before exact-substrate parent headroom is demonstrated; carrying a useful parent into a stricter direct apex-relative locality contract without prospective qualification; or relaxing a locality ceiling after observing the miss.
- Evidence that advances the program: a beyond-recipe method that exceeds the C53 blend ceiling on fresh blackbox events, or a local checkpoint that learns the shared transactional failure core while retaining broad conditional recovery and beating sample-more on two unseen procedural blocks; for the universal-curriculum line, the next checkpoint must intervene at short pre-failure decisions, match target exposure, and improve execution, induction, and probe scoring under the frozen absolute and strict control-relative local gates.
Charter
Show charter.md
Purpose
Install general agentic capability in Qwen3.5-4B via breadth-first expert iteration on a firewall-clean multi-family gym, arbitrated by the blackbox menagerie instrument.
Why This Is A Program
The corpus's central negative law is that installs are local — format-local (C14), depth-local (C21/C48), shift-local (C43), substrate-local (C45/C48). Every one of those findings came from training on a single narrow substrate. Breadth (many simultaneous, format-diverse, verifier-gated substrates) is the one untested variable, and menagerie — built as "the honest starting line install experiments must beat" and never yet targeted by any experiment — is the only instrument that can arbitrate whether an install moved something general rather than fitting a training distribution. This line hosts the gym itself, successive expert-iteration rounds, ablations (breadth vs matched single-substrate dose), and the transfer-ladder methodology — multiple experiments by construction.
Iteration Doctrine
Iteration speed is the research budget (user directive, 2026-07-09). Default to fast exploration rounds (≤ ~2 h: small harvest → train → menagerie quick base-vs-adapter on a fresh seed); escalate to the full registered regime and slower tiers only to confirm a recipe that fast rounds surfaced. A multi-hour-to-feedback loop is a design failure in this program.
Progress Signals
- A firewall-clean gym exists whose families pass oracle/random/degenerate selftests and whose harvests yield verified, naturally-closed thinking data.
- Adapter-vs-base deltas are reported on the same fresh menagerie seed, with the ~0.011 quick-tier noise floor respected and replication before claims.
- The transfer ladder is measured every round: trained-family held-out items vs held-out gym families vs menagerie aggregate.
- Negative results (locality survives breadth) are codified as claims with the same care as positive ones.
Boundaries
Training data provenance is the model's own verified outputs plus programmatic task generators — never another model. The benchmark firewall is absolute: menagerie is consumed via run.py CLI and aggregate scores only; gym content is invented against the public axis descriptions, never against benchmark items. This program owns the gym + iteration mechanism; it does not own single-lever questions (thinking budgets, confidence readouts) except as they compose into the install recipe — those belong to their existing programs.
Backlog
Show backlog.md
Next Experiments
- Completed qualification negative:
qwen35_4b_pareto_policy_integrationregenerated C54'sblendandapexpolicies and exercised the corrected paireddelta > 0rule.blendlost quick capability in both blocks (-0.0069,-0.0379; pooled-0.0224), whileapexwon deep capability (+0.0456, lower bound+0.0340) but missed six retention cells. The run stopped before teacher audit or MOPD; do not describe it as an integration failure. - Completed route-qualification negative:
qwen35_4b_same_prefix_advantage_routingreplaced tier labels with disjoint same-prefix outcomes. Deep and the combined router passed, but quick's soup-relative audit macro reversed from+0.2009to-0.0253; no MOPD, locality, or Menagerie event ran. Post-result diagnostics show conditional winner noise, not a missing+0.10threshold. - Completed capability negative:
qwen35_4b_deep_advantage_mopdfound replicated same-prefix deep advantage and passed exact locality, but primary seed 42 trailed deep by−0.006845pooled joint and soup best-of-eight by−0.169239; seeds 43/44 also trailed deep. Correct-teacher pressure beat wrong-teacher and non-advantage controls modestly, while retention/transfer passed. This is directional routing signal without source-frontier crossing, not a reason to repeat the NF4 recipe at larger scale. Benchmarking stopped. - Highest-value update-operator experiment: use direct-bf16 microtraining on a fresh self-contained substrate and require the causal update to survive its deployed merge and beat deep, ordinary interpolation, and matched-compute sampling before any full campaign. Preserve the same-prefix verifier and matched controls, but make train/deploy parity a hard admission gate.
- A later two-teacher attempt needs cross-fitted direct advantage prediction, uncertainty-aware adaptive allocation, and a third untouched route block. Permit zero quick allocation if it lacks independent conditional value; do not tune an observed-margin cutoff. This estimator work is downstream of the direct-bf16 update-kernel prerequisite.
- Completed negative:
qwen35_4b_repo_search_compress_bank— exact-token operator balance plus replay-minimized successful repository traces improved trained families 40/48→48/48 but regressed wholly held-out families 49/72→25/72 and failed locality. Do not repeat the one-patch success-only bank at another dose. Any successor needs a new intake, fresh families/seeds, and either (a) verifier-conditioned recovery transitions balanced at the state→action level or (b) external scaffold retrieval/execution that avoids a broad shared-weight policy edit. Require a synthetic failed-patch/failed-test recovery gate before training. - Completed locality-gated negative:
qwen35_4b_verifier_conditioned_recovery_bank. Conditional transition balance worked on fresh trained-family recovery (base 0.483, happy 0.817, action-only 0.850, reason 0.917), and full-dose recovery action-only passed locality at 0.098. The selected 5%-plan arm failed locality at 0.303 because highly off-policy plan-start tokens created a 42.1 pre-clip gradient and 29.5% larger delta norm. Transfer and Menagerie stayed sealed. Do not rerun this dose or reinterpret the exploratory action-only arm inside the result-bearing directory. - Completed policy-gated result:
qwen35_4b_recovery_reason_locality_interpolation. Every frozen mixture passed locality; λ=.18 reached 0.967 trained-family recovery at 0.104 drift, beating base/happy/action/full-reason. It still failed the frozen invalid-turn and immediate-rejected-transition gates, so confirmation, transfer, and Menagerie remained sealed. Do not extend the ladder or lower gates in that result directory. - Completed confirm-gated negative:
qwen35_4b_recovery_payload_budget_harness. The fixed λ=.18 candidate passed locality, calibration, and transfer dev, then tied action-only at 0.6875 on independent confirmation; every other gate passed and Menagerie stayed sealed. The larger payload and two-turn transition metric were validated, but reason mixing is complementary rather than uniformly superior. - Completed prospective infeasibility:
qwen35_4b_recovery_verifier_branch_tournament. On four new families, both sources scored 0.7375 and their union only 0.7500, tying action pass-if-either sample-more. Feasibility stopped the selector; confirm, banking, and Menagerie remained sealed. Do not tune the tie break or try another public source router—the two policies had only one exclusive win each. - Completed transfer-gated negative:
qwen35_4b_transaction_invariant_recovery_curriculumpassed locality and installed trained transaction families (0.817 versus parent 0.517 and replay-only 0.383), but unseen dev was only 0.719 versus parent/sample-more 0.703. It failed the +10/+5 and bootstrap bars; confirmation, broad retention, and Menagerie stayed sealed. Do not repeat generic transaction dose. - Completed calibration infeasibility:
qwen35_4b_validation_policy_counterexample_curriculumpassed locality at 0.109 but parent and matched extra-training control both scored 48/48 on the exact trained-family recovery block. Explicit contract + near-correct partial removed the historical residual; +15/+10 bars were impossible and candidate behavior, transfer, and Menagerie stayed sealed. Do not lower bars or expose the trained candidate post-stop. - Completed instrument-gated qualification:
qwen35_4b_semantic_policy_headroom_tournament. The formal verdict isINSTRUMENT_FAIL: answer-cap contacts were 12.08%/12.67% versus the frozen 5% ceiling. No post-failure axis qualified independently—negative and non-integer were 9/9 in both blocks, while blank had only one in-band shape per block and the shape changed. Do not train on these failed-test states. - Completed prerequisite stop:
qwen35_4b_counterfactual_evidence_acquisition_curriculumended atLINEAGE_LOCALITY_INFEASIBLEbefore interface behavior or training. The transaction-replay parent measured 0.110735 direct drift from C54 apex versus the frozen 0.10 ceiling, while entropy passed; every downstream white-box stage and Menagerie stayed sealed. The evidence-acquisition hypothesis remains untested. Do not relax the gate or swap the parent/anchor in this directory. Any successor must use fresh contexts and seeds, start from apex or a prospectively fixed apex-compatible parent, and separately requalify complete-loop retention plus acquisition headroom before copying the curriculum onto fresh skins. - Future retraining should calibrate plan dose by realized gradient/surprisal and avoid supervising plan starts already rank 1 or wildly off-policy lexical templates.
- Stopped experiment:
qwen35_4b_specialist_policy_integration— incumbent and compound-headroom gates passed, butferrier = 0.994made the mandatory tools specialist's frozen+0.10bar mathematically impossible. Zero specialist or MOPD updates ran. Do not lower the bar or extend this directory. Deferred predecessor repair: a harder disjoint-calibrated tools/provenance core could still make the original four-specialist design feasible, but it is no longer the best immediate test. The corrected successor shows that feasible teachers also need same-prefix advantage on the distillation distribution. Any revival must satisfy both prerequisites before best-of-k or training and use a new experiment with fresh confirmatory seeds.
Completed negative:
qwen35_4b_interactive_policy_curriculum— the run was already underway when the specialist experiment became active. Its full-sequence state-aware DAgger arm failed the mechanism gate (−25.3pp trained, −33.3pp untouched) through semantic-operator capture despite clean atom and closure guards. RL, controls, and Menagerie stopped. The active specialist experiment retained its copied machinery and registered mixed control, but its earlier feasibility stop meant the control never ran. Do not independently rerun this broad warm start.- Completed negative parent factorial:
qwen35_4b_universal_curriculum. The designed-only continuation is a specialization negative (+0.1406 versus base but three negative families and -0.1385 versus blend). The from-base replay union reached 0.692 local accuracy but failed parse (0.846 < 0.90) and cap (4 > 2) gates, so benchmark seed 78132 remained sealed. - Completed designed-arm negative with a stronger control:
qwen35_4b_universal_replay_anchor. The candidate passed local gates but scored 0.4238, belowblendby 0.0172 and replay refresh by 0.0613, with one family below base. Replay refresh scored 0.4851, +0.0441 overblend, with eight positive and two tied families. It is a new anchor, not a universal winner. - Completed exact-token local negative:
qwen35_4b_universal_low_density_token_matchtrained nested 0/40/80-row doses from authenticated replay refresh with exactly 1,429,053 forward tokens, 1,520 rows, and 190 steps per arm. On fresh seed 88004 every arm missed the 0.65 accuracy, 0.90 parse, and at-most-two cap bars; the best 80-row arm was 0.538/0.615/10 and benchmark seed 78134 remained sealed. - Completed exact-token mid-density negative:
qwen35_4b_universal_mid_density_token_matchtrained representative 0/160/240 doses with 1,520 rows, 190 updates, and exactly 1,405,510 forward tokens per arm. On fresh seed 88005, the 160-row arm improved replay from 17/26 to 19/26 accuracy, 18/26 to 23/26 parse, and 9 to 3 cap contacts, but missed the parse and cap gates by one case each. The 240-row arm reversed the accuracy gain and worsened both emission metrics. No arm advanced; benchmark seed 78135 remains sealed. - Completed close-weight mechanism negative:
qwen35_4b_universal_close_weight_token_matchcompared exact-token replay, ordinary fresh execute/induct continuation, and byte-identical continuation with 0.2→1.0 natural-close loss. Fresh seed 88006 scored parent 16/26 accuracy, 20/26 parse, 6 caps; replay 14/26, 18/26, 8; standard 15/26, 23/26, 3; and close 16/26, 23/26, 3. No treatment passed, and seed 78136 remains sealed. Retire close-span dose tuning: it did not improve parse/cap over ordinary training and both target arms remained 0/4 on execute/induct. Do not lower the gate, reuse seeds 88005/88006, or consume sealed seeds 78134/78135/78136. - Completed bounded-computation successor (terminal local negative, positive mechanism reading):
qwen35_4b_universal_fresh_surface_budget_commit_target_matchtrained the fresh-surface designed160 re-dose and the budget-commit ablation against a three-axis exact-exposure replay control and evaluated one frozen 104-task original-surface gate at seed 88013. designed_fresh 69/97/7 (correct/parsed/caps) beat parent 63/87/18 and replay 62/91/13 on all four preregistered strict comparisons with 31% shorter generation — the designed dose is surface-general and think-economical — but induct was 0/8 for every arm including the parent, so the induct >= 4/8 floor was structurally unpassable; no promotion, aggregate seed 78143 permanently sealed. budget_commit was at-or-below replay everywhere; the bounded-scan lesson is retired. The generic-dose line CLOSES here: its local gate cannot promote any continuation of this lineage (0/8 induct is the C38/C39 wall made exact at n=8). Do not reopen with lowered floors in an existing directory. - Completed goal-gap axis-curriculum successor (first local promotion; pilot negative):
qwen35_4b_goal_gap_axis_curriculum_target_matchinstalled its axis atoms (holdout 28/40 vs parent 22 / replay 18, retention byte-equal to parent) and consumed aggregate seed 78144: base 0.1085, axis 0.4223, parent 0.4644, replay_repeat 0.5081. Candidate vs base: +0.3138, 7 positive / 3 tie / 0 negative families, warren flipped; replay flipped rites and posted the line's best recorded aggregate. Pilot gate failed (lost aggregate to parent and replay); closed per contract. Standing facts for successors: menders is 0 and sirens exactly 0.500 for EVERY arm at EVERY seed at quick/tb1024; replay continuation has now compounded aggregate three consecutive times. - Completed stack trial (local negative on the breadth bar alone):
qwen35_4b_axis_replay_stack_medium_target_match— axis install REPLICATED on the new parent (24/40 vs 18/15; hygiene 9/10 twice across experiments; best-in-event termination), replay round two DRIFTED locally (parse 86, caps 18), and the protocol holdout tied at the parent ceiling for the second consecutive experiment, converting 3-of-4 into 3-of-3 and letting one control kind-fluke (explore 7/10) veto promotion. Seed 78145 sealed; medium pilot never ran. - Completed re-adjudication (corrected-bar negative; mechanism map final):
qwen35_4b_axis_stack_readjudication_medium_pilot— all four kinds detectable, candidate 22/40 vs 15/18 (third consecutive axis-total win), kind wins explore+hygiene only, protocol tied the parent a third time, tracefix trended to chance (4->2->1 of 10). Seed 78146 sealed. The corrected instrument worked; the deficit is CONTENT: hygiene/explore/termination install, tracefix/protocol do not. - Completed axis corpus v2 (kill rule fired; axis and dose-stacking closed):
qwen35_4b_axis_corpus_v2_staged_repair— demonstrated-search staged repair lessons did not install (bugfind tie 3/0/3, bugmend loss 3/4/2 vs parent/replay); the candidate tied its parent on the 50-row axis total (19) while the THIRD replay round won it outright (25) and the candidate lost retention (66 vs 71). Seed 78147 sealed. Standing laws: (1) the trace-repair axis is closed for this model at this dose — no v3 without a new mechanism argument; (2) third-dose interference — this adapter lineage is saturated for designed doses; fresh adapters from clean parents are required for any future dose, and replay continuation remains the strongest broad move. - Completed de-stack test (recovery confirmed; retention bands refused):
qwen35_4b_hygiene_explore_destack_medium— both recovery flags TRUE (explore 7/4/6, hygiene 8/4/5; axis 15/20 vs 11/8, the session's strongest axis result), so v2's stall was lineage interference, not content decay. The direct dose paid ten retention points (58 vs 68/66) and the gate correctly refused it; seed 78148 sealed. New law: replay interleaving between doses protects retention (the retention-safe dose-two precedent had a dedicated replay round at the dose boundary; this trial did not). - Completed interleaved-dose direct test (interleaving REFUTED; escalation fired):
qwen35_4b_interleaved_replay_dose_medium— hygiene won its sixth consecutive event (7/2/3) but explore lost (4/5/3) and retention broke against both controls (59 vs 68/69) despite the interleaved parent, reproducing the direct dose's cost. Seed 78149 sealed. The replay- interleaving law is refuted by direct test; the dose-recipe search is CLOSED by the frozen escalation rule. The retention-safe precedent's true cause is unidentified (corpus composition 160/4-kinds vs 80/2, lineage depth, or screen fortune). - Completed mechanism cell (REFUTED_INTRINSIC):
qwen35_4b_dose_diversity_mechanism_cell— retention 70/65/61/60 (parent/replay/axis160/hyg-exp): the diverse dose broke the band (-9), the known -10 reproduced, replay itself measured -5, and the sole retention-safe precedent is resolved as screen fortune. Hygiene 10/10 (7/7 all-time); axis160_direct best axis (26/40) and caps (5). LAW: designed doses cost ~5-10 retention points intrinsically at this vehicle; the trade is priced, real, and two-sided. - Completed rank-capacity cell (SCREEN_INSTABILITY):
qwen35_4b_rank_capacity_vehicle_cell— the r32 arm's known -9 re-measured at -5 on a fresh screen, tripping the preregistered instability guard; no capacity inference (r64 measured -7 retention, 19/40 axis, install_preserved false, but none of it is adjudicable at current bands). Pooled noise across four gates: same-composite retention deltas scatter +/-3-4 points; the screen's seed noise rivals the +/-5 band. - Completed calibration study (CALIBRATION_READ_COMPLETE):
qwen35_4b_retention_screen_calibration— the five published composites across four fresh screens (88022-88025, 20 authenticated eval runs, zero training). Governing delta-vs-parent SD 4.27 (adversarial review corrected the draft's level-SD estimand pre-freeze); recommended band 9; frozen protocolpooled_k3— every future retention adjudication pools THREE fresh screens and applies the +/-5 band to their mean (2 x 4.27/sqrt(3) = 4.9). All five historical tax readings (-9, -10, -10, -7, -5) sit inside measured noise; pooled deltas are only -3.75 (axis160_direct), -2.25 (hygiene_explore_direct), -0.75 (axis160_r64), -0.75 (replay_clean). The 5-10-point per-dose tax law REVISES to 1-4 points pooled; the trade is cheaper than priced. Vehicle descriptive: r64 -0.75 vs r32 -3.75 (+3.0, within noise) — capacity relief is plausible but unadjudicated. - Next funded branch (pick one, own intake): (a) vehicle resumption under pooled_k3 — adjudicate rank-64 vs rank-32 retention on three fresh pooled screens (eval-only if both arms stay published; decisive on the capacity question the rank cell paused); (b) the goal-gate path — a hygiene+explore-lesson dose from the best parent, gated under pooled_k3 retention plus the standard axis instrument, then the MEDIUM-tier four-model pilot (candidate > base/replay/parent aggregate; every-family -vs-base recorded as the goal gate). The revised 1-4-point tax makes (b) cheaper than previously priced.
- Completed tier forensics (CONSTANTS_ARE_INSTRUMENT_ARTIFACTS):
qwen35_4b_menders_sirens_tier_forensics— zero-GPU receipt analysis; the menders/sirens constants have three committed quick/tb1024 counterexamples and the all-ten-families goal gate passed 9/94 historical medium arm-events vs 1/84 quick (base never at a family ceiling at medium). The goal gate's venue is medium; a same-shape menders/sirens treatment at quick is dead. Caveat: medium passers were gym-trained; the universal line is unmeasured at medium. - Completed medium measurement (MEASUREMENT_READ_COMPLETE):
qwen35_4b_universal_medium_tier_measurement— seed 78150, tb1024; hygiene_explore 0.3379 leads (quick ordering inverted; tier-Pareto non-convexity replicates in the line); all three treated arms 8/10 strict wins vs base; hygiene_explore and replay_repeat ZERO losses, ties only at menders and rites (both 0.0). The recorded goal gate is two tie-flips wide. Sirens resolved at medium as the forensics predicted; base inside the historical envelope everywhere. - Completed two-tie install (NOT_PROMOTED, split):
qwen35_4b_feedback_loop_state_chain_install— u_statechain INSTALLED (11/20 fresh-instance holdout, strict over parent 7 and replay 10) but u_feedloop failed completely (0/20 in-domain, below untrained controls): the THIRD failed pedagogy at the menders-shaped skill. Retention -3.0 vs parent (inside the tax law) but -5.67 vs replay (0.67 outside the calibrated band; replay itself gained +2.67). pooled_k3's first live use measured SD 4.08 vs calibration 4.27. Seed 78151 permanently sealed. - REPAIR KILL RULE EXTENDED: asserted repair, demonstrated bounded search, and episode feedback-loop doses have ALL failed at the menders-shaped skill; no further small designed SFT dose of any pedagogy is believable there without a different mechanism CLASS (scale, external scaffolding, or non-SFT levers).
- Completed intermediate budget probe (second and FINAL BUDGET_GATE_STOP):
qwen35_4b_medium_intermediate_budget_probe— base refused at tb4096 exactly as at tb8192; per the frozen consequence the thinking-budget lever is CLOSED ENTIRELY for paired medium events (total cost: two seeds, zero exposed scores). Believable menders mechanisms remaining: dose SCALE (C43 precedent — the three failed attempts were all 80-160 rows; an 800+-row episode-feedback corpus tests scale as the mechanism class) and on-policy episode training. Each needs its own intake. - Completed budget probe (BUDGET_GATE_STOP):
qwen35_4b_medium_budget_probe_measurement— the gateway's hard wall budget refused base at medium/tb8192 before any treated arm ran (base-first order; no score; seed 78152 spent per the preregistered stop). The 8x lever is closed; tb2048-4096 remains the lever's only believable range, with hygiene_explore (230 s at tb1024, the slowest arm) as likely to bind as base. ONE further intermediate probe max; a second stop closes the lever entirely. - Completed statechain-only dose (PROMOTED + RITES_CONVERTED + PARENT_GOAL_GATE_PASS_RECORDED):
qwen35_4b_statechain_only_dose— first pooled_k3 promotion (axis 21/40 strict; retention in-band); at sealed 78154 the candidate beat base+replay but not its parent (-0.017), its rites 0.300 vs 0.100/0.100 is the first local-install→family conversion, and hygiene_explore_parent recorded the FIRST 10/10 all-families goal-gate pass (menders 0.017, warren 0.150 vs 0.100 — single-item margins). One seed ≠ a claim. - Completed enumerative-repair (FAILED_ON_ITS_OWN_TERMS):
qwen35_4b_enumerative_repair_protocol— the discipline INSTALLED at the program's starkest contrast (9/40 canonical-next vs 0/40 for both controls; bottleneck = long-prompt answer formatting, parseable 19/40; ordering 56% once legal-untried) but at 22.5% fidelity below the frozen 0.50 precondition, and did not convert at sealed 78162 (candidate menders 0.0; the untrained replay control drew 0.1; pilot lost to both controls). The pure-enumeration SFT route closes at this dose on its own preregistered terms. Install-vs-conversion hardens; protocols 5/5 on installs. A formatting-targeted variant needs new evidence. - Completed clean gym-mix (NOT_PROMOTED — mixture dilution):
qwen35_4b_clean_gym_mix_dose— 15/40 below both controls on its own holdout, no kind won, retention in-band (inert). Design rule hardened: ONE KIND PER DOSE at full concentration (50-60 rows/kind installs nothing; the statechain wins used all 160). Instrument rules: the mirage-abstain holdout ceilinged untrained (replay 10/13); the siren-episode holdout floored for everyone. Seed 78161 sealed. The three families need three separate concentrated cells. - REOPENED (owner steer, 2026-07-16): "no fundable intake" overstated — the closures cover the TESTED classes, not the idea space. Two new mechanism arguments funded, both derived from the repo's own laws: (a) CLEAN GYM-MIX DOSE (lifecycle 25; replaces a vetoed interpolation idea — owner: do not propagate the prefix lineage, recreate its gym-era ideas as fresh documented content instead) — a mixed 160-row dose on the clean parent covering the prefix's three strength axes with invented executable surfaces: episode-form injection resistance (the 7/7 hygiene mechanism extended to transcripts → sirens), the proven statechain machinery (→ rites), and provable-abstention constraint systems (→ mirage). Attacks the sweep FLOOR on fully clean ground. (b) ENUMERATIVE-REPAIR PROTOCOL (lifecycle 26) — every failed menders dose taught eliminative INFERENCE (closed); none taught systematic ENUMERATION, though C34 (brute search dominates) is a model-level law and protocols are the installable class (4-for-4). Teach: propose legal candidates in canonical order, one per turn, let the rerun verify, stop at success. A genuinely different mechanism the kill rules permit.
- Completed sweep-rate consolidation (CONSOLIDATED — erratum landed):
qwen35_4b_sweep_rate_consolidation— over ALL six recorded goal-gate readings the sweep rate is 2/6 (95% CI [0.04, 0.78]; the informal ~50% was a 2/4 window), menders blocking every miss at 0-margin, zero strict losses in 60 comparisons, aggregate 6/6. Visible errata in all carrier documents. The terminal claim rests on the complete record. - Completed clean-path extension (install 3-for-3; conversion scoped):
qwen35_4b_clean_path_statechain_extension— third statechain promotion on the third distinct parent (retention in-band every time), but converts_on_clean_lineage FALSE (rites 0.0 vs the original 0.300): the data-to-family converter is lineage-dependent, needing the C53 prefix's substrate (rites/sirens/mirage were its strengths). Clean model: 2.7x base, stages 1-7 fully receipted — the mission's reference artifact. Footnote: strict menders draw-win (0.017) at the sealed event; base drew its strongest hand ever (0.1234). - Completed zero-root rebuild (ZERO_ROOT_DEGRADED, mildly):
qwen35_4b_zero_root_lineage_rebuild— the documented stages alone carry ~90% of the transfer (0.3462 vs 0.3824 over base 0.0713; 7/10 strict wins, zero losses); the C53 prefix priced at ~0.036 aggregate concentrated in sirens/rites/mirage and SUPPRESSING chronicle/siftstack/stockade. Contamination-clean upgrade unavailable; prior readings carry the quantified scope. Mapped clean path: a zero-root lineage + the statechain converter, fully documented. - (superseded) QUEUED (zero-root rebuild, own intake — opened by the standalone directive's lineage walk): the hygiene_explore adapter chain warm-starts from the frozen C53-era
blendadapter (no committed creation receipt; gym-era training; vendored+hash-pinned in the confirmation cell). A zero-root rebuild — fresh adapter on the official base, the same six receipted stages replayed at their fixed seeds — answers how much of the recorded goal-gate pass the undocumented prefix is buying. It produces a DIFFERENT model (blend alone scored ~0.44 quick aggregate vs base 0.17), so it is a new experiment, not a reproduction; ~2.5-3h GPU. - Completed confirmation (AGGREGATE_ONLY):
qwen35_4b_goal_gate_confirmation— aggregate strict wins 3/3 fresh seeds (4/4 all-time, never close); goal gate 1/3 vs the 2/3 bar (78157 swept 10/10 — two full sweeps across four independent seeds; 78155 9/10 and 78156 8/10 with ZERO losses, blocked by menders 0-margin ties plus one warren tie). Demonstrated, not confirmed; the gate is localized to menders alone. First standalone-compliant cell (full lineage package). - Completed verifier probe (SIGNAL_ABSENT — the map completes):
qwen35_4b_repair_verifier_signal_probe— 2AFC repair verification at the coin-flip floor in both arms (51.5%/49.0%, below the pinned 53.25% shortcut ceiling, caps 7.5%): the model cannot recognize the repair it cannot produce, so on-policy training has no signal and the class closes by the frozen rule. EVERY mechanism class for menders is now closed by preregistered rule; the goal stands demonstrated (two 10/10 sweeps / four sealed seeds) and not confirmed; no fundable intake remains for the confirmed-goal within the constraint set. New intakes require genuinely new mechanism arguments. - Completed dose-scale null (NOT_PROMOTED + DOSE_SCALE_NULL):
qwen35_4b_menders_dose_scale— 10x dose, 8 formalisms: candidate 1/40 on fresh instances = both untrained controls exactly (guess floor); retention failed the parent band (-5.33); dose curve flat at floor (0/20 at 80 rows, control-level at 800); seed 78158 sealed. MENDERS IS CLOSED TO SFT at every tested point (3 pedagogies x 80-800 rows x 4-8 formalisms + budget levers). Remaining classes: on-policy episode training (own charter, heavy) or stand on demonstrated-not-confirmed. - (superseded) FUNDED NEXT (menders dose-scale intake): the one permitted mechanism class — 800+ episode-feedback rows (10x the failed doses; C43: partial installs were data-limited) from the hygiene_explore parent, calibrated gates, success criterion known exactly: any reliable nonzero menders yield completes the all-families gate whose every other family now holds on every seed.
- (superseded) FUNDED NEXT — THE CONFIRMATION CELL (nothing outranks it): per the confirmation law and the goal's own text, hygiene_explore vs base (and a same-backend matched-compute sample-more baseline arm) on MULTIPLE independent fresh sealed medium seeds at tb1024; the goal gate recorded per seed; pre-registered replication bar (e.g. aggregate + strict all-ten on a majority of seeds, with the menders/warren single-item fragility priced explicitly). Eval-only, no training. A confirmed pass completes the program goal's primary condition; a non-replication prices the 78154 pass honestly as a draw and the program continues from the 9/10 position with the conversion mechanism in hand.
- Queued (calibration notes attached): (a) replay-compounding line — measure whether iterated replay rounds keep gaining aggregate and where they saturate; the published
replay_repeatcomposite (0.5081) is the presumptive parent; cheap, believable, but cannot flip menders/sirens on current evidence. (b) menders/sirens forensics — explain the two frozen constants from public metadata and score behavior only (e.g., whether sirens' 0.500 is a fixed solvable/unsolvable item split at tb1024 and whether any config has ever scored menders > 0 at quick), BEFORE designing another treatment; two same-shape transfer attempts at menders (loomfix, tracefix) have failed, so a third is not believable without new mechanism evidence. Completed staged-search mechanism negative:
qwen35_4b_universal_search_scaffold_token_matchdecomposed two-step search into five independently scored executable lesson stages before a bounded full-search ledger. It starts from the authenticated close-weight near-miss, trains a new same-parent replay continuation, and reserves fresh local seed 88007 and conditional aggregate seed 78137. CPU feasibility and adversarial review passed: both arms have 320 rows, exactly 286814 forward tokens, zero skips, 40 updates, and 200 identical replay positions. Train replay first, publish/verify it, then train the sole scaffold candidate; local failure seals the benchmark. On fresh seed 88007, parent/replay/scaffold scored 18/16/16 correct, all parsed 23/26, and all had three caps. Scaffold was 0/2 execute, 0/2 induct, and 0/2 probe, failed five gates, and seed 78137 remains sealed. Do not repeat canonical two-op/two-branch lessons. A successor must use a new directory and fresh seeds to test variable-depth natural-language state tables plus independent hypothesis simulation/scoring and verified answer commitment.- Completed state-table mechanism negative:
qwen35_4b_universal_state_table_compiler_token_matchtrained variable-depth natural-language state tables, independent hypothesis scoring, repair, and commit against same-parent exact-token replay. Fresh seed 88008 scored parent/replay/ candidate 19/16/16 correct, 23/21/22 parsed, and 3/5/5 caps; target subtotals were 4/2/1 of six. Candidate failed five absolute gates and every relative gate, so seed 78138 remains sealed. Retire another idealized trace surface. Completed on-policy mechanism negative:
qwen35_4b_universal_on_policy_prefix_repair_token_matchfroze 288 fresh truth-audited tasks (48 each across six failure classes), an explicitly mergedclose_xivLLM deployment, exact generated-token prefix masking, and ten reachable failures per class. Reserved construction/rollout/training/local/ aggregate seeds are 77113/66113/47/88009/78139. Design receipt98c6a168...5638authorized the parent merge. That merge applied 128/128 nonzero LoRA modules and produced composite weight hash4933f2dd...eb373. The separately checkpointed parent event then completed 288/288 rollouts and 170252 sampled tokens at 849.9 tokens/s; rollout/receipt hashes are8010632f...3b17f/c6b98b79...74fa, with no generation rerun during wrapper recovery. Model-free mining found 230 reachable failures and cleared every fixed class quota with availability 46/48/35/24/36/41, selecting 10 each. Inventory/source hashes are7230af52...dfe7/30141538...d84b8; the selected prefixes contain 47123 masked tokens and are dominated by cap boundaries. The second review now freezes two 320-row arms at exactly 304313 forward tokens, zero skips, 40 updates, and 200 aligned replay positions; receipt hash iseb08026f...e0cfc. Candidate has 33421 fewer target tokens and lower absolute loss mass than replay, a required causal caveat. From pushed-green commita8529c04, replay trained 320/320 rows with zero skips and 40/40 updates; receipt/adapter hashes aref78f2069...d6de/bb59d3bd...5154d. After that checkpoint passed both workflows, the candidate independently trained 320/320 rows with zero skips and 40/40 updates; its receipt/adapter hashes are846d8107...7098/85811191...0f14. Publish and verify paired training before freezing the fresh local design. Seed 88009 now freezes 26 truth-audited rows with source/input/receipt hashes9682744e...acdee/ff407551...ce988/3982d5b8...6e85a, zero overlap against training and prior reserved local messages, and an identical merged-composite vLLM protocol for all arms. VerdictPASS_CONTROL_MERGEauthorizes control merge only; from pushed-green commit6dc0e677, that merge applied 128/128 nonzero modules. Its tracked receipt, external receipt, and full-weight hashes arebc78f332...d550/aa763255...45a3/7ab4c419...6e2e. From the resulting pushed-green commit619f1e53, the candidate merge also applied 128/128 nonzero modules; its tracked receipt, external receipt, and full-weight hashes are3deff026...438d/baa2027e...6d5a/376e2082...b528. Fresh seed 88009 then scored parent/replay/candidate 16/18/15 correct, 24/23/23 parsed, 2/3/3 caps, and 2/1/0 of six on execute+induct+probe. Candidate failed six absolute and all four relative checks; it won one task and lost four versus replay, with no per-skill count improvement. Local/promotion hashes areb4b333ca...b8c8/1e048e75...f5c; aggregate seed 78139 remains sealed. Retire long masked failure-prefix continuation. A successor must move to short pre-failure decision interventions and match supervised target exposure.- Completed counterfactual-restart mechanism negative:
qwen35_4b_universal_failure_selected_restart_target_matchuses the published stronger replay composite on 624 fresh tasks, four prospective failures per each of 13 universal skills, then removes the parent's failed trajectory and supervises a clean executable restart from the original prompt. Construction/rollout/ selection/training/local/aggregate seeds are 77114/66114/55114/48/88010/78140. Source/input hashes are81edc9ea...de304/25382689...0f5b; zero prompt overlap was found against predecessor sources and prior local seeds. The design requires exact equality of forward tokens, loss-bearing targets, and absolute loss mass against same-parent replay.PASS_PARENT_ROLLOUTauthorizes one parent event only. From pushed-green commit1744e753, that event completed 624/624 rows and 304013 sampled tokens at 879.9 tok/s with no recovery or rerun. Rollout/receipt hashes are4bf15134...1099f/1d35c63a...2b381; no benchmark data was read. After that collection was published green, model-free selection found 602 eligible and 228 hard-failure rows, cleared all 13 quotas, and froze 52 restarts: 40 hard failures plus 12 budget-only cases. Inventory/restart/selection/summary hashes arec19d3de7...66240/022b1ea4...d951f/567d6b02...b662/2e8a2192...e28ddf. The self-contained replay and inherited partition hashes are25a9595f...f0c2/abf8b505...0966f. Exact integral matching then passed: both 320-row arms have 297731 forward tokens, 126796 loss-bearing targets, absolute loss mass 27632.8, zero skips, and 200 aligned shared rows. Manifest/control/candidate/final-receipt hashes are7ba55045...91de1/7a8d4566...b5078/28deb20e...3190/52a761ef...170. Second-review verdictPASS_CONTROL_TRAININGauthorizes replay control only after this freeze is published green. From pushed-green commit821d50d4, replay control then trained 320/320 rows with zero skips and 40/40 updates; receipt/log/adapter hashes are3a9cc1ea...6d49/3bedc341...f25/5840757d...b1c. Publish this control checkpoint before the independently warm-started candidate. From pushed-green commit2c78e655, that candidate trained 320/320 rows with zero skips and 40/40 updates; receipt/log/adapter hashes are6aa5c3f1...9871/c8572c88...202a/2072c5c8...39bc. Publish paired training, then freeze explicit merges and fresh-local evaluation. Both exact seven-file composite trees authenticated. Fresh seed 88010 then scored parent/replay/candidate 17/16/15 correct, 21/22/25 parsed, 5/4/1 caps, and 2/2/0 of six on execute+induct+probe. Candidate failed its accuracy and all three target floors plus all four relative checks; local/promotion hashes are39fe68b9...de9e/4c381fbd...6759, and aggregate seed 78140 remains sealed. Retire another hand-authored oracle restart surface: it improved termination while losing semantic target competence. - Completed successful-sibling prerequisite stop:
qwen35_4b_universal_successful_sibling_target_matchfroze 624 fresh tasks and one authenticated greedy event before any sibling sampling. From green commit0038fba1, the parent completed 624/624 rows and 296259 sampled tokens at 859.6 tok/s with no recovery or rerun; raw/receipt hashes aree91313c0...f556/cee1f19d...4962. Model-free grading then found 227 hard failures, but the prospective four-per-skill prerequisite failed: count and route supplied zero and select supplied two. Inventory/receipt hashes are8e21caf8...d783/3397b773...2a6e; no sibling input was emitted and every downstream seed remains sealed. This does not test sibling distillation. The next result-separated design may reuse the immutable collection, define the ten skills with at least four failures as the residual treatment set, sample only those failures, and protect saturated skills through exact-exposure replay plus the unchanged all-skill local gate. Do not lower this experiment's quota or rescue it in place. Completed residual successful-sibling availability stop:
qwen35_4b_universal_residual_successful_sibling_target_matchcopied the immutable 624-task source and 227-failure inventory, treated only the ten skills with at least four hard failures (225 oracle-free rows), and completed the one same-parentn=16event at seed 66117 from published-green commitfc5a333b: 225 prompts, 3,600 outputs, 2,337,087 sampled tokens, 739.2 tok/s, no recovery or rerun (raw/receipt688c4f7e...c332/c3a3a297...f614). The frozen model-free selection from green checkpoint915a7c62qualified 855/3,600 siblings but only two of 46 induct failure tasks supplied any qualified sibling, below the four-task quota (all other skills ≥6). OutcomeSTOP_INSUFFICIENT_SUCCESSFUL_SIBLINGS; inventory/receipt hashes60c95b7a...083e/d3926daf...ad01; training/local/ aggregate seeds 50/88012/78142 were never consumed. Read together with the balanced predecessor: same-parent successful-sibling mining is closed as a curriculum source — the residual skill that most needs repair (induct) is the one whose successes the parent cannot sample (736 samples, 2 supported tasks). Do not reopen with largern, relaxed token ceilings, or a nine-skill treatment; the wall skill would be untreated by construction.- Experiment:
qwen35_4b_gauntlet_breadth_round1— build the 12-family gym, run round-1 expert iteration, first-ever menagerie-arbitrated install. - Experiment: round 2 re-harvest with the round-1 adapter (does the frontier move, does iteration compound or re-saturate).
- Experiment: breadth-vs-dose ablation — one family at matched total examples vs the full mixture (is breadth causal, or is it just dose?).
- Experiment: leave-one-axis-out mixture — train on 9 families, measure the left-out axis's gym family + menagerie per-family delta (which axes need in-axis data vs transfer in from the mixture).
Required Controls
- For the interactive-policy line: C53 blend incumbent, DAgger-only, compute-overmatched new-state SFT, shuffled trajectory rewards, exact oracle ceiling, family holdout, atom/closure retention, and matched-compute sampling.
- Semantic entropy/outcome variance may route state acquisition; it may not scale token loss or serve as a correctness reward (C52).
- Any future live-state warm start must gate the scarce
VERIFY/COMMIToperator rates and neighboring-policy/logit locality before trajectory RL; full-sequence correctness labels alone are insufficient. - Any compressed trajectory bank must balance conditional transitions, not only operator marginals: explicitly measure
failed_patch→changed_patch,failed_test→revision, andpassed_test→commit. Success-only minimization is not an agent-policy compression method. - For specialist integration: require all four specialists to beat sample-more, DAgger, extra SFT, and shuffled reward before MOPD; require correct-teacher continuation and exact-logit locality before integration; compare end-to-end matched joint RL, off-policy SFT, parameter merge, and KL-matched wrong routing; keep all benchmark seeds sealed until held-out compound transfer passes.
For any state-routed successor: route on a predeclared observable state key or a training-only verifier advantage, never on a hidden benchmark label; freeze the routing rule on disjoint calibration prefixes before producing updates.
- Baseline: base model, same fresh menagerie seed, same tier/decode, every event.
- Mechanism-falsifying control: held-out gym families (never trained) separate generic-protocol gains from axis-specific gains; parse/forced-close/horizon diagnostics reported alongside scores.
- Shift or robustness check: replication on a second fresh menagerie seed before any claim; confirm quick-tier conclusions on medium/slow.
Stop Conditions
Two consecutive rounds with menagerie quick delta inside the noise floor AND flat gym-internal held-out-family transfer would establish "locality survives breadth" — codify the negative claim, then pivot the program to targeted variants (think-economy-only mixture, abstention-only install) or retire.
- Completed:
qwen35_4b_think_ftpo_round2— confident-wrong-turn filtering plus positive-only uplift. P1/P2/P3 failed; true labels separated from shuffled locally but shared-weight collateral erased held-out capability. Candidate only after a new intake/design review: locality-first think-pivot steering. Compare +0.25 uplift with a context-gated last-layer/activation intervention and stop at P1 unless median non-target drift is ≤0.10. Do not fund n=32/gap=1.0 harvesting until a mechanism passes that preflight.
- Template-level hardening for future cells (from the count-dont-walk pre-GPU review; all inherited conventions, none blocking): scope rebuild_clean_chain's original-byte-compare to --verify-inputs so a cell is literally standalone with sibling dirs deleted; extend the normalized-hash pin (or a design-time pin) to eval_local_vllm.py; recompute canonical_next_counts inside authenticate_local_promotion; document the ledger crash-wedge recovery and add a lock or a single-invocation guard.
Experiments 64
- 2026-07-25 Real-Repo Agentic Instrument
Yes: 200 tasks over 15 libraries, each one verified two ways - the library's suite passes untouched, and deleting the target function actually breaks specific named tests. Three design traps had to be avoided, and each w
- 2026-07-18 Qwen35 4B WHY-Comment Install
The WHY idea paid off where it should — on writing correct functions — and did nothing where it shouldn't. We trained the 4B to write code with the causal reason for each line attached as an inline #WHY: comment (generat
- 2026-07-18 Qwen35 4B Self-Repair Install
The first bet that moved the needle at all — gently. We taught the 4B to debug by training on 504 examples of [buggy code + the real test-failure message] -> [diagnosis + fix], all self-generated by injecting bugs into c
- 2026-07-18 Qwen35 4B Repair + Why Stack
The obvious way to combine our two promising curricula - just train one model on both - backfired. We put the 504 self-repair rows and 504 WHY-comment rows into one 1008-row training set and trained a single adapter. Ins
- 2026-07-18 → Qwen35 4B WHY-Think Scale
This phase built and proved the machinery; the GPU training sweep has not run yet. The generator now emits, for every example, a real hidden reasoning trace generated mechanically from the program's shape - parse the spe
- 2026-07-18 → Qwen35 4B WHY Scale Ladder
This phase built and proved the machinery; the GPU training sweep has not run yet. The core blocker was that the original WHY generator saturated fast (about 75 distinct reasons, 438 distinct programs at 504 examples), s
- 2026-07-17 Qwen35 4b State Track Confirmation
The lift held up directionally, without becoming a slam dunk. Across six fresh sealed exams, running the same seed through both models so the noise cancels, the state-tracking model beat its parent on 4 of 6 with an aver
- 2026-07-17 Qwen35 4B Exec-Trace Install
The first bet at installing coding cognition came back flat. We trained the 4B on 400 self-generated, execution-verified program traces to install an accurate 'mental interpreter,' the idea being that a model that can si
- 2026-07-17 Count-Walk Replay Compound (Stage 8)
The believed-likelier outcome, delivered cleanly. 'Replay compounding' — retraining on the accumulated replay mixture — had lifted the aggregate score at every previous link in this model's build chain, so it was the saf
- 2026-07-17 Count-Walk Menders Confirmation
The answer the rule was built to force out, delivered without wiggle room. Across the four fresh exams the trained model solved a fix-the-procedure episode exactly once (plus one partial credit that the rules pre-declare
- 2026-07-17 → State-Track Installation (Stage 9)
The believed-unlikelier but hoped-for outcome landed. After the reliable 'just replay again' lever hit its ceiling, this tried a genuinely different lever: teach the model one new, universal skill — keeping a running tal
- 2026-07-16 Qwen35 4b Zero Root Lineage Rebuild
["The rebuild answered the provenance question with numbers. Retracing the six documented training steps from a truly blank starting adapter — same datasets, same seeds, same settings — produced a model with about ninety
- 2026-07-16 Sweep-Rate Consolidation (Erratum: 2/6, not ~50%)
Two sweeps in six readings — one in three, not one in two — with a wide honest confidence band (roughly 4 to 78 percent at 95%). The texture matters more than the point estimate: the model never lost a single family to t
- 2026-07-16 Repair-Verifier Signal Probe
The gate said no, cleanly. Handed two candidate repairs and the full failure evidence — a task solvable by mentally running each candidate through both trials — the model picked the working one 51.5 percent of the time w
- 2026-07-16 Qwen35 4b Menders Dose Scale
["The scale bet returned the cleanest possible no. After ten times the training data — eight hundred feedback-repair lessons across eight invented machines — the trained model solved exactly one of forty fresh test episo
- 2026-07-16 Enumerative Repair Protocol
A split with the sharpest teaching contrast yet. On the fresh exam, the trained model produced the next-in-order untried candidate on 9 of 40 puzzles while BOTH untrained comparison models scored exactly zero — nobody do
- 2026-07-16 Count-Don't-Walk Enumeration
Two results that point in opposite directions, both real. The taught compact phrasing did NOT take: at the local exam the trained model still thought all the way to its 1,024-token ceiling on most puzzles (25 of 40 answe
- 2026-07-16 Clean-Path Statechain Extension
["Split verdict with a clean lesson. The state-tracking dose installed for the THIRD time on its third different parent — the program's most reliable trained effect — and the clean-lineage model beat the untouched base b
- 2026-07-16 Clean Gym-Mix Dose
The mix failed cleanly and instructively. Splitting the standard 160-lesson budget across three skills — sixty trick-instruction episodes, fifty procedure chains, fifty answer-or-abstain puzzles — taught none of them: on
- 2026-07-15 Universal-Line Medium-Tier Measurement
The clean models' first medium-tier outing put all three at eight-of-ten family wins over the base — matching the best historical arms — and the top two lost NOTHING: they only tied on two families where both they and th
- 2026-07-15 Statechain-Only Dose
Three results in one event. First, the state-tracking dose passed its local gate cleanly — the skill installed again and this time forgetting stayed inside the calibrated margin. Second, on the real benchmark the trained
- 2026-07-15 Retention-Screen Calibration Study
The gap wobbles with a standard deviation of 4.3 tasks — the five-task pass/fail margin was barely one wobble wide, so single-quiz forgetting verdicts were close to coin flips. Every historical 'this model forgot 5-10 ta
- 2026-07-15 Rank-Capacity Vehicle Cell
The trial could not answer the capacity question — by design. Its built-in guard required the known nine-point forgetting case to reproduce before trusting any comparison, and on this fresh screen that case measured only
- 2026-07-15 Menders/Sirens + Tier Forensics
An artifact. The constants already have committed counterexamples on the line's own instrument, and the all-ten-families-at-once win happened 9 times in 94 historical medium-tier comparisons versus once in 84 quick-tier
- 2026-07-15 Medium Intermediate Budget Probe
Same outcome as the first probe, one setting lower: the benchmark's wall-clock referee refused the base model at four-times thinking allowance before any trained model ran. Per the plan written before the event, a second
- 2026-07-15 Medium Budget-Probe Measurement
The probe never got to ask its question. The benchmark's own referee enforces a wall-clock budget per model, and the untouched base — deliberately sent first because this risk was written into the plan — blew past it wit
- 2026-07-15 Interleaved-Replay Dose with Medium Pilot
The review round did not prevent forgetting: the dosed model lost nine to ten retained answers against both comparisons — almost exactly the cost of dosing directly — so the theory drawn from comparing old receipts is re
- 2026-07-15 Hygiene-Explore De-stacked Dose with Medium Pilot
On the fresh start, both reliable lessons came back decisively — the best skill-test result of the whole session (15 vs 11 and 8 of 20) — proving the earlier stall was about over-stacking practice on one model, not about
- 2026-07-15 Goal-Gate Confirmation
The replication returned a split answer. The overall improvement replicated without drama: the trained model beat the base decisively on all three fresh seeds, making four for four all-time. The perfect ten-family sweep
- 2026-07-15 Feedback-Loop + State-Chain Install
The dose split down the middle. The hidden-state-tracking half installed cleanly — on brand-new test instances the trained model tracked procedures better than both its parent and a matched control. The feedback-repair h
- 2026-07-15 Dose-Diversity Mechanism Cell
Variety does not protect memory: the larger varied practice set also cost nine retained answers, the known ten-point case reproduced exactly, and even pure review cost five on this fresh screen — so the one past case wit
- 2026-07-15 Axis Stack Re-adjudication with Medium Pilot
Judged fairly on brand-new tasks — with the ceiling quirk handled — the stacked model won the overall skill test for the third straight time (22 vs 15 and 18 of 40) with the cleanest finishing behavior again. But it won
- 2026-07-15 Axis Corpus V2 with Staged Repair
Even with lessons rebuilt from the autopsy — walking through the search step by step instead of asserting the answer — the repair skills did not stick (a tie and a loss against the comparisons), so the preregistered kill
- 2026-07-14 → 15 Fresh-Surface Budget-Commit Universal Curriculum
The rewritten practice set beat both comparison models on the big screen — more right answers (69 vs 63 and 62 of 104), many fewer run-on answers (7 vs 18 and 13), and 31 percent shorter output — on vocabulary it had nev
- 2026-07-14 → 15 Goal-Gap Axis Curriculum
The targeted practice worked on its own terms — the first screen pass in this program's history (28 vs 22 and 18 of 40 on unseen tasks, with zero forgetting). On the held-out benchmark it beat the untrained model by a wi
- 2026-07-14 → 15 Axis-on-Replay Stack with Medium Pilot
The stacked model kept the installed skills (24 vs 18 and 15 of 40 on unseen tasks, with the cleanest finishing behavior of any model) and lost nothing on the retained skills. But the screen required beating both compari
- 2026-07-14 Policy-Supported Successful-Sibling Universal Curriculum
This version cannot run its second step. Across 624 first attempts it found 227 failures, but counting and routing had none and selection had only two; the plan required four failures in every skill. It therefore stopped
- 2026-07-14 Natural-Language State-Table Universal Curriculum
Not known yet. CPU construction produced 80 truth-checked lessons and two 320-row streams with exactly 286,814 tokens each. All 48 smoke tests pass, but no model has been trained or evaluated.
- 2026-07-14 Search-Scaffold Universal Curriculum
No. On 26 fresh paired cases, the staged-search model tied replay at 16 correct and trailed its parent at 18. All three parsed 23 answers and hit the limit three times; the candidate solved none of four execute/induct ca
- 2026-07-14 Residual-Skill Successful-Sibling Universal Curriculum
Grading the 3,600 retries found 855 short fully-correct ones, and nine of ten weak skills had plenty. But rule-guessing (induction) yielded a usable correct retry on only 2 of its 46 failed tasks, below the required four
- 2026-07-14 On-Policy Failure-Prefix Universal Curriculum
No. The unmodified parent solved 16 of 26 fresh tasks and equal-compute replay solved 18, while training corrections after the model's own real failures solved only 15. The repair model got none of the six targeted execu
- 2026-07-14 Failure-Selected Counterfactual Restart Curriculum
The data-selection stage passed: from 624 fresh attempts, the frozen rule selected 52 clean restarts, exactly four for each of 13 skills. Forty correct accuracy or cap failures and twelve correct-but-too-long cases are n
- 2026-07-13 → 14 Close-Weighted Universal Commit Seam
No. On 26 fresh procedural cases, ordinary and close-weighted target training both produced 23 well-formed answers and three response-limit contacts. Close weighting scored 16 correct versus 15 for ordinary training, but
- 2026-07-13 Validation-policy counterexample curriculum
The test could not answer the question. Before the newly trained model was allowed to run, the starting model and the equal-size old-lesson control each fixed all 48 practice cases. The required gains were therefore impo
- 2026-07-13 Replay-Anchored Universal Curriculum Continuation
No at this dose. The designed mix passed its fresh synthetic screen but scored about 42% overall, below both the 44% mature policy and the 49% replay-only comparison. It also fell below base on one family. Replay-only wa
- 2026-07-13 Mid-Density Token-Matched Universal Curriculum
The 160-lesson stream improved fresh local accuracy from 17 to 19 of 26 cases, raised readable answers from 18 to 23, and cut answer-limit contacts from nine to three. It still missed the readability and length gates by
- 2026-07-13 Low-Density Token-Matched Universal Curriculum
No. Neither 40 nor 80 abstract lessons passed the fresh local screen. The stronger 80-lesson stream matched the inherited anchor at 54% accuracy and parsed 62% of answers, but still hit the answer cap 10 times; broad eva
- 2026-07-13 Semantic-policy headroom tournament
No rule produced a stable training target after a failed test: every one of 72 cases reached a correct patch. But before seeing test output, the model got zero of 54 ambiguous first patches fully right and never opened t
- 2026-07-13 Counterfactual evidence-acquisition curriculum
The run stopped at its first checkpoint-compatibility test. Across 48 frozen contexts, the starting checkpoint differed from the comparison anchor by 0.110735 centered logits, above the fixed 0.100000 limit. Entropy stay
- 2026-07-12 Transaction-invariant recovery curriculum
Partly. Focused training lifted practice repairs from 52% to 82% and preserved the test-and-revise loop. But on different transaction APIs it reached 72% versus 70% before training, far below the required ten-point gain.
- 2026-07-12 Verifier-conditioned recovery banking curriculum
Yes — replaying real failures taught recovery, lifting success from about half of cases to more than nine in ten. But the best version also added a little "explain your fix" text, and that supposedly tiny slice actually
- 2026-07-12 Qwen3.5-4B Same-Prefix Advantage Routing
Only when it replicates, and here one helper didn't. The test actually ran: the deep specialist's local edge held up on a fresh, separate batch, but the quick specialist looked strong on one batch of stuck points and the
- 2026-07-12 Repository search-compress-bank coding curriculum
No. It became flawless on the six codebase types it trained on, solving every task in exactly four clean steps. But on four never-seen codebase types its success collapsed from 68% to 35% — worse than just running the un
- 2026-07-12 Public-verifier recovery branch tournament
No. On four new problem types, both policies repaired 74%, and their combined best reached only 75%—exactly tied with two action-only attempts. The selector was correctly stopped before scoring.
- 2026-07-12 Locality-first recovery-reason interpolation
Yes on skill, no on shipping. One dial setting recovered from broken code 97% of the time, about 12 points above the act-only version and 15 above a matched-training baseline, while barely moving unrelated behavior. But
- 2026-07-12 Payload-capable recovery agent harness
No. It led on the first unseen task block at 71%, but confirmed at 69%, exactly tied with the action-only model and short of the required lead. The external benchmark stayed sealed.
- 2026-07-12 Qwen3.5-4B Pareto Policy Integration
No. Retested on fresh, uncontaminated tasks, the version built for quick work actually lost at quick work by about two points, while the version built for long work won at both quick AND long tasks. One version quietly d
- 2026-07-12 Qwen3.5-4B Deep-Advantage MOPD
No. The deep specialist really was the better teacher on selected states, and copying it worked slightly better than copying the wrong teacher or training on matched non-winning states. But after four rounds the resultin
- 2026-07-11 Entropy-routed think-pivot optimization round 2
No. Gently pulling the better word up at 155 hand-picked wrong-turn moments was cleaner than shoving the bad word down, and it carried real signal, beating a scrambled-label control by about 14 points on fresh coding rep
- 2026-07-11 Qwen3.5-4B Specialist Policy Integration
No — we never found out. Each of four required skills had to gain ten points before merging, but scores top out at 100. One tool-use skill already sat at 99.4%, so its target of 109.4% was impossible. The experiment stop
- 2026-07-11 Interactive policy curriculum: oracle DAgger to execution-reward RL
No — it backfired badly. Success on the practice tasks fell from about 60% to 35%, and on tasks never touched in training from about 69% to 35%. It kept editing fluently but stopped running its checks: the verify-and-com
- 2026-07-10 → 11 Think-block FTPO round 1: outcome-conditioned pivot steering as an agentic install recipe
No. Nudging at those forks made the model worse, not better — success on fresh tasks fell about 4 to 8 percent instead of clearing the 5-point gain hoped for. The tell: feeding it deliberately scrambled labels did nearly
- 2026-07-10 → 11 Gauntlet frontier: difficulty escalation past the breadth-install plateau
No, not both at once. Piling on more, harder, or more varied self-generated practice all stalled at the same ceiling, and even hand-written expert solutions the model could not discover on its own failed to move it. A ne
- 2026-07-09 → 10 Gauntlet round 1: breadth-first agentic expert iteration
Mostly the second. On a blind benchmark the model scored about 14 percent, with six task types near zero, but it had usually reasoned correctly. It simply hit its thinking limit, restarted explaining instead of writing t
Claims
- Confirmed C49 · INSTRUMENT HAZARD: vLLM 0.24 runtime LoRA is a SILENT NO-OP for Qwen3.5-4B PEFT adapters -- every vLLM adapter arm measures the base model; deploy installs as merged composite checkpoints and gate every adapter arm with an on-vs-off behavioral diff
- Promising C50 · Breadth-first expert iteration on a firewall-clean gym INSTALLS SUBSTRATE-GENERAL agentic competence: +0.22/+0.29 on blackbox menagerie quick (paired, deterministic) and +0.52 gym-wide including never-trained families -- the locality laws (C43/C45/C48) do not extend to this regime, and the causal lever was gradient placement at the answer-emission seam, not dose
- Negative C52 · Think-channel FTPO requires both outlier geometry AND parameter locality: near-parity pivots cause generic think-flow harm, while confident-wrong-turn filtering plus positive-only uplift preserves some label signal but still fails held-out capability because shared-weight collateral remains too large; exact loops are absent (~0.1%) at deployed budgets
- Promising C53 · THE SECOND WALL: the emission-policy install is a large ONE-TIME step to a robust menagerie ceiling (quick later broken to ~0.50 by convex mix composition; medium arm-means top out ~+0.31) — no variant of train-on-own-verified-outputs (dose, iteration, breadth, difficulty escalation, recovery supervision, deploy-budget matching) moves the blackbox band further, even as in-gym frontier competence installs
- Promising C54 · TIER-PARETO FRONTIER (corrected): novel serial-compute mechanisms (length-penalized compression advantage + skin-shuffle) lift the MEDIUM menagerie tier to the +0.32 line but do NOT decisively clear it — the early +0.345 read was favorable noise (n=3); pooled n=22 = +0.305 ± 0.010. No single Qwen3.5-4B model clears quick AND medium together by any method (training, capacity, data-interpolation, weight-space soup, expert iteration, tier-router, episode-mastery, or oracle-injection). Refined by C55: at the maxed 8192 budget the medium delta compresses further as the base catches up.
- Promising C55 · BUDGET-COMPRESSION LAW: maxing the menagerie think budget (all tiers → 8192, uncapped `huge` tier, max_model_len 65536) reveals the gym-installed advantage was PARTLY compensation for a budget-starved base. A deployment-time compute-response study first confirms the medium wall is SERIAL-COMPUTE (merged absolute medium score rises monotonically 0.337→0.436→0.518 at think budget 1024/2048/4096); then at the new canonical 8192 budget BASE leaps (quick 0.11→0.46, medium 0.13→0.36) and the merged-vs-base DELTA compresses from +0.33/+0.31 to +0.21/+0.15. The install still yields the best ABSOLUTE capability yet (merged 0.666 quick / 0.506 medium) but its MARGINAL value over a fairly-resourced base is ~+0.15–+0.21, not +0.32.
- Promising C56 · AXIS-STRUCTURED INSTALL COMPRESSION: at the maxed 8192 menagerie budget the two weakest axes DISSOCIATE — EXPLORATION is installable and transfers (gym burrowmaze mean +0.167 at 8192, L6 0.33->0.67; menagerie medium retain-delta +0.190 > the efficiency install's +0.146) while composed-rule INDUCTION is NOT (gym glyphgate L4-L6 stay ~0.0 before and after; trace-SFT even DEGRADES the easy induction the base could already do, L2 0.93->0.53). No single-4B install flavor clears the +0.32 conjunction at fair budget; decomposed by axis, the residual IS the executor-vs-inducer wall (C39/C44/C48), a serial-compute property of the fixed model, not a data or method gap. Answers C55's open next-test.
- Promising C60 · CODING THINK BLOCKS MUST BE HARVESTED, NOT AUTHORED: SFT on synthetic AST-templated <think> traces monotonically CRATERS a near-ceiling coder (HumanEval 147/164->129 at 2k rows, ->121 at 5k rows -- WORSE as train loss drops 5.85->2.70), because the base's 89.6% IS its native reasoning and templated traces are strictly worse. At matched recipe+weight, swapping synthetic->NATIVE (self-sampled, execution-verified) think is +35 HumanEval (113->148) and RETAINS coding (+1); native think survives FULL w=1.0 supervision (148). A 2x2 ablation isolates synthetic-think SUPERVISION as the dominant damage (+17 to +32 HE when removed), #WHY code-comments secondary (+7 HE / +13 MBPP). The -9 MBPP residual is training on easy short-trace problems globally SHORTENING thinking (base hits the 8192 budget on 106/200 MBPP, RFT on 21), hurting long-budget-dependent hard problems.
- Promising C61 · SINGLE-GPU AGENTIC RLVR IS FEASIBLE FOR Qwen3.5-4B, BUT NEEDS AN SFT WARM-START FOR SIGNAL: TRL 1.8 GRPO colocate closes the loop on one 24GB card (movable reward 0.028->0.38 in 3 steps -> the vLLM-served policy updates, so C49's runtime-LoRA no-op does NOT bite TRL colocate), with a required recipe (enforce_eager monkeypatch for the hybrid arch's CUDAGraph hang; bf16 model_init load not fp32; vllm_gpu_memory_utilization 0.55 + sleep_mode -> peak ~22.6GB). The full agentic loop (environment_factory CodingEnv, thinking-on, pi-mirroring read/write/list/bash tools, pytest reward) runs end-to-end. But the RAW base yields reward_std=0 (explores 1-2 reads ~120 tok then quits WITHOUT writing) and the 24GB card caps num_generations at ~4 (8 OOMs), so a narrow loop-discipline SFT warm-start is the prerequisite before RLVR learns.
- Confirmed C62 · NARROW POLICY EDITS PERTURB THE 0.606 pi-DEPLOYMENT WARM-START DOWNWARD; THE REMAINING HEADROOM IS A TERMINATION/FINISHING PROBLEM, NOT ABSENT CAPABILITY
- Confirmed C63 · SINCE EDITING THE 4B's POLICY REGRESSES IT, THE DEPLOYABLE AGENTIC-CODING LIFT COMES FROM INFERENCE-TIME EXECUTION SELECTION: best-of-3 beats single-shot +0.121 on the pi holdout with ZERO training
- Confirmed C64 · THE 4B DOES REAL-CODEBASE AGENTIC CODING: 0.70 single-shot / 0.91 execution-selected on real toolz functions via pi -- the '~0.00 on real repos' result was a HARNESS ARTIFACT