Structured Execution and Compilers
Represent tasks as executable, typed, latent, or stateful programs instead of direct final answers.
What we have learned
Seed Experiments
- qwen_structural_latent_compiler_expansion
- qwen_compiler_multiseed_reattribution
- qwen_typed_bytecode_expert_iteration
- latent_executor
- structured_slot_initializer_ladder
Key Result
qwen35_4b_state_carry_vs_state_bag (unclaimed; terminal LoRA
PILOT_MECHANISM_MISS): a valid, source-bound seed-7401 pilot tested whether rank-32 extra-call LoRA over two complete Qwen3.5 hybrid motifs could learn a serially inherited query-before-state representation against an equal-parameter/equal-compute Bag. All registered pilot cells and identities were complete, K=1 parity was exact, and the configured +0.05 gate was reachable. Carry nevertheless failed deep state formation: joint node+phase+checksum step accuracy was 0.00459 versus the frozen 0.40 gate and node accuracy was 0.0642. The small matched-depth answer effect (+0.043, pilot CI -0.008 to +0.094), unseen-K gain (+0.012, CI -0.035 to +0.059), and positive joint holdout (+0.051, CI +0.008 to +0.098) do not rescue a chance-like registered state. Swaps were also noncausal (donor-follow gain +0.008, CI -0.023 to +0.039; donor follow minus recipient preserve -0.055). Confirmation, edge cuts, and sample-more were correctly not run. This narrowed the result to the registered low-rank adaptation recipe did not form the required deep joint state and licensed the held-fixed full-rank successor below.qwen35_4b_state_carry_vs_state_bag_fullrank_delta (unclaimed; raw analyzer label
PILOT_STATE_FORMATION_MISS, post-result preregistration audit dispositionPILOT_PROMOTION_BLOCKED): the preregistered capacity control replaced rank-32 LoRA with 892,272,640 direct FP32 full-rank deltas on the same 62 extra-R linears while holding the parent rows, recurrence, loss, optimizer schedule, seeds, base K=1 path, and Carry/Bag comparison fixed. Live G0 passed with complete Adam state, exact K=1 before/after the optimizer, finite K=12, a bit-exact 3.571 GB checkpoint round trip, and 22.57 GiB reserved headroom. The complete matched pilot nevertheless formed essentially no registered state: joint step accuracy 0.00277 versus 0.40, node accuracy 0.0617. Carry lost to Bag by 0.0156 (CI -0.0664 to +0.0391), unseen-K gain was -0.00781 (CI -0.0625 to +0.0469), and swaps reduced donor following by 0.00781 (CI -0.0391 to +0.0156). The same pilot also failed non-capacity promotion requirements: its Carry-minus-Bag effect was not positive, and neither registered query-kind effect was positive (node 0.000, checksum -0.03125). Although all cells were complete, the answer gate was reachable, and the answer interface was valid, those simultaneous failures mean the state-formation miss was not isolated under the preregistered disposition table. The defensible terminal disposition is thereforePILOT_PROMOTION_BLOCKED, not a closure of LoRA capacity. A fresh RNG-matched three-seed LoRA-versus-full-rank state-formation adjudication is mandatory before drawing a rank conclusion; neither existing pilot is eligible for confirmation, edge cuts, or sample-more.qwen35_4b_state_formation_capacity_adjudication (unclaimed; in-progress producer verdict
LORA_JOINT_MISS_CONTROLS_REQUIRED): the fresh adjudication fixes the prior confounds with bit-identical shared state initialization, seed-matched dropout/order, exact K=1 bypass, setup positive controls, three fixed-final 1,500-step seeds, and a source-bound immutable analyzer. The rank-32 LoRA joint arm validly failed state formation: 0/57 required seed×split×depth cells passed 0.40, with maximum intact accuracy 0.0234375 and per-seed maxima 0.0234375/0.015625/0.015625. Trained, deep-extrapolation, and joint-shift categories all missed; adaptation contrast was uncertain. This is stronger evidence that the registered LoRA joint recipe fails than either predecessor, but it still does not identify rank as the cause. The frozen branch now mandates three LoRA state-only controls and three direct-full-shape joint arms before any rank-causal conclusion or sealed contrast.qwen35_4b_commit_slot_jacobian_value_transport (unclaimed; terminal
COMMIT_SLOT_SEAM_FAIL): a fixed latent answer interface repaired formatting—an alias was already the unmasked next token on 41/48 at cap 1,024—but semantic correctness remained task/alias concentrated. Real ordered thought was 15/48 versus 11/48 exact-length shuffled, with five mixed tasks versus six required. A structured slot can expose a decision without making that decision reliable; task-level confirmation remains mandatory.qwen35_4b_early_text_hypothesis_forking (unclaimed; terminal
INVALID_INTERFACE_PARSE): all 392 locked mechanics rows authenticated. Exact early bound text controlled direct one-operation execution broadly: systematic and length-matched deranged hints each achieved 84/96 execution of their own supplied operation, while deranged produced 0/96 of the registered target; all 24 operations and four contexts had support. This did not become a valid proposal controller. Every diagnostic arm exceeded the 0.05 answer-cap ceiling, duplicate/placebo also failed parse, and the independent noncausal full-program ceiling reached only 3/8 visible passes. Qualification stayed sealed. The next justified representation is a materialized residual state—candidate intermediate outputs plus the remaining target relation—not another opaque name, larger budget, or parser relaxation.qwen35_4b_partial_structure_search (unclaimed while ledger re-grade is open): type-only partial viability is oracle-useful but model-unreadable. A width-4 exact live-prefix beam retained a hidden solver on 12/12 dedicated depth-5 development tasks at 262,144x completed-leaf compression. Frozen Qwen3.5-4B thinking P(viable), however, was chance within task on 7,200 depth-4 children (AUROC 0.506, CI 0.470--0.543; recall@4 0.251) and significantly below no-think AUROC (delta -0.049, CI -0.090 to -0.010). Pooled AUROC 0.557 was a task-difficulty mirage; wrong-task visible examples were no worse. The gate correctly stopped depth-5 model search and banking. Separately, exact visible-only depth-5 brute covered 60/60 and selected 56/60 in 112 seconds on eight CPU workers, so the next question is the real depth-6 resource crossover, then a residualized rather than type-only state.
qwen35_4b_crosssubstrate_structure (claim C36): the recent structure findings are MODEL-LEVEL LAWS. C32 (wall-is-structure) + C34 (brute-search dominates) replicate on STRING (char edits) + REGISTER (int machine) + LIST: base ~0, structure-cov = concrete-cov, oracle-skelfill 1.0, random low, brute-deploy ~1.0 on all three. The fixed 4B is a value-computer not a deep-structure-proposer, across substrates.
qwen35_4b_structure_search_scaling (claim C35, re-graded): brute-full deploy stays 0.967 at depth 4 (vs 0.975 at depth 3), while the tested banked models' structure coverage is 0.10 and 0.51. The banked comparison crosses non-dose-matched models, so it is not a causal depth curve. Brute dominates the measured list-DSL cells through depth 4; depth 5 was projected, not tested, and is the open model-guided-search regime.
qwen35_4b_banking_installs_structure phase 2 (claim C34): end-to-end bank+value-fill deploy. bank-fill deploys 0.463 (= banked structure-cov, confirms C33) BUT brute-force structure enumeration + value-fill + execution-consensus deploys 0.975 (near-solves depth-3) WITHOUT the model. With the interpreter, free structure-search dominates; banking's structure is forward-pass-only. Extends C17 (selection free) to structure-search. Scope: brute wins because the 4096-skeleton space is enumerable.
qwen35_4b_banking_installs_structure (claim C33): banking installs STRUCTURE -- base op-sequence structure-coverage 0.00 -> banked 0.51 (held-out depth-3, generalizable). Banking converts the wall from structure-bound (base) to value-bound (banked struct 0.51 > concrete 0.36, value tax +0.15, fillable). Mechanistic closure of C32: banking = structure-installation. Unifies C22-24/C31/C32.
qwen35_4b_structure_or_values (claim C32): the compositional wall is STRUCTURE, not values. The model's STRUCTURE-coverage (right op-type sequence, any param) = its concrete coverage (value tax +0.000 at depth-3) -> failures are wrong-skeleton; oracle-skeletonfill=1.0 (values trivial given structure); random-skeletonfill low (DSL not value-fungible). Unifies C19/C25/C31; explains why tool-structure-seeds (C22)+banking were necessary. (op-seq generation fails at 0.00 = separate format handicap.)
- qwen35_4b_thinking_lookahead (claim C26): TEST-TIME thinking (on a model never trained to reason about this task) does NOT breach the lookahead wall — it amplifies recognition, not planning. (Scope: leaves open whether banking successful reasoning traces would install planning-via-thinking — the clean untested version.) C25 found the fixed 4B can't plan the first of 3 ops in one forward pass. Does thinking (serial test-time compute, the dormant C9 lever) breach it with no training? Channel-matched test (think→RANK vs no-think→RANK, parse-immune), headlined on STEP 1 (the only clean lookahead test). Step-1 stays at chance across budgets (0.025 → 0.050 → 0.075 at B=0/1024/2048; Wilson CIs overlap). But thinking's benefit scales inversely with lookahead distance: step-3 recognition (1 op away) 0.275 → 0.600, step-2 0 → 0.325, step-1 (3 away, real planning) ~flat. So thinking amplifies recognition, not planning — and internal-brute-force is refuted (step-1 would rise if the model could simulate the path; it doesn't). The juxtaposition with C25: banking lifted step-1 lookahead (0.013 → 0.138) while thinking does not — so for the planning gap, training is required; test-time compute alone can't elicit it. Reconciles with C23 (base think single-shot depth-3 = 0). Design hardened by an adversarial review. Limits: closed-set ranking, n=40, budgets ≤ 2048.
- qwen35_4b_latent_decomposition (claim C25, re-graded): the base next-op ranker is at/below chance only for the first move three operations from the goal; step 2 is weakly above chance and terminal recognition is stronger. Base-guided versus random search solved 1/80 versus 2/80 tasks, so the defensible conclusion is “no better than random,” not “worse.” Banking improved step-wise rankings and low-budget search (18/80 banked versus 2/80 random and 1/80 base), roughly matching brute's 23/80. The dose trend is supported at steps 2–3, not at step 1 (10/80 versus 11/80 for the two banked doses). This is a closed-menu behavioral guidance lift, not demonstrated internal planning; one beam and single adapter seeds bound it.
- qwen35_4b_depth_scaling_controls (claim C24): three follow-ups to C23 — no saturation, the gain is data-diversity, and the recipe repeats one rung deeper. (1) The depth-3 dose curve does NOT saturate through 1280 tool-pairs (1156 distinct functions): cov@16 climbs 0.00/0.087/0.212/0.375/0.537, deployable greedy@1 → 0.188; distinct functions grow near-linearly so it's real capacity. (2) A 2×2 at matched steps/mixture splits the gain: diversity (up40 0.163 → train_640 0.375, same compute) is cleanly significant; the pure compute effect (N=40 0.087 → up40 0.163, same 40 functions) is within noise — so C23's "data-limited" is data-DIVERSITY-limited. (3) The tool-search+banking recipe repeats one rung deeper, weakly: depth-4 cov@16 base 0.00 → scaffold transfer 0.067 → banked_d4 0.183 (~3×), but test-time-only (greedy flat 0.033) and marginally significant at n=60; no depth-3 forgetting (guardrail 0.425). Design hardened by an adversarial workflow review (scaffold-only baseline, distinct-fn counts, 0-leak d3 0/2305 & d4 0/318, true-depth-4). Limits: single seed, n=60–80 underpowers adjacent-dose/depth-4 significance, depth-4 single dose, 2560 dose dropped.
- qwen35_4b_depth3_dose_response (claim C23): the depth-3 install is DATA-LIMITED, not a representational cap — and it scales into deployable single-shot. C22 left open whether its weak depth-3 install was data-limited or capped. Bank N tool-found depth-3 pairs (N ∈ {40,160,640} nested, interpreter search over the 16-op DSL); eval on a frozen paired held-out set with 0 leakage (function-sig AND op-composition dedup → novel rules only). Depth-3 think coverage@16 rises MONOTONICALLY 0.00 → 0.087 → 0.212 → 0.375, no plateau; top-dose Wilson lower CI (0.28) > low-dose upper CI (0.17). The DEPLOYABLE install scales too: no-think coverage 0.00→0.338, no-think single-shot greedy@1 0.00→0.10 at N=640 (≈0 at C22's N=130). Depth-2 guardrail rose (scaffold intact). So the deep wall is a DATA bottleneck, not a hard cap: the thin depth-3 thread (C19) thickens with more explorer-found data and converts to deployable single-shot. Design hardened by an adversarial workflow review. Limits: single seed, fixed epochs (data~gradient confound), search-easy bias (untested past 640).
- qwen35_4b_tool_seeded_banking (claim C22): the C21 positive control — tool-seeded banking crosses the depth-3 wall self-banking couldn't, but weakly. Harvest depth-3 via an interpreter-backed explorer (CPU brute-search over the substrate's own 16-op DSL, no external model, 130/130 solved — what sampling gets ≈0 of), add to C21's exact depth-1+2 pairs, bank. On a frozen paired held-out set (behavioral dedup; design hardened by an adversarial multi-agent review): depth-3 think coverage@16 0.00 (0/40) → 0.125 (5/40 distinct novel tasks) — a significant unlock vs the hard 0/40 floor where C21 self-banking gave exactly 0. Validates the recipe: tools explore, banking installs. But CROSSED-BUT-WEAK — the install is test-time-dominated (no-think depth-3 0.025, greedy@1 0.00) vs depth-2 which installs deployably (greedy@1 0.15). New nuance: the installer's efficacy decays with depth (echoes C19). No free next rung (depth-4 stayed 0). Each rung must be seeded by the explorer.
- qwen35_4b_wall_climbing (claim C21): self-banking is coverage-seed-bounded — it can't climb the wall. Apex bootstrapping test: bank ONLY depth-1+2 self-solutions (130 pairs, 83 at depth 2, no depth-3 examples), does the banked model now sample depth-3? DEPTH-LOCAL. Depth-2 install works and generalizes to held-out tasks (coverage 0.12→0.36, tripled — clean C18 replication) but depth-3 coverage stays at exactly 0.00 (base 0.00 too) — a strong depth-2 composition skill does NOT length-generalize up. Banking installs only depths the base can already sample; it cannot bootstrap the frontier. Completes the wall picture: depth-3 is not represented (C19), not steerable (C20), not reachable by banking-shallow (C21). The only way up is to seed each rung externally — tool-augmented harvest (C12 decompose-search) → verify → bank. Sharpens C11-M4 into a hard cross-depth wall. Pre-registered P2 (unlock) refuted; P1/P3 held.
- qwen35_4b_activation_steering (claim C20): decodability ≠ steerability. Causal follow-up to C19: build mean-difference (ActAdd) directions for the first op from C19's cached activations and add them back to the residual stream during generation (forward hook). INERT — at depth 1 (cleanest direction, probe 0.99) steering toward the true op never beats baseline and only degrades at high strength; at depth 2 a faint predicted-direction whiff (+0.05, within noise of the random control, below the +0.10 pre-reg bar); null at earlier layers (8, 12) and on identification (0.03→0.03). All pre-registered predictions refuted. The latent signal C19 found is readable but not writable into behavior. Strengthens the throughline: test-time interventions (selection C17, steering C20) don't move the wall; only weight edits (banking C18) and tools (C12) do. Honest limit: a clean negative for standard ActAdd — patching / optimized vectors untested.
- qwen35_4b_latent_composition_probe (claim C19): first look INSIDE the wall. Linear probes on residual-stream activations (last identification-prompt token, all 33 layers, 1500 verified tasks) decode the composition's first operation. The wall's nature changes with depth: depth-1 first-op is decoded at 0.99 (rises to ~0.99 by layer 15) while the model names it 0.44 / generates it 0.68 → representation ≫ expression = latent capability; depth-2 probe 0.42 vs behavior ~0.13; depth-3 (the wall) probe 0.27 but the shuffled floor is 0.14, so the real signal (~0.13) ≈ behavior — the representation itself has thinned to a thread. So the wall is an EXPRESSION failure when shallow (info present, unexpressed) and a REPRESENTATION failure when deep (info not computed). Layer-0 stays at chance (signal is computed, not surface). Implication: steering has headroom at depth 1–2 but almost nothing to steer toward at the deep wall; explains why banking (C18) was necessary — it installs the representation the base lacks. Only proposal-installation, not test-time readout, crosses the deep wall.
- qwen35_4b_coverage_banking (claim C18): banking self-verified solutions does BOTH — concentrates AND expands. The correctly-aimed follow-through to C17 (only shifting the proposal distribution can beat sample-more). Harvest the fixed 4B's OWN execution-verified identification solutions (80 SFT pairs, no teacher), QLoRA-SFT single-shot, eval on DISJOINT held-out tasks (4 arms, base/banked × no-think/think). Depth 1: CONCENTRATION — think greedy@1 0.60→0.80, ceiling flat. Depth 2: EXPANSION — banked coverage@16 0.15→0.45 (3×) on held-out tasks: proposes correct compositions the base never sampled (unique-program count even drops — the proposal mass moved onto correct programs, C17's lever working). Depth 3–4: no move (7 / 0 training examples; wall holds). Bounded: doesn't beat think sample-more at k=1, but banking+sample-more > base+sample-more. To push the wall deeper you need verified deep examples plain sampling can't harvest → seed with tool-search (C12). Refuted its own concentration-only prediction (P3) in the optimistic direction.
- qwen35_4b_coverage_vs_selection (claim C17): the generation wall is COVERAGE, not selection. Pre-registered decomposition — draw K=32 identification samples/task (list + register, depths 1–4, 8 visible + 8 hidden examples), grade vs visible+hidden, compare selectors to the coverage ceiling. Selection is free: max(coverage − vfilter) = 0.00 in every cell; 90% of visible-passers also pass hidden, so an 8-example execution-filter, the model's own C10-verifier, and even a random pick among visible-consistent candidates all recover the full coverage ceiling identically. Single-shot undersells 2–5× (first@1→cov@32: list d2 0.10→0.30, register d2 0.15→0.60, d3 0.05→0.25) and sample+filter recovers it — but that IS sample-more. The coverage wall's depth is set by hypothesis-space size (list collapses at d3; register survives to d4 via a smaller op menu), mechanistically explaining C16's register floor as coverage-driven. Implication: you cannot beat sample-more by better selection — the lever is shifting the PROPOSAL distribution (C12 tool-search / C11-C12 banking). Refuted its own selection-centric predictions (P3, P4). Residual: overfit traps (visible-pass, hidden-fail) false-deploy at deep register — an abstention gap no example-filter catches.
- qwen35_4b_crossfamily_laws (claim C16): cross-substrate generality test of the C13–C15 ladder on two genuinely different fresh, execution-verified, collapse-rejected families (STRING char-edits, REGISTER 3-int machine) vs the LIST anchor, 100 verified tasks/family. Verdict SCOPED, and the split is the finding: two rungs are model-level LAWS — transcription/compiler (plan-given execution ~1.00 at every depth in all three families; the curves collapse to one line) and the generation wall (bare identification collapses with depth everywhere; trans−ident gap ≥0.84 at depth≥3) — so tools identify, the model compiles is substrate-general. But simulation fidelity is substrate-dependent (C15's decay constant was list-specific): register (compact state) is robust ~flat (0.92→0.72), list decays (1.00→0.56), string is floored near-zero (0.24→0.00). New sub-law: the wall's floor ≈ f(hypothesis-space size, simulability) — register alone (small op-menu + simulable) keeps a nonzero deep-ident floor (0.16/0.08). Promotes C13, narrows C15. (Caught a spurious string-sim-0.00 "law" — a quote-blind parser — before any scored run.)
qwen35_4b_depth_wall_anatomy (claim C13): pre-registered anatomy of the compositional wall. It is identification, not execution — plan-given execution 0.90–1.00 through depth 4 (zero execution deficit) while bare identification runs at ~2× over chance per composed op (odds fall ~30×/op; wall at depth 2), insensitive to op type, and barely helped by shown intermediates (segmentation deficit). Retro-explains C10/C11/C12 with one mechanism: the fixed 4B is a reliable compiler starved of hypothesis search — tools identify, the model compiles. Also: 40% of nominal depth-3 tasks were shallower-equivalent (min-depth audit; C12 corrected).
- qwen35_4b_decompose_compose_frontier (claim C12): the frontier is extendable without a teacher. A decompose-and-compose search (4B ranks next primitive → interpreter executes → recurse) cracks depth-3 monolithic sampling can't (0.125→0.40+, 3.4×); against the brute-force bar the model's guidance buys efficiency not coverage (planner-wall). And BANKING the search-found solutions (QLoRA-SFT, no teacher) extends the frontier into the weights (monolithic pass@5 0.125→0.237, depth-3 4×) — the bound M4 couldn't break. Answer to C11's open problem: tool-augmented search harvests frontier-exceeding solutions, banking pulls them into the model's distribution.
- qwen35_4b_neurosymbolic_repl_substrate (claim C11): a fresh, contamination-free procedural program-synthesis substrate (random primitive compositions, held-out-execution graded, oracle-solvable 100%) — a reusable asset for elicitation claims with no memorization confound. On it, a neurosymbolic execution-feedback REPL loop did NOT beat matched-compute sampling (M2), but self-training on the 4B's own verified solutions banked capability into held-out single-shot (M3, +0.095). Extends C1 (executable intermediates) toward execution feedback and self-correction training.
Current Read
Structured execution is one of the strongest imported signals. The next useful work is not another isolated positive run; it is controlled comparison of representations and supervision sources — and, per C11, on a CONTAMINATION-FREE substrate so that gains (and non-gains) are honestly measurable.
Replicated semantic commit interface (2026-07-12)
qwen35_4b_commit_slot_semantic_power_replication establishes a stable constrained compiler output seam on a fresh procedural depth-two substrate. At fixed cap 1,024, ordered thought independently beat an identical token-multiset shuffle by +13.57pp and +15.04pp across two disjoint 113-task stages, while the unrestricted next token was already an alias on 88.2% and 87.6% of rows. Syntax therefore exposes rather than manufactures most of the answer-mode state. The effect remains target heterogeneous and free-form close-only output remains poor. The subsequent task-held-out shared J-value measurement failed at chance (0.5021), below slot margin (0.5448) and equal- width non-J residual state (0.5292); midpoint and endpoint J rankings had opposite signs (0.6083 versus 0.3958). The interface is a stable output seam, not evidence for one phase-invariant scalar compiler state. A permitted phase-specific refit also failed: midpoint J 0.5375 versus non-J 0.6000 and margin 0.5396. Causal work stayed sealed and the midpoint-J successor is retired.
The next generative attempt also stops before compiler evaluation. qwen35_4b_jacobian_counterfactual_branching finds chance supplied-alias control (4/48 at every alpha) for balanced additive J edits at a 512-token native prefix, despite valid non-J geometry. The fixed answer slot is a semantic output seam, not evidence that an arbitrary preceding token is a writable compiler register. Explicit anchor tokens and donor- coordinate replacement are the remaining context-local hypothesis.
Fresh materialized-residual mechanics result (2026-07-13)
qwen35_4b_materialized_residual_sibling_search_fresh_replication finally produced the durable model evidence the parent incident lacked. All nine invocation transactions and 1,984 rows authenticated under a separately published lock, and three hostile result audits reproduced every score and gate. The result splits at the interface boundary. Generation is terminal MECHANICS_INTERFACE_INVALID: materialized/name/shuffled/echo suffixes parsed 12/7/12/20 of 52, every thought reached 512, and cap contact was 37/42/40/28; direct parsed 7/24 with 17 cap contacts. Materialized/name/shuffled/direct had zero visible successes, but this is only weak negative mechanism evidence because the registered ABI failed. The parse-immune cheap viability ranker is a clean negative: materialized recall@4 0.257 was below name-only 0.281, shuffled 0.323, listwise 0.271, and surface 0.375; its +0.149 gain over realized random narrowly missed +0.15 and it failed all absolute support/floor gates. Qualification and confirmation stayed sealed. Retire this cheap ranker and top-four branch; any remaining residual-generation test must first freeze an echo-qualified answer seam on separate calibration tasks.
Answer-seam factorial result (2026-07-14)
qwen35_4b_materialized_residual_answer_seam_factorial then tested that prerequisite under a reviewed, committed-green calibration lock. All five transactions and 240 outputs authenticated, but every one of the four think/no-think x freeform/PROGRAM: arms scored 0/48 strict parse and exact echo, yielding terminal NO_VALID_RESIDUAL_ANSWER_SEAM; mechanics and all protected reads stayed sealed. The failure is diagnostic rather than a broad copying miss. After the decision, removing only the exact terminal <|im_end|>\n suffix let the frozen parser accept 48/48 rows in both no-think arms, but only 38/48 think/PROGRAM: and 24/48 think/freeform because ten/five thinking rows had another close boundary. A looser expected-tail diagnostic was 48/48 for think/PROGRAM: and 29/48 think/freeform, but is not exact output. Sampled tokens showed tokenizer EOS 248046, newline 198, then registered HF EOS 248044. This does not repair the gate. It isolates a fresh next experiment: register first tokenizer EOS as the answer-stage deployment commit boundary, use fresh identities, retain strict pre-commit grammar and an HF-EOS control, and stop if that interface does not independently qualify.
Tokenizer-EOS answer-commit calibration (2026-07-14)
qwen35_4b_tokenizer_eos_answer_commit_factorial isolated the boundary cause on 48 fresh known-answer rows under an independently reviewed, committed-green implementation and lock. Both tokenizer-EOS no-think cells achieved 48/48 strict exactness and parseability with zero cap contacts; every matched HF-model-EOS cell achieved 0/48. All 192 paired boundary comparisons authenticated with shared prompts, seeds, thoughts where applicable, and identical sampled prefixes through the earlier stop. Thinking did not magnify the interface: the structured cell fell to 38/48 and freeform to 30/48, with 16 freeform cap contacts. The frozen decision is TOKENIZER_EOS_ONLY_INTERFACE_QUALIFIED, and the sole advancing winner is tokenizer_eos_no_think_program_slot. This establishes that termination-token identity is a causal part of Qwen3.5-4B's short structured-output interface; it does not establish residual-mechanics or capability gain. Mechanics and all protected labels remained sealed and require a second winner-bound published lock before access.
The later winner-bound run passed its independent transport gate (24/24 exact echo and parse, zero caps) and completed five durable transactions containing 4,056 outputs. It stopped before visible selection: replay authentication of the stored transport decision incorrectly reused the initial temporal gate after later invocations were complete. Hidden scoring and benchmark access remained sealed. Record this as terminal instrument failure, not capability evidence; a fresh-identity successor is required.
Replay-hardened tokenizer-EOS residual result (2026-07-14)
qwen35_4b_tokenizer_eos_residual_mechanics_fresh_replay supplies the clean adjudication. The fresh successor excluded every parent function, identity, prompt/token rendering, and seed domain; passed independent implementation review; and crossed separate calibration, mechanics-lock, and visible-publication CI boundaries. Calibration replicated the interface (48/48 in both no-think tokenizer-EOS cells; all HF-EOS controls 0/48), and transport was 24/24 exact/parse with zero caps. All 4,056 mechanics outputs and the repaired five-stage replay authenticated. Generation itself was healthy: parse was 99.78% direct, 99.13% materialized, 99.83% name-only, and 98.78% shuffled, with zero cap contacts and no direct-pool exhaustion.
The scientific result is a clean negative. Materialized, name-only, shuffled, sampled-token-matched direct, and logical-model-token-matched direct each had 0/24 deployable success and 0/24 oracle proposal coverage. This is not a selector or task-solvability artifact: exhaustive CPU evaluation of all 13,824 programs found 88 visible-consistent, hidden-correct programs, 1--9 on every task, for 24/24 coverage. Under this depth-three DSL and frozen answer seam, all-candidate semantic materialization did not move correct residual programs into Qwen3.5-4B's proposal support. Retire this prompt mechanism; future residual/Jacobian work must change the proposal-formation mechanism and show a forward causal coverage gain over matched sampling, not merely a readable coordinate, a looser selector, or more of the same prompt.
Materialized residual mechanics incident (2026-07-13)
qwen35_4b_materialized_residual_sibling_search does not update scientific belief. Its fresh exact-depth-three construction and controls passed model-free review, but attempt 1 stopped in cache preflight and attempt 2 lost 52 returned mechanics rows before durable persistence because a post-generation receipt collapsed model EOS 248044 and tokenizer EOS 248046. Independent review blocks replay of the terminal STARTED transaction. The materialized-residual hypothesis remains untested; only a new experiment with fresh identities/seeds and write-before-authentication quarantine may resume it. No claim is allocated.
Scorecard
- Program: charter
- Current read: structured intermediates remain strong, but tested latent interfaces usually fail before causal use. The replicated
First:seam is neither a stable J value coordinate nor an arbitrary writable last-thought register. Early concrete text demonstrates broad local routing—84/96 execution for both correct and deranged supplied operations across all 24 operations—but its interface failed and complete-program reachability was only 3/8. The durable fresh materialized-residual run split the question: free generation was interface-invalid, while the parse-immune cheap materialized ranker lost to every structured comparator. Tokenizer EOS plus no-think is a real short-output interface: two independent calibrations reached 48/48 in both prefix cells against 0/48 for every HF-EOS control, and fresh transport was 24/24. The replay-hardened mechanics result now cleanly rejects the prompt mechanism: all 4,056 outputs authenticated with 98.78--99.83% parse and zero caps, yet materialized and all matched controls had 0/24 selected success and 0/24 oracle proposal coverage while exhaustive search covered 24/24 tasks. A strict answer seam does not expose residual synthesis. Cheap viability/top-four ranking, parser repair, and all-candidate semantic-materialization prompting are retired. The fresh three-seed rank-32 LoRA joint adjudication also missed registered state in all 57 required cells (best 0.0234 versus 0.40); because adaptation contrast is uncertain, mandatory LoRA-state-only and matched direct-full-shape controls—not a rank conclusion—come next. - Best next experiments: complete the authorized state-formation Stage B with three LoRA state-only and three matched direct-full-shape joint seeds; separately measure the exact depth-6 resource crossover before learned pruning; and, for J-space, require a task-held-out correctness coordinate beyond margin/position/equal-width non-J controls before a separately preregistered same-prefix intervention must causally increase correct-proposal coverage over matched sampling.
- Strong anchors:
qwen35_4b_early_text_hypothesis_forking,qwen35_4b_partial_structure_search,qwen35_4b_crosssubstrate_structure,qwen35_4b_structure_search_scaling,qwen35_4b_commit_slot_semantic_power_replication,qwen35_4b_state_carry_vs_state_bag_fullrank_delta. - Avoid repeating: another opaque name/timing variant, this cheap materialized P(viable)/top-four ranker, the all-candidate semantic-materialization prompt with a looser selector or more samples, waiting for HF model EOS when tokenizer EOS is the registered answer commit, adding thinking to a short exact-output channel without a task-specific benefit, larger cap or parser relaxation after ABI failure, pooled-AUROC launch, model-guided search without measured brute wall time, or a rank comparison whose shared modules and dropout streams are shifted by parameter-construction RNG.
- Evidence that advances the program: causal ablation showing which intermediate structure transfers across family, length, or paraphrase shifts.
Charter
Show charter.md
Purpose
Discover when small models should stop emitting direct answers and instead emit, condition on, or be supervised through executable structure: typed slots, latent registers, bytecode, candidate programs, state traces, compiler heads, or differentiable runtimes.
Why This Is A Program
The imported experiments contain many successful variants, but the line is much broader than those prototypes. Future work should compare representations, supervision density, curricula, model sizes, and task substrates under a shared question: what structure turns small-model brittleness into reliable execution?
Progress Signals
- Longer-horizon execution improves without beam search or oracle selection.
- Paraphrase and distribution-shift splits preserve the same executable trace.
- Program/state supervision explains gains beyond raw data scale.
- Failures can be localized by operator, step, state prefix, or representation.
Boundaries
This program studies representation and execution. Selection between many candidates belongs primarily to Evidence-Conditioned Selection. Reusable banks of primitives belong primarily to Operator And Skill Inventories.
Backlog
Show backlog.md
Next Experiments
- The LoRA architecture counterfactual completed with valid
PILOT_MECHANISM_MISS; the held-fixed, zero-initialized 892M-parameter full-rank successor also completed, with raw joint-state accuracy 0.00277 versus the 0.40 gate, Carry minus Bag -0.0156, negative unseen-K scaling, and noncausal swaps. Its analyzer emittedPILOT_STATE_FORMATION_MISS, but post-result preregistration audit assignsPILOT_PROMOTION_BLOCKED: the run simultaneously failed the non-capacity requirements of positive Carry-minus-Bag and positive query-kind effects. It therefore does not isolate or close the rank/capacity question. Do not run either existing experiment's confirmation, edge-cut, or sample-more stages, and do not open the interface successor. - Continue the mandatory fresh RNG-matched state-formation adjudication. Its three-seed rank-32 LoRA joint arm validly emitted
LORA_JOINT_MISS_CONTROLS_REQUIRED: 0/57 required cells passed 0.40 and the best intact cell was 0.0234375, while the setup path, exact K=1 bypass, and positive controls passed. Execute the now-authorized Stage B exactly: seed-matched full-rank G0/positive controls, all three LoRA state-only fixed finals, and all three direct-full-shape joint fixed finals on the same rows, initialization, stochastic streams, and readouts. Do not attribute the miss to LoRA, open sealed contrast, or pivot representation/supervision until the registered Stage-B receipt. The first full-rank G0 exposed an immutable downstream path-consumer defect before model load; its exact-prefix recovery, archive, and retirement checkpoints are green. The recovered full-rank seed-7411 G0 now passes all registered feasibility gates, but that frozen wrapper's pathname-only retirement guard cannot hand off to the positive control after the successful G0 repopulates the canonical slot. The additive byte/status-aware handoff is published/green, and the full-rank all three full-rank G0/control pairs now pass independently: each intact control is 48/48 versus adaptation-disabled 0/48 after the exact 256-update, 4,096-presentation schedule. Publish the complete setup matrix, then run the three LoRA state-only and three full-rank joint fixed finals. This remains mechanics/setup evidence and does not authorize a LoRA rank conclusion. - Cross-program interface probe completed:
qwen35_4b_commit_slot_jacobian_value_transportshowed that a fixed latent answer slot repairs formatting but its semantic hint remains task/alias concentrated (15/48 real versus 11/48 shuffled at 1,024; five mixed tasks versus six required). A powered fresh replication must clear task-level uncertainty before the interface can license any J-space value work. - Powered cross-program replication completed:
qwen35_4b_commit_slot_semantic_power_replicationfixes the same slot/cap on 113+113 fresh tasks and requires both task-level ordered-content evidence and eight-alias semantic breadth before treating the latent interface as usable. Qualification passed with correct support across all 11 targets and choices across all 12 aliases; confirmation independently passed with 98/339 ordered versus 47/339 shuffled and support across 10/11 targets and all 12 choices. The licensed shared J-value measurement was then chance (0.502) and weaker than slot margin/non-J state, with midpoint/end phase reversal. Treat the slot as a stable output seam, not a stable scalar compiler-state coordinate; causal work is not licensed. A post-decision midpoint-only refit reached only 0.538 versus matched non-J 0.600, so do not replicate this J-axis interpretation. - Additive J directions do not turn the last native-thought token into a hypothesis register. The late explicit anchor was invalid, and early concrete text redirected direct execution across 24 operations but failed its interface. The first materialized-residual successor then ended without a durable, authenticated model result: one preflight abort and one terminal 52-row write-order/EOS-receipt incident. Retire opaque-name timing variants. Resume materialized residuals only in a new experiment with fresh tasks/ record IDs/seeds, write-before-authentication quarantine, the exact model/tokenizer EOS pair, and the already frozen candidate-blind, shuffled, exhaustive, and sampled/ logical-token-matched controls. That identity is now reserved as
qwen35_4b_materialized_residual_sibling_search_fresh_replication; fresh construction and two adversarial reviews passed with zero required parent identity/prompt intersections and zero model calls. The crash-safe mechanics implementation then passed three independent audits and model-free preparation with 1,984 requests, 676 unique identities, zero in all nine parent/terminal intersections, and zero model calls. Its first lock attempt failed closed on an incorrect historical source commit before any lock/raw/ model activity. A three-reviewer append-only V2 repair preserves every V1 payload byte and again records zero model calls. The audited live successor then completed all nine transactions. Generation was terminalMECHANICS_INTERFACE_INVALID: all thoughts hit cap and every arm missed the parse/cap ABI by a wide margin, so it cannot adjudicate residualization. The parse-immune cheap materialized ranker was a clean negative: recall@4 0.257, below name-only 0.281, shuffled 0.323, listwise 0.271, and surface 0.375. Retire that ranker and top-four branch. The fresh echo-gated answer-seam factorial has now completed under a reviewed lock: all four arms scored 0/48 strict parse/exact echo and mechanics stayed sealed. A post-decision diagnostic isolated the mismatch cleanly only for no-think: both no-think arms became 48/48 frozen-parser exact after terminal<|im_end|>\nremoval; think/PROGRAM:and think/freeform were only 38/48 and 24/48 because of extra close boundaries. Do not relax this result's parser. If residual generation continues, create one fresh answer-stage commit-boundary successor that stops at first tokenizer EOS, trims only that terminal token, retains current HF EOS as a control, uses new tasks/IDs/seeds, tests malformed/extra pre-commit bytes, and must independently clear the same >=90% echo/parse and <=5% cap gates before fresh transport and disjoint mechanics. That successor is now active asqwen35_4b_tokenizer_eos_answer_commit_factorial; its first-stop and strict- precommit calibration has now passed under a committed-green reviewed lock. Both tokenizer-EOS no-think cells were 48/48 exact/parse with zero cap contacts, while all matched HF-model-EOS controls were 0/48; thinking fell to 38/48 in the structured cell and 30/48 in freeform. Only the frozentokenizer_eos_no_think_program_slotwinner advanced from calibration under a second committed-green lock. This qualified an interface only; the residual-capability question remained unadjudicated. The winner-bound run subsequently passed transport 24/24 and durably completed all five transactions, but visible analysis failed because post-chain transport-decision replay reused the initial later-absent authorization gate. No hidden read occurred. Replace it with a fresh-identity successor that separates initial authorization from immutable replay under a new review and lock; do not rerun the result-bearing experiment. That successor is now registered asqwen35_4b_tokenizer_eos_residual_mechanics_fresh_replay, with fresh function fingerprints, identities, prompts/token sequences, seeds, ciphertext/key, and an enforced parent-sampled-bundle denylist. It completed cleanly: calibration again qualified no-think tokenizer EOS, transport was 24/24, every generation arm passed ABI, and all 4,056 outputs authenticated. Materialized and all four frozen comparators nevertheless had 0/24 selected success and 0/24 oracle proposal coverage, while exhaustive search covered 24/24 tasks. Retire the all-candidate semantic-materialization prompt branch. Do not spend another run on selector repair, parser relaxation, or a larger matched budget for the same mechanism. A J-space successor must first show a task-held-out correctness coordinate beyond margin/position/equal-width non-J controls, then causally increase correct-proposal coverage at the same prefix over matched sampling; stop at measurement if that forward gate fails. - Measure the exact behavioral quotient at fresh depth 6 before assuming model-guided pruning is economically needed; record wall time, memory, physical transitions, coverage, and selector success.
- If a real search wall appears, test a residualized partial state (feasible parameter domains, materialized prefix outputs, per-example target residuals) behind the same within-task AUROC and recall@beam gate.
- Repair the demonstrated visible-only selector gap over exact solver pools (60/60 coverage versus 56/60 selected) with frozen stability/simplicity/unlabeled-probe rules.
- Replicate the strongest structural compiler results across seeds, lengths, and operator mixes.
- Run a direct-text-program versus typed-bytecode versus latent-slot comparison on one shared task suite.
- Add adversarial paraphrase and compositional splits where direct prompt cues fail.
- Measure whether state-prefix supervision, final-answer supervision, or program-token supervision is the causal lift.
- Build a small diagnostic suite that every compiler-style experiment can run before claiming generalization.
Required Controls
- Direct answer baseline.
- Same model and data with unstructured output.
- Shuffled or corrupted state supervision when state traces are used.
- Length and family holdouts.
Stop Conditions
Retire a variant when it improves train or IID accuracy but cannot survive harder length/family/paraphrase splits after two controlled attempts.
Type-only absolute P(viable) at think@256 is retired as a search controller unless a materially richer state or interface first clears calibration; pooled AUROC alone must not reopen it.
Experiments 212
- 2026-07-27 → 28 Self-Written Verifier Fidelity
Fill this after the run. Separate deployable evidence from oracle/hidden evaluation.
- 2026-07-25 Real-Repo Agentic Instrument
Yes: 200 tasks over 15 libraries, each one verified two ways - the library's suite passes untouched, and deleting the target function actually breaks specific named tests. Three design traps had to be avoided, and each w
- 2026-07-19 Qwen35 4B Agentic RLVR Feasibility
Yes to the machinery: with three fixes (force eager mode so the hybrid architecture doesn't hang, load the model in bf16 not fp32, and split GPU memory 55% to the inference engine with sleep-mode) the loop CLOSES - a mov
- 2026-07-18 Qwen35 4B WHY-Comment Install
The WHY idea paid off where it should — on writing correct functions — and did nothing where it shouldn't. We trained the 4B to write code with the causal reason for each line attached as an inline #WHY: comment (generat
- 2026-07-18 Qwen35 4B Self-Repair Install
The first bet that moved the needle at all — gently. We taught the 4B to debug by training on 504 examples of [buggy code + the real test-failure message] -> [diagnosis + fix], all self-generated by injecting bugs into c
- 2026-07-18 Qwen35 4B Repair + Why Stack
The obvious way to combine our two promising curricula - just train one model on both - backfired. We put the 504 self-repair rows and 504 WHY-comment rows into one 1008-row training set and trained a single adapter. Ins
- 2026-07-18 → Qwen35 4B WHY-Think Scale
This phase built and proved the machinery; the GPU training sweep has not run yet. The generator now emits, for every example, a real hidden reasoning trace generated mechanically from the program's shape - parse the spe
- 2026-07-18 → Qwen35 4B WHY Scale Ladder
This phase built and proved the machinery; the GPU training sweep has not run yet. The core blocker was that the original WHY generator saturated fast (about 75 distinct reasons, 438 distinct programs at 504 examples), s
- 2026-07-17 Qwen35 4b State Track Confirmation
The lift held up directionally, without becoming a slam dunk. Across six fresh sealed exams, running the same seed through both models so the noise cancels, the state-tracking model beat its parent on 4 of 6 with an aver
- 2026-07-17 Qwen35 4B Exec-Trace Install
The first bet at installing coding cognition came back flat. We trained the 4B on 400 self-generated, execution-verified program traces to install an accurate 'mental interpreter,' the idea being that a model that can si
- 2026-07-17 Count-Walk Menders Confirmation
The answer the rule was built to force out, delivered without wiggle room. Across the four fresh exams the trained model solved a fix-the-procedure episode exactly once (plus one partial credit that the rules pre-declare
- 2026-07-17 Coding Fitness Harness (cognitive-core program)
Surprisingly good at writing single functions (HumanEval 76.2%, MBPP 56.5%) but weak at driving a multi-step coding task in a real agent loop (23%). The harness is validated: it agrees with an independent run to within 1
- 2026-07-17 → State-Track Installation (Stage 9)
The believed-unlikelier but hoped-for outcome landed. After the reliable 'just replay again' lever hit its ceiling, this tried a genuinely different lever: teach the model one new, universal skill — keeping a running tal
- 2026-07-16 Qwen35 4b Zero Root Lineage Rebuild
["The rebuild answered the provenance question with numbers. Retracing the six documented training steps from a truly blank starting adapter — same datasets, same seeds, same settings — produced a model with about ninety
- 2026-07-16 Sweep-Rate Consolidation (Erratum: 2/6, not ~50%)
Two sweeps in six readings — one in three, not one in two — with a wide honest confidence band (roughly 4 to 78 percent at 95%). The texture matters more than the point estimate: the model never lost a single family to t
- 2026-07-16 Clean-Path Statechain Extension
["Split verdict with a clean lesson. The state-tracking dose installed for the THIRD time on its third different parent — the program's most reliable trained effect — and the clean-lineage model beat the untouched base b
- 2026-07-16 Clean Gym-Mix Dose
The mix failed cleanly and instructively. Splitting the standard 160-lesson budget across three skills — sixty trick-instruction episodes, fifty procedure chains, fifty answer-or-abstain puzzles — taught none of them: on
- 2026-07-15 Statechain-Only Dose
Three results in one event. First, the state-tracking dose passed its local gate cleanly — the skill installed again and this time forgetting stayed inside the calibrated margin. Second, on the real benchmark the trained
- 2026-07-15 Goal-Gate Confirmation
The replication returned a split answer. The overall improvement replicated without drama: the trained model beat the base decisively on all three fresh seeds, making four for four all-time. The perfect ten-family sweep
- 2026-07-15 Feedback-Loop + State-Chain Install
The dose split down the middle. The hidden-state-tracking half installed cleanly — on brand-new test instances the trained model tracked procedures better than both its parent and a matched control. The feedback-repair h
- 2026-07-15 Dose-Diversity Mechanism Cell
Variety does not protect memory: the larger varied practice set also cost nine retained answers, the known ten-point case reproduced exactly, and even pure review cost five on this fresh screen — so the one past case wit
- 2026-07-15 Axis Corpus V2 with Staged Repair
Even with lessons rebuilt from the autopsy — walking through the search step by step instead of asserting the answer — the repair skills did not stick (a tie and a loss against the comparisons), so the preregistered kill
- 2026-07-14 → 15 Goal-Gap Axis Curriculum
The targeted practice worked on its own terms — the first screen pass in this program's history (28 vs 22 and 18 of 40 on unseen tasks, with zero forgetting). On the held-out benchmark it beat the untrained model by a wi
- 2026-07-14 → 15 Qwen3.5-4B Counterfactual Plan Reflection Transfer
Not known yet. The current checkpoint builds and checks fresh list, text, and register puzzles, plus a control where every plan is deliberately assigned to the wrong puzzle. No model has been loaded or trained.
- 2026-07-14 Natural-Language State-Table Universal Curriculum
Not known yet. CPU construction produced 80 truth-checked lessons and two 320-row streams with exactly 286,814 tokens each. All 48 smoke tests pass, but no model has been trained or evaluated.
- 2026-07-14 Search-Scaffold Universal Curriculum
No. On 26 fresh paired cases, the staged-search model tied replay at 16 correct and trailed its parent at 18. All three parsed 23 answers and hit the limit three times; the candidate solved none of four execute/induct ca
- 2026-07-14 Failure-Selected Counterfactual Restart Curriculum
The data-selection stage passed: from 624 fresh attempts, the frozen rule selected 52 clean restarts, exactly four for each of 13 skills. Forty correct accuracy or cap failures and twelve correct-but-too-long cases are n
- 2026-07-14 Qwen3.5-4B Tokenizer-EOS Residual Mechanics Fresh Replay
No result exists yet. The scaffold declares seven freshness controls and eight lifecycle controls, reads none of the parent's sampled bundles, and authorizes no model call until independent review and release gates pass.
- 2026-07-14 Qwen3.5-4B Tokenizer-EOS Answer Commit Factorial
The boundary worked cleanly: both no-thinking chat-end conditions produced 48 correct strict answers out of 48, while every matched later-end condition produced zero. The harder transport check also scored 24 of 24. But
- 2026-07-14 State-Formation Branch Authorization Recovery
Yes for the no-model safety check: all six controls passed, the failed first attempt has a third preserved copy, and the two retry-blocking source paths were retired only after that archive commit passed both checks. The
- 2026-07-14 State-Formation Branch Handoff Recovery
Yes, the narrow handoff passed every safety check, then all three full-size setup controls reached 48 of 48 with their learned path and zero of 48 with that path disabled.
- 2026-07-13 → 14 State-Formation Analysis Recovery
Yes, the narrow recovery check passed. It accepts only the one registered file location, still rejects unsafe shortcuts, and has not yet examined any model result.
- 2026-07-13 Validation-policy counterexample curriculum
The test could not answer the question. Before the newly trained model was allowed to run, the starting model and the equal-size old-lesson control each fixed all 48 practice cases. The required gains were therefore impo
- 2026-07-13 Replay-Anchored Universal Curriculum Continuation
No at this dose. The designed mix passed its fresh synthetic screen but scored about 42% overall, below both the 44% mature policy and the 49% replay-only comparison. It also fell below base on one family. Replay-only wa
- 2026-07-13 Mid-Density Token-Matched Universal Curriculum
The 160-lesson stream improved fresh local accuracy from 17 to 19 of 26 cases, raised readable answers from 18 to 23, and cut answer-limit contacts from nine to three. It still missed the readability and length gates by
- 2026-07-13 Low-Density Token-Matched Universal Curriculum
No. Neither 40 nor 80 abstract lessons passed the fresh local screen. The stronger 80-lesson stream matched the inherited anchor at 54% accuracy and parsed 62% of answers, but still hit the answer cap 10 times; broad eva
- 2026-07-13 Qwen3.5-4B: Installing Universal Features via Designed Synthetic Curricula
Not with either tested training geometry. Continuing from the strong policy raised fresh synthetic accuracy from 50% to 69% and beat the untrained base overall, but it scored 31% aggregate versus 45% for the strong start
- 2026-07-13 State-Formation Capacity Adjudication
Not run yet. The reviewed test starts with three compact-update training runs. If any required state check misses, six matched controls become mandatory; the final three runs open only if the full-size version also misse
- 2026-07-13 Full-Rank Extra-R Delta: State-Carry Versus State-Bag
The run worked mechanically, but it did not settle the question. All 892 million full-size update weights trained and fit comfortably, yet macro task-mean joint-state accuracy was only 0.28% against a 40% requirement. At
- 2026-07-13 Semantic-policy headroom tournament
No rule produced a stable training target after a failed test: every one of 72 cases reached a correct patch. But before seeing test output, the model got zero of 54 ambiguous first patches fully right and never opened t
- 2026-07-13 Qwen3.5-4B Semantic-Anchor Coordinate Branching
This run cannot establish that. The internal edit strongly changed the model's choice among candidate names, but none of 440 consequence outputs began with a valid answer token. Even in a restricted twelve-choice readout
- 2026-07-13 Qwen3.5-4B Materialized Residual Sibling Search Fresh Replication
The generation test cannot answer the question because every reasoning trace reached its length cap and the materialized arm parsed only 12 of 52 outputs, with zero successes. The separate one-token ranking test is a cle
- 2026-07-13 Qwen3.5-4B Materialized Residual Sibling Search
Not yet. The scientific design and every model-free construction check pass, but the model has not run. The frozen test contains 264 fresh functions, and all 38,596 planned prompt renderings fit their assigned context li
- 2026-07-13 Qwen3.5-4B Materialized Residual Answer-Seam Factorial
No registered style qualified: all four scored zero strict parses out of 48, so mechanics stayed sealed. Removing only the final chat-end marker and newline made both no-think styles exact on all 48 rows; thinking still
- 2026-07-13 Qwen3.5-4B Early Text Hypothesis Forking
The experiment design and model-free checks now pass, but the model has not run. Review expanded the first draft from twelve operation names to twenty-four fully specified operations, added two fair late-hint comparisons
- 2026-07-13 Qwen3.5-4B Counterfactual Order-Support Selector
Not reliably. The meaningful-order rule got 43 of 113 puzzles right, clearly better than first attempt or majority vote. But it was only two puzzles ahead of simply choosing the most decisive attempt, and a deliberately
- 2026-07-13 Qwen3.5-4B: The Compute-Optimal Confidence Policy (select + abstain + escalate)
Partly. Picking the answer with the highest single-token P(True) confidence is the best grader-free selector at every compute level, and abstaining below a confidence threshold trades coverage for accuracy cleanly — both
- 2026-07-12 Transaction-invariant recovery curriculum
Partly. Focused training lifted practice repairs from 52% to 82% and preserved the test-and-revise loop. But on different transaction APIs it reached 72% versus 70% before training, far below the required ten-point gain.
- 2026-07-12 Verifier-conditioned recovery banking curriculum
Yes — replaying real failures taught recovery, lifting success from about half of cases to more than nine in ten. But the best version also added a little "explain your fix" text, and that supposedly tiny slice actually
- 2026-07-12 State-Carry Versus State-Bag Counterfactual
The first matched test did not answer that architecture question because the low-rank update failed to learn the required running state. Carrying memory improved overall accuracy by only 4.3 points, with uncertainty span
- 2026-07-12 Qwen3.5-4B Same-Prefix Advantage Routing
Only when it replicates, and here one helper didn't. The test actually ran: the deep specialist's local edge held up on a fresh, separate batch, but the quick specialist looked strong on one batch of stuck points and the
- 2026-07-12 Repository search-compress-bank coding curriculum
No. It became flawless on the six codebase types it trained on, solving every task in exactly four clean steps. But on four never-seen codebase types its success collapsed from 68% to 35% — worse than just running the un
- 2026-07-12 Public-verifier recovery branch tournament
No. On four new problem types, both policies repaired 74%, and their combined best reached only 75%—exactly tied with two action-only attempts. The selector was correctly stopped before scoring.
- 2026-07-12 Payload-capable recovery agent harness
No. It led on the first unseen task block at 71%, but confirmed at 69%, exactly tied with the action-only model and short of the required lead. The external benchmark stayed sealed.
- 2026-07-12 Qwen3.5-4B Pareto Policy Integration
No. Retested on fresh, uncontaminated tasks, the version built for quick work actually lost at quick work by about two points, while the version built for long work won at both quick AND long tasks. One version quietly d
- 2026-07-12 Qwen3.5-4B Jacobian Value Transport
No. Editing one internal direction at a late layer flipped the concept the model said aloud on 18 of 24 tries, up from never, and far beating an ordinary readout-style edit that worked only 1 in 5 times. But when that co
- 2026-07-12 Qwen3.5-4B Deep-Advantage MOPD
No. The deep specialist really was the better teacher on selected states, and copying it worked slightly better than copying the wrong teacher or training on matched non-winning states. But after four rounds the resultin
- 2026-07-12 Qwen3.5-4B Commit-Slot Semantic Power Replication
The reasoning result is real, but the proposed internal value meter failed. Ordered scratch work beat the same words shuffled in two independent stages (about 29% versus 14% on confirmation). Yet a task-held-out model of
- 2026-07-12 Qwen3.5-4B Commit-Slot Jacobian Value Transport
Barely. Forcing the format fixed one problem outright: an allowed word was the model's top choice 85% of the time, versus 4% when it answered freely. But real step-by-step thinking beat the very same thought words scramb
- 2026-07-12 Qwen3.5-4B Balanced-Core Answer-Potential SFT
Not run yet — the reasoning bank is fully built (360 tasks, six competing selection rules staged) but no model has been trained or scored. The built test will fine-tune six copies, each fed reasoning chosen a different w
- 2026-07-11 Entropy-routed think-pivot optimization round 2
No. Gently pulling the better word up at 155 hand-picked wrong-turn moments was cleaner than shoving the bad word down, and it carried real signal, beating a scrambled-label control by about 14 points on fresh coding rep
- 2026-07-11 Interactive policy curriculum: oracle DAgger to execution-reward RL
No — it backfired badly. Success on the practice tasks fell from about 60% to 35%, and on tasks never touched in training from about 69% to 35%. It kept editing fluently but stopped running its checks: the verify-and-com
- 2026-07-10 → 11 Think-block FTPO round 1: outcome-conditioned pivot steering as an agentic install recipe
No. Nudging at those forks made the model worse, not better — success on fresh tasks fell about 4 to 8 percent instead of clearing the 5-point gain hoped for. The tell: feeding it deliberately scrambled labels did nearly
- 2026-07-10 Qwen3.5-4B Verified Macro Invention Long-Context Rerun
No. The model was not short on thinking space, it was stuck repeating itself, and more space only bought longer loops. Handed the blueprint, all 16 tasks finished cleanly. Asked to invent the program from examples alone,
- 2026-07-10 Qwen3.5-4B verified-macro exact CUDA-graph vLLM rerun
No, it just loops harder. At both budgets, all 48 reasoning attempts ran out of thinking room and had to be force-stopped, none finishing on its own. Widening the budget from 49,152 to 61,440 tokens pushed the share stuc
- 2026-07-10 Qwen3.5-4B verified-macro capacity-fit vLLM rerun
No. Two hidden limits bit first. Configuring the server for 64 simultaneous long prompts demanded about 2.4 million tokens of fast memory from a pool holding roughly 1 million, forcing constant re-reading. Cutting to 19
- 2026-07-10 Qwen3.5-4B Partial-Structure Recognition-Guided Search
No. Shown a half-finished program skeleton, the four-billion-parameter model's guess at whether it could still be completed was barely above a coin flip — about 51% correct, where 50% is pure chance — and letting it reas
- 2026-07-10 Qwen3.5-4B Answer-Potential Trace SFT
No. The score carried genuine signal — it tracked what the reasoning actually said, not its length or wording — but its top pick raised fresh-answer success only from about 13% to 20%, short of the useful margin needed t
- 2026-07-09 → 10 Gauntlet round 1: breadth-first agentic expert iteration
Mostly the second. On a blind benchmark the model scored about 14 percent, with six task types near zero, but it had usually reasoned correctly. It simply hit its thinking limit, restarted explaining instead of writing t
- 2026-07-09 Qwen3.5-4B Verified Macro Invention
No. Even handed the finished plan and asked only to re-express it — no problem to actually solve — the model got it exactly right just one time in four, short of the three-in-four bar. Every output looked flawless: valid
- 2026-07-08 → 09 Does the installable hypothesize-and-verify skill move the structure wall?
No. Fine-tuning the model on 1,476 worked guess-and-check traces doubled its success at the two-step depth it practiced on (lists jumped from 37% to 70%), yet did nothing one step deeper: three-step success stayed near 5
- 2026-07-07 → 08 Qwen3.5-4B: Can Confidence Replace the Verifier in the Banking Flywheel?
No. When training on answers checked by actually running the code lifted single-shot accuracy from 8% to 24%, confidence-filtered data — fifteen times purer than a random grab of the model's own outputs — landed right on
- 2026-07-07 Qwen3.5-4B: Does the Confidence Toolkit Survive on Real Code?
Yes, but not the obvious way. Averaging the model's certainty across every token of a program barely beats a plain majority vote among the tries. The real winner: make the model write the code, then answer one yes/no que
- 2026-07-06 Qwen3.5-4B: When Does the Model's Structure Beat Brute Search? (depth-4)
No. Adding one extra dial — a sixteen-times-larger space of combinations — did not flip things. Exhaustive search stayed near-perfect at about 97 percent, while the model's knack for guessing the right combination from m
- 2026-07-06 Qwen3.5-4B: Is the Wall Structure or Values? (skeleton-then-fill)
The steps. Handed the correct sequence of operations, a cheap number-search finished every single task, so the numbers were never the bottleneck. Left alone, the model almost never even lands the right sequence, and cred
- 2026-07-06 Qwen3.5-4B: Externalize the Latent Readout (probe-to-prompt)
Partly. Writing the concrete first step — exact value and all — into the prompt lifted the single-best-guess solve rate on two-step tasks sixfold, from 3% to 19%, where editing the model's internal state did nothing. But
- 2026-07-06 Qwen3.5-4B: Is the Parameter Latent? (probe the full first op)
It splits. The model genuinely computes the KIND of operation inside itself — reading its internal activity names the kind far better than the examples alone do (41% versus 27%, against 6% for blind guessing). But the sp
- 2026-07-06 Qwen3.5-4B: Do the Structure Findings Generalize? (string/register)
The wrong sequence. Across three completely different kinds of programming tasks, a 4-billion-parameter model almost never solved one on its own — 0 to 2 percent — and its misses were wrong-order, not right-order-wrong-n
- 2026-07-06 Qwen3.5-4B: Does Banking Install STRUCTURE?
Yes, then no. Training lifted a 4-billion-parameter model from never proposing the right step-sequence (0%) to getting it right about half the time (51%) on brand-new tasks—a real new skill, not memorized answers. But if
- 2026-07-05 Qwen3.5-4B: Thinking vs the Lookahead Wall
No. Asked to name the first of three moves toward a goal, the model stayed stuck at pure guessing — right about 1 time in 32 — no matter how long it thought, even with 2,048 tokens to think first. Yet recognizing a goal
- 2026-07-05 Qwen3.5-4B Latent Decomposition: be your own tool-search
No, not out of the box. One step from the goal it picks the right operation about eight times better than chance, but on the opening move three steps out it ranks correctly only at chance — recognition, not planning. As
- 2026-07-05 Qwen3.5-4B Depth Scaling & Controls: saturation, data-vs-compute, depth-4
Variety, decisively, and the gains never plateaued. Rereading the same 40 solved examples sixteen times over barely moved the solve rate (9% to 16%, within noise), but replacing them with more distinct examples at the ex
- 2026-07-05 Qwen3.5-4B: Do Banking and Thinking Stack?
It depends on how far the goal is. One move away, the two boosts stack almost perfectly: a plain model picks the right move 27.5% of the time, extra training lifts that to 52.5%, and adding thinking reaches 85% — the exa
- 2026-07-04 Qwen3.5-4B Tool-Seeded Banking: does tool-search + banking cross the depth-3 wall?
Partly. A brute-force search found working three-step programs the model never produces itself, and retraining on them lifted three-step success from a hard zero to 5 of 40 fresh, never-seen tasks when it can reason acro
- 2026-07-04 Qwen3.5-4B Depth-3 Dose-Response: data-limited or representational cap?
Just starved for examples. Feeding a fixed 4-billion-parameter model more search-found solutions lifted its solve rate on fresh three-layer puzzles from 0% to 38% when given sixteen tries — climbing steadily at every dos
- 2026-07-03 Qwen3.5-4B Wall Climbing: does banking shallow composition unlock deeper coverage?
No. Fine-tuning the model on the two-step solutions it could already produce tripled its two-step success on fresh tasks, from 12% to 36%. But its three-step success stayed at exactly zero, unchanged from before: both mo
- 2026-07-03 Qwen3.5-4B Latent Composition Probe: is the wall representational or expressive?
It depends on the number of steps. For a one-step recipe the first step is almost perfectly written on the internal scratch paper (99% readable), yet the model voices it only 44% of the time: it knows but stays silent. F
- 2026-07-03 Qwen3.5-4B Coverage vs Selection: anatomy of the generation wall
It never generates it. Whenever a correct program shows up among 32 tries, running each candidate against eight known examples finds it every single time — the model judging its own work, and even a random pick among sur
- 2026-07-03 Qwen3.5-4B Coverage Banking: does banking shift the proposal distribution?
Yes, but it depends on difficulty. On easy one-step problems it just pulls answers it already knew into its top guess (60% to 80%), gaining no new ground. On harder two-step problems something new happens: on brand-new t
- 2026-07-03 Qwen3.5-4B Activation Steering: is the latent first-op causally usable?
No. A simple reader picks the model's planned first operation out of its internal state almost perfectly — 99% of the time — yet pushing that exact signal back in during generation barely moves what the model does: at be
- 2026-07-02 Qwen3.5-4B Simulation Keystone Repair
No. Training made the model trace a multi-step process almost flawlessly — even on longer chains and steps it never studied, so it learned a genuine skill, not memorized answers. Yet every task that supposedly needs trac
- 2026-07-02 Qwen3.5-4B Depth-Wall Anatomy
It's figuring out the steps. Handed the exact sequence of operations, this 4-billion-parameter model writes correct code almost every time, even four steps deep, with zero execution deficit. Left to infer that sequence f
- 2026-07-02 Qwen3.5-4B Cross-Family Laws
No. Handed the exact steps, this fixed 4-billion-parameter model wrote correct code almost every time across three unrelated task types. But asked to infer the same procedure from example inputs and outputs alone, it fel
- 2026-07-02 Qwen3.5-4B Context Composition
Only when its answer survives. The fine-tuned skill is genuinely the sharpest — 95% correct when the model replies in the required form, beating the untrained model's 83% under the same step-by-step procedure. But the tr
- 2026-07-01 Qwen3.5-4B Decompose-and-Compose Frontier
Yes — but not because the model got smarter. Taking three steps one at a time, with a tool that runs each and shows the result, solves about 2 in 5 versus 1 in 8 in one shot. The catch: blindly trying all 23 operations d
- 2026-06-30 → 07-01 Qwen3.5-4B Neurosymbolic REPL Substrate + Failure Profile
No. Seeing its real error barely helped: the fix-it loop solved 29% of puzzles versus 34% for simply drawing five independent attempts and keeping the best, at equal compute; the error message itself added just two tasks
- 2026-06-30 Qwen3.5-4B Thinking Content vs Compute
It genuinely reasons; the boost is content, not compute. Blank filler of the same length, and the real thinking scrambled into nonsense, both scored like skipping thinking entirely, around 74 to 75 percent. Only coherent
- 2026-06-29 → 30 Qwen3.5-4B Thinking Separability Probe
Yes, but not for the reason you would expect. From one snapshot of internal activity, whether the model's own code is correct is readable well above a coin-flip, about 64 to 76 percent of the time. Thinking first sharpen
- 2026-06-29 → 30 Qwen3.5-4B Generator-Verifier Gap
Only after it thinks. Judging on sight, the model rubber-stamps 91% of its tries as correct while just 77% truly pass — barely a check, mostly agreeing with itself. Given room to reason first, it becomes a real critic: l
- 2026-06-28 Counterfactual Episodic ICL Posttraining
Yes. Untrained, a 4-billion-parameter model solved 23% of real text-transformation tasks perfectly; after this training, 57% — but only when it could see the prompt's examples. Scramble those examples and it fell to 17%,
- 2026-06-28 Counterexample-Guided Ephemeral Program
No. Answering each row directly solved three-quarters of tasks completely, while the best rule-program the method could pick solved only about four in ten — and a perfect picker that peeks at the answers did no better. T
- 2026-06-28 Qwen3.5-4B Tool State Policy LoRA
Surprisingly, yes, and with almost no learning. Trusting the program only when it passes a worked example and disagrees with the quick answer lifted accuracy from 56% to 66%, exactly matching the best any picker could re
- 2026-06-28 Qwen3.5-4B Live Tool DAgger
Only its cost, not its accuracy. Answering directly solved none of twelve unseen tasks; letting the model write and run code recovered two — about one in six — and spoiled nothing it already had right. But even a flawles
- 2026-06-28 Qwen3.5-4B Foofah Program Strategy Portfolio
Yes, but modestly, and which passing program you trust matters more than the programs. Asking directly for the finished table got 42% right. The rule the team locked in, commit only when two programs agree, reached just
- 2026-06-28 Qwen3.5-4B Foofah Adaptive Program Budget Router
Yes, and here is the twist: running all five programs on every task scored lower (56%) than the cheap rule (58%), because blanket spending overwrote one answer the quick pass already had right. The rule fires only when t
- 2026-06-28 Qwen3.5-4B Adaptive Tool Controller
Partly. One structural cue — does the direct answer have fewer columns than the raw data implies? — safely flags the reshaping tasks where a program helps, lifting accuracy from 42% to 50% with zero broken tasks. But it
- imported 2026-07-12 Sampled Query Filter Executor Experiment
Yes. Graded on just one sampled final answer per problem, the model rebuilt the entire set of still-possible number pairs it was never shown, capturing 94 to 98 percent of it, because holding that full set is the cheapes
- imported 2026-07-12 Qwen Slot Repair Distillation
No. A correct program almost always sits one or two edits away — a search that peeks at the answer lifts solve rates from about a quarter to roughly 86%. But the blind helper couldn't pick which edits to make: on reworde
- imported 2026-07-12 Qwen Register Trace Refiner
Rarely. Even a flawless picker that always grabbed the correct edit reached only 37% on plainly worded problems and 7% on reworded ones, because the correct program usually isn't among the roughly 1,300 nearby edits at a
- imported 2026-07-12 Qwen Register-Token Latent Compiler
Only for short chains. Up to twelve steps it builds the correct hidden program about nine times in ten, while stripped-down versions trained on the final answer alone never find the interface and stay at chance. But at t
- imported 2026-07-12 Qwen Register-Token Structured Runtime
Only up to a point. For chains of four to twelve steps the hidden program runs flawlessly, at 100 percent. But at 24 steps exact execution collapses to 25 percent — versus about 1 percent from pure guessing. The catch: e
- imported 2026-07-12 Qwen Python-Shaped Silent Executor
No. The silent hidden steps never cleared single digits: about 8% at best on the simplest programs, 4.5% on unseen longer ones, barely above a zero untrained model and near the 3% you would get by guessing. The tell: scr
- imported 2026-07-12 Qwen Progressive Repair Compiler
Partly. The correct fix sits in the candidate list almost nine times in ten, yet the judge finds it only about half the time, lifting exactly-correct programs from 30% to 49% against the 88% a flawless chooser would reac
- imported 2026-07-12 Qwen Learned Repair Verifier
Partly. Without ever seeing the true answer, the judge lifted correct execution from 30% to 47%. But a checker allowed to peek at the answer key found a correct fix already sitting in the candidate pile 88% of the time,
- imported 2026-07-12 Qwen Latent Beam Program Compiler
Yes, up to a point. It wrote exact programs that a fixed calculator ran perfectly at eight and twelve steps, and its answers matched those programs — so it truly computed rather than guessing, where chance is about one i
- imported 2026-07-12 Qwen Candidate-Trace Verifier
Yes. A small checker that reads each candidate's worked-out steps — never the true answer — lifted correctly-running programs from 30% to 54%, and to 56% when it cross-checks two wordings of one task. The twist: a picker
- imported 2026-07-12 Qwen 3.5 4B Verified Edit Closure
Yes, on the hardest unseen tasks. Testing small edits of the model's own near-miss program and keeping whichever passes the example cases raised fully-correct answers from 39% to 52%, while re-sampling the model and re-r
- imported 2026-07-12 Qwen 3.5 4B Typed Sketch Synthesis
It depends on difficulty. On the hardest problems, sketching the shape and letting a verified search fill the blanks lifted correct fixes from 33% to 78%, and a safe blend of both methods reached 88%. But on easy problem
- imported 2026-07-12 Qwen 3.5 4B Static Bridge Ceiling Breaker
Partly. Folding in just 60 slightly-harder "bridge" examples, a quarter of the training budget, more than doubled success on deeper, never-seen programs, from 20% to 44% fully repaired, with no loss on familiar skills. B
- imported 2026-07-12 Qwen 3.5 4B GraphIR Self Repair
No. The plain one-line formula fully solved 29% of brand-new tasks; the step-by-step diagram managed only 22%, and the fix-it pass recovered part of that gap to 24% — still behind. That pass genuinely works: on randomly
- imported 2026-07-12 Query Filter Executor Experiment
Yes, and that is the surprise. Trained only to name one final answer, the model taught itself to run each instruction in order and track every still-possible pair of the two hidden numbers. Given enough internal steps, i
- imported 2026-07-12 Latent Recurrent Executor Experiment
Yes, but only under strict conditions. Built to hold each running total and trained to hit every intermediate value, the network's exact-answer rate stayed under one percent until its private-step count reached the numbe
- imported 2026-07-12 Joint Register Executor Experiment
Yes, but only with the right memory. When the model held every allowed number-pair together and spent one thinking step per instruction, its confidence in the exactly correct answer set jumped from near-random (about 3%)
- imported 2026-07-12 Dense Supervision Ladder Experiment
It's the feedback. With the model held fixed, training it on only one sampled final answer left it weak; showing it the full odds of every possible answer at every step roughly doubled how often it solved the hardest 24-
- imported 2026-07-12 Dense Latent Query Executor Experiment
Partly. Giving the model enough internal thinking steps to walk through the whole program lifts accuracy sharply, and it clearly beats a model that reads everything in one glance. But the memory stays approximate: at its
- imported 2026-07-12 Belief Filter Executor Experiment
Yes. When the model runs one internal update per instruction, it lands over 91% of its confidence on the exact set of still-possible answers, versus under 14% when it stops before finishing. It even runs programs three t
- imported 2026-07-12 Adaptive Cognitive Kernel
No advantage. The self-rewiring is genuinely doing ordered work: scrambling the operation order collapses its step-by-step accuracy from about 12% to 2%, and switching the rewiring off cripples it. But it never beats a p
- 2026-06-27 Real Transform ABI Gate with Counterexamples
It depends on how messy the data is. For clean, spreadsheet-style pipeline jobs the toolbox covered every one (100%) and held firm even against deliberately tricky examples. For irregular date, ID, and text cleanup, cove
- 2026-06-27 Qwen Recursive Ephemeral Program Induction
Only when you check the rule first. On its own, a model writing and applying a reusable rule solved 40% of tasks perfectly versus 56% for plain row-by-row answering, and it broke six tasks direct answering had solved. Ad
- 2026-06-27 Qwen Real Task ABI Coverage Gate
It depends, and the split is sharp. A frozen kit of reusable office operations, with no training at all, assembled 84% of realistic tasks from stored parts alone, far above the 21% a bare kit managed, and it fully solved
- 2026-06-27 Qwen Public PROSE ABI Gate
No. The frozen toolkit fully solved only 19% of the 309 outside tasks. In 77% of them no recipe fit even the worked examples, so the toolkit lacked that operation entirely; under 4% overfit. Yet a small four-billion-para
- 2026-06-27 Noisy Row Program Crystallizer
No. Keeping the model's direct per-row answers fully solved half of the 40 tasks, while distilling those noisy answers into one fixed rule solved just 22.5% — worse even than a scrambled comparison at 25%. And a flawless
- 2026-06-27 Qwen Disagreement-Probe Program Induction
No. The disagreement quiz picked the same programs whether its judge answers were real, randomly assigned, or skipped entirely — all three landed at 64% of tasks fully solved. The only genuine gain came from a plain caut
- 2026-06-27 Qwen Active Crystallizer Public Gate
No. Using the model's votes to choose a rule worked on 25% of tasks — barely above the 22.5% you get from scrambled, meaningless votes, and it never beat the best rule the candidate pool could offer. The model answered i
- 2026-06-27 Qwen3.5-4B Transform ABI Compiler Pilot
Yes. After a light round of tuning, the model chose a recipe that produced the correct output on all 48 test tasks, matching a perfect answer key and beating the untuned model's 92%. It recovered the harder multi-step ch
- 2026-06-27 Qwen3.5-4B Independent Code ABI Coverage Gate
No, hardly any. The locked toolbox solved only about 14% of brand-new tasks, roughly 1 in 7, versus 37% on the familiar tasks it was shaped around. Reshuffling which tasks are unseen barely moves it, around 18%. And near
- 2026-06-27 Qwen3.5-4B Foofah Selective Program Fallback
Trust the program the moment it reproduces the visible worked examples. Doing that lifted exact-match accuracy from 55% to 62% across 250 table tasks, rescuing 18 answers the direct route got wrong while losing none it g
- 2026-06-27 Qwen3.5-4B Foofah Program Repair Agent
No. Guessing the answer directly won outright, solving 55% of unseen tables versus only 25% for the debugged program. But the program is a useful complement, not a replacement: it rescued 18 tables the direct guess botch
- 2026-06-27 Qwen3.5-4B Foofah Program Ensemble Consensus
No. The simplest rule won: run the first program that passes a single worked example. It solved 52% of tables versus 44% when the model just answered directly, rescuing 23 tables it had otherwise botched while breaking o
- 2026-06-27 Qwen3.5-4B Foofah Ephemeral Program Induction
No. Asking directly reshaped 55% of tables correctly; the write-and-test-a-program route managed just 15%. Even a magic chooser that always picked the right route each time would reach only 59% — four points above asking
- 2026-06-27 Qwen3.5-4B Foofah Direct vs ABI
Just ask the model. Directly generating the reshaped table got 55% of 250 table tasks exactly right, versus only 18% for the fixed-operation converter. The model even nailed 103 reshapes the converter could not even expr
- 2026-06-27 Qwen3.5-4B Code ABI Oracle Coverage Ladder
Yes, mostly. Snapping together at most three verified functions produced a passing solution for 84% of 160 small programming tasks, up from just 13% with a bare-bones library. But the library is a foundation, not a solve
- 2026-06-27 Qwen3.5-4B Code ABI Compiler Heldout Primitive Pilot
No. The block kit could build a working solution for 84% of the problems it was tuned against, but only 14% of brand-new ones — a 70-point collapse. Three fresh random batches of new problems all landed near 18%, so it w
- ~2026-06-27 Qwen 3.5 4B Balanced Discriminative Bridge
An even mix, clearly. Sixty evenly-spread ordinary examples across ten new program types lifted the model's success on unseen hard problems from 60% (with no examples at all) to 99% fully solved. Hand-picking only the tr
- 2026-06-26 Qwen Trace Procedure Depth Stress
Yes. Trained only on single-step tasks, the 4-billion-parameter model wrote six-step procedures that ran correctly 63% of the time once a plain step-follower executed them, versus 0% when the model tried to state the fin
- 2026-06-26 Qwen Program-Only Executable ABI
Yes. Teaching a small model to write a short runnable program lifted brand-new multi-step accuracy from about 44% (just stating an answer) to 73%. And when it wrote out its steps plus an answer, the steps ran correctly 9
- 2026-06-26 Qwen Large ABI Nested Compiler
Two things. Making the library four times larger, from 32 to 128 operations, did not hurt straight-line programs at all: both stayed perfect on 16-step chains. But branching is a separate skill. Models shown only straigh
- 2026-06-26 Qwen Extrapolation-Bound ABI
Yes. A model trained only on procedures up to three steps long reliably writes correct sixteen-step procedures, with accuracy climbing from 61% under single-step training to a perfect 100%. Surprisingly, adding longer tr
- 2026-06-26 Qwen Crystallized Trace ABI Tournament
Barely. On familiar inputs, writing out each step scored 94% versus 92% for answer-only, a two-point edge that cost three to six times more generated text. On genuinely new combinations of steps the model had never seen,
- 2026-06-26 Qwen Constrained ABI Parser
Yes, mostly. On the hardest six-step requests, blocking any invalid step as the model writes lifted correctly-running recipes from 60% to 75%, and it won on all five training runs. It also beat merely re-rolling until va
- 2026-06-26 Qwen Compositional Curriculum ABI
Yes. Adding a few two- and three-step examples lifted correct answers on unseen six-step problems from 72% to 83%, and on eight-step problems from 78% to 89%. But adding only two-step examples did nothing (72% stayed 72%
- 2026-06-26 Qwen3.5-4B Substrate Coverage Ladder
Only when a building-block was hand-carved for each specific problem. A shared library that reshapes old solutions to fit new ones solved zero of the nine stuck tasks, no better than nothing. Custom-written parts solved
- 2026-06-26 Qwen3.5-4B Reliability Exec OPSD Audit
No. All three cheap fixes failed. Ranking programs by the model's own confidence picked working code less often than just taking the first one that passed the visible sample tests (6 of 24 versus 8), and let more bugs sl
- 2026-06-26 Qwen3.5-4B Prefix Value Guided Search
No. Even when the half-finished drafts were graded with perfect knowledge of the hidden answer tests, finishing only the best-graded one solved the same share of problems as plain full-solution writing — 75% either way.
- 2026-06-25 Qwen Tail Repair Stability Critic
No. A correct fix sat among the candidate rewrites for about nine in ten programs, but the trained judge could not tell which rewrite was right from summary statistics alone, so it played safe and edited nothing, staying
- 2026-06-25 Qwen Readable Candidate Verifier
Barely. Training a picker on the model's reading of each fix's pseudocode plus its claimed output raised the share of tasks fixed correctly from 44% to 51%, closing only about 15% of the distance to a perfect picker's 90
- 2026-06-25 Qwen Candidate-Conditioned Trace Verifier
No. Doing nothing already solved about 44% of tasks, and always choosing a correct candidate could reach 90%. Yet no trained picker captured any of that headroom; the best merely tied doing nothing. Left unchecked, the m
- ~2026-06-25 Qwen3.5-4B Oracle-Distilled Semantic Verifier
Yes, but only on home turf. On the problem set it trained on, the trained judge picked a genuinely-correct program 81% of the time, up from 72% for the same model untrained and just 44% for grabbing the first candidate t
- ~2026-06-25 Qwen3.5-4B HumanEval Adaptive Evidence Budget
No. Every strategy landed at the same 16.7% correct-pick rate — running no tests, running all eight, and even a strategy allowed to peek at the right answer to time its stop. The trained model did learn to quit sooner, u
- 2026-06-25 Qwen3.5-4B Diversity-Keyed Coverage Gate
Mostly the second. Of 24 Python problems a 4-billion-parameter model missed on four tries, spending more and more varied sampling recovered 15, lifting the share solved from 70% to nearly 89%. Mixing three creativity set
- 2026-06-24 → 25 Qwen Compiler Multi-Seed Reattribution
No. The best schedule averaged 43% correct on plainly worded problems, yet the identical training swung from total failure to 81% just by changing the random starting number — so no schedule earns credit for the wins. Bo
- 2026-06-24 Qwen VM-ECHO Trace Distillation
Mostly no. The model got far better at predicting execution, with reading a running program's top value climbing from under 1 percent correct to 43 percent, but that rarely improved the programs it wrote. First-try accur
- 2026-06-24 Qwen VM-Agent ECHO QLoRA
Only the acting helped. Editing-and-running in a loop lifted the share of tasks solved from 10% at a blank start to 43%, beating a single one-shot guess near 37%. But adding a second job, predicting the program's output
- 2026-06-24 Qwen Structural Latent Compiler Expansion
Yes, mostly. The model fills fixed slots with operations a calculator runs — no text, no trying many guesses and picking one. After the learned short form was copied into bigger ones, it stayed perfectly correct on 8- an
- 2026-06-24 Qwen Structural Compiler Attribution Ablation
It's the practice schedule. Give the model its full size from the start, then feed examples easy-to-hard — 8 steps, then 16, then 24 — and it solves nearly every standard 24-step program (about 97%). The popular guess, g
- 2026-06-24 Qwen Search-Augmented Rollout Distillation
No. An automatic search found a verified correct fix for 98% of the dead-ends the model wandered into, yet retraining on those single fixes matched or trailed the simpler training on four of five test sets. The model cou
- 2026-06-24 Qwen Recurrent VM Repair Policy
Yes, but only partway. Letting the model run its program, read the output, and fix one line at a time roughly tripled accuracy, from about 11% to 34% on the main test and 16% to 40% on reworded prompts. But a perfect edi
- 2026-06-24 Qwen In-Policy VM-ECHO Distillation
Only partly. It became excellent at predicting mechanical outcomes, like how deep a program runs (about 92% right), but stayed no better than a coin flip at judging which program is actually correct. So it could not rank
- 2026-06-24 Qwen Fuyu VM GRPO-ECHO
No. Copying worked solutions alone solved about 10% of tasks; one round of reward-based self-coaching dropped that to 7-8%. The coaching produced encouraging signals — it could rank good fixes over bad ones 80% of the ti
- 2026-06-24 Qwen Dense-State DAgger VM Agent
Mostly yes. Teaching the model to make one edit at a time, then correcting it on the messes it made, roughly doubled accuracy on mixed tasks (22% to 41%) and won on four of five task types. But it lost on ordinary tasks
- 2026-06-24 Qwen Counterfactual Trace Preference Distillation
Barely. The self-grader learned to favor programs that run without crashing, but it identified the truly correct one only about 15% of the time, against a 41% best-possible ceiling. On fresh questions it picked worse tha
- 2026-06-24 Qwen Action-Conditioned VM-ECHO Policy Iteration
Barely. Learning to grade drafts by their run results nudged picking accuracy only from about 10% to 11%, far short of the 37% reachable by always choosing the best available draft. It reliably picked programs that ran,
- 2026-06-24 Qwen3.5-4B Oracle Process GRPO
Yes, but with sharp limits. Teaching the model to copy a flawless test-picker lifted its success from about a third of puzzles to 43%, versus roughly 31% for picking blindly. Layering on preference-based and reward-based
- 2026-06-24 Qwen3.5-4B Learned Active Trace Policy
It depends. On the main test set a simple even-splitting rule beat the trained picker after one extra input — 91 percent of programs fully correct versus 87 — and the best-possible choice reached 97 percent. The picker d
- 2026-06-24 Qwen3.5-4B Joint Shortlister Ladder
No. Across every version — untrained, trained, and with the glossary's descriptions scrambled — the model got both codes exactly right zero percent of the time, even when allowed sixteen guesses. Training pushed single-c
- 2026-06-24 Qwen3.5-4B Adaptive Evidence Budget Policy
Yes. Trained just for this single decision, the model matched the accuracy of running all ten checks — solving about 92 of every 100 tasks — while using only about five checks on average, near a perfect-hindsight stopper
- 2026-06-24 Qwen3.5-4B Active Counterexample Trace Selection
Yes, but choose them well. Committing on the visible examples alone left about a quarter of picks secretly wrong, even though every one passed all the examples shown. Requesting six new test cases where the surviving pro
- 2026-06-23 Qwen Typed Bytecode Expert Iteration
It depends. Training the model only on its own attempts that landed on the correct final answer lifted unaided first-try accuracy from 62% to 73% on fresh problems — a real gain that sticks when it writes programs alone.
- 2026-06-23 Qwen Semantic Prefix Value Model
No. Scoring each step by whether a correct answer is still reachable pushed the top pick to about 68 percent, level with plain confidence search and short of the 71 percent from scoring steps against the known correct pr
- 2026-06-23 Qwen Prefix-State Process Verifier
Barely. The judge got genuinely good at telling promising partial programs from dead ends, yet it lifted first-try accuracy only a few points, with hard problems going from 41% to 44%. The real gap: on hard problems a co
- 2026-06-23 Qwen On-Policy Repair-to-Compiler
Yes, but the self-correction isn't the secret ingredient. On brand-new requests the model climbed from about 29% correct programs to 99% after training on its own corrected drafts. The catch: training on plain correct pr
- 2026-06-23 Qwen Mixed-Domain Trace Verifier
Yes, partly. On fresh tasks the frozen model alone got 46% right; the proofreader lifted that to 57%, and a cross-check that compares reworded versions of the same task reached 61%. But a correct recipe was already among
- 2026-06-23 Qwen LoRA Typed-Bytecode Trace Compiler
Yes — but the win came from the teaching material, not from adapting the model. Fed fully worked recipes, it wrote a runnable recipe that reached the right answer about 68% of the time, versus only 15% when taught with f
- 2026-06-23 Qwen Iterative Repair Policy
Yes. The frozen model alone got about 30% of programs exactly right; editing one step at a time lifted that to 53% on fresh problems, closing roughly 38% of the distance to the best a perfect fixer could reach (89%). The
- 2026-06-23 Qwen Hidden VM On-Policy Canonical Repair
No. Training the model on automatically corrected recipes reached 61% on new tasks, versus 59% for plain training — a 2-point gap that is basically noise, and it left longer tasks no better. The corrections are genuinely
- 2026-06-23 Qwen Hidden VM Mixed Domains
Yes. Guessing the answer directly worked only about 15% of the time across six kinds of problems — arithmetic, dates, unit conversions, list totals, yes/no thresholds, and lookups. Having the model instead write a hidden
- 2026-06-23 Qwen Hidden VM Curriculum Repair
No. Feeding it nearby worksheets that merely land on the correct answer wrecked it. Plain step-by-step training scored 72% on the main mixed test; the same model after answer-chasing repair fell to 35%, and collapsed on
- 2026-06-23 Qwen Context-Conditioned Trace Verifier
Mostly no. A correct program was almost always sitting in the candidate pile, for 91 to 100 percent of questions, yet the judge barely helped: it nudged easy questions from about 69 to 70 percent and actually made the ha
- 2026-06-23 Qwen Complete-Program Trace Reranker
Barely, and it backfires on hard cases. A correct program sat in the candidate pile 91 to 100 percent of the time, yet the trained picker nudged easy prompts only from 69 to 70 percent and actively hurt longer ones, drop
- 2026-06-23 Qwen Budgeted Action-Value Compiler
Barely. On fresh problems the model drafts a correct program among its candidates 81% of the time but ranks it first only 67% of the time. A learned scorer that never sees the answer nudged that to just 70% — while a con
- 2026-06-22 Qwen Verifier-Guided Slot Repair
Yes, mostly, but with a catch. A small model copying 24-step calculations got only about 27% exactly right on its own. A checker that knows the correct running number after every step, allowed to swap one or two bad step
- 2026-06-22 Qwen Teacher-Distilled Slot Compiler
No. A model trained to copy numbers and operations out of text and run a 24-step calculation got 27% of final answers exactly right; adding the pointing signal landed at 28%, a tie. Worse, agreement between two rewording
- 2026-06-22 Qwen Checkpoint-Selected Scheduled-State Compiler
It depends. A light, steady dose of show-your-work coaching, kept on through the hardest problems, lifted correct answers on 24-step chains from 25% to 33% and nearly doubled agreement between two wordings of the same pr
- 2026-06-22 Qwen 3.5 4B Unsaturated Frontier Active Bridge
Spread evenly. Giving each of ten problem types the same six extra correction examples let the model fully fix 98% of hard cases. Piling those same examples onto whichever types it failed most reached only 85%, and starv
- 2026-06-22 Qwen 3.5 4B Model-In-Loop Counterexamples
No. Building practice cases from the model's actual wrong answers matched, but never beat, simply hand-picking the tricky categories in advance. Both lifted the hardest problems from 64% to a perfect 100% passing every h
- 2026-06-22 Qwen 3.5 4B Executable Program Posttraining
Yes, but with a catch. Shown worked-through reasoning in the prompt, the model fixed unseen problem types about three-quarters of the time, versus one-in-three when the prompt showed no steps. Strip out or scramble those
- 2026-06-22 Qwen 3.5 4B Counterexample-Directed DSL
It depends. Hand-picked examples lifted the model's single-best-guess repair rate from 51% to 58% over random examples. But when it generated several candidates and kept the best, the edge vanished (64% slipped to 61%).
- 2026-06-21 Structured Slot Initializer Ladder Experiment
Structure, not scale. A plain general-purpose network placed only 55.5% of its belief on the correct starting setup and gave barely half the possible values their own slot, doubling several onto the same one. A rule that
- 2026-06-21 Sparse Support Memory Executor Experiment
Only when its scratch memory held one slot for every possible starting value. With that, it answered every question correctly through the longest 24-step programs. Cut the memory roughly in half and accuracy fell to abou
- 2026-06-21 Qwen Trace Bootstrap Retention Experiment
Yes. Once step-by-step labels install the skill, training on final answers alone preserves and even sharpens it: 97% of the longest 24-step problems solved exactly, versus about 1 in 100 — no better than guessing — when
- 2026-06-21 Qwen Structured Bridge Experiment
Yes, but only when you show it the individual steps during training. A tiny translator turning the frozen model's read into calculator instructions solved chains far longer than it trained on: 96% correct at twelve steps
- 2026-06-21 Qwen State-Ladder Compiler
No. Grading the running total after every step never beat an identical model graded only on its final answer, and at full strength it collapsed on the hardest long programs. The real winner was the training schedule: sta
- 2026-06-21 Qwen Span-Free Compiler
Only when it is first taught where to look. Fed just the frozen model's raw internal notes, a plain reader stayed near random guessing (about 1 in 97). Adding training that also highlighted which spots held the numbers a
- 2026-06-21 Qwen Slot-Stability Compiler
Yes, mostly. The point-and-compute helper solved about 91% of short problems and around half of medium ones, while training the same model to just emit the final answer never beat random guessing, about 1 in 60, at any l
- 2026-06-21 Qwen Shared Parser Compiler
Only when every step is taught directly. Given step-by-step labels, the add-on rebuilds short programs well — nearly 4 in 5 four-step problems run exactly right — but accuracy fades to 39% at twelve steps and under 1% at
- 2026-06-21 Qwen Numeric-Copy Compiler
Yes. When the model just points to where each number and operation sits and copies the exact symbols for a hidden calculator to run, it solves four-step problems about 90% of the time. A version trained to write the answ
- 2026-06-21 Qwen LoRA Parser Compiler
Partly. With step-by-step coaching, a small four-billion-parameter model's hidden states became a readable program: it named the starting number every time and picked the right operation about 98 percent of the time, whi
- 2026-06-21 Learned Sparse Slot Executor Experiment
Yes, but only small. With eleven possible values and a ready-made scratchpad, the network kept the fully correct answer in view 95.5% of the time, even on longer chains than it trained on. Widen to thirty-one values and
- 2026-06-21 End-to-End Structured Slot Executor Experiment
Yes, but only when both halves carry built-in structure. On the numbers 0 to 30, the full model puts 98% of its confidence on the exactly correct final set of possibilities, nearly matching a version handed the answer, a
- 2026-06-21 Dense Teacher Distillation Experiment
No. Even with a flawless teacher revealing the exact set of still-possible answers at every step, the fixed-size memory learned only a rough approximation. The best version placed 52% of its confidence on the correct fin
- 2026-06-21 Cyclic Transition Ladder Experiment
It needs the matching wrap-around parts. A network built from clock-arithmetic moves stayed perfectly exact on programs three times longer than it practiced on. A plain generic network of the same size drifted down to ju
- 2026-06-19 → 20 Execution-Conditioned Repair LoRA Experiment
No. On bugs built from the same templates it practiced on, the fixer repaired all 60 of 60 cases, versus 11 of 60 with ordinary patch training and 6 of 60 with no training at all. But on bug types it never saw, every met
Claims
- Confirmed C1 · Structured intermediates improve small-model reliability
- Promising C11 · Self-training on verified self-solutions banks capability; test-time execution feedback does not (contamination-free substrate)
- Promising C12 · The fixed 4B's compositional frontier extends without a teacher via tool-augmented search + banking
- Promising C13 · The fixed 4B's compositional wall is broken multi-step mental simulation; transcription is intact
- Promising C14 · Repairing a broken primitive does not propagate: capability is format-local in the fixed 4B
- Promising C15 · Deployable capability = module x interface x procedure: context composes modules but cannot create generators
- Promising C16 · Cross-substrate: the compiler and the generation-wall are model-level LAWS; simulation fidelity is substrate-dependent (C15's decay constant was list-specific)
- Promising C17 · The generation wall is COVERAGE, not selection: with enough examples selection is free; only shifting the proposal distribution beats sample-more
- Promising C18 · Banking self-verified solutions does BOTH: concentrates coverage into single-shot (easy depth) and EXPANDS the coverage ceiling on held-out tasks (harder depth) -- shifting the proposal distribution, the lever C17 named
- Promising C19 · Inside the wall: the composition is linearly encoded but under-expressed at shallow depth (latent capability) and thins to a thread at the deep wall (information gap) -- the wall's nature changes with depth
- Promising C20 · Decodability != steerability: the latent first-op direction (C19) is readable but adding it back via activation steering does NOT change behavior -- test-time readout cannot elicit the wall's latent capability
- Promising C21 · Self-banking is coverage-seed-bounded: banking installs & expands WITHIN a depth but does NOT climb ACROSS depths -- the wall is not climbable by pure self-training
- Promising C22 · Tool-seeded banking crosses the depth-3 wall self-banking couldn't (validates C21 recipe: tools explore, banking installs) -- but the installer's efficacy DECAYS with depth (crossed-but-weak)
- Promising C23 · The depth-3 install is DATA-LIMITED, not a representational cap: tool-seeded banking scales monotonically with #solutions into DEPLOYABLE single-shot
- Promising C24 · Depth scaling & controls: no saturation through 1280 tool-pairs, the gain is data-DIVERSITY (not compute), and the tool-search+banking recipe repeats one rung deeper (weakly)
- Promising C25 · 'Be your own tool-search': base first-move ranking is at chance; banking improves step-wise next-op guidance at lookahead distance
- Promising C26 · TEST-TIME thinking (on a model never trained to reason about this task) does NOT breach the lookahead wall; it amplifies recognition, not planning
- Open C27 · Banking and TEST-TIME thinking stack additively on RECOGNITION but not on PLANNING (a baseline motivating bank-the-thoughts, NOT a claim that thinking can't help planning)
- Promising C28 · Banking correct decomposition PLANS installs deployable depth-3, but banking the model's OWN rejection-sampled thoughts does NOT (they are rationalizations) -- it is the plan QUALITY, not reasoning-as-such
- Promising C30 · Externalizing the latent readout (decode->prompt) elicits deployable depth-2 where steering (C20) failed -- but the decodable op-TYPE only narrows sampling; the PARAMETER is the deployable bottleneck
- Promising C31 · The op-TYPE is model-computed (latent, elicitable) but the PARAMETER is only read off surface I/O -- sharp localization of C30's deployable bottleneck
- Promising C32 · The compositional wall is STRUCTURE, not values: the model cannot propose the op-sequence (failures are wrong-skeleton), but once structure is known values are trivially searchable
- Promising C33 · Banking installs STRUCTURE: base op-sequence structure-coverage 0.00 -> banked 0.51 (held-out), converting the wall from structure-bound to value-bound
- Promising C34 · With the interpreter, brute-force structure-search + value-fill + execution-consensus near-solves depth-3 (0.975), dominating the model; banking's structure is a forward-pass-only asset
- Promising C35 · Brute search dominates tested banked models through depth 4 and remains operationally cheap at depth 5; learned crossover remains open
- Confirmed C36 · The recent structure findings are MODEL-LEVEL LAWS: wall-is-structure (C32) and brute-search-dominates (C34) replicate on string + register substrates
- Promising C37 · The compositional wall does NOT exist in language: the model chains depth-3+ multi-step SIMULATION in natural language near-perfectly -- the wall is formal-modality-specific, not a general multi-step limit
- Promising C38 · The structure-PROPOSAL wall persists in language: the model executes a given rule but cannot induce one -- proposal/induction is modality-general, while simulation is formal-specific (C37)
- Promising C48 · Hypothesize-and-verify SFT lifts taught depth 2 but not depth 3 at think@1024; cross-depth and cross-substrate transfer are null while the budget curve remains open