Research log Small Model Experimentation
GitHub

Structured Execution and Compilers

Represent tasks as executable, typed, latent, or stateful programs instead of direct final answers.

What we have learned

Seed Experiments

Key Result

  • qwen35_4b_state_carry_vs_state_bag (unclaimed; terminal LoRA PILOT_MECHANISM_MISS): a valid, source-bound seed-7401 pilot tested whether rank-32 extra-call LoRA over two complete Qwen3.5 hybrid motifs could learn a serially inherited query-before-state representation against an equal-parameter/equal-compute Bag. All registered pilot cells and identities were complete, K=1 parity was exact, and the configured +0.05 gate was reachable. Carry nevertheless failed deep state formation: joint node+phase+checksum step accuracy was 0.00459 versus the frozen 0.40 gate and node accuracy was 0.0642. The small matched-depth answer effect (+0.043, pilot CI -0.008 to +0.094), unseen-K gain (+0.012, CI -0.035 to +0.059), and positive joint holdout (+0.051, CI +0.008 to +0.098) do not rescue a chance-like registered state. Swaps were also noncausal (donor-follow gain +0.008, CI -0.023 to +0.039; donor follow minus recipient preserve -0.055). Confirmation, edge cuts, and sample-more were correctly not run. This narrowed the result to the registered low-rank adaptation recipe did not form the required deep joint state and licensed the held-fixed full-rank successor below.

  • qwen35_4b_state_carry_vs_state_bag_fullrank_delta (unclaimed; raw analyzer label PILOT_STATE_FORMATION_MISS, post-result preregistration audit disposition PILOT_PROMOTION_BLOCKED): the preregistered capacity control replaced rank-32 LoRA with 892,272,640 direct FP32 full-rank deltas on the same 62 extra-R linears while holding the parent rows, recurrence, loss, optimizer schedule, seeds, base K=1 path, and Carry/Bag comparison fixed. Live G0 passed with complete Adam state, exact K=1 before/after the optimizer, finite K=12, a bit-exact 3.571 GB checkpoint round trip, and 22.57 GiB reserved headroom. The complete matched pilot nevertheless formed essentially no registered state: joint step accuracy 0.00277 versus 0.40, node accuracy 0.0617. Carry lost to Bag by 0.0156 (CI -0.0664 to +0.0391), unseen-K gain was -0.00781 (CI -0.0625 to +0.0469), and swaps reduced donor following by 0.00781 (CI -0.0391 to +0.0156). The same pilot also failed non-capacity promotion requirements: its Carry-minus-Bag effect was not positive, and neither registered query-kind effect was positive (node 0.000, checksum -0.03125). Although all cells were complete, the answer gate was reachable, and the answer interface was valid, those simultaneous failures mean the state-formation miss was not isolated under the preregistered disposition table. The defensible terminal disposition is therefore PILOT_PROMOTION_BLOCKED, not a closure of LoRA capacity. A fresh RNG-matched three-seed LoRA-versus-full-rank state-formation adjudication is mandatory before drawing a rank conclusion; neither existing pilot is eligible for confirmation, edge cuts, or sample-more.

  • qwen35_4b_state_formation_capacity_adjudication (unclaimed; in-progress producer verdict LORA_JOINT_MISS_CONTROLS_REQUIRED): the fresh adjudication fixes the prior confounds with bit-identical shared state initialization, seed-matched dropout/order, exact K=1 bypass, setup positive controls, three fixed-final 1,500-step seeds, and a source-bound immutable analyzer. The rank-32 LoRA joint arm validly failed state formation: 0/57 required seed×split×depth cells passed 0.40, with maximum intact accuracy 0.0234375 and per-seed maxima 0.0234375/0.015625/0.015625. Trained, deep-extrapolation, and joint-shift categories all missed; adaptation contrast was uncertain. This is stronger evidence that the registered LoRA joint recipe fails than either predecessor, but it still does not identify rank as the cause. The frozen branch now mandates three LoRA state-only controls and three direct-full-shape joint arms before any rank-causal conclusion or sealed contrast.

  • qwen35_4b_commit_slot_jacobian_value_transport (unclaimed; terminal COMMIT_SLOT_SEAM_FAIL): a fixed latent answer interface repaired formatting—an alias was already the unmasked next token on 41/48 at cap 1,024—but semantic correctness remained task/alias concentrated. Real ordered thought was 15/48 versus 11/48 exact-length shuffled, with five mixed tasks versus six required. A structured slot can expose a decision without making that decision reliable; task-level confirmation remains mandatory.

  • qwen35_4b_early_text_hypothesis_forking (unclaimed; terminal INVALID_INTERFACE_PARSE): all 392 locked mechanics rows authenticated. Exact early bound text controlled direct one-operation execution broadly: systematic and length-matched deranged hints each achieved 84/96 execution of their own supplied operation, while deranged produced 0/96 of the registered target; all 24 operations and four contexts had support. This did not become a valid proposal controller. Every diagnostic arm exceeded the 0.05 answer-cap ceiling, duplicate/placebo also failed parse, and the independent noncausal full-program ceiling reached only 3/8 visible passes. Qualification stayed sealed. The next justified representation is a materialized residual state—candidate intermediate outputs plus the remaining target relation—not another opaque name, larger budget, or parser relaxation.

  • qwen35_4b_partial_structure_search (unclaimed while ledger re-grade is open): type-only partial viability is oracle-useful but model-unreadable. A width-4 exact live-prefix beam retained a hidden solver on 12/12 dedicated depth-5 development tasks at 262,144x completed-leaf compression. Frozen Qwen3.5-4B thinking P(viable), however, was chance within task on 7,200 depth-4 children (AUROC 0.506, CI 0.470--0.543; recall@4 0.251) and significantly below no-think AUROC (delta -0.049, CI -0.090 to -0.010). Pooled AUROC 0.557 was a task-difficulty mirage; wrong-task visible examples were no worse. The gate correctly stopped depth-5 model search and banking. Separately, exact visible-only depth-5 brute covered 60/60 and selected 56/60 in 112 seconds on eight CPU workers, so the next question is the real depth-6 resource crossover, then a residualized rather than type-only state.

  • qwen35_4b_crosssubstrate_structure (claim C36): the recent structure findings are MODEL-LEVEL LAWS. C32 (wall-is-structure) + C34 (brute-search dominates) replicate on STRING (char edits) + REGISTER (int machine) + LIST: base ~0, structure-cov = concrete-cov, oracle-skelfill 1.0, random low, brute-deploy ~1.0 on all three. The fixed 4B is a value-computer not a deep-structure-proposer, across substrates.

  • qwen35_4b_structure_search_scaling (claim C35, re-graded): brute-full deploy stays 0.967 at depth 4 (vs 0.975 at depth 3), while the tested banked models' structure coverage is 0.10 and 0.51. The banked comparison crosses non-dose-matched models, so it is not a causal depth curve. Brute dominates the measured list-DSL cells through depth 4; depth 5 was projected, not tested, and is the open model-guided-search regime.

  • qwen35_4b_banking_installs_structure phase 2 (claim C34): end-to-end bank+value-fill deploy. bank-fill deploys 0.463 (= banked structure-cov, confirms C33) BUT brute-force structure enumeration + value-fill + execution-consensus deploys 0.975 (near-solves depth-3) WITHOUT the model. With the interpreter, free structure-search dominates; banking's structure is forward-pass-only. Extends C17 (selection free) to structure-search. Scope: brute wins because the 4096-skeleton space is enumerable.

  • qwen35_4b_banking_installs_structure (claim C33): banking installs STRUCTURE -- base op-sequence structure-coverage 0.00 -> banked 0.51 (held-out depth-3, generalizable). Banking converts the wall from structure-bound (base) to value-bound (banked struct 0.51 > concrete 0.36, value tax +0.15, fillable). Mechanistic closure of C32: banking = structure-installation. Unifies C22-24/C31/C32.

  • qwen35_4b_structure_or_values (claim C32): the compositional wall is STRUCTURE, not values. The model's STRUCTURE-coverage (right op-type sequence, any param) = its concrete coverage (value tax +0.000 at depth-3) -> failures are wrong-skeleton; oracle-skeletonfill=1.0 (values trivial given structure); random-skeletonfill low (DSL not value-fungible). Unifies C19/C25/C31; explains why tool-structure-seeds (C22)+banking were necessary. (op-seq generation fails at 0.00 = separate format handicap.)

  • qwen35_4b_thinking_lookahead (claim C26): TEST-TIME thinking (on a model never trained to reason about this task) does NOT breach the lookahead wall — it amplifies recognition, not planning. (Scope: leaves open whether banking successful reasoning traces would install planning-via-thinking — the clean untested version.) C25 found the fixed 4B can't plan the first of 3 ops in one forward pass. Does thinking (serial test-time compute, the dormant C9 lever) breach it with no training? Channel-matched test (think→RANK vs no-think→RANK, parse-immune), headlined on STEP 1 (the only clean lookahead test). Step-1 stays at chance across budgets (0.025 → 0.050 → 0.075 at B=0/1024/2048; Wilson CIs overlap). But thinking's benefit scales inversely with lookahead distance: step-3 recognition (1 op away) 0.275 → 0.600, step-2 0 → 0.325, step-1 (3 away, real planning) ~flat. So thinking amplifies recognition, not planning — and internal-brute-force is refuted (step-1 would rise if the model could simulate the path; it doesn't). The juxtaposition with C25: banking lifted step-1 lookahead (0.013 → 0.138) while thinking does not — so for the planning gap, training is required; test-time compute alone can't elicit it. Reconciles with C23 (base think single-shot depth-3 = 0). Design hardened by an adversarial review. Limits: closed-set ranking, n=40, budgets ≤ 2048.
  • qwen35_4b_latent_decomposition (claim C25, re-graded): the base next-op ranker is at/below chance only for the first move three operations from the goal; step 2 is weakly above chance and terminal recognition is stronger. Base-guided versus random search solved 1/80 versus 2/80 tasks, so the defensible conclusion is “no better than random,” not “worse.” Banking improved step-wise rankings and low-budget search (18/80 banked versus 2/80 random and 1/80 base), roughly matching brute's 23/80. The dose trend is supported at steps 2–3, not at step 1 (10/80 versus 11/80 for the two banked doses). This is a closed-menu behavioral guidance lift, not demonstrated internal planning; one beam and single adapter seeds bound it.
  • qwen35_4b_depth_scaling_controls (claim C24): three follow-ups to C23 — no saturation, the gain is data-diversity, and the recipe repeats one rung deeper. (1) The depth-3 dose curve does NOT saturate through 1280 tool-pairs (1156 distinct functions): cov@16 climbs 0.00/0.087/0.212/0.375/0.537, deployable greedy@1 → 0.188; distinct functions grow near-linearly so it's real capacity. (2) A 2×2 at matched steps/mixture splits the gain: diversity (up40 0.163 → train_640 0.375, same compute) is cleanly significant; the pure compute effect (N=40 0.087 → up40 0.163, same 40 functions) is within noise — so C23's "data-limited" is data-DIVERSITY-limited. (3) The tool-search+banking recipe repeats one rung deeper, weakly: depth-4 cov@16 base 0.00 → scaffold transfer 0.067 → banked_d4 0.183 (~3×), but test-time-only (greedy flat 0.033) and marginally significant at n=60; no depth-3 forgetting (guardrail 0.425). Design hardened by an adversarial workflow review (scaffold-only baseline, distinct-fn counts, 0-leak d3 0/2305 & d4 0/318, true-depth-4). Limits: single seed, n=60–80 underpowers adjacent-dose/depth-4 significance, depth-4 single dose, 2560 dose dropped.
  • qwen35_4b_depth3_dose_response (claim C23): the depth-3 install is DATA-LIMITED, not a representational cap — and it scales into deployable single-shot. C22 left open whether its weak depth-3 install was data-limited or capped. Bank N tool-found depth-3 pairs (N ∈ {40,160,640} nested, interpreter search over the 16-op DSL); eval on a frozen paired held-out set with 0 leakage (function-sig AND op-composition dedup → novel rules only). Depth-3 think coverage@16 rises MONOTONICALLY 0.00 → 0.087 → 0.212 → 0.375, no plateau; top-dose Wilson lower CI (0.28) > low-dose upper CI (0.17). The DEPLOYABLE install scales too: no-think coverage 0.00→0.338, no-think single-shot greedy@1 0.00→0.10 at N=640 (≈0 at C22's N=130). Depth-2 guardrail rose (scaffold intact). So the deep wall is a DATA bottleneck, not a hard cap: the thin depth-3 thread (C19) thickens with more explorer-found data and converts to deployable single-shot. Design hardened by an adversarial workflow review. Limits: single seed, fixed epochs (data~gradient confound), search-easy bias (untested past 640).
  • qwen35_4b_tool_seeded_banking (claim C22): the C21 positive control — tool-seeded banking crosses the depth-3 wall self-banking couldn't, but weakly. Harvest depth-3 via an interpreter-backed explorer (CPU brute-search over the substrate's own 16-op DSL, no external model, 130/130 solved — what sampling gets ≈0 of), add to C21's exact depth-1+2 pairs, bank. On a frozen paired held-out set (behavioral dedup; design hardened by an adversarial multi-agent review): depth-3 think coverage@16 0.00 (0/40) → 0.125 (5/40 distinct novel tasks) — a significant unlock vs the hard 0/40 floor where C21 self-banking gave exactly 0. Validates the recipe: tools explore, banking installs. But CROSSED-BUT-WEAK — the install is test-time-dominated (no-think depth-3 0.025, greedy@1 0.00) vs depth-2 which installs deployably (greedy@1 0.15). New nuance: the installer's efficacy decays with depth (echoes C19). No free next rung (depth-4 stayed 0). Each rung must be seeded by the explorer.
  • qwen35_4b_wall_climbing (claim C21): self-banking is coverage-seed-bounded — it can't climb the wall. Apex bootstrapping test: bank ONLY depth-1+2 self-solutions (130 pairs, 83 at depth 2, no depth-3 examples), does the banked model now sample depth-3? DEPTH-LOCAL. Depth-2 install works and generalizes to held-out tasks (coverage 0.12→0.36, tripled — clean C18 replication) but depth-3 coverage stays at exactly 0.00 (base 0.00 too) — a strong depth-2 composition skill does NOT length-generalize up. Banking installs only depths the base can already sample; it cannot bootstrap the frontier. Completes the wall picture: depth-3 is not represented (C19), not steerable (C20), not reachable by banking-shallow (C21). The only way up is to seed each rung externally — tool-augmented harvest (C12 decompose-search) → verify → bank. Sharpens C11-M4 into a hard cross-depth wall. Pre-registered P2 (unlock) refuted; P1/P3 held.
  • qwen35_4b_activation_steering (claim C20): decodability ≠ steerability. Causal follow-up to C19: build mean-difference (ActAdd) directions for the first op from C19's cached activations and add them back to the residual stream during generation (forward hook). INERT — at depth 1 (cleanest direction, probe 0.99) steering toward the true op never beats baseline and only degrades at high strength; at depth 2 a faint predicted-direction whiff (+0.05, within noise of the random control, below the +0.10 pre-reg bar); null at earlier layers (8, 12) and on identification (0.03→0.03). All pre-registered predictions refuted. The latent signal C19 found is readable but not writable into behavior. Strengthens the throughline: test-time interventions (selection C17, steering C20) don't move the wall; only weight edits (banking C18) and tools (C12) do. Honest limit: a clean negative for standard ActAdd — patching / optimized vectors untested.
  • qwen35_4b_latent_composition_probe (claim C19): first look INSIDE the wall. Linear probes on residual-stream activations (last identification-prompt token, all 33 layers, 1500 verified tasks) decode the composition's first operation. The wall's nature changes with depth: depth-1 first-op is decoded at 0.99 (rises to ~0.99 by layer 15) while the model names it 0.44 / generates it 0.68 → representation ≫ expression = latent capability; depth-2 probe 0.42 vs behavior ~0.13; depth-3 (the wall) probe 0.27 but the shuffled floor is 0.14, so the real signal (~0.13) ≈ behavior — the representation itself has thinned to a thread. So the wall is an EXPRESSION failure when shallow (info present, unexpressed) and a REPRESENTATION failure when deep (info not computed). Layer-0 stays at chance (signal is computed, not surface). Implication: steering has headroom at depth 1–2 but almost nothing to steer toward at the deep wall; explains why banking (C18) was necessary — it installs the representation the base lacks. Only proposal-installation, not test-time readout, crosses the deep wall.
  • qwen35_4b_coverage_banking (claim C18): banking self-verified solutions does BOTH — concentrates AND expands. The correctly-aimed follow-through to C17 (only shifting the proposal distribution can beat sample-more). Harvest the fixed 4B's OWN execution-verified identification solutions (80 SFT pairs, no teacher), QLoRA-SFT single-shot, eval on DISJOINT held-out tasks (4 arms, base/banked × no-think/think). Depth 1: CONCENTRATION — think greedy@1 0.60→0.80, ceiling flat. Depth 2: EXPANSION — banked coverage@16 0.15→0.45 (3×) on held-out tasks: proposes correct compositions the base never sampled (unique-program count even drops — the proposal mass moved onto correct programs, C17's lever working). Depth 3–4: no move (7 / 0 training examples; wall holds). Bounded: doesn't beat think sample-more at k=1, but banking+sample-more > base+sample-more. To push the wall deeper you need verified deep examples plain sampling can't harvest → seed with tool-search (C12). Refuted its own concentration-only prediction (P3) in the optimistic direction.
  • qwen35_4b_coverage_vs_selection (claim C17): the generation wall is COVERAGE, not selection. Pre-registered decomposition — draw K=32 identification samples/task (list + register, depths 1–4, 8 visible + 8 hidden examples), grade vs visible+hidden, compare selectors to the coverage ceiling. Selection is free: max(coverage − vfilter) = 0.00 in every cell; 90% of visible-passers also pass hidden, so an 8-example execution-filter, the model's own C10-verifier, and even a random pick among visible-consistent candidates all recover the full coverage ceiling identically. Single-shot undersells 2–5× (first@1→cov@32: list d2 0.10→0.30, register d2 0.15→0.60, d3 0.05→0.25) and sample+filter recovers it — but that IS sample-more. The coverage wall's depth is set by hypothesis-space size (list collapses at d3; register survives to d4 via a smaller op menu), mechanistically explaining C16's register floor as coverage-driven. Implication: you cannot beat sample-more by better selection — the lever is shifting the PROPOSAL distribution (C12 tool-search / C11-C12 banking). Refuted its own selection-centric predictions (P3, P4). Residual: overfit traps (visible-pass, hidden-fail) false-deploy at deep register — an abstention gap no example-filter catches.
  • qwen35_4b_crossfamily_laws (claim C16): cross-substrate generality test of the C13C15 ladder on two genuinely different fresh, execution-verified, collapse-rejected families (STRING char-edits, REGISTER 3-int machine) vs the LIST anchor, 100 verified tasks/family. Verdict SCOPED, and the split is the finding: two rungs are model-level LAWS — transcription/compiler (plan-given execution ~1.00 at every depth in all three families; the curves collapse to one line) and the generation wall (bare identification collapses with depth everywhere; trans−ident gap ≥0.84 at depth≥3) — so tools identify, the model compiles is substrate-general. But simulation fidelity is substrate-dependent (C15's decay constant was list-specific): register (compact state) is robust ~flat (0.92→0.72), list decays (1.00→0.56), string is floored near-zero (0.24→0.00). New sub-law: the wall's floor ≈ f(hypothesis-space size, simulability) — register alone (small op-menu + simulable) keeps a nonzero deep-ident floor (0.16/0.08). Promotes C13, narrows C15. (Caught a spurious string-sim-0.00 "law" — a quote-blind parser — before any scored run.)
  • qwen35_4b_depth_wall_anatomy (claim C13): pre-registered anatomy of the compositional wall. It is identification, not execution — plan-given execution 0.90–1.00 through depth 4 (zero execution deficit) while bare identification runs at ~2× over chance per composed op (odds fall ~30×/op; wall at depth 2), insensitive to op type, and barely helped by shown intermediates (segmentation deficit). Retro-explains C10/C11/C12 with one mechanism: the fixed 4B is a reliable compiler starved of hypothesis search — tools identify, the model compiles. Also: 40% of nominal depth-3 tasks were shallower-equivalent (min-depth audit; C12 corrected).

  • qwen35_4b_decompose_compose_frontier (claim C12): the frontier is extendable without a teacher. A decompose-and-compose search (4B ranks next primitive → interpreter executes → recurse) cracks depth-3 monolithic sampling can't (0.125→0.40+, 3.4×); against the brute-force bar the model's guidance buys efficiency not coverage (planner-wall). And BANKING the search-found solutions (QLoRA-SFT, no teacher) extends the frontier into the weights (monolithic pass@5 0.125→0.237, depth-3 4×) — the bound M4 couldn't break. Answer to C11's open problem: tool-augmented search harvests frontier-exceeding solutions, banking pulls them into the model's distribution.
  • qwen35_4b_neurosymbolic_repl_substrate (claim C11): a fresh, contamination-free procedural program-synthesis substrate (random primitive compositions, held-out-execution graded, oracle-solvable 100%) — a reusable asset for elicitation claims with no memorization confound. On it, a neurosymbolic execution-feedback REPL loop did NOT beat matched-compute sampling (M2), but self-training on the 4B's own verified solutions banked capability into held-out single-shot (M3, +0.095). Extends C1 (executable intermediates) toward execution feedback and self-correction training.

Current Read

Structured execution is one of the strongest imported signals. The next useful work is not another isolated positive run; it is controlled comparison of representations and supervision sources — and, per C11, on a CONTAMINATION-FREE substrate so that gains (and non-gains) are honestly measurable.

Replicated semantic commit interface (2026-07-12)

qwen35_4b_commit_slot_semantic_power_replication establishes a stable constrained compiler output seam on a fresh procedural depth-two substrate. At fixed cap 1,024, ordered thought independently beat an identical token-multiset shuffle by +13.57pp and +15.04pp across two disjoint 113-task stages, while the unrestricted next token was already an alias on 88.2% and 87.6% of rows. Syntax therefore exposes rather than manufactures most of the answer-mode state. The effect remains target heterogeneous and free-form close-only output remains poor. The subsequent task-held-out shared J-value measurement failed at chance (0.5021), below slot margin (0.5448) and equal- width non-J residual state (0.5292); midpoint and endpoint J rankings had opposite signs (0.6083 versus 0.3958). The interface is a stable output seam, not evidence for one phase-invariant scalar compiler state. A permitted phase-specific refit also failed: midpoint J 0.5375 versus non-J 0.6000 and margin 0.5396. Causal work stayed sealed and the midpoint-J successor is retired.

The next generative attempt also stops before compiler evaluation. qwen35_4b_jacobian_counterfactual_branching finds chance supplied-alias control (4/48 at every alpha) for balanced additive J edits at a 512-token native prefix, despite valid non-J geometry. The fixed answer slot is a semantic output seam, not evidence that an arbitrary preceding token is a writable compiler register. Explicit anchor tokens and donor- coordinate replacement are the remaining context-local hypothesis.

Fresh materialized-residual mechanics result (2026-07-13)

qwen35_4b_materialized_residual_sibling_search_fresh_replication finally produced the durable model evidence the parent incident lacked. All nine invocation transactions and 1,984 rows authenticated under a separately published lock, and three hostile result audits reproduced every score and gate. The result splits at the interface boundary. Generation is terminal MECHANICS_INTERFACE_INVALID: materialized/name/shuffled/echo suffixes parsed 12/7/12/20 of 52, every thought reached 512, and cap contact was 37/42/40/28; direct parsed 7/24 with 17 cap contacts. Materialized/name/shuffled/direct had zero visible successes, but this is only weak negative mechanism evidence because the registered ABI failed. The parse-immune cheap viability ranker is a clean negative: materialized recall@4 0.257 was below name-only 0.281, shuffled 0.323, listwise 0.271, and surface 0.375; its +0.149 gain over realized random narrowly missed +0.15 and it failed all absolute support/floor gates. Qualification and confirmation stayed sealed. Retire this cheap ranker and top-four branch; any remaining residual-generation test must first freeze an echo-qualified answer seam on separate calibration tasks.

Answer-seam factorial result (2026-07-14)

qwen35_4b_materialized_residual_answer_seam_factorial then tested that prerequisite under a reviewed, committed-green calibration lock. All five transactions and 240 outputs authenticated, but every one of the four think/no-think x freeform/PROGRAM: arms scored 0/48 strict parse and exact echo, yielding terminal NO_VALID_RESIDUAL_ANSWER_SEAM; mechanics and all protected reads stayed sealed. The failure is diagnostic rather than a broad copying miss. After the decision, removing only the exact terminal <|im_end|>\n suffix let the frozen parser accept 48/48 rows in both no-think arms, but only 38/48 think/PROGRAM: and 24/48 think/freeform because ten/five thinking rows had another close boundary. A looser expected-tail diagnostic was 48/48 for think/PROGRAM: and 29/48 think/freeform, but is not exact output. Sampled tokens showed tokenizer EOS 248046, newline 198, then registered HF EOS 248044. This does not repair the gate. It isolates a fresh next experiment: register first tokenizer EOS as the answer-stage deployment commit boundary, use fresh identities, retain strict pre-commit grammar and an HF-EOS control, and stop if that interface does not independently qualify.

Tokenizer-EOS answer-commit calibration (2026-07-14)

qwen35_4b_tokenizer_eos_answer_commit_factorial isolated the boundary cause on 48 fresh known-answer rows under an independently reviewed, committed-green implementation and lock. Both tokenizer-EOS no-think cells achieved 48/48 strict exactness and parseability with zero cap contacts; every matched HF-model-EOS cell achieved 0/48. All 192 paired boundary comparisons authenticated with shared prompts, seeds, thoughts where applicable, and identical sampled prefixes through the earlier stop. Thinking did not magnify the interface: the structured cell fell to 38/48 and freeform to 30/48, with 16 freeform cap contacts. The frozen decision is TOKENIZER_EOS_ONLY_INTERFACE_QUALIFIED, and the sole advancing winner is tokenizer_eos_no_think_program_slot. This establishes that termination-token identity is a causal part of Qwen3.5-4B's short structured-output interface; it does not establish residual-mechanics or capability gain. Mechanics and all protected labels remained sealed and require a second winner-bound published lock before access.

The later winner-bound run passed its independent transport gate (24/24 exact echo and parse, zero caps) and completed five durable transactions containing 4,056 outputs. It stopped before visible selection: replay authentication of the stored transport decision incorrectly reused the initial temporal gate after later invocations were complete. Hidden scoring and benchmark access remained sealed. Record this as terminal instrument failure, not capability evidence; a fresh-identity successor is required.

Replay-hardened tokenizer-EOS residual result (2026-07-14)

qwen35_4b_tokenizer_eos_residual_mechanics_fresh_replay supplies the clean adjudication. The fresh successor excluded every parent function, identity, prompt/token rendering, and seed domain; passed independent implementation review; and crossed separate calibration, mechanics-lock, and visible-publication CI boundaries. Calibration replicated the interface (48/48 in both no-think tokenizer-EOS cells; all HF-EOS controls 0/48), and transport was 24/24 exact/parse with zero caps. All 4,056 mechanics outputs and the repaired five-stage replay authenticated. Generation itself was healthy: parse was 99.78% direct, 99.13% materialized, 99.83% name-only, and 98.78% shuffled, with zero cap contacts and no direct-pool exhaustion.

The scientific result is a clean negative. Materialized, name-only, shuffled, sampled-token-matched direct, and logical-model-token-matched direct each had 0/24 deployable success and 0/24 oracle proposal coverage. This is not a selector or task-solvability artifact: exhaustive CPU evaluation of all 13,824 programs found 88 visible-consistent, hidden-correct programs, 1--9 on every task, for 24/24 coverage. Under this depth-three DSL and frozen answer seam, all-candidate semantic materialization did not move correct residual programs into Qwen3.5-4B's proposal support. Retire this prompt mechanism; future residual/Jacobian work must change the proposal-formation mechanism and show a forward causal coverage gain over matched sampling, not merely a readable coordinate, a looser selector, or more of the same prompt.

Materialized residual mechanics incident (2026-07-13)

qwen35_4b_materialized_residual_sibling_search does not update scientific belief. Its fresh exact-depth-three construction and controls passed model-free review, but attempt 1 stopped in cache preflight and attempt 2 lost 52 returned mechanics rows before durable persistence because a post-generation receipt collapsed model EOS 248044 and tokenizer EOS 248046. Independent review blocks replay of the terminal STARTED transaction. The materialized-residual hypothesis remains untested; only a new experiment with fresh identities/seeds and write-before-authentication quarantine may resume it. No claim is allocated.

Scorecard

  • Program: charter
  • Current read: structured intermediates remain strong, but tested latent interfaces usually fail before causal use. The replicated First: seam is neither a stable J value coordinate nor an arbitrary writable last-thought register. Early concrete text demonstrates broad local routing—84/96 execution for both correct and deranged supplied operations across all 24 operations—but its interface failed and complete-program reachability was only 3/8. The durable fresh materialized-residual run split the question: free generation was interface-invalid, while the parse-immune cheap materialized ranker lost to every structured comparator. Tokenizer EOS plus no-think is a real short-output interface: two independent calibrations reached 48/48 in both prefix cells against 0/48 for every HF-EOS control, and fresh transport was 24/24. The replay-hardened mechanics result now cleanly rejects the prompt mechanism: all 4,056 outputs authenticated with 98.78--99.83% parse and zero caps, yet materialized and all matched controls had 0/24 selected success and 0/24 oracle proposal coverage while exhaustive search covered 24/24 tasks. A strict answer seam does not expose residual synthesis. Cheap viability/top-four ranking, parser repair, and all-candidate semantic-materialization prompting are retired. The fresh three-seed rank-32 LoRA joint adjudication also missed registered state in all 57 required cells (best 0.0234 versus 0.40); because adaptation contrast is uncertain, mandatory LoRA-state-only and matched direct-full-shape controls—not a rank conclusion—come next.
  • Best next experiments: complete the authorized state-formation Stage B with three LoRA state-only and three matched direct-full-shape joint seeds; separately measure the exact depth-6 resource crossover before learned pruning; and, for J-space, require a task-held-out correctness coordinate beyond margin/position/equal-width non-J controls before a separately preregistered same-prefix intervention must causally increase correct-proposal coverage over matched sampling.
  • Strong anchors: qwen35_4b_early_text_hypothesis_forking, qwen35_4b_partial_structure_search, qwen35_4b_crosssubstrate_structure, qwen35_4b_structure_search_scaling, qwen35_4b_commit_slot_semantic_power_replication, qwen35_4b_state_carry_vs_state_bag_fullrank_delta.
  • Avoid repeating: another opaque name/timing variant, this cheap materialized P(viable)/top-four ranker, the all-candidate semantic-materialization prompt with a looser selector or more samples, waiting for HF model EOS when tokenizer EOS is the registered answer commit, adding thinking to a short exact-output channel without a task-specific benefit, larger cap or parser relaxation after ABI failure, pooled-AUROC launch, model-guided search without measured brute wall time, or a rank comparison whose shared modules and dropout streams are shifted by parameter-construction RNG.
  • Evidence that advances the program: causal ablation showing which intermediate structure transfers across family, length, or paraphrase shifts.

Charter

Show charter.md

Purpose

Discover when small models should stop emitting direct answers and instead emit, condition on, or be supervised through executable structure: typed slots, latent registers, bytecode, candidate programs, state traces, compiler heads, or differentiable runtimes.

Why This Is A Program

The imported experiments contain many successful variants, but the line is much broader than those prototypes. Future work should compare representations, supervision density, curricula, model sizes, and task substrates under a shared question: what structure turns small-model brittleness into reliable execution?

Progress Signals

  • Longer-horizon execution improves without beam search or oracle selection.
  • Paraphrase and distribution-shift splits preserve the same executable trace.
  • Program/state supervision explains gains beyond raw data scale.
  • Failures can be localized by operator, step, state prefix, or representation.

Boundaries

This program studies representation and execution. Selection between many candidates belongs primarily to Evidence-Conditioned Selection. Reusable banks of primitives belong primarily to Operator And Skill Inventories.

Backlog

Show backlog.md

Next Experiments

  • The LoRA architecture counterfactual completed with valid PILOT_MECHANISM_MISS; the held-fixed, zero-initialized 892M-parameter full-rank successor also completed, with raw joint-state accuracy 0.00277 versus the 0.40 gate, Carry minus Bag -0.0156, negative unseen-K scaling, and noncausal swaps. Its analyzer emitted PILOT_STATE_FORMATION_MISS, but post-result preregistration audit assigns PILOT_PROMOTION_BLOCKED: the run simultaneously failed the non-capacity requirements of positive Carry-minus-Bag and positive query-kind effects. It therefore does not isolate or close the rank/capacity question. Do not run either existing experiment's confirmation, edge-cut, or sample-more stages, and do not open the interface successor.
  • Continue the mandatory fresh RNG-matched state-formation adjudication. Its three-seed rank-32 LoRA joint arm validly emitted LORA_JOINT_MISS_CONTROLS_REQUIRED: 0/57 required cells passed 0.40 and the best intact cell was 0.0234375, while the setup path, exact K=1 bypass, and positive controls passed. Execute the now-authorized Stage B exactly: seed-matched full-rank G0/positive controls, all three LoRA state-only fixed finals, and all three direct-full-shape joint fixed finals on the same rows, initialization, stochastic streams, and readouts. Do not attribute the miss to LoRA, open sealed contrast, or pivot representation/supervision until the registered Stage-B receipt. The first full-rank G0 exposed an immutable downstream path-consumer defect before model load; its exact-prefix recovery, archive, and retirement checkpoints are green. The recovered full-rank seed-7411 G0 now passes all registered feasibility gates, but that frozen wrapper's pathname-only retirement guard cannot hand off to the positive control after the successful G0 repopulates the canonical slot. The additive byte/status-aware handoff is published/green, and the full-rank all three full-rank G0/control pairs now pass independently: each intact control is 48/48 versus adaptation-disabled 0/48 after the exact 256-update, 4,096-presentation schedule. Publish the complete setup matrix, then run the three LoRA state-only and three full-rank joint fixed finals. This remains mechanics/setup evidence and does not authorize a LoRA rank conclusion.
  • Cross-program interface probe completed: qwen35_4b_commit_slot_jacobian_value_transport showed that a fixed latent answer slot repairs formatting but its semantic hint remains task/alias concentrated (15/48 real versus 11/48 shuffled at 1,024; five mixed tasks versus six required). A powered fresh replication must clear task-level uncertainty before the interface can license any J-space value work.
  • Powered cross-program replication completed: qwen35_4b_commit_slot_semantic_power_replication fixes the same slot/cap on 113+113 fresh tasks and requires both task-level ordered-content evidence and eight-alias semantic breadth before treating the latent interface as usable. Qualification passed with correct support across all 11 targets and choices across all 12 aliases; confirmation independently passed with 98/339 ordered versus 47/339 shuffled and support across 10/11 targets and all 12 choices. The licensed shared J-value measurement was then chance (0.502) and weaker than slot margin/non-J state, with midpoint/end phase reversal. Treat the slot as a stable output seam, not a stable scalar compiler-state coordinate; causal work is not licensed. A post-decision midpoint-only refit reached only 0.538 versus matched non-J 0.600, so do not replicate this J-axis interpretation.
  • Additive J directions do not turn the last native-thought token into a hypothesis register. The late explicit anchor was invalid, and early concrete text redirected direct execution across 24 operations but failed its interface. The first materialized-residual successor then ended without a durable, authenticated model result: one preflight abort and one terminal 52-row write-order/EOS-receipt incident. Retire opaque-name timing variants. Resume materialized residuals only in a new experiment with fresh tasks/ record IDs/seeds, write-before-authentication quarantine, the exact model/tokenizer EOS pair, and the already frozen candidate-blind, shuffled, exhaustive, and sampled/ logical-token-matched controls. That identity is now reserved as qwen35_4b_materialized_residual_sibling_search_fresh_replication; fresh construction and two adversarial reviews passed with zero required parent identity/prompt intersections and zero model calls. The crash-safe mechanics implementation then passed three independent audits and model-free preparation with 1,984 requests, 676 unique identities, zero in all nine parent/terminal intersections, and zero model calls. Its first lock attempt failed closed on an incorrect historical source commit before any lock/raw/ model activity. A three-reviewer append-only V2 repair preserves every V1 payload byte and again records zero model calls. The audited live successor then completed all nine transactions. Generation was terminal MECHANICS_INTERFACE_INVALID: all thoughts hit cap and every arm missed the parse/cap ABI by a wide margin, so it cannot adjudicate residualization. The parse-immune cheap materialized ranker was a clean negative: recall@4 0.257, below name-only 0.281, shuffled 0.323, listwise 0.271, and surface 0.375. Retire that ranker and top-four branch. The fresh echo-gated answer-seam factorial has now completed under a reviewed lock: all four arms scored 0/48 strict parse/exact echo and mechanics stayed sealed. A post-decision diagnostic isolated the mismatch cleanly only for no-think: both no-think arms became 48/48 frozen-parser exact after terminal <|im_end|>\n removal; think/PROGRAM: and think/freeform were only 38/48 and 24/48 because of extra close boundaries. Do not relax this result's parser. If residual generation continues, create one fresh answer-stage commit-boundary successor that stops at first tokenizer EOS, trims only that terminal token, retains current HF EOS as a control, uses new tasks/IDs/seeds, tests malformed/extra pre-commit bytes, and must independently clear the same >=90% echo/parse and <=5% cap gates before fresh transport and disjoint mechanics. That successor is now active as qwen35_4b_tokenizer_eos_answer_commit_factorial; its first-stop and strict- precommit calibration has now passed under a committed-green reviewed lock. Both tokenizer-EOS no-think cells were 48/48 exact/parse with zero cap contacts, while all matched HF-model-EOS controls were 0/48; thinking fell to 38/48 in the structured cell and 30/48 in freeform. Only the frozen tokenizer_eos_no_think_program_slot winner advanced from calibration under a second committed-green lock. This qualified an interface only; the residual-capability question remained unadjudicated. The winner-bound run subsequently passed transport 24/24 and durably completed all five transactions, but visible analysis failed because post-chain transport-decision replay reused the initial later-absent authorization gate. No hidden read occurred. Replace it with a fresh-identity successor that separates initial authorization from immutable replay under a new review and lock; do not rerun the result-bearing experiment. That successor is now registered as qwen35_4b_tokenizer_eos_residual_mechanics_fresh_replay, with fresh function fingerprints, identities, prompts/token sequences, seeds, ciphertext/key, and an enforced parent-sampled-bundle denylist. It completed cleanly: calibration again qualified no-think tokenizer EOS, transport was 24/24, every generation arm passed ABI, and all 4,056 outputs authenticated. Materialized and all four frozen comparators nevertheless had 0/24 selected success and 0/24 oracle proposal coverage, while exhaustive search covered 24/24 tasks. Retire the all-candidate semantic-materialization prompt branch. Do not spend another run on selector repair, parser relaxation, or a larger matched budget for the same mechanism. A J-space successor must first show a task-held-out correctness coordinate beyond margin/position/equal-width non-J controls, then causally increase correct-proposal coverage at the same prefix over matched sampling; stop at measurement if that forward gate fails.
  • Measure the exact behavioral quotient at fresh depth 6 before assuming model-guided pruning is economically needed; record wall time, memory, physical transitions, coverage, and selector success.
  • If a real search wall appears, test a residualized partial state (feasible parameter domains, materialized prefix outputs, per-example target residuals) behind the same within-task AUROC and recall@beam gate.
  • Repair the demonstrated visible-only selector gap over exact solver pools (60/60 coverage versus 56/60 selected) with frozen stability/simplicity/unlabeled-probe rules.
  • Replicate the strongest structural compiler results across seeds, lengths, and operator mixes.
  • Run a direct-text-program versus typed-bytecode versus latent-slot comparison on one shared task suite.
  • Add adversarial paraphrase and compositional splits where direct prompt cues fail.
  • Measure whether state-prefix supervision, final-answer supervision, or program-token supervision is the causal lift.
  • Build a small diagnostic suite that every compiler-style experiment can run before claiming generalization.

Required Controls

  • Direct answer baseline.
  • Same model and data with unstructured output.
  • Shuffled or corrupted state supervision when state traces are used.
  • Length and family holdouts.

Stop Conditions

Retire a variant when it improves train or IID accuracy but cannot survive harder length/family/paraphrase splits after two controlled attempts.

Type-only absolute P(viable) at think@256 is retired as a search controller unless a materially richer state or interface first clears calibration; pooled AUROC alone must not reopen it.

Experiments 212

  • 2026-07-27 → 28 Self-Written Verifier Fidelity

    Fill this after the run. Separate deployable evidence from oracle/hidden evaluation.

  • 2026-07-25 Real-Repo Agentic Instrument

    Yes: 200 tasks over 15 libraries, each one verified two ways - the library's suite passes untouched, and deleting the target function actually breaks specific named tests. Three design traps had to be avoided, and each w

  • 2026-07-19 Qwen35 4B Agentic RLVR Feasibility

    Yes to the machinery: with three fixes (force eager mode so the hybrid architecture doesn't hang, load the model in bf16 not fp32, and split GPU memory 55% to the inference engine with sleep-mode) the loop CLOSES - a mov

  • 2026-07-18 Qwen35 4B WHY-Comment Install

    The WHY idea paid off where it should — on writing correct functions — and did nothing where it shouldn't. We trained the 4B to write code with the causal reason for each line attached as an inline #WHY: comment (generat

  • 2026-07-18 Qwen35 4B Self-Repair Install

    The first bet that moved the needle at all — gently. We taught the 4B to debug by training on 504 examples of [buggy code + the real test-failure message] -> [diagnosis + fix], all self-generated by injecting bugs into c

  • 2026-07-18 Qwen35 4B Repair + Why Stack

    The obvious way to combine our two promising curricula - just train one model on both - backfired. We put the 504 self-repair rows and 504 WHY-comment rows into one 1008-row training set and trained a single adapter. Ins

  • 2026-07-18 → Qwen35 4B WHY-Think Scale

    This phase built and proved the machinery; the GPU training sweep has not run yet. The generator now emits, for every example, a real hidden reasoning trace generated mechanically from the program's shape - parse the spe

  • 2026-07-18 → Qwen35 4B WHY Scale Ladder

    This phase built and proved the machinery; the GPU training sweep has not run yet. The core blocker was that the original WHY generator saturated fast (about 75 distinct reasons, 438 distinct programs at 504 examples), s

  • 2026-07-17 Qwen35 4b State Track Confirmation

    The lift held up directionally, without becoming a slam dunk. Across six fresh sealed exams, running the same seed through both models so the noise cancels, the state-tracking model beat its parent on 4 of 6 with an aver

  • 2026-07-17 Qwen35 4B Exec-Trace Install

    The first bet at installing coding cognition came back flat. We trained the 4B on 400 self-generated, execution-verified program traces to install an accurate 'mental interpreter,' the idea being that a model that can si

  • 2026-07-17 Count-Walk Menders Confirmation

    The answer the rule was built to force out, delivered without wiggle room. Across the four fresh exams the trained model solved a fix-the-procedure episode exactly once (plus one partial credit that the rules pre-declare

  • 2026-07-17 Coding Fitness Harness (cognitive-core program)

    Surprisingly good at writing single functions (HumanEval 76.2%, MBPP 56.5%) but weak at driving a multi-step coding task in a real agent loop (23%). The harness is validated: it agrees with an independent run to within 1

  • 2026-07-17 → State-Track Installation (Stage 9)

    The believed-unlikelier but hoped-for outcome landed. After the reliable 'just replay again' lever hit its ceiling, this tried a genuinely different lever: teach the model one new, universal skill — keeping a running tal

  • 2026-07-16 Qwen35 4b Zero Root Lineage Rebuild

    ["The rebuild answered the provenance question with numbers. Retracing the six documented training steps from a truly blank starting adapter — same datasets, same seeds, same settings — produced a model with about ninety

  • 2026-07-16 Sweep-Rate Consolidation (Erratum: 2/6, not ~50%)

    Two sweeps in six readings — one in three, not one in two — with a wide honest confidence band (roughly 4 to 78 percent at 95%). The texture matters more than the point estimate: the model never lost a single family to t

  • 2026-07-16 Clean-Path Statechain Extension

    ["Split verdict with a clean lesson. The state-tracking dose installed for the THIRD time on its third different parent — the program's most reliable trained effect — and the clean-lineage model beat the untouched base b

  • 2026-07-16 Clean Gym-Mix Dose

    The mix failed cleanly and instructively. Splitting the standard 160-lesson budget across three skills — sixty trick-instruction episodes, fifty procedure chains, fifty answer-or-abstain puzzles — taught none of them: on

  • 2026-07-15 Statechain-Only Dose

    Three results in one event. First, the state-tracking dose passed its local gate cleanly — the skill installed again and this time forgetting stayed inside the calibrated margin. Second, on the real benchmark the trained

  • 2026-07-15 Goal-Gate Confirmation

    The replication returned a split answer. The overall improvement replicated without drama: the trained model beat the base decisively on all three fresh seeds, making four for four all-time. The perfect ten-family sweep

  • 2026-07-15 Feedback-Loop + State-Chain Install

    The dose split down the middle. The hidden-state-tracking half installed cleanly — on brand-new test instances the trained model tracked procedures better than both its parent and a matched control. The feedback-repair h

  • 2026-07-15 Dose-Diversity Mechanism Cell

    Variety does not protect memory: the larger varied practice set also cost nine retained answers, the known ten-point case reproduced exactly, and even pure review cost five on this fresh screen — so the one past case wit

  • 2026-07-15 Axis Corpus V2 with Staged Repair

    Even with lessons rebuilt from the autopsy — walking through the search step by step instead of asserting the answer — the repair skills did not stick (a tie and a loss against the comparisons), so the preregistered kill

  • 2026-07-14 → 15 Goal-Gap Axis Curriculum

    The targeted practice worked on its own terms — the first screen pass in this program's history (28 vs 22 and 18 of 40 on unseen tasks, with zero forgetting). On the held-out benchmark it beat the untrained model by a wi

  • 2026-07-14 → 15 Qwen3.5-4B Counterfactual Plan Reflection Transfer

    Not known yet. The current checkpoint builds and checks fresh list, text, and register puzzles, plus a control where every plan is deliberately assigned to the wrong puzzle. No model has been loaded or trained.

  • 2026-07-14 Natural-Language State-Table Universal Curriculum

    Not known yet. CPU construction produced 80 truth-checked lessons and two 320-row streams with exactly 286,814 tokens each. All 48 smoke tests pass, but no model has been trained or evaluated.

  • 2026-07-14 Search-Scaffold Universal Curriculum

    No. On 26 fresh paired cases, the staged-search model tied replay at 16 correct and trailed its parent at 18. All three parsed 23 answers and hit the limit three times; the candidate solved none of four execute/induct ca

  • 2026-07-14 Failure-Selected Counterfactual Restart Curriculum

    The data-selection stage passed: from 624 fresh attempts, the frozen rule selected 52 clean restarts, exactly four for each of 13 skills. Forty correct accuracy or cap failures and twelve correct-but-too-long cases are n

  • 2026-07-14 Qwen3.5-4B Tokenizer-EOS Residual Mechanics Fresh Replay

    No result exists yet. The scaffold declares seven freshness controls and eight lifecycle controls, reads none of the parent's sampled bundles, and authorizes no model call until independent review and release gates pass.

  • 2026-07-14 Qwen3.5-4B Tokenizer-EOS Answer Commit Factorial

    The boundary worked cleanly: both no-thinking chat-end conditions produced 48 correct strict answers out of 48, while every matched later-end condition produced zero. The harder transport check also scored 24 of 24. But

  • 2026-07-14 State-Formation Branch Authorization Recovery

    Yes for the no-model safety check: all six controls passed, the failed first attempt has a third preserved copy, and the two retry-blocking source paths were retired only after that archive commit passed both checks. The

  • 2026-07-14 State-Formation Branch Handoff Recovery

    Yes, the narrow handoff passed every safety check, then all three full-size setup controls reached 48 of 48 with their learned path and zero of 48 with that path disabled.

  • 2026-07-13 → 14 State-Formation Analysis Recovery

    Yes, the narrow recovery check passed. It accepts only the one registered file location, still rejects unsafe shortcuts, and has not yet examined any model result.

  • 2026-07-13 Validation-policy counterexample curriculum

    The test could not answer the question. Before the newly trained model was allowed to run, the starting model and the equal-size old-lesson control each fixed all 48 practice cases. The required gains were therefore impo

  • 2026-07-13 Replay-Anchored Universal Curriculum Continuation

    No at this dose. The designed mix passed its fresh synthetic screen but scored about 42% overall, below both the 44% mature policy and the 49% replay-only comparison. It also fell below base on one family. Replay-only wa

  • 2026-07-13 Mid-Density Token-Matched Universal Curriculum

    The 160-lesson stream improved fresh local accuracy from 17 to 19 of 26 cases, raised readable answers from 18 to 23, and cut answer-limit contacts from nine to three. It still missed the readability and length gates by

  • 2026-07-13 Low-Density Token-Matched Universal Curriculum

    No. Neither 40 nor 80 abstract lessons passed the fresh local screen. The stronger 80-lesson stream matched the inherited anchor at 54% accuracy and parsed 62% of answers, but still hit the answer cap 10 times; broad eva

  • 2026-07-13 Qwen3.5-4B: Installing Universal Features via Designed Synthetic Curricula

    Not with either tested training geometry. Continuing from the strong policy raised fresh synthetic accuracy from 50% to 69% and beat the untrained base overall, but it scored 31% aggregate versus 45% for the strong start

  • 2026-07-13 State-Formation Capacity Adjudication

    Not run yet. The reviewed test starts with three compact-update training runs. If any required state check misses, six matched controls become mandatory; the final three runs open only if the full-size version also misse

  • 2026-07-13 Full-Rank Extra-R Delta: State-Carry Versus State-Bag

    The run worked mechanically, but it did not settle the question. All 892 million full-size update weights trained and fit comfortably, yet macro task-mean joint-state accuracy was only 0.28% against a 40% requirement. At

  • 2026-07-13 Semantic-policy headroom tournament

    No rule produced a stable training target after a failed test: every one of 72 cases reached a correct patch. But before seeing test output, the model got zero of 54 ambiguous first patches fully right and never opened t

  • 2026-07-13 Qwen3.5-4B Semantic-Anchor Coordinate Branching

    This run cannot establish that. The internal edit strongly changed the model's choice among candidate names, but none of 440 consequence outputs began with a valid answer token. Even in a restricted twelve-choice readout

  • 2026-07-13 Qwen3.5-4B Materialized Residual Sibling Search Fresh Replication

    The generation test cannot answer the question because every reasoning trace reached its length cap and the materialized arm parsed only 12 of 52 outputs, with zero successes. The separate one-token ranking test is a cle

  • 2026-07-13 Qwen3.5-4B Materialized Residual Sibling Search

    Not yet. The scientific design and every model-free construction check pass, but the model has not run. The frozen test contains 264 fresh functions, and all 38,596 planned prompt renderings fit their assigned context li

  • 2026-07-13 Qwen3.5-4B Materialized Residual Answer-Seam Factorial

    No registered style qualified: all four scored zero strict parses out of 48, so mechanics stayed sealed. Removing only the final chat-end marker and newline made both no-think styles exact on all 48 rows; thinking still

  • 2026-07-13 Qwen3.5-4B Early Text Hypothesis Forking

    The experiment design and model-free checks now pass, but the model has not run. Review expanded the first draft from twelve operation names to twenty-four fully specified operations, added two fair late-hint comparisons

  • 2026-07-13 Qwen3.5-4B Counterfactual Order-Support Selector

    Not reliably. The meaningful-order rule got 43 of 113 puzzles right, clearly better than first attempt or majority vote. But it was only two puzzles ahead of simply choosing the most decisive attempt, and a deliberately

  • 2026-07-13 Qwen3.5-4B: The Compute-Optimal Confidence Policy (select + abstain + escalate)

    Partly. Picking the answer with the highest single-token P(True) confidence is the best grader-free selector at every compute level, and abstaining below a confidence threshold trades coverage for accuracy cleanly — both

  • 2026-07-12 Transaction-invariant recovery curriculum

    Partly. Focused training lifted practice repairs from 52% to 82% and preserved the test-and-revise loop. But on different transaction APIs it reached 72% versus 70% before training, far below the required ten-point gain.

  • 2026-07-12 Verifier-conditioned recovery banking curriculum

    Yes — replaying real failures taught recovery, lifting success from about half of cases to more than nine in ten. But the best version also added a little "explain your fix" text, and that supposedly tiny slice actually

  • 2026-07-12 State-Carry Versus State-Bag Counterfactual

    The first matched test did not answer that architecture question because the low-rank update failed to learn the required running state. Carrying memory improved overall accuracy by only 4.3 points, with uncertainty span

  • 2026-07-12 Qwen3.5-4B Same-Prefix Advantage Routing

    Only when it replicates, and here one helper didn't. The test actually ran: the deep specialist's local edge held up on a fresh, separate batch, but the quick specialist looked strong on one batch of stuck points and the

  • 2026-07-12 Repository search-compress-bank coding curriculum

    No. It became flawless on the six codebase types it trained on, solving every task in exactly four clean steps. But on four never-seen codebase types its success collapsed from 68% to 35% — worse than just running the un

  • 2026-07-12 Public-verifier recovery branch tournament

    No. On four new problem types, both policies repaired 74%, and their combined best reached only 75%—exactly tied with two action-only attempts. The selector was correctly stopped before scoring.

  • 2026-07-12 Payload-capable recovery agent harness

    No. It led on the first unseen task block at 71%, but confirmed at 69%, exactly tied with the action-only model and short of the required lead. The external benchmark stayed sealed.

  • 2026-07-12 Qwen3.5-4B Pareto Policy Integration

    No. Retested on fresh, uncontaminated tasks, the version built for quick work actually lost at quick work by about two points, while the version built for long work won at both quick AND long tasks. One version quietly d

  • 2026-07-12 Qwen3.5-4B Jacobian Value Transport

    No. Editing one internal direction at a late layer flipped the concept the model said aloud on 18 of 24 tries, up from never, and far beating an ordinary readout-style edit that worked only 1 in 5 times. But when that co

  • 2026-07-12 Qwen3.5-4B Deep-Advantage MOPD

    No. The deep specialist really was the better teacher on selected states, and copying it worked slightly better than copying the wrong teacher or training on matched non-winning states. But after four rounds the resultin

  • 2026-07-12 Qwen3.5-4B Commit-Slot Semantic Power Replication

    The reasoning result is real, but the proposed internal value meter failed. Ordered scratch work beat the same words shuffled in two independent stages (about 29% versus 14% on confirmation). Yet a task-held-out model of

  • 2026-07-12 Qwen3.5-4B Commit-Slot Jacobian Value Transport

    Barely. Forcing the format fixed one problem outright: an allowed word was the model's top choice 85% of the time, versus 4% when it answered freely. But real step-by-step thinking beat the very same thought words scramb

  • 2026-07-12 Qwen3.5-4B Balanced-Core Answer-Potential SFT

    Not run yet — the reasoning bank is fully built (360 tasks, six competing selection rules staged) but no model has been trained or scored. The built test will fine-tune six copies, each fed reasoning chosen a different w

  • 2026-07-11 Entropy-routed think-pivot optimization round 2

    No. Gently pulling the better word up at 155 hand-picked wrong-turn moments was cleaner than shoving the bad word down, and it carried real signal, beating a scrambled-label control by about 14 points on fresh coding rep

  • 2026-07-11 Interactive policy curriculum: oracle DAgger to execution-reward RL

    No — it backfired badly. Success on the practice tasks fell from about 60% to 35%, and on tasks never touched in training from about 69% to 35%. It kept editing fluently but stopped running its checks: the verify-and-com

  • 2026-07-10 → 11 Think-block FTPO round 1: outcome-conditioned pivot steering as an agentic install recipe

    No. Nudging at those forks made the model worse, not better — success on fresh tasks fell about 4 to 8 percent instead of clearing the 5-point gain hoped for. The tell: feeding it deliberately scrambled labels did nearly

  • 2026-07-10 Qwen3.5-4B Verified Macro Invention Long-Context Rerun

    No. The model was not short on thinking space, it was stuck repeating itself, and more space only bought longer loops. Handed the blueprint, all 16 tasks finished cleanly. Asked to invent the program from examples alone,

  • 2026-07-10 Qwen3.5-4B verified-macro exact CUDA-graph vLLM rerun

    No, it just loops harder. At both budgets, all 48 reasoning attempts ran out of thinking room and had to be force-stopped, none finishing on its own. Widening the budget from 49,152 to 61,440 tokens pushed the share stuc

  • 2026-07-10 Qwen3.5-4B verified-macro capacity-fit vLLM rerun

    No. Two hidden limits bit first. Configuring the server for 64 simultaneous long prompts demanded about 2.4 million tokens of fast memory from a pool holding roughly 1 million, forcing constant re-reading. Cutting to 19

  • 2026-07-10 Qwen3.5-4B Partial-Structure Recognition-Guided Search

    No. Shown a half-finished program skeleton, the four-billion-parameter model's guess at whether it could still be completed was barely above a coin flip — about 51% correct, where 50% is pure chance — and letting it reas

  • 2026-07-10 Qwen3.5-4B Answer-Potential Trace SFT

    No. The score carried genuine signal — it tracked what the reasoning actually said, not its length or wording — but its top pick raised fresh-answer success only from about 13% to 20%, short of the useful margin needed t

  • 2026-07-09 → 10 Gauntlet round 1: breadth-first agentic expert iteration

    Mostly the second. On a blind benchmark the model scored about 14 percent, with six task types near zero, but it had usually reasoned correctly. It simply hit its thinking limit, restarted explaining instead of writing t

  • 2026-07-09 Qwen3.5-4B Verified Macro Invention

    No. Even handed the finished plan and asked only to re-express it — no problem to actually solve — the model got it exactly right just one time in four, short of the three-in-four bar. Every output looked flawless: valid

  • 2026-07-08 → 09 Does the installable hypothesize-and-verify skill move the structure wall?

    No. Fine-tuning the model on 1,476 worked guess-and-check traces doubled its success at the two-step depth it practiced on (lists jumped from 37% to 70%), yet did nothing one step deeper: three-step success stayed near 5

  • 2026-07-07 → 08 Qwen3.5-4B: Can Confidence Replace the Verifier in the Banking Flywheel?

    No. When training on answers checked by actually running the code lifted single-shot accuracy from 8% to 24%, confidence-filtered data — fifteen times purer than a random grab of the model's own outputs — landed right on

  • 2026-07-07 Qwen3.5-4B: Does the Confidence Toolkit Survive on Real Code?

    Yes, but not the obvious way. Averaging the model's certainty across every token of a program barely beats a plain majority vote among the tries. The real winner: make the model write the code, then answer one yes/no que

  • 2026-07-06 Qwen3.5-4B: When Does the Model's Structure Beat Brute Search? (depth-4)

    No. Adding one extra dial — a sixteen-times-larger space of combinations — did not flip things. Exhaustive search stayed near-perfect at about 97 percent, while the model's knack for guessing the right combination from m

  • 2026-07-06 Qwen3.5-4B: Is the Wall Structure or Values? (skeleton-then-fill)

    The steps. Handed the correct sequence of operations, a cheap number-search finished every single task, so the numbers were never the bottleneck. Left alone, the model almost never even lands the right sequence, and cred

  • 2026-07-06 Qwen3.5-4B: Externalize the Latent Readout (probe-to-prompt)

    Partly. Writing the concrete first step — exact value and all — into the prompt lifted the single-best-guess solve rate on two-step tasks sixfold, from 3% to 19%, where editing the model's internal state did nothing. But

  • 2026-07-06 Qwen3.5-4B: Is the Parameter Latent? (probe the full first op)

    It splits. The model genuinely computes the KIND of operation inside itself — reading its internal activity names the kind far better than the examples alone do (41% versus 27%, against 6% for blind guessing). But the sp

  • 2026-07-06 Qwen3.5-4B: Do the Structure Findings Generalize? (string/register)

    The wrong sequence. Across three completely different kinds of programming tasks, a 4-billion-parameter model almost never solved one on its own — 0 to 2 percent — and its misses were wrong-order, not right-order-wrong-n

  • 2026-07-06 Qwen3.5-4B: Does Banking Install STRUCTURE?

    Yes, then no. Training lifted a 4-billion-parameter model from never proposing the right step-sequence (0%) to getting it right about half the time (51%) on brand-new tasks—a real new skill, not memorized answers. But if

  • 2026-07-05 Qwen3.5-4B: Thinking vs the Lookahead Wall

    No. Asked to name the first of three moves toward a goal, the model stayed stuck at pure guessing — right about 1 time in 32 — no matter how long it thought, even with 2,048 tokens to think first. Yet recognizing a goal

  • 2026-07-05 Qwen3.5-4B Latent Decomposition: be your own tool-search

    No, not out of the box. One step from the goal it picks the right operation about eight times better than chance, but on the opening move three steps out it ranks correctly only at chance — recognition, not planning. As

  • 2026-07-05 Qwen3.5-4B Depth Scaling & Controls: saturation, data-vs-compute, depth-4

    Variety, decisively, and the gains never plateaued. Rereading the same 40 solved examples sixteen times over barely moved the solve rate (9% to 16%, within noise), but replacing them with more distinct examples at the ex

  • 2026-07-05 Qwen3.5-4B: Do Banking and Thinking Stack?

    It depends on how far the goal is. One move away, the two boosts stack almost perfectly: a plain model picks the right move 27.5% of the time, extra training lifts that to 52.5%, and adding thinking reaches 85% — the exa

  • 2026-07-04 Qwen3.5-4B Tool-Seeded Banking: does tool-search + banking cross the depth-3 wall?

    Partly. A brute-force search found working three-step programs the model never produces itself, and retraining on them lifted three-step success from a hard zero to 5 of 40 fresh, never-seen tasks when it can reason acro

  • 2026-07-04 Qwen3.5-4B Depth-3 Dose-Response: data-limited or representational cap?

    Just starved for examples. Feeding a fixed 4-billion-parameter model more search-found solutions lifted its solve rate on fresh three-layer puzzles from 0% to 38% when given sixteen tries — climbing steadily at every dos

  • 2026-07-03 Qwen3.5-4B Wall Climbing: does banking shallow composition unlock deeper coverage?

    No. Fine-tuning the model on the two-step solutions it could already produce tripled its two-step success on fresh tasks, from 12% to 36%. But its three-step success stayed at exactly zero, unchanged from before: both mo

  • 2026-07-03 Qwen3.5-4B Latent Composition Probe: is the wall representational or expressive?

    It depends on the number of steps. For a one-step recipe the first step is almost perfectly written on the internal scratch paper (99% readable), yet the model voices it only 44% of the time: it knows but stays silent. F

  • 2026-07-03 Qwen3.5-4B Coverage vs Selection: anatomy of the generation wall

    It never generates it. Whenever a correct program shows up among 32 tries, running each candidate against eight known examples finds it every single time — the model judging its own work, and even a random pick among sur

  • 2026-07-03 Qwen3.5-4B Coverage Banking: does banking shift the proposal distribution?

    Yes, but it depends on difficulty. On easy one-step problems it just pulls answers it already knew into its top guess (60% to 80%), gaining no new ground. On harder two-step problems something new happens: on brand-new t

  • 2026-07-03 Qwen3.5-4B Activation Steering: is the latent first-op causally usable?

    No. A simple reader picks the model's planned first operation out of its internal state almost perfectly — 99% of the time — yet pushing that exact signal back in during generation barely moves what the model does: at be

  • 2026-07-02 Qwen3.5-4B Simulation Keystone Repair

    No. Training made the model trace a multi-step process almost flawlessly — even on longer chains and steps it never studied, so it learned a genuine skill, not memorized answers. Yet every task that supposedly needs trac

  • 2026-07-02 Qwen3.5-4B Depth-Wall Anatomy

    It's figuring out the steps. Handed the exact sequence of operations, this 4-billion-parameter model writes correct code almost every time, even four steps deep, with zero execution deficit. Left to infer that sequence f

  • 2026-07-02 Qwen3.5-4B Cross-Family Laws

    No. Handed the exact steps, this fixed 4-billion-parameter model wrote correct code almost every time across three unrelated task types. But asked to infer the same procedure from example inputs and outputs alone, it fel

  • 2026-07-02 Qwen3.5-4B Context Composition

    Only when its answer survives. The fine-tuned skill is genuinely the sharpest — 95% correct when the model replies in the required form, beating the untrained model's 83% under the same step-by-step procedure. But the tr

  • 2026-07-01 Qwen3.5-4B Decompose-and-Compose Frontier

    Yes — but not because the model got smarter. Taking three steps one at a time, with a tool that runs each and shows the result, solves about 2 in 5 versus 1 in 8 in one shot. The catch: blindly trying all 23 operations d

  • 2026-06-30 → 07-01 Qwen3.5-4B Neurosymbolic REPL Substrate + Failure Profile

    No. Seeing its real error barely helped: the fix-it loop solved 29% of puzzles versus 34% for simply drawing five independent attempts and keeping the best, at equal compute; the error message itself added just two tasks

  • 2026-06-30 Qwen3.5-4B Thinking Content vs Compute

    It genuinely reasons; the boost is content, not compute. Blank filler of the same length, and the real thinking scrambled into nonsense, both scored like skipping thinking entirely, around 74 to 75 percent. Only coherent

  • 2026-06-29 → 30 Qwen3.5-4B Thinking Separability Probe

    Yes, but not for the reason you would expect. From one snapshot of internal activity, whether the model's own code is correct is readable well above a coin-flip, about 64 to 76 percent of the time. Thinking first sharpen

  • 2026-06-29 → 30 Qwen3.5-4B Generator-Verifier Gap

    Only after it thinks. Judging on sight, the model rubber-stamps 91% of its tries as correct while just 77% truly pass — barely a check, mostly agreeing with itself. Given room to reason first, it becomes a real critic: l

  • 2026-06-28 Counterfactual Episodic ICL Posttraining

    Yes. Untrained, a 4-billion-parameter model solved 23% of real text-transformation tasks perfectly; after this training, 57% — but only when it could see the prompt's examples. Scramble those examples and it fell to 17%,

  • 2026-06-28 Counterexample-Guided Ephemeral Program

    No. Answering each row directly solved three-quarters of tasks completely, while the best rule-program the method could pick solved only about four in ten — and a perfect picker that peeks at the answers did no better. T

  • 2026-06-28 Qwen3.5-4B Tool State Policy LoRA

    Surprisingly, yes, and with almost no learning. Trusting the program only when it passes a worked example and disagrees with the quick answer lifted accuracy from 56% to 66%, exactly matching the best any picker could re

  • 2026-06-28 Qwen3.5-4B Live Tool DAgger

    Only its cost, not its accuracy. Answering directly solved none of twelve unseen tasks; letting the model write and run code recovered two — about one in six — and spoiled nothing it already had right. But even a flawles

  • 2026-06-28 Qwen3.5-4B Foofah Program Strategy Portfolio

    Yes, but modestly, and which passing program you trust matters more than the programs. Asking directly for the finished table got 42% right. The rule the team locked in, commit only when two programs agree, reached just

  • 2026-06-28 Qwen3.5-4B Foofah Adaptive Program Budget Router

    Yes, and here is the twist: running all five programs on every task scored lower (56%) than the cheap rule (58%), because blanket spending overwrote one answer the quick pass already had right. The rule fires only when t

  • 2026-06-28 Qwen3.5-4B Adaptive Tool Controller

    Partly. One structural cue — does the direct answer have fewer columns than the raw data implies? — safely flags the reshaping tasks where a program helps, lifting accuracy from 42% to 50% with zero broken tasks. But it

  • imported 2026-07-12 Sampled Query Filter Executor Experiment

    Yes. Graded on just one sampled final answer per problem, the model rebuilt the entire set of still-possible number pairs it was never shown, capturing 94 to 98 percent of it, because holding that full set is the cheapes

  • imported 2026-07-12 Qwen Slot Repair Distillation

    No. A correct program almost always sits one or two edits away — a search that peeks at the answer lifts solve rates from about a quarter to roughly 86%. But the blind helper couldn't pick which edits to make: on reworde

  • imported 2026-07-12 Qwen Register Trace Refiner

    Rarely. Even a flawless picker that always grabbed the correct edit reached only 37% on plainly worded problems and 7% on reworded ones, because the correct program usually isn't among the roughly 1,300 nearby edits at a

  • imported 2026-07-12 Qwen Register-Token Latent Compiler

    Only for short chains. Up to twelve steps it builds the correct hidden program about nine times in ten, while stripped-down versions trained on the final answer alone never find the interface and stay at chance. But at t

  • imported 2026-07-12 Qwen Register-Token Structured Runtime

    Only up to a point. For chains of four to twelve steps the hidden program runs flawlessly, at 100 percent. But at 24 steps exact execution collapses to 25 percent — versus about 1 percent from pure guessing. The catch: e

  • imported 2026-07-12 Qwen Python-Shaped Silent Executor

    No. The silent hidden steps never cleared single digits: about 8% at best on the simplest programs, 4.5% on unseen longer ones, barely above a zero untrained model and near the 3% you would get by guessing. The tell: scr

  • imported 2026-07-12 Qwen Progressive Repair Compiler

    Partly. The correct fix sits in the candidate list almost nine times in ten, yet the judge finds it only about half the time, lifting exactly-correct programs from 30% to 49% against the 88% a flawless chooser would reac

  • imported 2026-07-12 Qwen Learned Repair Verifier

    Partly. Without ever seeing the true answer, the judge lifted correct execution from 30% to 47%. But a checker allowed to peek at the answer key found a correct fix already sitting in the candidate pile 88% of the time,

  • imported 2026-07-12 Qwen Latent Beam Program Compiler

    Yes, up to a point. It wrote exact programs that a fixed calculator ran perfectly at eight and twelve steps, and its answers matched those programs — so it truly computed rather than guessing, where chance is about one i

  • imported 2026-07-12 Qwen Candidate-Trace Verifier

    Yes. A small checker that reads each candidate's worked-out steps — never the true answer — lifted correctly-running programs from 30% to 54%, and to 56% when it cross-checks two wordings of one task. The twist: a picker

  • imported 2026-07-12 Qwen 3.5 4B Verified Edit Closure

    Yes, on the hardest unseen tasks. Testing small edits of the model's own near-miss program and keeping whichever passes the example cases raised fully-correct answers from 39% to 52%, while re-sampling the model and re-r

  • imported 2026-07-12 Qwen 3.5 4B Typed Sketch Synthesis

    It depends on difficulty. On the hardest problems, sketching the shape and letting a verified search fill the blanks lifted correct fixes from 33% to 78%, and a safe blend of both methods reached 88%. But on easy problem

  • imported 2026-07-12 Qwen 3.5 4B Static Bridge Ceiling Breaker

    Partly. Folding in just 60 slightly-harder "bridge" examples, a quarter of the training budget, more than doubled success on deeper, never-seen programs, from 20% to 44% fully repaired, with no loss on familiar skills. B

  • imported 2026-07-12 Qwen 3.5 4B GraphIR Self Repair

    No. The plain one-line formula fully solved 29% of brand-new tasks; the step-by-step diagram managed only 22%, and the fix-it pass recovered part of that gap to 24% — still behind. That pass genuinely works: on randomly

  • imported 2026-07-12 Query Filter Executor Experiment

    Yes, and that is the surprise. Trained only to name one final answer, the model taught itself to run each instruction in order and track every still-possible pair of the two hidden numbers. Given enough internal steps, i

  • imported 2026-07-12 Latent Recurrent Executor Experiment

    Yes, but only under strict conditions. Built to hold each running total and trained to hit every intermediate value, the network's exact-answer rate stayed under one percent until its private-step count reached the numbe

  • imported 2026-07-12 Joint Register Executor Experiment

    Yes, but only with the right memory. When the model held every allowed number-pair together and spent one thinking step per instruction, its confidence in the exactly correct answer set jumped from near-random (about 3%)

  • imported 2026-07-12 Dense Supervision Ladder Experiment

    It's the feedback. With the model held fixed, training it on only one sampled final answer left it weak; showing it the full odds of every possible answer at every step roughly doubled how often it solved the hardest 24-

  • imported 2026-07-12 Dense Latent Query Executor Experiment

    Partly. Giving the model enough internal thinking steps to walk through the whole program lifts accuracy sharply, and it clearly beats a model that reads everything in one glance. But the memory stays approximate: at its

  • imported 2026-07-12 Belief Filter Executor Experiment

    Yes. When the model runs one internal update per instruction, it lands over 91% of its confidence on the exact set of still-possible answers, versus under 14% when it stops before finishing. It even runs programs three t

  • imported 2026-07-12 Adaptive Cognitive Kernel

    No advantage. The self-rewiring is genuinely doing ordered work: scrambling the operation order collapses its step-by-step accuracy from about 12% to 2%, and switching the rewiring off cripples it. But it never beats a p

  • 2026-06-27 Real Transform ABI Gate with Counterexamples

    It depends on how messy the data is. For clean, spreadsheet-style pipeline jobs the toolbox covered every one (100%) and held firm even against deliberately tricky examples. For irregular date, ID, and text cleanup, cove

  • 2026-06-27 Qwen Recursive Ephemeral Program Induction

    Only when you check the rule first. On its own, a model writing and applying a reusable rule solved 40% of tasks perfectly versus 56% for plain row-by-row answering, and it broke six tasks direct answering had solved. Ad

  • 2026-06-27 Qwen Real Task ABI Coverage Gate

    It depends, and the split is sharp. A frozen kit of reusable office operations, with no training at all, assembled 84% of realistic tasks from stored parts alone, far above the 21% a bare kit managed, and it fully solved

  • 2026-06-27 Qwen Public PROSE ABI Gate

    No. The frozen toolkit fully solved only 19% of the 309 outside tasks. In 77% of them no recipe fit even the worked examples, so the toolkit lacked that operation entirely; under 4% overfit. Yet a small four-billion-para

  • 2026-06-27 Noisy Row Program Crystallizer

    No. Keeping the model's direct per-row answers fully solved half of the 40 tasks, while distilling those noisy answers into one fixed rule solved just 22.5% — worse even than a scrambled comparison at 25%. And a flawless

  • 2026-06-27 Qwen Disagreement-Probe Program Induction

    No. The disagreement quiz picked the same programs whether its judge answers were real, randomly assigned, or skipped entirely — all three landed at 64% of tasks fully solved. The only genuine gain came from a plain caut

  • 2026-06-27 Qwen Active Crystallizer Public Gate

    No. Using the model's votes to choose a rule worked on 25% of tasks — barely above the 22.5% you get from scrambled, meaningless votes, and it never beat the best rule the candidate pool could offer. The model answered i

  • 2026-06-27 Qwen3.5-4B Transform ABI Compiler Pilot

    Yes. After a light round of tuning, the model chose a recipe that produced the correct output on all 48 test tasks, matching a perfect answer key and beating the untuned model's 92%. It recovered the harder multi-step ch

  • 2026-06-27 Qwen3.5-4B Independent Code ABI Coverage Gate

    No, hardly any. The locked toolbox solved only about 14% of brand-new tasks, roughly 1 in 7, versus 37% on the familiar tasks it was shaped around. Reshuffling which tasks are unseen barely moves it, around 18%. And near

  • 2026-06-27 Qwen3.5-4B Foofah Selective Program Fallback

    Trust the program the moment it reproduces the visible worked examples. Doing that lifted exact-match accuracy from 55% to 62% across 250 table tasks, rescuing 18 answers the direct route got wrong while losing none it g

  • 2026-06-27 Qwen3.5-4B Foofah Program Repair Agent

    No. Guessing the answer directly won outright, solving 55% of unseen tables versus only 25% for the debugged program. But the program is a useful complement, not a replacement: it rescued 18 tables the direct guess botch

  • 2026-06-27 Qwen3.5-4B Foofah Program Ensemble Consensus

    No. The simplest rule won: run the first program that passes a single worked example. It solved 52% of tables versus 44% when the model just answered directly, rescuing 23 tables it had otherwise botched while breaking o

  • 2026-06-27 Qwen3.5-4B Foofah Ephemeral Program Induction

    No. Asking directly reshaped 55% of tables correctly; the write-and-test-a-program route managed just 15%. Even a magic chooser that always picked the right route each time would reach only 59% — four points above asking

  • 2026-06-27 Qwen3.5-4B Foofah Direct vs ABI

    Just ask the model. Directly generating the reshaped table got 55% of 250 table tasks exactly right, versus only 18% for the fixed-operation converter. The model even nailed 103 reshapes the converter could not even expr

  • 2026-06-27 Qwen3.5-4B Code ABI Oracle Coverage Ladder

    Yes, mostly. Snapping together at most three verified functions produced a passing solution for 84% of 160 small programming tasks, up from just 13% with a bare-bones library. But the library is a foundation, not a solve

  • 2026-06-27 Qwen3.5-4B Code ABI Compiler Heldout Primitive Pilot

    No. The block kit could build a working solution for 84% of the problems it was tuned against, but only 14% of brand-new ones — a 70-point collapse. Three fresh random batches of new problems all landed near 18%, so it w

  • ~2026-06-27 Qwen 3.5 4B Balanced Discriminative Bridge

    An even mix, clearly. Sixty evenly-spread ordinary examples across ten new program types lifted the model's success on unseen hard problems from 60% (with no examples at all) to 99% fully solved. Hand-picking only the tr

  • 2026-06-26 Qwen Trace Procedure Depth Stress

    Yes. Trained only on single-step tasks, the 4-billion-parameter model wrote six-step procedures that ran correctly 63% of the time once a plain step-follower executed them, versus 0% when the model tried to state the fin

  • 2026-06-26 Qwen Program-Only Executable ABI

    Yes. Teaching a small model to write a short runnable program lifted brand-new multi-step accuracy from about 44% (just stating an answer) to 73%. And when it wrote out its steps plus an answer, the steps ran correctly 9

  • 2026-06-26 Qwen Large ABI Nested Compiler

    Two things. Making the library four times larger, from 32 to 128 operations, did not hurt straight-line programs at all: both stayed perfect on 16-step chains. But branching is a separate skill. Models shown only straigh

  • 2026-06-26 Qwen Extrapolation-Bound ABI

    Yes. A model trained only on procedures up to three steps long reliably writes correct sixteen-step procedures, with accuracy climbing from 61% under single-step training to a perfect 100%. Surprisingly, adding longer tr

  • 2026-06-26 Qwen Crystallized Trace ABI Tournament

    Barely. On familiar inputs, writing out each step scored 94% versus 92% for answer-only, a two-point edge that cost three to six times more generated text. On genuinely new combinations of steps the model had never seen,

  • 2026-06-26 Qwen Constrained ABI Parser

    Yes, mostly. On the hardest six-step requests, blocking any invalid step as the model writes lifted correctly-running recipes from 60% to 75%, and it won on all five training runs. It also beat merely re-rolling until va

  • 2026-06-26 Qwen Compositional Curriculum ABI

    Yes. Adding a few two- and three-step examples lifted correct answers on unseen six-step problems from 72% to 83%, and on eight-step problems from 78% to 89%. But adding only two-step examples did nothing (72% stayed 72%

  • 2026-06-26 Qwen3.5-4B Substrate Coverage Ladder

    Only when a building-block was hand-carved for each specific problem. A shared library that reshapes old solutions to fit new ones solved zero of the nine stuck tasks, no better than nothing. Custom-written parts solved

  • 2026-06-26 Qwen3.5-4B Reliability Exec OPSD Audit

    No. All three cheap fixes failed. Ranking programs by the model's own confidence picked working code less often than just taking the first one that passed the visible sample tests (6 of 24 versus 8), and let more bugs sl

  • 2026-06-26 Qwen3.5-4B Prefix Value Guided Search

    No. Even when the half-finished drafts were graded with perfect knowledge of the hidden answer tests, finishing only the best-graded one solved the same share of problems as plain full-solution writing — 75% either way.

  • 2026-06-25 Qwen Tail Repair Stability Critic

    No. A correct fix sat among the candidate rewrites for about nine in ten programs, but the trained judge could not tell which rewrite was right from summary statistics alone, so it played safe and edited nothing, staying

  • 2026-06-25 Qwen Readable Candidate Verifier

    Barely. Training a picker on the model's reading of each fix's pseudocode plus its claimed output raised the share of tasks fixed correctly from 44% to 51%, closing only about 15% of the distance to a perfect picker's 90

  • 2026-06-25 Qwen Candidate-Conditioned Trace Verifier

    No. Doing nothing already solved about 44% of tasks, and always choosing a correct candidate could reach 90%. Yet no trained picker captured any of that headroom; the best merely tied doing nothing. Left unchecked, the m

  • ~2026-06-25 Qwen3.5-4B Oracle-Distilled Semantic Verifier

    Yes, but only on home turf. On the problem set it trained on, the trained judge picked a genuinely-correct program 81% of the time, up from 72% for the same model untrained and just 44% for grabbing the first candidate t

  • ~2026-06-25 Qwen3.5-4B HumanEval Adaptive Evidence Budget

    No. Every strategy landed at the same 16.7% correct-pick rate — running no tests, running all eight, and even a strategy allowed to peek at the right answer to time its stop. The trained model did learn to quit sooner, u

  • 2026-06-25 Qwen3.5-4B Diversity-Keyed Coverage Gate

    Mostly the second. Of 24 Python problems a 4-billion-parameter model missed on four tries, spending more and more varied sampling recovered 15, lifting the share solved from 70% to nearly 89%. Mixing three creativity set

  • 2026-06-24 → 25 Qwen Compiler Multi-Seed Reattribution

    No. The best schedule averaged 43% correct on plainly worded problems, yet the identical training swung from total failure to 81% just by changing the random starting number — so no schedule earns credit for the wins. Bo

  • 2026-06-24 Qwen VM-ECHO Trace Distillation

    Mostly no. The model got far better at predicting execution, with reading a running program's top value climbing from under 1 percent correct to 43 percent, but that rarely improved the programs it wrote. First-try accur

  • 2026-06-24 Qwen VM-Agent ECHO QLoRA

    Only the acting helped. Editing-and-running in a loop lifted the share of tasks solved from 10% at a blank start to 43%, beating a single one-shot guess near 37%. But adding a second job, predicting the program's output

  • 2026-06-24 Qwen Structural Latent Compiler Expansion

    Yes, mostly. The model fills fixed slots with operations a calculator runs — no text, no trying many guesses and picking one. After the learned short form was copied into bigger ones, it stayed perfectly correct on 8- an

  • 2026-06-24 Qwen Structural Compiler Attribution Ablation

    It's the practice schedule. Give the model its full size from the start, then feed examples easy-to-hard — 8 steps, then 16, then 24 — and it solves nearly every standard 24-step program (about 97%). The popular guess, g

  • 2026-06-24 Qwen Search-Augmented Rollout Distillation

    No. An automatic search found a verified correct fix for 98% of the dead-ends the model wandered into, yet retraining on those single fixes matched or trailed the simpler training on four of five test sets. The model cou

  • 2026-06-24 Qwen Recurrent VM Repair Policy

    Yes, but only partway. Letting the model run its program, read the output, and fix one line at a time roughly tripled accuracy, from about 11% to 34% on the main test and 16% to 40% on reworded prompts. But a perfect edi

  • 2026-06-24 Qwen In-Policy VM-ECHO Distillation

    Only partly. It became excellent at predicting mechanical outcomes, like how deep a program runs (about 92% right), but stayed no better than a coin flip at judging which program is actually correct. So it could not rank

  • 2026-06-24 Qwen Fuyu VM GRPO-ECHO

    No. Copying worked solutions alone solved about 10% of tasks; one round of reward-based self-coaching dropped that to 7-8%. The coaching produced encouraging signals — it could rank good fixes over bad ones 80% of the ti

  • 2026-06-24 Qwen Dense-State DAgger VM Agent

    Mostly yes. Teaching the model to make one edit at a time, then correcting it on the messes it made, roughly doubled accuracy on mixed tasks (22% to 41%) and won on four of five task types. But it lost on ordinary tasks

  • 2026-06-24 Qwen Counterfactual Trace Preference Distillation

    Barely. The self-grader learned to favor programs that run without crashing, but it identified the truly correct one only about 15% of the time, against a 41% best-possible ceiling. On fresh questions it picked worse tha

  • 2026-06-24 Qwen Action-Conditioned VM-ECHO Policy Iteration

    Barely. Learning to grade drafts by their run results nudged picking accuracy only from about 10% to 11%, far short of the 37% reachable by always choosing the best available draft. It reliably picked programs that ran,

  • 2026-06-24 Qwen3.5-4B Oracle Process GRPO

    Yes, but with sharp limits. Teaching the model to copy a flawless test-picker lifted its success from about a third of puzzles to 43%, versus roughly 31% for picking blindly. Layering on preference-based and reward-based

  • 2026-06-24 Qwen3.5-4B Learned Active Trace Policy

    It depends. On the main test set a simple even-splitting rule beat the trained picker after one extra input — 91 percent of programs fully correct versus 87 — and the best-possible choice reached 97 percent. The picker d

  • 2026-06-24 Qwen3.5-4B Joint Shortlister Ladder

    No. Across every version — untrained, trained, and with the glossary's descriptions scrambled — the model got both codes exactly right zero percent of the time, even when allowed sixteen guesses. Training pushed single-c

  • 2026-06-24 Qwen3.5-4B Adaptive Evidence Budget Policy

    Yes. Trained just for this single decision, the model matched the accuracy of running all ten checks — solving about 92 of every 100 tasks — while using only about five checks on average, near a perfect-hindsight stopper

  • 2026-06-24 Qwen3.5-4B Active Counterexample Trace Selection

    Yes, but choose them well. Committing on the visible examples alone left about a quarter of picks secretly wrong, even though every one passed all the examples shown. Requesting six new test cases where the surviving pro

  • 2026-06-23 Qwen Typed Bytecode Expert Iteration

    It depends. Training the model only on its own attempts that landed on the correct final answer lifted unaided first-try accuracy from 62% to 73% on fresh problems — a real gain that sticks when it writes programs alone.

  • 2026-06-23 Qwen Semantic Prefix Value Model

    No. Scoring each step by whether a correct answer is still reachable pushed the top pick to about 68 percent, level with plain confidence search and short of the 71 percent from scoring steps against the known correct pr

  • 2026-06-23 Qwen Prefix-State Process Verifier

    Barely. The judge got genuinely good at telling promising partial programs from dead ends, yet it lifted first-try accuracy only a few points, with hard problems going from 41% to 44%. The real gap: on hard problems a co

  • 2026-06-23 Qwen On-Policy Repair-to-Compiler

    Yes, but the self-correction isn't the secret ingredient. On brand-new requests the model climbed from about 29% correct programs to 99% after training on its own corrected drafts. The catch: training on plain correct pr

  • 2026-06-23 Qwen Mixed-Domain Trace Verifier

    Yes, partly. On fresh tasks the frozen model alone got 46% right; the proofreader lifted that to 57%, and a cross-check that compares reworded versions of the same task reached 61%. But a correct recipe was already among

  • 2026-06-23 Qwen LoRA Typed-Bytecode Trace Compiler

    Yes — but the win came from the teaching material, not from adapting the model. Fed fully worked recipes, it wrote a runnable recipe that reached the right answer about 68% of the time, versus only 15% when taught with f

  • 2026-06-23 Qwen Iterative Repair Policy

    Yes. The frozen model alone got about 30% of programs exactly right; editing one step at a time lifted that to 53% on fresh problems, closing roughly 38% of the distance to the best a perfect fixer could reach (89%). The

  • 2026-06-23 Qwen Hidden VM On-Policy Canonical Repair

    No. Training the model on automatically corrected recipes reached 61% on new tasks, versus 59% for plain training — a 2-point gap that is basically noise, and it left longer tasks no better. The corrections are genuinely

  • 2026-06-23 Qwen Hidden VM Mixed Domains

    Yes. Guessing the answer directly worked only about 15% of the time across six kinds of problems — arithmetic, dates, unit conversions, list totals, yes/no thresholds, and lookups. Having the model instead write a hidden

  • 2026-06-23 Qwen Hidden VM Curriculum Repair

    No. Feeding it nearby worksheets that merely land on the correct answer wrecked it. Plain step-by-step training scored 72% on the main mixed test; the same model after answer-chasing repair fell to 35%, and collapsed on

  • 2026-06-23 Qwen Context-Conditioned Trace Verifier

    Mostly no. A correct program was almost always sitting in the candidate pile, for 91 to 100 percent of questions, yet the judge barely helped: it nudged easy questions from about 69 to 70 percent and actually made the ha

  • 2026-06-23 Qwen Complete-Program Trace Reranker

    Barely, and it backfires on hard cases. A correct program sat in the candidate pile 91 to 100 percent of the time, yet the trained picker nudged easy prompts only from 69 to 70 percent and actively hurt longer ones, drop

  • 2026-06-23 Qwen Budgeted Action-Value Compiler

    Barely. On fresh problems the model drafts a correct program among its candidates 81% of the time but ranks it first only 67% of the time. A learned scorer that never sees the answer nudged that to just 70% — while a con

  • 2026-06-22 Qwen Verifier-Guided Slot Repair

    Yes, mostly, but with a catch. A small model copying 24-step calculations got only about 27% exactly right on its own. A checker that knows the correct running number after every step, allowed to swap one or two bad step

  • 2026-06-22 Qwen Teacher-Distilled Slot Compiler

    No. A model trained to copy numbers and operations out of text and run a 24-step calculation got 27% of final answers exactly right; adding the pointing signal landed at 28%, a tie. Worse, agreement between two rewording

  • 2026-06-22 Qwen Checkpoint-Selected Scheduled-State Compiler

    It depends. A light, steady dose of show-your-work coaching, kept on through the hardest problems, lifted correct answers on 24-step chains from 25% to 33% and nearly doubled agreement between two wordings of the same pr

  • 2026-06-22 Qwen 3.5 4B Unsaturated Frontier Active Bridge

    Spread evenly. Giving each of ten problem types the same six extra correction examples let the model fully fix 98% of hard cases. Piling those same examples onto whichever types it failed most reached only 85%, and starv

  • 2026-06-22 Qwen 3.5 4B Model-In-Loop Counterexamples

    No. Building practice cases from the model's actual wrong answers matched, but never beat, simply hand-picking the tricky categories in advance. Both lifted the hardest problems from 64% to a perfect 100% passing every h

  • 2026-06-22 Qwen 3.5 4B Executable Program Posttraining

    Yes, but with a catch. Shown worked-through reasoning in the prompt, the model fixed unseen problem types about three-quarters of the time, versus one-in-three when the prompt showed no steps. Strip out or scramble those

  • 2026-06-22 Qwen 3.5 4B Counterexample-Directed DSL

    It depends. Hand-picked examples lifted the model's single-best-guess repair rate from 51% to 58% over random examples. But when it generated several candidates and kept the best, the edge vanished (64% slipped to 61%).

  • 2026-06-21 Structured Slot Initializer Ladder Experiment

    Structure, not scale. A plain general-purpose network placed only 55.5% of its belief on the correct starting setup and gave barely half the possible values their own slot, doubling several onto the same one. A rule that

  • 2026-06-21 Sparse Support Memory Executor Experiment

    Only when its scratch memory held one slot for every possible starting value. With that, it answered every question correctly through the longest 24-step programs. Cut the memory roughly in half and accuracy fell to abou

  • 2026-06-21 Qwen Trace Bootstrap Retention Experiment

    Yes. Once step-by-step labels install the skill, training on final answers alone preserves and even sharpens it: 97% of the longest 24-step problems solved exactly, versus about 1 in 100 — no better than guessing — when

  • 2026-06-21 Qwen Structured Bridge Experiment

    Yes, but only when you show it the individual steps during training. A tiny translator turning the frozen model's read into calculator instructions solved chains far longer than it trained on: 96% correct at twelve steps

  • 2026-06-21 Qwen State-Ladder Compiler

    No. Grading the running total after every step never beat an identical model graded only on its final answer, and at full strength it collapsed on the hardest long programs. The real winner was the training schedule: sta

  • 2026-06-21 Qwen Span-Free Compiler

    Only when it is first taught where to look. Fed just the frozen model's raw internal notes, a plain reader stayed near random guessing (about 1 in 97). Adding training that also highlighted which spots held the numbers a

  • 2026-06-21 Qwen Slot-Stability Compiler

    Yes, mostly. The point-and-compute helper solved about 91% of short problems and around half of medium ones, while training the same model to just emit the final answer never beat random guessing, about 1 in 60, at any l

  • 2026-06-21 Qwen Shared Parser Compiler

    Only when every step is taught directly. Given step-by-step labels, the add-on rebuilds short programs well — nearly 4 in 5 four-step problems run exactly right — but accuracy fades to 39% at twelve steps and under 1% at

  • 2026-06-21 Qwen Numeric-Copy Compiler

    Yes. When the model just points to where each number and operation sits and copies the exact symbols for a hidden calculator to run, it solves four-step problems about 90% of the time. A version trained to write the answ

  • 2026-06-21 Qwen LoRA Parser Compiler

    Partly. With step-by-step coaching, a small four-billion-parameter model's hidden states became a readable program: it named the starting number every time and picked the right operation about 98 percent of the time, whi

  • 2026-06-21 Learned Sparse Slot Executor Experiment

    Yes, but only small. With eleven possible values and a ready-made scratchpad, the network kept the fully correct answer in view 95.5% of the time, even on longer chains than it trained on. Widen to thirty-one values and

  • 2026-06-21 End-to-End Structured Slot Executor Experiment

    Yes, but only when both halves carry built-in structure. On the numbers 0 to 30, the full model puts 98% of its confidence on the exactly correct final set of possibilities, nearly matching a version handed the answer, a

  • 2026-06-21 Dense Teacher Distillation Experiment

    No. Even with a flawless teacher revealing the exact set of still-possible answers at every step, the fixed-size memory learned only a rough approximation. The best version placed 52% of its confidence on the correct fin

  • 2026-06-21 Cyclic Transition Ladder Experiment

    It needs the matching wrap-around parts. A network built from clock-arithmetic moves stayed perfectly exact on programs three times longer than it practiced on. A plain generic network of the same size drifted down to ju

  • 2026-06-19 → 20 Execution-Conditioned Repair LoRA Experiment

    No. On bugs built from the same templates it practiced on, the fixer repaired all 60 of 60 cases, versus 11 of 60 with ordinary patch training and 6 of 60 with no training at all. But on bug types it never saw, every met

Claims

Queued proposals 4