Research log Small Model Experimentation
GitHub

Interpretability and Diagnostics

Measure why methods work or fail through attribution, probes, pressure audits, and controlled ablations.

What we have learned

Seed Experiments

Current Read

Diagnostics should become standard infrastructure. They are how future agents avoid retesting the same mistaken explanations.

  • qwen35_4b_probe_to_prompt (claim C30): EXTERNALIZING the latent readout (decode C19's first-op probe -> inject as a PROMPT hint) elicits deployable depth-2 (oracle_full 6x) where steering (C20) was inert -- the first test-time lever to move the wall. But the decodable op-TYPE only narrows sampling; the PARAMETER is the deployable bottleneck, so the type-only probe nets to zero. Graded by depth (fades at depth-3 thread). Layer-0 leak control at chance.

  • qwen35_4b_probe_the_parameter (claim C31): sharp localization of C30 -- the op-TYPE is MODEL-LATENT (residual probe 0.41 > external-I/O baseline 0.27) but the PARAMETER is SURFACE-READABLE (external I/O 0.53 >= probe 0.49; and surface-hint deploys 0.027 > probe-hint 0.014). The forward pass computes the type (elicitable) but only reads the param off surface I/O. Real surface control = external classifier on raw I/O features (the last-token layer-0 probe is degenerate under RoPE).

  • qwen35_4b_jacobian_value_transport (unclaimed while the ledger re-grade is open): an averaged token-Jacobian coordinate is strongly writable at one late layer but does not transport. On an untouched 24-item confirmation split, the layer-24 intervention changed the direct concept report on 18/24 items (75%; random 0/24, logit-lens 5/24), while an arbitrary prompt-local consequence changed on 0/24 at every tested layer. Its direct margin crossed zero with intervention strength while the consequence margin stayed flat. G0 correctly cancelled thought-prefix value and task patching. This separates three properties that prior probe work often conflated: decodability, local writability, and downstream transport. Scope: the random control was not exact realized-delta-norm matched, so direct J specificity remains provisional; the transport failure and adjacency failure are unaffected.

  • qwen35_4b_context_local_jacobian_clamp (unclaimed while the ledger re-grade is open; terminal INVALID_CONTROL): the corrected early selected-token clamp produced the full semantic-transport signature on 48 untouched mappings. All-24 J changed direct key and mapped digit on 48/48; pair J changed the digit on 47/48; wrong-donor J changed 48/48 to its own digit; concept logit lens and random changed 0/48. Full donors also exposed a sharp causal window: four bands through 16–20 were 24/24, bands 20–24 and 24–28 were inert. Yet one of 96 random rows missed the frozen realized-norm tolerance (1.155e-5 > 1e-5), and bf16 rounding left up to 5.7% realized J-span projection despite pre-cast orthogonality. The result cannot be promoted. It sharply prioritizes a fresh quantization-aware control replication and keeps native-thinking continuation gated.

  • qwen35_4b_jacobian_transport_control_replication (unclaimed while the ledger re-grade is open; terminal REPLICATED_J_TRANSPORT): fresh exact-control replication resolves the parent invalidity. All-24 J changed 48/48 direct keys and 48/48 separately computed mapped digits. Two independent random arms and the concept logit lens changed 0/48; wrong-donor J produced its own key/digit 48/48 and the registered target 0/48; pair J reached 46/48 consequences. All 480 calibration and 960 confirmation post-bf16 control-layer rows met relative norm <=1e-5 and realized J-span projection <=0.01. Both paired bootstrap intervals were [1,1]. This establishes an oracle, context-local, causally consumed concept state on the procedural lookup substrate and unlocks native-thought work. It does not install capability: target donor identity remains supplied.

  • qwen35_4b_native_thought_jacobian_value_transport (unclaimed; terminal NO_NATURAL_SEAM): the first licensed native-thought successor stopped before value fitting. All 48/48 frozen 160-token traces on 16 fresh identifiable tasks hit the thought cap without natural close; parse, success, and mixed-task counts were zero. Model smoke also found up to 0.0625 historical-token activation drift across different suffix lengths. This is not evidence against J-space value: the natural answer seam was unreachable. It requires a fresh cap-selection/confirmation experiment and dynamic per-length control geometry before causal patching.

  • qwen35_4b_native_thought_seam_budget_ladder (unclaimed; terminal NO_BUDGET_SELECTED): the separate frozen 256/512/1024 selector also found no natural seam. All 48/48 selection traces used all 1,024 thought tokens, so natural close, parse, usable-prefix, and mixed-task counts were zero at every paired rung; the 24-task confirmation split remained unopened. All rows passed the cached-forward audit. A post-decision token diagnostic found 0/48 exact short-period 256-token tail loops, so loop breaking is not licensed. This still precedes J-space value. It closes the natural-cap branch on this workload and redirects the next test to an explicitly deployed forced-commit policy with C51 parse/headroom gates and per-length post-bf16 controls.

  • qwen35_4b_forced_commit_jacobian_value_transport (unclaimed; terminal FORCED_COMMIT_SEAM_FAIL): explicitly deploying the close token did not create a usable emission seam. Across 48 paired traces, forced parse was 6/48, 8/48, and 9/48 at caps 256/512/1024; exact success was 1/48 at every cap; only one task mixed outcomes; and 41--46/48 answers hit the 16-token answer cap. A post-decision EOS-tolerant parser raised parse only to 7/11/10 and correctness to 1/2/2, leaving every gate failed. Confirmation, value, and causal splits stayed sealed. This localizes the next interface change: close is a mode delimiter, not an answer slot. A fresh syntax-only First: slot may test semantic alias choice while close-only remains control.

  • qwen35_4b_commit_slot_jacobian_value_transport (unclaimed; terminal COMMIT_SLOT_SEAM_FAIL): fixed First: syntax repaired answer mode but did not earn the semantic-variation gate. At cap 1,024, real ordered thought reached 15/48 versus 4/16 no-thought and 11/48 exact-token- multiset shuffle, clearing both frozen accuracy-gap bars, while an alias was already the unmasked top token on 41/48. Yet only five tasks mixed correct and incorrect traces versus six required; task-bootstrap intervals for both gains crossed zero, and effects were alias/task concentrated. Correct-alias mentions were not explanatory, and three post-decision label-free logit residuals all underperformed 15/48. Confirmation and every J stage stayed sealed. Power a fresh fixed-1,024 task-level replication before changing cap or fitting value.

  • qwen35_4b_commit_slot_semantic_power_replication (unclaimed; terminal seam POWERED_COMMIT_SLOT_SEAM_REPLICATED): the powered repair independently passed twice on 113 fresh tasks per stage. Qualification scored 92/339 ordered versus 46/339 shuffled with task lower +8.85pp; confirmation scored 98/339 versus 47/339 with task lower +9.44pp. No-thought was 11/113 then 8/113; mixed tasks were 32 then 31; unrestricted top-is-alias remained 88.2%/87.6%. Correct-answer mention strata did not explain success. A deterministic audit found ordered-only versus shuffle-only paired wins of 60:14 and 64:13. The seam is now usable for native prefix-value work, but target identity remains heterogeneous (one confirmation target 0/30; shuffle favored two targets). Any J result must add task-held-out signal beyond correct-alias activity, slot margin, and alias identity, with live-length post-bf16 controls. The licensed value run was complete but terminal NO_PREFIX_J_VALUE: shared J AUC 0.5021 (task lower 0.4417), below slot margin 0.5448 and equal-width non-J 0.5292. Midpoint prospective AUC reached 0.6083, but endpoint AUC reversed to 0.3958 and every incremental uncertainty bound crossed zero. The shuffled null passed at 0.5061, so this is a scientific negative, not invalid control. Causal data stayed sealed. A phase-specific midpoint hypothesis requires a distinct fresh replication and cannot rescue this shared-coordinate failure. The allowed post-decision audit then retired that hypothesis: midpoint-only J was 0.5375 (lower 0.4417), below matched non-J state 0.6000 and tied margin 0.5396; endpoint-only J was 0.4292. Phase-fit coefficient cosine was -0.0681. Do not spend fresh tasks or causal data on a midpoint J-axis replication.

  • qwen35_4b_counterfactual_order_support_selector tested the non-scalar forward counterfactual next. Mean per-alias ordered- minus-exact-shuffle probability reached 0.381 accuracy, beating majority 0.292 with paired lower +0.027, and recovered eight correct answers outside the three argmax choices. But it was only +0.018 over minimum entropy and +0.027 over max confidence with negative uncertainty bounds; a gold-balanced task-mismatched shuffle reached 0.389. Terminal NO_ORDER_SUPPORT_SELECTOR keeps confirmation absent. The causal group contrast is informative but not a task-specific value readout.

  • qwen35_4b_jacobian_counterfactual_branching moves intervention before commit but stops at NO_NATIVE_J_BRANCH_CONTROL. After five outcome-blind smoke/repair receipts, every live non-J control passes (max norm error 9.60e-6; J-span projection 0.00950). Yet centered additive J writes at the last token of four 512-token thoughts select their supplied alias only 4/48 at alpha 0.5, 1, and 2—exactly chance and identical to non-J. Mean target-probability lift peaks at 0.00566 versus 0.15 required. No correctness or continuation opens. Donor-coordinate context-local transport remains positive; additive direction transport at an arbitrary native token is now cleanly negative.

  • qwen35_4b_semantic_anchor_coordinate_branching is terminal INVALID_MECHANICS_CONTROL (parse), not a clean J-transport negative. All 880 numeric and 2,240 intervention rows match calibration after canonical identity sorting, and the constrained direct distribution is strongly writable: text/full donor reach 43/44, donor J 42/44, and wrong- donor J writes its own alias 42/44. But unrestricted parse is 56/880 and zero consequence rows parse; constrained consequence is only 6/44 text/full and 5/44 donor J, with +0.00170 probability lift. Post-run audit also found the two cyclic maps cancel, making alias -> result label fixed across tasks. No continuation opened. Retire the late opaque one-token interface and treat direct writes only as a diagnostic; this result cannot establish a general negative about native state/J transport.

Scorecard

  • Program: charter
  • Current read: early synthetic donor-coordinate J transport and the cap-1,024 semantic seam replicate, but shared/midpoint J value, terminal counterfactual selection, additive native branching, and the late semantic-anchor bridge do not yield a usable controller. Early concrete text supplies local causal routing. Tokenizer EOS plus no-think is a replicated 48/48 short-output seam against 0/48 HF-EOS controls, and mechanics transport is 24/24. The replay-hardened residual comparison is now cleanly negative: materialized and every matched comparator had 0/24 oracle proposal coverage despite healthy ABI and 24/24 exhaustive task solvability. Semantic addressing and a strict token-native commit boundary are real locally; semantic materialization does not create proposal competence.
  • Best next experiment: retire native J token/layer/scale, opaque-name timing, cheap materialized viability, and the all-candidate semantic-materialization prompt. If J-space continues, use a fresh measurement-first correctness coordinate inside <think> with task-held-out and equal-width non-J controls, then require a separate same-prefix intervention to increase correct-proposal coverage over matched sampling. Use a readable J coordinate as diagnosis only until that forward causal gate passes.
  • Strong anchors: qwen_structural_compiler_attribution_ablation, qwen35_4b_probe_to_prompt, qwen35_4b_jacobian_value_transport, qwen35_4b_context_local_jacobian_clamp, qwen35_4b_jacobian_transport_control_replication, qwen35_4b_native_thought_jacobian_value_transport, qwen35_4b_native_thought_seam_budget_ladder, qwen35_4b_forced_commit_jacobian_value_transport, qwen35_4b_commit_slot_jacobian_value_transport, qwen35_4b_commit_slot_semantic_power_replication.
  • Avoid repeating: probes that do not change the next experiment, next-token writing tests presented as reasoning transport, component permutations without a composed-map test, post-hoc parser repair, promoting a perfect point estimate past a failed control gate, or presenting an oracle donor as a deployable gain.
  • Evidence that advances the program: a diagnostic predicts which variants will fail before the final metric is observed.

Charter

Show charter.md

Purpose

Explain why methods work or fail through controlled ablations, attribution, pressure audits, trace diagnostics, state-prefix analysis, and failure taxonomies.

Why This Is A Program

The repository should not only collect wins. It should make each result more inspectable so future experiments build on mechanisms, not vibes.

Progress Signals

  • Diagnostics predict failure before full runs.
  • Ablations distinguish competing explanations.
  • Reports include failure slices, not just averages.
  • Diagnostic tools become reusable across programs.

Boundaries

This program studies explanation and measurement. It supports every other program.

Backlog

Show backlog.md

Next Experiments

  • qwen35_4b_native_thought_seam_budget_ladder is terminal NO_BUDGET_SELECTED: all 48/48 traces contacted even the 1,024 ceiling, with zero natural closes at every rung. Confirmation was correctly unopened. Do not add a larger natural-close rung or treat these rows as completed thoughts.
  • qwen35_4b_forced_commit_jacobian_value_transport is terminal FORCED_COMMIT_SEAM_FAIL: forced-only parse was 12.5%--18.8%, exact success 1/48 at every cap, and 85%--96% of answers exhausted 16 tokens, usually by restarting analysis. An EOS-tolerant parser diagnostic stayed <=22.9%. No cap, confirmation, value fit, or J causal outcome opened.
  • qwen35_4b_commit_slot_jacobian_value_transport is terminal COMMIT_SLOT_SEAM_FAIL: fixed syntax repaired answer mode and the 1,024 arm beat no-thought/shuffled by +6.25pp/+8.33pp, but only five tasks mixed outcomes versus six required and task-bootstrap intervals crossed zero. J value stayed sealed.
  • Replicated seam, terminal value negative: qwen35_4b_commit_slot_semantic_power_replication froze cap 1,024 and uses the calculated 113 fresh tasks per seam stage (~80% power at the parent ordered-over-shuffled effect), a task-bootstrap lower bound, 28 mixed tasks, and eight-alias success/choice support. It keeps no-thought as a +3pp point gate and opened value only after an identical untouched pass. Qualification passed (92/339 real versus 46/339 shuffled; lower +8.85pp), then confirmation independently passed (98/339 versus 47/339; lower +9.44pp). Correct/chosen breadth was 11/12 then 10/12. Resume task-held-out prefix-value work only with alias-identity, correct-logit, margin, and dynamic-length controls. The one value run then returned NO_PREFIX_J_VALUE: shared J AUC 0.502 (lower 0.442), below slot margin 0.545 and equal-width non-J 0.529. Midpoint prospective AUC was 0.608 but endpoint reversed to 0.396, so causal work stayed sealed. The permitted phase-specific audit reduced midpoint-only J to 0.538 (lower 0.442), below matched non-J 0.600 and tied margin 0.540; retire this J-axis successor and redirect to mechanisms that exploit coherent thought without assuming scalar J certainty.
  • Raw ordered-minus-exact-shuffle probability also failed as a vector-valued commit selector: 0.381 beat majority but not confidence/entropy robustly, and exact task matching did not beat an oracle-balanced mismatch. Do not tune the score; move Jacobian/counterfactual work before the commit to alter proposals or continuations under matched compute.
  • Balanced additive J branching is stopped before continuation: supplied-target control is chance at every norm-anchored alpha despite exact live controls. Any final native branch successor must restore the positive mechanism's explicit semantic anchor token and donor-coordinate replacement, with full- activation, text-hypothesis, additive-J, and non-J controls. Do not sweep larger alphas or layers on this last-token interface.
  • The final late semantic-anchor bridge is terminal invalid/unreachable, not a transport pass: unrestricted parse is 56/880, consequence parse is 0/440, text/full/donor-J constrained consequence is 6/44, 6/44, and 5/44, and the intended rotating alias-to-label composition was accidentally fixed. Do not repair it in place or run another native J token/layer/scale variant. A fresh deployable successor may move concrete text hypotheses before reasoning and test full continuations against matched-compute sampling; J is diagnostic context, not an intervention arm required for that successor.
  • If a future native thought-state transport experiment passes, train a non-oracle prefix controller and require a replicated held-out capability gain over frozen Qwen3.5-4B and matched-compute sampling. Oracle donor selection is a mechanism control, not the endpoint.
  • Build a standard failure-slicing template by operator, family, length, parse status, and evidence state.
  • Add attribution and ablation reports for high-performing compiler and selector lines.
  • Compare token-pressure and execution-pressure diagnostics across tasks.
  • Create small diagnostic probes that can run before expensive training.
  • Track when diagnostics change the next experiment, not just describe a result.

Required Controls

  • Ablation tied to a named hypothesis.
  • Negative examples and false positives included.
  • Diagnostic result connected to a decision.

Stop Conditions

Do not add diagnostics that cannot change an experiment decision or falsify an explanation.

Do not treat a coordinate that controls the next reported token as a reasoning variable until a separately computed consequence changes under a matched control.

Do not relabel the replicated 48/48 oracle context-local transport result as a capability gain. Target identity and clean donor coordinates are supplied; a native, non-oracle controller must earn deployment separately.

Experiments 91

  • 2026-07-27 → 28 Self-Written Verifier Fidelity

    Fill this after the run. Separate deployable evidence from oracle/hidden evaluation.

  • 2026-07-19 Qwen35 4B Agentic RLVR Feasibility

    Yes to the machinery: with three fixes (force eager mode so the hybrid architecture doesn't hang, load the model in bf16 not fp32, and split GPU memory 55% to the inference engine with sleep-mode) the loop CLOSES - a mov

  • 2026-07-18 Qwen35 4B Self-Repair Install

    The first bet that moved the needle at all — gently. We taught the 4B to debug by training on 504 examples of [buggy code + the real test-failure message] -> [diagnosis + fix], all self-generated by injecting bugs into c

  • 2026-07-16 Repair-Verifier Signal Probe

    The gate said no, cleanly. Handed two candidate repairs and the full failure evidence — a task solvable by mentally running each candidate through both trials — the model picked the working one 51.5 percent of the time w

  • 2026-07-15 Retention-Screen Calibration Study

    The gap wobbles with a standard deviation of 4.3 tasks — the five-task pass/fail margin was barely one wobble wide, so single-quiz forgetting verdicts were close to coin flips. Every historical 'this model forgot 5-10 ta

  • 2026-07-15 Rank-Capacity Vehicle Cell

    The trial could not answer the capacity question — by design. Its built-in guard required the known nine-point forgetting case to reproduce before trusting any comparison, and on this fresh screen that case measured only

  • 2026-07-15 Interleaved-Replay Dose with Medium Pilot

    The review round did not prevent forgetting: the dosed model lost nine to ten retained answers against both comparisons — almost exactly the cost of dosing directly — so the theory drawn from comparing old receipts is re

  • 2026-07-15 Dose-Diversity Mechanism Cell

    Variety does not protect memory: the larger varied practice set also cost nine retained answers, the known ten-point case reproduced exactly, and even pure review cost five on this fresh screen — so the one past case wit

  • 2026-07-15 Axis Stack Re-adjudication with Medium Pilot

    Judged fairly on brand-new tasks — with the ceiling quirk handled — the stacked model won the overall skill test for the third straight time (22 vs 15 and 18 of 40) with the cleanest finishing behavior again. But it won

  • 2026-07-15 Axis Corpus V2 with Staged Repair

    Even with lessons rebuilt from the autopsy — walking through the search step by step instead of asserting the answer — the repair skills did not stick (a tie and a loss against the comparisons), so the preregistered kill

  • 2026-07-14 → 15 Qwen3.5-4B Counterfactual Plan Reflection Transfer

    Not known yet. The current checkpoint builds and checks fresh list, text, and register puzzles, plus a control where every plan is deliberately assigned to the wrong puzzle. No model has been loaded or trained.

  • 2026-07-14 Policy-Supported Successful-Sibling Universal Curriculum

    This version cannot run its second step. Across 624 first attempts it found 227 failures, but counting and routing had none and selection had only two; the plan required four failures in every skill. It therefore stopped

  • 2026-07-14 Natural-Language State-Table Universal Curriculum

    Not known yet. CPU construction produced 80 truth-checked lessons and two 320-row streams with exactly 286,814 tokens each. All 48 smoke tests pass, but no model has been trained or evaluated.

  • 2026-07-14 Failure-Selected Counterfactual Restart Curriculum

    The data-selection stage passed: from 624 fresh attempts, the frozen rule selected 52 clean restarts, exactly four for each of 13 skills. Forty correct accuracy or cap failures and twelve correct-but-too-long cases are n

  • 2026-07-13 Validation-policy counterexample curriculum

    The test could not answer the question. Before the newly trained model was allowed to run, the starting model and the equal-size old-lesson control each fixed all 48 practice cases. The required gains were therefore impo

  • 2026-07-13 Qwen3.5-4B: Installing Universal Features via Designed Synthetic Curricula

    Not with either tested training geometry. Continuing from the strong policy raised fresh synthetic accuracy from 50% to 69% and beat the untrained base overall, but it scored 31% aggregate versus 45% for the strong start

  • 2026-07-13 State-Formation Capacity Adjudication

    Not run yet. The reviewed test starts with three compact-update training runs. If any required state check misses, six matched controls become mandatory; the final three runs open only if the full-size version also misse

  • 2026-07-13 Semantic-policy headroom tournament

    No rule produced a stable training target after a failed test: every one of 72 cases reached a correct patch. But before seeing test output, the model got zero of 54 ambiguous first patches fully right and never opened t

  • 2026-07-13 Qwen3.5-4B Semantic-Anchor Coordinate Branching

    This run cannot establish that. The internal edit strongly changed the model's choice among candidate names, but none of 440 consequence outputs began with a valid answer token. Even in a restricted twelve-choice readout

  • 2026-07-13 Qwen3.5-4B Materialized Residual Sibling Search Fresh Replication

    The generation test cannot answer the question because every reasoning trace reached its length cap and the materialized arm parsed only 12 of 52 outputs, with zero successes. The separate one-token ranking test is a cle

  • 2026-07-13 Qwen3.5-4B Materialized Residual Sibling Search

    Not yet. The scientific design and every model-free construction check pass, but the model has not run. The frozen test contains 264 fresh functions, and all 38,596 planned prompt renderings fit their assigned context li

  • 2026-07-13 Qwen3.5-4B Jacobian Counterfactual Branching

    No. Across all three allowed strengths, the meaningful nudge made its assigned answer win only 4 of 48 times—exactly the one-in-twelve chance rate and identical to a generic nudge. The probabilities barely moved even at

  • 2026-07-13 Qwen3.5-4B Early Text Hypothesis Forking

    The experiment design and model-free checks now pass, but the model has not run. Review expanded the first draft from twelve operation names to twenty-four fully specified operations, added two fair late-hint comparisons

  • 2026-07-13 Qwen3.5-4B Counterfactual Order-Support Selector

    Not reliably. The meaningful-order rule got 43 of 113 puzzles right, clearly better than first attempt or majority vote. But it was only two puzzles ahead of simply choosing the most decisive attempt, and a deliberately

  • 2026-07-13 Qwen3.5-4B: The Compute-Optimal Confidence Policy (select + abstain + escalate)

    Partly. Picking the answer with the highest single-token P(True) confidence is the best grader-free selector at every compute level, and abstaining below a confidence threshold trades coverage for accuracy cleanly — both

  • 2026-07-12 Transaction-invariant recovery curriculum

    Partly. Focused training lifted practice repairs from 52% to 82% and preserved the test-and-revise loop. But on different transaction APIs it reached 72% versus 70% before training, far below the required ten-point gain.

  • 2026-07-12 Verifier-conditioned recovery banking curriculum

    Yes — replaying real failures taught recovery, lifting success from about half of cases to more than nine in ten. But the best version also added a little "explain your fix" text, and that supposedly tiny slice actually

  • 2026-07-12 Repository search-compress-bank coding curriculum

    No. It became flawless on the six codebase types it trained on, solving every task in exactly four clean steps. But on four never-seen codebase types its success collapsed from 68% to 35% — worse than just running the un

  • 2026-07-12 Public-verifier recovery branch tournament

    No. On four new problem types, both policies repaired 74%, and their combined best reached only 75%—exactly tied with two action-only attempts. The selector was correctly stopped before scoring.

  • 2026-07-12 Qwen3.5-4B Native-Thought Seam Budget Ladder

    No. Across 48 tries on simple list-transformation puzzles, and at every thinking budget up to 1,024 tokens, the model closed its reasoning and produced an answer exactly zero times. It always burned the whole budget stil

  • 2026-07-12 Qwen3.5-4B Native-Thought Jacobian Value Transport

    We could not even reach the test. The whole plan needs the model to finish reasoning and write an answer worth grading. But on all 48 attempts at simple two-step list puzzles, it rambled straight into its 160-token think

  • 2026-07-12 Qwen3.5-4B Jacobian Value Transport

    No. Editing one internal direction at a late layer flipped the concept the model said aloud on 18 of 24 tries, up from never, and far beating an ordinary readout-style edit that worked only 1 in 5 times. But when that co

  • 2026-07-12 Qwen3.5-4B Jacobian Transport Control Replication

    Yes. Editing a handful of internal numbers at the early moment the model names its concept made it answer as a completely different concept on all 48 fresh test items, and a separately computed digit that depends on that

  • 2026-07-12 Qwen3.5-4B Forced-Commit Jacobian Value Transport

    No. Inserting the model's own "done thinking" marker cut the reasoning off but almost never flipped it into answer mode. Across three thinking budgets, only 13% to 19% of forced stops produced anything readable, and just

  • 2026-07-12 Qwen3.5-4B Context-Local Jacobian Clamp

    Yes, but only when the edit lands on the earlier token that first stores the word. There the model looked up the swapped word's digit on all 48 fresh puzzles, up from zero without the edit, and a wrong-word swap produced

  • 2026-07-12 Qwen3.5-4B Commit-Slot Semantic Power Replication

    The reasoning result is real, but the proposed internal value meter failed. Ordered scratch work beat the same words shuffled in two independent stages (about 29% versus 14% on confirmation). Yet a task-held-out model of

  • 2026-07-12 Qwen3.5-4B Commit-Slot Jacobian Value Transport

    Barely. Forcing the format fixed one problem outright: an allowed word was the model's top choice 85% of the time, versus 4% when it answered freely. But real step-by-step thinking beat the very same thought words scramb

  • 2026-07-11 Interactive policy curriculum: oracle DAgger to execution-reward RL

    No — it backfired badly. Success on the practice tasks fell from about 60% to 35%, and on tasks never touched in training from about 69% to 35%. It kept editing fluently but stopped running its checks: the verify-and-com

  • 2026-07-10 Qwen3.5-4B Verified Macro Invention Long-Context Rerun

    No. The model was not short on thinking space, it was stuck repeating itself, and more space only bought longer loops. Handed the blueprint, all 16 tasks finished cleanly. Asked to invent the program from examples alone,

  • 2026-07-10 Qwen3.5-4B verified-macro exact CUDA-graph vLLM rerun

    No, it just loops harder. At both budgets, all 48 reasoning attempts ran out of thinking room and had to be force-stopped, none finishing on its own. Widening the budget from 49,152 to 61,440 tokens pushed the share stuc

  • 2026-07-10 Qwen3.5-4B verified-macro capacity-fit vLLM rerun

    No. Two hidden limits bit first. Configuring the server for 64 simultaneous long prompts demanded about 2.4 million tokens of fast memory from a pool holding roughly 1 million, forcing constant re-reading. Cutting to 19

  • 2026-07-09 → 10 Gauntlet round 1: breadth-first agentic expert iteration

    Mostly the second. On a blind benchmark the model scored about 14 percent, with six task types near zero, but it had usually reasoned correctly. It simply hit its thinking limit, restarted explaining instead of writing t

  • 2026-07-09 Qwen3.5-4B Verified Macro Invention

    No. Even handed the finished plan and asked only to re-express it — no problem to actually solve — the model got it exactly right just one time in four, short of the three-in-four bar. Every output looked flawless: valid

  • 2026-07-08 → 09 Does the installable hypothesize-and-verify skill move the structure wall?

    No. Fine-tuning the model on 1,476 worked guess-and-check traces doubled its success at the two-step depth it practiced on (lists jumped from 37% to 70%), yet did nothing one step deeper: three-step success stayed near 5

  • 2026-07-08 Qwen3.5-4B: Does Code Confidence Replicate on HumanEval?

    Ask it directly. Averaging the model's confidence across every code token barely helps, nudging correct picks from 77% (blind guessing) to just 79%. But handing the model its own code and reading its confidence in one ye

  • 2026-07-07 → 08 Qwen3.5-4B: Can Confidence Replace the Verifier in the Banking Flywheel?

    No. When training on answers checked by actually running the code lifted single-shot accuracy from 8% to 24%, confidence-filtered data — fifteen times purer than a random grab of the model's own outputs — landed right on

  • 2026-07-07 Qwen3.5-4B: Does the Confidence Toolkit Survive on Real Code?

    Yes, but not the obvious way. Averaging the model's certainty across every token of a program barely beats a plain majority vote among the tries. The real winner: make the model write the code, then answer one yes/no que

  • 2026-07-06 Qwen3.5-4B: Externalize the Latent Readout (probe-to-prompt)

    Partly. Writing the concrete first step — exact value and all — into the prompt lifted the single-best-guess solve rate on two-step tasks sixfold, from 3% to 19%, where editing the model's internal state did nothing. But

  • 2026-07-06 Qwen3.5-4B: Is the Parameter Latent? (probe the full first op)

    It splits. The model genuinely computes the KIND of operation inside itself — reading its internal activity names the kind far better than the examples alone do (41% versus 27%, against 6% for blind guessing). But the sp

  • 2026-07-03 Qwen3.5-4B Coverage Banking: does banking shift the proposal distribution?

    Yes, but it depends on difficulty. On easy one-step problems it just pulls answers it already knew into its top guess (60% to 80%), gaining no new ground. On harder two-step problems something new happens: on brand-new t

  • 2026-06-30 Qwen3.5-4B Verifier vs Visible Selector Showdown

    No. When you can run even a single example test on each candidate, that filter alone lifts the share of shipped programs that fully work from 77% to 85%. Adding a free, instant self-confidence rating reaches 87% — matchi

  • 2026-06-30 Qwen3.5-4B Thinking-Budget Scaling

    Yes, but with a sweet spot. Turning on the model's built-in reasoning lifted the accuracy of the answer you would actually ship from 76% to 91% on basic Python tasks — a genuine capability gain, not just a wider pool of

  • 2026-06-29 → 30 Qwen3.5-4B Thinking Separability Probe

    Yes, but not for the reason you would expect. From one snapshot of internal activity, whether the model's own code is correct is readable well above a coin-flip, about 64 to 76 percent of the time. Thinking first sharpen

  • 2026-06-29 → 30 Qwen3.5-4B Generator-Verifier Gap

    Only after it thinks. Judging on sight, the model rubber-stamps 91% of its tries as correct while just 77% truly pass — barely a check, mostly agreeing with itself. Given room to reason first, it becomes a real critic: l

  • 2026-06-28 Qwen Oracle-Distilled Acquisition Policy

    No. The trained picker does read real signal: it beats revealing nothing (50 to 57 percent of tasks fully solved) and crushes a version fed scrambled answers (27 percent). But it lost to simply grabbing a varied spread o

  • imported 2026-07-12 Qwen Register Trace Refiner

    Rarely. Even a flawless picker that always grabbed the correct edit reached only 37% on plainly worded problems and 7% on reworded ones, because the correct program usually isn't among the roughly 1,300 nearby edits at a

  • imported 2026-07-12 Qwen Progressive Repair Compiler

    Partly. The correct fix sits in the candidate list almost nine times in ten, yet the judge finds it only about half the time, lifting exactly-correct programs from 30% to 49% against the 88% a flawless chooser would reac

  • imported 2026-07-12 Qwen Learned Repair Verifier

    Partly. Without ever seeing the true answer, the judge lifted correct execution from 30% to 47%. But a checker allowed to peek at the answer key found a correct fix already sitting in the candidate pile 88% of the time,

  • imported 2026-07-12 Qwen Candidate-Trace Verifier

    Yes. A small checker that reads each candidate's worked-out steps — never the true answer — lifted correctly-running programs from 30% to 54%, and to 56% when it cross-checks two wordings of one task. The twist: a picker

  • imported 2026-07-12 Qwen 3.5 4B Verified Edit Closure

    Yes, on the hardest unseen tasks. Testing small edits of the model's own near-miss program and keeping whichever passes the example cases raised fully-correct answers from 39% to 52%, while re-sampling the model and re-r

  • imported 2026-07-12 Qwen 3.5 4B GraphIR Self Repair

    No. The plain one-line formula fully solved 29% of brand-new tasks; the step-by-step diagram managed only 22%, and the fix-it pass recovered part of that gap to 24% — still behind. That pass genuinely works: on randomly

  • 2026-06-27 Qwen Verified Skill Memory RAG

    No. Handing the model a matched, verified example made it slightly worse, not better: every value came out correct on 47.5% of jobs, versus 50% when it worked through each value alone with no example. The lookup itself w

  • 2026-06-27 Pairwise Table Judge

    No. Letting the model run a knockout comparison among its own candidate tables did not help: fully-correct tables slipped from 50% (just trusting its first attempt) to 47.5%. On the hard cases where a correct table exist

  • 2026-06-27 Full-Table Consistency Reranker

    No. The trained scorer got every row right on half the tasks, exactly what you get by just keeping the model's first answer, and barely better than scoring tables at random. A flawless table was reachable on five more ta

  • 2026-06-27 Qwen3.5-4B Foofah Program Ensemble Consensus

    No. The simplest rule won: run the first program that passes a single worked example. It solved 52% of tables versus 44% when the model just answered directly, rescuing 23 tables it had otherwise botched while breaking o

  • 2026-06-27 Qwen3.5-4B Code ABI Oracle Coverage Ladder

    Yes, mostly. Snapping together at most three verified functions produced a passing solution for 84% of 160 small programming tasks, up from just 13% with a bare-bones library. But the library is a foundation, not a solve

  • 2026-06-26 Qwen3.5-4B Verified Algorithm Retrieval Adaptation

    Yes, but modestly. Fetching the most similar already-solved problem and adapting its code cracked 3 of the 8 problems that repeated tries kept missing, lifting overall success from 67% to 79%. Matching by meaning clearly

  • 2026-06-26 Qwen3.5-4B Substrate Coverage Ladder

    Only when a building-block was hand-carved for each specific problem. A shared library that reshapes old solutions to fit new ones solved zero of the nine stuck tasks, no better than nothing. Custom-written parts solved

  • 2026-06-26 Qwen3.5-4B Retrieval Adapt Verify Scale

    It splits. Pulling a similar solved problem from a 364-entry library and rewriting it recovered 8 of 24 stuck tasks — a third, double the 4 that a random-library control found. But the only safe check, the visible sample

  • 2026-06-26 Qwen3.5-4B Reliability Exec OPSD Audit

    No. All three cheap fixes failed. Ranking programs by the model's own confidence picked working code less often than just taking the first one that passed the visible sample tests (6 of 24 versus 8), and let more bugs sl

  • 2026-06-26 Qwen3.5-4B OPSD Pressure Locality Audit

    No. At the exact spots where correct code diverges from code that passes surface tests but is secretly wrong, the reference hint adds essentially nothing — scoring no better than a scrambled, meaningless hint. The hint o

  • 2026-06-24 → 26 Qwen3.5-4B Oracle Probe Synthesis MDP

    It's the menu. Just curating which eight test inputs the model chose from raised success from about 43% to 49% — a bigger jump than any training gave. Supervised coaching added a bit more (48% to 51%); preference- and re

  • 2026-06-25 Qwen Readable Candidate Verifier

    Barely. Training a picker on the model's reading of each fix's pseudocode plus its claimed output raised the share of tasks fixed correctly from 44% to 51%, closing only about 15% of the distance to a perfect picker's 90

  • 2026-06-25 Qwen Candidate-Conditioned Trace Verifier

    No. Doing nothing already solved about 44% of tasks, and always choosing a correct candidate could reach 90%. Yet no trained picker captured any of that headroom; the best merely tied doing nothing. Left unchecked, the m

  • 2026-06-25 Qwen3.5-4B Verifier-Guided Self-Improvement Report

    No. After a four-billion-parameter model retrained on its own test-passing code, its success on unseen problems barely moved, going from about 65% to about 65% (a hair lower). Retraining made its several attempts look mo

  • 2026-06-25 Qwen3.5-4B Real Sample Verify Commit

    It's the writing. When a 4-billion-parameter model writes several code attempts per problem, choosing is not the bottleneck. On easier problems nearly every attempt is already correct, so the best result sits near 97%. O

  • ~2026-06-25 Qwen3.5-4B Oracle-Distilled Semantic Verifier

    Yes, but only on home turf. On the problem set it trained on, the trained judge picked a genuinely-correct program 81% of the time, up from 72% for the same model untrained and just 44% for grabbing the first candidate t

  • ~2026-06-25 Qwen3.5-4B HumanEval Adaptive Evidence Budget

    No. Every strategy landed at the same 16.7% correct-pick rate — running no tests, running all eight, and even a strategy allowed to peek at the right answer to time its stop. The trained model did learn to quit sooner, u

  • 2026-06-24 Qwen Structural Compiler Attribution Ablation

    It's the practice schedule. Give the model its full size from the start, then feed examples easy-to-hard — 8 steps, then 16, then 24 — and it solves nearly every standard 24-step program (about 97%). The popular guess, g

  • 2026-06-24 Qwen3.5-4B Sketch Coverage Shift Probe

    Only when the template explicitly names the new operation. On tasks needing a brand-new operation, hand-written templates that named and typed it kept the correct program every single time; auto-generated templates kept

  • 2026-06-24 Qwen3.5-4B Oracle Process GRPO

    Yes, but with sharp limits. Teaching the model to copy a flawless test-picker lifted its success from about a third of puzzles to 43%, versus roughly 31% for picking blindly. Layering on preference-based and reward-based

  • 2026-06-24 Qwen3.5-4B Adaptive Evidence Budget Policy

    Yes. Trained just for this single decision, the model matched the accuracy of running all ten checks — solving about 92 of every 100 tasks — while using only about five checks on average, near a perfect-hindsight stopper

  • 2026-06-23 Qwen Typed Bytecode Expert Iteration

    It depends. Training the model only on its own attempts that landed on the correct final answer lifted unaided first-try accuracy from 62% to 73% on fresh problems — a real gain that sticks when it writes programs alone.

  • 2026-06-23 Qwen Prefix-State Process Verifier

    Barely. The judge got genuinely good at telling promising partial programs from dead ends, yet it lifted first-try accuracy only a few points, with hard problems going from 41% to 44%. The real gap: on hard problems a co

  • 2026-06-23 Qwen On-Policy Repair-to-Compiler

    Yes, but the self-correction isn't the secret ingredient. On brand-new requests the model climbed from about 29% correct programs to 99% after training on its own corrected drafts. The catch: training on plain correct pr

  • 2026-06-23 Qwen Mixed-Domain Trace Verifier

    Yes, partly. On fresh tasks the frozen model alone got 46% right; the proofreader lifted that to 57%, and a cross-check that compares reworded versions of the same task reached 61%. But a correct recipe was already among

  • 2026-06-23 Qwen Hidden VM Curriculum Repair

    No. Feeding it nearby worksheets that merely land on the correct answer wrecked it. Plain step-by-step training scored 72% on the main mixed test; the same model after answer-chasing repair fell to 35%, and collapsed on

  • 2026-06-23 Qwen Context-Conditioned Trace Verifier

    Mostly no. A correct program was almost always sitting in the candidate pile, for 91 to 100 percent of questions, yet the judge barely helped: it nudged easy questions from about 69 to 70 percent and actually made the ha

  • 2026-06-23 Qwen Complete-Program Trace Reranker

    Barely, and it backfires on hard cases. A correct program sat in the candidate pile 91 to 100 percent of the time, yet the trained picker nudged easy prompts only from 69 to 70 percent and actively hurt longer ones, drop

  • 2026-06-22 Qwen Verifier-Guided Slot Repair

    Yes, mostly, but with a catch. A small model copying 24-step calculations got only about 27% exactly right on its own. A checker that knows the correct running number after every step, allowed to swap one or two bad step

  • 2026-06-22 Qwen Teacher-Distilled Slot Compiler

    No. A model trained to copy numbers and operations out of text and run a 24-step calculation got 27% of final answers exactly right; adding the pointing signal landed at 28%, a tie. Worse, agreement between two rewording

Claims

Queued proposals 4