Benchmark Generalization
Stress whether mechanisms transfer across substrates, families, lengths, distributions, and real tasks.
What we have learned
Seed Experiments
- factor_recombination_ladder
- feature_factorized_rule_diversity
- targeted_bridge_allocation
- qwen35_4b_sketch_coverage_shift_probe
Current Read
The imported tracks include strong shift probes. Future work should make shift evaluation a cross-program norm rather than a special case.
qwen35_4b_specialist_policy_integration: the registered primitive-transfer and compound-composition evaluation was not opened. A saturated mandatory tools core made the upstream
+0.10specialist bar impossible, so the run consumed no confirmatory or benchmark seeds. This preserves the evaluation instrument and demonstrates that every training core needs a feasibility check before a generalization claim is scheduled.qwen35_4b_pareto_policy_integration: C54's menagerie quick/medium labels were tested as proposed teacher roles on a fresh procedural quick/deep proxy without opening a benchmark seed. The quick
blendrole failed both paired blocks and the deepapexrole missed retention. This does not refute C54 on its own instrument; it shows that instrument ranking is not automatically a transportable local teacher rule. Future integration claims need separate gates for proxy-state teacher advantage and held-out instrument generalization.qwen35_4b_language_reasoning_wall (claim C37): the compositional wall does NOT exist in language. The model chains depth-3+ multi-step SIMULATION in natural language near-perfectly (no-think), unlike the depth-3 formal-composition wall -- the wall is formal-modality-specific, not a general multi-step limit. Made-up-relation control confirms it is MODALITY not a semantic prior. Formal-dict triggers code-mode. Tests SIMULATION (C13), not the C32/C36 proposal wall.
qwen35_4b_language_proposal_wall (claim C38): the structure-PROPOSAL wall persists in language. Depth-1 dissociation -- the model EXECUTES a given rule (0.86) but cannot INDUCE one from examples (0.00 no-think; 0.50 think). The model is an executor, not an inducer, in language as in formal domains. Complement to C37 (simulation intact): the wall's two components dissociate by modality -- execution formal-specific, induction modality-general.
qwen35_4b_icl_retrieval_vs_induction (claim C39): in-context learning is RETRIEVAL of familiar structure, not INDUCTION of novel structure. The model EXECUTES a novel rule perfectly (0.97) but cannot INDUCE it from examples (0.12=chance), while it induces a familiar rule far better (0.45); more examples don't rescue novel induction. Unifies the arc: executor/retriever of pretrained structure, not inducer of novel structure.
qwen35_4b_metacognitive_boundary (claim C40): the model knows when it will fail IMPLICITLY (answer-token probability predicts per-item correctness at AUROC 0.95 within a surface-matched cell, >> surface baseline 0.61) but NOT EXPLICITLY (self-verification P(True) 0.46 = chance; verbalized confidence a constant 100). Deployable: read logits for a confidence/abstain signal, never the self-report.
qwen35_4b_confidence_guided_compute (claim C41): beat sample-more with the model's own uncertainty. Confidence-select (argmax P(answer), verification-free) 0.62 beats flat self-consistency 0.48 at every budget; max P(answer) predicts solvability (AUROC 0.83) for abstention. Turns C40's calibrated confidence into a deployable compute tool; selection+abstention (not allocation) is the win.
qwen35_4b_code_confidence (claim C46, MBPP leg): C40/C41 transfer to real code, but the signal changes. Sequence mean-logprob dilutes over long programs; the deployable readout is single-token P(True). MBPP: P(True)-select 0.762 vs public-output majority 0.721 and random 0.696. Visible execution still wins when tests exist, so confidence is the verifier-free selection/abstention lever.
qwen35_4b_humaneval_code_confidence (claim C46, HumanEval replication): the same P(True) selector wins on all 164 HumanEval tasks with no public probes: P(True) 0.835 vs mean-logprob 0.787 and random 0.766, oracle 0.872. Greedy solvability AUROC is 0.862 for P(True), supporting the cross-benchmark code-confidence law.
qwen35_4b_error_localization (claim C42): the model can localize its own errors -- per-step confidence dips exactly at the first slip (surviving de-trending; single-slip localization 0.56 >> position-prior 0.36). C40's metacognition is step-resolved; deployable targeted repair (redo from the located step, cheaper than redo-all).
qwen35_4b_native_thought_seam_budget_ladder (unclaimed setup negative): the MBPP-derived 512--1024 native-thinking scale did not transport into autonomous termination on fresh list induction. All 48/48 traces reached 1,024 without close, versus prior benchmark-dependent closure. This is not an accuracy comparison; it demonstrates that termination calibration itself is a workload-specific interface property. Future thought-state mechanisms must calibrate on their actual task/prompt/backend.
- qwen35_4b_materialized_residual_answer_seam_factorial (unclaimed interface negative): termination identity also matters within one pinned backend. All 240 calibration outputs authenticated and every arm scored 0/48 strict parse. Tokenizer EOS plus newline before registered HF EOS completely explains the two no-think cells: suffix-only removal yields 48/48 frozen-parser exactness in both. It does not completely explain thinking: suffix-only scores were 38/48 and 24/48, ten/five rows had extra closes, and 18 think/freeform rows capped without tokenizer EOS. None of this changes the terminal gate. Every workload/interface must register and test its actual token-native commit boundary; “EOS” is not a fungible label.
Scorecard
- Program: charter
- Current read: many mechanisms look good in-family; the repository needs standard transfer stress before strategic claims harden. C46 shows the confidence toolkit survives MBPP->HumanEval only after the signal is re-expressed as a single-token P(True) readout. The Pareto qualification negative adds the reverse warning: a held-out instrument ranking can be real on that instrument yet fail to define a useful teacher ordering on the clean training proxy. Native-thinking termination is also non-portable: a 512--1024 scale that often closed on MBPP produced 0/48 natural closes at 1,024 on fresh list induction.
- Best next experiment: compositional-grammar induction as the C45 stress test, plus a small cross-program generalization suite used by compiler, selector, memory, and adaptation work.
- Strong anchors:
factor_recombination_ladder,feature_factorized_rule_diversity,targeted_bridge_allocation. - Avoid repeating: reporting only IID or narrow held-out splits for a mechanism meant to generalize.
- Evidence that advances the program: transfer across substrate, family, length, and real-task variants with a clear failure taxonomy.
Charter
Show charter.md
Purpose
Stress whether mechanisms survive new substrates, longer horizons, held-out families, bridge compositions, real tasks, and distribution shifts.
Why This Is A Program
Without a dedicated generalization line, every successful local experiment risks becoming a narrow benchmark trick. This program keeps the corpus honest.
Progress Signals
- Methods are evaluated on held-out primitives, families, lengths, or domains.
- Gains replicate across at least two substrates.
- Failure under shift is documented and converted into new hypotheses.
- Benchmark construction itself is versioned and reusable.
Boundaries
This program evaluates transfer. It does not own any single mechanism.
Backlog
Show backlog.md
Next Experiments
- Stopped cross-program test:
qwen35_4b_specialist_policy_integrationkept its no-new-exposure compound and confirmatory distributions sealed, but stopped before training because one mandatory specialist target exceeded the score ceiling. A new experiment may reuse the generalization design only with a harder independently calibrated tools/provenance core and fresh frozen confirmatory seeds. - Completed proxy-transport negative:
qwen35_4b_pareto_policy_integrationkept every benchmark seed sealed and tested C54's external quick/medium labels on a fresh procedural quick/deep proxy. The required crossover did not reproduce:blendlost both quick blocks. Future distillation work must distinguish instrument-specific ranking from teacher advantage on the training-state distribution; neither can stand in for the other. - Completed termination-transport negative:
qwen35_4b_native_thought_seam_budget_ladderfound 0/48 natural closes at 1,024 on fresh list induction, so an MBPP-era budget scale is not a portable answer-seam guarantee. Calibrate termination on every new workload; the next forced-commit experiment must keep selection and confirmation tasks fresh and treat the commit action as part of the interface under test. - Completed termination-identity negative:
qwen35_4b_materialized_residual_answer_seam_factorialauthenticated exact no-think outputs and expected thinking answer tails followed by tokenizer EOS/newline, but its HF-EOS interface correctly failed strict parsing. Test tokenizer EOS versus HF EOS explicitly on fresh calibration rows; never inherit a termination token by name alone or repair a result after observing which boundary the model used. - Build a common shift taxonomy: length, family, primitive, composition, prompt, format, and real-task shift.
- Re-run top mechanisms on at least one non-original substrate.
- Add bridge-composition and held-out-primitive splits to new experiments by default.
- Create tiny smoke suites for fast sanity and larger challenge suites for claims.
- Track which claims are single-substrate versus cross-substrate.
- Run the C45 follow-up as compositional-grammar induction, not a flat non-affine menu: train reasoning-SFT on condition x action depth-1 rules, then test held-out combinations, held-out productions, and held-out composition-depth (depth-2 nested/two-action chains) as separate endpoints. Gate every family with execute-given-rule ceilings, a token-budget/truncation curve, and automated example-set sufficiency checks so failures distinguish search/composition limits from non-executable or underdetermined episodes.
Required Controls
- IID split.
- At least one held-out split.
- Baseline repeated on every split.
- Report both row-level and task-level metrics when applicable.
Stop Conditions
Do not promote a result to shared strategy if it has only been shown on a single easy or IID split.
Experiments 131
- 2026-07-27 → 28 Self-Written Verifier Fidelity
Fill this after the run. Separate deployable evidence from oracle/hidden evaluation.
- 2026-07-18 → Qwen35 4B WHY-Think Scale
This phase built and proved the machinery; the GPU training sweep has not run yet. The generator now emits, for every example, a real hidden reasoning trace generated mechanically from the program's shape - parse the spe
- 2026-07-18 → Qwen35 4B WHY Scale Ladder
This phase built and proved the machinery; the GPU training sweep has not run yet. The core blocker was that the original WHY generator saturated fast (about 75 distinct reasons, 438 distinct programs at 504 examples), s
- 2026-07-17 Coding Fitness Harness (cognitive-core program)
Surprisingly good at writing single functions (HumanEval 76.2%, MBPP 56.5%) but weak at driving a multi-step coding task in a real agent loop (23%). The harness is validated: it agrees with an independent run to within 1
- 2026-07-17 → State-Track Installation (Stage 9)
The believed-unlikelier but hoped-for outcome landed. After the reliable 'just replay again' lever hit its ceiling, this tried a genuinely different lever: teach the model one new, universal skill — keeping a running tal
- 2026-07-15 Axis Stack Re-adjudication with Medium Pilot
Judged fairly on brand-new tasks — with the ceiling quirk handled — the stacked model won the overall skill test for the third straight time (22 vs 15 and 18 of 40) with the cleanest finishing behavior again. But it won
- 2026-07-15 Axis Corpus V2 with Staged Repair
Even with lessons rebuilt from the autopsy — walking through the search step by step instead of asserting the answer — the repair skills did not stick (a tie and a loss against the comparisons), so the preregistered kill
- 2026-07-14 → 15 Fresh-Surface Budget-Commit Universal Curriculum
The rewritten practice set beat both comparison models on the big screen — more right answers (69 vs 63 and 62 of 104), many fewer run-on answers (7 vs 18 and 13), and 31 percent shorter output — on vocabulary it had nev
- 2026-07-14 → 15 Goal-Gap Axis Curriculum
The targeted practice worked on its own terms — the first screen pass in this program's history (28 vs 22 and 18 of 40 on unseen tasks, with zero forgetting). On the held-out benchmark it beat the untrained model by a wi
- 2026-07-14 Policy-Supported Successful-Sibling Universal Curriculum
This version cannot run its second step. Across 624 first attempts it found 227 failures, but counting and routing had none and selection had only two; the plan required four failures in every skill. It therefore stopped
- 2026-07-14 Natural-Language State-Table Universal Curriculum
Not known yet. CPU construction produced 80 truth-checked lessons and two 320-row streams with exactly 286,814 tokens each. All 48 smoke tests pass, but no model has been trained or evaluated.
- 2026-07-14 Search-Scaffold Universal Curriculum
No. On 26 fresh paired cases, the staged-search model tied replay at 16 correct and trailed its parent at 18. All three parsed 23 answers and hit the limit three times; the candidate solved none of four execute/induct ca
- 2026-07-14 Residual-Skill Successful-Sibling Universal Curriculum
Grading the 3,600 retries found 855 short fully-correct ones, and nine of ten weak skills had plenty. But rule-guessing (induction) yielded a usable correct retry on only 2 of its 46 failed tasks, below the required four
- 2026-07-14 On-Policy Failure-Prefix Universal Curriculum
No. The unmodified parent solved 16 of 26 fresh tasks and equal-compute replay solved 18, while training corrections after the model's own real failures solved only 15. The repair model got none of the six targeted execu
- 2026-07-14 Failure-Selected Counterfactual Restart Curriculum
The data-selection stage passed: from 624 fresh attempts, the frozen rule selected 52 clean restarts, exactly four for each of 13 skills. Forty correct accuracy or cap failures and twelve correct-but-too-long cases are n
- 2026-07-14 State-Formation Branch Authorization Recovery
Yes for the no-model safety check: all six controls passed, the failed first attempt has a third preserved copy, and the two retry-blocking source paths were retired only after that archive commit passed both checks. The
- 2026-07-14 State-Formation Branch Handoff Recovery
Yes, the narrow handoff passed every safety check, then all three full-size setup controls reached 48 of 48 with their learned path and zero of 48 with that path disabled.
- 2026-07-13 → 14 Close-Weighted Universal Commit Seam
No. On 26 fresh procedural cases, ordinary and close-weighted target training both produced 23 well-formed answers and three response-limit contacts. Close weighting scored 16 correct versus 15 for ordinary training, but
- 2026-07-13 → 14 State-Formation Analysis Recovery
Yes, the narrow recovery check passed. It accepts only the one registered file location, still rejects unsafe shortcuts, and has not yet examined any model result.
- 2026-07-13 Validation-policy counterexample curriculum
The test could not answer the question. Before the newly trained model was allowed to run, the starting model and the equal-size old-lesson control each fixed all 48 practice cases. The required gains were therefore impo
- 2026-07-13 Replay-Anchored Universal Curriculum Continuation
No at this dose. The designed mix passed its fresh synthetic screen but scored about 42% overall, below both the 44% mature policy and the 49% replay-only comparison. It also fell below base on one family. Replay-only wa
- 2026-07-13 Mid-Density Token-Matched Universal Curriculum
The 160-lesson stream improved fresh local accuracy from 17 to 19 of 26 cases, raised readable answers from 18 to 23, and cut answer-limit contacts from nine to three. It still missed the readability and length gates by
- 2026-07-13 Low-Density Token-Matched Universal Curriculum
No. Neither 40 nor 80 abstract lessons passed the fresh local screen. The stronger 80-lesson stream matched the inherited anchor at 54% accuracy and parsed 62% of answers, but still hit the answer cap 10 times; broad eva
- 2026-07-13 Qwen3.5-4B: Installing Universal Features via Designed Synthetic Curricula
Not with either tested training geometry. Continuing from the strong policy raised fresh synthetic accuracy from 50% to 69% and beat the untrained base overall, but it scored 31% aggregate versus 45% for the strong start
- 2026-07-13 Semantic-policy headroom tournament
No rule produced a stable training target after a failed test: every one of 72 cases reached a correct patch. But before seeing test output, the model got zero of 54 ambiguous first patches fully right and never opened t
- 2026-07-13 Qwen3.5-4B Semantic-Anchor Coordinate Branching
This run cannot establish that. The internal edit strongly changed the model's choice among candidate names, but none of 440 consequence outputs began with a valid answer token. Even in a restricted twelve-choice readout
- 2026-07-13 Qwen3.5-4B Materialized Residual Answer-Seam Factorial
No registered style qualified: all four scored zero strict parses out of 48, so mechanics stayed sealed. Removing only the final chat-end marker and newline made both no-think styles exact on all 48 rows; thinking still
- 2026-07-13 Counterfactual evidence-acquisition curriculum
The run stopped at its first checkpoint-compatibility test. Across 48 frozen contexts, the starting checkpoint differed from the comparison anchor by 0.110735 centered logits, above the fixed 0.100000 limit. Entropy stay
- 2026-07-12 Transaction-invariant recovery curriculum
Partly. Focused training lifted practice repairs from 52% to 82% and preserved the test-and-revise loop. But on different transaction APIs it reached 72% versus 70% before training, far below the required ten-point gain.
- 2026-07-12 Verifier-conditioned recovery banking curriculum
Yes — replaying real failures taught recovery, lifting success from about half of cases to more than nine in ten. But the best version also added a little "explain your fix" text, and that supposedly tiny slice actually
- 2026-07-12 Qwen3.5-4B Same-Prefix Advantage Routing
Only when it replicates, and here one helper didn't. The test actually ran: the deep specialist's local edge held up on a fresh, separate batch, but the quick specialist looked strong on one batch of stuck points and the
- 2026-07-12 Repository search-compress-bank coding curriculum
No. It became flawless on the six codebase types it trained on, solving every task in exactly four clean steps. But on four never-seen codebase types its success collapsed from 68% to 35% — worse than just running the un
- 2026-07-12 Qwen3.5-4B Native-Thought Seam Budget Ladder
No. Across 48 tries on simple list-transformation puzzles, and at every thinking budget up to 1,024 tokens, the model closed its reasoning and produced an answer exactly zero times. It always burned the whole budget stil
- 2026-07-12 Qwen3.5-4B Deep-Advantage MOPD
No. The deep specialist really was the better teacher on selected states, and copying it worked slightly better than copying the wrong teacher or training on matched non-winning states. But after four rounds the resultin
- 2026-07-11 Qwen3.5-4B Specialist Policy Integration
No — we never found out. Each of four required skills had to gain ten points before merging, but scores top out at 100. One tool-use skill already sat at 99.4%, so its target of 109.4% was impossible. The experiment stop
- 2026-07-11 Interactive policy curriculum: oracle DAgger to execution-reward RL
No — it backfired badly. Success on the practice tasks fell from about 60% to 35%, and on tasks never touched in training from about 69% to 35%. It kept editing fluently but stopped running its checks: the verify-and-com
- 2026-07-10 Qwen3.5-4B Verified Macro Invention Long-Context Rerun
No. The model was not short on thinking space, it was stuck repeating itself, and more space only bought longer loops. Handed the blueprint, all 16 tasks finished cleanly. Asked to invent the program from examples alone,
- 2026-07-10 Qwen3.5-4B Partial-Structure Recognition-Guided Search
No. Shown a half-finished program skeleton, the four-billion-parameter model's guess at whether it could still be completed was barely above a coin flip — about 51% correct, where 50% is pure chance — and letting it reas
- 2026-07-09 Qwen3.5-4B Verified Macro Invention
No. Even handed the finished plan and asked only to re-express it — no problem to actually solve — the model got it exactly right just one time in four, short of the three-in-four bar. Every output looked flawless: valid
- 2026-07-08 Qwen3.5-4B: Does Code Confidence Replicate on HumanEval?
Ask it directly. Averaging the model's confidence across every code token barely helps, nudging correct picks from 77% (blind guessing) to just 79%. But handing the model its own code and reading its confidence in one ye
- 2026-07-07 Qwen3.5-4B: Does the Model Know When It Will Fail?
Yes, but only in its numbers, never its words. The probability it quietly places on the digit it writes sorts right answers from wrong ones almost perfectly — 95 out of 100, versus 50 for a coin flip — and beats guessing
- 2026-07-07 Qwen3.5-4B: Does the Compositional Wall Exist in Language?
No. Written as ordinary sentences, the model follows a chain of invented names near-perfectly through four hops (94 to 100 percent correct), far above the roughly 4 percent a blind guess earns. The three-step ceiling is
- 2026-07-07 Qwen3.5-4B: Does the Structure-Proposal Wall Exist in Language?
No. Handed the rule outright, the model applies it correctly 86% of the time. Asked to infer that same one-step rule from worked examples, it scores 0% — below even the roughly 6% that pure guessing would earn. Letting i
- 2026-07-07 Qwen3.5-4B: Is In-Context Learning Retrieval or Induction?
Only patterns it already has. Told a scrambled counting order outright, the model applied it almost perfectly — 97 percent right. But shown examples of that same order and asked to work it out, it scored 12 percent, no b
- 2026-07-07 Qwen3.5-4B: Can the Model Localize Its Own Errors in Multi-Step Reasoning?
Yes, and the dip lands on the exact step, not just late in the chain. Confidence naturally climbs the deeper the model goes, so "least sure" could just mean "last step." Correcting for that, the least-confident step is t
- 2026-07-07 Qwen3.5-4B: Beating Sample-More with the Model's Own Uncertainty
No. Picking the most common answer wastes the extra tries: accuracy stays flat near 48% however many you draw, because the model keeps confidently repeating the same wrong rule. Picking the answer it was surest of instea
- 2026-07-07 Qwen3.5-4B: Does the Confidence Toolkit Survive on Real Code?
Yes, but not the obvious way. Averaging the model's certainty across every token of a program barely beats a plain majority vote among the tries. The real winner: make the model write the code, then answer one yes/no que
- 2026-07-06 Qwen3.5-4B: Learn from Your Own Failures (DPO)
No. The model already ranked its own correct answer above its wrong one 81% of the time — a sharp internal judge. But training it to favor the correct ones destroyed its writing: single-best-guess solving peaked near 5%
- 2026-07-05 Qwen3.5-4B: Bank the Thoughts
No. Training the model on its own successful reasoning taught it nothing beyond showing it the bare answers — both solved about 9 in 100 fresh three-step problems. But a short hand-written plan that builds the solution f
- 2026-07-02 Qwen3.5-4B Cross-Family Laws
No. Handed the exact steps, this fixed 4-billion-parameter model wrote correct code almost every time across three unrelated task types. But asked to infer the same procedure from example inputs and outputs alone, it fel
- 2026-06-30 Qwen3.5-4B Overthinking Content Ladder
No, the opposite. Giving this small coding model more room to reason made genuine, ordered reasoning matter more, not less. Scrambling its own reasoning steps into random order cost about 10 points of solved tasks at the
- 2026-06-28 Qwen Support-Contrastive Meta-ICL
Not by default. One tuned model scored 49% whether its worked examples were intact or scrambled, meaning it had memorized the task and ignored the examples entirely. A contrastive objective forced genuine reading: 50% wi
- 2026-06-28 Qwen Oracle-Distilled Acquisition Policy
No. The trained picker does read real signal: it beats revealing nothing (50 to 57 percent of tasks fully solved) and crushes a version fed scrambled answers (27 percent). But it lost to simply grabbing a varied spread o
- 2026-06-28 Qwen Learned Active Interrogation
No. Unlocking four extra answer keys did raise fully-solved tasks from 63% to 70%, and scrambling those answers sank it to 47%, so real labels genuinely matter. But the model's deliberate picks scored exactly the same 70
- 2026-06-28 Counterfactual ICL Public Multiseed Gate
Yes, but not for the reason expected. Tuning tripled whole-task accuracy on real text tasks, from 20% to about 61%, stable across three training runs — and the model genuinely reads its examples: remove them and it colla
- 2026-06-28 Counterfactual Episodic ICL Posttraining
Yes. Untrained, a 4-billion-parameter model solved 23% of real text-transformation tasks perfectly; after this training, 57% — but only when it could see the prompt's examples. Scramble those examples and it fell to 17%,
- 2026-06-28 Counterexample-Guided Ephemeral Program
No. Answering each row directly solved three-quarters of tasks completely, while the best rule-program the method could pick solved only about four in ten — and a perfect picker that peeks at the answers did no better. T
- 2026-06-28 Qwen3.5-4B Foofah Strategy Discovery Live
Barely, and nothing genuinely new. The self-written recipes lift exactly-solved tables from 42% to 46%, edging out plain prompting. But a ready-made library of solved programs already handled 56 to 58%, and every table t
- 2026-06-28 Qwen3.5-4B Foofah Program Strategy Portfolio
Yes, but modestly, and which passing program you trust matters more than the programs. Asking directly for the finished table got 42% right. The rule the team locked in, commit only when two programs agree, reached just
- 2026-06-28 Qwen3.5-4B Foofah Adaptive Program Budget Router
Yes, and here is the twist: running all five programs on every task scored lower (56%) than the cheap rule (58%), because blanket spending overwrote one answer the quick pass already had right. The rule fires only when t
- 2026-06-28 Qwen3.5-4B Adaptive Tool Controller
Partly. One structural cue — does the direct answer have fewer columns than the raw data implies? — safely flags the reshaping tasks where a program helps, lifting accuracy from 42% to 50% with zero broken tasks. But it
- imported 2026-07-12 Qwen Slot Repair Distillation
No. A correct program almost always sits one or two edits away — a search that peeks at the answer lifts solve rates from about a quarter to roughly 86%. But the blind helper couldn't pick which edits to make: on reworde
- imported 2026-07-12 Qwen Register-Token Latent Compiler
Only for short chains. Up to twelve steps it builds the correct hidden program about nine times in ten, while stripped-down versions trained on the final answer alone never find the interface and stay at chance. But at t
- imported 2026-07-12 Qwen Register-Token Structured Runtime
Only up to a point. For chains of four to twelve steps the hidden program runs flawlessly, at 100 percent. But at 24 steps exact execution collapses to 25 percent — versus about 1 percent from pure guessing. The catch: e
- imported 2026-07-12 Qwen Progressive Repair Compiler
Partly. The correct fix sits in the candidate list almost nine times in ten, yet the judge finds it only about half the time, lifting exactly-correct programs from 30% to 49% against the 88% a flawless chooser would reac
- imported 2026-07-12 Qwen Candidate-Trace Verifier
Yes. A small checker that reads each candidate's worked-out steps — never the true answer — lifted correctly-running programs from 30% to 54%, and to 56% when it cross-checks two wordings of one task. The twist: a picker
- imported 2026-07-12 Qwen 3.5 4B Verified Edit Closure
Yes, on the hardest unseen tasks. Testing small edits of the model's own near-miss program and keeping whichever passes the example cases raised fully-correct answers from 39% to 52%, while re-sampling the model and re-r
- imported 2026-07-12 Qwen 3.5 4B Typed Sketch Synthesis
It depends on difficulty. On the hardest problems, sketching the shape and letting a verified search fill the blanks lifted correct fixes from 33% to 78%, and a safe blend of both methods reached 88%. But on easy problem
- imported 2026-07-12 Qwen 3.5 4B Static Bridge Ceiling Breaker
Partly. Folding in just 60 slightly-harder "bridge" examples, a quarter of the training budget, more than doubled success on deeper, never-seen programs, from 20% to 44% fully repaired, with no loss on familiar skills. B
- imported 2026-07-12 Qwen 3.5 4B GraphIR Self Repair
No. The plain one-line formula fully solved 29% of brand-new tasks; the step-by-step diagram managed only 22%, and the fix-it pass recovered part of that gap to 24% — still behind. That pass genuinely works: on randomly
- imported 2026-07-12 Dense Supervision Ladder Experiment
It's the feedback. With the model held fixed, training it on only one sampled final answer left it weak; showing it the full odds of every possible answer at every step roughly doubled how often it solved the hardest 24-
- 2026-06-27 Real Transform ABI Gate with Counterexamples
It depends on how messy the data is. For clean, spreadsheet-style pipeline jobs the toolbox covered every one (100%) and held firm even against deliberately tricky examples. For irregular date, ID, and text cleanup, cove
- 2026-06-27 Qwen Verified Skill Memory RAG
No. Handing the model a matched, verified example made it slightly worse, not better: every value came out correct on 47.5% of jobs, versus 50% when it worked through each value alone with no example. The lookup itself w
- 2026-06-27 Qwen Recursive Task Decomposition
No. Forcing a 4-billion-parameter model to commit to one written recipe lowered accuracy: it got 59% of individual cases right versus 69% when answering each fresh, and 43% of whole tasks perfectly versus 47%. The recipe
- 2026-06-27 Qwen Recursive Ephemeral Program Induction
Only when you check the rule first. On its own, a model writing and applying a reusable rule solved 40% of tasks perfectly versus 56% for plain row-by-row answering, and it broke six tasks direct answering had solved. Ad
- 2026-06-27 Qwen Public PROSE ABI Gate
No. The frozen toolkit fully solved only 19% of the 309 outside tasks. In 77% of them no recipe fit even the worked examples, so the toolkit lacked that operation entirely; under 4% overfit. Yet a small four-billion-para
- 2026-06-27 Pairwise Table Judge
No. Letting the model run a knockout comparison among its own candidate tables did not help: fully-correct tables slipped from 50% (just trusting its first attempt) to 47.5%. On the hard cases where a correct table exist
- 2026-06-27 Full-Table Consistency Reranker
No. The trained scorer got every row right on half the tasks, exactly what you get by just keeping the model's first answer, and barely better than scoring tables at random. A flawless table was reachable on five more ta
- 2026-06-27 Qwen Episodic Soft-Prompt Task Vectors
No. Tuning the primer tokens on each task's own examples solved exactly the same share of tasks perfectly as plain instructions: just over half either way. The tuning genuinely worked, beating a version trained on delibe
- 2026-06-27 Qwen Disagreement-Probe Program Induction
No. The disagreement quiz picked the same programs whether its judge answers were real, randomly assigned, or skipped entirely — all three landed at 64% of tasks fully solved. The only genuine gain came from a plain caut
- 2026-06-27 Counterexample-Guided Consistency Projection
No. Distilling the model's messy guesses into one reliable formula solved only 20% of tasks completely, versus 50% when it simply transformed each row directly, a 30-point drop. For 29 of the 40 tasks no formula even fit
- 2026-06-27 Qwen Batched Transduction Consistency
No. Answering every entry in one combined pass got the whole task right 45 percent of the time, versus 50 percent when each entry was handled alone, a five-point loss. Bundling never rescued a task the solo approach had
- 2026-06-27 Qwen Active Example Acquisition
Barely, and the real lesson is the downside. Letting the model request the single example it was most unsure about lifted fully-solved tasks from 20 to 21 out of 30. But three random extras dropped it to 19, and feeding
- 2026-06-27 Qwen Active Crystallizer Public Gate
No. Using the model's votes to choose a rule worked on 25% of tasks — barely above the 22.5% you get from scrambled, meaningless votes, and it never beat the best rule the candidate pool could offer. The model answered i
- 2026-06-27 Qwen3.5-4B Transform ABI Compiler Pilot
Yes. After a light round of tuning, the model chose a recipe that produced the correct output on all 48 test tasks, matching a perfect answer key and beating the untuned model's 92%. It recovered the harder multi-step ch
- 2026-06-27 Qwen3.5-4B Foofah Selective Program Fallback
Trust the program the moment it reproduces the visible worked examples. Doing that lifted exact-match accuracy from 55% to 62% across 250 table tasks, rescuing 18 answers the direct route got wrong while losing none it g
- 2026-06-27 Qwen3.5-4B Foofah Program Repair Agent
No. Guessing the answer directly won outright, solving 55% of unseen tables versus only 25% for the debugged program. But the program is a useful complement, not a replacement: it rescued 18 tables the direct guess botch
- 2026-06-27 Qwen3.5-4B Foofah Program Ensemble Consensus
No. The simplest rule won: run the first program that passes a single worked example. It solved 52% of tables versus 44% when the model just answered directly, rescuing 23 tables it had otherwise botched while breaking o
- 2026-06-27 Qwen3.5-4B Foofah External Transform Gate
No. Across 250 real spreadsheet-cleanup tasks, the frozen set of moves could rebuild the correct table for only 18% of them. And the weakness is not picking the right sequence: simply grabbing the first move-sequence tha
- 2026-06-27 Qwen3.5-4B Foofah Ephemeral Program Induction
No. Asking directly reshaped 55% of tables correctly; the write-and-test-a-program route managed just 15%. Even a magic chooser that always picked the right route each time would reach only 59% — four points above asking
- 2026-06-27 Qwen3.5-4B Foofah Direct vs ABI
Just ask the model. Directly generating the reshaped table got 55% of 250 table tasks exactly right, versus only 18% for the fixed-operation converter. The model even nailed 103 reshapes the converter could not even expr
- 2026-06-27 Qwen3.5-4B Code ABI Oracle Coverage Ladder
Yes, mostly. Snapping together at most three verified functions produced a passing solution for 84% of 160 small programming tasks, up from just 13% with a bare-bones library. But the library is a foundation, not a solve
- ~2026-06-27 Qwen 3.5 4B Balanced Discriminative Bridge
An even mix, clearly. Sixty evenly-spread ordinary examples across ten new program types lifted the model's success on unseen hard problems from 60% (with no examples at all) to 99% fully solved. Hand-picking only the tr
- 2026-06-26 Qwen Trace Procedure Depth Stress
Yes. Trained only on single-step tasks, the 4-billion-parameter model wrote six-step procedures that ran correctly 63% of the time once a plain step-follower executed them, versus 0% when the model tried to state the fin
- 2026-06-26 Qwen Program-Only Executable ABI
Yes. Teaching a small model to write a short runnable program lifted brand-new multi-step accuracy from about 44% (just stating an answer) to 73%. And when it wrote out its steps plus an answer, the steps ran correctly 9
- 2026-06-26 Qwen Crystallized Trace ABI Tournament
Barely. On familiar inputs, writing out each step scored 94% versus 92% for answer-only, a two-point edge that cost three to six times more generated text. On genuinely new combinations of steps the model had never seen,
- 2026-06-26 Qwen Constrained ABI Parser
Yes, mostly. On the hardest six-step requests, blocking any invalid step as the model writes lifted correctly-running recipes from 60% to 75%, and it won on all five training runs. It also beat merely re-rolling until va
- 2026-06-26 Qwen Compositional Curriculum ABI
Yes. Adding a few two- and three-step examples lifted correct answers on unseen six-step problems from 72% to 83%, and on eight-step problems from 78% to 89%. But adding only two-step examples did nothing (72% stayed 72%
- 2026-06-26 Qwen3.5-4B Substrate Coverage Ladder
Only when a building-block was hand-carved for each specific problem. A shared library that reshapes old solutions to fit new ones solved zero of the nine stuck tasks, no better than nothing. Custom-written parts solved
- 2026-06-25 Qwen Tail Repair Stability Critic
No. A correct fix sat among the candidate rewrites for about nine in ten programs, but the trained judge could not tell which rewrite was right from summary statistics alone, so it played safe and edited nothing, staying
- 2026-06-25 Qwen Readable Candidate Verifier
Barely. Training a picker on the model's reading of each fix's pseudocode plus its claimed output raised the share of tasks fixed correctly from 44% to 51%, closing only about 15% of the distance to a perfect picker's 90
- 2026-06-25 Qwen Candidate-Conditioned Trace Verifier
No. Doing nothing already solved about 44% of tasks, and always choosing a correct candidate could reach 90%. Yet no trained picker captured any of that headroom; the best merely tied doing nothing. Left unchecked, the m
- ~2026-06-25 Qwen3.5-4B HumanEval Adaptive Evidence Budget
No. Every strategy landed at the same 16.7% correct-pick rate — running no tests, running all eight, and even a strategy allowed to peek at the right answer to time its stop. The trained model did learn to quit sooner, u
- 2026-06-24 → 25 Qwen Compiler Multi-Seed Reattribution
No. The best schedule averaged 43% correct on plainly worded problems, yet the identical training swung from total failure to 81% just by changing the random starting number — so no schedule earns credit for the wins. Bo
- 2026-06-24 Qwen Structural Latent Compiler Expansion
Yes, mostly. The model fills fixed slots with operations a calculator runs — no text, no trying many guesses and picking one. After the learned short form was copied into bigger ones, it stayed perfectly correct on 8- an
- 2026-06-24 Qwen Structural Compiler Attribution Ablation
It's the practice schedule. Give the model its full size from the start, then feed examples easy-to-hard — 8 steps, then 16, then 24 — and it solves nearly every standard 24-step program (about 97%). The popular guess, g
- 2026-06-24 Qwen3.5-4B Sketch Coverage Shift Probe
Only when the template explicitly names the new operation. On tasks needing a brand-new operation, hand-written templates that named and typed it kept the correct program every single time; auto-generated templates kept
- 2026-06-24 Qwen3.5-4B Oracle Process GRPO
Yes, but with sharp limits. Teaching the model to copy a flawless test-picker lifted its success from about a third of puzzles to 43%, versus roughly 31% for picking blindly. Layering on preference-based and reward-based
- 2026-06-24 Qwen3.5-4B Joint Shortlister Ladder
No. Across every version — untrained, trained, and with the glossary's descriptions scrambled — the model got both codes exactly right zero percent of the time, even when allowed sixteen guesses. Training pushed single-c
- 2026-06-24 Qwen3.5-4B Adaptive Evidence Budget Policy
Yes. Trained just for this single decision, the model matched the accuracy of running all ten checks — solving about 92 of every 100 tasks — while using only about five checks on average, near a perfect-hindsight stopper
- 2026-06-23 Qwen Semantic Prefix Value Model
No. Scoring each step by whether a correct answer is still reachable pushed the top pick to about 68 percent, level with plain confidence search and short of the 71 percent from scoring steps against the known correct pr
- 2026-06-23 Qwen Prefix-State Process Verifier
Barely. The judge got genuinely good at telling promising partial programs from dead ends, yet it lifted first-try accuracy only a few points, with hard problems going from 41% to 44%. The real gap: on hard problems a co
- 2026-06-23 Qwen Mixed-Domain Trace Verifier
Yes, partly. On fresh tasks the frozen model alone got 46% right; the proofreader lifted that to 57%, and a cross-check that compares reworded versions of the same task reached 61%. But a correct recipe was already among
- 2026-06-23 Qwen LoRA Typed-Bytecode Trace Compiler
Yes — but the win came from the teaching material, not from adapting the model. Fed fully worked recipes, it wrote a runnable recipe that reached the right answer about 68% of the time, versus only 15% when taught with f
- 2026-06-23 Qwen Hidden VM Curriculum Repair
No. Feeding it nearby worksheets that merely land on the correct answer wrecked it. Plain step-by-step training scored 72% on the main mixed test; the same model after answer-chasing repair fell to 35%, and collapsed on
- 2026-06-23 Qwen Budgeted Action-Value Compiler
Barely. On fresh problems the model drafts a correct program among its candidates 81% of the time but ranks it first only 67% of the time. A learned scorer that never sees the answer nudged that to just 70% — while a con
- 2026-06-22 Targeted Bridge Allocation
Barely. Piling examples on the hardest never-seen combinations repaired 33% of test cases versus 28% for spreading them evenly — a three-case edge on sixty tests, and no better than the same budget spent on easy combinat
- 2026-06-22 Qwen Checkpoint-Selected Scheduled-State Compiler
It depends. A light, steady dose of show-your-work coaching, kept on through the hardest problems, lifted correct answers on 24-step chains from 25% to 33% and nearly doubled agreement between two wordings of the same pr
- 2026-06-22 Qwen 3.5 4B Unsaturated Frontier Active Bridge
Spread evenly. Giving each of ten problem types the same six extra correction examples let the model fully fix 98% of hard cases. Piling those same examples onto whichever types it failed most reached only 85%, and starv
- 2026-06-22 Qwen 3.5 4B Model-In-Loop Counterexamples
No. Building practice cases from the model's actual wrong answers matched, but never beat, simply hand-picking the tricky categories in advance. Both lifted the hardest problems from 64% to a perfect 100% passing every h
- 2026-06-22 Qwen 3.5 4B Executable Program Posttraining
Yes, but with a catch. Shown worked-through reasoning in the prompt, the model fixed unseen problem types about three-quarters of the time, versus one-in-three when the prompt showed no steps. Strip out or scramble those
- 2026-06-22 Qwen 3.5 4B Counterexample-Directed DSL
It depends. Hand-picked examples lifted the model's single-best-guess repair rate from 51% to 58% over random examples. But when it generated several candidates and kept the best, the edge vanished (64% slipped to 61%).
- 2026-06-21 → 22 Bridge-Dose Recombination Curriculum
Yes, but only with the right examples. With none, the model solved just 7% of never-seen skill pairings; adding as few as two to four exact worked examples of each pairing lifted that to about 30%, a nearly five-fold jum
- 2026-06-21 Structured Slot Initializer Ladder Experiment
Structure, not scale. A plain general-purpose network placed only 55.5% of its belief on the correct starting setup and gave barely half the possible values their own slot, doubling several onto the same one. A rule that
- 2026-06-21 Qwen Structured Bridge Experiment
Yes, but only when you show it the individual steps during training. A tiny translator turning the frozen model's read into calculator instructions solved chains far longer than it trained on: 96% correct at twelve steps
- 2026-06-21 Qwen State-Ladder Compiler
No. Grading the running total after every step never beat an identical model graded only on its final answer, and at full strength it collapsed on the hardest long programs. The real winner was the training schedule: sta
- 2026-06-21 Qwen Span-Free Compiler
Only when it is first taught where to look. Fed just the frozen model's raw internal notes, a plain reader stayed near random guessing (about 1 in 97). Adding training that also highlighted which spots held the numbers a
- 2026-06-21 Qwen Shared Parser Compiler
Only when every step is taught directly. Given step-by-step labels, the add-on rebuilds short programs well — nearly 4 in 5 four-step problems run exactly right — but accuracy fades to 39% at twelve steps and under 1% at
- 2026-06-21 Feature-Factorized Rule Diversity
No. All three practice diets fixed only about 1 in 5 brand-new bug combinations, so mixing bought nothing over drilling either kind alone. What actually mattered was showing worked, step-by-step repair reasoning during t
- 2026-06-21 Factor Recombination Ladder
No. Trained on worked solutions, the model fixed about 81% of bugs when two skills were paired the way it saw in training, but only 8 to 10% when the same familiar skills were paired in a new way. Adding skill labels lif
- 2026-06-21 Cyclic Transition Ladder Experiment
It needs the matching wrap-around parts. A network built from clock-arithmetic moves stayed perfectly exact on programs three times longer than it practiced on. A plain generic network of the same size drifted down to ju
Claims
- Confirmed C6 · Controls are the difference between a result and a story
- Promising C37 · The compositional wall does NOT exist in language: the model chains depth-3+ multi-step SIMULATION in natural language near-perfectly -- the wall is formal-modality-specific, not a general multi-step limit
- Promising C38 · The structure-PROPOSAL wall persists in language: the model executes a given rule but cannot induce one -- proposal/induction is modality-general, while simulation is formal-specific (C37)
- Promising C39 · In-context learning is RETRIEVAL of familiar structure, not INDUCTION of novel structure: the model executes a novel rule perfectly (0.97) but cannot induce it from examples (0.12=chance)
- Promising C40 · The model knows when it will fail IMPLICITLY (answer-token probability, within-cell AUROC 0.95) but NOT EXPLICITLY (self-verification P(True) at chance, verbalized confidence a constant 100)
- Promising C41 · Beat sample-more with the model's own uncertainty: confidence-select (argmax P(answer), verification-free) beats self-consistency, which is flat; max P(answer) predicts solvability (AUROC 0.83) for abstention
- Promising C42 · The model can localize its own errors: per-step confidence dips exactly at the first slip (surviving de-trending), enabling targeted repair -- C40's metacognition is step-resolved
- Promising C43 · SFT PARTIALLY lifts the induction wall but does not cleanly install the skill: data-limited (0.087->0.40) yet below the execute ceiling, procedure-specific (weak OOF transfer), and catastrophically forgets execution
- Promising C44 · The forward-pass induction wall is a SERIAL-COMPUTE limit, not a knowledge limit: reasoning-SFT induces held-out rules PERFECTLY via generation (1.00) but at CHANCE in one forward pass (0.01) -- the CoT is 100% load-bearing
- Promising C45 · GENERAL induction-via-reasoning IS installable: a general hypothesize-and-verify CoT trained multi-family transfers to a HELD-OUT rule family (a=7: 0.905, as high as in-family) -- resolving C44's shift-specificity
- Promising C46 · The confidence toolkit TRANSFERS to real code (MBPP + HumanEval) -- but the hierarchy INVERTS: the single-token P(True) judge readout, not sequence mean-logprob, is the program-level confidence
- Promising C50 · Breadth-first expert iteration on a firewall-clean gym INSTALLS SUBSTRATE-GENERAL agentic competence: +0.22/+0.29 on blackbox menagerie quick (paired, deterministic) and +0.52 gym-wide including never-trained families -- the locality laws (C43/C45/C48) do not extend to this regime, and the causal lever was gradient placement at the answer-emission seam, not dose
- Promising C53 · THE SECOND WALL: the emission-policy install is a large ONE-TIME step to a robust menagerie ceiling (quick later broken to ~0.50 by convex mix composition; medium arm-means top out ~+0.31) — no variant of train-on-own-verified-outputs (dose, iteration, breadth, difficulty escalation, recovery supervision, deploy-budget matching) moves the blackbox band further, even as in-gym frontier competence installs