Program scorecards
Rendered from knowledge/program_scorecards.md
On this page
- Structured Execution And Compilers
- Evidence-Conditioned Selection
- Active Evidence Acquisition
- Algorithmic Memory And Retrieval
- Operator And Skill Inventories
- Posttraining And Adaptation
- Process Control And Tool Use
- Benchmark Generalization
- Interpretability And Diagnostics
- Reliability And Safety
- Collective Experimentation Infrastructure
- Test-Time Reasoning Budget
- Agentic Breadth Installation
Use these scorecards before opening a new experiment. They are intentionally short: the goal is to route ideas, avoid duplicate variants, and choose the next result that would most change the repository's shared beliefs.
For evidence-linked durable claims, use claims/index.md.
Structured Execution And Compilers
- Program: charter
- Current read: structured intermediates remain strong, but tested latent interfaces usually fail before causal use. The replicated
First:seam is neither a stable J value coordinate nor an arbitrary writable last-thought register. Early concrete text demonstrates broad local routing—84/96 execution for both correct and deranged supplied operations across all 24 operations—but its interface failed and complete-program reachability was only 3/8. The durable fresh materialized-residual run split the question: free generation was interface-invalid, while the parse-immune cheap materialized ranker lost to every structured comparator. Tokenizer EOS plus no-think is a real short-output interface: two independent calibrations reached 48/48 in both prefix cells against 0/48 for every HF-EOS control, and fresh transport was 24/24. The replay-hardened mechanics result now cleanly rejects the prompt mechanism: all 4,056 outputs authenticated with 98.78--99.83% parse and zero caps, yet materialized and all matched controls had 0/24 selected success and 0/24 oracle proposal coverage while exhaustive search covered 24/24 tasks. A strict answer seam does not expose residual synthesis. Cheap viability/top-four ranking, parser repair, and all-candidate semantic-materialization prompting are retired. The fresh three-seed rank-32 LoRA joint adjudication also missed registered state in all 57 required cells (best 0.0234 versus 0.40); because adaptation contrast is uncertain, mandatory LoRA-state-only and matched direct-full-shape controls—not a rank conclusion—come next. - Best next experiments: complete the authorized state-formation Stage B with three LoRA state-only and three matched direct-full-shape joint seeds; separately measure the exact depth-6 resource crossover before learned pruning; and, for J-space, require a task-held-out correctness coordinate beyond margin/position/equal-width non-J controls before a separately preregistered same-prefix intervention must causally increase correct-proposal coverage over matched sampling.
- Strong anchors:
qwen35_4b_early_text_hypothesis_forking,qwen35_4b_partial_structure_search,qwen35_4b_crosssubstrate_structure,qwen35_4b_structure_search_scaling,qwen35_4b_commit_slot_semantic_power_replication,qwen35_4b_state_carry_vs_state_bag_fullrank_delta. - Avoid repeating: another opaque name/timing variant, this cheap materialized P(viable)/top-four ranker, the all-candidate semantic-materialization prompt with a looser selector or more samples, waiting for HF model EOS when tokenizer EOS is the registered answer commit, adding thinking to a short exact-output channel without a task-specific benefit, larger cap or parser relaxation after ABI failure, pooled-AUROC launch, model-guided search without measured brute wall time, or a rank comparison whose shared modules and dropout streams are shifted by parameter-construction RNG.
- Evidence that advances the program: causal ablation showing which intermediate structure transfers across family, length, or paraphrase shifts.
Evidence-Conditioned Selection
- Program: charter
- Current read: confidence can rank completed candidates, but value signals do not automatically become controllers. The coherent-order delta beat majority but not confidence/entropy; additive-J branching failed at chance; a late semantic anchor wrote names without a valid consequence interface; and early concrete text routed direct operations but failed before proposal selection or any matched-sampling comparison.
- Best next experiment: close the exact-pool visible-selector gap; for training-time policy routing, fit direct cross-fitted advantages and confirm a frozen rule on a third block rather than retrying statewise argmax.
- Strong anchors:
qwen35_4b_partial_structure_search,qwen35_4b_generator_verifier_gap,qwen35_4b_code_confidence,qwen35_4b_answer_potential_trace_sft,qwen35_4b_same_prefix_advantage_routing. - Avoid repeating: pooled-AUROC confidence claims, type-only partial judges, answer-only potential over cap-bound traces, posthoc winner margins, raw ordered-minus-shuffle commit-logit tuning, or gains that hide abstention/commit-rate changes.
- Evidence that advances the program: deployable selection gains under family-held-out candidate pools and adversarial visible examples.
Active Evidence Acquisition
- Program: charter
- Current read: active examples and probes are promising, but the acquisition objective must be tied to downstream decision quality.
- Best next experiment: compare active probe policies by downstream selector lift under a fixed evidence budget.
- Strong anchors:
qwen_active_example_acquisition,qwen35_4b_active_counterexample_trace_selection,qwen35_4b_learned_active_trace_policy. - Avoid repeating: optimizing probe informativeness without showing that decisions improve.
- Evidence that advances the program: budget-normalized gains from probes that do not rely on expected answers at deployment time.
Algorithmic Memory And Retrieval
- Program: charter
- Current read: memory helps when it supplies verifiable candidates, constraints, or tests; naive context stuffing is weak.
- Best next experiment: compare retrieved examples, retrieved algorithms, retrieved tests, and retrieved failure cases on one task family.
- Strong anchors:
qwen_verified_skill_memory_rag,qwen35_4b_verified_algorithm_retrieval_adaptation,learned_sparse_slot_executor. - Avoid repeating: retrieval demos without negative retrieval controls or verification of the retrieved artifact.
- Evidence that advances the program: memory improves transfer while random or mismatched memory fails under the same budget.
Operator And Skill Inventories
- Program: charter
- Current read: inventories can scale coverage, but only if search and shortlisting remain reliable as the bank grows.
- Best next experiment: stress a skill/operator shortlister as distractor inventory size increases.
- Strong anchors:
qwen35_4b_operator_inventory_search_pilot,qwen35_4b_operator_inventory_scaling_stress,qwen35_4b_inventory_shortlister_training. - Avoid repeating: small-bank wins that do not test distractors, type collisions, or compositional reuse.
- Evidence that advances the program: graceful degradation curves and recovery strategies for large noisy inventories.
Posttraining And Adaptation
- Program: charter
- Current read: adaptation can reshape behavior, but the beyond-C53 line now exposes seven launch boundaries: use deployable outcome-valid targets; preserve neighboring behavior; make every absolute gate feasible; verify teachers on the student's same-prefix distribution; estimate conditional teacher advantage without winner's curse; preserve verifier-conditioned recovery; and prove that the training update survives the deployed merge. Deep-only routed MOPD passed route, locality, retention, transfer, and directional teacher/control tests, yet all three seeds trailed deep and primary lost to interpolation and sample-more. The NF4 objective gain reversed after bf16 merge.
- Best next experiment: direct-bf16 deployment-parity microtraining with a hard merged-checkpoint gate against deep, interpolation, and matched-compute sampling. Only after it passes should two-teacher work add cross-fitted direct advantages, adaptive allocation (including zero quick), and a third untouched block.
- Strong anchors:
qwen35_4b_gauntlet_breadth_round1,qwen35_4b_gauntlet_frontier,qwen35_4b_bank_the_thoughts,qwen35_4b_answer_potential_trace_sft,qwen35_4b_think_ftpo_round2,qwen35_4b_specialist_policy_integration,qwen35_4b_pareto_policy_integration,qwen35_4b_same_prefix_advantage_routing,qwen35_4b_deep_advantage_mopd,qwen35_4b_repo_search_compress_bank. - Avoid repeating: unauditable adapters, hidden-label wins without frozen alternatives, infeasible mandatory arms, coarse teacher labels, four-branch statewise argmax as a two-teacher labeler, posthoc route margins, non-local sparse steering, or success-only banks that erase recovery.
- Evidence that advances the program: a trained behavior beats strong frozen/tool baselines without hidden-label leakage.
Process Control And Tool Use
- Program: charter
- Current read: small models need explicit control policies for tools, budgets, and commit/repair decisions. Exact operator marginals are not enough: a compact repository bank preserved commit after pass but learned zero patch recovery after 48 failed tests, regressing held-out success 49/72→25/72.
- Best next experiment: compare STOP/MORE, commit/repair, and tool-choice policies with strict cost accounting and explicit failed-patch/failed-test transition gates; a recovery curriculum must contain changed second actions, not only successful terminal traces.
- Strong anchors:
qwen35_4b_adaptive_tool_controller,qwen35_4b_tool_state_policy_lora,qwen35_4b_adaptive_evidence_budget_policy,qwen35_4b_repo_search_compress_bank. - Avoid repeating: tool-use demos that omit the no-tool, fixed-tool, and random-tool baselines, or banks that balance action counts while deleting verifier-rejection contingencies.
- Evidence that advances the program: policy lift survives latency, token, and tool-call ceilings.
Benchmark Generalization
- Program: charter
- Current read: many mechanisms look good in-family; the repository needs standard transfer stress before strategic claims harden. C46 shows the confidence toolkit survives MBPP->HumanEval only after the signal is re-expressed as a single-token P(True) readout. The Pareto qualification negative adds the reverse warning: a held-out instrument ranking can be real on that instrument yet fail to define a useful teacher ordering on the clean training proxy. Native-thinking termination is also non-portable: a 512--1024 scale that often closed on MBPP produced 0/48 natural closes at 1,024 on fresh list induction.
- Best next experiment: compositional-grammar induction as the C45 stress test, plus a small cross-program generalization suite used by compiler, selector, memory, and adaptation work.
- Strong anchors:
factor_recombination_ladder,feature_factorized_rule_diversity,targeted_bridge_allocation. - Avoid repeating: reporting only IID or narrow held-out splits for a mechanism meant to generalize.
- Evidence that advances the program: transfer across substrate, family, length, and real-task variants with a clear failure taxonomy.
Interpretability And Diagnostics
- Program: charter
- Current read: early synthetic donor-coordinate J transport and the cap-1,024 semantic seam replicate, but shared/midpoint J value, terminal counterfactual selection, additive native branching, and the late semantic-anchor bridge do not yield a usable controller. Early concrete text supplies local causal routing. Tokenizer EOS plus no-think is a replicated 48/48 short-output seam against 0/48 HF-EOS controls, and mechanics transport is 24/24. The replay-hardened residual comparison is now cleanly negative: materialized and every matched comparator had 0/24 oracle proposal coverage despite healthy ABI and 24/24 exhaustive task solvability. Semantic addressing and a strict token-native commit boundary are real locally; semantic materialization does not create proposal competence.
- Best next experiment: retire native J token/layer/scale, opaque-name timing, cheap materialized viability, and the all-candidate semantic-materialization prompt. If J-space continues, use a fresh measurement-first correctness coordinate inside
<think>with task-held-out and equal-width non-J controls, then require a separate same-prefix intervention to increase correct-proposal coverage over matched sampling. Use a readable J coordinate as diagnosis only until that forward causal gate passes. - Strong anchors:
qwen_structural_compiler_attribution_ablation,qwen35_4b_probe_to_prompt,qwen35_4b_jacobian_value_transport,qwen35_4b_context_local_jacobian_clamp,qwen35_4b_jacobian_transport_control_replication,qwen35_4b_native_thought_jacobian_value_transport,qwen35_4b_native_thought_seam_budget_ladder,qwen35_4b_forced_commit_jacobian_value_transport,qwen35_4b_commit_slot_jacobian_value_transport,qwen35_4b_commit_slot_semantic_power_replication. - Avoid repeating: probes that do not change the next experiment, next-token writing tests presented as reasoning transport, component permutations without a composed-map test, post-hoc parser repair, promoting a perfect point estimate past a failed control gate, or presenting an oracle donor as a deployable gain.
- Evidence that advances the program: a diagnostic predicts which variants will fail before the final metric is observed.
Reliability And Safety
- Program: charter
- Current read: reliability depends on precision, abstention, artifact hygiene, and explicit hidden-label boundaries.
- Best next experiment: create a reliability scorecard applied to selectors, verifiers, and tool controllers.
- Strong anchors:
qwen35_4b_reliability_exec_opsd_audit,qwen35_4b_real_sample_verify_commit,qwen_readable_candidate_verifier. - Avoid repeating: impressive accuracy reports without commit-rate, abstention, artifact, and leakage checks.
- Evidence that advances the program: reliability metrics expose regressions that ordinary accuracy hides.
Collective Experimentation Infrastructure
- Program: charter
- Current read: the repository now has program structure, validation, CI, and collaboration templates; the next lift is faster research navigation.
- Best next experiment: measure whether a new agent can find prior evidence and propose a non-duplicate experiment faster using scorecards and intake records.
- Strong anchors:
knowledge/research_program_index.md,knowledge/claims/initial_claims.md,docs/quality_gates.md. - Avoid repeating: adding process that slows pilots without improving memory or decision quality.
- Evidence that advances the program: navigation artifacts reduce duplicate proposals and improve citation of prior evidence.
Test-Time Reasoning Budget
- Program: charter
- Current read: native thinking is a real coherent-content lever, but its value geometry and write sites are workload/phase-specific. Fixed
First:exposes a replicated semantic state, shared/midpoint value and last-token additive branching fail, and early concrete text routes one-operation execution without reaching a valid full-program controller. The answer seam qualifies no-think tokenizer EOS at 48/48 in both prefix cells against 0/48 HF EOS, while thinking is worse at 38/48 structured and 30/48 freeform. Replay-hardened mechanics then passed transport and ABI but produced 0/24 oracle proposal coverage for materialized and every matched comparator. The short no-think seam controls emission; it does not unlock residual composition. - Best next experiment: retain the separate 16k+ loop-control line. For J-space, measure whether any within-
<think>coordinate predicts eventual correctness across held-out tasks beyond margin/position/equal-width non-J controls, and only then test a same-prefix causal branch intervention against matched sampling. - Strong anchors:
qwen35_4b_thinking_content_vs_compute,qwen35_4b_overthinking_content_ladder,qwen35_4b_answer_potential_trace_sft,qwen35_4b_native_thought_seam_budget_ladder,qwen35_4b_forced_commit_jacobian_value_transport,qwen35_4b_commit_slot_jacobian_value_transport,qwen35_4b_commit_slot_semantic_power_replication,qwen35_4b_think_ftpo_round2. - Avoid repeating: thinking-budget wins without content controls, calibration on a different workload class, cap-bound score interpretation, larger-N harvesting before termination/locality works, or treating high varentropy as a monotone “push harder” signal.
- Evidence that advances the program: a controller or distillation that Pareto-beats fixed budgets, and a content control that isolates genuine reasoning from compute + scaffold + token-presence.
Agentic Breadth Installation
- Program: charter
- Current read: breadth-first expert iteration remains the first blackbox-arbitrated install (+0.223/+0.294 menagerie quick; C49/C50), but C53 closes same-recipe scaling at a robust second wall. Transaction action supervision locally installed proposal structure without transfer. Direct failed-test semantic revision is now descriptively near-saturated: all 72 headroom cases reached a correct patch, while inferred rejected-state first-patch correctness was 0/54 and test inspection before patching 0/72. Structure, evidence acquisition before proposal, and explicit policy editing are separate units. The headroom result is formally instrument-failed because answer-cap contacts exceeded its frozen ceiling. The counterfactual evidence-acquisition strategy remains open: its first curriculum stopped at
LINEAGE_LOCALITY_INFEASIBLE(0.110735 direct start-to-apex drift versus 0.10, entropy retained) before behavior or training, so it is not a capability negative. The universal-curriculum line now has three consecutive procedural-interface negatives: canonical staged search scored 16/26, natural-language state tables scored 16/26, and on-policy failure-prefix correction scored 15/26 against replay's 18/26. The last arm was 0/6 on execute+induct+probe despite 230 reachable parent failures, so on-policy substrate did not overcome late-prefix teacher forcing. - Active experiments: none on this exact MOPD line; the deep-advantage experiment is finished negative after sealed confirmation and its frozen stop prevented benchmark exposure.
qwen35_4b_counterfactual_evidence_acquisition_curriculumis finished at the preregistered lineage prerequisite; an apex-rooted or prospectively parent-qualified successor has not opened. The latest universal on-policy prefix experiment is terminal negative; no aggregate seed was consumed. - Strong anchors:
qwen35_4b_gauntlet_breadth_round1,qwen35_4b_gauntlet_frontier(C53/C54),qwen35_4b_specialist_policy_integration,qwen35_4b_pareto_policy_integration,qwen35_4b_same_prefix_advantage_routing,qwen35_4b_deep_advantage_mopd,qwen35_4b_verifier_conditioned_recovery_bank,qwen35_4b_recovery_reason_locality_interpolation,qwen35_4b_recovery_payload_budget_harness,qwen35_4b_recovery_verifier_branch_tournament,qwen35_4b_transaction_invariant_recovery_curriculum,qwen35_4b_validation_policy_counterexample_curriculum,qwen35_4b_counterfactual_evidence_acquisition_curriculum,qwen35_4b_universal_search_scaffold_token_match,qwen35_4b_universal_state_table_compiler_token_match,qwen35_4b_universal_on_policy_prefix_repair_token_match,qwen35_4b_think_ftpo_round2(C52),qwen35_4b_interactive_policy_curriculum,qwen35_4b_repo_search_compress_bank. - Avoid repeating: un-gated vLLM runtime LoRA; near-self-distillation; filtering deployment-critical force-closes; backend-mixed scores; infeasible specialists; external tier labels; four-branch statewise argmax or posthoc route margins; scaling NF4 MOPD before direct-bf16 merge survival is proven; non-local thought LoRA; unbalanced DAgger; success-only banks without recovery transitions; more generic transaction families; canonical two-operation/two-branch search scaffolds; idealized state-table traces; long masked failure-prefix continuation with reduced target exposure; capability-production training before exact-substrate parent headroom is demonstrated; carrying a useful parent into a stricter direct apex-relative locality contract without prospective qualification; or relaxing a locality ceiling after observing the miss.
- Evidence that advances the program: a beyond-recipe method that exceeds the C53 blend ceiling on fresh blackbox events, or a local checkpoint that learns the shared transactional failure core while retaining broad conditional recovery and beating sample-more on two unseen procedural blocks; for the universal-curriculum line, the next checkpoint must intervene at short pre-failure decisions, match target exposure, and improve execution, induction, and probe scoring under the frozen absolute and strict control-relative local gates.