Research log Small Model Experimentation
GitHub

Program scorecards

Rendered from knowledge/program_scorecards.md

On this page
  1. Structured Execution And Compilers
  2. Evidence-Conditioned Selection
  3. Active Evidence Acquisition
  4. Algorithmic Memory And Retrieval
  5. Operator And Skill Inventories
  6. Posttraining And Adaptation
  7. Process Control And Tool Use
  8. Benchmark Generalization
  9. Interpretability And Diagnostics
  10. Reliability And Safety
  11. Collective Experimentation Infrastructure
  12. Test-Time Reasoning Budget
  13. Agentic Breadth Installation

Use these scorecards before opening a new experiment. They are intentionally short: the goal is to route ideas, avoid duplicate variants, and choose the next result that would most change the repository's shared beliefs.

For evidence-linked durable claims, use claims/index.md.

Structured Execution And Compilers

  • Program: charter
  • Current read: structured intermediates remain strong, but tested latent interfaces usually fail before causal use. The replicated First: seam is neither a stable J value coordinate nor an arbitrary writable last-thought register. Early concrete text demonstrates broad local routing—84/96 execution for both correct and deranged supplied operations across all 24 operations—but its interface failed and complete-program reachability was only 3/8. The durable fresh materialized-residual run split the question: free generation was interface-invalid, while the parse-immune cheap materialized ranker lost to every structured comparator. Tokenizer EOS plus no-think is a real short-output interface: two independent calibrations reached 48/48 in both prefix cells against 0/48 for every HF-EOS control, and fresh transport was 24/24. The replay-hardened mechanics result now cleanly rejects the prompt mechanism: all 4,056 outputs authenticated with 98.78--99.83% parse and zero caps, yet materialized and all matched controls had 0/24 selected success and 0/24 oracle proposal coverage while exhaustive search covered 24/24 tasks. A strict answer seam does not expose residual synthesis. Cheap viability/top-four ranking, parser repair, and all-candidate semantic-materialization prompting are retired. The fresh three-seed rank-32 LoRA joint adjudication also missed registered state in all 57 required cells (best 0.0234 versus 0.40); because adaptation contrast is uncertain, mandatory LoRA-state-only and matched direct-full-shape controls—not a rank conclusion—come next.
  • Best next experiments: complete the authorized state-formation Stage B with three LoRA state-only and three matched direct-full-shape joint seeds; separately measure the exact depth-6 resource crossover before learned pruning; and, for J-space, require a task-held-out correctness coordinate beyond margin/position/equal-width non-J controls before a separately preregistered same-prefix intervention must causally increase correct-proposal coverage over matched sampling.
  • Strong anchors: qwen35_4b_early_text_hypothesis_forking, qwen35_4b_partial_structure_search, qwen35_4b_crosssubstrate_structure, qwen35_4b_structure_search_scaling, qwen35_4b_commit_slot_semantic_power_replication, qwen35_4b_state_carry_vs_state_bag_fullrank_delta.
  • Avoid repeating: another opaque name/timing variant, this cheap materialized P(viable)/top-four ranker, the all-candidate semantic-materialization prompt with a looser selector or more samples, waiting for HF model EOS when tokenizer EOS is the registered answer commit, adding thinking to a short exact-output channel without a task-specific benefit, larger cap or parser relaxation after ABI failure, pooled-AUROC launch, model-guided search without measured brute wall time, or a rank comparison whose shared modules and dropout streams are shifted by parameter-construction RNG.
  • Evidence that advances the program: causal ablation showing which intermediate structure transfers across family, length, or paraphrase shifts.

Evidence-Conditioned Selection

  • Program: charter
  • Current read: confidence can rank completed candidates, but value signals do not automatically become controllers. The coherent-order delta beat majority but not confidence/entropy; additive-J branching failed at chance; a late semantic anchor wrote names without a valid consequence interface; and early concrete text routed direct operations but failed before proposal selection or any matched-sampling comparison.
  • Best next experiment: close the exact-pool visible-selector gap; for training-time policy routing, fit direct cross-fitted advantages and confirm a frozen rule on a third block rather than retrying statewise argmax.
  • Strong anchors: qwen35_4b_partial_structure_search, qwen35_4b_generator_verifier_gap, qwen35_4b_code_confidence, qwen35_4b_answer_potential_trace_sft, qwen35_4b_same_prefix_advantage_routing.
  • Avoid repeating: pooled-AUROC confidence claims, type-only partial judges, answer-only potential over cap-bound traces, posthoc winner margins, raw ordered-minus-shuffle commit-logit tuning, or gains that hide abstention/commit-rate changes.
  • Evidence that advances the program: deployable selection gains under family-held-out candidate pools and adversarial visible examples.

Active Evidence Acquisition

Algorithmic Memory And Retrieval

  • Program: charter
  • Current read: memory helps when it supplies verifiable candidates, constraints, or tests; naive context stuffing is weak.
  • Best next experiment: compare retrieved examples, retrieved algorithms, retrieved tests, and retrieved failure cases on one task family.
  • Strong anchors: qwen_verified_skill_memory_rag, qwen35_4b_verified_algorithm_retrieval_adaptation, learned_sparse_slot_executor.
  • Avoid repeating: retrieval demos without negative retrieval controls or verification of the retrieved artifact.
  • Evidence that advances the program: memory improves transfer while random or mismatched memory fails under the same budget.

Operator And Skill Inventories

Posttraining And Adaptation

  • Program: charter
  • Current read: adaptation can reshape behavior, but the beyond-C53 line now exposes seven launch boundaries: use deployable outcome-valid targets; preserve neighboring behavior; make every absolute gate feasible; verify teachers on the student's same-prefix distribution; estimate conditional teacher advantage without winner's curse; preserve verifier-conditioned recovery; and prove that the training update survives the deployed merge. Deep-only routed MOPD passed route, locality, retention, transfer, and directional teacher/control tests, yet all three seeds trailed deep and primary lost to interpolation and sample-more. The NF4 objective gain reversed after bf16 merge.
  • Best next experiment: direct-bf16 deployment-parity microtraining with a hard merged-checkpoint gate against deep, interpolation, and matched-compute sampling. Only after it passes should two-teacher work add cross-fitted direct advantages, adaptive allocation (including zero quick), and a third untouched block.
  • Strong anchors: qwen35_4b_gauntlet_breadth_round1, qwen35_4b_gauntlet_frontier, qwen35_4b_bank_the_thoughts, qwen35_4b_answer_potential_trace_sft, qwen35_4b_think_ftpo_round2, qwen35_4b_specialist_policy_integration, qwen35_4b_pareto_policy_integration, qwen35_4b_same_prefix_advantage_routing, qwen35_4b_deep_advantage_mopd, qwen35_4b_repo_search_compress_bank.
  • Avoid repeating: unauditable adapters, hidden-label wins without frozen alternatives, infeasible mandatory arms, coarse teacher labels, four-branch statewise argmax as a two-teacher labeler, posthoc route margins, non-local sparse steering, or success-only banks that erase recovery.
  • Evidence that advances the program: a trained behavior beats strong frozen/tool baselines without hidden-label leakage.

Process Control And Tool Use

  • Program: charter
  • Current read: small models need explicit control policies for tools, budgets, and commit/repair decisions. Exact operator marginals are not enough: a compact repository bank preserved commit after pass but learned zero patch recovery after 48 failed tests, regressing held-out success 49/72→25/72.
  • Best next experiment: compare STOP/MORE, commit/repair, and tool-choice policies with strict cost accounting and explicit failed-patch/failed-test transition gates; a recovery curriculum must contain changed second actions, not only successful terminal traces.
  • Strong anchors: qwen35_4b_adaptive_tool_controller, qwen35_4b_tool_state_policy_lora, qwen35_4b_adaptive_evidence_budget_policy, qwen35_4b_repo_search_compress_bank.
  • Avoid repeating: tool-use demos that omit the no-tool, fixed-tool, and random-tool baselines, or banks that balance action counts while deleting verifier-rejection contingencies.
  • Evidence that advances the program: policy lift survives latency, token, and tool-call ceilings.

Benchmark Generalization

  • Program: charter
  • Current read: many mechanisms look good in-family; the repository needs standard transfer stress before strategic claims harden. C46 shows the confidence toolkit survives MBPP->HumanEval only after the signal is re-expressed as a single-token P(True) readout. The Pareto qualification negative adds the reverse warning: a held-out instrument ranking can be real on that instrument yet fail to define a useful teacher ordering on the clean training proxy. Native-thinking termination is also non-portable: a 512--1024 scale that often closed on MBPP produced 0/48 natural closes at 1,024 on fresh list induction.
  • Best next experiment: compositional-grammar induction as the C45 stress test, plus a small cross-program generalization suite used by compiler, selector, memory, and adaptation work.
  • Strong anchors: factor_recombination_ladder, feature_factorized_rule_diversity, targeted_bridge_allocation.
  • Avoid repeating: reporting only IID or narrow held-out splits for a mechanism meant to generalize.
  • Evidence that advances the program: transfer across substrate, family, length, and real-task variants with a clear failure taxonomy.

Interpretability And Diagnostics

  • Program: charter
  • Current read: early synthetic donor-coordinate J transport and the cap-1,024 semantic seam replicate, but shared/midpoint J value, terminal counterfactual selection, additive native branching, and the late semantic-anchor bridge do not yield a usable controller. Early concrete text supplies local causal routing. Tokenizer EOS plus no-think is a replicated 48/48 short-output seam against 0/48 HF-EOS controls, and mechanics transport is 24/24. The replay-hardened residual comparison is now cleanly negative: materialized and every matched comparator had 0/24 oracle proposal coverage despite healthy ABI and 24/24 exhaustive task solvability. Semantic addressing and a strict token-native commit boundary are real locally; semantic materialization does not create proposal competence.
  • Best next experiment: retire native J token/layer/scale, opaque-name timing, cheap materialized viability, and the all-candidate semantic-materialization prompt. If J-space continues, use a fresh measurement-first correctness coordinate inside <think> with task-held-out and equal-width non-J controls, then require a separate same-prefix intervention to increase correct-proposal coverage over matched sampling. Use a readable J coordinate as diagnosis only until that forward causal gate passes.
  • Strong anchors: qwen_structural_compiler_attribution_ablation, qwen35_4b_probe_to_prompt, qwen35_4b_jacobian_value_transport, qwen35_4b_context_local_jacobian_clamp, qwen35_4b_jacobian_transport_control_replication, qwen35_4b_native_thought_jacobian_value_transport, qwen35_4b_native_thought_seam_budget_ladder, qwen35_4b_forced_commit_jacobian_value_transport, qwen35_4b_commit_slot_jacobian_value_transport, qwen35_4b_commit_slot_semantic_power_replication.
  • Avoid repeating: probes that do not change the next experiment, next-token writing tests presented as reasoning transport, component permutations without a composed-map test, post-hoc parser repair, promoting a perfect point estimate past a failed control gate, or presenting an oracle donor as a deployable gain.
  • Evidence that advances the program: a diagnostic predicts which variants will fail before the final metric is observed.

Reliability And Safety

  • Program: charter
  • Current read: reliability depends on precision, abstention, artifact hygiene, and explicit hidden-label boundaries.
  • Best next experiment: create a reliability scorecard applied to selectors, verifiers, and tool controllers.
  • Strong anchors: qwen35_4b_reliability_exec_opsd_audit, qwen35_4b_real_sample_verify_commit, qwen_readable_candidate_verifier.
  • Avoid repeating: impressive accuracy reports without commit-rate, abstention, artifact, and leakage checks.
  • Evidence that advances the program: reliability metrics expose regressions that ordinary accuracy hides.

Collective Experimentation Infrastructure

  • Program: charter
  • Current read: the repository now has program structure, validation, CI, and collaboration templates; the next lift is faster research navigation.
  • Best next experiment: measure whether a new agent can find prior evidence and propose a non-duplicate experiment faster using scorecards and intake records.
  • Strong anchors: knowledge/research_program_index.md, knowledge/claims/initial_claims.md, docs/quality_gates.md.
  • Avoid repeating: adding process that slows pilots without improving memory or decision quality.
  • Evidence that advances the program: navigation artifacts reduce duplicate proposals and improve citation of prior evidence.

Test-Time Reasoning Budget

  • Program: charter
  • Current read: native thinking is a real coherent-content lever, but its value geometry and write sites are workload/phase-specific. Fixed First: exposes a replicated semantic state, shared/midpoint value and last-token additive branching fail, and early concrete text routes one-operation execution without reaching a valid full-program controller. The answer seam qualifies no-think tokenizer EOS at 48/48 in both prefix cells against 0/48 HF EOS, while thinking is worse at 38/48 structured and 30/48 freeform. Replay-hardened mechanics then passed transport and ABI but produced 0/24 oracle proposal coverage for materialized and every matched comparator. The short no-think seam controls emission; it does not unlock residual composition.
  • Best next experiment: retain the separate 16k+ loop-control line. For J-space, measure whether any within-<think> coordinate predicts eventual correctness across held-out tasks beyond margin/position/equal-width non-J controls, and only then test a same-prefix causal branch intervention against matched sampling.
  • Strong anchors: qwen35_4b_thinking_content_vs_compute, qwen35_4b_overthinking_content_ladder, qwen35_4b_answer_potential_trace_sft, qwen35_4b_native_thought_seam_budget_ladder, qwen35_4b_forced_commit_jacobian_value_transport, qwen35_4b_commit_slot_jacobian_value_transport, qwen35_4b_commit_slot_semantic_power_replication, qwen35_4b_think_ftpo_round2.
  • Avoid repeating: thinking-budget wins without content controls, calibration on a different workload class, cap-bound score interpretation, larger-N harvesting before termination/locality works, or treating high varentropy as a monotone “push harder” signal.
  • Evidence that advances the program: a controller or distillation that Pareto-beats fixed budgets, and a content control that isolates genuine reasoning from compute + scaffold + token-presence.

Agentic Breadth Installation