Research log Small Model Experimentation
GitHub

Qwen3.5-4B Balanced-Core Answer-Potential SFT

Does answer likelihood pick the best reasoning to

The one idea you need

Imagine a student who solves one problem many different ways. To make a study example, do you keep the write-up that best predicts the correct final answer, or just the shortest complete one? This test bakes six such picking rules into six small models.

The question

When choosing which of a model's own worked-out reasonings to train it on, does picking by how well each predicts the right answer beat simpler rules?

What we found

Not run yet — the reasoning bank is fully built (360 tasks, six competing selection rules staged) but no model has been trained or scored. The built test will fine-tune six copies, each fed reasoning chosen a different way, then measure which answers fresh sealed problems most accurately. A key rival is simply keeping the shortest complete reasoning, which may quietly win.

Why it matters

If ranking reasoning by answer likelihood does not beat just keeping the shortest complete thought, teams can skip expensive scoring machinery and use a far cheaper selection rule with no loss.

Reasoning traces banked360 taskscomplete pool ready to train on, none deficient
Ways of picking reasoning tested6each trains its own model copy, including a shortest-thought control
Bar a method must clear3 pointsaccuracy gain over the best simple baseline required to advance
Task families balanced3equal share of the training pool, no fourth added
On this page
  1. Results at a glance
  2. Overview
  3. Report
    1. Summary
    2. Plain-Language Question
    3. Method
    4. Results
    5. Controls
    6. Oracle Versus Deployable Evidence
    7. Interpretation
    8. Next Experiments
    9. Artifact Manifest
  4. Experiment log
  5. Data files
  6. Reproduce
  7. Related

Results at a glance 1

Preserved candidate thoughts versus the balanced-core hard stop Operational progress only, not a capability result. The second bar is fixed at 360 tasks x N=64; pivot branches and additional families are prohibited.

Candidate thoughts · Balanced-core harvest state →

01k2k3k331-task checkpoint331-task checkpoint21.18k360-task hard stop360-task hard stop23.04k
Data table
Balanced-core harvest stateindependent candidate thoughts
331-task checkpoint21.18k
360-task hard stop23.04k

Numbers from experiments/qwen35_4b_balanced_core_answer_potential_sft/README.md and configs/default.yaml

In the author’s words from the Overview · “Results”

The success control is much narrower: only 97 unique R1-successful traces from 58 tasks in four of nine cells exist, so deterministic row matching repeats each source seven or eight times. It remains a useful ordinary rejection-sampling control, but not a task-balanced alternative treatment. The six datasets total 34,446,994 forward tokens at the frozen two epochs. Rescaling the preserved training stress envelope gives roughly 9.6--18.1 GPU-hours for the full matrix before merge/evaluation, which exceeds the user's current time budget. No training was started. A second selection invocation reproduced every dataset and the tracked manifest byte-for-byte; the manifest SHA-256 is 27d4b0b4b1120381a48cb3cd14ddd06f7630a5b8bee9bb43225fb0f7300acfa2.

Overview

Status

Prospective resource-constrained follow-up to qwen35_4b_long_horizon_answer_potential_sft. The original balanced-funnel design was frozen before the remaining 29 harvest tasks, any training-pool scoring, any SFT, or any held-out evaluation. A selector balance defect was then discovered only after all candidate scores existed; its repair is transparently classified as a post-score/partial-rollout, pre-official-selection implementation deviation, not a prospective amendment. Partial R1 success labels were subsequently inspected for cost planning before the deviation was committed; they did not determine the repair. At that boundary no official SFT dataset, adapter, or held-out outcome existed. Selection is now banked behind the committed seal, but no adapter or capability result exists yet.

This experiment is closed as a preserved, compute-stopped negative boundary. The exact six-arm matrix was estimated at 9.6--18.1 GPU-hours before merge and evaluation, beyond the accepted budget, so no SFT was started. Its deterministic selections are not evidence that answer-potential training works and remain reusable only through a separately preregistered follow-up.

The preserved parent is linked here.

This fork preserves the sunk cost of 331 complete, atomic task shards while imposing a hard compute funnel: finish exactly 360 balanced tasks, score only the independent N=64 pool, skip pivot branching, train six discriminating arms, and run a small mandatory evaluation before any optional expansion.

Research Programs

  • Primary: posttraining_and_adaptation.
  • Secondary: evidence_conditioned_selection and test_time_reasoning_budget.
  • Closest near-duplicate: qwen35_4b_long_horizon_answer_potential_sft, whose original nine-family, 95,040-candidate protocol remains frozen and unfinished after calibration and 331/1,080 train tasks.
  • Other anchors: C51 (cap-bound answer potential), C28 (own successful thoughts can be rationalizations), C50 (the answer-emission seam matters), and C24 (banking gains are driven by distinct data rather than repeated exposure).

Question

On a balanced three-family pool of complete, naturally closed Qwen3.5-4B thoughts, does banking traces selected by canonical-answer likelihood produce better fresh behavior than banking:

  1. length-matched random natural thoughts;
  2. R1 answer-success rejection samples;
  3. the two shortest eligible thoughts; or
  4. the same potential-selected thought multiset reassigned to other tasks?

The second treatment asks whether joint likelihood of the close/answer boundary plus the correct answer is better than answer-only likelihood.

Why This Is A Separate Experiment

The parent experiment already exposed calibration results and partial-harvest runtime. Its preregistration cannot honestly be rewritten. This fork is prospectively frozen after those observations and scopes its claim accordingly.

Observed parent evidence that informs, but cannot confirm, this design:

  • 8,640 calibration traces yielded answer-gain AUROC 0.597 and joint-gain AUROC 0.678;
  • top-one answer/joint selections improved R4 answer-rollout success by +6.84/+6.25 points over seeded random;
  • negative length was stronger (AUROC 0.690 and top-one success 26.56%); and
  • the first 331 train tasks required 97,883,041 thought tokens, making the nine-family schedule too slow.

Therefore shortest-natural is a mandatory control, calibration is treated only as design input, and all claims come from new SFT outcomes on sealed evaluation tasks.

Model, Firewall, And Inherited Data

  • Only Qwen/Qwen3.5-4B, revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
  • Fresh procedural tasks copied into this experiment from the firewall-clean parent harness.
  • No content under benchmarks/ is read, imported, or used for training.
  • Canonical answers are training-time curation instruments and never appear in deployment prompts.
  • The 331 inherited task shards remain immutable and are imported by recorded index SHA-256 plus each shard's existing SHA-256 receipt. The remaining 29 use the exact same per-task N=64 vLLM protocol.

The inherited source index is /workspace/large_artifacts/qwen35_4b_long_horizon_answer_potential_sft/pools/train_independent/index.json at SHA-256 da09176ddf05918712913b4c66ca893f47ed141986f8ed37ef289a63dc37fb63. It contains 331 tasks, 21,184 traces, 97,883,041 sampled thought tokens, 20,917 natural closes, and four exact periodic loops.

Balanced Core

The core is the first complete family blocks in the parent's pre-existing train order, not families chosen by score or correctness:

familylevelstasks per leveltaskscandidates
caravan1--3401207,680
foundry_ledger1--3401207,680
runeward1--3401207,680
total36023,040

The only remaining generation is 29 runeward level-3 tasks. The hard stop is 360 tasks: no fourth family, no pivot branches, and no adaptive enlargement based on outcomes.

Natural-Thought Protocol

Independent thoughts use temperature 1.0, top-p 0.95, top-k 20, and a 12,288-token natural-close allowance. A non-loop allowance contact receives one exact-prefix continuation of at most 2,048 tokens. Only traces that emit their own </think>, are not exact periodic loops, and fit the 16,000-token training record are eligible. No incomplete thought is force-closed into SFT.

If a task has fewer than two eligible traces after N=64, it receives up to four deterministic N=16 top-up batches. A task still deficient is excluded symmetrically from all trace arms.

Fast Canonical Scoring

For task prompt x, complete thought z, boundary b = </think>\n\nANSWER:, and canonical answer y*:

answer_gain(z) = log p(y* | x, z, b) - log p(y* | x, empty, b)

joint_gain(z)  = log p(b, y* | x, z) - log p(b, y* | x, empty)

The initial candidate instrument used the experiment-local vLLM exact targeted-token readout. It teacher-forces the observed prefix, reads raw target-token log probabilities, never constrains the sampled token, and bypasses only the unused vocabulary-rank reduction. Before bulk scoring, 32 fixed inherited calibration traces must match the Transformers bf16 full-prefix scorer within 0.15 mean nats/token for answer likelihood, joint likelihood, both empty baselines, and both gains. Any parity failure blocks scoring.

Only the canonical boundary is scored at train scale. The parent's calibration already measured canonical versus one-newline rank stability (task-macro Kendall tau-b 0.841), so recomputing the second format would spend substantial prefix work without changing the frozen selector.

The broadened 32-row gate failed before bulk scoring (maximum 0.692447 versus 0.15), exposing known batch-sensitive long-prefix logits. The threshold was not relaxed. Per the dated preregistration amendment, vLLM is retired for train likelihoods and every canonical answer/joint score is now computed by the single-context Transformers bf16 reference uniformly. Generative comparisons remain vLLM-only.

Selection And SFT Arms

Selection is within task and keeps two full traces per arm:

armtrace targetpurpose
random_naturaleligible trace nearest each answer-treatment lengthlong-thought/style control
success_rftR1-successful trace nearest each answer-treatment lengthordinary rejection sampling
shortest_naturaltwo shortest eligible tracesstrongest observed calibration control
answer_potentialanswer-gain quality first, structural diversity secondoriginal treatment
joint_potentialjoint-gain quality first, structural diversity secondclose/commit treatment
potential_shuffleanswer-treatment multiset reassigned within family/level/lengthtask-specific content control

Potential selection retains the top 12 by score, takes the best, then the structurally most distant trace within 0.25 nats per answer token. If that band contains no unused second trace, it deterministically uses the second-ranked member of the same frozen top 12 so every balanced task still contributes two rows. It never rewards brevity. The shortest arm is intentionally not token-matched; its lower token dose is part of the mechanism being tested and is reported.

The pre-selection audit found that this fallback is rare for answer potential (5/360 tasks) but common for joint potential (244/360 tasks). The joint arm is therefore explicitly interpreted as a best-plus-diverse-or- second-ranked hybrid, and results must be stratified by selection mode and score gap rather than described as a uniformly near-best-diverse treatment. This disclosure narrows the selector claim; a commit-and-evidence seal prevents any further selector change after the completed score bank was inspected.

All arms otherwise use identical QLoRA settings: rank 32, alpha 64, dropout 0.05, two epochs, learning rate 2e-4, batch 1 x gradient accumulation 16, maximum length 16,000, prompt loss 0, thought loss 0.5, and boundary/answer loss 1.0. Rows and optimizer exposure are matched; success_rft is deterministically oversampled only when it has fewer eligible rows. Adapters and merged checkpoints remain outside git.

Staged Evaluation

All evaluation is natural-thinking vLLM with a 12,288-token allowance. Merged checkpoints must first produce a real same-prompt behavioral difference from base.

Mandatory Stage A evaluates base plus all six arms greedily on sealed subsets fixed by task metadata:

splitconstructiontasks
core IID3 train families x L1--L3 x 20180
core hard3 train families x L4 x 2060
held familybrinework/spindle x L1--L3 x first 1060

Primary metric: core-IID exact-answer accuracy. Report paired 10,000-resample task bootstraps, parse rate, natural-close rate, family macro, thought lengths, and actual forward tokens.

For each potential treatment, the strongest trace baseline is the highest-accuracy member of random_natural, success_rft, and shortest_natural. Stage B triggers only if the treatment:

  • beats that baseline by at least 0.03 core-IID accuracy with paired 95% lower bound above zero;
  • beats potential_shuffle pointwise in the aggregate;
  • loses no more than 0.02 parse rate or family macro; and
  • has a mathematically reachable registered gate.

If no treatment triggers, the experiment stops after Stage A. If one triggers, Stage B runs all-arm greedy evaluation on the full inherited IID/hard/held/rendering splits, k=8 only for base, the winning treatment, its strongest baseline, shortest-natural, and shuffle, and training-seed-43 replication for the treatment and strongest baseline. A mission-level positive additionally requires the trained method to beat base sample-more at matched actual forward tokens.

Verdicts

  • CORE_BANKING_NEGATIVE: neither potential arm clears the Stage-A trigger.
  • POTENTIAL_BANKING_POSITIVE: a potential arm clears Stage A and the full Stage-B comparison while preserving interface metrics.
  • REPLICATED_BANKING_POSITIVE: seed 43 has the same sign and the pooled paired interval excludes zero.
  • MISSION_POSITIVE: replicated positive plus a matched-compute win over base sample-more.
  • SHORTEST_BANKING_LEADS: shortest-natural is the strongest trace arm; this supports a compression or optimization mechanism, not answer-potential selection.

The three-family result cannot support a nine-family claim. A null is a core-scope null, and any broader follow-up must be a new experiment.

Run

.venv-vllm/bin/python experiments/qwen35_4b_balanced_core_answer_potential_sft/scripts/run.py --stage smoke
.venv-vllm/bin/python experiments/qwen35_4b_balanced_core_answer_potential_sft/scripts/run.py --stage full

The granular path is import -> harvest -> parity -> score -> rollouts -> evidence-seal -> select -> train -> merge -> deployment-probe -> evaluate-stage-a -> analyze-stage-a, followed only conditionally by Stage B. evidence-seal is a one-time retrospective attestation for the legacy indexes; select remains blocked until that seal and the post-score deviation are committed in the machine amendment receipt.

full is resume-only across these commit boundaries: it intentionally stops if the evidence/amendment seal is not committed, and a fresh selection is not trainable until the byte-identical tracked SFT manifest and selection summary are committed and pushed. This prevents a one-process run from selecting and immediately training on an unreviewed dataset. The current execution stops after selection in any case pending the user's compute choice.

Results

No capability result yet. The balanced bank is complete: 360 tasks, 23,040 traces, 108,759,239 sampled thought tokens, 22,681 natural closes, four loops, and zero deficient tasks. The candidate vLLM scoring instrument failed its strict cross-backend gate before bulk scoring; the reference-scoring amendment above was frozen before any training score or outcome. Exact single-context reference scoring then completed for all 22,681 eligible traces in 17,296 seconds, and R1 completed one answer rollout for every scored trace in 10,915 seconds. All 360 raw/score/R1 shards, hashes, task scopes, source links, trace joins, and eligibility sets passed the read-only pre-seal audit. The retrospective evidence seal is now committed-bound: its pre-attestation hashes, post-seal index hashes, operation contracts, and post-score deviation disclosure are recorded in machine-readable receipts.

Official selection then produced exactly 720 rows for each of six arms. Answer, joint, shuffle, random, and shortest each cover all 360 tasks; selected thoughts retain their natural lengths, reaching 14,325 tokens. The success control is much narrower: only 97 unique R1-successful traces from 58 tasks in four of nine cells exist, so deterministic row matching repeats each source seven or eight times. It remains a useful ordinary rejection-sampling control, but not a task-balanced alternative treatment.

The six datasets total 34,446,994 forward tokens at the frozen two epochs. Rescaling the preserved training stress envelope gives roughly 9.6--18.1 GPU-hours for the full matrix before merge/evaluation, which exceeds the user's current time budget. No training was started. A second selection invocation reproduced every dataset and the tracked manifest byte-for-byte; the manifest SHA-256 is 27d4b0b4b1120381a48cb3cd14ddd06f7630a5b8bee9bb43225fb0f7300acfa2.

Artifacts

  • idea_intake.md: routing, novelty, and post-calibration boundary
  • reports/preregistration.md: original frozen protocol plus dated amendments/deviations
  • reports/design_review.md: adversarial review and applied fixes
  • configs/default.yaml: exact counts, seeds, and gates
  • reports/artifact_manifest.yaml: inherited pool, external scores, adapters, and checkpoints
  • runs/preselection_amendment_receipt.json: commit-bound code, evidence, and deviation boundary
  • runs/preselection_evidence_seal.json: exact pre/post index identity and absence checks at seal time
  • data/sft_manifest.json: exact six-arm selected-dataset hashes, counts, costs, and provenance
  • runs/selection_summary.json: byte-identical tracked copy of the official selection manifest
  • external root: /workspace/large_artifacts/qwen35_4b_balanced_core_answer_potential_sft

Report

Rendered from reports/report.md

Summary

Experiment finished at its selection-only stop. The balanced raw pool is complete and the failed candidate scoring instrument was replaced prospectively by its single-context reference. A later selector-balance repair is explicitly a post-score, pre-official-selection deviation and is machine-sealed before selection; no SFT ran and no trained capability result was observed.

Plain-Language Question

If we sample many complete ways the model thinks through a problem, can the correct answer's likelihood tell us which reasoning to teach back—or is simply choosing the shortest complete reasoning better?

Method

Finish a checksum-preserved 360-task, three-family N=64 bank; compare answer-potential, joint-potential, random, successful, shortest, and task-shuffled full-thought SFT; evaluate every arm on fresh core, harder, and family-held tasks before any optional expansion.

Results

Operational results: 360/360 tasks, 23,040 traces, 108,759,239 sampled thought tokens, 22,681 natural closes, four loops, and no top-ups. The task-diverse joint HF/vLLM gate failed at 0.692447 > 0.15 before any bulk score. The frozen threshold was preserved; vLLM likelihood scoring was retired in favor of the single-context Transformers reference. Exact scoring completed for all 22,681 eligible traces in 17,296 seconds, followed by one R1 answer rollout per trace in 10,915 seconds. A read-only whole-bank validation passes exact scope, artifact, source-link, trace-join, and eligibility-set checks. The resulting retrospective seal binds the original and final index hashes, exact operation contracts, frozen code/data, and deviation disclosure. Official selection has now run; capability results remain pending because SFT has not.

Exact reference scoring subsequently completed for all 360 tasks and 22,681 eligible traces. Applying the original helper in memory exposed an unintended 116-task filter; because those scores were already observed, the balance fallback is an exploratory post-score deviation. Partial R1 labels were subsequently inspected for cost planning before commit but did not determine the fallback; no official SFT row, adapter, or held-out outcome informed it.

Official selection contains 720 rows per arm and is byte-deterministic on rerun. Five arms cover all 360 tasks; the success-RFT control has 97 unique successful source traces from 58 tasks and repeats each source seven or eight times to reach matched optimizer exposure. Selected potential thoughts are not length-capped at 512: answer/joint maxima are 14,240/14,325 tokens. The frozen two-epoch matrix is 34,446,994 forward tokens, with an estimated 9.6--18.1 GPU-hour training envelope before merge/evaluation. No adapter, merge, deployment probe, or evaluation artifact exists.

Controls

Six exact selected datasets are banked. Random-natural, shortest-natural, success-RFT, and task-shuffled potential remain the controls; their training has not started.

Oracle Versus Deployable Evidence

Reference answers curate training traces only. The primary result will be autonomous natural-thinking exact accuracy; answer likelihood is not itself a deployment metric.

Interpretation

Selection alone is not a capability result. The full frozen matrix is too slow for the current budget, and the success control's narrow task support must be considered when choosing a smaller prospective fork. This experiment is closed; any such fork must receive a new experiment directory and prospective boundary.

Next Experiments

A lower-cost prospective fork may be created after an explicit compute/design choice. The frozen full-matrix claim is unchanged; no subset result may be relabeled as its confirmatory verdict.

Artifact Manifest

See artifact_manifest.yaml.

Experiment log 10

Show the running log (10 entries, 2026-07-12 → 14)

2026-07-12 — Resource-Constrained Fork

The parent long-horizon experiment was paused after 331/1,080 train tasks because observed throughput made the remaining nine-family schedule incompatible with the user's time budget. No saved shard was lost: 21,184 traces and 97,883,041 sampled thought tokens are present behind atomic SHA-256 receipts.

The user selected the balanced-core funnel. Repository lifecycle rules require a new experiment because calibration results are already visible. This fork declares those observations, selects the already-leading three complete family blocks, adds shortest-natural as the strongest calibration control, replaces slow HF train scoring with a broader parity-gated exact vLLM readout, drops pivot branches, and makes full evaluation conditional.

At this boundary the GPU is idle. No remaining harvest task, training score, R1 train rollout, SFT update, or evaluation generation has run under this experiment.

2026-07-12 — Immutable Design Anchor

  • Prospective README, preregistration, adversarial review, full restartable harness, frozen data, and 40 passing CPU tests were committed at original c847615f before any experiment GPU call.
  • The configured guard now points to that commit and its three exact file digests. It fails before model load if ancestry or content identity changes.
  • After the anchor, the README's first relative link was moved below the generated summary paragraph so the repository catalog resolves it from the correct directory. This is a navigation-only repair; the guard still verifies that byte-exact prospective design.

2026-07-12 — Concurrent-Main Rebase

  • Rebasing over three concurrent site commits changed the design anchor from original c847615f to cb3d64e3. All three frozen-file SHA-256 values remained identical.
  • The configured ancestry pointer was re-anchored to the rebased commit without changing any design text, threshold, split, arm, or code. No experiment GPU call had run.

2026-07-12 — Balanced Harvest And Scorer Instrument Stop

  • Imported all 331 parent shards at the frozen index digest, then completed exactly 29 Runeward-L3 tasks. Final pool: 360 tasks, 23,040 traces, 108,759,239 sampled thought tokens, 22,681 natural closes, four exact loops, finite priors on every trace, and no task requiring a top-up.
  • The registered task-diverse 32-row joint parity gate then failed closed at 0.692447 > 0.15. No training score, R1 train rollout, selection, adapter update, or evaluation had run.
  • Inspection of the instrument receipt showed the known long-prefix batch-sensitivity boundary: answer gain max 0.147865, joint-likelihood mean-token max 0.054477, empty-answer max 0.156281, and parent-normalized joint-gain max 0.692447. No threshold or row was changed.
  • Added the pre-outcome amendment to retire vLLM bulk likelihoods and use the single-context Transformers bf16 reference uniformly. The failed receipt remains evidence; all generation and later evaluation stay on vLLM.

2026-07-13 — Post-Score, Pre-Official-Selection Balance Deviation

  • Exact scoring completed for all 360 tasks and 22,681 natural traces. Before selection or training, a read-only preflight found that the near-best diversity helper could return one row and silently remove the entire task from every arm.
  • On the frozen scores, unchanged behavior would have retained only 116 tasks, distributed 23 Caravan, 71 Foundry Ledger, and 22 Runeward. This violates the declared balanced-core estimand.
  • Because the complete candidate score bank and the induced imbalance were observed before this repair was committed, it violates the preregistration's amendment-timing rule. It is a post-score deviation, not a prospective amendment; no later seal can restore that status. Partial incomplete-R1 labels were later inspected for cost planning before commit, but did not determine the repair. Official selections, adapters, and held-out outcomes remained unseen.
  • Repaired the contradiction by keeping near-best diversity when available and otherwise taking the deterministic second-ranked trace from the same frozen top-12. Added fail-closed assertions for 360 total tasks, 40 per family/level cell, and 720 rows per arm. No selection artifact or adapter existed.
  • The same audit found that Trainer seeding occurred after LoRA construction. Added global pre-model seeding, an immediate pre-adapter seed reset, receipt provenance, and CPU ordering tests before any training run.
  • Hardened training, merge, deployment-probe, and evaluation restart contracts before any adapter existed: exact two-epoch exposure/steps, initial LoRA digest, training-lock/final-artifact hashes, complete 128-pair merge enforcement, deployed-file fingerprints, and prompt/sampling/runner-bound generation caches.
  • Replaced the shuffle control's cyclic heuristic with exact minimum-cost forbidden-edge assignment and made Stage-A baseline ties conservative and explicit.
  • On the frozen score bank, the balance fallback applies to 5/360 answer tasks and 244/360 joint tasks. Answer fallback gaps have median 3.409 and maximum 7.378 nats/answer-token; joint fallback gaps have median 0.893, mean 1.410, p90 2.954, and maximum 12.128. The joint arm is consequently a registered hybrid treatment; its result will be stratified by mode/gap.
  • Shuffle rows now namespace both target and actual-source selection/quality provenance, and their unprefixed audit fields describe the trace actually trained. Training receipts are invalidated before artifact replacement, exclude themselves from artifact hashes, and are installed atomically for safe restart.

2026-07-13 — Training Cost Re-estimate

  • In-memory selection over all 360 completed score shards (without writing official selection artifacts) gives 32,187,564 two-epoch forward tokens across the five rollout-independent arms. Current-path stress receipts imply 9.0--16.9 GPU-hours for those five; allowing the unfinished success arm gives a provisional six-arm range of roughly 9.3--20.7 GPU-hours.
  • This materially exceeds the user's stated time constraint. Finish and bank the already-running R1 rollout and exact selection, but do not start SFT until choosing between the frozen full matrix and a smaller prospective follow-up. No licensed shortcut exists inside this frozen experiment before mandatory Stage A.

2026-07-13 — Exact Evidence Bank Complete

  • Single-context bf16 SDPA scoring finished 360/360 tasks and 22,681 eligible rows in 17,296 seconds.
  • R1 finished 360/360 tasks and exactly one rollout for each of the same 22,681 trace IDs in 10,915 seconds.
  • Before any seal write, the full read-only validator confirmed all three exact task scopes, every shard hash, raw-to-score/R1 source links, per-shard task identity, unique trace IDs, exact score/R1 joins, and the natural-close/non-loop eligibility set. The raw pool has 23,040 rows; score and R1 each have 22,681.
  • Selection remains absent and blocked pending the committed post-score/partial-rollout deviation seal.

2026-07-13 — Pre-Selection Evidence Boundary Sealed

  • The seal first repeated the complete read-only raw/score/R1 validation, then added only retrospective operation-contract attestations to the three legacy indexes. It records that these contracts were not emitted by the original generation/scoring processes.
  • Pre-attestation index SHA-256 values are 6aeae76f...24d3, c0fed08b...db8, and 9a9ab75b...71f for raw, exact scores, and R1. Final sealed index SHA-256 values are f635d060...18fd, e2b0a402...5740, and b116eea5...eea0 respectively; exact full values live in the tracked machine receipts and artifact manifest.
  • The amendment receipt binds rebased commit 0a6ccf6d68c79bf80705f48a3de58ad06a0a57ec, every transitive selection/training dependency, all procedural task data, the three final evidence indexes, and the explicit post-score/partial-rollout deviation disclosure. It also proves that official selection and adapter artifacts were absent at seal time.
  • A second seal invocation was byte-idempotent across all three external indexes and all four tracked seal receipts. Official selection remains absent until this boundary is committed and pushed.

2026-07-13 — Official Selection Banked, Training Paused

  • After the evidence boundary was pushed and both GitHub workflows passed, official selection retained all 360 tasks at exactly 40 per family/level cell and wrote 720 rows for every registered arm. The five task-conditioned arms each cover 360 tasks with 720 distinct source traces.
  • Answer potential used 355 near-best-diverse and five fallback-second selections in addition to 360 best rows. Joint potential used 116 near-best-diverse and 244 fallback-second selections in addition to 360 best rows. The official order-statistic joint fallback p90 is 3.075; the earlier 2.954 planning figure used an interpolated percentile. Median, mean, and maximum remain 0.893, 1.410, and 12.128.
  • The success-RFT control has only 97 unique successful source traces from 58 tasks, all in four cells: Caravan L1/L2 and Foundry Ledger L1/L3. Deterministic oversampling repeats 56 sources seven times and 41 sources eight times. This is a narrow-support rejection-sampling control, not balanced task coverage.
  • Full-thought lengths remain uncapped by the retired 512-token pilot limit. Answer/joint selected thoughts reach 14,240/14,325 tokens; their medians are 3,968/4,420 tokens.
  • Exact frozen two-epoch cost is 34,446,994 forward tokens: 7,129,440 answer, 7,620,122 joint, 7,129,440 shuffle, 6,999,860 random, 3,308,702 shortest, and 2,259,430 success. The existing stress envelope implies about 9.6--18.1 GPU-hours before merge/evaluation, so no SFT was started.
  • A second official selection invocation left all six deterministic gzip SHA-256 values, the manifest, the selection summary, and the design receipt byte-identical. Both tracked manifest copies have SHA-256 27d4b0b4b1120381a48cb3cd14ddd06f7630a5b8bee9bb43225fb0f7300acfa2.

2026-07-14 — administrative closure at the selection boundary

  • The experiment is finished as compute-stopped, not successful. The frozen six-arm matrix remains estimated at 9.6--18.1 GPU-hours before merge/evaluation; no SFT, adapter, merge, deployment probe, or capability evaluation exists, and selection alone supports no capability claim.
  • The deterministic bank is preserved, but any lower-cost continuation or restart must be a new preregistered experiment. Leaving this boundary in-progress incorrectly implied an active run.

Data files 5

Result tables and metrics copied from the experiment folder — preview inline or open the raw file.

Reproduce

Smoke test

.venv-vllm/bin/python experiments/qwen35_4b_balanced_core_answer_potential_sft/scripts/run.py --stage smoke

Full run

.venv-vllm/bin/python experiments/qwen35_4b_balanced_core_answer_potential_sft/scripts/run.py --stage full

Run steps are documented inside the experiment folder (README and scripts).

Browse the experiment folder on GitHub ↗