Qwen3.5-4B Balanced-Core Answer-Potential SFT
The one idea you need
Imagine a student who solves one problem many different ways. To make a study example, do you keep the write-up that best predicts the correct final answer, or just the shortest complete one? This test bakes six such picking rules into six small models.
The question
When choosing which of a model's own worked-out reasonings to train it on, does picking by how well each predicts the right answer beat simpler rules?
What we found
Not run yet — the reasoning bank is fully built (360 tasks, six competing selection rules staged) but no model has been trained or scored. The built test will fine-tune six copies, each fed reasoning chosen a different way, then measure which answers fresh sealed problems most accurately. A key rival is simply keeping the shortest complete reasoning, which may quietly win.
Why it matters
If ranking reasoning by answer likelihood does not beat just keeping the shortest complete thought, teams can skip expensive scoring machinery and use a far cheaper selection rule with no loss.
On this page
Results at a glance 1
Candidate thoughts · Balanced-core harvest state →
Data table
| Balanced-core harvest state | independent candidate thoughts |
|---|---|
| 331-task checkpoint | 21.18k |
| 360-task hard stop | 23.04k |
Numbers from experiments/qwen35_4b_balanced_core_answer_potential_sft/README.md and configs/default.yaml
In the author’s words from the Overview · “Results”
The success control is much narrower: only 97 unique R1-successful traces from 58 tasks in four of nine cells exist, so deterministic row matching repeats each source seven or eight times. It remains a useful ordinary rejection-sampling control, but not a task-balanced alternative treatment. The six datasets total 34,446,994 forward tokens at the frozen two epochs. Rescaling the preserved training stress envelope gives roughly 9.6--18.1 GPU-hours for the full matrix before merge/evaluation, which exceeds the user's current time budget. No training was started. A second selection invocation reproduced every dataset and the tracked manifest byte-for-byte; the manifest SHA-256 is 27d4b0b4b1120381a48cb3cd14ddd06f7630a5b8bee9bb43225fb0f7300acfa2.
Overview
Status
Prospective resource-constrained follow-up to qwen35_4b_long_horizon_answer_potential_sft. The original balanced-funnel design was frozen before the remaining 29 harvest tasks, any training-pool scoring, any SFT, or any held-out evaluation. A selector balance defect was then discovered only after all candidate scores existed; its repair is transparently classified as a post-score/partial-rollout, pre-official-selection implementation deviation, not a prospective amendment. Partial R1 success labels were subsequently inspected for cost planning before the deviation was committed; they did not determine the repair. At that boundary no official SFT dataset, adapter, or held-out outcome existed. Selection is now banked behind the committed seal, but no adapter or capability result exists yet.
This experiment is closed as a preserved, compute-stopped negative boundary. The exact six-arm matrix was estimated at 9.6--18.1 GPU-hours before merge and evaluation, beyond the accepted budget, so no SFT was started. Its deterministic selections are not evidence that answer-potential training works and remain reusable only through a separately preregistered follow-up.
The preserved parent is linked here.
This fork preserves the sunk cost of 331 complete, atomic task shards while imposing a hard compute funnel: finish exactly 360 balanced tasks, score only the independent N=64 pool, skip pivot branching, train six discriminating arms, and run a small mandatory evaluation before any optional expansion.
Research Programs
- Primary:
posttraining_and_adaptation. - Secondary:
evidence_conditioned_selectionandtest_time_reasoning_budget. - Closest near-duplicate:
qwen35_4b_long_horizon_answer_potential_sft, whose original nine-family, 95,040-candidate protocol remains frozen and unfinished after calibration and 331/1,080 train tasks. - Other anchors: C51 (cap-bound answer potential), C28 (own successful thoughts can be rationalizations), C50 (the answer-emission seam matters), and C24 (banking gains are driven by distinct data rather than repeated exposure).
Question
On a balanced three-family pool of complete, naturally closed Qwen3.5-4B thoughts, does banking traces selected by canonical-answer likelihood produce better fresh behavior than banking:
- length-matched random natural thoughts;
- R1 answer-success rejection samples;
- the two shortest eligible thoughts; or
- the same potential-selected thought multiset reassigned to other tasks?
The second treatment asks whether joint likelihood of the close/answer boundary plus the correct answer is better than answer-only likelihood.
Why This Is A Separate Experiment
The parent experiment already exposed calibration results and partial-harvest runtime. Its preregistration cannot honestly be rewritten. This fork is prospectively frozen after those observations and scopes its claim accordingly.
Observed parent evidence that informs, but cannot confirm, this design:
- 8,640 calibration traces yielded answer-gain AUROC 0.597 and joint-gain AUROC 0.678;
- top-one answer/joint selections improved R4 answer-rollout success by +6.84/+6.25 points over seeded random;
- negative length was stronger (AUROC 0.690 and top-one success 26.56%); and
- the first 331 train tasks required 97,883,041 thought tokens, making the nine-family schedule too slow.
Therefore shortest-natural is a mandatory control, calibration is treated only as design input, and all claims come from new SFT outcomes on sealed evaluation tasks.
Model, Firewall, And Inherited Data
- Only
Qwen/Qwen3.5-4B, revision851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a. - Fresh procedural tasks copied into this experiment from the firewall-clean parent harness.
- No content under
benchmarks/is read, imported, or used for training. - Canonical answers are training-time curation instruments and never appear in deployment prompts.
- The 331 inherited task shards remain immutable and are imported by recorded index SHA-256 plus each shard's existing SHA-256 receipt. The remaining 29 use the exact same per-task N=64 vLLM protocol.
The inherited source index is /workspace/large_artifacts/qwen35_4b_long_horizon_answer_potential_sft/pools/train_independent/index.json at SHA-256 da09176ddf05918712913b4c66ca893f47ed141986f8ed37ef289a63dc37fb63. It contains 331 tasks, 21,184 traces, 97,883,041 sampled thought tokens, 20,917 natural closes, and four exact periodic loops.
Balanced Core
The core is the first complete family blocks in the parent's pre-existing train order, not families chosen by score or correctness:
| family | levels | tasks per level | tasks | candidates |
|---|---|---|---|---|
caravan | 1--3 | 40 | 120 | 7,680 |
foundry_ledger | 1--3 | 40 | 120 | 7,680 |
runeward | 1--3 | 40 | 120 | 7,680 |
| total | 360 | 23,040 |
The only remaining generation is 29 runeward level-3 tasks. The hard stop is 360 tasks: no fourth family, no pivot branches, and no adaptive enlargement based on outcomes.
Natural-Thought Protocol
Independent thoughts use temperature 1.0, top-p 0.95, top-k 20, and a 12,288-token natural-close allowance. A non-loop allowance contact receives one exact-prefix continuation of at most 2,048 tokens. Only traces that emit their own </think>, are not exact periodic loops, and fit the 16,000-token training record are eligible. No incomplete thought is force-closed into SFT.
If a task has fewer than two eligible traces after N=64, it receives up to four deterministic N=16 top-up batches. A task still deficient is excluded symmetrically from all trace arms.
Fast Canonical Scoring
For task prompt x, complete thought z, boundary b = </think>\n\nANSWER:, and canonical answer y*:
answer_gain(z) = log p(y* | x, z, b) - log p(y* | x, empty, b)
joint_gain(z) = log p(b, y* | x, z) - log p(b, y* | x, empty)The initial candidate instrument used the experiment-local vLLM exact targeted-token readout. It teacher-forces the observed prefix, reads raw target-token log probabilities, never constrains the sampled token, and bypasses only the unused vocabulary-rank reduction. Before bulk scoring, 32 fixed inherited calibration traces must match the Transformers bf16 full-prefix scorer within 0.15 mean nats/token for answer likelihood, joint likelihood, both empty baselines, and both gains. Any parity failure blocks scoring.
Only the canonical boundary is scored at train scale. The parent's calibration already measured canonical versus one-newline rank stability (task-macro Kendall tau-b 0.841), so recomputing the second format would spend substantial prefix work without changing the frozen selector.
The broadened 32-row gate failed before bulk scoring (maximum 0.692447 versus 0.15), exposing known batch-sensitive long-prefix logits. The threshold was not relaxed. Per the dated preregistration amendment, vLLM is retired for train likelihoods and every canonical answer/joint score is now computed by the single-context Transformers bf16 reference uniformly. Generative comparisons remain vLLM-only.
Selection And SFT Arms
Selection is within task and keeps two full traces per arm:
| arm | trace target | purpose |
|---|---|---|
random_natural | eligible trace nearest each answer-treatment length | long-thought/style control |
success_rft | R1-successful trace nearest each answer-treatment length | ordinary rejection sampling |
shortest_natural | two shortest eligible traces | strongest observed calibration control |
answer_potential | answer-gain quality first, structural diversity second | original treatment |
joint_potential | joint-gain quality first, structural diversity second | close/commit treatment |
potential_shuffle | answer-treatment multiset reassigned within family/level/length | task-specific content control |
Potential selection retains the top 12 by score, takes the best, then the structurally most distant trace within 0.25 nats per answer token. If that band contains no unused second trace, it deterministically uses the second-ranked member of the same frozen top 12 so every balanced task still contributes two rows. It never rewards brevity. The shortest arm is intentionally not token-matched; its lower token dose is part of the mechanism being tested and is reported.
The pre-selection audit found that this fallback is rare for answer potential (5/360 tasks) but common for joint potential (244/360 tasks). The joint arm is therefore explicitly interpreted as a best-plus-diverse-or- second-ranked hybrid, and results must be stratified by selection mode and score gap rather than described as a uniformly near-best-diverse treatment. This disclosure narrows the selector claim; a commit-and-evidence seal prevents any further selector change after the completed score bank was inspected.
All arms otherwise use identical QLoRA settings: rank 32, alpha 64, dropout 0.05, two epochs, learning rate 2e-4, batch 1 x gradient accumulation 16, maximum length 16,000, prompt loss 0, thought loss 0.5, and boundary/answer loss 1.0. Rows and optimizer exposure are matched; success_rft is deterministically oversampled only when it has fewer eligible rows. Adapters and merged checkpoints remain outside git.
Staged Evaluation
All evaluation is natural-thinking vLLM with a 12,288-token allowance. Merged checkpoints must first produce a real same-prompt behavioral difference from base.
Mandatory Stage A evaluates base plus all six arms greedily on sealed subsets fixed by task metadata:
| split | construction | tasks |
|---|---|---|
| core IID | 3 train families x L1--L3 x 20 | 180 |
| core hard | 3 train families x L4 x 20 | 60 |
| held family | brinework/spindle x L1--L3 x first 10 | 60 |
Primary metric: core-IID exact-answer accuracy. Report paired 10,000-resample task bootstraps, parse rate, natural-close rate, family macro, thought lengths, and actual forward tokens.
For each potential treatment, the strongest trace baseline is the highest-accuracy member of random_natural, success_rft, and shortest_natural. Stage B triggers only if the treatment:
- beats that baseline by at least 0.03 core-IID accuracy with paired 95% lower bound above zero;
- beats
potential_shufflepointwise in the aggregate; - loses no more than 0.02 parse rate or family macro; and
- has a mathematically reachable registered gate.
If no treatment triggers, the experiment stops after Stage A. If one triggers, Stage B runs all-arm greedy evaluation on the full inherited IID/hard/held/rendering splits, k=8 only for base, the winning treatment, its strongest baseline, shortest-natural, and shuffle, and training-seed-43 replication for the treatment and strongest baseline. A mission-level positive additionally requires the trained method to beat base sample-more at matched actual forward tokens.
Verdicts
CORE_BANKING_NEGATIVE: neither potential arm clears the Stage-A trigger.POTENTIAL_BANKING_POSITIVE: a potential arm clears Stage A and the full Stage-B comparison while preserving interface metrics.REPLICATED_BANKING_POSITIVE: seed 43 has the same sign and the pooled paired interval excludes zero.MISSION_POSITIVE: replicated positive plus a matched-compute win over base sample-more.SHORTEST_BANKING_LEADS: shortest-natural is the strongest trace arm; this supports a compression or optimization mechanism, not answer-potential selection.
The three-family result cannot support a nine-family claim. A null is a core-scope null, and any broader follow-up must be a new experiment.
Run
.venv-vllm/bin/python experiments/qwen35_4b_balanced_core_answer_potential_sft/scripts/run.py --stage smoke
.venv-vllm/bin/python experiments/qwen35_4b_balanced_core_answer_potential_sft/scripts/run.py --stage fullThe granular path is import -> harvest -> parity -> score -> rollouts -> evidence-seal -> select -> train -> merge -> deployment-probe -> evaluate-stage-a -> analyze-stage-a, followed only conditionally by Stage B. evidence-seal is a one-time retrospective attestation for the legacy indexes; select remains blocked until that seal and the post-score deviation are committed in the machine amendment receipt.
full is resume-only across these commit boundaries: it intentionally stops if the evidence/amendment seal is not committed, and a fresh selection is not trainable until the byte-identical tracked SFT manifest and selection summary are committed and pushed. This prevents a one-process run from selecting and immediately training on an unreviewed dataset. The current execution stops after selection in any case pending the user's compute choice.
Results
No capability result yet. The balanced bank is complete: 360 tasks, 23,040 traces, 108,759,239 sampled thought tokens, 22,681 natural closes, four loops, and zero deficient tasks. The candidate vLLM scoring instrument failed its strict cross-backend gate before bulk scoring; the reference-scoring amendment above was frozen before any training score or outcome. Exact single-context reference scoring then completed for all 22,681 eligible traces in 17,296 seconds, and R1 completed one answer rollout for every scored trace in 10,915 seconds. All 360 raw/score/R1 shards, hashes, task scopes, source links, trace joins, and eligibility sets passed the read-only pre-seal audit. The retrospective evidence seal is now committed-bound: its pre-attestation hashes, post-seal index hashes, operation contracts, and post-score deviation disclosure are recorded in machine-readable receipts.
Official selection then produced exactly 720 rows for each of six arms. Answer, joint, shuffle, random, and shortest each cover all 360 tasks; selected thoughts retain their natural lengths, reaching 14,325 tokens. The success control is much narrower: only 97 unique R1-successful traces from 58 tasks in four of nine cells exist, so deterministic row matching repeats each source seven or eight times. It remains a useful ordinary rejection-sampling control, but not a task-balanced alternative treatment.
The six datasets total 34,446,994 forward tokens at the frozen two epochs. Rescaling the preserved training stress envelope gives roughly 9.6--18.1 GPU-hours for the full matrix before merge/evaluation, which exceeds the user's current time budget. No training was started. A second selection invocation reproduced every dataset and the tracked manifest byte-for-byte; the manifest SHA-256 is 27d4b0b4b1120381a48cb3cd14ddd06f7630a5b8bee9bb43225fb0f7300acfa2.
Artifacts
idea_intake.md: routing, novelty, and post-calibration boundaryreports/preregistration.md: original frozen protocol plus dated amendments/deviationsreports/design_review.md: adversarial review and applied fixesconfigs/default.yaml: exact counts, seeds, and gatesreports/artifact_manifest.yaml: inherited pool, external scores, adapters, and checkpointsruns/preselection_amendment_receipt.json: commit-bound code, evidence, and deviation boundaryruns/preselection_evidence_seal.json: exact pre/post index identity and absence checks at seal timedata/sft_manifest.json: exact six-arm selected-dataset hashes, counts, costs, and provenanceruns/selection_summary.json: byte-identical tracked copy of the official selection manifest- external root:
/workspace/large_artifacts/qwen35_4b_balanced_core_answer_potential_sft
Report
Rendered from reports/report.md
Summary
Experiment finished at its selection-only stop. The balanced raw pool is complete and the failed candidate scoring instrument was replaced prospectively by its single-context reference. A later selector-balance repair is explicitly a post-score, pre-official-selection deviation and is machine-sealed before selection; no SFT ran and no trained capability result was observed.
Plain-Language Question
If we sample many complete ways the model thinks through a problem, can the correct answer's likelihood tell us which reasoning to teach back—or is simply choosing the shortest complete reasoning better?
Method
Finish a checksum-preserved 360-task, three-family N=64 bank; compare answer-potential, joint-potential, random, successful, shortest, and task-shuffled full-thought SFT; evaluate every arm on fresh core, harder, and family-held tasks before any optional expansion.
Results
Operational results: 360/360 tasks, 23,040 traces, 108,759,239 sampled thought tokens, 22,681 natural closes, four loops, and no top-ups. The task-diverse joint HF/vLLM gate failed at 0.692447 > 0.15 before any bulk score. The frozen threshold was preserved; vLLM likelihood scoring was retired in favor of the single-context Transformers reference. Exact scoring completed for all 22,681 eligible traces in 17,296 seconds, followed by one R1 answer rollout per trace in 10,915 seconds. A read-only whole-bank validation passes exact scope, artifact, source-link, trace-join, and eligibility-set checks. The resulting retrospective seal binds the original and final index hashes, exact operation contracts, frozen code/data, and deviation disclosure. Official selection has now run; capability results remain pending because SFT has not.
Exact reference scoring subsequently completed for all 360 tasks and 22,681 eligible traces. Applying the original helper in memory exposed an unintended 116-task filter; because those scores were already observed, the balance fallback is an exploratory post-score deviation. Partial R1 labels were subsequently inspected for cost planning before commit but did not determine the fallback; no official SFT row, adapter, or held-out outcome informed it.
Official selection contains 720 rows per arm and is byte-deterministic on rerun. Five arms cover all 360 tasks; the success-RFT control has 97 unique successful source traces from 58 tasks and repeats each source seven or eight times to reach matched optimizer exposure. Selected potential thoughts are not length-capped at 512: answer/joint maxima are 14,240/14,325 tokens. The frozen two-epoch matrix is 34,446,994 forward tokens, with an estimated 9.6--18.1 GPU-hour training envelope before merge/evaluation. No adapter, merge, deployment probe, or evaluation artifact exists.
Controls
Six exact selected datasets are banked. Random-natural, shortest-natural, success-RFT, and task-shuffled potential remain the controls; their training has not started.
Oracle Versus Deployable Evidence
Reference answers curate training traces only. The primary result will be autonomous natural-thinking exact accuracy; answer likelihood is not itself a deployment metric.
Interpretation
Selection alone is not a capability result. The full frozen matrix is too slow for the current budget, and the success control's narrow task support must be considered when choosing a smaller prospective fork. This experiment is closed; any such fork must receive a new experiment directory and prospective boundary.
Next Experiments
A lower-cost prospective fork may be created after an explicit compute/design choice. The frozen full-matrix claim is unchanged; no subset result may be relabeled as its confirmatory verdict.
Artifact Manifest
See artifact_manifest.yaml.
Experiment log 10
Show the running log (10 entries, 2026-07-12 → 14)
2026-07-12 — Resource-Constrained Fork
The parent long-horizon experiment was paused after 331/1,080 train tasks because observed throughput made the remaining nine-family schedule incompatible with the user's time budget. No saved shard was lost: 21,184 traces and 97,883,041 sampled thought tokens are present behind atomic SHA-256 receipts.
The user selected the balanced-core funnel. Repository lifecycle rules require a new experiment because calibration results are already visible. This fork declares those observations, selects the already-leading three complete family blocks, adds shortest-natural as the strongest calibration control, replaces slow HF train scoring with a broader parity-gated exact vLLM readout, drops pivot branches, and makes full evaluation conditional.
At this boundary the GPU is idle. No remaining harvest task, training score, R1 train rollout, SFT update, or evaluation generation has run under this experiment.
2026-07-12 — Immutable Design Anchor
- Prospective README, preregistration, adversarial review, full restartable harness, frozen data, and 40 passing CPU tests were committed at original
c847615fbefore any experiment GPU call. - The configured guard now points to that commit and its three exact file digests. It fails before model load if ancestry or content identity changes.
- After the anchor, the README's first relative link was moved below the generated summary paragraph so the repository catalog resolves it from the correct directory. This is a navigation-only repair; the guard still verifies that byte-exact prospective design.
2026-07-12 — Concurrent-Main Rebase
- Rebasing over three concurrent site commits changed the design anchor from original
c847615ftocb3d64e3. All three frozen-file SHA-256 values remained identical. - The configured ancestry pointer was re-anchored to the rebased commit without changing any design text, threshold, split, arm, or code. No experiment GPU call had run.
2026-07-12 — Balanced Harvest And Scorer Instrument Stop
- Imported all 331 parent shards at the frozen index digest, then completed exactly 29 Runeward-L3 tasks. Final pool: 360 tasks, 23,040 traces, 108,759,239 sampled thought tokens, 22,681 natural closes, four exact loops, finite priors on every trace, and no task requiring a top-up.
- The registered task-diverse 32-row joint parity gate then failed closed at 0.692447 > 0.15. No training score, R1 train rollout, selection, adapter update, or evaluation had run.
- Inspection of the instrument receipt showed the known long-prefix batch-sensitivity boundary: answer gain max 0.147865, joint-likelihood mean-token max 0.054477, empty-answer max 0.156281, and parent-normalized joint-gain max 0.692447. No threshold or row was changed.
- Added the pre-outcome amendment to retire vLLM bulk likelihoods and use the single-context Transformers bf16 reference uniformly. The failed receipt remains evidence; all generation and later evaluation stay on vLLM.
2026-07-13 — Post-Score, Pre-Official-Selection Balance Deviation
- Exact scoring completed for all 360 tasks and 22,681 natural traces. Before selection or training, a read-only preflight found that the near-best diversity helper could return one row and silently remove the entire task from every arm.
- On the frozen scores, unchanged behavior would have retained only 116 tasks, distributed 23 Caravan, 71 Foundry Ledger, and 22 Runeward. This violates the declared balanced-core estimand.
- Because the complete candidate score bank and the induced imbalance were observed before this repair was committed, it violates the preregistration's amendment-timing rule. It is a post-score deviation, not a prospective amendment; no later seal can restore that status. Partial incomplete-R1 labels were later inspected for cost planning before commit, but did not determine the repair. Official selections, adapters, and held-out outcomes remained unseen.
- Repaired the contradiction by keeping near-best diversity when available and otherwise taking the deterministic second-ranked trace from the same frozen top-12. Added fail-closed assertions for 360 total tasks, 40 per family/level cell, and 720 rows per arm. No selection artifact or adapter existed.
- The same audit found that Trainer seeding occurred after LoRA construction. Added global pre-model seeding, an immediate pre-adapter seed reset, receipt provenance, and CPU ordering tests before any training run.
- Hardened training, merge, deployment-probe, and evaluation restart contracts before any adapter existed: exact two-epoch exposure/steps, initial LoRA digest, training-lock/final-artifact hashes, complete 128-pair merge enforcement, deployed-file fingerprints, and prompt/sampling/runner-bound generation caches.
- Replaced the shuffle control's cyclic heuristic with exact minimum-cost forbidden-edge assignment and made Stage-A baseline ties conservative and explicit.
- On the frozen score bank, the balance fallback applies to 5/360 answer tasks and 244/360 joint tasks. Answer fallback gaps have median 3.409 and maximum 7.378 nats/answer-token; joint fallback gaps have median 0.893, mean 1.410, p90 2.954, and maximum 12.128. The joint arm is consequently a registered hybrid treatment; its result will be stratified by mode/gap.
- Shuffle rows now namespace both target and actual-source selection/quality provenance, and their unprefixed audit fields describe the trace actually trained. Training receipts are invalidated before artifact replacement, exclude themselves from artifact hashes, and are installed atomically for safe restart.
2026-07-13 — Training Cost Re-estimate
- In-memory selection over all 360 completed score shards (without writing official selection artifacts) gives 32,187,564 two-epoch forward tokens across the five rollout-independent arms. Current-path stress receipts imply 9.0--16.9 GPU-hours for those five; allowing the unfinished success arm gives a provisional six-arm range of roughly 9.3--20.7 GPU-hours.
- This materially exceeds the user's stated time constraint. Finish and bank the already-running R1 rollout and exact selection, but do not start SFT until choosing between the frozen full matrix and a smaller prospective follow-up. No licensed shortcut exists inside this frozen experiment before mandatory Stage A.
2026-07-13 — Exact Evidence Bank Complete
- Single-context bf16 SDPA scoring finished 360/360 tasks and 22,681 eligible rows in 17,296 seconds.
- R1 finished 360/360 tasks and exactly one rollout for each of the same 22,681 trace IDs in 10,915 seconds.
- Before any seal write, the full read-only validator confirmed all three exact task scopes, every shard hash, raw-to-score/R1 source links, per-shard task identity, unique trace IDs, exact score/R1 joins, and the natural-close/non-loop eligibility set. The raw pool has 23,040 rows; score and R1 each have 22,681.
- Selection remains absent and blocked pending the committed post-score/partial-rollout deviation seal.
2026-07-13 — Pre-Selection Evidence Boundary Sealed
- The seal first repeated the complete read-only raw/score/R1 validation, then added only retrospective operation-contract attestations to the three legacy indexes. It records that these contracts were not emitted by the original generation/scoring processes.
- Pre-attestation index SHA-256 values are
6aeae76f...24d3,c0fed08b...db8, and9a9ab75b...71ffor raw, exact scores, and R1. Final sealed index SHA-256 values aref635d060...18fd,e2b0a402...5740, andb116eea5...eea0respectively; exact full values live in the tracked machine receipts and artifact manifest. - The amendment receipt binds rebased commit
0a6ccf6d68c79bf80705f48a3de58ad06a0a57ec, every transitive selection/training dependency, all procedural task data, the three final evidence indexes, and the explicit post-score/partial-rollout deviation disclosure. It also proves that official selection and adapter artifacts were absent at seal time. - A second seal invocation was byte-idempotent across all three external indexes and all four tracked seal receipts. Official selection remains absent until this boundary is committed and pushed.
2026-07-13 — Official Selection Banked, Training Paused
- After the evidence boundary was pushed and both GitHub workflows passed, official selection retained all 360 tasks at exactly 40 per family/level cell and wrote 720 rows for every registered arm. The five task-conditioned arms each cover 360 tasks with 720 distinct source traces.
- Answer potential used 355 near-best-diverse and five fallback-second selections in addition to 360 best rows. Joint potential used 116 near-best-diverse and 244 fallback-second selections in addition to 360 best rows. The official order-statistic joint fallback p90 is 3.075; the earlier 2.954 planning figure used an interpolated percentile. Median, mean, and maximum remain 0.893, 1.410, and 12.128.
- The success-RFT control has only 97 unique successful source traces from 58 tasks, all in four cells: Caravan L1/L2 and Foundry Ledger L1/L3. Deterministic oversampling repeats 56 sources seven times and 41 sources eight times. This is a narrow-support rejection-sampling control, not balanced task coverage.
- Full-thought lengths remain uncapped by the retired 512-token pilot limit. Answer/joint selected thoughts reach 14,240/14,325 tokens; their medians are 3,968/4,420 tokens.
- Exact frozen two-epoch cost is 34,446,994 forward tokens: 7,129,440 answer, 7,620,122 joint, 7,129,440 shuffle, 6,999,860 random, 3,308,702 shortest, and 2,259,430 success. The existing stress envelope implies about 9.6--18.1 GPU-hours before merge/evaluation, so no SFT was started.
- A second official selection invocation left all six deterministic gzip SHA-256 values, the manifest, the selection summary, and the design receipt byte-identical. Both tracked manifest copies have SHA-256
27d4b0b4b1120381a48cb3cd14ddd06f7630a5b8bee9bb43225fb0f7300acfa2.
2026-07-14 — administrative closure at the selection boundary
- The experiment is finished as compute-stopped, not successful. The frozen six-arm matrix remains estimated at 9.6--18.1 GPU-hours before merge/evaluation; no SFT, adapter, merge, deployment probe, or capability evaluation exists, and selection alone supports no capability claim.
- The deterministic bank is preserved, but any lower-cost continuation or restart must be a new preregistered experiment. Leaving this boundary
in-progressincorrectly implied an active run.
Data files 5
Result tables and metrics copied from the experiment folder — preview inline or open the raw file.
runs/scorer_parity_joint_32.json30 kBruns/selection_summary.json9.5 kBruns/train_independent_scores_summary.json312 Bruns/train_independent_summary.json397 Bruns/train_rollouts_r1_summary.json299 B
Reproduce
Smoke test
.venv-vllm/bin/python experiments/qwen35_4b_balanced_core_answer_potential_sft/scripts/run.py --stage smokeFull run
.venv-vllm/bin/python experiments/qwen35_4b_balanced_core_answer_potential_sft/scripts/run.py --stage fullRun steps are documented inside the experiment folder (README and scripts).