Goal-Gate Confirmation
The one idea you need
One sealed benchmark event just recorded the program's stated goal for the first time: a clean-trained model beat the untouched base on every one of ten families at once. But two of those family wins rode on single-item margins, and this program's own laws say one seed is a draw until it replicates. This experiment is the replication: the same two models, three brand-new sealed seeds, and a verdict written before the event.
The question
Does the all-ten-families win replicate across three fresh sealed seeds?
What we found
The replication returned a split answer. The overall improvement replicated without drama: the trained model beat the base decisively on all three fresh seeds, making four for four all-time. The perfect ten-family sweep repeated on one of the three — two full sweeps across four independent seeds — but the pre-written rule demanded two of three, so the formal verdict is aggregate-only. The near-misses are the striking part: nine-of-ten and eight-of-ten with ZERO losses, blocked purely by zero-zero ties at the debugging-style family that has resisted every teaching method. The goal now hangs on exactly one family, and the queued larger-dose experiment knows precisely what it must produce: any reliable nonzero score there. [Erratum 2026-07-16: the 'two of four seeds' framing omitted the earlier 78,150 reading; over all six recorded readings the sweep rate is 2/6 — see the sweep-rate consolidation cell.]
Why it matters
This is the program goal's final gate: a confirmed pass completes the demonstration-plus-confirmation chain the goal demands; a non-replication prices the discovery honestly and hands the program back its 9-of-10 position with the newly proven skill-to-family conversion mechanism.
On this page
Results at a glance 1
How to read
One bar per seed: how many of the ten families the trained model strictly won.
Takeaway → Ten, nine, and eight family wins with zero losses anywhere — the misses are all zero-zero ties at one stubborn family.
Data table
| sealed medium seed (tb 1024) | families strictly won | families lost |
|---|---|---|
| 78154 (discovery) | 10 | 0 |
| 78155 | 9 | 0 |
| 78156 | 8 | 0 |
| 78157 | 10 | 0 |
Numbers from experiments/qwen35_4b_goal_gate_confirmation/runs/benchmark/confirmation_readout.json
Technical framing
Strict family wins vs base per sealed seed (of 10) — AGGREGATE_ONLY under the frozen ordered verdict: aggregate strict wins on all three confirmation seeds (0.3287/0.3737/0.3837 vs base 0.0586/0.1122/0.0982; 4/4 all-time with the discovery), goal gate 1/3 vs the required 2/3. Every non-passing seed had ZERO losses — blocked by menders 0-margin ties (both) and one warren tie (warren WON +0.267 elsewhere). Two full 10/10 sweeps across four independent sealed seeds; the goal is localized to exactly one family. First standalone-compliant cell: full six-stage lineage package (copied datasets, fixed-seed manifest, vendored trainers/merger/root) with the readout provenance-anchored end-to-end.
In the author’s words from the Overview · “Results”
All six runs clean (both arms authenticated per invocation, within budget, per-seed ledger opened/closed, readout provenance-anchored) (table on the experiment page). Verdict per the frozen ordered partition: AGGREGATE_ONLY — aggregate strict wins 3/3 (plus the discovery, 4/4 all-time), goal-gate passes 1/3 against the required 2/3.
Overview
The mandatory replication of the program's recorded 10/10: base versus the hygiene_explore composite on three independent fresh sealed medium seeds, with an ordered confirmation verdict and the discovery seed reported but never counted.
Research Program
- Program:
agentic_breadth_installation. - Program question: can synthetic curricula install general capability that lifts the held-out aggregate without a negative family?
- Prior anchors: the statechain dose's sealed event (hygiene_explore_parent goal_gate_pass TRUE at 78,154 — aggregate 0.3663 vs 0.0800, all ten families strictly above, menders 0.017 and warren 0.050 — 0.150 vs 0.100 — on single-item margins); the confirmation law; the tier forensics (menders was draw-dependent, never a wall).
Question
Does the all-families pass replicate? CONFIRMED requires the aggregate to win on all three fresh seeds and the 10/10 gate on at least two.
Setup
- Arms:
base(b654e033…/26d8ee48…) andhygiene_explore(9eb653d7…/e2112344…), full trees recomputed once per runner invocation, before any gateway call (wording matches the code: authentication is per-invocation, not per-seed; closed seeds' receipts stay sha-pinned in the ledger). - Event: three sealed seeds 78,155/78,156/78,157, medium, tb 1,024, per-seed write-ahead ledger whose closed records sha-pin the sealed summary AND both per-arm gateway receipts (the readout refuses anything the ledger did not pin), implementation signature anchored to the discovery event.
- Recovery:
--resumeis the single recovery path. A crashed (opened) seed reuses its preserved receipts; a crash between a seed's summary write and its closed-record append is recovered by deterministic byte-identical regeneration of the summary (divergence refuses loudly with both digests); an unopened seed refuses to run over pre-existing event files (clean slate). - Verdict: CONFIRMED / AGGREGATE_ONLY / NOT_REPLICATED (ordered, total); fragility margins reported per seed; no promotion logic anywhere.
- Standalone lineage (owner directive 2026-07-15): the hygiene_explore composite's complete reproduction package lives in this cell —
data/lineage/(six ordered SFT dataset copies + fixed-seed recipe manifest),scripts/lineage_trainers/+scripts/merge_adapter.py(byte-identical trainer/merger copies), the frozen root adapter vendored atlarge_artifacts/qwen35_4b_goal_gate_confirmation/lineage_root/blend(hard provenance boundary: no committed creation receipt), andscripts/rebuild_lineage.py(GPU replay + sha verification;--verify-inputsruns in smoke).
Run
PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_goal_gate_confirmation/scripts/run.py --smoke
.venv/bin/python -B experiments/qwen35_4b_goal_gate_confirmation/scripts/run.py --stage benchmarkResults
All six runs clean (both arms authenticated per invocation, within budget, per-seed ledger opened/closed, readout provenance-anchored):
| seed | base | hygiene_explore | goal gate | blockers |
|---|---|---|---|---|
| 78,155 | 0.0586 | 0.3287 | 9/10 | menders tie (0.0 margin); warren WON +0.267 |
| 78,156 | 0.1122 | 0.3737 | 8/10 | menders + warren ties; zero losses |
| 78,157 | 0.0982 | 0.3837 | PASS 10/10 | — |
| (78,154 discovery, not counted) | 0.0800 | 0.3663 | PASS 10/10 | — |
Verdict per the frozen ordered partition: AGGREGATE_ONLY — aggregate strict wins 3/3 (plus the discovery, 4/4 all-time), goal-gate passes 1/3 against the required 2/3.
Interpretation
The replication sharpens rather than overturns the discovery. The aggregate transfer is unconditional — never close on any seed. The all-families sweep is real but draw-gated at exactly one family: every non-passing seed carried ZERO losses, and menders blocked with a 0.0 margin (both arms at zero) on both failing seeds while warren, the other discovery-day fragility, WON by +0.267 on one seed and tied once. Two full 10/10 sweeps in four independent sealed seeds is a model within one item-draw of the goal on every roll — and a preregistered bar honestly not met. The program position: the goal's primary condition is demonstrated but not confirmed at the frozen majority bar; menders is the single binding family, and the queued dose-scale intake (the one mechanism class the kill rules permit) is now aimed at a precisely-known target: any reliable nonzero menders yield completes the gate.
Knowledgebase Update
- Program evidence updated: the AGGREGATE_ONLY verdict, the second sweep, and the menders 0-margin localization recorded.
- Program backlog updated: the menders dose-scale intake is the funded successor; the zero-root rebuild stays queued.
- Claim ledger updated: no confirmed claim; the demonstrated-not-confirmed position stated exactly.
Artifacts
data/design_receipt.json: seeds/tier/budget/models/gateway/discovery pins + the standalone lineage-package pins.data/lineage/: the six stage datasets andlineage_manifest.json(the complete fixed-seed recipe; produced shas are verification aids).reports/preregistration.md,reports/benchmark_design_review.md: contract and authorization.
Erratum (2026-07-16)
The sweep-rate framing in this document ("two full sweeps across four independent sealed seeds", ~50%) reflects the 78,154–78,157 window and omits the earlier 78,150 reading (8/10, menders+rites ties). Over ALL six recorded goal-gate readings the rate is 2/6 (exact 95% CI [0.04, 0.78]), with menders blocking every miss. See experiments/qwen35_4b_sweep_rate_consolidation for the consolidated record; the per-seed facts in this document are unchanged.
Report
Rendered from reports/report.md
Summary
The three-seed replication closed AGGREGATE_ONLY. The aggregate transfer is unconditional — hygiene_explore strictly beat base on all three fresh sealed seeds (0.3287/0.3737/0.3837 vs 0.0586/0.1122/0.0982; with the discovery, 4/4 all-time, never close). The all-families sweep replicated once: seed 78,157 passed 10/10, making two full sweeps across four independent sealed seeds; the frozen 2/3 majority bar failed because 78,155 (9/10) and 78,156 (8/10) were blocked entirely by ties — menders at a 0.0 margin on both, warren once (while warren WON +0.267 on the other) — with zero strict losses anywhere. The verdict localizes the goal to a single family: any reliable nonzero menders yield completes the gate. Every reading is provenance-anchored (receipt shas pinned in closed ledger records; the readout refuses any break in the sealed chain).
Research Program Fit
The program goal demands demonstration and confirmation; the confirmation law demands independent seeds. This cell is both, at the gateway's highest supported tier.
Method
See the preregistration.
Results
runs/benchmark/confirmation_readout.json: verdict AGGREGATE_ONLY; per-seed goal gates 9/10, 8/10, 10/10 (PASS); aggregate strict wins 3/3; fragility — menders margin 0.0 on both failing seeds, warren +0.267/tie/win; discovery seed reported (0.3663 vs 0.0800, 10/10) and never counted.
Controls
Both arms published and tree-authenticated once per runner invocation, before any gateway call; per-seed write-ahead ledger whose closed records sha-pin the sealed summary and both per-arm receipts (the readout reads verdict inputs only through those pins); implementation-signature equality across all six runs and against the pinned discovery summary; fragility margins preregistered.
Oracle Versus Deployable Evidence
Gateway aggregates and public family scores only; benchmarks/ never read.
Next Stage
Closed. The menders dose-scale intake (the one permitted mechanism class) is the funded successor with a precisely-known target; the zero-root lineage rebuild stays queued.
Artifact Manifest
Two composite pins external with committed receipts; everything else in-repo.
Erratum (2026-07-16)
The sweep-rate framing in this document ("two full sweeps across four independent sealed seeds", ~50%) reflects the 78,154–78,157 window and omits the earlier 78,150 reading (8/10, menders+rites ties). Over ALL six recorded goal-gate readings the rate is 2/6 (exact 95% CI [0.04, 0.78]), with menders blocking every miss. See experiments/qwen35_4b_sweep_rate_consolidation for the consolidated record; the per-seed facts in this document are unchanged.
Experiment log 5
Show the running log (5 entries, 2026-07-15)
2026-07-15 — Model-free design freeze
- Opened as the mandatory successor to the recorded 10/10 at seed 78,154: three independent sealed medium seeds (78,155/78,156/78,157), two authenticated arms (base, hygiene_explore), K-seed write-ahead ledger, ordered confirmation verdict (CONFIRMED / AGGREGATE_ONLY / NOT_REPLICATED) with the discovery seed reported but never counted, per-seed fragility readings on the menders/warren margins, implementation-signature equality anchored to the discovery event.
- No model event has run; nothing trains in this cell.
2026-07-15 — Standalone-reproducibility retrofit (owner directive)
- Retrofitted the new standalone gate (AGENTS.md / docs/quality_gates.md, commits 02284dc3/1b4248a0) before the freeze: the hygiene_explore composite's complete model-reproduction package now lives in this cell.
data/lineage/carries byte-identical copies of the six ordered SFT datasets (stage 3 preserved asstage03_close_xi__targeted_standard.jsonl— the source filename does not match the arm name) pluslineage_manifest.json(fixed seeds 42/43/44/47/51/55, full hyperparameters incl. stage 3's targeted close overrides and stages 1-2's missing close channel, per-stage produced shas as verification aids, final merge onto the raw HF base).scripts/lineage_trainers/carries the three trainer variants (400e4b85… / 10b4914c… / 0cfb126f…) andscripts/merge_adapter.py(cb9af8b4…) byte-identically. The frozenblendroot adapter (weights ad2ef4fa…, no committed creation receipt — a hard provenance boundary, documented as such) is vendored into this cell's ownlarge_artifacts/…/lineage_root/blend(~181 MiB, six files hash-pinned).scripts/rebuild_lineage.pyreplays stages 1→6 plus the merge with per-stage sha verification; its--verify-inputsmode (no GPU) is wired intorun.py --smoke, and the design receipt now pins the whole package (regenerated; sha changed accordingly). The rebuilt merge is verified onmodel.safetensors(e2112344…) and the content files — the published tree sha additionally covers a merge receipt embedding a machine-local absolute adapter path, recorded honestly in the manifest. - Still no model event; the GPU rebuild path has not been executed.
2026-07-15 — Standalone retrofit (owner directive) before freeze
- The owner's standalone-reproducibility directive landed mid-build and this cell is the first to comply: six lineage datasets copied in byte-identically with a fixed-seed manifest, three trainer variants and the merger vendored into scripts/, rebuild_lineage.py with a no-GPU verify-inputs mode wired into smoke, and the C53-era root adapter vendored into this cell's own artifact storage with its provenance boundary (no committed creation receipt) stated plainly in the preregistration.
- Lineage fact established by the receipts walk: every stage and every merge uses the raw official Qwen3.5-4B revision as base; the lineage is carried entirely by one warm-started LoRA adapter; the root predates committed receipts. 123 tests green; smoke green; receipt regenerated (3864e812…) and --check byte-identical twice.
2026-07-15 — Adversarial review: the verdict chain hardened pre-freeze
- Three lenses; the lineage package audited clean against the source receipts (zero findings). Two MAJORs confirmed in the K-seed machinery and fixed: the verdict inputs are now provenance-anchored end-to-end (receipt shas pinned into closed ledger records; the readout refuses any break in the sealed chain; smoke verifies receipt hashes), and the summary-write/closed-append crash window closes via byte-equal reconciliation with a documented recovery. Four minors fixed including the honest 216-combination verdict-partition enumeration.
- 146 tests green; smoke green; receipt 66c19b24… --check twice; PASS_BENCHMARK_EVENT granted.
2026-07-15 — The three-seed event and closure
- CI green on the freeze; the six runs executed in the frozen seed-major order; every closed ledger record carries both receipt pins; the readout verified the full provenance chain before rendering.
- Verdict AGGREGATE_ONLY: aggregate strict wins on all three seeds (0.3287/0.3737/0.3837 vs 0.0586/0.1122/0.0982); goal gate 1/3 (78,157 swept 10/10; 78,155 read 9/10 blocked by a menders 0-margin tie with warren WON at +0.267; 78,156 read 8/10 blocked by menders and warren ties; zero losses anywhere).
- Position: two full sweeps across four independent sealed seeds; demonstrated, not confirmed at the frozen 2/3 bar; menders is the single binding family (0.0 margin on every failing seed). The dose-scale intake aims at a precisely-known target.
Data files 3
Result tables and metrics copied from the experiment folder — preview inline or open the raw file.
runs/benchmark/medium_tb1024_seed78155_confirmation/summary.json2.3 kBruns/benchmark/medium_tb1024_seed78156_confirmation/summary.json2.3 kBruns/benchmark/medium_tb1024_seed78157_confirmation/summary.json2.3 kB
Reproduce
Smoke test
PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_goal_gate_confirmation/scripts/run.py --smokeFull run
checkpointed scripts/run.py --stage benchmark onlyRun steps are documented inside the experiment folder (README and scripts).