Count-Walk Menders Confirmation
The one idea you need
The previous experiment produced a first: on one sealed exam, its trained model solved a fix-the-broken-procedure puzzle (the 'menders' family) that all three of its comparison models scored exactly zero on. That pattern - trained model above zero, every comparison at zero - was the pre-declared positive outcome, and it had never happened before in this program. But there is a catch the team wrote down immediately: exams are drawn fresh each time, and roughly one untrained model in ten also lands a single menders solve on any given draw. One good draw can be luck. The only honest way to tell luck from skill is to run the same head-to-head again on fresh sealed exams and count.
The question
If the same four models sit four fresh sealed exams, does the trained model keep solving menders puzzles while the three comparison models keep scoring zero - or does the one win from last time dissolve into ordinary exam-to-exam noise?
What we found
The answer the rule was built to force out, delivered without wiggle room. Across the four fresh exams the trained model solved a fix-the-procedure episode exactly once (plus one partial credit that the rules pre-declared doesn't count) — and on that same exam, the comparison model trained WITHOUT the special lessons solved one too. One hit when two were required, and a dead tie against a control, is the pre-written middle verdict: no claim. The clean interpretation is that occasional single-episode solves are background luck this family hands out to roughly one run in ten — the pre-registered noise rate — and the earlier headline result (the trained model scoring while all three controls sat at zero) was most likely that luck landing photogenically. The rule also pre-committed the consequence: no more exam seeds for this comparison; any future attempt at this family must be a genuinely different design, not a re-roll. One quietly encouraging descriptive note: the trained model posted the best overall score on two of the four exams (0.398 and 0.392, its two best readings ever), though those readings carry no claim.
Why it matters
The program's one unbroken habit is pricing good news honestly: every favorable draw must survive fresh sealed exams before anyone may claim it. Menders is the single exam family that has blocked the program's headline goal on every previous attempt, so a confirmed movement there would be the first durable capability gain on the hardest surface - and a clean failure to replicate is nearly as valuable, because it closes the line and frees the budget for a different mechanism.
On this page
Results at a glance 1
menders score · sealed event seed →
Data table
| sealed event seed | count_walk menders | replay_ctl7 menders | zero_root_parent menders | base menders |
|---|---|---|---|---|
| 78164 | 0.1 | 0.1 | 0 | 0 |
| 78165 | 0 | 0 | 0 | 0 |
| 78166 | 0 | 0 | 0 | 0 |
| 78167 | 0.0167 | 0 | 0 | 0 |
Numbers from experiments/qwen35_4b_count_walk_menders_confirmation/runs/benchmark/confirmation_readout.json
In the author’s words from the Report · “Summary”
hits_c = 1 (rule required ≥ 2) and episode totals tie replay 1-1 (strict dominance required) → the preregistered middle verdict with its frozen claim text: no claim; further spending on this contrast requires a mechanism-differentiated NEW design, not more seeds of the same. The honest reading: the untreated control's episode is the ~10% background noise process the preregistration priced, and lifecycle 27's MECHANISM_ANSWER (78163) is now best read as that coincidence landing photogenically. Menders remains without a confirmed mover. … Read the full result →
Overview
The mandatory multi-seed confirmation of lifecycle 27's MECHANISM_ANSWER: at sealed seed 78,163 the count_walk composite drew menders 0.1 while base, zero_root_parent, and replay_ctl7 all drew exactly 0.0 — the preregistered positive branch, first in program history. Eval-only: four fresh sealed medium seeds, four authenticated pre-existing arms per seed, one frozen integer-exact replication rule, no training anywhere.
Research Program
- Program:
agentic_breadth_installation - Program question: can DESIGNED synthetic curricula install universal, transferable agentic skills into the one 4B — provably, on fully documented lineage?
- Prior anchors: lifecycle 27 (
qwen35_4b_count_dont_walk_enumeration— the prior event: count_walk menders 0.1 vs all controls 0.0 at seed 78,163, aggregates 0.3312 / 0.3298 / 0.2950 / base 0.0753), lifecycle 26 (qwen35_4b_enumerative_repair_protocol— its replay control drew menders 0.1 at seed 78,162, proving untreated arms can draw), and the goal-gate confirmation law (a favorable draw is priced by fresh sealed seeds, never by re-reading).
Question
Does the seed-78163 menders pattern — candidate above zero while every control sits at exactly zero — replicate across four fresh sealed seeds, or does it close as seed noise? Single-episode menders draws by untreated arms have already happened once in nine recorded medium events, so under the observed noise rate one MECHANISM_ANSWER draw has non-trivial probability; only replication can separate a real rate difference from a favorable roll.
Hypothesis
If the count-dont-walk dose genuinely moved menders, the candidate hits episodes at a per-event rate materially above the program's 0.10 arm-event noise rate and accumulates more episodes than every control over four events. Honest prior: the mechanism the dose TAUGHT was already refuted locally (the candidate still thinks to the 1,024-token cap; fidelity 7/40 vs the 0.50 bar), so this cell tests the capability movement itself, mechanism-agnostic, and a NOT_REPLICATED close is the likelier branch under the noise model.
Setup
- Model: Qwen/Qwen3.5-4B (revision
851bf6e8…), always; nothing trains or merges. - Arms (four pre-existing committed composites, full tree+weights sha256 recomputed and matched fail-closed at event time against design-time constants — no TODO pins):
base(tree26d8ee48…, weightsb654e033…),zero_root_parent(tree414f5829…, weights6e9aad25…, lifecycle 22's committed lineage merge receipte906caea…),replay_ctl7(tree044a4599…, weightsc5035b4d…, committed receipt3f65b4c6…),count_walk(treed5fdc55c…, weightsddd7bc4b…, committed receipt840edca0…). - Event: four fresh sealed seeds 78,164 / 78,165 / 78,166 / 78,167 (grep-fresh in seed contexts at design time; audit in the preregistration), tier medium, think budget 1,024, four arms per seed in frozen order, seed-major; per-seed write-ahead opened/closed ledger, closed records sha-pin the summary and all four gateway receipts;
--resumeis the single recovery path with byte-identical deterministic summary regeneration; the implementation signature is checked LIVE before each seed's first gateway call (pre-consumption — a drifted suite refuses before any GPU run) and all sixteen receipts must equal the prior event's pinned block. FROZEN REPLICATION RULE (integer-exact, over the four NEW events only; 78,163 is prior evidence, never pooled; review amendment A1+A2): an event counts as a hit only if it contains at least one FULL menders episode (score contributes int(10*s + 1e-9) episodes, FLOOR semantics — partial-credit draws on the k/60 lattice are recorded but never counted, as hits or episodes); hits_c = new events with candidate FULL-EPISODE count > 0; E per arm = sum of int(10*score + 1e-9) episodes. REPLICATED iff hits_c >= 2 AND E_c strictly exceeds EVERY control's total; NOT_REPLICATED iff hits_c == 0; AMBIGUOUS otherwise. No fourth state. The three frozen consequences:
- REPLICATED — "the count_walk composite solves menders episodes at a rate no
control matches; the first confirmed menders capability movement in the program." 2. NOT_REPLICATED — "the 78163 reading closes as seed noise; the count-dont-walk dose did not durably move menders; the expression-cost law stands; the composite remains a documented artifact (at a true per-event hit rate of 0.3 this outcome retains probability ≈ 0.24 — the closure is a preregistered funding decision, not a nonexistence proof)." 3. AMBIGUOUS — "no claim; further spending on this contrast requires a mechanism-differentiated NEW design, not more seeds of the same."
- Honest priors (arithmetic in the preregistration; recomputed by
scripts/power_analysis.py --checkin smoke and tests, which enforce every printed number): under the FULL-EPISODE null (design-time audit over all 9 recorded medium/tb1024 sealed events, 29 arm-events, 3 full-episode draws; the 2 partial draws are rule-invisible, recorded-only) the false-REPLICATED probability is 0.0450 at the headline p = 0.10 and 0.0475 at the exact p = 3/29 (exact fraction printed by the script); the counterfactual ceiling — if every raw-positive draw were promoted to a full episode, which the frozen conversion forbids — is 0.0947 at p = 5/29. Power of hits_c >= 2 is 0.5248 / 0.6875 / 0.8735 at candidate hit rates 0.4 / 0.5 / 0.65, and full REPLICATED power (with dominance) is 0.4717 / 0.6289 / 0.8230 (both unchanged by the amendment); NOT_REPLICATED retains probability 0.2401 at q = 0.3. - Standalone boundary (review amendment B1): this cell produces no model but EVALUATES three non-base composites, so it carries the complete in-cell reproduction package per
docs/quality_gates.md(including eval-only cells):data/lineage/(six ordered zero-root stage datasets + the extendedlineage_manifest.json+ lifecycle 22's provenance receipts), the stage-7 production inputs (data/count_walk.jsonl,data/replay_ctl7.jsonl,data/sft_count_walk.jsonl,data/sft_blend.jsonl,data/stream_token_receipt.json), the byte-identical production scripts (scripts/lineage_trainers/,scripts/train_think.py,scripts/merge_adapter.py,scripts/train_trial.py,scripts/merge_trained_arm.py,scripts/rebuild_clean_chain.py), andscripts/rebuild_lineage.py(stages 1-6 rebuild the zero-root parent; stage 7 trains both arms at fixed seed 85 and merges;--verify-inputsruns in smoke and tests). The four committed provenance documents remain copied byte-identically intodata/provenance/as verification aids; the measurement gateway stays shared perdocs/quality_gates.md. - Hidden-label boundary: the benchmark suite's contents are never parsed or read as data; only the trusted aggregate gateway (sha
53cf6533…) runs, and the pre-consumption implementation check hashes suite bytes exclusively through that gateway's own inventory functions.
Run
Smoke (no GPU, no writes):
PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_count_walk_menders_confirmation/scripts/run.py --smokeFull (the only stage; requires the committed PASS_BENCHMARK_EVENT review and clean pushed main):
.venv/bin/python -B experiments/qwen35_4b_count_walk_menders_confirmation/scripts/run.py --stage benchmarkLineage package verification alone (no GPU, no writes; also runs inside smoke):
.venv/bin/python -B experiments/qwen35_4b_count_walk_menders_confirmation/scripts/rebuild_lineage.py --verify-inputsOps: crash recovery for torn artifacts
A hard crash can tear (partially write) exactly one derived artifact — the trailing ledger line, a per-seed summary.json, or the terminal confirmation_readout.json. Each of these is a deterministic pure function of the authenticated gateway receipts and the frozen pins, so the recovery is always the same:
- Audit the event directory, then DELETE only the torn artifact (never edit it in place).
- Re-run with
--resume: the artifact regenerates BYTE-IDENTICALLY from the preserved receipts, the byte-equality reconciliation re-anchors it, and the ledger close proceeds.
NEVER edit a receipt, summary, ledger line, or readout by hand — every one of them is sha-pinned at close time and a hand edit permanently fails the chain. Per-arm gateway receipts are written atomically by the trusted gateway (temp-file + rename), so a torn receipt should not occur; if one somehow does, audit it, delete it, and --resume re-runs only that arm through the gateway (receipts are gateway outputs, not deterministic regenerations — the re-run consumes no new seed because the seed's opened record already exists). A preserved <arm>.failure.json always requires an explicit audit-and-delete before any retry.
Results
Pending: the model-free construction is frozen; the four-seed sealed event runs behind its review verdict. The terminal artifact will be runs/benchmark/confirmation_readout.json with the frozen three-state verdict.
Interpretation
Pending the sealed events. Whatever the draw, the frozen claims above are the only sentences this cell may emit; a REPLICATED verdict speaks about the composite as built, never about the refuted count-don't-walk expression mechanism.
Knowledgebase Update
- Program evidence updated: pending the readout.
- Program backlog updated: pending the readout.
- Claim ledger updated: pending the readout (design-only work manufactures no claim).
Artifacts
scripts/run_benchmark.py— the hardened four-seed sixteen-run event runner (k-seed write-ahead ledger, byte-equal crash reconciliation, fail-closed arm authentication, pre-consumption implementation check, the frozen full-episode replication rule).scripts/check_benchmark.py— ledger-anchored readout writer/verifier.scripts/power_analysis.py— the exact preregistered power arithmetic (both alphas, the counterfactual ceiling, the NOT_REPLICATED retention).scripts/rebuild_lineage.py— the in-cell standalone rebuild path (stages 1-6 zero-root parent; stage 7 both arms at seed 85;--verify-inputs).scripts/run.py—--smokeand--stage benchmarkonly.data/lineage/— the copied ordered stage datasets, the extendedlineage_manifest.json, and lifecycle 22's provenance receipts.data/count_walk.jsonl,data/replay_ctl7.jsonl,data/sft_count_walk.jsonl,data/sft_blend.jsonl,data/stream_token_receipt.json— the stage-7 production inputs, byte-identical copies.scripts/lineage_trainers/,scripts/train_think.py,scripts/merge_adapter.py,scripts/train_trial.py,scripts/merge_trained_arm.py,scripts/rebuild_clean_chain.py— the byte-identical production script copies.data/provenance/— byte-identical verification copies of the four committed provenance documents.reports/preregistration.md— the frozen contract (with the recorded pre-event review amendments).reports/artifact_manifest.yaml
Report
Rendered from reports/report.md
Summary
AMBIGUOUS — no claim, by the frozen rule. Across the four fresh sealed events (78164-78167) the count_walk candidate hit one full menders episode (78164) plus one partial (78167, recorded-never-counted per the review-amended floor semantics); the replay control ALSO hit a full episode at 78164. hits_c = 1 (rule required ≥ 2) and episode totals tie replay 1-1 (strict dominance required) → the preregistered middle verdict with its frozen claim text: no claim; further spending on this contrast requires a mechanism-differentiated NEW design, not more seeds of the same. The honest reading: the untreated control's episode is the ~10% background noise process the preregistration priced, and lifecycle 27's MECHANISM_ANSWER (78163) is now best read as that coincidence landing photogenically. Menders remains without a confirmed mover. Descriptive only: count_walk topped the aggregate at 2 of 4 seeds (0.398 at 78164 and 0.392 at 78167 — its best readings on record), and lost narrowly at the other two (parent 0.3492 vs 0.3373; replay 0.3283 vs 0.3269).
Research Program Fit
The program's confirmation law: a favorable draw is priced by fresh sealed seeds, never by re-reading the discovery. This cell does for the first menders movement what the goal-gate confirmation did for the 10/10 sweep, at the same instrument (medium/tb1024, trusted gateway only).
Method
See reports/preregistration.md — frozen identities, the seed-freshness audit, the replication rule, the power arithmetic, the provenance boundary, and the recorded pre-event review amendments. The runner (scripts/run_benchmark.py) enforces the review verdict, clean pushed main, gateway sha, fail-closed tree/weights authentication of all four arms, the k-seed write-ahead ledger with byte-equal crash reconciliation, and implementation-signature equality against the pinned prior event — checked live BEFORE each seed's first gateway call (pre-consumption) and again across all sixteen receipts.
Results
Pending: no seed has been consumed. The terminal artifact will be runs/benchmark/confirmation_readout.json; every verdict input is provenance-anchored (receipt shas pinned in closed ledger records; the readout refuses any break in the sealed chain).
Controls
Three control arms per event (base, the zero-root parent, and the exposure-matched replay control from the same lifecycle-27 training pair), all authenticated by full tree+weights sha256 against design-time constants; descriptive per-family tables, goal gates, and candidate-vs-control deltas are recorded per event and never gate.
Oracle Versus Deployable Evidence
Gateway aggregates and public family scores only; benchmarks/ contents never parsed or read as data. Menders episode counts derive from public family scores via the frozen floor conversion int(10*score + 1e-9): partial-credit draws are recorded as raw positives but never counted as hits or episodes.
Interpretation
Deferred to the frozen consequence set: REPLICATED claims a menders-rate difference for the composite as built (mechanism-agnostic — lifecycle 27 already refuted its taught expression route locally); NOT_REPLICATED closes 78,163 as seed noise and leaves the expression-cost law standing (at a true per-event hit rate of 0.3 this outcome retains probability ≈ 0.24 — the closure is a preregistered funding decision, not a nonexistence proof); AMBIGUOUS forbids further seeds on this contrast in favor of a mechanism-differentiated new design.
Next Experiments
Determined by the verdict, per the frozen claims; no successor is funded from this cell's design phase.
Artifact Manifest
Four composite pins are external with committed receipts (verification copies in data/provenance/); reproduction of the three non-base composites is IN-CELL via the copied lineage package and scripts/rebuild_lineage.py (review amendment B1); everything else is in-repo. See reports/artifact_manifest.yaml.
Experiment log 3
Show the running log (3 entries, 2026-07-17)
2026-07-17 — pre-event review amendments (no seed consumed)
The adversarial review of the frozen design (frozen at bd253e48) returned 3 MAJOR + 4 minor findings; all were applied inside the legitimate pre-event amendment window — the ledger does not exist, no gateway call has ever run, and every change precedes the benchmark design review verdict. Full provenance in reports/preregistration.md, "Review amendments" section.
- A1+A2 — one coherent full-episode semantics. Episode conversion moved from
round(10*score)to FLOORint(10*score + 1e-9)(a k/60-lattice partial-credit draw contributes ZERO episodes unless it crosses a full 0.1 step; a new unit test sweeps all 61 lattice points via the float k/60 representation and matchesint(k/6)exactly). Hits redefined: an event is a hit only if the arm's FULL-EPISODE count is > 0 — partial-only events are recorded descriptively (raw_positiveper event and per arm) but are neither hits nor episodes. The rule now coincides exactly with the priced model. Power arithmetic restated on the full-episode null: alpha 0.0450 at the headline p = 0.10 AND 0.0475 at the exact p = 3/29 (exact fraction 11885589964581732052992/250246473680347348787521 = 0.04749553426180864, printed and--check-enforced); p = 5/29 (0.0947) retained strictly as a counterfactual ceiling; hits>=2 and REPLICATED power numbers verified unchanged (0.5248/0.6875/0.8735 and 0.4717/0.6289/0.8230). - B1 — standalone doctrine for an eval-only cell. Copied byte-identically from lifecycle 27: the entire
data/lineage/package (six stage datasets, manifest, seven provenance receipts), the stage-7 production inputs (count_walk.jsonl,replay_ctl7.jsonl,sft_count_walk.jsonl,sft_blend.jsonl,stream_token_receipt.json), and the production scripts (lineage_trainers/×3,train_think.py,merge_adapter.py,rebuild_clean_chain.py,train_trial.py,merge_trained_arm.py). Extended the copiedlineage_manifest.jsonwith astage7_confirmation_armsblock: both arm streams (shas recomputed from the copies — replay_ctl794e8259e…, count_walk71291542…), training seed 85, trainer/merger shas, and the final composite tree/weights pins this cell authenticates. Addedscripts/rebuild_lineage.py(stages 1-6 rebuild the zero-root parent; stage 7 trains the two arms with the train_trial.py recipe at seed 85 and merges via the merge_trained_arm.py merge); its--verify-inputschecks every copied file against the manifest shas and runs green in smoke and a new unit test. Reproduction path is now IN-CELL; receipt copies remain verification aids. - Minor 1. Design-time audit corrected to 9 recorded medium/tb1024 sealed events (78,150/78,154/78,155/78,156/78,157/78,159/78,160/78,162/78,163); the 29 arm-event count was already correct.
- Minor 2. The implementation-signature equality check now ALSO runs pre-consumption, before each seed's FIRST gateway arm (live signature via the trusted gateway's own hash-only inventory functions) — a drifted suite refuses before any GPU run or opened record; the post-arm check is kept.
- Minor 3. NOT_REPLICATED consequence text now carries "(at a true per-event hit rate of 0.3 this outcome retains probability ≈ 0.24 — the closure is a preregistered funding decision, not a nonexistence proof)" everywhere it is stated; 0.2401 is
--check-enforced. - Minor 4. Torn-ledger / partial-receipt manual recovery documented in the README ops section (delete the torn artifact;
--resumeregenerates byte-identically; never edit receipts by hand). - Tests updated for the new semantics (lattice sweep, partial-only NOT_REPLICATED branch, raw-positive records, lineage package); full suite green; smoke green;
power_analysis.py --checkgreen;rebuild_lineage.py --verify-inputsgreen. No GPU stage run; no seed consumed.
2026-07-17 — design freeze (lifecycle 28, eval-only)
- Scaffolded as the mandatory confirmation cell for lifecycle 27's MECHANISM_ANSWER (count_walk menders 0.1 vs base / zero_root_parent / replay_ctl7 all 0.0 at sealed seed 78,163).
- Seed-freshness audit: 78,164 / 78,165 / 78,166 / 78,167 verified grep-fresh in seed contexts across the repo (every raw numeric hit is a float/sha substring in unrelated data files); benchmark seeds previously spent through 78,163; no substitution required.
- Frozen the integer-exact two-directional replication rule (REPLICATED / NOT_REPLICATED / AMBIGUOUS, no fourth state) with all three claims worded in the preregistration, and the exact power arithmetic (false-REPLICATED 0.0450 under the p=0.10 noise model, sensitivity 0.0947; power of hits_c >= 2: 0.5248 / 0.6875 / 0.8735 at q = 0.4 / 0.5 / 0.65; full REPLICATED power 0.4717 / 0.6289 / 0.8230), recomputed fail-closed by
scripts/power_analysis.py --check. - Cloned and adapted the hardened runner machinery: fail-closed tree+weights authentication of the four pre-existing composites (constants baked at design time, no TODO pins), lifecycle-27 merge receipt + lifecycle-22 zero-root provenance authentication, gateway sha pin, k-seed write-ahead opened/closed ledger with byte-equal crash reconciliation, implementation-signature equality against the pinned prior event, ledger-anchored terminal readout.
- Copied the four committed provenance documents byte-identically into
data/provenance/as verification aids; composite reproduction remains lifecycle 27's / lifecycle 22's own standalone rebuild path (this cell produces no model); the measurement gateway stays shared per docs/quality_gates.md. - Unit tests: replication-rule truth table (including the E_c tie branch), ledger open/close/reconcile/double-consume refusals, arm-authentication failure paths, frozen constants, readout schema, finiteness guards, power arithmetic. Smoke green; no GPU stage run; no seed consumed.
- Next checkpoint: adversarial benchmark design review (
reports/benchmark_design_review.mdwith the literal PASS_BENCHMARK_EVENT verdict) before--stage benchmarkcan consume any seed.
2026-07-17 — Four-seed event complete: AMBIGUOUS; cell closed
- Events 78164-78167 (16 runs, all within budget, paired comparison valid, implementation signature identical across all sixteen receipts and the prior event). Menders per event: 78164 candidate 0.1 AND replay control 0.1 (one full episode each); 78165/78166 all arms 0.0; 78167 candidate 0.0167 partial (recorded, never counted — the review's floor-semantics fix operating as designed).
- Frozen rule: hits_c = 1 (< 2) and episode totals candidate 1 vs replay 1 (tie = no dominance) → AMBIGUOUS. Frozen claim applies: no claim; further spending on this contrast requires a mechanism-differentiated NEW design, not more seeds of the same.
- Honest reading: the replay control's full episode at 78164 is the noise process the preregistration priced (background arm-event rate ~0.10); the 78163 MECHANISM_ANSWER is now best read as that coincidence. Menders remains without a confirmed mover. Descriptive: count_walk topped the aggregate at 78164 (0.398) and 78167 (0.392), lost narrowly at 78165 (parent 0.3492 vs 0.3373) and 78166 (replay 0.3283 vs 0.3269) — single-seed readings, never gating.
Data files 4
Result tables and metrics copied from the experiment folder — preview inline or open the raw file.
runs/benchmark/medium_tb1024_seed78164_confirmation/summary.json3.9 kBruns/benchmark/medium_tb1024_seed78165_confirmation/summary.json4.0 kBruns/benchmark/medium_tb1024_seed78166_confirmation/summary.json4.0 kBruns/benchmark/medium_tb1024_seed78167_confirmation/summary.json4.0 kB
Reproduce
Smoke test
PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_count_walk_menders_confirmation/scripts/run.py --smokeFull run
.venv/bin/python -B experiments/qwen35_4b_count_walk_menders_confirmation/scripts/run.py --stage benchmarkRun steps are documented inside the experiment folder (README and scripts).