Research log Small Model Experimentation
GitHub

Count-Walk Menders Confirmation

Ambiguous: the menders hit did not replicate — a control drew the same score, so no claim is made

The one idea you need

The previous experiment produced a first: on one sealed exam, its trained model solved a fix-the-broken-procedure puzzle (the 'menders' family) that all three of its comparison models scored exactly zero on. That pattern - trained model above zero, every comparison at zero - was the pre-declared positive outcome, and it had never happened before in this program. But there is a catch the team wrote down immediately: exams are drawn fresh each time, and roughly one untrained model in ten also lands a single menders solve on any given draw. One good draw can be luck. The only honest way to tell luck from skill is to run the same head-to-head again on fresh sealed exams and count.

The question

If the same four models sit four fresh sealed exams, does the trained model keep solving menders puzzles while the three comparison models keep scoring zero - or does the one win from last time dissolve into ordinary exam-to-exam noise?

What we found

The answer the rule was built to force out, delivered without wiggle room. Across the four fresh exams the trained model solved a fix-the-procedure episode exactly once (plus one partial credit that the rules pre-declared doesn't count) — and on that same exam, the comparison model trained WITHOUT the special lessons solved one too. One hit when two were required, and a dead tie against a control, is the pre-written middle verdict: no claim. The clean interpretation is that occasional single-episode solves are background luck this family hands out to roughly one run in ten — the pre-registered noise rate — and the earlier headline result (the trained model scoring while all three controls sat at zero) was most likely that luck landing photogenically. The rule also pre-committed the consequence: no more exam seeds for this comparison; any future attempt at this family must be a genuinely different design, not a re-roll. One quietly encouraging descriptive note: the trained model posted the best overall score on two of the four exams (0.398 and 0.392, its two best readings ever), though those readings carry no claim.

Why it matters

The program's one unbroken habit is pricing good news honestly: every favorable draw must survive fresh sealed exams before anyone may claim it. Menders is the single exam family that has blocked the program's headline goal on every previous attempt, so a confirmed movement there would be the first durable capability gain on the hardest surface - and a clean failure to replicate is nearly as valuable, because it closes the line and frees the budget for a different mechanism.

Candidate hits1 of 4 eventsrule required ≥ 2; partial at 78167 recorded, not counted
Control tie1 – 1replay control also solved an episode at 78164
VerdictAMBIGUOUSfrozen: no claim; no more seeds of this design
Aggregate (desc.)0.398 / 0.392count_walk tops 2 of 4 seeds — descriptive only
On this page
  1. Results at a glance
  2. Overview
  3. Report
    1. Summary
    2. Research Program Fit
    3. Method
    4. Results
    5. Controls
    6. Oracle Versus Deployable Evidence
    7. Interpretation
    8. Next Experiments
    9. Artifact Manifest
  4. Experiment log
  5. Data files
  6. Reproduce
  7. Related

Results at a glance 1

Four-seed confirmation: the menders reading does not replicate (AMBIGUOUS, no claim) The frozen integer rule read out mechanically: the candidate hit one full episode in four events (needed ≥2) and the replay control hit one too (78164, both 0.1), tying episode totals 1–1 — no dominance. The 78167 partial draw (0.0167) is recorded but never counted, per the review-amended floor semantics that replaced banker's rounding (which could have manufactured phantom episodes). Verdict AMBIGUOUS with the preregistered claim text: no claim; further spending on this contrast requires a mechanism-differentiated NEW design, not more seeds of the same. Honest reading: the untreated replay control drawing a full episode is exactly the ~10% background noise process the preregistration priced, so the prior event's MECHANISM_ANSWER (78163: candidate 0.1 vs all controls 0.0) is now best read as that coincidence. Menders remains without a confirmed mover. Descriptive only: count_walk topped the aggregate at 2 of 4 seeds (0.398 at 78164, 0.392 at 78167).

menders score · sealed event seed →

00.0250.050.0750.10.10.100781640000781650000781660.016700078167
Data table
sealed event seedcount_walk mendersreplay_ctl7 menderszero_root_parent mendersbase menders
781640.10.100
781650000
781660000
781670.0167000

Numbers from experiments/qwen35_4b_count_walk_menders_confirmation/runs/benchmark/confirmation_readout.json

In the author’s words from the Report · “Summary”

hits_c = 1 (rule required ≥ 2) and episode totals tie replay 1-1 (strict dominance required) → the preregistered middle verdict with its frozen claim text: no claim; further spending on this contrast requires a mechanism-differentiated NEW design, not more seeds of the same. The honest reading: the untreated control's episode is the ~10% background noise process the preregistration priced, and lifecycle 27's MECHANISM_ANSWER (78163) is now best read as that coincidence landing photogenically. Menders remains without a confirmed mover. … Read the full result →

Overview

The mandatory multi-seed confirmation of lifecycle 27's MECHANISM_ANSWER: at sealed seed 78,163 the count_walk composite drew menders 0.1 while base, zero_root_parent, and replay_ctl7 all drew exactly 0.0 — the preregistered positive branch, first in program history. Eval-only: four fresh sealed medium seeds, four authenticated pre-existing arms per seed, one frozen integer-exact replication rule, no training anywhere.

Research Program

  • Program: agentic_breadth_installation
  • Program question: can DESIGNED synthetic curricula install universal, transferable agentic skills into the one 4B — provably, on fully documented lineage?
  • Prior anchors: lifecycle 27 (qwen35_4b_count_dont_walk_enumeration — the prior event: count_walk menders 0.1 vs all controls 0.0 at seed 78,163, aggregates 0.3312 / 0.3298 / 0.2950 / base 0.0753), lifecycle 26 (qwen35_4b_enumerative_repair_protocol — its replay control drew menders 0.1 at seed 78,162, proving untreated arms can draw), and the goal-gate confirmation law (a favorable draw is priced by fresh sealed seeds, never by re-reading).

Question

Does the seed-78163 menders pattern — candidate above zero while every control sits at exactly zero — replicate across four fresh sealed seeds, or does it close as seed noise? Single-episode menders draws by untreated arms have already happened once in nine recorded medium events, so under the observed noise rate one MECHANISM_ANSWER draw has non-trivial probability; only replication can separate a real rate difference from a favorable roll.

Hypothesis

If the count-dont-walk dose genuinely moved menders, the candidate hits episodes at a per-event rate materially above the program's 0.10 arm-event noise rate and accumulates more episodes than every control over four events. Honest prior: the mechanism the dose TAUGHT was already refuted locally (the candidate still thinks to the 1,024-token cap; fidelity 7/40 vs the 0.50 bar), so this cell tests the capability movement itself, mechanism-agnostic, and a NOT_REPLICATED close is the likelier branch under the noise model.

Setup

  • Model: Qwen/Qwen3.5-4B (revision 851bf6e8…), always; nothing trains or merges.
  • Arms (four pre-existing committed composites, full tree+weights sha256 recomputed and matched fail-closed at event time against design-time constants — no TODO pins): base (tree 26d8ee48…, weights b654e033…), zero_root_parent (tree 414f5829…, weights 6e9aad25…, lifecycle 22's committed lineage merge receipt e906caea…), replay_ctl7 (tree 044a4599…, weights c5035b4d…, committed receipt 3f65b4c6…), count_walk (tree d5fdc55c…, weights ddd7bc4b…, committed receipt 840edca0…).
  • Event: four fresh sealed seeds 78,164 / 78,165 / 78,166 / 78,167 (grep-fresh in seed contexts at design time; audit in the preregistration), tier medium, think budget 1,024, four arms per seed in frozen order, seed-major; per-seed write-ahead opened/closed ledger, closed records sha-pin the summary and all four gateway receipts; --resume is the single recovery path with byte-identical deterministic summary regeneration; the implementation signature is checked LIVE before each seed's first gateway call (pre-consumption — a drifted suite refuses before any GPU run) and all sixteen receipts must equal the prior event's pinned block.
  • FROZEN REPLICATION RULE (integer-exact, over the four NEW events only; 78,163 is prior evidence, never pooled; review amendment A1+A2): an event counts as a hit only if it contains at least one FULL menders episode (score contributes int(10*s + 1e-9) episodes, FLOOR semantics — partial-credit draws on the k/60 lattice are recorded but never counted, as hits or episodes); hits_c = new events with candidate FULL-EPISODE count > 0; E per arm = sum of int(10*score + 1e-9) episodes. REPLICATED iff hits_c >= 2 AND E_c strictly exceeds EVERY control's total; NOT_REPLICATED iff hits_c == 0; AMBIGUOUS otherwise. No fourth state. The three frozen consequences:

    1. REPLICATED — "the count_walk composite solves menders episodes at a rate no

    control matches; the first confirmed menders capability movement in the program." 2. NOT_REPLICATED — "the 78163 reading closes as seed noise; the count-dont-walk dose did not durably move menders; the expression-cost law stands; the composite remains a documented artifact (at a true per-event hit rate of 0.3 this outcome retains probability ≈ 0.24 — the closure is a preregistered funding decision, not a nonexistence proof)." 3. AMBIGUOUS — "no claim; further spending on this contrast requires a mechanism-differentiated NEW design, not more seeds of the same."

  • Honest priors (arithmetic in the preregistration; recomputed by scripts/power_analysis.py --check in smoke and tests, which enforce every printed number): under the FULL-EPISODE null (design-time audit over all 9 recorded medium/tb1024 sealed events, 29 arm-events, 3 full-episode draws; the 2 partial draws are rule-invisible, recorded-only) the false-REPLICATED probability is 0.0450 at the headline p = 0.10 and 0.0475 at the exact p = 3/29 (exact fraction printed by the script); the counterfactual ceiling — if every raw-positive draw were promoted to a full episode, which the frozen conversion forbids — is 0.0947 at p = 5/29. Power of hits_c >= 2 is 0.5248 / 0.6875 / 0.8735 at candidate hit rates 0.4 / 0.5 / 0.65, and full REPLICATED power (with dominance) is 0.4717 / 0.6289 / 0.8230 (both unchanged by the amendment); NOT_REPLICATED retains probability 0.2401 at q = 0.3.
  • Standalone boundary (review amendment B1): this cell produces no model but EVALUATES three non-base composites, so it carries the complete in-cell reproduction package per docs/quality_gates.md (including eval-only cells): data/lineage/ (six ordered zero-root stage datasets + the extended lineage_manifest.json + lifecycle 22's provenance receipts), the stage-7 production inputs (data/count_walk.jsonl, data/replay_ctl7.jsonl, data/sft_count_walk.jsonl, data/sft_blend.jsonl, data/stream_token_receipt.json), the byte-identical production scripts (scripts/lineage_trainers/, scripts/train_think.py, scripts/merge_adapter.py, scripts/train_trial.py, scripts/merge_trained_arm.py, scripts/rebuild_clean_chain.py), and scripts/rebuild_lineage.py (stages 1-6 rebuild the zero-root parent; stage 7 trains both arms at fixed seed 85 and merges; --verify-inputs runs in smoke and tests). The four committed provenance documents remain copied byte-identically into data/provenance/ as verification aids; the measurement gateway stays shared per docs/quality_gates.md.
  • Hidden-label boundary: the benchmark suite's contents are never parsed or read as data; only the trusted aggregate gateway (sha 53cf6533…) runs, and the pre-consumption implementation check hashes suite bytes exclusively through that gateway's own inventory functions.

Run

Smoke (no GPU, no writes):

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_count_walk_menders_confirmation/scripts/run.py --smoke

Full (the only stage; requires the committed PASS_BENCHMARK_EVENT review and clean pushed main):

.venv/bin/python -B experiments/qwen35_4b_count_walk_menders_confirmation/scripts/run.py --stage benchmark

Lineage package verification alone (no GPU, no writes; also runs inside smoke):

.venv/bin/python -B experiments/qwen35_4b_count_walk_menders_confirmation/scripts/rebuild_lineage.py --verify-inputs

Ops: crash recovery for torn artifacts

A hard crash can tear (partially write) exactly one derived artifact — the trailing ledger line, a per-seed summary.json, or the terminal confirmation_readout.json. Each of these is a deterministic pure function of the authenticated gateway receipts and the frozen pins, so the recovery is always the same:

  1. Audit the event directory, then DELETE only the torn artifact (never edit it in place).
  2. Re-run with --resume: the artifact regenerates BYTE-IDENTICALLY from the preserved receipts, the byte-equality reconciliation re-anchors it, and the ledger close proceeds.

NEVER edit a receipt, summary, ledger line, or readout by hand — every one of them is sha-pinned at close time and a hand edit permanently fails the chain. Per-arm gateway receipts are written atomically by the trusted gateway (temp-file + rename), so a torn receipt should not occur; if one somehow does, audit it, delete it, and --resume re-runs only that arm through the gateway (receipts are gateway outputs, not deterministic regenerations — the re-run consumes no new seed because the seed's opened record already exists). A preserved <arm>.failure.json always requires an explicit audit-and-delete before any retry.

Results

Pending: the model-free construction is frozen; the four-seed sealed event runs behind its review verdict. The terminal artifact will be runs/benchmark/confirmation_readout.json with the frozen three-state verdict.

Interpretation

Pending the sealed events. Whatever the draw, the frozen claims above are the only sentences this cell may emit; a REPLICATED verdict speaks about the composite as built, never about the refuted count-don't-walk expression mechanism.

Knowledgebase Update

  • Program evidence updated: pending the readout.
  • Program backlog updated: pending the readout.
  • Claim ledger updated: pending the readout (design-only work manufactures no claim).

Artifacts

  • scripts/run_benchmark.py — the hardened four-seed sixteen-run event runner (k-seed write-ahead ledger, byte-equal crash reconciliation, fail-closed arm authentication, pre-consumption implementation check, the frozen full-episode replication rule).
  • scripts/check_benchmark.py — ledger-anchored readout writer/verifier.
  • scripts/power_analysis.py — the exact preregistered power arithmetic (both alphas, the counterfactual ceiling, the NOT_REPLICATED retention).
  • scripts/rebuild_lineage.py — the in-cell standalone rebuild path (stages 1-6 zero-root parent; stage 7 both arms at seed 85; --verify-inputs).
  • scripts/run.py--smoke and --stage benchmark only.
  • data/lineage/ — the copied ordered stage datasets, the extended lineage_manifest.json, and lifecycle 22's provenance receipts.
  • data/count_walk.jsonl, data/replay_ctl7.jsonl, data/sft_count_walk.jsonl, data/sft_blend.jsonl, data/stream_token_receipt.json — the stage-7 production inputs, byte-identical copies.
  • scripts/lineage_trainers/, scripts/train_think.py, scripts/merge_adapter.py, scripts/train_trial.py, scripts/merge_trained_arm.py, scripts/rebuild_clean_chain.py — the byte-identical production script copies.
  • data/provenance/ — byte-identical verification copies of the four committed provenance documents.
  • reports/preregistration.md — the frozen contract (with the recorded pre-event review amendments).
  • reports/artifact_manifest.yaml

Report

Rendered from reports/report.md

Summary

AMBIGUOUS — no claim, by the frozen rule. Across the four fresh sealed events (78164-78167) the count_walk candidate hit one full menders episode (78164) plus one partial (78167, recorded-never-counted per the review-amended floor semantics); the replay control ALSO hit a full episode at 78164. hits_c = 1 (rule required ≥ 2) and episode totals tie replay 1-1 (strict dominance required) → the preregistered middle verdict with its frozen claim text: no claim; further spending on this contrast requires a mechanism-differentiated NEW design, not more seeds of the same. The honest reading: the untreated control's episode is the ~10% background noise process the preregistration priced, and lifecycle 27's MECHANISM_ANSWER (78163) is now best read as that coincidence landing photogenically. Menders remains without a confirmed mover. Descriptive only: count_walk topped the aggregate at 2 of 4 seeds (0.398 at 78164 and 0.392 at 78167 — its best readings on record), and lost narrowly at the other two (parent 0.3492 vs 0.3373; replay 0.3283 vs 0.3269).

Research Program Fit

The program's confirmation law: a favorable draw is priced by fresh sealed seeds, never by re-reading the discovery. This cell does for the first menders movement what the goal-gate confirmation did for the 10/10 sweep, at the same instrument (medium/tb1024, trusted gateway only).

Method

See reports/preregistration.md — frozen identities, the seed-freshness audit, the replication rule, the power arithmetic, the provenance boundary, and the recorded pre-event review amendments. The runner (scripts/run_benchmark.py) enforces the review verdict, clean pushed main, gateway sha, fail-closed tree/weights authentication of all four arms, the k-seed write-ahead ledger with byte-equal crash reconciliation, and implementation-signature equality against the pinned prior event — checked live BEFORE each seed's first gateway call (pre-consumption) and again across all sixteen receipts.

Results

Pending: no seed has been consumed. The terminal artifact will be runs/benchmark/confirmation_readout.json; every verdict input is provenance-anchored (receipt shas pinned in closed ledger records; the readout refuses any break in the sealed chain).

Controls

Three control arms per event (base, the zero-root parent, and the exposure-matched replay control from the same lifecycle-27 training pair), all authenticated by full tree+weights sha256 against design-time constants; descriptive per-family tables, goal gates, and candidate-vs-control deltas are recorded per event and never gate.

Oracle Versus Deployable Evidence

Gateway aggregates and public family scores only; benchmarks/ contents never parsed or read as data. Menders episode counts derive from public family scores via the frozen floor conversion int(10*score + 1e-9): partial-credit draws are recorded as raw positives but never counted as hits or episodes.

Interpretation

Deferred to the frozen consequence set: REPLICATED claims a menders-rate difference for the composite as built (mechanism-agnostic — lifecycle 27 already refuted its taught expression route locally); NOT_REPLICATED closes 78,163 as seed noise and leaves the expression-cost law standing (at a true per-event hit rate of 0.3 this outcome retains probability ≈ 0.24 — the closure is a preregistered funding decision, not a nonexistence proof); AMBIGUOUS forbids further seeds on this contrast in favor of a mechanism-differentiated new design.

Next Experiments

Determined by the verdict, per the frozen claims; no successor is funded from this cell's design phase.

Artifact Manifest

Four composite pins are external with committed receipts (verification copies in data/provenance/); reproduction of the three non-base composites is IN-CELL via the copied lineage package and scripts/rebuild_lineage.py (review amendment B1); everything else is in-repo. See reports/artifact_manifest.yaml.

Experiment log 3

Show the running log (3 entries, 2026-07-17)

2026-07-17 — pre-event review amendments (no seed consumed)

The adversarial review of the frozen design (frozen at bd253e48) returned 3 MAJOR + 4 minor findings; all were applied inside the legitimate pre-event amendment window — the ledger does not exist, no gateway call has ever run, and every change precedes the benchmark design review verdict. Full provenance in reports/preregistration.md, "Review amendments" section.

  • A1+A2 — one coherent full-episode semantics. Episode conversion moved from round(10*score) to FLOOR int(10*score + 1e-9) (a k/60-lattice partial-credit draw contributes ZERO episodes unless it crosses a full 0.1 step; a new unit test sweeps all 61 lattice points via the float k/60 representation and matches int(k/6) exactly). Hits redefined: an event is a hit only if the arm's FULL-EPISODE count is > 0 — partial-only events are recorded descriptively (raw_positive per event and per arm) but are neither hits nor episodes. The rule now coincides exactly with the priced model. Power arithmetic restated on the full-episode null: alpha 0.0450 at the headline p = 0.10 AND 0.0475 at the exact p = 3/29 (exact fraction 11885589964581732052992/250246473680347348787521 = 0.04749553426180864, printed and --check-enforced); p = 5/29 (0.0947) retained strictly as a counterfactual ceiling; hits>=2 and REPLICATED power numbers verified unchanged (0.5248/0.6875/0.8735 and 0.4717/0.6289/0.8230).
  • B1 — standalone doctrine for an eval-only cell. Copied byte-identically from lifecycle 27: the entire data/lineage/ package (six stage datasets, manifest, seven provenance receipts), the stage-7 production inputs (count_walk.jsonl, replay_ctl7.jsonl, sft_count_walk.jsonl, sft_blend.jsonl, stream_token_receipt.json), and the production scripts (lineage_trainers/ ×3, train_think.py, merge_adapter.py, rebuild_clean_chain.py, train_trial.py, merge_trained_arm.py). Extended the copied lineage_manifest.json with a stage7_confirmation_arms block: both arm streams (shas recomputed from the copies — replay_ctl7 94e8259e…, count_walk 71291542…), training seed 85, trainer/merger shas, and the final composite tree/weights pins this cell authenticates. Added scripts/rebuild_lineage.py (stages 1-6 rebuild the zero-root parent; stage 7 trains the two arms with the train_trial.py recipe at seed 85 and merges via the merge_trained_arm.py merge); its --verify-inputs checks every copied file against the manifest shas and runs green in smoke and a new unit test. Reproduction path is now IN-CELL; receipt copies remain verification aids.
  • Minor 1. Design-time audit corrected to 9 recorded medium/tb1024 sealed events (78,150/78,154/78,155/78,156/78,157/78,159/78,160/78,162/78,163); the 29 arm-event count was already correct.
  • Minor 2. The implementation-signature equality check now ALSO runs pre-consumption, before each seed's FIRST gateway arm (live signature via the trusted gateway's own hash-only inventory functions) — a drifted suite refuses before any GPU run or opened record; the post-arm check is kept.
  • Minor 3. NOT_REPLICATED consequence text now carries "(at a true per-event hit rate of 0.3 this outcome retains probability ≈ 0.24 — the closure is a preregistered funding decision, not a nonexistence proof)" everywhere it is stated; 0.2401 is --check-enforced.
  • Minor 4. Torn-ledger / partial-receipt manual recovery documented in the README ops section (delete the torn artifact; --resume regenerates byte-identically; never edit receipts by hand).
  • Tests updated for the new semantics (lattice sweep, partial-only NOT_REPLICATED branch, raw-positive records, lineage package); full suite green; smoke green; power_analysis.py --check green; rebuild_lineage.py --verify-inputs green. No GPU stage run; no seed consumed.

2026-07-17 — design freeze (lifecycle 28, eval-only)

  • Scaffolded as the mandatory confirmation cell for lifecycle 27's MECHANISM_ANSWER (count_walk menders 0.1 vs base / zero_root_parent / replay_ctl7 all 0.0 at sealed seed 78,163).
  • Seed-freshness audit: 78,164 / 78,165 / 78,166 / 78,167 verified grep-fresh in seed contexts across the repo (every raw numeric hit is a float/sha substring in unrelated data files); benchmark seeds previously spent through 78,163; no substitution required.
  • Frozen the integer-exact two-directional replication rule (REPLICATED / NOT_REPLICATED / AMBIGUOUS, no fourth state) with all three claims worded in the preregistration, and the exact power arithmetic (false-REPLICATED 0.0450 under the p=0.10 noise model, sensitivity 0.0947; power of hits_c >= 2: 0.5248 / 0.6875 / 0.8735 at q = 0.4 / 0.5 / 0.65; full REPLICATED power 0.4717 / 0.6289 / 0.8230), recomputed fail-closed by scripts/power_analysis.py --check.
  • Cloned and adapted the hardened runner machinery: fail-closed tree+weights authentication of the four pre-existing composites (constants baked at design time, no TODO pins), lifecycle-27 merge receipt + lifecycle-22 zero-root provenance authentication, gateway sha pin, k-seed write-ahead opened/closed ledger with byte-equal crash reconciliation, implementation-signature equality against the pinned prior event, ledger-anchored terminal readout.
  • Copied the four committed provenance documents byte-identically into data/provenance/ as verification aids; composite reproduction remains lifecycle 27's / lifecycle 22's own standalone rebuild path (this cell produces no model); the measurement gateway stays shared per docs/quality_gates.md.
  • Unit tests: replication-rule truth table (including the E_c tie branch), ledger open/close/reconcile/double-consume refusals, arm-authentication failure paths, frozen constants, readout schema, finiteness guards, power arithmetic. Smoke green; no GPU stage run; no seed consumed.
  • Next checkpoint: adversarial benchmark design review (reports/benchmark_design_review.md with the literal PASS_BENCHMARK_EVENT verdict) before --stage benchmark can consume any seed.

2026-07-17 — Four-seed event complete: AMBIGUOUS; cell closed

  • Events 78164-78167 (16 runs, all within budget, paired comparison valid, implementation signature identical across all sixteen receipts and the prior event). Menders per event: 78164 candidate 0.1 AND replay control 0.1 (one full episode each); 78165/78166 all arms 0.0; 78167 candidate 0.0167 partial (recorded, never counted — the review's floor-semantics fix operating as designed).
  • Frozen rule: hits_c = 1 (< 2) and episode totals candidate 1 vs replay 1 (tie = no dominance) → AMBIGUOUS. Frozen claim applies: no claim; further spending on this contrast requires a mechanism-differentiated NEW design, not more seeds of the same.
  • Honest reading: the replay control's full episode at 78164 is the noise process the preregistration priced (background arm-event rate ~0.10); the 78163 MECHANISM_ANSWER is now best read as that coincidence. Menders remains without a confirmed mover. Descriptive: count_walk topped the aggregate at 78164 (0.398) and 78167 (0.392), lost narrowly at 78165 (parent 0.3492 vs 0.3373) and 78166 (replay 0.3283 vs 0.3269) — single-seed readings, never gating.

Data files 4

Result tables and metrics copied from the experiment folder — preview inline or open the raw file.

Reproduce

Smoke test

PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -B experiments/qwen35_4b_count_walk_menders_confirmation/scripts/run.py --smoke

Full run

.venv/bin/python -B experiments/qwen35_4b_count_walk_menders_confirmation/scripts/run.py --stage benchmark

Run steps are documented inside the experiment folder (README and scripts).

Browse the experiment folder on GitHub ↗