Qwen3.5-4B Tokenizer-EOS Residual Mechanics Fresh Replay
The one idea you need
The parent experiment fixed where a short answer ends, but its final analysis mixed up two timing rules: the first transport check must have no later stages, while a later replay must accept a fully completed history. This successor separates those rules and starts from entirely fresh tasks.
The question
Can a clean replay finally test whether concrete intermediate consequences help this small model solve three-step programs better than spending the same effort on ordinary sampling?
What we found
No result exists yet. The scaffold declares seven freshness controls and eight lifecycle controls, reads none of the parent's sampled bundles, and authorizes no model call until independent review and release gates pass.
Why it matters
A valid rerun could answer the capability question the parent left open without learning from its 4,056 quarantined outputs or weakening the matched-effort baseline.
On this page
Results at a glance 1
How to read
The first two bars count declared freshness and lifecycle checks. The last two bars show model calls and sampled outputs, both at zero.
Takeaway → This is readiness evidence only: the successor has a stricter plan but no scientific result.
Data table
| scaffold receipt field | model-free scaffold count |
|---|---|
| freshness controls | 7 |
| lifecycle controls | 8 |
| model calls | 0 |
| sampled outputs | 0 |
Numbers from experiments/qwen35_4b_tokenizer_eos_residual_mechanics_fresh_replay/runs/smoke/summary.json
Technical framing
Fresh replay scaffold declares controls with zero experimental access — Readiness evidence only. Independent design review remains pending; the smoke read no parent sampled bundle, hidden file, or benchmark and authorizes no model call.
In the author’s words from the Overview · “Results”
Fresh three-round independent review of exact pushed-green commit 50fd804b returned PASS_IMPLEMENTATION: 145/145 protected-safe tests passed, the full 48/48 predecessor mutation matrix failed closed, both static launchers rebuilt byte-identically, and model/GPU/protected access remained zero. Canonical calibration and mechanics machine receipts were then committed, pushed, and green before calibration-lock publication. The authenticated static launcher subsequently minted the calibration-only implementation lock with SHA-256 b220466742071e3fa02a698925251132cd7bb05535a8ac54d0ce819ab90b733e. … Read the full result →
Overview
This fresh-identity successor preserves the parent's qualified tokenizer-EOS interface and frozen residual-mechanics science while separating initial transport authorization from post-chain replay. It forbids every parent sampled bundle and requires fresh functions, prompts, identities, seeds, ciphertext, and key before testing capability against matched-compute sampling.
Research Program
- Program:
structured_execution_and_compilers - Program question: can explicit, executable intermediate structure expose capabilities that direct generation and matched-compute sampling miss?
- Scientific parent:
qwen35_4b_tokenizer_eos_answer_commit_factorial. - Recovery-pattern anchor:
qwen35_4b_materialized_residual_sibling_search_fresh_replication. - Interface anchor:
qwen35_4b_materialized_residual_answer_seam_factorial.
Question
Can the already-qualified tokenizer-EOS/no-think/PROGRAM: interface turn materialized first-operation consequences into a residual capability gain over name-only, semantically shuffled, and taskwise matched-compute direct sampling on genuinely fresh exact-depth-three tasks?
The parent never answered this question. Its transport gate passed 24/24, but visible analysis failed after generation because historical replay reused the initial later-absent authorization invariant. Its 4,056 sampled outputs are scientifically quarantined and may not be imported, inspected, or rescored.
Hypothesis
Materializing each candidate first operation's public state-to-target relation should reduce a depth-three inverse problem to a depth-two residual completion. If that mechanism is causal rather than surface prompting, it must outperform name-only and token-preserving shuffled relations, and it must beat ordinary full-program sampling at both sampled-token and logical-model-token taskwise first-over budgets. A distinct initial-authorization API and immutable historical-replay API should remove the parent's instrument failure without changing any scientific arm, metric, threshold, or selector.
Setup
- Model: only
Qwen/Qwen3.5-4B, revision851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, bf16. - Backend: pinned experiment-local vLLM 0.24.0+cu129 for every model arm; no backend mixing, training, adapter, teacher, or other model.
- Dataset/task source: fresh procedural exact-depth-three tasks under seed block
2026140800--2026140806, excluding every parent function fingerprint, task/record/request identity, rendered prompt/token sequence, and sampling seed. Alias semantics remain fixed A--X to preserve the frozen science. - Train/eval split: no training. Fresh calibration and transport rows are known-answer interface instruments; mechanics visible and hidden partitions remain disjoint.
- Baseline: candidate-blind full-program sampling matched taskwise at the first sampled-token and logical-model-token overtake points.
- Controls: all-24 name-only relations and all-24 token-preserving semantic shuffles; HF-model-EOS, thinking, and freeform calibration controls remain.
- Primary metric: hidden eight-example task success after a visible/probe-only selector, compared with every frozen control and matched-compute baseline.
- Oracle-only metrics: proposal coverage and exhaustive CPU ceiling, disclosed separately and never used for selection.
- Hidden-label boundary: no hidden/gold/key read until a visible selection is committed, pushed to
main, and green. The parent raw sampled bundles are a permanently forbidden input.
Run
Model-free smoke only:
python3 -B experiments/qwen35_4b_tokenizer_eos_residual_mechanics_fresh_replay/scripts/run.py --smokeThe independent design review authorizes model-free construction only. The construction uses the pinned tokenizer and model-config metadata solely to prove rendered-token freshness; it does not load model weights or use a GPU. No model/GPU access or live command is authorized before implementation review, exact-commit green CI, and separately published locks.
Results
No residual-capability result exists yet. Fresh model-free construction passed:
- 72/72 new exact-depth-three functions and public instances are disjoint from the parent.
- 72 task IDs, 2,952 causal request/seed-key identities, 5,904 derived runner seeds, 1,824 distinct identity-free prompts, and 3,648 think/no-think rendered prompt-token sequences have zero parent intersection.
- Calibration/mechanics remain exactly 48/24 with the frozen strata and alias balance.
- Mechanics gold exists only as fresh AES-256-GCM ciphertext. The new ignored key was created once; completed-state authentication was tested without opening or hashing it.
- Construction and tokenizer receipts record zero model calls, sampled model outputs, hidden reads, benchmark reads, and parent raw-bundle reads.
An adversarial audit found that the first parent exporter reached one transitively imported protocol source without declaring it. The corrected exporter installs a repository read firewall and reconstructs the complete parent inventory from exactly eight authenticated administrative files. All scientific collision domains are byte-for-byte unchanged. The repaired receipt chain is recorded in runs/construction/parent_collision_receipt_repair.json:
- parent collision manifest:
450eb55d41b09aabff33967d5c75e6315e41cf093bd043004045a8ae5d0d07ef; - construction receipt:
fd7f09c22c468fd7678f16c5bc294461cc432382a1310c4c7d7aa3feff665115; - preoutcome receipt:
42b3710a3e52cbad9e8c00110410574ac233775eebf3c8fa5f7ed70e1568887f; - tokenizer receipt:
ea6bf73d06d94112ead421d2c408ab2cc1cd5e28079fd6cc17f93079bb373896.
The earlier hashes beginning 72fa, 7e1f, 41c7, and da57 are retained inside that migration receipt as superseded administrative bindings.
The repaired model-free implementation candidate passes 147/147 tests. Its production path now has distinct APIs for initial transport authorization and historical transport replay. Initial authorization requires every later invocation to be absent; historical replay first authenticates the complete five-invocation chain. Both paths reauthenticate after semantic replay, and all transport comparisons use recursive exact JSON types, so nested Boolean/integer aliases, descendant injection, and in-replay chain mutation fail closed. An unmocked synthetic lifecycle runs the real transport, all four descendant transactions, historical replay, and visible selection with no hidden access or recovery generation.
The independent review of exact commit 98e9e9f6 returned HOLD_IMPLEMENTATION for the undeclared source, exact-type alias, and temporal reauthentication defects. Those findings are repaired here, but that verdict did not transfer: a fresh independent review of the new exact pushed-green commit was required before release.
A release-path audit also found that the calibration critical inventory named an implementation-review report that had never been scaffolded in the successor. A tracked HOLD_IMPLEMENTATION placeholder and regression test now ensure the exact reviewed commit contains that required path; the placeholder cannot authorize a lock.
Fresh three-round independent review of exact pushed-green commit 50fd804b returned PASS_IMPLEMENTATION: 145/145 protected-safe tests passed, the full 48/48 predecessor mutation matrix failed closed, both static launchers rebuilt byte-identically, and model/GPU/protected access remained zero. Canonical calibration and mechanics machine receipts were then committed, pushed, and green before calibration-lock publication.
The authenticated static launcher subsequently minted the calibration-only implementation lock with SHA-256 b220466742071e3fa02a698925251132cd7bb05535a8ac54d0ce819ab90b733e. It binds reviewed implementation 50fd804b, review-release commit 21f964e6, all 31 calibration critical files, the exact Qwen3.5-4B revision/backend, and the frozen mechanics Git blobs. The lock records zero prior model requests, sampled outputs, or protected reads. Calibration remained sealed until this lock commit was pushed and green in both exact workflows.
After lock commit e21aa1bb passed exact workflows 29336526045 and 29336525980, the sealed calibration completed and authenticated all 192 boundary pairs / 384 answer requests. Decision SHA-256 is 7a52bb4f097bffd220ac9d54f22a07e474b1600f822004c05a66645b896166eb:
- tokenizer-EOS no-think
PROGRAM:: 48/48 exact and parse; - tokenizer-EOS no-think freeform: 48/48 exact and parse;
- tokenizer-EOS think512
PROGRAM:: 34/48; - tokenizer-EOS think512 freeform: 26/48;
- all four matched HF-model-EOS cells: 0/48.
The fixed tokenizer-priority rule selected tokenizer_eos_no_think_program_slot. This is interface qualification, not a capability claim. Calibration read no mechanics, hidden, or benchmark files. Calibration decision commit 3b5bcdfc passed exact workflows 29337076993 and 29337076973. The separately reviewed static mechanics launcher then minted the winner-bound mechanics lock with SHA-256 5ce13cfec79fca9faa91a3edfc7854751474774abce5f76ca3cd51732d780e2f. It binds reviewed implementation 50fd804b, release commit 3b5bcdfc, the authenticated calibration decision, frozen mechanics code, exact five-arm sampling, and tokenizer_eos_no_think_program_slot. It records zero mechanics generation requests or sampled outputs before lock. No mechanics model call is authorized until this lock is committed, pushed, and green in both exact workflows.
Mechanics lock commit d02ebf56 passed exact workflows 29337577321 and 29337577892. After a fail-closed wait for the static launcher's anonymous GitHub verifier quota to reset, the sealed run used only the pinned model and backend. Transport passed 24/24 exact and parse with zero cap contacts. All 4,056 outputs then completed and the five-stage transaction chain authenticated. Visible selection SHA-256 is c64dd163c617216a50c03a0e6b536057088cb7bbd7042e69b04649e7e3c965ed.
All generation arms passed the ABI gate with zero cap contacts. Parse rates were 2,299/2,304 direct, 571/576 materialized, 575/576 name-only, and 569/576 shuffled. The sampled-token direct match used 16--17 candidates per task and the logical-model-token match used 30--34; no direct pool exhausted. The frozen selector found no visible-passing candidate for materialized, name-only, shuffled, or either matched-direct comparator on any of 24 tasks, so all five deployable arms abstain on all tasks. The non-primary 96-sample full-direct diagnostic found two visible-passing candidates on one task and selected one. The visible receipt records no hidden or benchmark read. It must be committed, pushed, and green before the hidden key may be opened once.
Visible checkpoint commit 5bce0231 passed exact workflows 29341369041 and 29341369091. The separately authorized one-time hidden stage then returned TOKENIZER_EOS_MATERIALIZED_RESIDUAL_LARGE_EFFECT_PILOT_FAIL. Materialized, name-only, shuffled, sampled-token-matched direct, and logical-model-token-matched direct each scored 0/24 selected success and 0/24 oracle proposal coverage; every gain was 0.0 and materialized first-operation support was empty. This is not task impossibility: exhaustive CPU evaluation of all 13,824 programs found 1--9 visible-consistent, hidden-correct programs per task, 88 total, for 24/24 coverage. Hidden-result SHA-256 is 71090626bcea3f8fb0ef1d3b802f2535107bfba6419a0056c3fae7bb5f1245a6. The authorized stage read only the fresh ciphertext/key, read no benchmark, and made no model request. The visible publication authorization binds file bytes c64dd163...c965ed; the hidden scorer separately records canonical-object digest a62bd73e...04a03. They are intentionally different hash domains.
Interpretation
The tokenizer-EOS answer boundary is a real and highly reliable output interface, but it does not make residual synthesis available. Materializing every first operation's public consequences changed prompt content and compute without moving the generated proposal support: the treatment and all frozen controls missed every hidden-correct program, despite exhaustive proof that each task had multiple reachable solutions. The bottleneck is proposal formation, not parsing, termination, selector quality, direct-pool exhaustion, or task unsatisfiability.
This clean pilot retires this exact all-candidate semantic-materialization prompt branch. It does not prove that every residual representation or Jacobian-derived intervention must fail. A successor must alter the mechanism, not merely sample more, loosen the selector, repeat the tokenizer boundary, or raise the same prompt's budget. In particular, a J-space successor should first show that a pre-answer coordinate predicts executable correctness across held- out tasks beyond ordinary margins and equal-width residual controls, then show that an intervention at the same prefix causally increases correct-proposal coverage over matched sampling. A readable coordinate without forward causal proposal lift is diagnostic only.
Knowledgebase Update
- Program evidence updated: clean null recorded; semantic materialization retired as a proposal-generation lever under this interface.
- Program backlog updated: next work must change the mechanism and require forward causal proposal lift, not selector repair or more of the same prompt.
- Claim ledger updated: no claim ID allocated.
Artifacts
src/scripts/configs/data/runs/analysis/reports/reports/artifact_manifest.yamlreports/preregistration.mdreports/design_review.md
Report
Rendered from reports/report.md
Summary
The fresh replay is a clean large-effect pilot fail. Construction, calibration, transport, generation ABI, transaction replay, visible selection, and the hidden boundary all passed. Materialized residuals and every matched comparator then had 0/24 selected success and 0/24 oracle proposal coverage, while exhaustive CPU search covered 24/24 tasks. The failed mechanism is proposal formation, not the answer interface, selector, or task generator.
Research Program Fit
Method
The constructor consumes one authenticated hash-only parent manifest. It rejects parent common-function/public fingerprints before assignment, creates fresh tasks and requests, hashes identity-free prompts and pinned-tokenizer renderings, and seals mechanics gold once with a new ignored AES-256-GCM key. Completed reruns authenticate tracked outputs without reading that key.
Results
- Zero overlap across 72 functions, 72 public instances, 72 task IDs, 2,952 request/seed-key identities, 5,904 derived runner seeds, 1,824 prompts, and 3,648 rendered token sequences.
- Exact 48 calibration / 24 mechanics task geometry and frozen strata.
CONSTRUCTION_PASS; tokenizer receiptTOKENIZER_GRAMMAR_PROMPT_FRESHNESS_PASSover 28,800 registered grammar rows.- Zero model calls, sampled outputs, protected reads, benchmark reads, and parent sampled-bundle reads.
- The repaired implementation candidate passes 147/147 model-free tests, including an unmocked complete production lifecycle through visible selection and exact rejection of initial transport replay after descendants exist. It also rejects nested Boolean/integer aliases, descendant creation during initial replay, and chain mutation during historical replay.
- The parent collision exporter is constrained by an exact eight-file read firewall. Its administrative receipt was migrated without changing any scientific collision domain or reading the key, sampled outputs, parent raw bundles, or benchmarks.
- Exact commit
98e9e9f6remains underHOLD_IMPLEMENTATION; the repaired bytes later received a separate exact-commit independent review. - The calibration review report is now pre-scaffolded as a tracked, non-authorizing HOLD artifact because the lock's reviewed critical inventory requires that path to exist in the implementation commit.
- Fresh three-round review of exact pushed-green commit
50fd804breturnedPASS_IMPLEMENTATION: 145/145 protected-safe tests and 48/48 predecessor mutation probes passed with zero model/GPU/protected access. Machine review receipts were subsequently published as a separate checkpoint. - The authenticated static launcher minted calibration-only implementation lock
b2204667...b733e, binding the reviewed bytes, receipt release, exact inputs/runtime, and frozen mechanics blobs with zero prior model/protected access. Exact lock-commit CI passed before calibration. - After exact lock CI, calibration authenticated all 192 boundary pairs and selected
tokenizer_eos_no_think_program_slot: both no-think tokenizer-EOS forms were 48/48, think512 forms were 34/48 and 26/48, and all HF-model-EOS controls were 0/48. This qualifies transport; it is not capability evidence. - After calibration decision commit
3b5bcdfcpassed both exact workflows, the reviewed launcher minted mechanics lock5ce13cfe...780e2f. It binds the qualified interface and five frozen mechanics arms and records zero mechanics requests or sampled outputs. Mechanics stays unauthorized until the lock commit itself passes exact CI. - Mechanics lock commit
d02ebf56passed exact CI before the sealed run. Transport was 24/24 exact and parse with zero cap contacts. The complete chain authenticated 4,056 outputs: 2,304 direct, 576 per suffix arm, and 24 transport. - Generation passed the ABI gate: parse was 99.78% direct, 99.13% materialized, 99.83% name-only, and 98.78% shuffled, with zero cap contacts. Matched direct pools used 16--17 sampled-token rows and 30--34 logical-token rows per task without exhaustion.
- No deployable arm produced a candidate that passed all eight visible rows on any task. Materialized, name-only, shuffled, sampled-matched direct, and logical-matched direct therefore each froze 0/24 selections. Only the non-primary full 96-sample direct diagnostic selected on one task. Visible selection SHA-256 is
c64dd163...c965ed; it records no hidden or benchmark read. - Visible checkpoint
5bce0231passed exact workflows29341369041and29341369091before hidden access. The one-time hidden stage returnedTOKENIZER_EOS_MATERIALIZED_RESIDUAL_LARGE_EFFECT_PILOT_FAILwith 0/24 selected success and 0/24 oracle proposal coverage for materialized, name-only, shuffled, sampled-matched direct, and logical-matched direct. Every materialized gain was 0.0 and first-operation support was empty. - Exhaustive evaluation of all 13,824 programs found 88 visible-consistent, hidden-correct programs: 1--9 per task and 24/24 task coverage. The scorer therefore had valid solutions for every task; the sampled proposal pools did not reach any. Hidden-result SHA-256 is
71090626...f1245a6. Visible publication authorization binds indented file bytesc64dd163...c965ed, whereas hidden scoring records canonical-object digesta62bd73e...04a03; these are distinct, consistent hash domains.
Controls
- Name-only and token-preserving shuffled suffix prompts distinguish semantic consequences from candidate naming and prompt length/order.
- Direct full-program pools are matched taskwise at the first sampled-token and logical-model-token overtake points; neither pool exhausted.
- The non-primary 96-sample direct pool is diagnostic only and does not affect the gate.
- Exact tokenizer-EOS transport and strict ABI checks rule out parser or termination collapse. Exhaustive CPU search rules out unsatisfiable tasks.
Oracle Versus Deployable Evidence
Deployable selected accuracy is 0.0 for all five arms because every arm abstained on all 24 tasks. Oracle proposal coverage is independently also 0.0 for all five arms. Thus this is not the common selector-gap pattern in which a correct candidate exists but cannot be chosen. By contrast, exhaustive search has 1.0 coverage and is explicitly report-only; it demonstrates task solvability without being credited as model capability.
Interpretation
Tokenizer EOS plus no-think is a reliable short-output interface, not a latent residual solver. Exposing every candidate first-operation consequence did not move correct programs into the model's proposal support at matched compute. Because oracle coverage is zero, selector repair cannot rescue this variant.
The conclusion is local to this frozen depth-three DSL, prompt ABI, seed block, and one-sample-per-first-operation treatment. It does not rule out different residual representations, learned state interfaces, or causal Jacobian interventions. It does rule out repeating this semantic-materialization prompt with looser parsing, the same selector, or a larger matched budget as if that were a new mechanism.
Next Experiments
- Retire this all-candidate semantic-materialization prompt branch.
- For J-space work, preregister a measurement-first gate on fresh tasks: determine whether a within-
<think>Jacobian coordinate predicts eventual executable correctness across task-held-out splits beyond answer margin, token position, and equal-width non-J residual controls. - Only if that gate passes, intervene at the same prefix and require a forward increase in correct-proposal coverage over taskwise matched sampling. A readable coordinate without causal proposal lift remains diagnostic.
Artifact Manifest
The manifest explicitly quarantines the parent's sampled bundles and records the fresh ignored key as the only required external artifact.
Experiment log 10
Show the running log (10 entries)
Scaffold
Created as a new experiment scaffold.
- Registered as the fresh-identity successor to
qwen35_4b_tokenizer_eos_answer_commit_factorialafter its 24/24 transport pass ended in post-chain replay instrument failure. - Froze the scientific parent and prohibited import, inspection, or rescoring of its 4,056 sampled mechanics outputs.
- Drafted distinct initial-authorization versus historical-replay invariants, fresh seed/identity domains, parent-collision controls, and an unmocked full lifecycle test requirement. No model call is authorized.
- Independent review of exact pushed-green commit
035c3f8creturnedPASS_DESIGN_FOR_MODEL_FREE_CONSTRUCTION_ONLY. Construction must prove the parent default-deny boundary, every declared zero-intersection domain, fresh ciphertext/key without plaintext or reread, exact depth-three balance, and zero model/GPU/protected access. Pinned-tokenizer access is limited to its required rendered-token collision proof. All model calls remain sealed.
Fresh construction
- Corrected a pre-run collision-export defect: deduplicating parent prompts by causal ID omitted two suffix representations because all three controls intentionally share IDs. Regenerated the hash-only manifest over every row, increasing complete prompt/token domains from 672/1,344 to 1,824/3,648.
- Replaced direct parent payload reads with one authenticated hash-only manifest and enforced exact schema/count/type validation.
- Excluded all 72 parent common-function fingerprints before split assignment, plus every public identity, request/seed identity, derived seed, prompt, and rendered token sequence.
- Preserved exact 48/24 split geometry, minimum depth three, mechanics strata, and calibration alias balance.
- Created fresh ciphertext and an ignored key once. Partial construction is terminal; completed-state validation never opens or hashes the key.
- Construction, tokenizer grammar receipt, and protocol smoke passed with zero model calls, sampled outputs, hidden/benchmark reads, or parent raw reads.
Model-free implementation candidate
- Rebuilt fresh-path static calibration and mechanics launchers and bound their reproducible hashes in both bootstraps and tests.
- Split transport into
authorize_initial_transportandauthenticate_historical_transport. The former requires all descendants absent; the latter requires exact authentication of the complete descendant chain before replaying transport. - Made partial-generation restart authenticate the immutable transport-decision hash through descendant transaction receipts instead of re-running the initial later-absent gate.
- Replaced pre-live mechanics tests' dependency on a nonexistent fresh calibration outcome with an explicitly synthetic model-free fixture.
- Added a real five-transaction historical-prefix lifecycle and an unmocked production-stage lifecycle through visible selection. The complete suite is 145/145 with zero model/GPU/hidden/benchmark access.
Adversarial implementation HOLD and repair
- Three-round independent review of exact pushed-green commit
98e9e9f6returnedHOLD_IMPLEMENTATION. It reproduced nested1/truetransport aliases and introduced or mutated descendants during semantic replay without a post-analysis authentication. - A separate science review found the parent collision exporter called a helper that transitively read earlier-lineage task/request payloads beyond the seven declared administrative sources.
- Replaced that helper with direct reconstruction of the scientific parent's own task/request inventory under a repository read firewall. The exporter now declares exactly eight source files, including the transitively imported protocol module. Every scientific collision domain remains exactly equal to the old authenticated Git object.
- Migrated only the parent-manifest hash and its downstream construction, preoutcome, and tokenizer bindings. The outcome-blind repair receipt records old/new hashes and zero model, sampled-output, hidden, key, benchmark, or parent-raw reads.
- Added recursive exact-type transport comparison plus post-semantic reauthentication for both initial and historical replay. Regression probes now reject nested integer-to-Boolean mutation, descendant injection, and complete-chain mutation.
- The repaired model-free suite was 146/146. The prior HOLD remains controlling until a fresh independent review passes the new exact pushed-green commit; no lock, model, GPU, or hidden access is authorized.
- A dry audit of lock publication then found a successor-scaffolding omission: the calibration critical inventory required
reports/calibration_implementation_review.md, but the path did not exist in the candidate commit. Added a tracked non-authorizing HOLD placeholder and a test that requires it to exist, remain in the bootstrap inventory, and be Git-tracked. The suite is now 147/147; review of the superseded commit was stopped and must restart on the next pushed-green SHA.
Exact implementation PASS
- Exact commit
50fd804bce7222fcce19d79e6b695bbb78a15c04passed both required workflows (29334944189,29334944084). - Fresh independent review completed three rounds and returned
PASS_IMPLEMENTATION. Its protected-safe suite passed 145/145; the operator suite passed 147/147. It rejected 48/48 predecessor transaction mutations, authenticated initial and historical replay on both sides of semantic analysis, and verified restart/recovery without resampling. - Review accounting recorded zero model requests, sampled outputs reviewed, GPU calls, hidden reads, qualification reads, confirmation reads, and benchmark reads.
- The PASS authorizes only publication of machine-readable calibration and mechanics review receipts. Locks remain absent and all model calls remain sealed until the receipts and subsequent lock commits are separately pushed and green.
Calibration implementation lock
- After the review-receipt commit
21f964e6passed exact workflows29336107856and29336108419, invoked the static calibration launcher's model-freelockstage from a cleanorigin/mainworktree. - Minted calibration-only lock
b220466742071e3fa02a698925251132cd7bb05535a8ac54d0ce819ab90b733e. It binds implementation50fd804b, review receipt21f964e6, 31 critical files, exact calibration inputs and sampling, and the frozen mechanics Git blobs. - The lock records zero experimental requests, sampled outputs, hidden reads, qualification reads, confirmation reads, and benchmark reads. No model call is authorized until the lock commit itself is pushed and green.
Fresh calibration
- Lock commit
e21aa1bbpassed exact workflows29336526045and29336525980; GPU availability was 48,639 MiB free with no compute process. - Ran the sealed static calibration launcher with only
Qwen/Qwen3.5-4B@851bf6e8and the pinned vLLM 0.24.0+cu129 backend. The live preflight passed before 0 experimental requests / 0 sampled outputs and has SHA-25643ab69bb...e6d5d5. - All 192 boundary pairs and 384 answer requests authenticated. No-think tokenizer-EOS
PROGRAM:and freeform were each 48/48 exact/parse; think512PROGRAM:was 34/48 and think512 freeform 26/48. Every matched HF-model-EOS cell was 0/48. - The fixed priority selected
tokenizer_eos_no_think_program_slot, yieldingTOKENIZER_EOS_ONLY_INTERFACE_QUALIFIED. Decision SHA-256 is7a52bb4f097bffd220ac9d54f22a07e474b1600f822004c05a66645b896166eb. - Calibration read no mechanics, hidden, qualification, confirmation, or benchmark file. Mechanics remains sealed until this decision is committed, pushed, green, and bound by the separately reviewed mechanics lock.
Mechanics implementation lock
- Calibration decision commit
3b5bcdfcpassed exact workflows29337076993and29337076973before mechanics lock construction. - Invoked the reviewed static mechanics launcher's model-free
lockstage and minted SHA-2565ce13cfec79fca9faa91a3edfc7854751474774abce5f76ca3cd51732d780e2f. - The lock binds implementation
50fd804b, release commit3b5bcdfc, the authenticated calibration decision/lock, all frozen mechanics blobs, exact transport/direct/control sampling, and selected interfacetokenizer_eos_no_think_program_slot. - It records zero experimental mechanics generation requests and sampled outputs before lock. Mechanics remains sealed until this lock commit is pushed and green in both exact workflows.
Fresh mechanics and visible selection
- Mechanics lock commit
d02ebf56passed exact workflows29337577321and29337577892. A first launch on a concurrently advancedmainfailed closed because its Publish workflow was still running; two later attempts failed closed before model load because the static launcher's credential-free GitHub verifier pool was exhausted. No preflight or generation artifact was created by those attempts. The unchanged launcher ran after the recorded quota reset. - Loaded only
Qwen/Qwen3.5-4B@851bf6e8in bf16 on pinned vLLM 0.24.0+cu129. Live preflight SHA-256 is8d2a693c20e783e24c490a0e98c821e00d024f87cfeb071f0fec4048975bf879. - Fresh transport passed 24/24 exact and parse, 12/12 in both arity strata, with zero cap contacts. Transport decision SHA-256 is
0036aaab7f9af12d0fb4a94272ace7f02bb2642d6f3e9d9bcedb494a2687a65b. - Completed and authenticated 4,056 outputs: 24 transport, 2,304 direct, and 576 each materialized, name-only, and shuffled. The complete five-stage chain passed historical replay with no recovery generation.
- All generation arms passed the ABI gate with zero cap contacts. Parse was 2,299/2,304 direct, 571/576 materialized, 575/576 name-only, and 569/576 shuffled. Sampled-token direct matching required 16--17 rows per task; logical-token matching required 30--34. No direct pool exhausted.
- Visible selection froze with SHA-256
c64dd163c617216a50c03a0e6b536057088cb7bbd7042e69b04649e7e3c965ed. Materialized, name-only, shuffled, sampled-matched direct, and logical-matched direct each had zero visible-passing candidates and zero selections on all 24 tasks. The non-primary 96-sample direct diagnostic had two visible-passing candidates on one task and selected one. - Visible analysis read no hidden or benchmark file. Hidden scoring remains unauthorized until this exact visible checkpoint is committed, pushed, and green in both workflows.
Hidden scoring and terminal decision
- Visible checkpoint commit
5bce0231passed exact workflows29341369041and29341369091. Invoked the sealedscore-hiddenstage once; it authenticated that exact commit and visible-selection SHA before opening the fresh key. No generation call occurred. - Terminal decision is
TOKENIZER_EOS_MATERIALIZED_RESIDUAL_LARGE_EFFECT_PILOT_FAIL. Selected success was 0/24 for materialized, name-only, shuffled, sampled-token-matched direct, and logical-model-token-matched direct. - Oracle proposal coverage was also 0/24 for every arm; every materialized gain was 0.0 and materialized first-operation support was empty. The result therefore cannot be repaired by changing only the visible selector.
- Exhaustive CPU evaluation of all 13,824 programs found 1--9 visible-consistent, hidden-correct programs per task, 88 total, and covered 24/24 tasks. This separates proposal-distribution failure from task impossibility.
- Hidden-result SHA-256 is
71090626bcea3f8fb0ef1d3b802f2535107bfba6419a0056c3fae7bb5f1245a6. The authorized stage read only the fresh ciphertext/key, read no benchmark, and made no model request. Visible publication authorization binds file-byte SHA-256c64dd163...c965ed; hidden scoring separately records canonical- object digesta62bd73e...04a03.
Data files 4
Result tables and metrics copied from the experiment folder — preview inline or open the raw file.
runs/construction/summary.json18 kBruns/mechanics/hidden_result.json36 kBruns/protocol_smoke/summary.json3.4 kBruns/smoke/summary.json521 B
Reproduce
Smoke test
python3 -B experiments/qwen35_4b_tokenizer_eos_residual_mechanics_fresh_replay/scripts/run.py --smokeRun steps are documented inside the experiment folder (README and scripts).