Research programs
Durable lines of inquiry. Each program page carries its charter, the evidence gathered so far, and every experiment that advanced it.
Structured Execution and Compilers
Represent tasks as executable, typed, latent, or stateful programs instead of direct final answers.
212 experiments · 29 claims · 4 queued
Benchmark Generalization
Stress whether mechanisms transfer across substrates, families, lengths, distributions, and real tasks.
131 experiments · 13 claims · 6 queued
Evidence-Conditioned Selection
Turn candidate-pool coverage into deployable decisions under visible evidence.
107 experiments · 11 claims · 6 queued
Posttraining and Adaptation
Change small-model behavior through LoRA, DPO, distillation, GRPO, DAgger, and other update mechanisms.
103 experiments · 30 claims · 4 queued
Interpretability and Diagnostics
Measure why methods work or fail through attribution, probes, pressure audits, and controlled ablations.
91 experiments · 10 claims · 4 queued
Reliability and Safety
Improve precision, abstention, robustness, reproducibility, and safe artifact handling.
85 experiments · 7 claims · 10 queued
Agentic Breadth Installation
Install general agentic capability in Qwen3.5-4B via breadth-first expert iteration on a firewall-clean multi-family gym, arbitrated by the blackbox menagerie instrument.
64 experiments · 12 claims · 1 queued
Process Control and Tool Use
Train or evaluate small models as controllers over tools, verifiers, budgets, and intermediate actions.
58 experiments · 3 claims · 8 queued
Algorithmic Memory and Retrieval
Use libraries of verified algorithms, examples, traces, skills, or failures as reusable memory.
30 experiments · 1 claim · 4 queued
Test-Time Reasoning Budget
Study the native thinking-token budget as a first-class controllable test-time-compute axis for Qwen3.5-4B, which the corpus has universally disabled.
25 experiments · 12 claims · 2 queued
Active Evidence Acquisition
Choose or synthesize the next visible example, probe, test, or trace that collapses uncertainty.
16 experiments · 1 claim · 5 queued
Operator and Skill Inventories
Grow, search, shortlist, compose, and stress-test reusable operator and skill banks.
13 experiments · 1 claim · 3 queued
Collective Experimentation Infrastructure
Make the repository itself better at spawning, comparing, validating, and remembering many independent lines.
4 experiments · 2 claims · 8 queued