Research log Small Model Experimentation
GitHub

Research programs

Durable lines of inquiry. Each program page carries its charter, the evidence gathered so far, and every experiment that advanced it.

Structured Execution and Compilers

Represent tasks as executable, typed, latent, or stateful programs instead of direct final answers.

212 experiments · 29 claims · 4 queued

Benchmark Generalization

Stress whether mechanisms transfer across substrates, families, lengths, distributions, and real tasks.

131 experiments · 13 claims · 6 queued

Evidence-Conditioned Selection

Turn candidate-pool coverage into deployable decisions under visible evidence.

107 experiments · 11 claims · 6 queued

Posttraining and Adaptation

Change small-model behavior through LoRA, DPO, distillation, GRPO, DAgger, and other update mechanisms.

103 experiments · 30 claims · 4 queued

Interpretability and Diagnostics

Measure why methods work or fail through attribution, probes, pressure audits, and controlled ablations.

91 experiments · 10 claims · 4 queued

Reliability and Safety

Improve precision, abstention, robustness, reproducibility, and safe artifact handling.

85 experiments · 7 claims · 10 queued

Agentic Breadth Installation

Install general agentic capability in Qwen3.5-4B via breadth-first expert iteration on a firewall-clean multi-family gym, arbitrated by the blackbox menagerie instrument.

64 experiments · 12 claims · 1 queued

Process Control and Tool Use

Train or evaluate small models as controllers over tools, verifiers, budgets, and intermediate actions.

58 experiments · 3 claims · 8 queued

Algorithmic Memory and Retrieval

Use libraries of verified algorithms, examples, traces, skills, or failures as reusable memory.

30 experiments · 1 claim · 4 queued

Test-Time Reasoning Budget

Study the native thinking-token budget as a first-class controllable test-time-compute axis for Qwen3.5-4B, which the corpus has universally disabled.

25 experiments · 12 claims · 2 queued

Active Evidence Acquisition

Choose or synthesize the next visible example, probe, test, or trace that collapses uncertainty.

16 experiments · 1 claim · 5 queued

Operator and Skill Inventories

Grow, search, shortlist, compose, and stress-test reusable operator and skill banks.

13 experiments · 1 claim · 3 queued

Collective Experimentation Infrastructure

Make the repository itself better at spawning, comparing, validating, and remembering many independent lines.

4 experiments · 2 claims · 8 queued