NanoHorizon
LiveTrain the strongest Craftax policy in 30 minutes from a Qwen3.5-2B starter. Compare receipt-backed submissions, achievement frequencies, and replayable policy traces.
Open challenge →Synth Research
We study how language-model agents improve, operate across long horizons, and produce evidence that can be independently verified. We publish open challenges, evaluations, systems, reports, and field notes.
01
Open research problems with fixed contracts, public starters, held-out evaluation, and receipt-backed leaderboards.
Train the strongest Craftax policy in 30 minutes from a Qwen3.5-2B starter. Compare receipt-backed submissions, achievement frequencies, and replayable policy traces.
Open challenge →Hill-climb Banking77 exact-label classification with fixed development data, a sealed heldout, and reproducible GPT-OSS-20B training receipts.
Open challenge →02
Frozen benchmark contracts, comparable costs, and rollout-level evidence for understanding how models and agents actually behave.
Executable policies across frozen game environments, paired with source, replayable visual rollouts, and measured candidate results.
Open evaluation →Long-horizon survival, crafting, exploration, and combat with cost curves and replayable terminal traces.
Open evaluation →Fog-of-war coordination across single- and multi-agent parties, measured against inference cost.
Open evaluation →A 730-day operational benchmark with strict turn, budget, and output contracts.
Open evaluation →03
Methods for improving prompts, policies, and long-horizon agents through measured search and replay.
Synthesis in progress
Reflective prompt evolution, Pareto candidate coverage, and compute scaling across stable task contracts.
Active research
Archive-based optimization that returns to useful trajectory states before exploring the next frontier skill.
Report
Go-Explore Long-Horizon Optimizer (GELO) is a checkpoint-native optimizer for language agents in sparse, long-horizon environments. GELO maintains an archive of high-value trajectory states, opens scoped theme searches from those checkpoints, and promotes candidates only when replayed heldout evidence improves the archive. In Craftax, NetHack, and Crafter, GELO turns undifferentiated prompt mutation into targeted frontier search over skills such as furnace placement, coal and iron recovery, and dungeon coverage.
Jun 11, 2026
Report
Scaling train-time compute for GEPA across public task containers — same-container comparisons, coverage curves, and proposer scaling evidence.
Jun 2, 2026
04
Technical reports on the open and internal systems that make agent research reproducible, inspectable, and extensible.
Public building blocks from Synth Laboratories.
View the GitHub organization ↗Runnable GEPA, HealthBench eval, local MLX SFT on Craftax, and CISPO examples with measured receipts.
Explore the repository ↗An Apache-2.0 Rust optimization platform with a public GEPA engine and Python API, plus the SDK, CLI, and runbooks for hosted GELO jobs.
Explore the repository ↗An MIT-licensed Python SDK and HTTP task contract for datasets, mutable prompt programs, rollouts, scoring, traces, checkpoints, and resume.
Explore the repository ↗05
Working observations, negative results, reproductions, and early findings that are useful before they become formal reports.