Task packages are organized by implementation family and aligned with best_results/. The launcher auto-discovers anything at datasets/<family>/<subtask>/init_program.{py,cpp,rs,...}.
- Catalogue — what's here
- Designing a task — the file contract + worked example
| Task area | Family | Subtasks | Language | Setup |
|---|---|---|---|---|
| 🪐 Quantum circuit compilation | qubit_routing |
1 | Rust | Rust toolchain |
znaa |
1 | Python | family venv | |
| 🛰️ Astrodynamics | astrodynamics |
5 | Python | family venv + NAIF kernels |
| ⚡ GPU kernel optimization | gpukernel |
1 | CUDA / Triton | GPU + server |
| 🧮 Algorithm engineering | ahc |
2 | C++ in Docker | Docker |
numerical_tasks |
1 | C++ via Python | g++ + Eigen | |
| 📐 Mathematics — extremal analysis | erdos |
1 | Python | none |
autocorrelation |
3 | Python | none | |
| 🧩 Combinatorial construction | circle_packing |
2 | Python | none |
hadamard_maximal_det |
1 | Python | none | |
sums_diffs |
1 | Python | none | |
| 🧬 Data science | scaling_law |
4 | Python | HuggingFace cache |
zapbench |
1 | Python | GPU + task venv + dataset | |
open_problems_bio |
1 | Python | bundled venv + dataset |
First-time pick: any Setup: none row. For setup-heavy families:
uv run python scripts/prepare_task.py --list
uv run python scripts/prepare_task.py --check
uv run python scripts/prepare_task.py --task scaling_lawEach family has its own README.md with task-specific assumptions and a run command.
One task = one directory with three files: the instruction states the constraints, the seed implements the entry function, the evaluator checks constraints and scores.
datasets/<family>/<task>/
├── init_program.{py|cpp|rs|...} required — seed, with EVOLVE-BLOCK markers
├── evaluator.py required — scores a candidate
└── <task>.txt required — problem statement for the LLM
datasets/<family>/
├── requirements.txt optional — packages allowed in the evolved code
├── pyproject.toml + venv/ optional — family-local Python env (auto-detected)
├── data_manifest.json optional — data this family needs
└── README.md optional — family-level notes
The region between EVOLVE-BLOCK-START and EVOLVE-BLOCK-END is the part that gets evolved. Everything else is the fixed harness — imports, the entry function, anything the evaluator depends on.
# EVOLVE-BLOCK-START
import numpy as np
def construct_circles():
... # ← the LLM edits this
# EVOLVE-BLOCK-END
def run_code(): # fixed entry point the evaluator calls
return construct_circles()Requirements: markers paired; entry function in the fixed region; the seed runs and earns a finite, non-zero score. C++ / Rust seeds use the same markers via the host language's comment syntax.
Runs the candidate in an isolated subprocess, checks constraints, recomputes the score. The framework reads only combined_score from the returned dict.
Six structural parts. Only parts 3 and 4 are task-specific; copy an existing evaluator and edit those two.
| # | Part | Per-task? |
|---|---|---|
| 1 | Configuration: TIMEOUT_SECONDS, concurrency, memory fraction. |
shared |
| 2 | Exception classes: EvaluatorTimeoutError, MemoryLimitExceededError. |
shared |
| 3 | validate_solution(sol) — every hard constraint with tolerances. |
per task |
| 4 | compute_score(sol) — recompute the score from the solution. |
per task |
| 5 | run_with_timeout(path, …) — subprocess + timeout + memory cap + BLAS thread cap. |
shared |
| 6 | evaluate(path) — entry; must never raise. |
shared |
def evaluate(path: str) -> dict:
# success
return {"combined_score": 12.34, "validity": 1.0, "eval_time": 8.7}
# failure
return {"combined_score": 0.0, "validity": 0.0, "error": "Timeout: ..."}Design rules:
- Higher-is-better score. Minimising
q? Use1 / (eps + q)(autocorrelation, erdos) or a reference ratio (hadamard). - Score must discriminate. A 0/1 verdict degenerates into random search; aggregate over cases (ahc, kernelbench) or normalise (kissing-number-style).
- Hack-proof: recompute the score from the solution; never trust a self-reported value.
evaluate()never raises. Wrap intry/exceptand returncombined_score: 0on failure.
Shown verbatim to the model. Cover: the problem, every constraint, the objective (max/min + how combined_score is computed), resource limits (especially per-evaluation timeout in seconds), any reshaping / discretisation, and the program interface.
Always include the sentence
Do this by evolving the code between
# EVOLVE-BLOCK-STARTand# EVOLVE-BLOCK-END.
so the code extractor knows the evolved region.
| File | Purpose |
|---|---|
requirements.txt |
Not installed by SimpleTES; package names go into the prompt so the LLM does not hallucinate dependencies. |
venv/ (or <family>/venv/) |
Task-local Python env, auto-detected. Override with --eval-venv <path>. |
<family>/pyproject.toml + uv.lock |
Reproducible family-level env. |
data_manifest.json |
Files the task needs; scripts/prepare_task.py --task <family> runs declared commands, --check verifies. |
Minimal data_manifest.json:
{
"prepare_commands": [
{"command": ["bash", "setup.sh"], "cwd": ".", "description": "Build local deps"}
],
"required_files": ["my_deps/built_artifact"]
}Don't commit large data; declare it.
End-to-end square-root task.
datasets/my_demo/square_root/
├── init_program.py
├── evaluator.py
└── square_root.txt
init_program.py:
# EVOLVE-BLOCK-START
def sqrt(x: float) -> float:
return x / 2 # bad baseline; the LLM will fix it
# EVOLVE-BLOCK-END
def run_code():
return sqrtevaluator.py:
import importlib.util, math, sys, time
TIMEOUT_SECONDS = 30
def evaluate(filepath: str) -> dict:
try:
t0 = time.time()
spec = importlib.util.spec_from_file_location("cand", filepath)
mod = importlib.util.module_from_spec(spec)
sys.modules["cand"] = mod
spec.loader.exec_module(mod)
sqrt = mod.run_code()
targets = [0.25, 1.0, 2.0, 9.0, 1024.0, 1e-6]
errors = [abs(sqrt(x) - math.sqrt(x)) for x in targets]
return {
"combined_score": -sum(errors), # higher = better → negate error
"validity": 1.0,
"eval_time": time.time() - t0,
"max_error": max(errors),
}
except Exception as e:
return {"combined_score": 0.0, "validity": 0.0, "eval_time": 0.0, "error": str(e)}square_root.txt:
Improve sqrt(x) so it approximates math.sqrt as closely as possible on the
listed targets. You may use the math module but not math.sqrt.
Do this by evolving the code between # EVOLVE-BLOCK-START and # EVOLVE-BLOCK-END.
The time limit for each program evaluation is 30 seconds.
Run:
uv run python main.py \
--init-program datasets/my_demo/square_root/init_program.py \
--evaluator datasets/my_demo/square_root/evaluator.py \
--instruction datasets/my_demo/square_root/square_root.txt \
--model gemini/gemini-2.0-flash \
--max-generations 30-
<task>.txt,run_code(), andvalidate_solutionagree on the interface. Timeout values in.txt,TIMEOUT_SECONDS, and--eval-timeoutall match. - EVOLVE-BLOCK markers paired; entry function in the fixed region; seed earns a finite, non-zero score locally.
- Evaluator has subprocess isolation, hard timeout, memory cap.
validate_solutioncovers every hard constraint with tolerances. Score is recomputed from the solution.evaluate()never raises. -
combined_scoreis higher-is-better and discriminates (not 0/1). -
main.py ... --max-generations 5completes locally. If the family has setup steps,scripts/prepare_task.py --checkis clean. - Family
README.mdcovers problem, scoring, and environment.