Skip to content

Safe parallel debugging

Codex edited this page Aug 22, 2026 · 1 revision

Safe parallel debugging

Safe parallel debugging is the clearest product-level demonstration of GenOS. It turns one failing program into several isolated futures, then promotes a fix only after the selected future can be reproduced from the original snapshot.

Safe parallel debugging demo

The scenario

The fixture contains a boundary bug. Instead of mutating the only working directory repeatedly, the demo declares three candidate mutations and evaluates each against the same five-test gate.

failing world W0
      │ snapshot
      ▼
snapshot S0
      │ fork
      ├─ candidate A ─ apply mutation ─ 4/5 tests ─ reject
      ├─ candidate B ─ apply mutation ─ 5/5 tests ─ candidate winner
      └─ candidate C ─ apply mutation ─ 3/5 tests ─ reject
                                      │
                                      ▼
                         restore S0 → replay B → compare → promote

Why snapshot before proposing fixes?

Every candidate must start from the same observable baseline. Without a snapshot, an earlier attempt might leave files, caches, generated artifacts, or state that contaminates later attempts.

The snapshot also becomes the source for replay: the system can prove that the selected mutation, reapplied to the original state, produces the same winning world.

Why isolated candidate worlds?

Each candidate is evaluated in a distinct directory-backed world. This prevents sibling file writes from leaking into one another and makes the test results comparable.

Isolation here is a controlled filesystem/world property. The deterministic fixture does not claim to sandbox arbitrary hostile processes or external network side effects.

Why use the same gate?

All candidates run the same five tests. The winner is therefore selected by declared evidence rather than by model rhetoric or branch order.

The gate records:

  • candidate identity and mutation;
  • test count and pass/fail result;
  • runtime observation;
  • replay status;
  • world equality check;
  • promotion decision;
  • model call, token, and model cost fields.

The fixture makes no model call, so model calls, tokens, and model cost are measured as zero. This proves runtime mechanics—not model quality.

Replay and exact comparison

After candidate B passes, GenOS does not simply copy its directory into the baseline. The workflow:

  1. restores the original snapshot;
  2. reapplies B's recorded mutation;
  3. runs the gate again;
  4. compares the replayed world with the candidate winner;
  5. promotes only when the equality and test checks pass.

This catches candidates that passed because of hidden or accidental branch state.

Evidence artifacts

Run:

./examples/safe-debugging-demo/run-demo.sh

Then inspect:

Artifact Purpose
artifacts/latest.json Machine-readable summary of candidates, winner, replay, and promotion
artifacts/events.jsonl Append-only sequence of snapshot, fork, mutation, test, replay, and promotion events
Studio evidence view Human-readable timeline and branch comparison

The repeated benchmark under benchmarks/safe-debugging/ runs the deterministic fixture multiple times and publishes raw samples. It is not a cross-framework benchmark.

From fixture to real agent use

With a model in the loop, the same control structure can wrap generated candidates:

LLM proposes branches
        ↓
GenOS snapshots and isolates execution
        ↓
tests/invariants evaluate outcomes
        ↓
recorded mutation and output are replayed where supported
        ↓
policy or human review authorizes promotion

Model generation remains nondeterministic; the surrounding baseline, branch identity, world changes, receipts, and promotion rule become explicit and inspectable.

Clone this wiki locally