Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
40 commits
Select commit Hold shift + click to select a range
10cc649
wire: scan past a tag's siblings without recursing
CMGS Sep 14, 2026
3eac2a9
mcp: keep exec output when the call fails, cap it, and reject a negat…
CMGS Sep 14, 2026
0a485c2
sdk/go: release the relay when a watch ends on its own
CMGS Sep 14, 2026
0652d6e
sdk/python: enforce the exec cutoff inside the guest
CMGS Sep 14, 2026
3b42d68
config: keep an explicit warm 0 under warm_max; refuse a wildcard adv…
CMGS Sep 14, 2026
a7a1b55
server: stop a preview forward after one hop, keep encoded slashes, e…
CMGS Sep 14, 2026
a4aa227
review: take the pool-key digest off the warm claim path; small simpl…
CMGS Sep 14, 2026
965297f
silkd: bound the find read on the handle, fail a session without entr…
CMGS Sep 14, 2026
a6d0faf
review: silkd and boot-init simplifications
CMGS Sep 14, 2026
05b0bda
review: comment register and budget sweep; adapter method order
CMGS Sep 14, 2026
0bc4ded
pool: refuse capture verbs on an archived claim; keep rollbacks and e…
CMGS Sep 14, 2026
87ffe99
egress: keep the upstream connection for a closing guest; keep a clam…
CMGS Sep 14, 2026
098804d
server: serve a held claim under a stale owner name; omit an unspecif…
CMGS Sep 14, 2026
0cd543c
mcp: bound read_file to a regular file under the exec cap
CMGS Sep 14, 2026
017c97b
sdk/go: release the relay when a pty exits
CMGS Sep 14, 2026
c209a6d
sdk/python: a wall-clock timeout on run and exec
CMGS Sep 14, 2026
bd83542
silkd: bound the replace read through one handle
CMGS Sep 14, 2026
5ae1fd0
review: simplify-lens follow-ups in silkd and boot-init
CMGS Sep 14, 2026
5f0f600
docs: align the pages with the code
CMGS Sep 14, 2026
64b578c
sdk/go: stream a file into a writer; keep reading after a failed stdi…
CMGS Sep 14, 2026
dc453a8
mcp: bound read_file by the bytes it reads
CMGS Sep 14, 2026
7f3bdab
review: gofumpt the mcp tools table inside its var block
CMGS Sep 14, 2026
7659b82
pool: check for an archived claim under the transition lock
CMGS Sep 14, 2026
d66e083
sdk/langchain: refuse the command when the claim used the whole call …
CMGS Sep 14, 2026
f8f7d64
silkd: read only a regular file
CMGS Sep 14, 2026
235d7e4
mcp: refuse a non-regular file before streaming it
CMGS Sep 14, 2026
7d2b509
pool: read the VM name for the fork and checkpoint usage events under…
CMGS Sep 14, 2026
2a90c24
review: lint follow-ups
CMGS Sep 14, 2026
3cc28e7
silkd: check the file kind on the opened descriptor
CMGS Sep 14, 2026
120c83a
egress: track both halves of a tunnel so Close ends it
CMGS Sep 14, 2026
3fa209d
pool: drop the consumed snapshot when a release beats the wake to the…
CMGS Sep 14, 2026
e8bcac7
pool: evict a template's record lock with the record
CMGS Sep 14, 2026
1f275a0
docs: align the remaining pages with the code
CMGS Sep 14, 2026
5a82908
docs: tighten four sentences from the last review round
CMGS Sep 14, 2026
e85823c
pool: take the release re-check off the claim path
CMGS Sep 14, 2026
4b8db0b
review: drop the no-op Fetch release and the dead pty exit guard
CMGS Sep 14, 2026
e6fc001
pool: write the archive delete marker outside the manager mutex
CMGS Sep 14, 2026
ef9a3f6
docs: a bare egress rule admits the SOCKS5 door
CMGS Sep 14, 2026
b34ec98
sdk/python: enforce run deadline through upgrade and send
CMGS Sep 14, 2026
1616e17
sdk/python: wrap the exec send call
CMGS Sep 14, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,6 +81,7 @@ make help # this list
make lint test # Rust: boot/init + silkd (fmt --check, clippy -D warnings, tests)
make go-lint # Go: protocol/wire + sandboxd + sdk/go + e2e + mcp, GOOS linux AND darwin
make go-test # Go: go test -race across the Go modules
make sh-lint # shellcheck every tracked shell script
make sandboxd # build dist/sandboxd
make boot # kernel + initramfs artifact image (docker)
# KERNEL_MIRROR=… if kernel.org tarball paths 404 locally
Expand Down Expand Up @@ -116,7 +117,9 @@ TEMPLATE=rt:24.04 scripts/sandboxd-e2e.sh
## CI

- `silkd.yml` / `sandboxd.yml` — Rust and Go test+lint suites
- `boot-init.yml` — the boot/init crate's own fmt+clippy+test gate
- `python.yml` — ruff + pytest for the three Python packages
- `shell.yml` — shellcheck over every tracked shell script
- `images.yml` — the single image entry point: on a push touching
`boot/**`, `silkd/**`, `protocol/**`, or `os-image/**` it builds the
changed carriers (via `build-boot.yml` / `build-silkd.yml`,
Expand All @@ -133,7 +136,7 @@ On a fresh repo run build-boot first — images build `FROM` the boot artifact.

```
cloud-hypervisor
→ vmlinux (PVH ELF, everything =y, no decompress stage)
→ guest kernel (amd64: PVH ELF vmlinux, everything =y, no decompress stage; arm64: flat Image)
→ uncompressed ~1.5MB cpio: /init = sandbox-init (static Rust)
→ resolve virtio-blk serials via sysfs (2ms poll, no udev)
→ mount EROFS layers → overlayfs + ext4 COW → switch_root
Expand Down
11 changes: 5 additions & 6 deletions boot/init/src/boot.rs
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,7 @@ pub fn run() -> ! {
let _ = sys::mount("proc", "/proc", Some("proc"), sys::MNT_SECURE, None);
let _ = sys::mount("sysfs", "/sys", Some("sysfs"), sys::MNT_SECURE, None);

// start marker, visible at production loglevel where the kernel's own boot lines are suppressed.
// printed at the production loglevel, which suppresses the kernel's own boot lines.
println!("sandbox-init: start at {}s", uptime());

let cmdline = fs::read_to_string("/proc/cmdline").unwrap_or_default();
Expand Down Expand Up @@ -193,9 +193,9 @@ fn poll_slots(
}

fn read_nic_mac(device: &str) -> Option<String> {
let mac = fs::read_to_string(format!("/sys/class/net/{device}/address")).ok()?;
let mac = mac.trim_end();
(!mac.is_empty() && mac != "00:00:00:00:00:00").then(|| mac.to_string())
let mut mac = fs::read_to_string(format!("/sys/class/net/{device}/address")).ok()?;
mac.truncate(mac.trim_end().len());
(!mac.is_empty() && mac != "00:00:00:00:00:00").then_some(mac)
}

fn mkdir_all(path: &str) -> Result<(), String> {
Expand Down Expand Up @@ -233,15 +233,14 @@ fn scan_serials(ids: &[&str], found: &mut [Option<String>]) {
format!("/sys/block/{name}/serial"),
format!("/sys/block/{name}/device/serial"),
];
let device = format!("/dev/{name}");
for path in paths {
let Ok(serial) = fs::read_to_string(&path) else {
continue;
};
if serial.trim_end().is_empty() {
continue;
}
record_serial(ids, found, serial.trim_end(), &device);
record_serial(ids, found, serial.trim_end(), &format!("/dev/{name}"));
break;
}
}
Expand Down
2 changes: 1 addition & 1 deletion boot/init/src/sys.rs
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ use std::io;
use std::ptr;
use std::time::Duration;

/// nosuid|nodev|noexec for the kernel API filesystems.
/// Mount flags for the kernel API filesystems.
pub const MNT_SECURE: libc::c_ulong = libc::MS_NOSUID | libc::MS_NODEV | libc::MS_NOEXEC;

pub fn mount(
Expand Down
6 changes: 4 additions & 2 deletions docs/benchmarks.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ the stop line.
| **warm pool hit** | ownership transfer of a pre-booted, probed VM; no VM lifecycle work on the request path |
| **clone from golden** | restore a full VM (memory + disk) from a golden snapshot, reseed entropy/machine identity, re-probe readiness |
| **cold boot** | boot from the template image: kernel + initramfs + rootfs assembly + init to a probed silkd |
| **burst** | `BURST_N` clone-tier claims issued concurrently — per-claim latency under restore contention plus the batch wall clock. Runs last so its churn cannot contaminate the RTT and throughput windows |
| **burst** | `BURST_N` clone-tier claims issued concurrently — per-claim latency under restore contention plus the batch wall clock. Runs after the RTT and throughput windows so its churn cannot contaminate them |

The harness also reports **warm refill recovery**: after fully draining the
warm pool it times the refill loop rebuilding to target — the number bounded
Expand Down Expand Up @@ -72,7 +72,9 @@ table stamped with the host evidence: virtualization
Knobs (environment variables): `WARM`/`WARM_N` (warm-pool depth and burst
size — the burst must stay within the depth, or refill loses the race and
the tail measures clones), `CLONE_N`, `COLD_N`, `RPC_N`, `PULL_MB`,
`PULL_N`, `BURST_N` (concurrent clone-claim burst; 0 skips the stage).
`PULL_N`, `BURST_N` (concurrent clone-claim burst; 0 skips the stage),
`ADDR`/`TOKEN` (the throwaway daemon's listen address and token; change them
to run beside a live sandboxd).

Boot anatomy (where inside the cold tier the milliseconds go — kernel,
initramfs phases, rootfs handoff) has its own harness:
Expand Down
5 changes: 3 additions & 2 deletions docs/browser.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,8 +66,9 @@ branch from.
`sandbox-id:port`, which Chrome's DevTools allowlist rejects. Use
`ProxyPort`/`DialPort` for CDP; preview URLs serve human-facing HTTP the
workload chooses to expose (a live-view page, a screenshot server).
- Guest env knobs on the unit: `CDP_PORT` (default 9222),
`CHROMIUM_FLAGS` (extra flags).
- Guest env knobs read by `/usr/local/bin/chromium-cdp`: `CDP_PORT` (default
9222) and `CHROMIUM_FLAGS` (extra flags); set them with a `chromium.service`
drop-in, the unit itself declares no environment.
- No stealth build: headless Chromium is fingerprintable; this flavor
targets automation, not anti-bot evasion.
- One browser per sandbox by design — the VM is the isolation and
Expand Down
6 changes: 4 additions & 2 deletions docs/cluster.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,8 @@
A cluster is a set of sandboxd nodes joined through a
[hashicorp/memberlist](https://github.com/hashicorp/memberlist) SWIM mesh.
Gossip carries only placement hints — per-pool warm counts, promoted-template
hashes, available volume names, and each node's data-plane address.
hashes, available volume names, each node's data-plane address, and a digest
of the cluster-invariant config.
Per-sandbox state never leaves its owning node, so a stale view costs at most
one extra redirect, never correctness. A single node with no seeds is a valid
mesh of one.
Expand Down Expand Up @@ -65,7 +66,8 @@ membership is node-local and deliberately excluded from the cluster config
digest. Nodes gossip only their currently available catalog names: host paths
and access lists never leave the node. After config load the set appears on the
next gossip tick; later image distribution or removal is detected the same way.
The node epoch bumps only when the advertised name set changes.
The node epoch bumps only when one of the gossiped sets — warm counts,
template hashes, volume names — actually changes.

A writable name (`writable: true`) still needs its catalog entry — name,
access list, and the `writable` flag — declared identically on every node
Expand Down
15 changes: 8 additions & 7 deletions docs/deploy.md
Original file line number Diff line number Diff line change
Expand Up @@ -99,7 +99,7 @@ sandboxd reads one JSON file (`-config`, default
| `restore_mode` | unset | clone and wake-restore memory mode: `copy`, `ondemand`, or `mmap`; use `mmap` for dense pools |
| `no_direct_io` | false | use buffered writable disks for Cloud Hypervisor cold boots and clones; recommended for dense ephemeral pools to avoid direct-I/O CoW journal contention |
| `no_balloon` | false | boot pool and template VMs without the virtio-balloon (cocoon otherwise returns 25% of guest memory to the host); clones inherit it from the golden. A guest that thrashes before deflate-on-OOM fires — a 16G build tier running a large typecheck — needs its whole memory |
| `advertise_addr` | = `listen` | the host:port clients reach this node at; returned as a claim's owner address and gossiped to peers. Must be routable when `listen` is a wildcard |
| `advertise_addr` | = `listen` | the host:port clients reach this node at; returned as a claim's owner address and gossiped to peers. Must be routable when `listen` is a wildcard; a node with `mesh` set refuses to load while it names an unspecified host |
| `bridges` / `networks` | unset | egress-lane attachment: a list of host bridge devices, or a list of CNI conflist names. Mutually exclusive; with neither set the node serves only the no-network lane. A Linux bridge holds at most 1024 ports (kernel `BR_MAX_PORTS`), so an N-entry list raises the node's egress ceiling to N×1024 — VMs spread over the list by a stable hash of the VM name, so size it with headroom (the spread is statistical, not exact). `bridges` keeps the raw TAP-on-bridge attachment (taps in the root netns, no per-VM network namespace or CNI plugin execution); `networks` runs the CNI chain per VM. [Guarded egress](egress.md) on the egress lane (an egress-lane pool policy or any tenant policy) needs `bridges` and rejects a CNI network at load; none-lane pool policies ride the proxy on either |
| `volumes` | unset | node-local catalog of operator-managed dataset images: `[ {"name":"imagenet","path":"/srv/datasets/imagenet.img","directio":"off","tenants":["acme"]}, {"name":"scratch-db","path":"/srv/datasets/scratch.img","writable":true} ]`. Names match `^[a-z][a-z0-9_-]{0,19}$` and cannot start with `cocoon-`; paths are absolute; `directio` is `on`, `off`, or `auto` and defaults to `off` for both read-only and writable entries. `tenants` is an optional access list: empty means every authenticated scope, while every listed name must exist in the node's `tenants` config; root always has access. `writable` (default `false`) lets a claim request `mode: "rw"` on that entry — see [Dataset volumes](#dataset-volumes). The catalog is intentionally not part of the cluster digest |
| `secrets` | unset | node-side credentials the egress proxy injects by name: `[{"name": "gh", "header": "Authorization", "value_env": "GH_TOKEN"}]`. A pool or tenant rule references the name; the value comes from the environment, never this file. See [egress](egress.md) |
Expand All @@ -118,14 +118,14 @@ sandboxd reads one JSON file (`-config`, default
| `checkpoint_ttl_hours` | 0 (keep forever) | ages out checkpoints older than this; the sweep runs hourly and at startup. Explicit deletes never wait for it. Must be nonzero and match fleet-wide when `checkpoint_peer_heal` is on — it is the expiry eligibility point for a healed replica a delete broadcast missed, after which its next successful hourly sweep removes it; persistent sweep failure extends retention until one succeeds, so it is not a hard ceiling |
| `checkpoint_peer_heal` | false | on a cluster, lets a node pull a checkpoint it lacks from a peer — found via a live probe, not gossip — rather than failing the branch; see [placement lifecycle](cluster.md#checkpoints-on-a-cluster). Three requirements, all enforced at config load: a nonempty `api_token` (the blob transfer between peers authenticates with it; without one the raw record stream would be open), `mesh.cluster_key` set (the pull presents the fleet `api_token` to an address learned from the peer probe, so the gossip layer carrying that address must itself be authenticated), and `checkpoint_ttl_hours` nonzero (a replica a delete broadcast missed becomes eligible for expiry after it, and its next successful hourly sweep removes it — so it is the finite eligibility point, not an exact ceiling). A shared checkpoint store (`checkpoint_store` kind `s3`) ignores this setting — every node already resolves every checkpoint directly, so there is nothing to heal |
| `warm_max` (pool entry) | 0 (static) | turns on the demand-adaptive watermark for that pool: the warm target rises from `warm` toward `warm_max` while claims arrive faster than the measured provision lead covers, and decays back over ~a minute of silence |
| `warmup` (pool entry) | unset | argv run in the golden VM after readiness and before its snapshot, so the files it touches are page-cache-resident in every clone, and again in every clone before it joins the warm pool, so those pages are already faulted into the restored VM when the first command runs — e.g. `["node", "-e", "0"]` on a Node flavor. It runs under the engine's 2-minute command timeout in silkd's base environment (`PATH`, `TERM`, and the guest image's proxy variables wherever nothing routes directly — the none lane and the locked bridge egress lane — with the relay not yet armed, since arming happens at claim); a non-zero exit or a timeout fails the golden build, so the pool stays unfilled until the config is fixed. Config-owned like `egress`: `PUT /v1/pools` rejects it, and a golden built with a different warmup is rebuilt |
| `warmup` (pool entry) | unset | argv run in the golden VM after readiness and before its snapshot, so the files it touches are page-cache-resident in every clone, and again in every clone before it joins the warm pool, so those pages are already faulted into the restored VM when the first command runs — e.g. `["node", "-e", "0"]` on a Node flavor. It runs under the engine's 2-minute command timeout in silkd's base environment (`PATH`, `TERM`, and the guest image's proxy variables wherever nothing routes directly — the none lane and the locked bridge egress lane — with the proxy not yet serving, since a door pre-bound at refill only starts serving at claim); a non-zero exit or a timeout fails the golden build, so the pool stays unfilled until the config is fixed. Config-owned like `egress`: `PUT /v1/pools` rejects it, and a golden built with a different warmup is rebuilt |
| `max_claims` | 0 (unlimited) | node-wide cap on live claims; claim/fork/branch requests beyond it answer 429 with the pool state unharmed (on a cluster, normal warm-candidate placement applies, with volume claims limited to candidates holding every requested volume) |
| `audit_log` | false | append every relayed request frame's op + addressing fields (never payloads) to `<data_dir>/audit.jsonl`, size-rotated with one `.1` backup. Records are `{t, id, op}` plus whichever addressing fields the op carries (`argv`, `path`, `dest`, `from`, `to`, `url`, `session`, `port`), plus `method` (`GET`, `CONNECT`, `SOCKS5`, …), `decision` and `secret` (the ref name, never its value) on `egress` records; preview accesses record as op `preview`, one per request. A request frame whose first line exceeds 4 KiB is skipped, never truncated |
| `audit_log` | false | append every relayed request frame's op + addressing fields (never payloads) to `<data_dir>/audit.jsonl`, size-rotated with one `.1` backup. Records are `{t, id, op}` plus whichever addressing fields the op carries (`argv`, `path`, `dest`, `from`, `to`, `url`, `session`, `port`), plus `method` (`GET`, `CONNECT`, `SOCKS5`, …), `decision` and `secret` (the ref name, never its value) on `egress` records; preview accesses record as op `preview`, one per request. A request frame whose first line exceeds 4 KiB records as op `oversized` with no addressing fields |
| `idle_hibernate_seconds` | 0 (off) | node-wide idle policy for unpooled claims (template/checkpoint claims): a none-lane claim is hibernated once it has had no open data-plane connection (relay, buffered exec, preview) and no egress request in flight for this long; the clock restarts when the last connection closes, so a long command is never cut short. The next call that reaches the guest wakes it transparently. Per-pool `idle_hibernate_seconds` does the same for that pool's claims; pooled keys ignore the node-wide value, and egress pools reject it because they cannot resume safely. Opt in deliberately: a wake costs latency and the snapshot, so callers with their own idle logic must not pay twice |
| `archive_after_seconds` | 0 (off) | tier below hibernation: a hibernated claim idle this long is checkpointed to the store and its local VM dropped, freeing the node entirely; the next call that reaches the guest restores it transparently (a checkpoint restore's latency) with a fresh server-default 5m lease. Requires `idle_hibernate_seconds > 0` and must exceed it. Node-wide for unpooled keys; per-pool overrides for that pool |
| `archive_delete_after_seconds` | 0 (keep) | purge an archived claim's store checkpoint this long after it was archived, reclaiming storage; the claim is then gone for good. On archive this retention window replaces the live claim deadline; 0 clears the deadline so the archive is kept forever. Same node-wide/per-pool split |
| `mesh` | unset | join a cluster ([Clusters](cluster.md)); unset = single node |
| `pools[]` | — | warm pools, keyed by `(template, net, size)`. `warm` defaults to 4; `net` is `none` or `egress`; `size` is a tier, below. Retune online without a restart via [`PUT /v1/pools`](sandboxd-api.md#put-v1pools) — omitted pools drain. This is the **first-boot seed**: once a node takes a `PUT /v1/pools`, the applied set persists to `<data_dir>/pools.json` and overrides this section on every later boot (a startup log notes it); delete `pools.json` to return to config-owned pools. Egress stays config-owned either way. See [state ownership](cluster.md#state-ownership) |
| `pools[]` | — | warm pools, keyed by `(template, net, size)`. `warm` defaults to 4 unless `warm_max` is set, so `warm: 0` under a `warm_max` starts the pool empty and grows it on demand; `net` is `none` or `egress`; `size` is a tier, below. Retune online without a restart via [`PUT /v1/pools`](sandboxd-api.md#put-v1pools) — omitted pools drain. This is the **first-boot seed**: once a node takes a `PUT /v1/pools`, the applied set persists to `<data_dir>/pools.json` and overrides this section on every later boot (a startup log notes it); delete `pools.json` to return to config-owned pools. Egress stays config-owned either way. See [state ownership](cluster.md#state-ownership) |

Size tiers (free-form CPU/memory is deliberately not accepted — it would
fragment the warm pools):
Expand Down Expand Up @@ -312,9 +312,10 @@ here validates on load:
"pools": [
{"template": "rt:24.04", "net": "none", "size": "small", "warm": 4, "warm_max": 12},
{"template": "rt:24.04", "net": "egress", "size": "medium", "warm": 2,
"egress": {"allow": [
"egress": {"socks5": true, "allow": [
{"host": "api.github.com", "methods": ["GET", "POST"], "secret": "gh", "intercept": true},
{"host": "*.googleapis.com"}
{"host": "*.googleapis.com"},
{"host": "imap.example.com", "ports": [993]}
]}}
]
}
Expand Down Expand Up @@ -463,7 +464,7 @@ exclusion, and a clean read-only claim afterward.
HTTP port under a signed, expiring shareable URL. The whole mechanism is in
sandboxd:

- **Minting** (`sb.PreviewURL(port, ttl)`): the owner node signs a token
- **Minting** (`sb.PreviewURL(ctx, port, ttl)`): the owner node signs a token
embedding `{sandbox, port, owner, exp}` with `preview_secret`; the URL's
life is clamped to the claim's lease.
- **Serving**: any node's preview listener verifies the token (no shared
Expand Down
Loading