Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions docs/deploy.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,6 +86,7 @@ sandboxd reads one JSON file (`-config`, default
| `cocoon_bin` | `cocoon` | cocoon CLI binary |
| `restore_mode` | unset | clone and wake-restore memory mode: `copy`, `ondemand`, or `mmap`; use `mmap` for dense pools |
| `no_direct_io` | false | use buffered writable disks for Cloud Hypervisor cold boots and clones; recommended for dense ephemeral pools to avoid direct-I/O CoW journal contention |
| `no_balloon` | false | boot pool and template VMs without the virtio-balloon (cocoon otherwise returns 25% of guest memory to the host); clones inherit it from the golden. A guest that thrashes before deflate-on-OOM fires — a 16G build tier running a large typecheck — needs its whole memory |
| `advertise_addr` | = `listen` | the host:port clients reach this node at; returned as a claim's owner address and gossiped to peers. Must be routable when `listen` is a wildcard |
| `bridges` / `networks` | unset | egress-lane attachment: a list of host bridge devices, or a list of CNI conflist names. Mutually exclusive; with neither set the node serves only the no-network lane. A Linux bridge holds at most 1024 ports (kernel `BR_MAX_PORTS`), so an N-entry list raises the node's egress ceiling to N×1024 — VMs spread over the list by a stable hash of the VM name, so size it with headroom (the spread is statistical, not exact). `bridges` keeps the raw TAP-on-bridge attachment (taps in the root netns, no per-VM network namespace or CNI plugin execution); `networks` runs the CNI chain per VM. [Guarded egress](egress.md) needs `bridges` and rejects a CNI network at load |
| `volumes` | unset | node-local catalog of operator-managed dataset images: `[ {"name":"imagenet","path":"/srv/datasets/imagenet.img","directio":"off","tenants":["acme"]}, {"name":"scratch-db","path":"/srv/datasets/scratch.img","writable":true} ]`. Names match `^[a-z][a-z0-9_-]{0,19}$` and cannot start with `cocoon-`; paths are absolute; `directio` is `on`, `off`, or `auto` and defaults to `off` for both read-only and writable entries. `tenants` is an optional access list: empty means every authenticated scope, while every listed name must exist in the node's `tenants` config; root always has access. `writable` (default `false`) lets a claim request `mode: "rw"` on that entry — see [Dataset volumes](#dataset-volumes). The catalog is intentionally not part of the cluster digest |
Expand All @@ -104,6 +105,7 @@ sandboxd reads one JSON file (`-config`, default
| `checkpoint_ttl_hours` | 0 (keep forever) | ages out checkpoints older than this; the sweep runs hourly and at startup. Explicit deletes never wait for it. Must be nonzero and match fleet-wide when `checkpoint_peer_heal` is on — it is the expiry eligibility point for a healed replica a delete broadcast missed, after which its next successful hourly sweep removes it; persistent sweep failure extends retention until one succeeds, so it is not a hard ceiling |
| `checkpoint_peer_heal` | false | on a cluster, lets a node pull a checkpoint it lacks from a peer — found via a live probe, not gossip — rather than failing the branch; see [placement lifecycle](cluster.md#checkpoints-on-a-cluster). Three requirements, all enforced at config load: a nonempty `api_token` (the blob transfer between peers authenticates with it; without one the raw record stream would be open), `mesh.cluster_key` set (the pull presents the fleet `api_token` to an address learned from the peer probe, so the gossip layer carrying that address must itself be authenticated), and `checkpoint_ttl_hours` nonzero (a replica a delete broadcast missed becomes eligible for expiry after it, and its next successful hourly sweep removes it — so it is the finite eligibility point, not an exact ceiling). A shared checkpoint store (`checkpoint_store` kind `s3`) ignores this setting — every node already resolves every checkpoint directly, so there is nothing to heal |
| `warm_max` (pool entry) | 0 (static) | turns on the demand-adaptive watermark for that pool: the warm target rises from `warm` toward `warm_max` while claims arrive faster than the measured provision lead covers, and decays back over ~a minute of silence |
| `warmup` (pool entry) | unset | argv run in the golden VM after readiness and before its snapshot, so the files it touches are page-cache-resident in every clone — e.g. `["node", "-e", "0"]` on a Node flavor. It runs under the engine's 2-minute command timeout with only `PATH` in its environment; a non-zero exit or a timeout fails the golden build, so the pool stays unfilled until the config is fixed. Config-owned like `egress`: `PUT /v1/pools` rejects it, and a golden built with a different warmup is rebuilt |
| `max_claims` | 0 (unlimited) | node-wide cap on live claims; claim/fork/branch requests beyond it answer 429 with the pool state unharmed (on a cluster, normal warm-candidate placement applies, with volume claims limited to candidates holding every requested volume) |
| `audit_log` | false | append every relayed request frame's op + addressing fields (never payloads) to `<data_dir>/audit.jsonl`, size-rotated with one `.1` backup. Records are `{t, id, op}` plus whichever addressing fields the op carries (`argv`, `path`, `dest`, `from`, `to`, `url`, `session`, `port`), plus `decision` and `secret` (the ref name, never its value) on `egress` records; preview accesses record as op `preview`, one per request. A request frame whose first line exceeds 4 KiB is skipped, never truncated |
| `idle_hibernate_seconds` | 0 (off) | node-wide idle policy for unpooled claims (template/checkpoint claims): a none-lane claim with no data-plane connection for this long is hibernated; the next call that reaches the guest wakes it transparently. Per-pool `idle_hibernate_seconds` does the same for that pool's claims; pooled keys ignore the node-wide value, and egress pools reject it because they cannot resume safely. Opt in deliberately: a wake costs latency and the snapshot, so callers with their own idle logic must not pay twice |
Expand Down
22 changes: 21 additions & 1 deletion docs/sandboxd-api.md
Original file line number Diff line number Diff line change
Expand Up @@ -325,7 +325,8 @@ targets online — no restart, live claims untouched:
Pools omitted from the list are drained: their unclaimed warm VMs are
destroyed and the pool entry retires. `net`/`size` default like a claim's.
Answers the fresh `GET /v1/info` payload. 400 bad key, negative warm/idle,
`warm_max` below `warm`, or duplicate pool; 401 bad api token; 409 egress
`warm_max` below `warm`, duplicate pool, or a config-owned `egress`/`warmup`
field; 401 bad api token; 409 egress
pool on a node without an egress attachment.

## POST /v1/drain
Expand Down Expand Up @@ -515,6 +516,25 @@ the connection is a byte-for-byte relay to the guest's silkd (one silkd RPC
per connection — see [silkd](silkd.md)). 426 without the upgrade header, 404
unknown sandbox, 502 guest unreachable.

## POST /v1/sandboxes/{id}/exec

Auth: the sandbox's own token. A buffered exec for clients that cannot hold
an upgraded connection — plain JSON in, plain JSON out, so it multiplexes over
HTTP/2 through a TLS proxy:

```json
{"argv": ["node", "-v"], "cwd": "/work", "env": {"CI": "1"}, "timeout_seconds": 60}
```

→ `200 {"exit_code": 0, "stdout": "v22.23.2\n", "stderr": ""}`. The command
runs to completion (no stdin, no streaming, no detach — use the relay for
those); `timeout_seconds` 0 means no limit beyond the request itself, and a
timeout closes the guest connection, which kills the command, then answers
504. 400 empty `argv`, negative `timeout_seconds`, or a silkd `bad_request`;
401 missing bearer token; 404 unknown sandbox or wrong token; 413 when
stdout+stderr exceed 8 MiB; 502 guest unreachable or any other silkd error.
A hibernated sandbox wakes transparently like on the relay.

## GET /v1/sandboxes/{id}/owner

Auth: the sandbox's own token. Answers `{"owner_addr": "host:port"}` when
Expand Down
4 changes: 3 additions & 1 deletion e2e/fakeengine_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ type fakeEngine struct {

func newFakeEngine(dir string) *fakeEngine {
return &fakeEngine{
real: engine.New("cocoon", nil, nil, false, ""),
real: engine.New("cocoon", nil, nil, false, false, ""),
dir: dir,
listeners: map[string]io.Closer{},
socks: map[string]string{},
Expand Down Expand Up @@ -110,6 +110,8 @@ func (f *fakeEngine) DialGuestPort(context.Context, string, uint16) (net.Conn, e

func (f *fakeEngine) InstallCACert(context.Context, string, []byte) error { return nil }

func (f *fakeEngine) Warmup(context.Context, string, []string) error { return nil }

func (f *fakeEngine) DiskAttach(_ context.Context, vmName string, spec engine.VolumeSpec) error {
f.mu.Lock()
defer f.mu.Unlock()
Expand Down
9 changes: 9 additions & 0 deletions sandboxd/config/config.go
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,9 @@ type PoolSpec struct {
// Egress is this pool's allow-list, intersected with the tenant's; nil denies all egress.
Egress *egress.Policy `json:"egress,omitempty"`

// Warmup runs in the golden VM before its snapshot, so every clone starts with its page cache.
Warmup []string `json:"warmup,omitempty"`

// IdleHibernateSeconds, when >0, hibernates idle claims after that many seconds.
IdleHibernateSeconds int `json:"idle_hibernate_seconds,omitempty"`

Expand All @@ -69,6 +72,9 @@ func (s PoolSpec) ValidateLimits() error {
if s.Net == types.NetEgress && s.IdleHibernateSeconds > 0 {
return fmt.Errorf("idle_hibernate_seconds is not supported for egress pools")
}
if slices.Contains(s.Warmup, "") {
return fmt.Errorf("warmup must not contain an empty argument")
}
return validateArchiveWindow(s.IdleHibernateSeconds, s.ArchiveAfterSeconds, s.ArchiveDeleteAfterSeconds)
}

Expand Down Expand Up @@ -171,6 +177,9 @@ type Config struct {
// NoDirectIO enables buffered writable disks for cold boots and clones.
NoDirectIO bool `json:"no_direct_io,omitempty"`

// NoBalloon boots VMs without the virtio-balloon, so a guest keeps its whole memory.
NoBalloon bool `json:"no_balloon,omitempty"`

// APIToken, when set, guards claim and info.
APIToken string `json:"api_token,omitempty"` //nolint:gosec // config field, not a hardcoded credential

Expand Down
29 changes: 22 additions & 7 deletions sandboxd/engine/cloneargs_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ func TestCloneArgsRestoreMode(t *testing.T) {
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
e := New("cocoon", []string{"br0"}, nil, false, tc.mode)
e := New("cocoon", []string{"br0"}, nil, false, false, tc.mode)
for _, args := range [][]string{
e.cloneArgs("/goldens/g1", "sbx-1", tc.key),
e.cloneSnapArgs("ck_1", "sbx-1", tc.key),
Expand All @@ -45,7 +45,7 @@ func TestLifecycleArgsApplyDirectIOPolicy(t *testing.T) {
key := types.PoolKey{Template: "rt:24.04", Net: types.NetNone, Size: types.SizeMedium}
for _, noDirectIO := range []bool{false, true} {
t.Run(strconv.FormatBool(noDirectIO), func(t *testing.T) {
e := New("cocoon", nil, nil, noDirectIO, "")
e := New("cocoon", nil, nil, noDirectIO, false, "")
want := "--no-direct-io=" + strconv.FormatBool(noDirectIO)
cold := e.runColdArgs("sbx-1", key)
for _, args := range [][]string{
Expand All @@ -71,8 +71,8 @@ func TestEgressVMsSpreadOverEveryConfiguredShard(t *testing.T) {
e *Engine
flag string
}{
"networks": {New("cocoon", nil, shards, false, ""), "--network"},
"bridges": {New("cocoon", shards, nil, false, ""), "--bridge"},
"networks": {New("cocoon", nil, shards, false, false, ""), "--network"},
"bridges": {New("cocoon", shards, nil, false, false, ""), "--bridge"},
} {
t.Run(name, func(t *testing.T) {
counts := map[string]int{}
Expand Down Expand Up @@ -107,13 +107,28 @@ func TestNetArgsHonorsTheLaneAndTheAttachment(t *testing.T) {
none := types.PoolKey{Template: "rt:24.04", Net: types.NetNone, Size: types.SizeMedium}
egress := types.PoolKey{Template: "rt:24.04", Net: types.NetEgress, Size: types.SizeMedium}

if args := New("cocoon", nil, []string{"cni"}, false, "").netArgs("sbx-1", none, false); len(args) != 0 {
if args := New("cocoon", nil, []string{"cni"}, false, false, "").netArgs("sbx-1", none, false); len(args) != 0 {
t.Errorf("none lane took an attachment: %v", args)
}
if args := New("cocoon", []string{"br0"}, nil, false, "").netArgs("sbx-1", egress, false); !slices.Equal(args, []string{"--bridge", "br0"}) {
if args := New("cocoon", []string{"br0"}, nil, false, false, "").netArgs("sbx-1", egress, false); !slices.Equal(args, []string{"--bridge", "br0"}) {
t.Errorf("bridge lane args = %v", args)
}
if args := New("cocoon", nil, []string{"cocoon-dhcp"}, false, "").netArgs("sbx-1", egress, false); !slices.Equal(args, []string{"--network", "cocoon-dhcp"}) {
if args := New("cocoon", nil, []string{"cocoon-dhcp"}, false, false, "").netArgs("sbx-1", egress, false); !slices.Equal(args, []string{"--network", "cocoon-dhcp"}) {
t.Errorf("single-network args = %v", args)
}
}

func TestNoBalloonReachesColdBootsOnly(t *testing.T) {
key := types.PoolKey{Template: "rt:24.04", Net: types.NetNone, Size: types.SizeSmall}
for _, noBalloon := range []bool{false, true} {
t.Run(strconv.FormatBool(noBalloon), func(t *testing.T) {
e := New("cocoon", nil, nil, false, noBalloon, "")
if got := slices.Contains(e.runColdArgs("sbx-1", key), "--no-balloon"); got != noBalloon {
t.Errorf("cold args carry --no-balloon = %v, want %v", got, noBalloon)
}
if slices.Contains(e.cloneArgs("/goldens/g1", "sbx-1", key), "--no-balloon") {
t.Error("clone args carry --no-balloon; clones inherit it from the golden")
}
})
}
}
8 changes: 6 additions & 2 deletions sandboxd/engine/engine.go
Original file line number Diff line number Diff line change
Expand Up @@ -65,12 +65,13 @@ type Engine struct {
bridges []string
networks []string
noDirectIO bool
noBalloon bool
restoreMode types.RestoreMode
}

// New returns a cocoon engine with node-wide network and disk policy.
func New(bin string, bridges, networks []string, noDirectIO bool, restoreMode types.RestoreMode) *Engine {
return &Engine{bin: bin, bridges: bridges, networks: networks, noDirectIO: noDirectIO, restoreMode: restoreMode}
func New(bin string, bridges, networks []string, noDirectIO, noBalloon bool, restoreMode types.RestoreMode) *Engine {
return &Engine{bin: bin, bridges: bridges, networks: networks, noDirectIO: noDirectIO, noBalloon: noBalloon, restoreMode: restoreMode}
}

// Version reports cocoon's version string: a "vX.Y.Z" release or a "master-<sha>" dev build.
Expand Down Expand Up @@ -338,6 +339,9 @@ func (e *Engine) restoreArgs() []string {
func (e *Engine) runColdArgs(name string, key types.PoolKey) []string {
spec, _ := key.Size.Spec()
args := []string{"vm", "run", argName, name, argOutput, formatJSON, "--cpu", strconv.Itoa(spec.CPU), "--memory", spec.Memory, e.directIOArg()}
if e.noBalloon {
args = append(args, "--no-balloon")
}
args = append(args, e.netArgs(name, key, true)...)
return append(args, key.Template)
}
Expand Down
16 changes: 8 additions & 8 deletions sandboxd/engine/engine_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ func TestDialSilkdConsumesOnlyHandshake(t *testing.T) {

listenMuxer(t, path, "OK 2048\nX")

conn, err := New("cocoon", nil, nil, false, "").DialSilkd(t.Context(), path)
conn, err := New("cocoon", nil, nil, false, false, "").DialSilkd(t.Context(), path)
if err != nil {
t.Fatalf("DialSilkd: %v", err)
}
Expand All @@ -41,7 +41,7 @@ func TestDialSilkdRejectedHandshake(t *testing.T) {
path := sockPath(t)
listenMuxer(t, path, "ERR no guest listener\n")

_, err := New("cocoon", nil, nil, false, "").DialSilkd(t.Context(), path)
_, err := New("cocoon", nil, nil, false, false, "").DialSilkd(t.Context(), path)
if err == nil || !strings.Contains(err.Error(), "no guest listener") {
t.Errorf("got %v, want handshake rejection", err)
}
Expand All @@ -51,7 +51,7 @@ func TestProbeSucceeds(t *testing.T) {
path := sockPath(t)
listenMuxer(t, path, "OK 2048\n", infoFrame)

if err := New("cocoon", nil, nil, false, "").Probe(t.Context(), path, 2*time.Second); err != nil {
if err := New("cocoon", nil, nil, false, false, "").Probe(t.Context(), path, 2*time.Second); err != nil {
t.Errorf("Probe: %v", err)
}
}
Expand All @@ -60,7 +60,7 @@ func TestInfoRoundTripRejectsErrorFrame(t *testing.T) {
path := sockPath(t)
listenMuxer(t, path, "OK 2048\n", errFrame)

err := New("cocoon", nil, nil, false, "").infoRoundTrip(t.Context(), path)
err := New("cocoon", nil, nil, false, false, "").infoRoundTrip(t.Context(), path)
if err == nil || !strings.Contains(err.Error(), `info reply type "error"`) {
t.Errorf("got %v, want error-frame rejection", err)
}
Expand All @@ -70,7 +70,7 @@ func TestProbeRetriesPastFailures(t *testing.T) {
path := sockPath(t)
listenMuxer(t, path, "OK 2048\n", errFrame, errFrame, infoFrame)

if err := New("cocoon", nil, nil, false, "").Probe(t.Context(), path, 2*time.Second); err != nil {
if err := New("cocoon", nil, nil, false, false, "").Probe(t.Context(), path, 2*time.Second); err != nil {
t.Errorf("Probe: %v", err)
}
}
Expand All @@ -79,7 +79,7 @@ func TestProbeRetriesUntilListenerAppears(t *testing.T) {
path := sockPath(t)
done := make(chan error, 1)
go func() {
done <- New("cocoon", nil, nil, false, "").Probe(t.Context(), path, 2*time.Second)
done <- New("cocoon", nil, nil, false, false, "").Probe(t.Context(), path, 2*time.Second)
}()

time.Sleep(60 * time.Millisecond)
Expand Down Expand Up @@ -117,15 +117,15 @@ func TestDialGuestPortCtxCancel(t *testing.T) {

ctx, cancel := context.WithTimeout(t.Context(), 150*time.Millisecond)
defer cancel()
if _, err := New("cocoon", nil, nil, false, "").DialGuestPort(ctx, path, 8080); !errors.Is(err, context.DeadlineExceeded) {
if _, err := New("cocoon", nil, nil, false, false, "").DialGuestPort(ctx, path, 8080); !errors.Is(err, context.DeadlineExceeded) {
t.Errorf("got %v, want context.DeadlineExceeded", err)
}
}

func TestProbeTimeout(t *testing.T) {
path := sockPath(t)

err := New("cocoon", nil, nil, false, "").Probe(t.Context(), path, 150*time.Millisecond)
err := New("cocoon", nil, nil, false, false, "").Probe(t.Context(), path, 150*time.Millisecond)
if err == nil || !strings.Contains(err.Error(), "silkd probe") {
t.Errorf("got %v, want probe timeout", err)
}
Expand Down
2 changes: 1 addition & 1 deletion sandboxd/engine/installca.go
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ const (
caCertGuestPath = "/usr/local/share/ca-certificates/sandbox-egress.crt"
caBundlePath = "/etc/ssl/certs/ca-certificates.crt"
// guestExecPATH is set because silkd starts the guest command with an empty environment.
guestExecPATH = "/usr/sbin:/usr/bin:/sbin:/bin"
guestExecPATH = "/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
)

// InstallCACert makes the guest trust the cluster root without update-ca-certificates.
Expand Down
6 changes: 3 additions & 3 deletions sandboxd/engine/installca_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ import (
func TestInstallCACertWritesCertAndUpdates(t *testing.T) {
path := sockPath(t)
fake := serveFakeSilkd(t, path)
if err := New("cocoon", nil, nil, false, "").InstallCACert(t.Context(), path, []byte("CERT-PEM")); err != nil {
if err := New("cocoon", nil, nil, false, false, "").InstallCACert(t.Context(), path, []byte("CERT-PEM")); err != nil {
t.Fatalf("InstallCACert: %v", err)
}
fake.mu.Lock()
Expand Down Expand Up @@ -48,7 +48,7 @@ func TestInstallCACertNonzeroExitFails(t *testing.T) {
path := sockPath(t)
fake := serveFakeSilkd(t, path)
fake.execCode = 3
err := New("cocoon", nil, nil, false, "").InstallCACert(t.Context(), path, []byte("x"))
err := New("cocoon", nil, nil, false, false, "").InstallCACert(t.Context(), path, []byte("x"))
if err == nil || !strings.Contains(err.Error(), "exit code 3") {
t.Errorf("got %v, want exit code 3 failure", err)
}
Expand All @@ -58,7 +58,7 @@ func TestInstallCACertWriteErrorFrameFails(t *testing.T) {
path := sockPath(t)
fake := serveFakeSilkd(t, path)
fake.writeErr = "disk full"
err := New("cocoon", nil, nil, false, "").InstallCACert(t.Context(), path, []byte("x"))
err := New("cocoon", nil, nil, false, false, "").InstallCACert(t.Context(), path, []byte("x"))
if err == nil || !strings.Contains(err.Error(), "disk full") {
t.Errorf("got %v, want fs_write error-frame failure", err)
}
Expand Down
Loading