Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 6 additions & 1 deletion boot/init/src/boot.rs
Original file line number Diff line number Diff line change
Expand Up @@ -233,11 +233,16 @@ fn scan_serials(ids: &[&str], found: &mut [Option<String>]) {
format!("/sys/block/{name}/serial"),
format!("/sys/block/{name}/device/serial"),
];
let device = format!("/dev/{name}");
for path in paths {
let Ok(serial) = fs::read_to_string(&path) else {
continue;
};
record_serial(ids, found, serial.trim_end(), &format!("/dev/{name}"));
if serial.trim_end().is_empty() {
continue;
}
record_serial(ids, found, serial.trim_end(), &device);
break;
}
}
}
Expand Down
4 changes: 1 addition & 3 deletions boot/init/src/cfg.rs
Original file line number Diff line number Diff line change
Expand Up @@ -128,9 +128,7 @@ fn debug_token(val: &str) -> bool {
val.is_empty() || val == "1"
}

/// Kernel ip= fields: client:server:gw:netmask:hostname:device:autoconf[:dns0[:dns1]].
/// Shorthand forms (ip=dhcp, ip=off) and malformed params are ignored — the
/// baked DHCP .network fallback then covers the NIC, like the old hook did.
/// Kernel ip= fields: client:server:gw:netmask:hostname:device:autoconf[:dns0[:dns1]]; anything else is ignored.
fn parse_ip_param(val: &str) -> Option<IpParam> {
let f: Vec<&str> = val.split(':').collect();
if f.len() < 7 || f[0].is_empty() || f[5].is_empty() {
Expand Down
1 change: 0 additions & 1 deletion boot/init/src/main.rs
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,6 @@ fn main() {
boot::run()
}

/// Keeps `cargo test` runnable on non-Linux dev hosts (cfg logic tests).
#[cfg(not(target_os = "linux"))]
fn main() {
eprintln!("sandbox-init is Linux-only");
Expand Down
2 changes: 2 additions & 0 deletions docs/deploy.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,6 +86,7 @@ sandboxd reads one JSON file (`-config`, default
| `cocoon_bin` | `cocoon` | cocoon CLI binary |
| `restore_mode` | unset | clone and wake-restore memory mode: `copy`, `ondemand`, or `mmap`; use `mmap` for dense pools |
| `no_direct_io` | false | use buffered writable disks for Cloud Hypervisor cold boots and clones; recommended for dense ephemeral pools to avoid direct-I/O CoW journal contention |
| `no_balloon` | false | boot pool and template VMs without the virtio-balloon (cocoon otherwise returns 25% of guest memory to the host); clones inherit it from the golden. A guest that thrashes before deflate-on-OOM fires — a 16G build tier running a large typecheck — needs its whole memory |
| `advertise_addr` | = `listen` | the host:port clients reach this node at; returned as a claim's owner address and gossiped to peers. Must be routable when `listen` is a wildcard |
| `bridges` / `networks` | unset | egress-lane attachment: a list of host bridge devices, or a list of CNI conflist names. Mutually exclusive; with neither set the node serves only the no-network lane. A Linux bridge holds at most 1024 ports (kernel `BR_MAX_PORTS`), so an N-entry list raises the node's egress ceiling to N×1024 — VMs spread over the list by a stable hash of the VM name, so size it with headroom (the spread is statistical, not exact). `bridges` keeps the raw TAP-on-bridge attachment (taps in the root netns, no per-VM network namespace or CNI plugin execution); `networks` runs the CNI chain per VM. [Guarded egress](egress.md) needs `bridges` and rejects a CNI network at load |
| `volumes` | unset | node-local catalog of operator-managed dataset images: `[ {"name":"imagenet","path":"/srv/datasets/imagenet.img","directio":"off","tenants":["acme"]}, {"name":"scratch-db","path":"/srv/datasets/scratch.img","writable":true} ]`. Names match `^[a-z][a-z0-9_-]{0,19}$` and cannot start with `cocoon-`; paths are absolute; `directio` is `on`, `off`, or `auto` and defaults to `off` for both read-only and writable entries. `tenants` is an optional access list: empty means every authenticated scope, while every listed name must exist in the node's `tenants` config; root always has access. `writable` (default `false`) lets a claim request `mode: "rw"` on that entry — see [Dataset volumes](#dataset-volumes). The catalog is intentionally not part of the cluster digest |
Expand All @@ -104,6 +105,7 @@ sandboxd reads one JSON file (`-config`, default
| `checkpoint_ttl_hours` | 0 (keep forever) | ages out checkpoints older than this; the sweep runs hourly and at startup. Explicit deletes never wait for it. Must be nonzero and match fleet-wide when `checkpoint_peer_heal` is on — it is the expiry eligibility point for a healed replica a delete broadcast missed, after which its next successful hourly sweep removes it; persistent sweep failure extends retention until one succeeds, so it is not a hard ceiling |
| `checkpoint_peer_heal` | false | on a cluster, lets a node pull a checkpoint it lacks from a peer — found via a live probe, not gossip — rather than failing the branch; see [placement lifecycle](cluster.md#checkpoints-on-a-cluster). Three requirements, all enforced at config load: a nonempty `api_token` (the blob transfer between peers authenticates with it; without one the raw record stream would be open), `mesh.cluster_key` set (the pull presents the fleet `api_token` to an address learned from the peer probe, so the gossip layer carrying that address must itself be authenticated), and `checkpoint_ttl_hours` nonzero (a replica a delete broadcast missed becomes eligible for expiry after it, and its next successful hourly sweep removes it — so it is the finite eligibility point, not an exact ceiling). A shared checkpoint store (`checkpoint_store` kind `s3`) ignores this setting — every node already resolves every checkpoint directly, so there is nothing to heal |
| `warm_max` (pool entry) | 0 (static) | turns on the demand-adaptive watermark for that pool: the warm target rises from `warm` toward `warm_max` while claims arrive faster than the measured provision lead covers, and decays back over ~a minute of silence |
| `warmup` (pool entry) | unset | argv run in the golden VM after readiness and before its snapshot, so the files it touches are page-cache-resident in every clone — e.g. `["node", "-e", "0"]` on a Node flavor. It runs under the engine's 2-minute command timeout in silkd's base environment (`PATH`, `TERM`, and the node's proxy variables on the none lane); a non-zero exit or a timeout fails the golden build, so the pool stays unfilled until the config is fixed. Config-owned like `egress`: `PUT /v1/pools` rejects it, and a golden built with a different warmup is rebuilt |
| `max_claims` | 0 (unlimited) | node-wide cap on live claims; claim/fork/branch requests beyond it answer 429 with the pool state unharmed (on a cluster, normal warm-candidate placement applies, with volume claims limited to candidates holding every requested volume) |
| `audit_log` | false | append every relayed request frame's op + addressing fields (never payloads) to `<data_dir>/audit.jsonl`, size-rotated with one `.1` backup. Records are `{t, id, op}` plus whichever addressing fields the op carries (`argv`, `path`, `dest`, `from`, `to`, `url`, `session`, `port`), plus `decision` and `secret` (the ref name, never its value) on `egress` records; preview accesses record as op `preview`, one per request. A request frame whose first line exceeds 4 KiB is skipped, never truncated |
| `idle_hibernate_seconds` | 0 (off) | node-wide idle policy for unpooled claims (template/checkpoint claims): a none-lane claim with no data-plane connection for this long is hibernated; the next call that reaches the guest wakes it transparently. Per-pool `idle_hibernate_seconds` does the same for that pool's claims; pooled keys ignore the node-wide value, and egress pools reject it because they cannot resume safely. Opt in deliberately: a wake costs latency and the snapshot, so callers with their own idle logic must not pay twice |
Expand Down
25 changes: 24 additions & 1 deletion docs/sandboxd-api.md
Original file line number Diff line number Diff line change
Expand Up @@ -325,7 +325,8 @@ targets online — no restart, live claims untouched:
Pools omitted from the list are drained: their unclaimed warm VMs are
destroyed and the pool entry retires. `net`/`size` default like a claim's.
Answers the fresh `GET /v1/info` payload. 400 bad key, negative warm/idle,
`warm_max` below `warm`, or duplicate pool; 401 bad api token; 409 egress
`warm_max` below `warm`, duplicate pool, or a config-owned `egress`/`warmup`
field; 401 bad api token; 409 egress
pool on a node without an egress attachment.

## POST /v1/drain
Expand Down Expand Up @@ -515,6 +516,28 @@ the connection is a byte-for-byte relay to the guest's silkd (one silkd RPC
per connection — see [silkd](silkd.md)). 426 without the upgrade header, 404
unknown sandbox, 502 guest unreachable.

## POST /v1/sandboxes/{id}/exec

Auth: the sandbox's own token. A buffered exec for clients that cannot hold
an upgraded connection — plain JSON in, plain JSON out, so it multiplexes over
HTTP/2 through a TLS proxy:

```json
{"argv": ["node", "-v"], "cwd": "/work", "env": {"CI": "1"}, "timeout_seconds": 60}
```

→ `200 {"exit_code": 0, "stdout": "v22.23.2\n", "stderr": ""}`. The command
runs to completion (no stdin, no streaming, no detach — use the relay for
those); `timeout_seconds` 0 means no limit beyond the request itself. When
the node gives up on a started command — timeout, client gone, output cap —
it kills the child through silkd before answering, so nothing keeps running
behind a 504. Output comes back as JSON strings: bytes that are not valid
UTF-8 are replaced with U+FFFD, so binary output belongs on the relay. 400
empty `argv`, negative `timeout_seconds`, an unknown field, or a silkd
`bad_request`; 401 missing bearer token; 404 unknown sandbox or wrong token;
413 when stdout+stderr exceed 8 MiB; 502 guest unreachable or any other
silkd error. A hibernated sandbox wakes transparently like on the relay.

## GET /v1/sandboxes/{id}/owner

Auth: the sandbox's own token. Answers `{"owner_addr": "host:port"}` when
Expand Down
8 changes: 4 additions & 4 deletions e2e/cmd/smoke/main.go
Original file line number Diff line number Diff line change
Expand Up @@ -376,11 +376,11 @@ func smokeFork(ctx context.Context, sb *sandbox.Sandbox) error {
}
for i := range children {
if _, err := children[1-i].Stat(ctx, fmt.Sprintf("/work/child-%d.txt", i)); err == nil {
return fmt.Errorf("child %d write visible in sibling — shared disk?", i)
return fmt.Errorf("child %d write visible in sibling: shared disk", i)
}
}
if _, err := sb.Stat(ctx, "/work/child-0.txt"); err == nil {
return errors.New("child write visible in parent — shared disk?")
return errors.New("child write visible in parent: shared disk")
}
return nil
}
Expand Down Expand Up @@ -693,8 +693,8 @@ func want(got, exp string) error {
}

func isSilkdKind(err error, kind string) bool {
var er *wire.ErrorResp
return errors.As(err, &er) && er.Kind == kind
er, ok := errors.AsType[*wire.ErrorResp](err)
return ok && er.Kind == kind
}

func lspWrite(w io.Writer, body string) error {
Expand Down
4 changes: 2 additions & 2 deletions e2e/cmd/volumesmoke/main.go
Original file line number Diff line number Diff line change
Expand Up @@ -105,11 +105,11 @@ func run(addr, token, template, volume, rwVolume, probe string) error {
}

_, err = sbA.Exec(ctx, "touch", path.Join(mountA, ".sandboxd-write-probe"))
var exitErr *sandbox.ExitError
if err == nil {
return errors.New("write to read-only volume succeeded")
}
if !errors.As(err, &exitErr) || !strings.Contains(strings.ToLower(exitErr.Stderr), "read-only file system") {
exitErr, ok := errors.AsType[*sandbox.ExitError](err)
if !ok || !strings.Contains(strings.ToLower(exitErr.Stderr), "read-only file system") {
return fmt.Errorf("write failed without EROFS: %w", err)
}

Expand Down
4 changes: 3 additions & 1 deletion e2e/fakeengine_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ type fakeEngine struct {

func newFakeEngine(dir string) *fakeEngine {
return &fakeEngine{
real: engine.New("cocoon", nil, nil, false, ""),
real: engine.New("cocoon", nil, nil, false, false, ""),
dir: dir,
listeners: map[string]io.Closer{},
socks: map[string]string{},
Expand Down Expand Up @@ -110,6 +110,8 @@ func (f *fakeEngine) DialGuestPort(context.Context, string, uint16) (net.Conn, e

func (f *fakeEngine) InstallCACert(context.Context, string, []byte) error { return nil }

func (f *fakeEngine) Warmup(context.Context, string, []string) error { return nil }

func (f *fakeEngine) DiskAttach(_ context.Context, vmName string, spec engine.VolumeSpec) error {
f.mu.Lock()
defer f.mu.Unlock()
Expand Down
24 changes: 18 additions & 6 deletions mcp/e2e.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,9 @@

class McpClient:
def __init__(self, argv):
self.proc = subprocess.Popen(argv, stdin=subprocess.PIPE, stdout=subprocess.PIPE)
self.proc = subprocess.Popen(
argv, stdin=subprocess.PIPE, stdout=subprocess.PIPE
)
self.seq = 0

def call(self, method, params=None):
Expand Down Expand Up @@ -52,12 +54,18 @@ def main() -> int:
parser.add_argument("--template", default="rt:24.04")
args = parser.parse_args()

mcp = McpClient([args.bin, "-addr", args.addr, "-token", args.token, "-template", args.template])
mcp = McpClient(
[args.bin, "-addr", args.addr, "-token", args.token, "-template", args.template]
)
try:
init = mcp.call("initialize", {"protocolVersion": "2024-11-05", "capabilities": {}})
init = mcp.call(
"initialize", {"protocolVersion": "2024-11-05", "capabilities": {}}
)
assert init["serverInfo"]["name"] == "sandbox-mcp", init
tools = {t["name"] for t in mcp.call("tools/list")["tools"]}
assert {"create_sandbox", "exec", "checkpoint", "branch_checkpoint"} <= tools, tools
assert {"create_sandbox", "exec", "checkpoint", "branch_checkpoint"} <= tools, (
tools
)
print(f" initialize + tools/list ok ({len(tools)} tools)")

sandbox_id = mcp.tool("create_sandbox")["sandbox_id"]
Expand All @@ -67,11 +75,15 @@ def main() -> int:

mcp.tool("write_file", sandbox_id=sandbox_id, path="/root/m.txt", content="v1")
assert mcp.tool("read_file", sandbox_id=sandbox_id, path="/root/m.txt") == "v1"
names = {e["name"] for e in mcp.tool("list_dir", sandbox_id=sandbox_id, path="/root")}
names = {
e["name"] for e in mcp.tool("list_dir", sandbox_id=sandbox_id, path="/root")
}
assert "m.txt" in names, names
print(" files ok")

ckpt = mcp.tool("checkpoint", sandbox_id=sandbox_id, name="mcp-step")["checkpoint_id"]
ckpt = mcp.tool("checkpoint", sandbox_id=sandbox_id, name="mcp-step")[
"checkpoint_id"
]
mcp.tool("write_file", sandbox_id=sandbox_id, path="/root/m.txt", content="v2")
branch = mcp.tool("branch_checkpoint", checkpoint_id=ckpt)["sandbox_id"]
assert mcp.tool("read_file", sandbox_id=branch, path="/root/m.txt") == "v1"
Expand Down
36 changes: 18 additions & 18 deletions mcp/server.go
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,24 @@ const (
defaultToolTTL = time.Hour
)

type rpcRequest struct {
ID json.RawMessage `json:"id"`
Method string `json:"method"`
Params json.RawMessage `json:"params"`
}

type rpcResponse struct {
JSONRPC string `json:"jsonrpc"`
ID json.RawMessage `json:"id"`
Result any `json:"result,omitempty"`
Error *rpcErrorBody `json:"error,omitempty"`
}

type rpcErrorBody struct {
Code int `json:"code"`
Message string `json:"message"`
}

// server owns one sandboxd client and the handles minted over this stdio
// session: MCP tools address sandboxes and checkpoints by id, so the live
// handles (with their tokens) stay here.
Expand Down Expand Up @@ -179,24 +197,6 @@ func (s *server) dropCkpt(id string) {
delete(s.ckpts, id)
}

type rpcRequest struct {
ID json.RawMessage `json:"id"`
Method string `json:"method"`
Params json.RawMessage `json:"params"`
}

type rpcResponse struct {
JSONRPC string `json:"jsonrpc"`
ID json.RawMessage `json:"id"`
Result any `json:"result,omitempty"`
Error *rpcErrorBody `json:"error,omitempty"`
}

type rpcErrorBody struct {
Code int `json:"code"`
Message string `json:"message"`
}

func result(id json.RawMessage, v any) rpcResponse {
return rpcResponse{JSONRPC: "2.0", ID: id, Result: v}
}
Expand Down
10 changes: 5 additions & 5 deletions protocol/wire/frame.go
Original file line number Diff line number Diff line change
Expand Up @@ -182,7 +182,7 @@ func (Ps) Op() string { return "ps" }
// Kill signals a process; nil Signal means SIGKILL.
type Kill struct {
PID uint32 `json:"pid"`
Signal *int32 `json:"signal,omitempty"`
Signal *int32 `json:"signal,omitzero"`
}

func (Kill) Op() string { return "kill" }
Expand Down Expand Up @@ -238,7 +238,7 @@ func (StdinClose) Op() string { return "stdin_close" }
// defaults.
type FsWrite struct {
Path string `json:"path"`
Mode *uint32 `json:"mode,omitempty"`
Mode *uint32 `json:"mode,omitzero"`
}

func (FsWrite) Op() string { return "fs_write" }
Expand Down Expand Up @@ -366,7 +366,7 @@ type GitClone struct {
URL string `json:"url"`
Path string `json:"path"`
Branch string `json:"branch,omitempty"`
Depth uint32 `json:"depth,omitempty"`
Depth uint32 `json:"depth,omitzero"`
Auth string `json:"auth,omitempty"`
}

Expand Down Expand Up @@ -524,7 +524,7 @@ type ProcInfo struct {
Argv []string `json:"argv"`
Detached bool `json:"detached"`
State string `json:"state"`
ExitCode *int32 `json:"exit_code,omitempty"`
ExitCode *int32 `json:"exit_code,omitzero"`
StartedAtEpochSecs uint64 `json:"started_at_epoch_secs"`
}

Expand Down Expand Up @@ -630,7 +630,7 @@ type GitStatusResult struct {
Ahead uint32 `json:"ahead"`
Behind uint32 `json:"behind"`
Files []GitFileStatus `json:"files"`
Truncated bool `json:"truncated,omitempty"`
Truncated bool `json:"truncated,omitzero"`
}

func (GitStatusResult) RespType() string { return "git_status_result" }
Expand Down
1 change: 0 additions & 1 deletion sandboxd/ca.go
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,6 @@ import (
"github.com/cocoonstack/sandbox/sandboxd/egress"
)

// runCA is the operator PKI tool: mint the cluster root, then per-node intermediates.
func runCA(args []string) error {
if len(args) == 0 {
return fmt.Errorf("usage: sandboxd ca {init|issue-intermediate}")
Expand Down
Loading