diff --git a/DESIGN.md b/DESIGN.md new file mode 100644 index 0000000..0c74b2f --- /dev/null +++ b/DESIGN.md @@ -0,0 +1,413 @@ +# Canopy design + +Canopy is a Git hosting service embedded in the Cellule framework. The Directory Cell owns names and account identity; each repository UUID selects its own Repository Cell containing Git authority, access policy and collaboration state. Native Git prepares and serves wire protocol data through a disposable cache. A successful write depends on durable Cell publication, so rebuilding that cache does not discard acknowledged repository state. + +This document explains the system design, its authority and transaction boundaries, request and control flow, persistence, recovery, resource ownership and evolution. It is intended for engineers changing Canopy or operating a deployment. It describes the inspected Canopy revision `9438bb865959fb975d5349ba8b9908b461653821` and the Cellule dependency pinned by Cargo to `161067f5a21703b3e257024bcb64e565fd9657b4`, as of October 4, 2026. Several older documents record different pins; this guide uses the current declarations and executable registration. Diagrams show implementation boundaries, not measured production capacity. + +Use the [browsable gallery](diagram/canopy-architecture/index.html) to inspect the diagrams. Every diagram is a standalone SVG with a matching PNG at twice its logical resolution. The [persisted contracts](docs/contracts.md) define detailed protocol and storage behavior; the [delivery plan](docs/delivery-plan.md), [roadmap](ROADMAP.md) and [performance plan](docs/performance-plan.md) remain the records of qualification and open release gates. + +## Design scope and implementation status + +This is a description of the inspected implementation and its design implications. It does not introduce a new storage format or declare the service production ready. The snapshot includes an active Git hosting path and a newer packed-storage subsystem whose complete production integration remains open. + +| Area | Status in this design | Boundary | +| --- | --- | --- | +| Git hosting and collaboration | Implemented serving path | Smart HTTP, optional SSH, browse/API, Git LFS and SQL-backed repository features | +| Cellule integration | Implemented SQL Cells | Directory and Repository modules, typed calls, ownership fencing, durable outcomes and restoration | +| Multi-node routing | Implemented | Receiving gateways can call a live remote Cell owner through signed HTTPS | +| Archived pack storage | Implemented optimization | Per-object SQL authority remains the serving model | +| Immutable catalogs and ref-state roots | Implemented primitives with incomplete integration | Fresh-schema startup, all producers/readers and production hard cutover remain open | +| Composed repository capabilities | Proposed | KV, Queue, Workflow, Blob, Cron, Timer and Effects are not installed in today's Repository Cells | +| Production capacity and complete fault qualification | Open gates | A local result does not establish a deployment-wide resource or availability guarantee | + +The design preserves one Canopy binary, native Git, a managed local workspace, shared immutable storage and HTTPS ingress. It does not depend on a separate repository microservice, a second authoritative Git filesystem, or a distributed SQL transaction spanning every repository. The distinction between current behavior and future work is particularly important when reading the packed-storage and capability proposals. + +## Document map + +- [Goals and constraints](#goals-and-constraints) and [authority and invariants](#authority-and-invariants) explain the rules that implementation changes must preserve. +- [System components](#system-components) and [modeling Canopy on Cellule](#modeling-canopy-on-cellule) explain placement, contracts and stable identities. +- [Durable commands](#durable-command-execution), [routing](#request-routing-and-residency), [push](#git-push-and-replay), and [reads and LFS](#fetch-browse-and-git-lfs) trace execution. +- [Recovery](#owner-recovery-and-operational-control) and [collaboration](#collaboration-and-final-policy-checks) explain ownership changes and final policy decisions. +- [Storage evolution](#packed-storage-and-framework-capability-evolution), [resource ownership](#resource-admission-and-request-lifetimes), [security](#security-and-trust-boundaries), and [failure handling](#failure-handling) describe boundaries across those flows. +- [Tradeoffs](#design-decisions-and-tradeoffs), [verification](#verification-and-change-obligations), and [diagram maintenance](#regenerate-and-verify-the-diagrams) explain how to assess and maintain the design. + +## Diagram index + +| Diagram | What it explains | SVG | PNG | +| --- | --- | --- | --- | +| 01 | Public clients, node components and storage authority | [System overview](diagram/canopy-architecture/01-system-overview.svg) | [PNG](diagram/canopy-architecture/01-system-overview@2x.png) | +| 02 | How Canopy models Directory and Repository Cells on Cellule | [Framework mapping](diagram/canopy-architecture/02-canopy-on-cellule.svg) | [PNG](diagram/canopy-architecture/02-canopy-on-cellule@2x.png) | +| 03 | Commands, SQLite, LTX and the durable acknowledgement boundary | [Durable command](diagram/canopy-architecture/03-durable-command.svg) | [PNG](diagram/canopy-architecture/03-durable-command@2x.png) | +| 04 | Authentication, residency, local routing and signed peer routing | [Request routing](diagram/canopy-architecture/04-request-routing.svg) | [PNG](diagram/canopy-architecture/04-request-routing@2x.png) | +| 05 | Push preparation, object ingestion, policy checks and exact replay | [Git push](diagram/canopy-architecture/05-git-push.svg) | [PNG](diagram/canopy-architecture/05-git-push@2x.png) | +| 06 | Fetch, browse, LFS and current physical storage choices | [Reads and LFS](diagram/canopy-architecture/06-read-and-lfs.svg) | [PNG](diagram/canopy-architecture/06-read-and-lfs@2x.png) | +| 07 | Startup, leases, takeover, exact restore, drain and backup | [Owner recovery](diagram/canopy-architecture/07-owner-recovery.svg) | [PNG](diagram/canopy-architecture/07-owner-recovery@2x.png) | +| 08 | Repository features and merge publication controls | [Domain and policy](diagram/canopy-architecture/08-domain-and-policy.svg) | [PNG](diagram/canopy-architecture/08-domain-and-policy@2x.png) | +| 09 | Packed catalog primitives, remaining integration and capability proposals | [Storage evolution](diagram/canopy-architecture/09-packed-storage-evolution.svg) | [PNG](diagram/canopy-architecture/09-packed-storage-evolution@2x.png) | + +## Goals and constraints + +The hosting design combines stock Git behavior with repository-local durable policy. Git processes prepare or serve bytes; authoritative commands decide which changes become repository state. A successful acknowledgement must survive loss of the serving node's local files. + +| Goal | Design mechanism | Constraint or consequence | +| --- | --- | --- | +| Preserve acknowledged repository state | Fenced Cell authority and verified SQLite recovery roots | Local commit or object upload alone cannot authorize a success reply | +| Keep repository identity stable | UUID-derived Cell targets separate from owner/name aliases | Name changes and owner movement must not redirect stale operations to another repository | +| Check policy at publication | Refs, ACLs, graph certificates, rules and outcomes share a Repository Cell | Native preparation cannot authorize a later write by itself | +| Let any gateway receive traffic | Local or signed peer Cell bindings | Cell execution can move independently of native Git and body streaming | +| Recover from local disk loss | Restore accepted Cell roots and rebuild native caches | Cache contents and object listings cannot select authority | +| Bound admitted work | Residency, account, transfer, disk and native resource ownership | Admission accounting still needs OS containment and workload qualification | +| Preserve existing Git behavior | Stock Git workers and canonical object verification | Compatibility must be tested for each supported protocol and object format | + +There is no fixed logical repository-size quota used as a substitute for resource admission. SQLite representation limits, request codec bounds, batch sizes, disk budgets and worker admission are different constraints. None is evidence that the node can serve a particular number of repositories or developers. See the [implementation limits](docs/implementation.md) and [native resource contract](docs/design/native-resource-admission.md). + +## Authority and invariants + +### Terms and state ownership + +| Term | Meaning | Authority boundary | +| --- | --- | --- | +| Directory Cell | Shared name, account and credential state | One fixed SQL shard within the configured tenant/application | +| Repository Cell | Durable Git and collaboration state for one repository UUID | One independently fenced Cell | +| Cell target | Tenant, application, namespace and partition | Stable address independent of the node that currently owns it | +| Owner fence | Authority for a particular owner incarnation/session | Prevents a superseded owner from publishing authoritative state | +| Mutation identity | Identity of one Cellule command attempt and its recorded result | Reused when resolving or retrying the same command | +| Push identity | Product-level UUID bound to actor and encoded input digest | Names the complete HTTP push across several Cell commands | +| Receipt | Proof of an accepted per-Cell write position | Can constrain a later read of that Cell | +| Recovery root | Exact immutable SQLite recovery state selected by Cell authority | Publication makes a prepared root authoritative | +| Native cache | Local private refs and verified object files used by Git workers | Rebuildable data with explicit disk/process ownership | + +SQLite rows and immutable body references are durable repository state. The object store carries both Cellule control/recovery data and Canopy product bodies, but those have different roles. Conditional control updates select authority; content verification establishes the identity of immutable bytes. A stored object is not automatically reachable, published or safe to collect. + +### Rules that changes must preserve + +1. Each repository UUID has one durable Cell identity. A name lookup, cache directory or receiving node cannot replace that identity. +2. A Cell's accepted root is selected through fenced authority. Neither the newest-looking local database nor the result of a bucket listing can select recovery state. +3. A durable command records its domain mutation and replay outcome in the same SQLite transaction. A successful caller response follows the configured durability gate. +4. Ref publication rechecks current repository authorization, graph requirements, policy and expected ref versions. Preparation evidence does not bypass that final decision. +5. Unknown outcomes retain their original mutation/push identities. A timeout is not proof that the operation failed. +6. External bytes are verified and available before authoritative metadata points to them. Unreferenced prepared bytes do not expose repository refs. +7. Request pins, disk charges and native claims remain owned until their work and cleanup have actually settled. Cancellation alone is not release evidence. +8. Directory and Repository Cells are separate transaction domains. Multi-step creation and discovery need explicit coordination and rechecks. + +These rules follow the [persisted contracts](docs/contracts.md), [application registration](crates/canopy-server/src/lib.rs), [runtime assembly](crates/canopy-server/src/server/mod.rs) and [native ownership contract](docs/design/native-process-ownership.md). They apply to the active serving model and must remain true through a future storage cutover. + +## System components + +![Canopy components and storage boundaries](diagram/canopy-architecture/01-system-overview.svg) + +The `canopy` binary runs one Rust service. Axum handles HTTP, JSON APIs, the embedded browser and Git LFS; an optional russh listener handles SSH Git operations. `RepositoryManager` resolves identities, admits repository transitions, binds local or remote Cell clients and retains request pins. `GitGateway` turns repository state into a native Git workspace and translates native results into authoritative Cell operations. `LfsService` verifies and publishes LFS bodies independently of native Git. + +The Rust workspace has three crates. `canopy-git-format` supplies object kinds, SHA-1/SHA-256 identities and canonical hashing. `canopy-object-storage` owns immutable body and artifact operations. `canopy-server` composes those crates with Cellule and owns product schemas, protocols, authorization and deployment lifecycle. These are linked components of the service, not separately deployed microservices. See the [workspace map](docs/workspace.md) and [server assembly](crates/canopy-server/src/server/mod.rs). + +### Component responsibilities + +| Component | Responsibility | State it may publish | +| --- | --- | --- | +| HTTP and SSH ingress | Admit protocol input, authenticate and manage transport lifetime | Through registered product operations | +| RepositoryManager | Resolve targets, own resident bindings, transition locks and request pins | Catalog/acquisition/release through Cellule and name coordination through Directory | +| GitGateway | Prepare native work, ingest verified objects, certify closure and finalize pushes | Repository commands; private native refs stay provisional | +| Repository HTTP and browse services | Expose repository and collaboration operations | Registered commands with current actor/precondition checks | +| LfsService | Publish and verify LFS bodies, metadata and advisory locks | Authorized Repository Cell metadata after body verification | +| CanopyApplication and modules | Describe topology, schema, operation IDs/codecs and handlers | Compiled application contracts | +| CellNode and runtime | Supervise Cell execution, durable outputs, ownership and recovery | Fenced Cell control and published roots | +| Object storage | Retain immutable recovery and product bytes | Conditional control records and verified immutable objects | + +### Execution and control placement + +The request data path includes input spooling, native Git, verified readers and body streaming. The ownership control path includes release selection, catalog records, node enrollment, Cell acquisition and conditional root publication. They meet when a product operation calls a registered Cell command or query. This separation allows a receiving gateway to keep its native process and socket while authoritative SQL execution happens on another node. + +The Directory is shared state, so account and name operations do not scale by creating a new Directory for every repository. Repository state is partitioned by UUID, allowing unrelated repositories to have different owners. The architecture does not imply that one hot repository can have multiple concurrent authoritative SQL writers. + +## Modeling Canopy on Cellule + +![Domain modules, Cell targets and framework components](diagram/canopy-architecture/02-canopy-on-cellule.svg) + +`CanopyApplication::register` installs `DirectoryModule` and `RepositoryModule`, then declares their SQL Cell types. Modules describe schema migrations, operation IDs, codecs, code digests and limits. A compiled registry binds these contracts to the application and selected release. Product wrappers invoke registered commands and queries through typed handles. See [application and repository registration](crates/canopy-server/src/lib.rs) and [Directory registration](crates/canopy-server/src/directory/mod.rs). + +| Canopy concept | Cellule representation | Consequence | +| --- | --- | --- | +| Deployment identity | Tenant and application IDs | Scope every durable target | +| Directory | SQL namespace `[0x48; 16]`, fixed shard zero | One shared name and identity authority per tenant/application | +| Repository | SQL namespace `[0x47; 16]`, entity partition derived from UUID | One independently owned Cell for each repository | +| Directory operations | `SqlCell` plus credential commands | Account state and name changes stay inside the Directory boundary | +| Repository operations | `SqlCell` plus typed domain commands | Ref publication can check ACL, graph and policy in one transaction | +| Mutation | `MutationIdentity` and recorded outcome | Retries preserve one logical command identity | +| Observed write position | `Receipt` | A later read can require that per-Cell position | +| Node lifecycle | `CellNode` and installed runtime/replica facilities | Readiness and shutdown depend on supervised framework state | + +The repository partition is a prefix byte followed by a domain-separated 32-byte digest; `RepositoryCell::new` verifies that the supplied UUID derives the supplied target. A rename changes the Directory mapping while retaining the repository UUID and Cell identity. Creating a repository uses a durable pending name reservation, initializes the Cell and then marks the name ready. Those steps coordinate separate Cells; they do not form a distributed SQL transaction. + +`cellule-app` compiles topology and exposes handles. `cellule-host` supervises the node. `cellule-runtime` manages actors, SQL execution, request outcomes and fenced authority. `cellule-ltx` captures SQLite changes and verifies restoration; `cellule-store` supplies bounded object operations and conditional updates. Canopy supplies public ingress, user authorization, credentials and deployment policy. The [Cellule framework architecture at the pinned revision](https://github.com/crabbuild/cellule/blob/161067f5a21703b3e257024bcb64e565fd9657b4/docs/architecture.md) documents these boundaries. + +### Module contracts and typed handles + +A module is the executable contract installed into a Cell. Its operation descriptors specify IDs, codec versions, schema compatibility and input/output bounds; its code identity covers the authoritative implementation. `SqlCell` and `SqlCell` bind calls to these registered contracts. Product wrappers add domain operations while leaving routing, command identity, receipts and durable output handling to Cellule. + +The application declaration and runtime assembly serve different purposes. `CanopyApplication` declares the available Cell types; server startup installs runtime/replica facilities, enrolls the node and binds actual targets. Adding a Rust module or schema table does not automatically make an operation callable on the production path. Registration and release admission must agree with the handlers that execute and restore it. + +### Repository creation and alias changes + +Creation first reserves the owner/name with a canonical UUID and a pending state in Directory. It then initializes that UUID's Repository Cell and marks the reservation ready. A request interrupted between these phases must preserve the existing reservation rather than allocate a different identity for the same operation. The pending/ready protocol coordinates the two durable boundaries. + +Rename compares the expected UUID and moves the ready Directory name while retaining repository identity. Administrative operations that carry repository identity must reject name reuse that would otherwise retarget a stale write. The Repository Cell continues to own its object format, owner identity and repository policy. See [creation and discovery](docs/contracts.md#repository-creation-and-authorized-discovery). + +## Durable command execution + +![Command execution and the durability gate](diagram/canopy-architecture/03-durable-command.svg) + +A command executes at the current fenced owner of one Cell. Its SQLite transaction records the domain change and the request outcome together. On the object-store path assembled here, LTX captures the committed WAL boundary, verifies and publishes immutable recovery bytes, and supplies a proposed root. A conditional authority update publishes that exact root under the owner fence. The caller receives `Committed` and a receipt after this durability gate. + +Authority fencing prevents a former owner from publishing new authoritative state after ownership changes. A transport timeout can occur after a command committed; resolving its original identity distinguishes a completed outcome from an unstarted operation. A receipt is scoped to a Cell, owner incarnation and commit sequence. It can constrain a subsequent query, but it cannot create an atomic transaction across Directory and Repository Cells. Cellule also documents an optional follower-log durability path; the diagram describes Canopy's inspected object-store setup. + +### Commit and acknowledgement ordering + +The local SQLite transaction is an execution boundary. The accepted authority revision is the durable publication boundary. Between them, LTX must capture the exact WAL endpoint and immutable recovery bytes must become available. The authority update binds that recovery root to the valid owner fence. The caller must not receive a durable success acknowledgement solely because local SQL committed or uploads completed. + +The recorded outcome travels with the same restored SQLite state as the domain mutation. If a node disappears after publication but before its reply arrives, a new owner can recover both. If a proposed root was uploaded but never selected, that upload does not supersede the last accepted root. This is why failure handling uses command resolution and authority state rather than interpreting transport errors as domain results. + +### Read consistency and optimistic preconditions + +A receipt-bound query asks to observe at least the relevant published position of one Cell. It does not make a Directory lookup and a Repository query atomic. Discovery therefore rechecks current repository access, and final writes carry expected identities and versions into their own transaction. + +Git ref snapshots use a shared generation with HEAD. Continuation pages must retain that generation; a changed generation invalidates the scan. Ref writes compare both the expected OID and monotonic version. Deleting and recreating a name cannot make an old plan valid again merely because the OID matches. These are separate mechanisms: receipts constrain observation, generations bind a multi-page read, and versions fence optimistic writes. See [ref snapshots](docs/contracts.md#ref-snapshots-and-default-branch). + +## Request routing and residency + +![Routing through authentication, residency and owner resolution](diagram/canopy-architecture/04-request-routing.svg) + +HTTP credentials, SSH keys and LFS grants are checked using Directory state. Ready owner/name entries resolve to stable repository UUIDs. Repository routing then binds a `CellClient::local` to a resident actor or a `CellClient::peer` to the current remote owner. Canopy's peer transport signs requests to `POST /internal/cell` and verifies the enrolled sender, signature, release, expiry and principal over HTTPS. It is Canopy's own transport over runtime peer contracts; the optional Cellule peer adapter is not a Canopy dependency. See [peer routing](crates/canopy-server/src/server/peer.rs). + +The receiving node can run native Git and stream product bodies while Cell calls go to another node. Public requests therefore do not require sticky load-balancer sessions. Node advertisement and Cell ownership are separate controls: the advertisement proves a node session is live; the Cell control record identifies authority for a particular Cell. + +Cold and remote transitions use bounded account admission and a lock for that repository. The configured residency limit counts locally bound repository gateways, including remote bindings. When no slot is free, an inactive gateway can be evicted; local ownership must be released before its slot is reused. Requests and response streams pin their route. Supervised activation and cleanup continue after client cancellation. Full admission or ownership movement can return a retryable 503. See [residency](crates/canopy-server/src/server/residency/mod.rs) and [admission](crates/canopy-server/src/admission.rs). + +### Routing decisions + +| Observed state | Routing action | Required condition | +| --- | --- | --- | +| Resident local binding | Call the local Cell client | The binding remains valid and request-pinned | +| Live owner on another node | Bind a signed peer client | Current authority identifies an enrolled live owner and HTTPS endpoint | +| New or idle target | Admit local acquisition/bootstrap | Catalog/release checks and runtime capacity permit it | +| Expired owner | Fence and restore under new ownership | Recover exactly the root selected by authority | +| Full pinned resident set or busy movement | Return retryable capacity/unavailability | Do not steal an active slot or infer owner death from latency | + +A remote binding owns a local gateway entry but not the remote SQLite workspace. Evicting that binding does not release the remote Cell. When remote authority changes, subsequent routing refreshes the binding using current ownership. Node discovery is a routing input, while Cell control remains the ownership authority. + +## Git push and replay + +![Push from encoded input to saved durable response](diagram/canopy-architecture/05-git-push.svg) + +The gateway spools input to admitted scratch storage and binds the logical push UUID to the account, repository context and encoded request digest. `begin_push` detects completed operations before gzip decoding and native preparation. Reusing an ID with different bytes or an actor produces a conflict. Upload reception can overlap, but the current gateway serializes its ID check, native push work and publication phase through a mutex. See [gateway control flow](crates/canopy-server/src/git_gateway/mod.rs) and [owned preflight](crates/canopy-server/src/git_gateway/preflight.rs). + +New attempts capture consistent refs and policy, run native `receive-pack` against private refs and derive the actual accepted changes. Ingestion verifies canonical object identities, publishes required external bytes, persists bounded object batches and certifies typed graph closure. The gateway stages the response and ref plan. `CompletePush` rechecks current permissions, branch rules, expected OIDs and monotonic ref versions, then commits accepted refs, the generation change and the canonical response pointer together. Cellule publishes the durable result before the client sees success. See [push execution](crates/canopy-server/src/git_gateway/push.rs), [saved outcomes](crates/canopy-server/src/push/mod.rs) and [shared ref guards](crates/canopy-server/src/refs.rs). + +There are two identity layers: the push UUID identifies the complete wire operation; Cellule mutation identities identify its individual durable commands. A retry of a completed push replays the saved result and does not reapply old refs over later repository changes. Native partial acceptance is preserved, while stock Git `--atomic` requests atomic validation. Preparation failure cannot publish private refs, although verified unreferenced objects can remain. An uncertain final publication must be resolved rather than reported as a definite refusal. + +### Preparation and final publication + +| Phase | Work | Authoritative effect | +| --- | --- | --- | +| Input identity | Spool encoded input and bind actor, UUID and digest | Establish the product operation and detect completed replay | +| Native preparation | Build a private ref snapshot and run receive-pack | Produce provisional objects and the actual native report | +| Verified ingestion | Check canonical hashes, store object metadata/bodies and certify typed graph closure | Make verified objects available without publishing the final ref plan | +| Staging | Retain ordered ref changes and exact response data | Prepare a bounded final command | +| CompletePush | Recheck permissions, policy, OIDs and versions; commit refs, generation and response pointer | Publish the accepted repository transition atomically | +| Durable output | Publish the exact SQLite root through Cellule | Authorize the successful report to the client | + +Preparation can overlap unrelated operations, but the active gateway's push mutex serializes its native/ref-publication phase. A single push may involve several object batches and Cell commands; it is not one long distributed transaction over all uploaded bytes. The final command is the point that joins accepted refs with the canonical saved outcome. + +Graph closure validates typed dependencies: branch tips are commits; commit trees/parents, tree entries and tag targets require the appropriate reachable objects. Gitlinks refer to another repository and do not require local object presence. Certification supports final publication without repeatedly traversing already certified history. It does not replace every metadata or portability check performed by `git fsck`. + +### Retry scope and transport differences + +HTTP exposes the product push UUID so a caller can recover an uncertain report under the same input binding. A completed retry returns the saved status, headers and verified body instead of running receive-pack again. Reusing that UUID with a different actor or encoded digest conflicts. This replay must not overwrite refs changed by later independent pushes. + +SSH uses the same object ingestion and durable ref-publication rules but does not expose the HTTP push retry-ID contract. Documentation and clients must not promise that opening a new SSH command recovers the exact result of an earlier disconnected command. See [SSH transport](docs/contracts.md#ssh-transport) and [lost push reply recovery](docs/operations.md#recover-a-lost-push-reply). + +## Fetch browse and Git LFS + +![Read and storage paths](diagram/canopy-architecture/06-read-and-lfs.svg) + +Fetch reads generation-consistent refs and HEAD, validates wants against current repository reachability, hydrates selected verified bodies and lets native `upload-pack` generate the response. Warm verified object files are shared across private ref snapshots. Git v2 capability discovery avoids object hydration; ref discovery prepares ref targets and tag chains. Blobless fetch omits ordinary blobs, and supported filters guide further body selection. Browse APIs use repository metadata and verified object readers for trees, blobs, history and comparisons. See [fetch](crates/canopy-server/src/git_gateway/fetch.rs), [hydration](crates/canopy-server/src/git_gateway/hydration.rs) and [Git reads](crates/canopy-server/src/git_read/mod.rs). + +The active [repository schema](crates/canopy-server/src/schema.sql) stores per-object kind, size, digest and storage choice. Inline objects are bounded at 768 KiB; larger structural objects use SQL chunks. Large loose blobs use immutable external bodies. The serving code also archives pack/index bodies and can store packed-blob references while retaining per-object SQL authority. This active optimization is distinct from the future immutable catalog design. + +LFS transfers use HTTP even when SSH issues the authorization grant. Upload hashes and publishes bounded immutable parts, verifies the complete body and then commits authorized metadata. Download uses the manifest digest pinned in SQLite and verifies requested parts, including tail range responses. Locks are advisory Repository Cell records. See [LFS upload](crates/canopy-server/src/lfs/upload.rs) and [LFS read](crates/canopy-server/src/lfs/read.rs). Canopy's immutable body store is product code; it is not a registered Cellule Blob capability. + +### Storage representations + +| Representation | Durable metadata | Byte location | Verification role | +| --- | --- | --- | --- | +| Inline Git object | Kind, size, digest and OID | Repository SQLite row | Recompute canonical object identity | +| Chunked structural object | Bound upload/chunk metadata | Repository SQLite chunks | Verify complete object before publication | +| External loose blob | Size, digests and immutable body reference | Canopy object storage | Check body bytes against SQL metadata and Git identity | +| Packed blob | Per-object SQL record and archive binding | Immutable archived pack/index | Authenticate archive/object mapping and object bytes | +| LFS object | SHA-256, size and manifest binding | Immutable LFS body parts | Verify manifest-bound parts and requested ranges | +| SQLite recovery data | Authority-selected root and LTX metadata | Cellule immutable recovery storage | Restore the exact accepted Cell state | + +Each representation has an authoritative reference and a verified reader. A native cache may contain an additional copy, but cache presence does not create an object record or authorize a ref. The independent content checks matter because an object-store key or physical pack offset is only a location, not proof of the bytes' identity. + +### Snapshot lifetime + +An admitted fetch retains a coherent ref generation and its cache/input ownership through streaming. Later writes can advance the repository while that request serves its selected immutable snapshot. Discovery does not need to wait for an unrelated history hydration job. Authorization still precedes access; cached refs and warm objects cannot grant read permissions. The absence of production collection currently keeps unreferenced immutable bytes available, but a future collector must explicitly protect active readers and their retained roots. + +## Owner recovery and operational control + +![Node control and recovery](diagram/canopy-architecture/07-owner-recovery.svg) + +Startup locks the managed workspace, probes provider behavior, compiles the application, checks release admission and enrolls a signed node lease. Readiness depends on valid framework state and leases. A new Cell bootstraps; an idle Cell restores under acquired authority; a dead owner requires lease expiry and a fenced takeover. Recovery restores the root pinned by Cell authority, verifies required bytes and restores the outcome ledger before activation. Local databases and object listings cannot select a newer-looking root. Native caches rebuild afterward. See [acquisition and startup](crates/canopy-server/src/server/mod.rs) and [workspace management](crates/canopy-server/src/server/workspace/mod.rs). + +Graceful shutdown stops ingress, finishes admitted work, drains Cells and withdraws the node advertisement. Deployment maintenance closes release admission and uses an operation UUID until drain and recovery are proven complete; an explicit end reopens the release. Backup copies pinned roots and referenced external bodies into a disjoint prefix, verifies that independent copy and restores into an unused reserved destination. The current path is a same-provider copy. These controls do not complete schema migration, cross-provider export or collection. See the [operations runbook](docs/operations.md). + +### Startup and release admission + +Startup constructs and validates native admission, locks the local workspace, checks provider capabilities and the selected compiled release, then enrolls the signed node session. It checks release readiness around enrollment and before exposing ingress. Release changes or failed lease validation can close readiness and request supervised shutdown. A Cell's catalog and registered handlers must be compatible with the selected release before acquisition. + +Maintenance admission, local file exclusion and ownership fencing are distinct controls. Maintenance can close new work across the release; the workspace lock prevents competing local runtimes; Cell authority fences the writer. A released listener or expired heartbeat alone does not prove that every admitted native or Cell operation finished. + +### Drain and workspace reuse + +One supervisor retains startup, admitted work, native processes, Cell drain, leases and workspace exclusion. Canceling the caller's startup/shutdown wait requests or observes cleanup; it does not discard that supervisor. Native admission closes, tracked ingress/work joins, native owners drain, then Cellule shuts down. Only confirmed drain authorizes workspace release and advertisement withdrawal. + +If drain or cleanup is uncertain, the workspace fence or native claim remains retained. On Unix, worker ownership includes inherited completion/lock descriptors so participating descendants cannot outlive the accounting boundary unnoticed. Linux listener guards additionally prevent a fork-inherited listener from surviving a completed handoff. These mechanisms do not claim general containment of every possible helper on every OS. See [native process ownership](docs/design/native-process-ownership.md) and [resource admission](docs/design/native-resource-admission.md#shutdown-boundary). + +### Backup and restoration boundaries + +A backup is a verified independent copy of pinned recovery roots and referenced product bodies, with its own operation identity and destination. Restoring an empty reserved deployment from that copy differs from reconstructing a live owner's local cache. The former must preserve the copied release/schema and destination admission; the latter follows current Cell authority. Maintenance recoveries for explicitly supported retained contracts do not imply a general old-release migration or cross-provider restore path. + +## Collaboration and final policy checks + +![Repository components and merge control](diagram/canopy-architecture/08-domain-and-policy.svg) + +The Repository Cell stores ACLs and visibility alongside issues, PRs, reviews, line discussions, commit checks and branch rules. This placement lets final commands check current policy against the same state that they mutate. Directory listings only suggest candidate repositories; the product rechecks actual Repository Cell access before exposing them. + +For a merge, preparation captures exact base/head revisions and may use native `merge-tree` or `commit-tree` to construct candidate objects. `MergePull` checks the actor, revisions, current branch state, required checks, reviews and unresolved discussions before changing the branch and PR state. The candidate remains provisional until that command publishes. Check results are API records tied to commits and attempts; they do not imply that Canopy currently runs CI through a Cellule Workflow capability. See [merge command](crates/canopy-server/src/pulls/merge/command.rs), [candidates](crates/canopy-server/src/pulls/candidates/mod.rs) and [branch rules](crates/canopy-server/src/branch_rules/command.rs). + +### Why policy belongs beside refs + +An ACL grant, branch rule, review or check can change after native preparation starts. Co-locating these facts with refs lets the final Repository Cell command examine their current versions and exact reviewed revisions in the same transaction that changes the branch. A ready merge candidate describes prepared bytes; it does not grant a standing right to merge them later. + +Issues, comments and PR records use repository-local identities and optimistic versions. Check attempts retain reporter/context policy and terminal results. These records participate in repository recovery, so owner movement does not split collaboration history from the Git branch state it governs. External runners, durable event delivery and autonomous workflow execution are separate integration work. + +## Packed storage and framework capability evolution + +![Active model, new primitives and remaining work](diagram/canopy-architecture/09-packed-storage-evolution.svg) + +The newer `packs` subsystem implements immutable native-pack metadata, canonical directory runs, leveled indexes, source roots, catalogs, ref-state roots and exact outcome artifacts. Staging and preparation retain attempts and input custody; trusted certificates bind verified work to the repository, owner fence and base. A bounded publication coordinator dispatches registered commands by class and account. Short final commands publish catalog/ref state and outcomes while rechecking authority and policy. Uploaded artifacts remain preparation until authorized publication selects them. + +The [publication module](crates/canopy-server/src/packs/publication/mod.rs) explicitly states that its commands are not registered on the legacy serving path. The [implementation status](docs/large-repository-implementation-status.md) lists production startup, HTTP, SSH, generated producers/readers and the fresh-schema hard cutover as open work, along with recovery, collection, backup and capacity qualification. The atlas therefore separates implemented primitives from a complete replacement serving system. + +A separate [repository capability proposal](docs/repository-cell-primitives.md) aims to compose SQL, KV, Queue, Workflow, Blob, Cron, Timer and Effects within one repository identity, fence and recovery boundary. Current `RepositoryModule` declares `CatalogRole::Sql` and has empty workflow/activity inventories. Autonomous per-repository work and composed capabilities remain acceptance goals. Framework support for a primitive does not mean Canopy has wired that primitive into its Repository Cells. + +### Hard cutover requirements + +The new model shifts large inventories, ref snapshots and exact outcomes into immutable artifacts while retaining short authoritative SQL publication commands. Prepared certificates bind verified artifacts to the repository, owner fence and base state. Service-owned staging, checkpoints and recovery records must retain the original command inputs across uncertainty; rebuilding a similar-looking command is not exact recovery. + +Completing this model requires compatible registration and recovery admission, production HTTP/SSH producers, all readers, a fresh-schema format selection, retained-root collection, isolated restore, maintenance and capacity qualification. The current implementation status, rather than the presence of individual types or local tests, determines whether each gate is closed. The active archived-pack optimization must not be presented as the completed immutable-catalog cutover. + +### Composed capabilities are a separate design + +Installing additional primitives into a repository requires explicit capability metadata, registry validation, typed APIs, durable work inspection and runner/lifecycle integration under the same target and fence. A SQL table named queue is not sufficient. Current release/acquisition logic must account for every installed primitive's live work before moving or retiring the Cell. + +Canopy's product-level blob/LFS store and check records do not instantiate Cellule Blob or Workflow capabilities. The proposal describes the desired composition; it must be rechecked against the exact framework revision selected for implementation. Historical dependency pins in that proposal are not the pin of the serving snapshot documented here. + +## Resource admission and request lifetimes + +Canopy uses several independent admission boundaries. They protect different retained resources and cannot be replaced by one count of open HTTP requests. + +| Boundary | What it owns | Release condition | +| --- | --- | --- | +| Repository residency | Local gateway slot, route binding and transition state | No request pins; confirmed local Cell release and cleanup when applicable | +| Account/transition admission | Accepted activation or product work | Admitted operation and supervised cleanup settle | +| Transfer admission | Git/LFS input and output work | Body/trailers finish or transport cleanup completes | +| Local disk budget | Spools, cache generations and temporary body files | Confirmed deletion/reclamation, retaining charges on uncertain cleanup | +| Native resource vector | Process slots, CPU admission units, memory claims and descriptor claims | Process ownership/reaping and inherited completion drain are proven | +| Packed preparation/publication services | Bounded retained inputs, commands, results and observation state | Exact terminal result or safe explicit lifecycle transition | + +The node constructs one `NativeResources` pool from required `native_limits` configuration and shares it across gateways/services. A claim admits the complete vector atomically or returns typed exhaustion. Foreground and reserved maintenance shares are disjoint; an unused reservation is not silently borrowed. Every shared Git process requires a private permit. This prevents independent repositories from each assuming the whole node budget is available. + +The configured CPU units and memory bytes are admission estimates. They do not throttle CPU or cap actual RSS. OS/container limits, native helper behavior and full-history workloads need separate qualification. Node-native admission also does not implement per-account vector fairness by itself. The [configuration example](config.example.json) illustrates a profile; it is not a throughput or memory guarantee. + +A streamed response retains the route pin and relevant ownership until EOF, trailers, error or disconnect cleanup. Native cancellation transfers ownership to the independent reaper; signaling a process group cannot immediately return capacity. Unknown cleanup can quarantine credits and keep drain pending. New admission receives capacity/unavailability responses instead of creating an unbounded retained waiting queue. See [native resource admission](docs/design/native-resource-admission.md) and [disk/request ownership](docs/contracts.md#repository-routing-and-residency). + +## Security and trust boundaries + +### Public requests and repository policy + +The Directory owns accounts, token digests, scopes, expirations, SSH key enrollment and administrative audit. HTTP authentication checks credential and account state; SSH authentication requires verified key possession in addition to a stored active key. Repository read/write/admin decisions remain in current Repository Cell state. The immutable repository owner and individual collaborator grants are different from fleet node identity. + +Directory credential-management commands repeat their authorization in their transaction. Repository commands repeat repository policy at their final write. Account disablement or token expiry prevents subsequent admission, while an already authenticated request can finish according to its documented lifetime and current repository ACL checks. A cached name, browser view or discovery candidate cannot restore revoked permissions. Anonymous reads require current public visibility, and forbidden repository metadata is hidden according to the API contract. + +### Internal peers and infrastructure + +The peer adapter authenticates enrolled node sessions, not end-user credentials. It verifies the signed Canopy node principal/action, tenant/application/namespace, release and expiry, and routes using authoritative ownership plus the matching signed advertisement. Outbound HTTPS verifies the endpoint certificate and hostname; a configured private CA adds trust. Redirects, environment proxies and transport retries are disabled so untrusted fields cannot redirect Cell commands or silently replace an uncertain mutation. + +Canopy uses signed requests over HTTPS; this adapter does not use Cellule's separate mTLS helper. The object store, TLS terminator and enrolled signing keys are trusted fleet infrastructure. A compromised enrolled node has internal Cell authority and is outside the end-user ACL boundary. Content hashes authenticate bytes; they do not replace TLS, credential protection or deployment access controls. See [peer routing contracts](docs/contracts.md#peer-routing) and [peer implementation](crates/canopy-server/src/server/peer.rs). + +## Failure handling + +The key distinction is whether failure occurs before execution, during preparation, or after a command may have published. Only authoritative resolution can turn an unknown command into a known result or known absence. + +| Failure or race | Design response | State that must remain protected | +| --- | --- | --- | +| Admission is full or movement is busy | Return retryable unavailability before accepting more work | Existing pins, slots and resource claims | +| A live Cell owner is remote | Route signed Cell calls to that owner | Stable target and current ownership | +| Owner/session expires | Fence takeover and restore the accepted root | Acknowledged SQL state and outcome ledger | +| Local SQL commits but root publication is unproven | Do not issue durable success; resolve the original identity | Last accepted authority root | +| Root publishes but the reply is lost | Resolve/replay the original command or product push | Exact saved outcome and later independent ref changes | +| External upload succeeds but final publication fails | Leave bytes provisional/unreferenced | Authoritative refs and metadata | +| ACL, rule or ref changes during preparation | Final command rechecks and can refuse/conflict | Current policy and monotonic ref versions | +| A ref generation changes during pagination | Discard the partial scan and retry within its bounded policy | One coherent refs/HEAD snapshot | +| A cache or body cannot be verified | Do not expose unverified bytes as repository state | Canonical object identity and selected durable references | +| Client cancels accepted work | Retain supervised operation/native cleanup until settlement | Request identity, route pins, disk charges and process claims | +| Shutdown cannot prove drain | Retain exclusion/claims; do not reuse the workspace as clean | Local worker and authority safety | +| Backup or maintenance reply is ambiguous | Continue using the original operation UUID and status protocol | Exact operation ownership and isolated destination | + +An API conflict and a transport timeout therefore require different handling. A known conflict provides a domain decision; an unknown reply requires resolution. Changing the logical operation ID to bypass uncertainty can create a second operation and violates the replay contract. Detailed operator procedures live in the [operations runbook](docs/operations.md). + +## Design decisions and tradeoffs + +These are implications of the current architecture, rather than additional production promises. + +| Decision | Benefit | Cost or limit | +| --- | --- | --- | +| One SQL Cell per repository | Ref, ACL and collaboration decisions can share a transaction | One hot repository retains a single authoritative execution boundary | +| Shared Directory separate from repositories | Names and accounts have one durable identity authority | Creation/discovery need coordination across Cells; Directory is shared work | +| Native Git over private rebuildable caches | Uses stock wire/pack behavior while keeping durable state independent of local files | Hydration, scratch, native helpers and resource ownership add work | +| Durable exact outcomes | Unknown replies can resolve without reapplying old mutations | Retained outcomes, staged input and exact identity need lifecycle/retention rules | +| Immutable external bodies with SQL references | Large bytes need not live inline in SQLite | Upload and SQL publication are separate phases; orphan collection needs root proof | +| Any gateway with signed peer calls | Public routing need not be sticky to the SQL owner | Adds peer authentication, remote latency and uncertainty handling | +| Final policy rechecks | Prepared work cannot bypass later authorization/ref changes | Preparation may be discarded after substantial native work | +| Explicit admission and maintenance reservations | Retained resource demand is accounted before launch | Estimates require measurement; quarantined claims reduce available capacity | +| Fresh-schema packed-storage cutover | Can establish one coherent producer/reader/authority format | Cannot claim completion until every registration, reader and recovery gate is satisfied | + +## Verification and change obligations + +Changes should identify which invariant and transaction boundary they affect before modifying code. Protocol or storage changes update the [persisted contracts](docs/contracts.md) alongside focused tests; compatibility and fault results record the exact revision, provider, tools, workload and hardware. The [delivery plan](docs/delivery-plan.md) determines release acceptance, while the [performance plan](docs/performance-plan.md) separates measured results from targets. + +| Change area | Required design evidence | +| --- | --- | +| Module registration or codec/schema | Compatible descriptors and exact release/code identity; explicit upgrade or fresh-prefix policy | +| Ref/push publication | SHA-1/SHA-256 client behavior, stale versions, policy changes, rollback and exact replay | +| Routing or ownership | Live remote binding, competing acquisition, lease loss, fenced takeover and exact restoration | +| Streaming or native work | Cancellation, disconnect, descendant/reaper ownership, disk accounting and honest drain | +| External bytes or pack readers | Canonical identity, corruption/range failures, immutable reference binding and recovery | +| Backup, retention or cutover | Complete retained root set, independent restore, concurrent readers/writers and fault injection | +| Capacity claim | Whole-operation CPU/RSS/disk/descriptor behavior, mixed workload and recorded environment | + +Operational visibility is also an open release concern. The design needs low-cardinality signals for admission/refusals, active/transitioning Cells, native claims/drain, lease/release health, object-store failures, disk use and operation latency. Existing health/tracing or bounded native counters must not be described as a completed metrics/alerting contract. See [operational visibility](ROADMAP.md#r06-operational-visibility-and-runbooks). + +Safe versioned migrations, complete provider/fault qualification, conservative collection, cross-provider backup, full observability, repeatable deployment, immutable-storage hard cutover and composed capabilities remain separate acceptance work. A change should close a gate only with its stated evidence; it must not widen timeouts, invent fixed logical size caps or reinterpret admission estimates as measured capacity. + +## Regenerate and verify the diagrams + +The [generator](diagram/canopy-architecture/generate.py) contains the SVG source layouts, labels and source map. It uses embedded styles and system monospace fonts, with no external rendering dependency in the SVG files. + +```sh +python3 diagram/canopy-architecture/generate.py +python3 diagram/canopy-architecture/render_pngs.py +python3 diagram/canopy-architecture/build_gallery.py +``` + +The PNG export helper requires `rsvg-convert` on `PATH` and renders at twice the SVG viewBox dimensions. The gallery embeds the SVGs directly and works offline. Sequence diagrams draw activation bars behind dashed lifelines, with message arrows on top. The atlas is documentation only; repository runtime behavior is unchanged. diff --git a/README.md b/README.md index e9636fa..2ab98bf 100644 --- a/README.md +++ b/README.md @@ -14,6 +14,7 @@ This page keeps the project overview, setup path, and first repository workflow. | You want to… | Start with | | --- | --- | | Understand the storage model | [How Canopy stores a repository](#how-canopy-stores-a-repository) | +| Understand the design and control flow | [Canopy design](DESIGN.md) and [architecture gallery](diagram/canopy-architecture/index.html) | | Navigate the Rust crates | [Rust workspace](docs/workspace.md) | | Run a development node | [Run a local node](#run-a-local-node) | | Start an isolated local server and test a Kubernetes-sized repository | [Local evaluation and real-repository benchmark](deploy/local-evaluation.md) | diff --git a/diagram/canopy-architecture/01-system-overview.svg b/diagram/canopy-architecture/01-system-overview.svg new file mode 100644 index 0000000..0d67446 --- /dev/null +++ b/diagram/canopy-architecture/01-system-overview.svg @@ -0,0 +1,6 @@ + +Canopy system architectureCanopy is an embedded Cellule application. Public ingress and native Git run in Canopy, while each durable Directory or Repository Cell has one fenced owner. A receiving node may call a remote owner. The object store holds authority records, SQLite recovery roots and product body objects. Native Git caches are disposable. + + +CANOPY NODE / ONE RUST SERVICEDURABLE CELL AUTHORITYGit clientsSmart HTTP / SSHclone, fetch, pushBrowser / APIEmbedded web UIJSON collaboration APIGit LFSHTTP bodies / locksSSH can issue grantsIngressAxum HTTP routerOptional russh listenerAccount authenticationRepositoryManagerName resolutionResidency + request pinsLocal / peer Cell handlesProduct servicesGitGateway + LfsService + repository_httpBrowse, issues, PRs, reviews, checks, policyNative Git work runs on the receiving nodeCellule integrationCanopyApplication + CellNodeTyped SQL / domain commandsFenced ownership + durable output gateDirectory CellOne fixed SQL shard per tenantNames -> stable repository UUIDAccounts, tokens, keys, auditRepository CellOne SQL Cell per repository UUIDRefs, objects, ACL, collaborationOutcomes, closure and policyDisposable Git cachePrivate refs + verified shared objectsNative receive/upload-pack workersCan be rebuilt after local disk lossShared S3-compatible object storeCellule: catalog, authority, node leases, LTX roots and immutable SQLite recovery bytesCanopy: external Git blobs, archived pack/index bytes and Git LFS bodiesConditional writes select authority; verified immutable bytes supply recoveryresolve / authtyped operationsnative workdurable publication / external body I/OhydrateOwnership ruleGit success isacknowledged only afterthe Cell commits andpasses its durabilitygate.public clientsproduct logicframeworkdurable authorityrebuildable local filesCanopy system architectureClients enter any node; Cell authority and immutable storage preserve repository state.CURRENT SERVING PATH + \ No newline at end of file diff --git a/diagram/canopy-architecture/01-system-overview@2x.png b/diagram/canopy-architecture/01-system-overview@2x.png new file mode 100644 index 0000000..16384d3 Binary files /dev/null and b/diagram/canopy-architecture/01-system-overview@2x.png differ diff --git a/diagram/canopy-architecture/02-canopy-on-cellule.svg b/diagram/canopy-architecture/02-canopy-on-cellule.svg new file mode 100644 index 0000000..b01a539 --- /dev/null +++ b/diagram/canopy-architecture/02-canopy-on-cellule.svg @@ -0,0 +1,6 @@ + +How Canopy models its domain on CelluleCanopyApplication registers DirectoryModule and RepositoryModule as SQL Cell types. Directory uses a fixed shard; repository UUIDs derive entity partitions. Product wrappers hold typed SqlCell handles. Cellule owns execution and persistence mechanics, and Canopy owns protocols and authorization. Other framework primitives are available but are not composed into Canopy Repository Cells. + + +APPLICATION DECLARATION / COMPILED AT STARTUPRUNTIME ADDRESSING / STABLE ACROSS OWNER MOVEMENTEMBEDDED CELLULE FRAMEWORK / THE PINNED DEPENDENCYCanopyApplicationCellApplication::registerDirectoryModule + RepositoryModuleTwo declared SQL Cell typesDirectoryModuleCatalogRole::Sql / one fixed shardNamespace = [0x48; 16]Names, accounts, credentialsRepositoryModuleCatalogRole::Sql / entity partitionsNamespace = [0x47; 16]Git + ACL + collaboration commandsDirectory targetTenant + application + namespacepartition_for_shard(0)DirectoryCell wraps SqlCell<DirectoryModule>Repository targetTenant + application + namespace + UUID-derived partitionCellType::entity_partition -> 33-byte partitionRepositoryCell wraps SqlCell<RepositoryModule>UUID selects one Cellcellule-app + hostDescriptor / registry / typed handlesCellNode lifecycle + readinessCanopy installs facilities and leasescellule-runtimePer-Cell actor + SQL workerAuthority fence + request ledgerQueries, commands, peer contractscellule-ltx + storeManaged SQLite WAL captureVerified roots + exact restorationBounded I/O + conditional updatesDependency pin: 161067f5a21703b3e257024bcb64e565fd9657b4cellule-types is transitive. Canopy supplies its own signed HTTPS peer transport.Framework capabilitiesCellule includes SQL, KV, Queue, Workflow, Blob, Cron, Timer and Effects.Capabilities use typed APIs and explicit topology.Canopy capability boundaryCanopy registers SQL Cells today. Multi-capability Repository Cells, repositoryqueues and autonomous workflows remain a proposal.domain declarationsframework bindingfenced executioncapabilities not wired into CanopyHow Canopy models its domain on CelluleDomain code supplies schema and policy; the embedded framework supplies identity, execution and recovery.CURRENT SERVING PATH + \ No newline at end of file diff --git a/diagram/canopy-architecture/02-canopy-on-cellule@2x.png b/diagram/canopy-architecture/02-canopy-on-cellule@2x.png new file mode 100644 index 0000000..5844fb8 Binary files /dev/null and b/diagram/canopy-architecture/02-canopy-on-cellule@2x.png differ diff --git a/diagram/canopy-architecture/03-durable-command.svg b/diagram/canopy-architecture/03-durable-command.svg new file mode 100644 index 0000000..90ae9cc --- /dev/null +++ b/diagram/canopy-architecture/03-durable-command.svg @@ -0,0 +1,6 @@ + +Cellule command execution and durabilityThis is the object-store durability path assembled by Canopy. The Cell owner records domain state and the request outcome together, captures SQLite through LTX and publishes the exact root using a fenced conditional authority update. Only then does the caller receive Committed and a receipt. A Cellule follower-log mode exists, but this diagram does not claim Canopy enables it. + + +Canopy callerTyped handleCell ownerActor + fenceManaged SQLiteState + ledgerLTX / object storeImmutable rootsCellAuthorityControl CAS1 Command + stable MutationIdentity2 Validate owner incarnation / lease / compatible code3 Execute domain mutation + request outcome in one transaction4 Commit SQLite and capture its exact WAL boundary5 Verify capture; upload immutable chunks and proposed root6 Return proposed recovery root7 Conditional publish of the exact root under owner fence8 Accepted authority revision = durable publication proof9 Committed<Output> + ReceiptLost reply or rejected fenceA timeout does not prove failure. Resolve the original request identity; retry thesame logical command only under its existing identity.Receipt and scopeA receipt-bound query must observe the published per-Cell position. Directory andRepository commands do not form one cross-Cell SQL transaction.invocationlocal transactionimmutable bytesauthority publicationdurable acknowledgementCellule command execution and durabilityOne command changes one Cell. The domain mutation and its recorded answer share a SQLite transaction.CURRENT SERVING PATH + \ No newline at end of file diff --git a/diagram/canopy-architecture/03-durable-command@2x.png b/diagram/canopy-architecture/03-durable-command@2x.png new file mode 100644 index 0000000..49130de Binary files /dev/null and b/diagram/canopy-architecture/03-durable-command@2x.png differ diff --git a/diagram/canopy-architecture/04-request-routing.svg b/diagram/canopy-architecture/04-request-routing.svg new file mode 100644 index 0000000..2665549 --- /dev/null +++ b/diagram/canopy-architecture/04-request-routing.svg @@ -0,0 +1,6 @@ + +Request routing and repository residencyDirectory state authenticates accounts and resolves ready repository names. RepositoryManager binds a stable target to either a local actor or a signed peer client. Cold transitions are admitted and supervised; remote-owner loss triggers acquisition on demand. Repository authorization is checked using current durable state. Response pins prevent eviction during streaming. + + +Incoming requestHTTP / SSH / API / Git LFSDirectory CellValidate token / key / LFS grantResolve ready owner/name -> UUIDRepositoryManagerDerive target; pin resident routeOtherwise admit a transitionLive ownerelsewhere?Remote Cell bindingCellClient::peer -> /internal/cellSigned request + enrolled node keyTLS, release, expiry, principal checksLocal acquisitionBootstrap new / acquire idle CellExpired owner: fence + restore exact rootCellClient::local -> resident actorRepository operationsRecheck current ACL / visibilityRun domain command or native GitRetain route pin through responseYesNoDiscovery is a hintDirectory listing candidates are checkedagainst current Repository Cell access. Cachednames and UI state cannot grant permissions.Bounded transitionsCold/remote work uses supervised admission anda per-repository transition lock. Differentrepositories can activate concurrently.Capacity and cancellationAn unavailable slot or in-progress movement can return 503. Streamed responses anddetached admitted work retain pins until they finish.Native work placementRemote binding moves Cell calls. Git workers, body streaming and the disposable cachecan remain on the receiving gateway.requestidentity / authorityadmissionlocal or signed remote executionRequest routing and repository residencyAuthentication, name lookup, account admission and owner resolution precede repository work.CURRENT SERVING PATH + \ No newline at end of file diff --git a/diagram/canopy-architecture/04-request-routing@2x.png b/diagram/canopy-architecture/04-request-routing@2x.png new file mode 100644 index 0000000..a1d774b Binary files /dev/null and b/diagram/canopy-architecture/04-request-routing@2x.png differ diff --git a/diagram/canopy-architecture/05-git-push.svg b/diagram/canopy-architecture/05-git-push.svg new file mode 100644 index 0000000..d129131 --- /dev/null +++ b/diagram/canopy-architecture/05-git-push.svg @@ -0,0 +1,6 @@ + +Git push from wire input to durable refsAfter outer authentication, GitGateway spools and hashes the encoded request and binds a logical push ID to the actor and digest. It can replay a completed push before decoding and native work. New attempts prepare a private ref snapshot, run receive-pack, ingest canonical objects, certify graph closure and stage the report and ref plan. CompletePush checks current policy and expected ref versions and commits accepted refs with the canonical response pointer. Cellule gates acknowledgement on durable publication. + + +FINAL DOMAIN TRANSACTION: ACCEPTED REFS + GENERATION + OUTCOME POINTER / AUDITGit clientreceive-packGitGatewayReceiving nodeNative GitDisposable refsRepository CellVia local / peer clientObject storeVerified body bytes1 Spool encoded input; bind push UUID + actor + request digest2 begin_push: claim logical ID or find completed response3 Completed ID replays saved result; fresh ID continues4 Read consistent refs, generation and branch policy5 Decode / validate; receive-pack in private cache6 Native report + actual accepted ref differences7 Upload verified external blobs or archived pack/index bytes8 Persist canonical object records in bounded batches9 Certify typed graph closure; stage response + ref plan10 CompletePush: recheck ACL, branch rules, OIDs and versions11 Cellule publishes SQLite root12 Fenced durable proof13 Read canonical saved response14 Git report + push identityTwo identities matterThe HTTP push UUID identifies the whole wire operation. Cellule MutationIdentityidentifies each durable command within that operation.Failure behaviorPreparation changes disposable refs only. Objects may remain unreferenced afterrefusal. An uncertain final publication is resolved, never turned into a falsenegative report.Native partial acceptance is preserved; Git --atomic can request all-or-none validation. Current gateway push work uses a mutex.wire protocolprivate native statedurable Cell callsbody / root publicationGit push from wire input to durable refsCurrent serving path: native Git prepares private state; CompletePush publishes refs and the saved result.CURRENT SERVING PATH + \ No newline at end of file diff --git a/diagram/canopy-architecture/05-git-push@2x.png b/diagram/canopy-architecture/05-git-push@2x.png new file mode 100644 index 0000000..19538a3 Binary files /dev/null and b/diagram/canopy-architecture/05-git-push@2x.png differ diff --git a/diagram/canopy-architecture/06-read-and-lfs.svg b/diagram/canopy-architecture/06-read-and-lfs.svg new file mode 100644 index 0000000..153d3ec --- /dev/null +++ b/diagram/canopy-architecture/06-read-and-lfs.svg @@ -0,0 +1,6 @@ + +Fetch, browse and Git LFS data pathsFetch takes consistent refs, validates requested object reachability, hydrates selected verified objects and uses native upload-pack to stream the wire response. Browser APIs read through Cell metadata and verified body readers. LFS uploads first publish and verify immutable bytes, then commit authorized metadata; downloads verify manifest-bound parts. The active serving schema supports inline, chunked, external and packed storage records, distinct from the incomplete immutable catalog replacement. + + +GIT CLONE / FETCHBROWSE / JSON READS / LFSCURRENT SERVING STORAGE / AUTHORITATIVE METADATA REMAINS PER OBJECTAuthorize and selectCurrent ACL / public visibilityValidate wants against live refsSnapshot and hydrateGeneration-consistent refs + HEADVerify required objects / bodiesApply supported partial-clone filterNative upload-packPrivate ref snapshotShared verified object filesStream generated pack to clientDiscovery fast pathGit v2 capability discovery needs no object hydration. Ref discoveryprepares ref targets and annotated tag chains.Selective fetchBlobless fetch omits ordinary blobs. Selected wants, ref targets and structuralhistory determine hydration; warm verified objects are reused.Repository browserRead trees, blobs, history and diffsUse Repository Cell metadataLoad bodies from verified readersNo receive-pack is neededLFS uploadAuthorize batch / PUTHash SHA-256 + bounded partsPublish immutable body, verifyThen commit authorized metadataLFS downloadAuthorize against current CellRead manifest pinned in SQLiteVerify requested parts / digestsSupport tail Range / 206 responseLFS locks are advisory records in the Repository Cell. SSH grants still deliver LFS bytes over HTTP.Repository SQLiteobjects: kind, size, digest, storageinline bodies <= 768 KiBLarge non-blob bodies: SQL chunksrefs + graph edges / certificatesLFS metadata + locksImmutable external bytesLarge loose Git blobs + manifestsArchived pack and index bodiesPacked blobs reference approved packsGit LFS bodies + part digestsReaders verify canonical identitiesLocal cacheHydrated Git files / installed packsInsertion cursor + ref generationsPrivate native process workspaceDisk and native-worker admissionDisposable after recoveryreferencesverify / hydrateThis active archive path differs from the future immutable catalog / ref-root hard cutover in diagram 09.Read authorityAn uploaded pack, cached object, stale listing or UI view cannot independently authorize an object read. Current Cell access and verified published metadata govern visibility.stock Git readsbrowse / selectiondurable state and bodiesrebuildable work filesFetch, browse and Git LFS data pathsReads verify identities against Cell metadata. LFS body delivery bypasses native Git.CURRENT SERVING PATH + \ No newline at end of file diff --git a/diagram/canopy-architecture/06-read-and-lfs@2x.png b/diagram/canopy-architecture/06-read-and-lfs@2x.png new file mode 100644 index 0000000..6d662bf Binary files /dev/null and b/diagram/canopy-architecture/06-read-and-lfs@2x.png differ diff --git a/diagram/canopy-architecture/07-owner-recovery.svg b/diagram/canopy-architecture/07-owner-recovery.svg new file mode 100644 index 0000000..043d65c --- /dev/null +++ b/diagram/canopy-architecture/07-owner-recovery.svg @@ -0,0 +1,6 @@ + +Owner lifecycle and exact recoveryCanopy locks its local runtime workspace, verifies storage capabilities, validates the selected release and enrolls a signed node lease before serving. New or cold Cells bootstrap, acquire idle authority or fence an expired owner and restore the authority-pinned SQLite root. Recovery preserves the request outcome ledger. Native caches rebuild afterward. Shutdown and maintenance supervise all admitted work; backup copies durable roots and referenced external bodies into an independent prefix. + + +STARTUP AND SERVINGCOLD ACTIVATION / NODE LOSS / LOCAL DISK LOSSDRAIN, MAINTENANCE AND INDEPENDENT BACKUPStartup preflightLock managed local workspaceProbe conditional writes / ranged I/OCompile and check selected releaseEnroll and renew nodeSigned live node advertisementNodeLeaseGuard + compatible registryStart Directory + on-demand reposServe with fencingReady only while leases are validPer-Cell authority / actor / SQLCancel ingress when lease failsRead catalog and controlValidate compiled code and schemaLive owner: route to that ownerIdle or expired: acquire authorityFence and choose exact rootClaim expired node for takeoverUse CellAuthority-pinned rootDo not elect a root by listing keysVerify and restore SQLiteLTX chunk checksums / endpointsRecover exact state + outcome ledgerInvalid bytes: fail without activationActivate successorSame CellTarget / repository UUIDNew fenced owner / incarnationRecovered refs, ACL and outcomesRebuild Git cacheHydrate verified published objectsRecreate private ref snapshotsResume Git/API/LFS requestsCache loss is recoverableCell publication protects acknowledged state.Files from an old workspace never become asubstitute recovery root.Recovery is on demand. An unclean owner loss waits for lease expiry; acquisition and restore stay supervised.Graceful shutdownStop public ingress; finish admitted workDrain Cells + close native descendantsWithdraw advertisement / release diskWorkspace stays locked through cleanupDeployment maintenanceClose release admission with operation IDFleet drains; prove all Cells settledRecovery worker fences expired ownersExplicit end reopens same releaseBackup and restoreCopy pinned roots + referenced bodiesVerify independent disjoint prefixRestore into unused reserved prefixSame provider; preserve identitynode / deployment controlowner fenceverified recovery bytessame Cell / new ownerOwner lifecycle and exact recoveryA repository keeps its identity when its owner changes. Only authority can select the recovery root.CURRENT SERVING PATH + \ No newline at end of file diff --git a/diagram/canopy-architecture/07-owner-recovery@2x.png b/diagram/canopy-architecture/07-owner-recovery@2x.png new file mode 100644 index 0000000..6a1c24b Binary files /dev/null and b/diagram/canopy-architecture/07-owner-recovery@2x.png differ diff --git a/diagram/canopy-architecture/08-domain-and-policy.svg b/diagram/canopy-architecture/08-domain-and-policy.svg new file mode 100644 index 0000000..d6d9fd5 --- /dev/null +++ b/diagram/canopy-architecture/08-domain-and-policy.svg @@ -0,0 +1,6 @@ + +Repository components and collaboration controlThe Repository Cell colocates Git refs and graph metadata with ACLs, visibility, issues, PRs, reviews, line threads, check attempts, branch rules and operation outcomes. Merge preparation may run native Git outside the authority boundary, but MergePull rechecks current exact revisions, policy and required evidence before atomically changing the branch and PR state. Directory and Repository state remain separate Cells. + + +ONE REPOSITORY CELL / ONE UUID / ONE FENCED WRITEREXAMPLE: MERGE A PULL REQUESTDIRECTORY AND REPOSITORY ARE SEPARATE TRANSACTION DOMAINSGit identity and historyObject format, objects and graphRefs, versions, HEAD, generationAccess and policyOwner / members / visibilityBranch rules and version guardsIssues and discussionsIssues, comments and editsRequest records / paginationPull requests and reviewsExact base/head comparisonsReviews, threads and resolutionChecks and candidatesCommit-bound check attemptsMerge / squash / rebase candidatesDurable operation outcomesPush reports / certificate auditMutation replay and receiptsPrepare candidateAuthenticate and capture base/headNative merge-tree / commit-treePersist candidate objects + closurePublish with MergePullRecheck actor and exact revisionsRequired checks / reviews / threadsBranch policy + ref expectationsCommit and acknowledgeUpdate branch ref + generationClose PR + save merge outcomeCellule gates durable resultPreparation is provisionalA candidate or passed check cannot grant publication. The final commandchecks the current branch and exact candidate/head bindings.CI and web boundariesChecks are product records submitted through the API. Canopy does not install aRepository Cell workflow runner for CI today.Directory CellGlobal accounts / token scopes / SSH keysNames and discovery candidates; pending -> readyRepository CellActual membership, visibility and write policyEach sensitive transaction checks its own authorityUUIDdurable domain statefinal authorization / guardscollaboration logicdurable outcomeRepository components and collaboration controlGit state, authorization and collaboration meet at one Repository Cell transaction boundary.CURRENT SERVING PATH + \ No newline at end of file diff --git a/diagram/canopy-architecture/08-domain-and-policy@2x.png b/diagram/canopy-architecture/08-domain-and-policy@2x.png new file mode 100644 index 0000000..6cb9c1b Binary files /dev/null and b/diagram/canopy-architecture/08-domain-and-policy@2x.png differ diff --git a/diagram/canopy-architecture/09-packed-storage-evolution.svg b/diagram/canopy-architecture/09-packed-storage-evolution.svg new file mode 100644 index 0000000..c842387 --- /dev/null +++ b/diagram/canopy-architecture/09-packed-storage-evolution.svg @@ -0,0 +1,6 @@ + +Packed storage evolution and remaining integrationThe current product still uses per-object SQL authority, with an active archived-pack optimization. The new subsystem moves canonical inventories, graph metadata, ref snapshots and exact outcomes into verified immutable artifacts and short Cell publication commands. Preparation, publication dispatch, certificates, policy pages, recovery records and readers exist, but the full serving path is not registered and the hard cutover remains incomplete. Multi-capability repository Cells are a separate proposal. + + +ACTIVE SERVING MODELNEW PACKED CATALOG MODEL / IMPLEMENTED BUILDING BLOCKSPer-object SQL authorityCanonical objects + edges + closure in Repository CellInline / chunks / external / archived packed blobsCompletePush publishes refs and stored responseCurrent gateway preparationSerialized native push phase within each gatewayVerified immutable body / archive uploadsSQL ingestion batches and graph certificationPrivate native preparationRetain exact input + native resultOwned staging / preparation attemptsPack + index + metadata artifactsVerify canonical identity and closureImmutable catalog metadataSorted directory runs + leveled indexSource roots identify physical packsCatalog binds directory + sourcesImmutable ref-state / outcome rootsPublication coordinatorBounded class / account dispatchRegister exact SDK command identityBind owner fence / attempt / baseTrusted certificates + ref policy pagesShort Cell publicationCheck current ACL / policy / generationPublish catalog + ref root + outcomeRetain generation / recovery factsCellule still supplies durability gatePinned catalog readersVerify catalog / source descriptorsResolve OID -> pack locationIndependent bounded reader facilitiesDo not infer access from pack membershipcertified rootpacks/publication/mod.rs explicitly says these commands are not registered on the legacy repository serving path.Required production workWire startup, HTTP, SSH and generated producers/readers; select the fresh schema;complete recovery, collection, backup and capacity qualification.Separate Cell capability proposalAdding KV, Queue, Workflow, Blob, Cron, Timer and Effects to the same Repository Cellneeds framework composition and autonomous-progress acceptance gates.What does not change in the intended designThe repository UUID, one fenced authoritative publisher, final policy checks and durable outcome acknowledgement remain. Immutable uploaded artifacts are preparation until an authorizedCell command publishes them.active serving pathnew implemented primitivesimmutable metadata / bytesintegration and qualification openPacked storage evolution and remaining integrationThese primitives exist in code; the replacement production request path and fresh-schema cutover remain open.INCOMPLETE PRODUCT INTEGRATION + \ No newline at end of file diff --git a/diagram/canopy-architecture/09-packed-storage-evolution@2x.png b/diagram/canopy-architecture/09-packed-storage-evolution@2x.png new file mode 100644 index 0000000..bdfbffe Binary files /dev/null and b/diagram/canopy-architecture/09-packed-storage-evolution@2x.png differ diff --git a/diagram/canopy-architecture/README.md b/diagram/canopy-architecture/README.md new file mode 100644 index 0000000..0014f59 --- /dev/null +++ b/diagram/canopy-architecture/README.md @@ -0,0 +1,31 @@ +# Canopy architecture diagrams + +Read the root [Canopy design document](../../DESIGN.md) for system goals, Cellule integration, authority boundaries, request and control flows, resource ownership, security, failure handling and design tradeoffs. It is the source of the gallery's detailed explanations. + +Open the [browsable gallery](index.html) to inspect the nine diagrams. Each diagram has a standalone SVG and a PNG at twice its logical resolution. Diagrams 01–08 describe the inspected serving architecture; diagram 09 distinguishes implemented packed-storage primitives from incomplete production integration and proposed repository capabilities. See [design scope and implementation status](../../DESIGN.md#design-scope-and-implementation-status). + +| Diagram | What it explains | SVG | PNG | +| --- | --- | --- | --- | +| 01 | Public clients, node components and storage authority | [System overview](01-system-overview.svg) | [PNG](01-system-overview@2x.png) | +| 02 | How Canopy models Directory and Repository Cells on Cellule | [Framework mapping](02-canopy-on-cellule.svg) | [PNG](02-canopy-on-cellule@2x.png) | +| 03 | Commands, SQLite, LTX and the durable acknowledgement boundary | [Durable command](03-durable-command.svg) | [PNG](03-durable-command@2x.png) | +| 04 | Authentication, residency, local routing and signed peer routing | [Request routing](04-request-routing.svg) | [PNG](04-request-routing@2x.png) | +| 05 | Push preparation, object ingestion, policy checks and exact replay | [Git push](05-git-push.svg) | [PNG](05-git-push@2x.png) | +| 06 | Fetch, browse, LFS and current physical storage choices | [Reads and LFS](06-read-and-lfs.svg) | [PNG](06-read-and-lfs@2x.png) | +| 07 | Startup, leases, takeover, exact restore, drain and backup | [Owner recovery](07-owner-recovery.svg) | [PNG](07-owner-recovery@2x.png) | +| 08 | Repository features and merge publication controls | [Domain and policy](08-domain-and-policy.svg) | [PNG](08-domain-and-policy@2x.png) | +| 09 | Packed catalog primitives, remaining integration and capability proposals | [Storage evolution](09-packed-storage-evolution.svg) | [PNG](09-packed-storage-evolution@2x.png) | + +## Regenerate the diagrams and gallery + +Run from the repository root: + +```sh +python3 diagram/canopy-architecture/generate.py +python3 diagram/canopy-architecture/render_pngs.py +python3 diagram/canopy-architecture/build_gallery.py +``` + +[generate.py](generate.py) owns SVG layouts, labels and source maps. [render_pngs.py](render_pngs.py) requires `rsvg-convert` on `PATH`. [build_gallery.py](build_gallery.py) embeds the SVGs and matching sections from [DESIGN.md](../../DESIGN.md); it links repository files relative to the gallery. The generated gallery works offline, with network access needed only for external source links. + +Sequence diagrams draw activation bars behind dashed lifelines and message arrows. When updating the design, keep the diagram labels, source snapshot, implementation boundaries and generated gallery consistent. See [verification and change obligations](../../DESIGN.md#verification-and-change-obligations). diff --git a/diagram/canopy-architecture/build_gallery.py b/diagram/canopy-architecture/build_gallery.py new file mode 100644 index 0000000..3a0db07 --- /dev/null +++ b/diagram/canopy-architecture/build_gallery.py @@ -0,0 +1,89 @@ +#!/usr/bin/env python3 +"""Build the offline HTML atlas from the diagrams and root design document.""" +from pathlib import Path +from html import escape +import json +import re +from urllib.parse import urlsplit + +ROOT = Path(__file__).resolve().parent +CANOPY_REV = '9438bb865959fb975d5349ba8b9908b461653821' +CELLULE_REV = '161067f5a21703b3e257024bcb64e565fd9657b4' +diagrams = json.loads((ROOT/'manifest.json').read_text()) +guide = (ROOT.parents[1]/'DESIGN.md').read_text() +sections = dict(section.split('\n', 1) for section in + re.split(r'^## ', guide, flags=re.M)[1:]) +design_sections = { + '01-system-overview': 'System components', + '02-canopy-on-cellule': 'Modeling Canopy on Cellule', + '03-durable-command': 'Durable command execution', + '04-request-routing': 'Request routing and residency', + '05-git-push': 'Git push and replay', + '06-read-and-lfs': 'Fetch browse and Git LFS', + '07-owner-recovery': 'Owner recovery and operational control', + '08-domain-and-policy': 'Collaboration and final policy checks', + '09-packed-storage-evolution': 'Packed storage and framework capability evolution', +} +prose = {} +for slug, heading in design_sections.items(): + blocks = [] + for paragraph in sections[heading].split('\n\n'): + paragraph = paragraph.strip() + if not paragraph or paragraph.startswith(('![', '```')): + continue + if paragraph.startswith('|'): + rows = [[escape(cell.strip()) for cell in row.strip().strip('|').split('|')] + for row in paragraph.splitlines()] + header = ''+''.join(''+cell+'' for cell in rows[0])+'' + body = ''+''.join(''+''.join(''+cell+'' for cell in row)+'' + for row in rows[2:])+'' + blocks.append('
'+header+body+'
') + elif paragraph.startswith('### '): + blocks.append('

'+escape(paragraph[4:])+'

') + else: + blocks.append('

'+escape(paragraph)+'

') + prose[slug] = '\n'.join(blocks) + +def inline_markup(value): + value = re.sub(r'`([^`]+)`',r'\1',value) + def link(match): + destination = match.group(2) + if not urlsplit(destination).scheme and not destination.startswith('//'): + destination = ('../../DESIGN.md' if destination.startswith('#') else '../../') + destination + return f'{match.group(1)}' + return re.sub(r'\[([^\]]+)\]\(([^)]+)\)',link,value) + +nav = ''.join(f'{i:02}{escape(d["title"])}' for i,d in enumerate(diagrams,1)) +articles = [] +for i,d in enumerate(diagrams): + slug = d['slug'] + svg = (ROOT/f'{slug}.svg').read_text() + ids = re.findall(r'\bid="([^"]+)"',svg) + for original in ids: + svg = svg.replace(f'id="{original}"',f'id="{slug}-{original}"').replace(f'url(#{original})',f'url(#{slug}-{original})') + svg = svg.replace('aria-labelledby="title desc"',f'aria-labelledby="{slug}-title {slug}-desc"') + sources = ''.join(f'
  • {escape(s)}
  • ' for s in d['sources']) + articles.append(f'''
    +
    {i+1:02}

    {escape(d['title'])}

    +

    {escape(d['explanation'])}

    +
    100%SVGPNG @2x
    {svg}
    +
    Read the explanation and source code
    {inline_markup(prose[slug])}

    Read the complete design section

    Source map at the inspected revision

      {sources}
    +
    ''') + +html = '''Canopy architecture and Cellule integration +
    +

    Canopy architecture and Cellule integration

    Trace a request from a Git client or browser to its authoritative repository state. See which components Canopy owns, what Cellule supplies, and how writes survive owner changes and local disk loss.

    +

    Source snapshot: October 4, 2026 · Canopy 9438bb8 · Cellule 161067f

    +
    Diagrams 01–08 describe the current serving architecture. Diagram 09 separates newer implemented packed-storage primitives from incomplete production integration. Composed repository queues and workflows remain a separate proposal. These diagrams do not assert production capacity.
    +
    Product boundaryCanopy owns ingress, Git, auth, collaboration and operational policy.
    Framework boundaryCellule supplies stable targets, fenced execution, durable outcomes and restoration.
    Authority boundaryCells and published roots preserve state. Native Git files are rebuildable.
    +

    Use + and − to inspect a diagram; drag the scrollbar to pan. Fit restores the overview. Each figure has standalone SVG and PNG exports and expandable explanations with source links.

    +'''+''.join(articles)+'''
    +''' +(ROOT/'index.html').write_text(html) +print(f'Built self-contained gallery with {len(diagrams)} diagrams') diff --git a/diagram/canopy-architecture/generate.py b/diagram/canopy-architecture/generate.py new file mode 100644 index 0000000..b89cdab --- /dev/null +++ b/diagram/canopy-architecture/generate.py @@ -0,0 +1,339 @@ +#!/usr/bin/env python3 +"""Regenerate the source-based Canopy architecture atlas (stdlib only).""" +from pathlib import Path +from html import escape as esc +import json +import textwrap + +OUT = Path(__file__).resolve().parent +COLORS = { + 'cyan': ('#083344', '#22d3ee'), 'green': ('#064e3b', '#34d399'), + 'violet': ('#4c1d95', '#a78bfa'), 'amber': ('#78350f', '#fbbf24'), + 'rose': ('#881337', '#fb7185'), 'orange': ('#7c2d12', '#fb923c'), + 'slate': ('#1e293b', '#94a3b8'), 'blue': ('#1e3a8a', '#60a5fa'), +} +ATLAS = [] + +class Diagram: + def __init__(self, slug, title, subtitle, height=850, width=1100, status='CURRENT SERVING PATH', sequence=False): + self.slug, self.title, self.subtitle = slug, title, subtitle + self.w, self.h, self.status = width, height, status + self.sequence = sequence + self.regions, self.edges, self.nodes, self.labels = [], [], [], [] + self.activations, self.lifelines = [], [] + self.foreground_edges = [] + self.boxes = [] + + def txt(self, x, y, value, size=9, color='#94a3b8', anchor='start', weight=400, layer=None): + (self.labels if layer is None else layer).append( + f'{esc(value)}') + + def region(self, x, y, w, h, label, color='amber'): + self.regions.append(f'') + self.txt(x+16, y+20, label, color=COLORS[color][1], weight=600, layer=self.regions) + + def box(self, x, y, w, h, title, lines=(), color='green', dashed=False): + fill, stroke = COLORS[color] + self.boxes.append((x,y,w,h,title)) + self.nodes.append(f'') + self.nodes.append(f'') + titles = title if isinstance(title, list) else [title] + ytext = y+25 + for line in titles: + assert len(line)*7.2 <= w-18, (self.slug, title, 'title too wide') + self.txt(x+w/2, ytext, line, 12, '#f1f5f9', 'middle', 600) + ytext += 17 + if lines: + ytext += 3 + for line in lines: + assert len(line)*5.4 <= w-18, (self.slug, line, 'body too wide') + self.txt(x+w/2, ytext, line, 9, '#b4c1d3', 'middle') + ytext += 15 + assert ytext-15 <= y+h-10, (self.slug, title, 'text too tall') + + def note(self, x, y, w, title, text, color='slate', h=None): + lines = textwrap.wrap(text, int((w-30)/5.4), break_long_words=False, break_on_hyphens=False) + self.box(x,y,w,h or 48+15*len(lines),title,lines,color) + + def path(self, points, label=None, color='slate', dashed=False, lx=None, ly=None, foreground=False): + stroke = COLORS[color][1] + d = 'M '+' L '.join(f'{x},{y}' for x,y in points) + # Every sequence message belongs above both activation bars and lifelines. + layer = self.foreground_edges if self.sequence or foreground else self.edges + layer.append(f'') + if label: + self.txt(lx if lx is not None else (points[0][0]+points[-1][0])/2, + ly if ly is not None else (points[0][1]+points[-1][1])/2-9, + label, 8, stroke, 'middle') + + def lifeline(self, x, start, end): + self.lifelines.append(f'') + + def activation(self, x, y, height, color): + self.activations.append(f'') + + def diamond(self, cx, cy, w, h, lines): + self.nodes.append(f'') + self.nodes.append(f'') + for i,line in enumerate(lines): + self.txt(cx,cy+(i-(len(lines)-1)/2)*15+4,line,10,'#f1f5f9','middle',600) + + def legend(self, entries, y=None): + y = y or self.h-40 + x = 35 + for color,label in entries: + self.labels.append(f'') + self.txt(x+16,y,label,8) + x += len(label)*4.8+46 + + def save(self, explanation, sources): + for x,y,w,h,t in self.boxes: + assert x>=30 and y>=85 and x+w<=self.w-30 and y+h<=self.h-60, (self.slug,t,'outside drawing area') + markers = ''.join(f'' for c,p in COLORS.items()) + heading = f'{esc(self.title)}{esc(self.subtitle)}{esc(self.status)}' + svg = f''' +{esc(self.title)}{esc(explanation)} +{markers} + +{''.join(self.regions)}{''.join(self.edges)}{''.join(self.activations)}{''.join(self.nodes)}{''.join(self.lifelines)}{''.join(self.foreground_edges)}{''.join(self.labels)}{heading} +''' + (OUT/f'{self.slug}.svg').write_text(svg) + ATLAS.append(dict(slug=self.slug,title=self.title,explanation=explanation,sources=sources)) + + +# 1. Product components and storage authority. +d = Diagram('01-system-overview','Canopy system architecture','Clients enter any node; Cell authority and immutable storage preserve repository state.', 900) +d.region(230,95,465,550,'CANOPY NODE / ONE RUST SERVICE') +d.region(735,95,335,330,'DURABLE CELL AUTHORITY','violet') +d.box(35,145,155,75,'Git clients',['Smart HTTP / SSH','clone, fetch, push'],'cyan') +d.box(35,285,155,75,'Browser / API',['Embedded web UI','JSON collaboration API'],'cyan') +d.box(35,425,155,75,'Git LFS',['HTTP bodies / locks','SSH can issue grants'],'cyan') +d.box(255,145,185,90,'Ingress',['Axum HTTP router','Optional russh listener','Account authentication'],'cyan') +d.box(480,145,185,90,'RepositoryManager',['Name resolution','Residency + request pins','Local / peer Cell handles'],'orange') +d.box(255,290,410,100,'Product services',['GitGateway + LfsService + repository_http','Browse, issues, PRs, reviews, checks, policy','Native Git work runs on the receiving node']) +d.box(255,450,410,100,'Cellule integration',['CanopyApplication + CellNode','Typed SQL / domain commands','Fenced ownership + durable output gate'],'blue') +d.box(765,145,275,95,'Directory Cell',['One fixed SQL shard per tenant','Names -> stable repository UUID','Accounts, tokens, keys, audit'],'violet') +d.box(765,300,275,95,'Repository Cell',['One SQL Cell per repository UUID','Refs, objects, ACL, collaboration','Outcomes, closure and policy'],'violet') +d.box(765,485,275,100,'Disposable Git cache',['Private refs + verified shared objects','Native receive/upload-pack workers','Can be rebuilt after local disk loss'],'slate',True) +d.box(255,685,785,100,'Shared S3-compatible object store',['Cellule: catalog, authority, node leases, LTX roots and immutable SQLite recovery bytes','Canopy: external Git blobs, archived pack/index bytes and Git LFS bodies','Conditional writes select authority; verified immutable bytes supply recovery'],'violet') +d.path([(190,182),(255,182)],color='cyan') +d.path([(190,322),(212,322),(212,200),(255,200)],color='cyan') +d.path([(190,462),(220,462),(220,215),(255,215)],color='cyan') +d.path([(440,185),(480,185)],color='orange') +d.path([(665,175),(765,175)],'resolve / auth','violet') +d.path([(665,205),(714,205),(714,345),(765,345)],color='violet') +d.path([(575,235),(575,290)],color='green') +d.path([(460,390),(460,450)],'typed operations','blue') +d.path([(665,340),(715,340),(715,525),(765,525)],'native work','slate',lx=718,ly=453) +d.path([(665,500),(705,500),(705,375),(765,375)],color='blue') +d.path([(450,550),(450,685)],'durable publication / external body I/O','violet',lx=520,ly=628) +d.path([(1000,395),(1000,440),(1057,440),(1057,650),(850,650),(850,685)],color='violet') +d.path([(922,395),(922,445),(902,445),(902,485)],'hydrate','slate',True,lx=958,ly=460) +d.note(35,690,175,'Ownership rule','Git success is acknowledged only after the Cell commits and passes its durability gate.', 'amber') +d.legend([('cyan','public clients'),('green','product logic'),('blue','framework'),('violet','durable authority'),('slate','rebuildable local files')]) +d.save('Canopy is an embedded Cellule application. Public ingress and native Git run in Canopy, while each durable Directory or Repository Cell has one fenced owner. A receiving node may call a remote owner. The object store holds authority records, SQLite recovery roots and product body objects. Native Git caches are disposable.', ['crates/canopy-server/src/server/mod.rs','crates/canopy-server/src/server/residency/mod.rs','crates/canopy-server/src/git_gateway/mod.rs','crates/canopy-server/src/lib.rs','crates/canopy-server/src/schema.sql']) + +# 2. Exact mapping from Canopy domain to framework topology. +d=Diagram('02-canopy-on-cellule','How Canopy models its domain on Cellule','Domain code supplies schema and policy; the embedded framework supplies identity, execution and recovery.',1020) +d.region(35,95,1030,185,'APPLICATION DECLARATION / COMPILED AT STARTUP','green') +d.box(60,140,260,105,'CanopyApplication',['CellApplication::register','DirectoryModule + RepositoryModule','Two declared SQL Cell types']) +d.box(360,140,315,105,'DirectoryModule',['CatalogRole::Sql / one fixed shard','Namespace = [0x48; 16]','Names, accounts, credentials']) +d.box(715,140,325,105,'RepositoryModule',['CatalogRole::Sql / entity partitions','Namespace = [0x47; 16]','Git + ACL + collaboration commands']) +d.region(35,330,1030,190,'RUNTIME ADDRESSING / STABLE ACROSS OWNER MOVEMENT','blue') +d.box(60,380,410,105,'Directory target',['Tenant + application + namespace','partition_for_shard(0)','DirectoryCell wraps SqlCell'],'blue') +d.box(545,380,495,105,'Repository target',['Tenant + application + namespace + UUID-derived partition','CellType::entity_partition -> 33-byte partition','RepositoryCell wraps SqlCell'],'blue') +d.path([(518,245),(518,310),(265,310),(265,380)],color='blue') +d.path([(875,245),(875,380)],'UUID selects one Cell','blue',lx=949,ly=312) +d.path([(265,485),(265,605)],color='blue') +d.path([(795,485),(795,540),(545,540),(545,605)],color='blue') +d.region(35,560,1030,230,'EMBEDDED CELLULE FRAMEWORK / THE PINNED DEPENDENCY','amber') +d.box(60,605,290,95,'cellule-app + host',['Descriptor / registry / typed handles','CellNode lifecycle + readiness','Canopy installs facilities and leases'],'blue') +d.box(390,605,310,95,'cellule-runtime',['Per-Cell actor + SQL worker','Authority fence + request ledger','Queries, commands, peer contracts'],'orange') +d.box(740,605,300,95,'cellule-ltx + store',['Managed SQLite WAL capture','Verified roots + exact restoration','Bounded I/O + conditional updates'],'violet') +d.path([(350,650),(390,650)],color='orange') +d.path([(700,650),(740,650)],color='violet') +d.txt(60,746,'Dependency pin: 161067f5a21703b3e257024bcb64e565fd9657b4',9,'#e2e8f0') +d.txt(60,767,'cellule-types is transitive. Canopy supplies its own signed HTTPS peer transport.',9) +d.note(35,840,490,'Framework capabilities','Cellule includes SQL, KV, Queue, Workflow, Blob, Cron, Timer and Effects. Capabilities use typed APIs and explicit topology.', 'blue',115) +d.note(575,840,490,'Canopy capability boundary','Canopy registers SQL Cells today. Multi-capability Repository Cells, repository queues and autonomous workflows remain a proposal.', 'amber',115) +d.legend([('green','domain declarations'),('blue','framework binding'),('orange','fenced execution'),('amber','capabilities not wired into Canopy')]) +d.save('CanopyApplication registers DirectoryModule and RepositoryModule as SQL Cell types. Directory uses a fixed shard; repository UUIDs derive entity partitions. Product wrappers hold typed SqlCell handles. Cellule owns execution and persistence mechanics, and Canopy owns protocols and authorization. Other framework primitives are available but are not composed into Canopy Repository Cells.', ['crates/canopy-server/src/lib.rs','crates/canopy-server/src/directory/mod.rs','crates/canopy-server/Cargo.toml','docs/repository-cell-primitives.md']) + +# 3. Framework command execution and the real acknowledgement boundary. +d=Diagram('03-durable-command','Cellule command execution and durability','One command changes one Cell. The domain mutation and its recorded answer share a SQLite transaction.',1040,width=1140,sequence=True) +xs=[115,340,565,795,1025] +actors=[('Canopy caller','Typed handle','cyan'),('Cell owner','Actor + fence','orange'),('Managed SQLite','State + ledger','green'),('LTX / object store','Immutable roots','violet'),('CellAuthority','Control CAS','amber')] +for x,(t,s,c) in zip(xs,actors): + d.box(x-85,110,170,65,t,[s],c) + d.lifeline(x,175,835) +for x,y,h,c in [(340,220,550,'orange'),(565,370,95,'green'),(795,510,95,'violet'),(1025,650,65,'amber')]: + d.activation(x,y,h,c) +msgs=[(0,1,225,'1 Command + stable MutationIdentity','cyan',False), + (1,4,290,'2 Validate owner incarnation / lease / compatible code','amber',False), + (1,2,375,'3 Execute domain mutation + request outcome in one transaction','green',False), + (2,1,445,'4 Commit SQLite and capture its exact WAL boundary','green',True), + (1,3,525,'5 Verify capture; upload immutable chunks and proposed root','violet',False), + (3,1,595,'6 Return proposed recovery root','violet',True), + (1,4,660,'7 Conditional publish of the exact root under owner fence','amber',False), + (4,1,730,'8 Accepted authority revision = durable publication proof','amber',True), + (1,0,795,'9 Committed + Receipt','blue',True)] +# Sequence layers: activation bars, dashed lifelines, then message arrows. +for a,b,y,l,c,ret in msgs:d.path([(xs[a],y),(xs[b],y)],l,c,ret) +d.note(35,865,490,'Lost reply or rejected fence','A timeout does not prove failure. Resolve the original request identity; retry the same logical command only under its existing identity.', 'rose',110) +d.note(575,865,490,'Receipt and scope','A receipt-bound query must observe the published per-Cell position. Directory and Repository commands do not form one cross-Cell SQL transaction.', 'blue',110) +d.legend([('cyan','invocation'),('green','local transaction'),('violet','immutable bytes'),('amber','authority publication'),('blue','durable acknowledgement')]) +d.save('This is the object-store durability path assembled by Canopy. The Cell owner records domain state and the request outcome together, captures SQLite through LTX and publishes the exact root using a fenced conditional authority update. Only then does the caller receive Committed and a receipt. A Cellule follower-log mode exists, but this diagram does not claim Canopy enables it.', ['crates/canopy-server/src/server/mod.rs','crates/canopy-server/src/lib.rs'],) + +# 4. Request routing with admission and takeover branches. +d=Diagram('04-request-routing','Request routing and repository residency','Authentication, name lookup, account admission and owner resolution precede repository work.',1190) +d.box(400,105,300,65,'Incoming request',['HTTP / SSH / API / Git LFS'],'cyan') +d.box(400,230,300,85,'Directory Cell',['Validate token / key / LFS grant','Resolve ready owner/name -> UUID'],'violet') +d.box(400,375,300,85,'RepositoryManager',['Derive target; pin resident route','Otherwise admit a transition'],'orange') +d.diamond(550,575,210,85,['Live owner','elsewhere?']) +d.box(60,695,350,100,'Remote Cell binding',['CellClient::peer -> /internal/cell','Signed request + enrolled node key','TLS, release, expiry, principal checks'],'blue') +d.box(690,695,350,100,'Local acquisition',['Bootstrap new / acquire idle Cell','Expired owner: fence + restore exact root','CellClient::local -> resident actor'],'blue') +d.box(400,900,300,95,'Repository operations',['Recheck current ACL / visibility','Run domain command or native Git','Retain route pin through response']) +d.path([(550,170),(550,230)],color='cyan') +d.path([(550,315),(550,375)],color='violet') +d.path([(550,460),(550,532)],color='orange') +d.path([(445,575),(235,575),(235,695)],'Yes','blue',lx=330,ly=565) +d.path([(655,575),(865,575),(865,695)],'No','blue',lx=758,ly=565) +d.path([(235,795),(235,850),(485,850),(485,900)],color='blue') +d.path([(865,795),(865,850),(615,850),(615,900)],color='blue') +d.note(35,230,285,'Discovery is a hint','Directory listing candidates are checked against current Repository Cell access. Cached names and UI state cannot grant permissions.', 'slate',135) +d.note(785,375,280,'Bounded transitions','Cold/remote work uses supervised admission and a per-repository transition lock. Different repositories can activate concurrently.', 'amber',135) +d.note(35,1040,490,'Capacity and cancellation','An unavailable slot or in-progress movement can return 503. Streamed responses and detached admitted work retain pins until they finish.', 'amber',85) +d.note(575,1040,490,'Native work placement','Remote binding moves Cell calls. Git workers, body streaming and the disposable cache can remain on the receiving gateway.', 'slate',85) +d.legend([('cyan','request'),('violet','identity / authority'),('orange','admission'),('blue','local or signed remote execution')]) +d.save('Directory state authenticates accounts and resolves ready repository names. RepositoryManager binds a stable target to either a local actor or a signed peer client. Cold transitions are admitted and supervised; remote-owner loss triggers acquisition on demand. Repository authorization is checked using current durable state. Response pins prevent eviction during streaming.', ['crates/canopy-server/src/server/residency/mod.rs','crates/canopy-server/src/server/peer.rs','crates/canopy-server/src/repository_http/authorization.rs','crates/canopy-server/src/server/discovery.rs']) + +# 5. Production receive-pack sequence, including archives on the serving path. +d=Diagram('05-git-push','Git push from wire input to durable refs','Current serving path: native Git prepares private state; CompletePush publishes refs and the saved result.',1300,width=1140,sequence=True) +xs=[110,335,560,800,1030] +for x,(t,s,c) in zip(xs,[('Git client','receive-pack','cyan'),('GitGateway','Receiving node','green'),('Native Git','Disposable refs','slate'),('Repository Cell','Via local / peer client','blue'),('Object store','Verified body bytes','violet')]): + d.box(x-80,110,160,65,t,[s],c) + d.lifeline(x,175,1045) +msgs=[(0,1,225,'1 Spool encoded input; bind push UUID + actor + request digest','cyan',False), + (1,3,285,'2 begin_push: claim logical ID or find completed response','blue',False), + (3,1,345,'3 Completed ID replays saved result; fresh ID continues','blue',True), + (1,3,405,'4 Read consistent refs, generation and branch policy','blue',False), + (1,2,465,'5 Decode / validate; receive-pack in private cache','green',False), + (2,1,525,'6 Native report + actual accepted ref differences','slate',True), + (1,4,585,'7 Upload verified external blobs or archived pack/index bytes','violet',False), + (1,3,645,'8 Persist canonical object records in bounded batches','blue',False), + (1,3,705,'9 Certify typed graph closure; stage response + ref plan','blue',False), + (1,3,765,'10 CompletePush: recheck ACL, branch rules, OIDs and versions','green',False), + (3,4,835,'11 Cellule publishes SQLite root','violet',False), + (4,3,895,'12 Fenced durable proof','violet',True), + (3,1,955,'13 Read canonical saved response','blue',True), + (1,0,1015,'14 Git report + push identity','cyan',True)] +for a,b,y,l,c,ret in msgs:d.path([(xs[a],y),(xs[b],y)],l,c,ret) +d.region(35,720,1030,100,'FINAL DOMAIN TRANSACTION: ACCEPTED REFS + GENERATION + OUTCOME POINTER / AUDIT','green') +d.note(35,1080,490,'Two identities matter','The HTTP push UUID identifies the whole wire operation. Cellule MutationIdentity identifies each durable command within that operation.', 'amber',125) +d.note(575,1080,490,'Failure behavior','Preparation changes disposable refs only. Objects may remain unreferenced after refusal. An uncertain final publication is resolved, never turned into a false negative report.', 'rose',125) +d.txt(35,1230,'Native partial acceptance is preserved; Git --atomic can request all-or-none validation. Current gateway push work uses a mutex.',9) +d.legend([('cyan','wire protocol'),('slate','private native state'),('blue','durable Cell calls'),('violet','body / root publication')]) +d.save('After outer authentication, GitGateway spools and hashes the encoded request and binds a logical push ID to the actor and digest. It can replay a completed push before decoding and native work. New attempts prepare a private ref snapshot, run receive-pack, ingest canonical objects, certify graph closure and stage the report and ref plan. CompletePush checks current policy and expected ref versions and commits accepted refs with the canonical response pointer. Cellule gates acknowledgement on durable publication.', ['crates/canopy-server/src/git_gateway/mod.rs','crates/canopy-server/src/git_gateway/preflight.rs','crates/canopy-server/src/git_gateway/push.rs','crates/canopy-server/src/push/mod.rs','crates/canopy-server/src/refs.rs']) + +# 6. Separate read and LFS paths, plus exact physical storage choices. +d=Diagram('06-read-and-lfs','Fetch, browse and Git LFS data paths','Reads verify identities against Cell metadata. LFS body delivery bypasses native Git.',1220) +d.region(35,95,1030,330,'GIT CLONE / FETCH','cyan') +d.box(60,150,270,85,'Authorize and select',['Current ACL / public visibility','Validate wants against live refs'],'cyan') +d.box(390,150,310,85,'Snapshot and hydrate',['Generation-consistent refs + HEAD','Verify required objects / bodies','Apply supported partial-clone filter']) +d.box(760,150,280,85,'Native upload-pack',['Private ref snapshot','Shared verified object files','Stream generated pack to client'],'slate') +d.path([(330,192),(390,192)],color='cyan') +d.path([(700,192),(760,192)],color='green') +d.note(60,290,440,'Discovery fast path','Git v2 capability discovery needs no object hydration. Ref discovery prepares ref targets and annotated tag chains.', 'blue',95) +d.note(570,290,470,'Selective fetch','Blobless fetch omits ordinary blobs. Selected wants, ref targets and structural history determine hydration; warm verified objects are reused.', 'slate',95) +d.region(35,475,1030,245,'BROWSE / JSON READS / LFS','green') +d.box(60,530,280,115,'Repository browser',['Read trees, blobs, history and diffs','Use Repository Cell metadata','Load bodies from verified readers','No receive-pack is needed']) +d.box(400,530,280,115,'LFS upload',['Authorize batch / PUT','Hash SHA-256 + bounded parts','Publish immutable body, verify','Then commit authorized metadata'],'violet') +d.box(740,530,300,115,'LFS download',['Authorize against current Cell','Read manifest pinned in SQLite','Verify requested parts / digests','Support tail Range / 206 response'],'violet') +d.txt(60,689,'LFS locks are advisory records in the Repository Cell. SSH grants still deliver LFS bytes over HTTP.',9) +d.region(35,770,1030,270,'CURRENT SERVING STORAGE / AUTHORITATIVE METADATA REMAINS PER OBJECT','violet') +d.box(60,825,290,155,'Repository SQLite',['objects: kind, size, digest, storage','inline bodies <= 768 KiB','Large non-blob bodies: SQL chunks','refs + graph edges / certificates','LFS metadata + locks'],'violet') +d.box(390,825,310,155,'Immutable external bytes',['Large loose Git blobs + manifests','Archived pack and index bodies','Packed blobs reference approved packs','Git LFS bodies + part digests','Readers verify canonical identities'],'violet') +d.box(740,825,300,155,'Local cache',['Hydrated Git files / installed packs','Insertion cursor + ref generations','Private native process workspace','Disk and native-worker admission','Disposable after recovery'],'slate',True) +d.path([(350,903),(390,903)],'references','violet') +d.path([(700,903),(740,903)],'verify / hydrate','slate',True) +d.txt(60,1014,'This active archive path differs from the future immutable catalog / ref-root hard cutover in diagram 09.',9,'#fbbf24') +d.note(35,1080,1030,'Read authority','An uploaded pack, cached object, stale listing or UI view cannot independently authorize an object read. Current Cell access and verified published metadata govern visibility.', 'amber',65) +d.legend([('cyan','stock Git reads'),('green','browse / selection'),('violet','durable state and bodies'),('slate','rebuildable work files')],y=1180) +d.save('Fetch takes consistent refs, validates requested object reachability, hydrates selected verified objects and uses native upload-pack to stream the wire response. Browser APIs read through Cell metadata and verified body readers. LFS uploads first publish and verify immutable bytes, then commit authorized metadata; downloads verify manifest-bound parts. The active serving schema supports inline, chunked, external and packed storage records, distinct from the incomplete immutable catalog replacement.', ['crates/canopy-server/src/git_gateway/fetch.rs','crates/canopy-server/src/git_gateway/hydration.rs','crates/canopy-server/src/git_gateway/discovery.rs','crates/canopy-server/src/git_read/mod.rs','crates/canopy-server/src/lfs/upload.rs','crates/canopy-server/src/lfs/read.rs','crates/canopy-server/src/schema.sql']) + +# 7. Lifecycle controls, fencing and exact restoration. +d=Diagram('07-owner-recovery','Owner lifecycle and exact recovery','A repository keeps its identity when its owner changes. Only authority can select the recovery root.',1150) +d.region(35,95,1030,180,'STARTUP AND SERVING','amber') +d.box(60,140,290,95,'Startup preflight',['Lock managed local workspace','Probe conditional writes / ranged I/O','Compile and check selected release'],'amber') +d.box(390,140,310,95,'Enroll and renew node',['Signed live node advertisement','NodeLeaseGuard + compatible registry','Start Directory + on-demand repos'],'amber') +d.box(740,140,300,95,'Serve with fencing',['Ready only while leases are valid','Per-Cell authority / actor / SQL','Cancel ingress when lease fails'],'green') +d.path([(350,187),(390,187)],color='amber') +d.path([(700,187),(740,187)],color='amber') +d.region(35,325,1030,455,'COLD ACTIVATION / NODE LOSS / LOCAL DISK LOSS','blue') +d.box(60,380,290,100,'Read catalog and control',['Validate compiled code and schema','Live owner: route to that owner','Idle or expired: acquire authority'],'blue') +d.box(390,380,310,100,'Fence and choose exact root',['Claim expired node for takeover','Use CellAuthority-pinned root','Do not elect a root by listing keys'],'orange') +d.box(740,380,300,100,'Verify and restore SQLite',['LTX chunk checksums / endpoints','Recover exact state + outcome ledger','Invalid bytes: fail without activation'],'violet') +d.box(390,590,310,105,'Activate successor',['Same CellTarget / repository UUID','New fenced owner / incarnation','Recovered refs, ACL and outcomes'],'blue') +d.box(740,590,300,105,'Rebuild Git cache',['Hydrate verified published objects','Recreate private ref snapshots','Resume Git/API/LFS requests'],'slate',True) +d.path([(350,430),(390,430)],color='blue') +d.path([(700,430),(740,430)],color='violet') +d.path([(890,480),(890,540),(545,540),(545,590)],color='violet') +d.path([(700,642),(740,642)],color='slate',dashed=True) +d.note(60,590,290,'Cache loss is recoverable','Cell publication protects acknowledged state. Files from an old workspace never become a substitute recovery root.', 'rose',105) +d.txt(60,749,'Recovery is on demand. An unclean owner loss waits for lease expiry; acquisition and restore stay supervised.',9) +d.region(35,830,1030,220,'DRAIN, MAINTENANCE AND INDEPENDENT BACKUP','amber') +d.box(60,880,290,110,'Graceful shutdown',['Stop public ingress; finish admitted work','Drain Cells + close native descendants','Withdraw advertisement / release disk','Workspace stays locked through cleanup'],'amber') +d.box(390,880,310,110,'Deployment maintenance',['Close release admission with operation ID','Fleet drains; prove all Cells settled','Recovery worker fences expired owners','Explicit end reopens same release'],'amber') +d.box(740,880,300,110,'Backup and restore',['Copy pinned roots + referenced bodies','Verify independent disjoint prefix','Restore into unused reserved prefix','Same provider; preserve identity'],'violet') +d.legend([('amber','node / deployment control'),('orange','owner fence'),('violet','verified recovery bytes'),('blue','same Cell / new owner')]) +d.save('Canopy locks its local runtime workspace, verifies storage capabilities, validates the selected release and enrolls a signed node lease before serving. New or cold Cells bootstrap, acquire idle authority or fence an expired owner and restore the authority-pinned SQLite root. Recovery preserves the request outcome ledger. Native caches rebuild afterward. Shutdown and maintenance supervise all admitted work; backup copies durable roots and referenced external bodies into an independent prefix.', ['crates/canopy-server/src/server/mod.rs','crates/canopy-server/src/server/lifecycle.rs','crates/canopy-server/src/server/workspace/mod.rs','crates/canopy-server/src/deployment/recovery.rs','crates/canopy-server/src/deployment/backup/mod.rs','docs/operations.md']) + +# 8. Domain composition and example merge control flow. +d=Diagram('08-domain-and-policy','Repository components and collaboration control','Git state, authorization and collaboration meet at one Repository Cell transaction boundary.',1130) +d.region(35,95,1030,350,'ONE REPOSITORY CELL / ONE UUID / ONE FENCED WRITER','violet') +for x,y,title,lines,c in [ + (60,150,'Git identity and history',['Object format, objects and graph','Refs, versions, HEAD, generation'],'violet'), + (400,150,'Access and policy',['Owner / members / visibility','Branch rules and version guards'],'rose'), + (740,150,'Issues and discussions',['Issues, comments and edits','Request records / pagination'],'green'), + (60,295,'Pull requests and reviews',['Exact base/head comparisons','Reviews, threads and resolution'],'green'), + (400,295,'Checks and candidates',['Commit-bound check attempts','Merge / squash / rebase candidates'],'green'), + (740,295,'Durable operation outcomes',['Push reports / certificate audit','Mutation replay and receipts'],'blue')]:d.box(x,y,300,100,title,lines,c) +d.region(35,495,1030,335,'EXAMPLE: MERGE A PULL REQUEST','green') +d.box(60,550,300,105,'Prepare candidate',['Authenticate and capture base/head','Native merge-tree / commit-tree','Persist candidate objects + closure']) +d.box(400,550,300,105,'Publish with MergePull',['Recheck actor and exact revisions','Required checks / reviews / threads','Branch policy + ref expectations'],'rose') +d.box(740,550,300,105,'Commit and acknowledge',['Update branch ref + generation','Close PR + save merge outcome','Cellule gates durable result'],'blue') +d.path([(360,602),(400,602)],color='green') +d.path([(700,602),(740,602)],color='blue') +d.note(60,710,440,'Preparation is provisional','A candidate or passed check cannot grant publication. The final command checks the current branch and exact candidate/head bindings.', 'rose',85) +d.note(570,710,470,'CI and web boundaries','Checks are product records submitted through the API. Canopy does not install a Repository Cell workflow runner for CI today.', 'amber',85) +d.region(35,880,1030,175,'DIRECTORY AND REPOSITORY ARE SEPARATE TRANSACTION DOMAINS','orange') +d.box(60,930,440,85,'Directory Cell',['Global accounts / token scopes / SSH keys','Names and discovery candidates; pending -> ready'],'violet') +d.box(570,930,470,85,'Repository Cell',['Actual membership, visibility and write policy','Each sensitive transaction checks its own authority'],'violet') +d.path([(500,974),(570,974)],'UUID','orange') +d.legend([('violet','durable domain state'),('rose','final authorization / guards'),('green','collaboration logic'),('blue','durable outcome')]) +d.save('The Repository Cell colocates Git refs and graph metadata with ACLs, visibility, issues, PRs, reviews, line threads, check attempts, branch rules and operation outcomes. Merge preparation may run native Git outside the authority boundary, but MergePull rechecks current exact revisions, policy and required evidence before atomically changing the branch and PR state. Directory and Repository state remain separate Cells.', ['crates/canopy-server/src/schema.sql','crates/canopy-server/src/pulls/merge/command.rs','crates/canopy-server/src/pulls/candidates/mod.rs','crates/canopy-server/src/branch_rules/command.rs','crates/canopy-server/src/server/discovery.rs']) + +# 9. Implemented primitives versus production cutover and capability proposal. +d=Diagram('09-packed-storage-evolution','Packed storage evolution and remaining integration','These primitives exist in code; the replacement production request path and fresh-schema cutover remain open.',1220,status='INCOMPLETE PRODUCT INTEGRATION') +d.region(35,95,1030,185,'ACTIVE SERVING MODEL','green') +d.box(60,140,440,95,'Per-object SQL authority',['Canonical objects + edges + closure in Repository Cell','Inline / chunks / external / archived packed blobs','CompletePush publishes refs and stored response']) +d.box(570,140,470,95,'Current gateway preparation',['Serialized native push phase within each gateway','Verified immutable body / archive uploads','SQL ingestion batches and graph certification']) +d.region(35,330,1030,475,'NEW PACKED CATALOG MODEL / IMPLEMENTED BUILDING BLOCKS','blue') +d.box(60,380,290,110,'Private native preparation',['Retain exact input + native result','Owned staging / preparation attempts','Pack + index + metadata artifacts','Verify canonical identity and closure'],'slate') +d.box(390,380,310,110,'Immutable catalog metadata',['Sorted directory runs + leveled index','Source roots identify physical packs','Catalog binds directory + sources','Immutable ref-state / outcome roots'],'violet') +d.box(740,380,300,110,'Publication coordinator',['Bounded class / account dispatch','Register exact SDK command identity','Bind owner fence / attempt / base','Trusted certificates + ref policy pages'],'orange') +d.box(390,590,310,115,'Short Cell publication',['Check current ACL / policy / generation','Publish catalog + ref root + outcome','Retain generation / recovery facts','Cellule still supplies durability gate'],'blue') +d.box(740,590,300,115,'Pinned catalog readers',['Verify catalog / source descriptors','Resolve OID -> pack location','Independent bounded reader facilities','Do not infer access from pack membership'],'violet') +d.path([(350,435),(390,435)],color='violet') +d.path([(700,435),(740,435)],color='orange') +d.path([(890,490),(890,540),(545,540),(545,590)],color='orange') +d.path([(700,648),(740,648)],'certified root','violet') +d.txt(60,758,'packs/publication/mod.rs explicitly says these commands are not registered on the legacy repository serving path.',9,'#fbbf24') +d.note(35,855,490,'Required production work','Wire startup, HTTP, SSH and generated producers/readers; select the fresh schema; complete recovery, collection, backup and capacity qualification.', 'amber',130) +d.note(575,855,490,'Separate Cell capability proposal','Adding KV, Queue, Workflow, Blob, Cron, Timer and Effects to the same Repository Cell needs framework composition and autonomous-progress acceptance gates.', 'amber',130) +d.note(35,1035,1030,'What does not change in the intended design','The repository UUID, one fenced authoritative publisher, final policy checks and durable outcome acknowledgement remain. Immutable uploaded artifacts are preparation until an authorized Cell command publishes them.', 'blue',90) +d.legend([('green','active serving path'),('blue','new implemented primitives'),('violet','immutable metadata / bytes'),('amber','integration and qualification open')]) +d.save('The current product still uses per-object SQL authority, with an active archived-pack optimization. The new subsystem moves canonical inventories, graph metadata, ref snapshots and exact outcomes into verified immutable artifacts and short Cell publication commands. Preparation, publication dispatch, certificates, policy pages, recovery records and readers exist, but the full serving path is not registered and the hard cutover remains incomplete. Multi-capability repository Cells are a separate proposal.', ['crates/canopy-server/src/packs/publication/mod.rs','crates/canopy-server/src/packs/publication/coordinator.rs','crates/canopy-server/src/packs/catalog/mod.rs','crates/canopy-server/src/packs/directory/mod.rs','crates/canopy-server/src/packs/ref_state/mod.rs','docs/large-repository-implementation-status.md','docs/large-team-scalability.md','docs/repository-cell-primitives.md']) + +(OUT/'manifest.json').write_text(json.dumps(ATLAS,indent=2)+'\n') +print(f'Generated {len(ATLAS)} SVG diagrams in {OUT}') diff --git a/diagram/canopy-architecture/index.html b/diagram/canopy-architecture/index.html new file mode 100644 index 0000000..d639a43 --- /dev/null +++ b/diagram/canopy-architecture/index.html @@ -0,0 +1,169 @@ +Canopy architecture and Cellule integration +
    +

    Canopy architecture and Cellule integration

    Trace a request from a Git client or browser to its authoritative repository state. See which components Canopy owns, what Cellule supplies, and how writes survive owner changes and local disk loss.

    +

    Source snapshot: October 4, 2026 · Canopy 9438bb8 · Cellule 161067f

    +
    Diagrams 01–08 describe the current serving architecture. Diagram 09 separates newer implemented packed-storage primitives from incomplete production integration. Composed repository queues and workflows remain a separate proposal. These diagrams do not assert production capacity.
    +
    Product boundaryCanopy owns ingress, Git, auth, collaboration and operational policy.
    Framework boundaryCellule supplies stable targets, fenced execution, durable outcomes and restoration.
    Authority boundaryCells and published roots preserve state. Native Git files are rebuildable.
    +

    Use + and − to inspect a diagram; drag the scrollbar to pan. Fit restores the overview. Each figure has standalone SVG and PNG exports and expandable explanations with source links.

    +
    +
    01

    Canopy system architecture

    +

    Canopy is an embedded Cellule application. Public ingress and native Git run in Canopy, while each durable Directory or Repository Cell has one fenced owner. A receiving node may call a remote owner. The object store holds authority records, SQLite recovery roots and product body objects. Native Git caches are disposable.

    +
    100%SVGPNG @2x
    +Canopy system architectureCanopy is an embedded Cellule application. Public ingress and native Git run in Canopy, while each durable Directory or Repository Cell has one fenced owner. A receiving node may call a remote owner. The object store holds authority records, SQLite recovery roots and product body objects. Native Git caches are disposable. + + +CANOPY NODE / ONE RUST SERVICEDURABLE CELL AUTHORITYGit clientsSmart HTTP / SSHclone, fetch, pushBrowser / APIEmbedded web UIJSON collaboration APIGit LFSHTTP bodies / locksSSH can issue grantsIngressAxum HTTP routerOptional russh listenerAccount authenticationRepositoryManagerName resolutionResidency + request pinsLocal / peer Cell handlesProduct servicesGitGateway + LfsService + repository_httpBrowse, issues, PRs, reviews, checks, policyNative Git work runs on the receiving nodeCellule integrationCanopyApplication + CellNodeTyped SQL / domain commandsFenced ownership + durable output gateDirectory CellOne fixed SQL shard per tenantNames -> stable repository UUIDAccounts, tokens, keys, auditRepository CellOne SQL Cell per repository UUIDRefs, objects, ACL, collaborationOutcomes, closure and policyDisposable Git cachePrivate refs + verified shared objectsNative receive/upload-pack workersCan be rebuilt after local disk lossShared S3-compatible object storeCellule: catalog, authority, node leases, LTX roots and immutable SQLite recovery bytesCanopy: external Git blobs, archived pack/index bytes and Git LFS bodiesConditional writes select authority; verified immutable bytes supply recoveryresolve / authtyped operationsnative workdurable publication / external body I/OhydrateOwnership ruleGit success isacknowledged only afterthe Cell commits andpasses its durabilitygate.public clientsproduct logicframeworkdurable authorityrebuildable local filesCanopy system architectureClients enter any node; Cell authority and immutable storage preserve repository state.CURRENT SERVING PATH +
    +
    Read the explanation and source code

    The canopy binary runs one Rust service. Axum handles HTTP, JSON APIs, the embedded browser and Git LFS; an optional russh listener handles SSH Git operations. RepositoryManager resolves identities, admits repository transitions, binds local or remote Cell clients and retains request pins. GitGateway turns repository state into a native Git workspace and translates native results into authoritative Cell operations. LfsService verifies and publishes LFS bodies independently of native Git.

    +

    The Rust workspace has three crates. canopy-git-format supplies object kinds, SHA-1/SHA-256 identities and canonical hashing. canopy-object-storage owns immutable body and artifact operations. canopy-server composes those crates with Cellule and owns product schemas, protocols, authorization and deployment lifecycle. These are linked components of the service, not separately deployed microservices. See the workspace map and server assembly.

    +

    Component responsibilities

    +
    ComponentResponsibilityState it may publish
    HTTP and SSH ingressAdmit protocol input, authenticate and manage transport lifetimeThrough registered product operations
    RepositoryManagerResolve targets, own resident bindings, transition locks and request pinsCatalog/acquisition/release through Cellule and name coordination through Directory
    GitGatewayPrepare native work, ingest verified objects, certify closure and finalize pushesRepository commands; private native refs stay provisional
    Repository HTTP and browse servicesExpose repository and collaboration operationsRegistered commands with current actor/precondition checks
    LfsServicePublish and verify LFS bodies, metadata and advisory locksAuthorized Repository Cell metadata after body verification
    CanopyApplication and modulesDescribe topology, schema, operation IDs/codecs and handlersCompiled application contracts
    CellNode and runtimeSupervise Cell execution, durable outputs, ownership and recoveryFenced Cell control and published roots
    Object storageRetain immutable recovery and product bytesConditional control records and verified immutable objects
    +

    Execution and control placement

    +

    The request data path includes input spooling, native Git, verified readers and body streaming. The ownership control path includes release selection, catalog records, node enrollment, Cell acquisition and conditional root publication. They meet when a product operation calls a registered Cell command or query. This separation allows a receiving gateway to keep its native process and socket while authoritative SQL execution happens on another node.

    +

    The Directory is shared state, so account and name operations do not scale by creating a new Directory for every repository. Repository state is partitioned by UUID, allowing unrelated repositories to have different owners. The architecture does not imply that one hot repository can have multiple concurrent authoritative SQL writers.

    Read the complete design section

    Source map at the inspected revision

    +
    +
    02

    How Canopy models its domain on Cellule

    +

    CanopyApplication registers DirectoryModule and RepositoryModule as SQL Cell types. Directory uses a fixed shard; repository UUIDs derive entity partitions. Product wrappers hold typed SqlCell handles. Cellule owns execution and persistence mechanics, and Canopy owns protocols and authorization. Other framework primitives are available but are not composed into Canopy Repository Cells.

    +
    100%SVGPNG @2x
    +How Canopy models its domain on CelluleCanopyApplication registers DirectoryModule and RepositoryModule as SQL Cell types. Directory uses a fixed shard; repository UUIDs derive entity partitions. Product wrappers hold typed SqlCell handles. Cellule owns execution and persistence mechanics, and Canopy owns protocols and authorization. Other framework primitives are available but are not composed into Canopy Repository Cells. + + +APPLICATION DECLARATION / COMPILED AT STARTUPRUNTIME ADDRESSING / STABLE ACROSS OWNER MOVEMENTEMBEDDED CELLULE FRAMEWORK / THE PINNED DEPENDENCYCanopyApplicationCellApplication::registerDirectoryModule + RepositoryModuleTwo declared SQL Cell typesDirectoryModuleCatalogRole::Sql / one fixed shardNamespace = [0x48; 16]Names, accounts, credentialsRepositoryModuleCatalogRole::Sql / entity partitionsNamespace = [0x47; 16]Git + ACL + collaboration commandsDirectory targetTenant + application + namespacepartition_for_shard(0)DirectoryCell wraps SqlCell<DirectoryModule>Repository targetTenant + application + namespace + UUID-derived partitionCellType::entity_partition -> 33-byte partitionRepositoryCell wraps SqlCell<RepositoryModule>UUID selects one Cellcellule-app + hostDescriptor / registry / typed handlesCellNode lifecycle + readinessCanopy installs facilities and leasescellule-runtimePer-Cell actor + SQL workerAuthority fence + request ledgerQueries, commands, peer contractscellule-ltx + storeManaged SQLite WAL captureVerified roots + exact restorationBounded I/O + conditional updatesDependency pin: 161067f5a21703b3e257024bcb64e565fd9657b4cellule-types is transitive. Canopy supplies its own signed HTTPS peer transport.Framework capabilitiesCellule includes SQL, KV, Queue, Workflow, Blob, Cron, Timer and Effects.Capabilities use typed APIs and explicit topology.Canopy capability boundaryCanopy registers SQL Cells today. Multi-capability Repository Cells, repositoryqueues and autonomous workflows remain a proposal.domain declarationsframework bindingfenced executioncapabilities not wired into CanopyHow Canopy models its domain on CelluleDomain code supplies schema and policy; the embedded framework supplies identity, execution and recovery.CURRENT SERVING PATH +
    +
    Read the explanation and source code

    CanopyApplication::register installs DirectoryModule and RepositoryModule, then declares their SQL Cell types. Modules describe schema migrations, operation IDs, codecs, code digests and limits. A compiled registry binds these contracts to the application and selected release. Product wrappers invoke registered commands and queries through typed handles. See application and repository registration and Directory registration.

    +
    Canopy conceptCellule representationConsequence
    Deployment identityTenant and application IDsScope every durable target
    DirectorySQL namespace [0x48; 16], fixed shard zeroOne shared name and identity authority per tenant/application
    RepositorySQL namespace [0x47; 16], entity partition derived from UUIDOne independently owned Cell for each repository
    Directory operationsSqlCell<DirectoryModule> plus credential commandsAccount state and name changes stay inside the Directory boundary
    Repository operationsSqlCell<RepositoryModule> plus typed domain commandsRef publication can check ACL, graph and policy in one transaction
    MutationMutationIdentity and recorded outcomeRetries preserve one logical command identity
    Observed write positionReceiptA later read can require that per-Cell position
    Node lifecycleCellNode and installed runtime/replica facilitiesReadiness and shutdown depend on supervised framework state
    +

    The repository partition is a prefix byte followed by a domain-separated 32-byte digest; RepositoryCell::new verifies that the supplied UUID derives the supplied target. A rename changes the Directory mapping while retaining the repository UUID and Cell identity. Creating a repository uses a durable pending name reservation, initializes the Cell and then marks the name ready. Those steps coordinate separate Cells; they do not form a distributed SQL transaction.

    +

    cellule-app compiles topology and exposes handles. cellule-host supervises the node. cellule-runtime manages actors, SQL execution, request outcomes and fenced authority. cellule-ltx captures SQLite changes and verifies restoration; cellule-store supplies bounded object operations and conditional updates. Canopy supplies public ingress, user authorization, credentials and deployment policy. The Cellule framework architecture at the pinned revision documents these boundaries.

    +

    Module contracts and typed handles

    +

    A module is the executable contract installed into a Cell. Its operation descriptors specify IDs, codec versions, schema compatibility and input/output bounds; its code identity covers the authoritative implementation. SqlCell<RepositoryModule> and SqlCell<DirectoryModule> bind calls to these registered contracts. Product wrappers add domain operations while leaving routing, command identity, receipts and durable output handling to Cellule.

    +

    The application declaration and runtime assembly serve different purposes. CanopyApplication declares the available Cell types; server startup installs runtime/replica facilities, enrolls the node and binds actual targets. Adding a Rust module or schema table does not automatically make an operation callable on the production path. Registration and release admission must agree with the handlers that execute and restore it.

    +

    Repository creation and alias changes

    +

    Creation first reserves the owner/name with a canonical UUID and a pending state in Directory. It then initializes that UUID's Repository Cell and marks the reservation ready. A request interrupted between these phases must preserve the existing reservation rather than allocate a different identity for the same operation. The pending/ready protocol coordinates the two durable boundaries.

    +

    Rename compares the expected UUID and moves the ready Directory name while retaining repository identity. Administrative operations that carry repository identity must reject name reuse that would otherwise retarget a stale write. The Repository Cell continues to own its object format, owner identity and repository policy. See creation and discovery.

    Read the complete design section

    Source map at the inspected revision

    +
    +
    03

    Cellule command execution and durability

    +

    This is the object-store durability path assembled by Canopy. The Cell owner records domain state and the request outcome together, captures SQLite through LTX and publishes the exact root using a fenced conditional authority update. Only then does the caller receive Committed and a receipt. A Cellule follower-log mode exists, but this diagram does not claim Canopy enables it.

    +
    100%SVGPNG @2x
    +Cellule command execution and durabilityThis is the object-store durability path assembled by Canopy. The Cell owner records domain state and the request outcome together, captures SQLite through LTX and publishes the exact root using a fenced conditional authority update. Only then does the caller receive Committed and a receipt. A Cellule follower-log mode exists, but this diagram does not claim Canopy enables it. + + +Canopy callerTyped handleCell ownerActor + fenceManaged SQLiteState + ledgerLTX / object storeImmutable rootsCellAuthorityControl CAS1 Command + stable MutationIdentity2 Validate owner incarnation / lease / compatible code3 Execute domain mutation + request outcome in one transaction4 Commit SQLite and capture its exact WAL boundary5 Verify capture; upload immutable chunks and proposed root6 Return proposed recovery root7 Conditional publish of the exact root under owner fence8 Accepted authority revision = durable publication proof9 Committed<Output> + ReceiptLost reply or rejected fenceA timeout does not prove failure. Resolve the original request identity; retry thesame logical command only under its existing identity.Receipt and scopeA receipt-bound query must observe the published per-Cell position. Directory andRepository commands do not form one cross-Cell SQL transaction.invocationlocal transactionimmutable bytesauthority publicationdurable acknowledgementCellule command execution and durabilityOne command changes one Cell. The domain mutation and its recorded answer share a SQLite transaction.CURRENT SERVING PATH +
    +
    Read the explanation and source code

    A command executes at the current fenced owner of one Cell. Its SQLite transaction records the domain change and the request outcome together. On the object-store path assembled here, LTX captures the committed WAL boundary, verifies and publishes immutable recovery bytes, and supplies a proposed root. A conditional authority update publishes that exact root under the owner fence. The caller receives Committed<Output> and a receipt after this durability gate.

    +

    Authority fencing prevents a former owner from publishing new authoritative state after ownership changes. A transport timeout can occur after a command committed; resolving its original identity distinguishes a completed outcome from an unstarted operation. A receipt is scoped to a Cell, owner incarnation and commit sequence. It can constrain a subsequent query, but it cannot create an atomic transaction across Directory and Repository Cells. Cellule also documents an optional follower-log durability path; the diagram describes Canopy's inspected object-store setup.

    +

    Commit and acknowledgement ordering

    +

    The local SQLite transaction is an execution boundary. The accepted authority revision is the durable publication boundary. Between them, LTX must capture the exact WAL endpoint and immutable recovery bytes must become available. The authority update binds that recovery root to the valid owner fence. The caller must not receive a durable success acknowledgement solely because local SQL committed or uploads completed.

    +

    The recorded outcome travels with the same restored SQLite state as the domain mutation. If a node disappears after publication but before its reply arrives, a new owner can recover both. If a proposed root was uploaded but never selected, that upload does not supersede the last accepted root. This is why failure handling uses command resolution and authority state rather than interpreting transport errors as domain results.

    +

    Read consistency and optimistic preconditions

    +

    A receipt-bound query asks to observe at least the relevant published position of one Cell. It does not make a Directory lookup and a Repository query atomic. Discovery therefore rechecks current repository access, and final writes carry expected identities and versions into their own transaction.

    +

    Git ref snapshots use a shared generation with HEAD. Continuation pages must retain that generation; a changed generation invalidates the scan. Ref writes compare both the expected OID and monotonic version. Deleting and recreating a name cannot make an old plan valid again merely because the OID matches. These are separate mechanisms: receipts constrain observation, generations bind a multi-page read, and versions fence optimistic writes. See ref snapshots.

    Read the complete design section

    Source map at the inspected revision

    +
    +
    04

    Request routing and repository residency

    +

    Directory state authenticates accounts and resolves ready repository names. RepositoryManager binds a stable target to either a local actor or a signed peer client. Cold transitions are admitted and supervised; remote-owner loss triggers acquisition on demand. Repository authorization is checked using current durable state. Response pins prevent eviction during streaming.

    +
    100%SVGPNG @2x
    +Request routing and repository residencyDirectory state authenticates accounts and resolves ready repository names. RepositoryManager binds a stable target to either a local actor or a signed peer client. Cold transitions are admitted and supervised; remote-owner loss triggers acquisition on demand. Repository authorization is checked using current durable state. Response pins prevent eviction during streaming. + + +Incoming requestHTTP / SSH / API / Git LFSDirectory CellValidate token / key / LFS grantResolve ready owner/name -> UUIDRepositoryManagerDerive target; pin resident routeOtherwise admit a transitionLive ownerelsewhere?Remote Cell bindingCellClient::peer -> /internal/cellSigned request + enrolled node keyTLS, release, expiry, principal checksLocal acquisitionBootstrap new / acquire idle CellExpired owner: fence + restore exact rootCellClient::local -> resident actorRepository operationsRecheck current ACL / visibilityRun domain command or native GitRetain route pin through responseYesNoDiscovery is a hintDirectory listing candidates are checkedagainst current Repository Cell access. Cachednames and UI state cannot grant permissions.Bounded transitionsCold/remote work uses supervised admission anda per-repository transition lock. Differentrepositories can activate concurrently.Capacity and cancellationAn unavailable slot or in-progress movement can return 503. Streamed responses anddetached admitted work retain pins until they finish.Native work placementRemote binding moves Cell calls. Git workers, body streaming and the disposable cachecan remain on the receiving gateway.requestidentity / authorityadmissionlocal or signed remote executionRequest routing and repository residencyAuthentication, name lookup, account admission and owner resolution precede repository work.CURRENT SERVING PATH +
    +
    Read the explanation and source code

    HTTP credentials, SSH keys and LFS grants are checked using Directory state. Ready owner/name entries resolve to stable repository UUIDs. Repository routing then binds a CellClient::local to a resident actor or a CellClient::peer to the current remote owner. Canopy's peer transport signs requests to POST /internal/cell and verifies the enrolled sender, signature, release, expiry and principal over HTTPS. It is Canopy's own transport over runtime peer contracts; the optional Cellule peer adapter is not a Canopy dependency. See peer routing.

    +

    The receiving node can run native Git and stream product bodies while Cell calls go to another node. Public requests therefore do not require sticky load-balancer sessions. Node advertisement and Cell ownership are separate controls: the advertisement proves a node session is live; the Cell control record identifies authority for a particular Cell.

    +

    Cold and remote transitions use bounded account admission and a lock for that repository. The configured residency limit counts locally bound repository gateways, including remote bindings. When no slot is free, an inactive gateway can be evicted; local ownership must be released before its slot is reused. Requests and response streams pin their route. Supervised activation and cleanup continue after client cancellation. Full admission or ownership movement can return a retryable 503. See residency and admission.

    +

    Routing decisions

    +
    Observed stateRouting actionRequired condition
    Resident local bindingCall the local Cell clientThe binding remains valid and request-pinned
    Live owner on another nodeBind a signed peer clientCurrent authority identifies an enrolled live owner and HTTPS endpoint
    New or idle targetAdmit local acquisition/bootstrapCatalog/release checks and runtime capacity permit it
    Expired ownerFence and restore under new ownershipRecover exactly the root selected by authority
    Full pinned resident set or busy movementReturn retryable capacity/unavailabilityDo not steal an active slot or infer owner death from latency
    +

    A remote binding owns a local gateway entry but not the remote SQLite workspace. Evicting that binding does not release the remote Cell. When remote authority changes, subsequent routing refreshes the binding using current ownership. Node discovery is a routing input, while Cell control remains the ownership authority.

    Read the complete design section

    Source map at the inspected revision

    +
    +
    05

    Git push from wire input to durable refs

    +

    After outer authentication, GitGateway spools and hashes the encoded request and binds a logical push ID to the actor and digest. It can replay a completed push before decoding and native work. New attempts prepare a private ref snapshot, run receive-pack, ingest canonical objects, certify graph closure and stage the report and ref plan. CompletePush checks current policy and expected ref versions and commits accepted refs with the canonical response pointer. Cellule gates acknowledgement on durable publication.

    +
    100%SVGPNG @2x
    +Git push from wire input to durable refsAfter outer authentication, GitGateway spools and hashes the encoded request and binds a logical push ID to the actor and digest. It can replay a completed push before decoding and native work. New attempts prepare a private ref snapshot, run receive-pack, ingest canonical objects, certify graph closure and stage the report and ref plan. CompletePush checks current policy and expected ref versions and commits accepted refs with the canonical response pointer. Cellule gates acknowledgement on durable publication. + + +FINAL DOMAIN TRANSACTION: ACCEPTED REFS + GENERATION + OUTCOME POINTER / AUDITGit clientreceive-packGitGatewayReceiving nodeNative GitDisposable refsRepository CellVia local / peer clientObject storeVerified body bytes1 Spool encoded input; bind push UUID + actor + request digest2 begin_push: claim logical ID or find completed response3 Completed ID replays saved result; fresh ID continues4 Read consistent refs, generation and branch policy5 Decode / validate; receive-pack in private cache6 Native report + actual accepted ref differences7 Upload verified external blobs or archived pack/index bytes8 Persist canonical object records in bounded batches9 Certify typed graph closure; stage response + ref plan10 CompletePush: recheck ACL, branch rules, OIDs and versions11 Cellule publishes SQLite root12 Fenced durable proof13 Read canonical saved response14 Git report + push identityTwo identities matterThe HTTP push UUID identifies the whole wire operation. Cellule MutationIdentityidentifies each durable command within that operation.Failure behaviorPreparation changes disposable refs only. Objects may remain unreferenced afterrefusal. An uncertain final publication is resolved, never turned into a falsenegative report.Native partial acceptance is preserved; Git --atomic can request all-or-none validation. Current gateway push work uses a mutex.wire protocolprivate native statedurable Cell callsbody / root publicationGit push from wire input to durable refsCurrent serving path: native Git prepares private state; CompletePush publishes refs and the saved result.CURRENT SERVING PATH +
    +
    Read the explanation and source code

    The gateway spools input to admitted scratch storage and binds the logical push UUID to the account, repository context and encoded request digest. begin_push detects completed operations before gzip decoding and native preparation. Reusing an ID with different bytes or an actor produces a conflict. Upload reception can overlap, but the current gateway serializes its ID check, native push work and publication phase through a mutex. See gateway control flow and owned preflight.

    +

    New attempts capture consistent refs and policy, run native receive-pack against private refs and derive the actual accepted changes. Ingestion verifies canonical object identities, publishes required external bytes, persists bounded object batches and certifies typed graph closure. The gateway stages the response and ref plan. CompletePush rechecks current permissions, branch rules, expected OIDs and monotonic ref versions, then commits accepted refs, the generation change and the canonical response pointer together. Cellule publishes the durable result before the client sees success. See push execution, saved outcomes and shared ref guards.

    +

    There are two identity layers: the push UUID identifies the complete wire operation; Cellule mutation identities identify its individual durable commands. A retry of a completed push replays the saved result and does not reapply old refs over later repository changes. Native partial acceptance is preserved, while stock Git --atomic requests atomic validation. Preparation failure cannot publish private refs, although verified unreferenced objects can remain. An uncertain final publication must be resolved rather than reported as a definite refusal.

    +

    Preparation and final publication

    +
    PhaseWorkAuthoritative effect
    Input identitySpool encoded input and bind actor, UUID and digestEstablish the product operation and detect completed replay
    Native preparationBuild a private ref snapshot and run receive-packProduce provisional objects and the actual native report
    Verified ingestionCheck canonical hashes, store object metadata/bodies and certify typed graph closureMake verified objects available without publishing the final ref plan
    StagingRetain ordered ref changes and exact response dataPrepare a bounded final command
    CompletePushRecheck permissions, policy, OIDs and versions; commit refs, generation and response pointerPublish the accepted repository transition atomically
    Durable outputPublish the exact SQLite root through CelluleAuthorize the successful report to the client
    +

    Preparation can overlap unrelated operations, but the active gateway's push mutex serializes its native/ref-publication phase. A single push may involve several object batches and Cell commands; it is not one long distributed transaction over all uploaded bytes. The final command is the point that joins accepted refs with the canonical saved outcome.

    +

    Graph closure validates typed dependencies: branch tips are commits; commit trees/parents, tree entries and tag targets require the appropriate reachable objects. Gitlinks refer to another repository and do not require local object presence. Certification supports final publication without repeatedly traversing already certified history. It does not replace every metadata or portability check performed by git fsck.

    +

    Retry scope and transport differences

    +

    HTTP exposes the product push UUID so a caller can recover an uncertain report under the same input binding. A completed retry returns the saved status, headers and verified body instead of running receive-pack again. Reusing that UUID with a different actor or encoded digest conflicts. This replay must not overwrite refs changed by later independent pushes.

    +

    SSH uses the same object ingestion and durable ref-publication rules but does not expose the HTTP push retry-ID contract. Documentation and clients must not promise that opening a new SSH command recovers the exact result of an earlier disconnected command. See SSH transport and lost push reply recovery.

    Read the complete design section

    Source map at the inspected revision

    +
    +
    06

    Fetch, browse and Git LFS data paths

    +

    Fetch takes consistent refs, validates requested object reachability, hydrates selected verified objects and uses native upload-pack to stream the wire response. Browser APIs read through Cell metadata and verified body readers. LFS uploads first publish and verify immutable bytes, then commit authorized metadata; downloads verify manifest-bound parts. The active serving schema supports inline, chunked, external and packed storage records, distinct from the incomplete immutable catalog replacement.

    +
    100%SVGPNG @2x
    +Fetch, browse and Git LFS data pathsFetch takes consistent refs, validates requested object reachability, hydrates selected verified objects and uses native upload-pack to stream the wire response. Browser APIs read through Cell metadata and verified body readers. LFS uploads first publish and verify immutable bytes, then commit authorized metadata; downloads verify manifest-bound parts. The active serving schema supports inline, chunked, external and packed storage records, distinct from the incomplete immutable catalog replacement. + + +GIT CLONE / FETCHBROWSE / JSON READS / LFSCURRENT SERVING STORAGE / AUTHORITATIVE METADATA REMAINS PER OBJECTAuthorize and selectCurrent ACL / public visibilityValidate wants against live refsSnapshot and hydrateGeneration-consistent refs + HEADVerify required objects / bodiesApply supported partial-clone filterNative upload-packPrivate ref snapshotShared verified object filesStream generated pack to clientDiscovery fast pathGit v2 capability discovery needs no object hydration. Ref discoveryprepares ref targets and annotated tag chains.Selective fetchBlobless fetch omits ordinary blobs. Selected wants, ref targets and structuralhistory determine hydration; warm verified objects are reused.Repository browserRead trees, blobs, history and diffsUse Repository Cell metadataLoad bodies from verified readersNo receive-pack is neededLFS uploadAuthorize batch / PUTHash SHA-256 + bounded partsPublish immutable body, verifyThen commit authorized metadataLFS downloadAuthorize against current CellRead manifest pinned in SQLiteVerify requested parts / digestsSupport tail Range / 206 responseLFS locks are advisory records in the Repository Cell. SSH grants still deliver LFS bytes over HTTP.Repository SQLiteobjects: kind, size, digest, storageinline bodies <= 768 KiBLarge non-blob bodies: SQL chunksrefs + graph edges / certificatesLFS metadata + locksImmutable external bytesLarge loose Git blobs + manifestsArchived pack and index bodiesPacked blobs reference approved packsGit LFS bodies + part digestsReaders verify canonical identitiesLocal cacheHydrated Git files / installed packsInsertion cursor + ref generationsPrivate native process workspaceDisk and native-worker admissionDisposable after recoveryreferencesverify / hydrateThis active archive path differs from the future immutable catalog / ref-root hard cutover in diagram 09.Read authorityAn uploaded pack, cached object, stale listing or UI view cannot independently authorize an object read. Current Cell access and verified published metadata govern visibility.stock Git readsbrowse / selectiondurable state and bodiesrebuildable work filesFetch, browse and Git LFS data pathsReads verify identities against Cell metadata. LFS body delivery bypasses native Git.CURRENT SERVING PATH +
    +
    Read the explanation and source code

    Fetch reads generation-consistent refs and HEAD, validates wants against current repository reachability, hydrates selected verified bodies and lets native upload-pack generate the response. Warm verified object files are shared across private ref snapshots. Git v2 capability discovery avoids object hydration; ref discovery prepares ref targets and tag chains. Blobless fetch omits ordinary blobs, and supported filters guide further body selection. Browse APIs use repository metadata and verified object readers for trees, blobs, history and comparisons. See fetch, hydration and Git reads.

    +

    The active repository schema stores per-object kind, size, digest and storage choice. Inline objects are bounded at 768 KiB; larger structural objects use SQL chunks. Large loose blobs use immutable external bodies. The serving code also archives pack/index bodies and can store packed-blob references while retaining per-object SQL authority. This active optimization is distinct from the future immutable catalog design.

    +

    LFS transfers use HTTP even when SSH issues the authorization grant. Upload hashes and publishes bounded immutable parts, verifies the complete body and then commits authorized metadata. Download uses the manifest digest pinned in SQLite and verifies requested parts, including tail range responses. Locks are advisory Repository Cell records. See LFS upload and LFS read. Canopy's immutable body store is product code; it is not a registered Cellule Blob capability.

    +

    Storage representations

    +
    RepresentationDurable metadataByte locationVerification role
    Inline Git objectKind, size, digest and OIDRepository SQLite rowRecompute canonical object identity
    Chunked structural objectBound upload/chunk metadataRepository SQLite chunksVerify complete object before publication
    External loose blobSize, digests and immutable body referenceCanopy object storageCheck body bytes against SQL metadata and Git identity
    Packed blobPer-object SQL record and archive bindingImmutable archived pack/indexAuthenticate archive/object mapping and object bytes
    LFS objectSHA-256, size and manifest bindingImmutable LFS body partsVerify manifest-bound parts and requested ranges
    SQLite recovery dataAuthority-selected root and LTX metadataCellule immutable recovery storageRestore the exact accepted Cell state
    +

    Each representation has an authoritative reference and a verified reader. A native cache may contain an additional copy, but cache presence does not create an object record or authorize a ref. The independent content checks matter because an object-store key or physical pack offset is only a location, not proof of the bytes' identity.

    +

    Snapshot lifetime

    +

    An admitted fetch retains a coherent ref generation and its cache/input ownership through streaming. Later writes can advance the repository while that request serves its selected immutable snapshot. Discovery does not need to wait for an unrelated history hydration job. Authorization still precedes access; cached refs and warm objects cannot grant read permissions. The absence of production collection currently keeps unreferenced immutable bytes available, but a future collector must explicitly protect active readers and their retained roots.

    Read the complete design section

    Source map at the inspected revision

    +
    +
    07

    Owner lifecycle and exact recovery

    +

    Canopy locks its local runtime workspace, verifies storage capabilities, validates the selected release and enrolls a signed node lease before serving. New or cold Cells bootstrap, acquire idle authority or fence an expired owner and restore the authority-pinned SQLite root. Recovery preserves the request outcome ledger. Native caches rebuild afterward. Shutdown and maintenance supervise all admitted work; backup copies durable roots and referenced external bodies into an independent prefix.

    +
    100%SVGPNG @2x
    +Owner lifecycle and exact recoveryCanopy locks its local runtime workspace, verifies storage capabilities, validates the selected release and enrolls a signed node lease before serving. New or cold Cells bootstrap, acquire idle authority or fence an expired owner and restore the authority-pinned SQLite root. Recovery preserves the request outcome ledger. Native caches rebuild afterward. Shutdown and maintenance supervise all admitted work; backup copies durable roots and referenced external bodies into an independent prefix. + + +STARTUP AND SERVINGCOLD ACTIVATION / NODE LOSS / LOCAL DISK LOSSDRAIN, MAINTENANCE AND INDEPENDENT BACKUPStartup preflightLock managed local workspaceProbe conditional writes / ranged I/OCompile and check selected releaseEnroll and renew nodeSigned live node advertisementNodeLeaseGuard + compatible registryStart Directory + on-demand reposServe with fencingReady only while leases are validPer-Cell authority / actor / SQLCancel ingress when lease failsRead catalog and controlValidate compiled code and schemaLive owner: route to that ownerIdle or expired: acquire authorityFence and choose exact rootClaim expired node for takeoverUse CellAuthority-pinned rootDo not elect a root by listing keysVerify and restore SQLiteLTX chunk checksums / endpointsRecover exact state + outcome ledgerInvalid bytes: fail without activationActivate successorSame CellTarget / repository UUIDNew fenced owner / incarnationRecovered refs, ACL and outcomesRebuild Git cacheHydrate verified published objectsRecreate private ref snapshotsResume Git/API/LFS requestsCache loss is recoverableCell publication protects acknowledged state.Files from an old workspace never become asubstitute recovery root.Recovery is on demand. An unclean owner loss waits for lease expiry; acquisition and restore stay supervised.Graceful shutdownStop public ingress; finish admitted workDrain Cells + close native descendantsWithdraw advertisement / release diskWorkspace stays locked through cleanupDeployment maintenanceClose release admission with operation IDFleet drains; prove all Cells settledRecovery worker fences expired ownersExplicit end reopens same releaseBackup and restoreCopy pinned roots + referenced bodiesVerify independent disjoint prefixRestore into unused reserved prefixSame provider; preserve identitynode / deployment controlowner fenceverified recovery bytessame Cell / new ownerOwner lifecycle and exact recoveryA repository keeps its identity when its owner changes. Only authority can select the recovery root.CURRENT SERVING PATH +
    +
    Read the explanation and source code

    Startup locks the managed workspace, probes provider behavior, compiles the application, checks release admission and enrolls a signed node lease. Readiness depends on valid framework state and leases. A new Cell bootstraps; an idle Cell restores under acquired authority; a dead owner requires lease expiry and a fenced takeover. Recovery restores the root pinned by Cell authority, verifies required bytes and restores the outcome ledger before activation. Local databases and object listings cannot select a newer-looking root. Native caches rebuild afterward. See acquisition and startup and workspace management.

    +

    Graceful shutdown stops ingress, finishes admitted work, drains Cells and withdraws the node advertisement. Deployment maintenance closes release admission and uses an operation UUID until drain and recovery are proven complete; an explicit end reopens the release. Backup copies pinned roots and referenced external bodies into a disjoint prefix, verifies that independent copy and restores into an unused reserved destination. The current path is a same-provider copy. These controls do not complete schema migration, cross-provider export or collection. See the operations runbook.

    +

    Startup and release admission

    +

    Startup constructs and validates native admission, locks the local workspace, checks provider capabilities and the selected compiled release, then enrolls the signed node session. It checks release readiness around enrollment and before exposing ingress. Release changes or failed lease validation can close readiness and request supervised shutdown. A Cell's catalog and registered handlers must be compatible with the selected release before acquisition.

    +

    Maintenance admission, local file exclusion and ownership fencing are distinct controls. Maintenance can close new work across the release; the workspace lock prevents competing local runtimes; Cell authority fences the writer. A released listener or expired heartbeat alone does not prove that every admitted native or Cell operation finished.

    +

    Drain and workspace reuse

    +

    One supervisor retains startup, admitted work, native processes, Cell drain, leases and workspace exclusion. Canceling the caller's startup/shutdown wait requests or observes cleanup; it does not discard that supervisor. Native admission closes, tracked ingress/work joins, native owners drain, then Cellule shuts down. Only confirmed drain authorizes workspace release and advertisement withdrawal.

    +

    If drain or cleanup is uncertain, the workspace fence or native claim remains retained. On Unix, worker ownership includes inherited completion/lock descriptors so participating descendants cannot outlive the accounting boundary unnoticed. Linux listener guards additionally prevent a fork-inherited listener from surviving a completed handoff. These mechanisms do not claim general containment of every possible helper on every OS. See native process ownership and resource admission.

    +

    Backup and restoration boundaries

    +

    A backup is a verified independent copy of pinned recovery roots and referenced product bodies, with its own operation identity and destination. Restoring an empty reserved deployment from that copy differs from reconstructing a live owner's local cache. The former must preserve the copied release/schema and destination admission; the latter follows current Cell authority. Maintenance recoveries for explicitly supported retained contracts do not imply a general old-release migration or cross-provider restore path.

    Read the complete design section

    Source map at the inspected revision

    +
    +
    08

    Repository components and collaboration control

    +

    The Repository Cell colocates Git refs and graph metadata with ACLs, visibility, issues, PRs, reviews, line threads, check attempts, branch rules and operation outcomes. Merge preparation may run native Git outside the authority boundary, but MergePull rechecks current exact revisions, policy and required evidence before atomically changing the branch and PR state. Directory and Repository state remain separate Cells.

    +
    100%SVGPNG @2x
    +Repository components and collaboration controlThe Repository Cell colocates Git refs and graph metadata with ACLs, visibility, issues, PRs, reviews, line threads, check attempts, branch rules and operation outcomes. Merge preparation may run native Git outside the authority boundary, but MergePull rechecks current exact revisions, policy and required evidence before atomically changing the branch and PR state. Directory and Repository state remain separate Cells. + + +ONE REPOSITORY CELL / ONE UUID / ONE FENCED WRITEREXAMPLE: MERGE A PULL REQUESTDIRECTORY AND REPOSITORY ARE SEPARATE TRANSACTION DOMAINSGit identity and historyObject format, objects and graphRefs, versions, HEAD, generationAccess and policyOwner / members / visibilityBranch rules and version guardsIssues and discussionsIssues, comments and editsRequest records / paginationPull requests and reviewsExact base/head comparisonsReviews, threads and resolutionChecks and candidatesCommit-bound check attemptsMerge / squash / rebase candidatesDurable operation outcomesPush reports / certificate auditMutation replay and receiptsPrepare candidateAuthenticate and capture base/headNative merge-tree / commit-treePersist candidate objects + closurePublish with MergePullRecheck actor and exact revisionsRequired checks / reviews / threadsBranch policy + ref expectationsCommit and acknowledgeUpdate branch ref + generationClose PR + save merge outcomeCellule gates durable resultPreparation is provisionalA candidate or passed check cannot grant publication. The final commandchecks the current branch and exact candidate/head bindings.CI and web boundariesChecks are product records submitted through the API. Canopy does not install aRepository Cell workflow runner for CI today.Directory CellGlobal accounts / token scopes / SSH keysNames and discovery candidates; pending -> readyRepository CellActual membership, visibility and write policyEach sensitive transaction checks its own authorityUUIDdurable domain statefinal authorization / guardscollaboration logicdurable outcomeRepository components and collaboration controlGit state, authorization and collaboration meet at one Repository Cell transaction boundary.CURRENT SERVING PATH +
    +
    Read the explanation and source code

    The Repository Cell stores ACLs and visibility alongside issues, PRs, reviews, line discussions, commit checks and branch rules. This placement lets final commands check current policy against the same state that they mutate. Directory listings only suggest candidate repositories; the product rechecks actual Repository Cell access before exposing them.

    +

    For a merge, preparation captures exact base/head revisions and may use native merge-tree or commit-tree to construct candidate objects. MergePull checks the actor, revisions, current branch state, required checks, reviews and unresolved discussions before changing the branch and PR state. The candidate remains provisional until that command publishes. Check results are API records tied to commits and attempts; they do not imply that Canopy currently runs CI through a Cellule Workflow capability. See merge command, candidates and branch rules.

    +

    Why policy belongs beside refs

    +

    An ACL grant, branch rule, review or check can change after native preparation starts. Co-locating these facts with refs lets the final Repository Cell command examine their current versions and exact reviewed revisions in the same transaction that changes the branch. A ready merge candidate describes prepared bytes; it does not grant a standing right to merge them later.

    +

    Issues, comments and PR records use repository-local identities and optimistic versions. Check attempts retain reporter/context policy and terminal results. These records participate in repository recovery, so owner movement does not split collaboration history from the Git branch state it governs. External runners, durable event delivery and autonomous workflow execution are separate integration work.

    Read the complete design section

    Source map at the inspected revision

    +
    +
    09

    Packed storage evolution and remaining integration

    +

    The current product still uses per-object SQL authority, with an active archived-pack optimization. The new subsystem moves canonical inventories, graph metadata, ref snapshots and exact outcomes into verified immutable artifacts and short Cell publication commands. Preparation, publication dispatch, certificates, policy pages, recovery records and readers exist, but the full serving path is not registered and the hard cutover remains incomplete. Multi-capability repository Cells are a separate proposal.

    +
    100%SVGPNG @2x
    +Packed storage evolution and remaining integrationThe current product still uses per-object SQL authority, with an active archived-pack optimization. The new subsystem moves canonical inventories, graph metadata, ref snapshots and exact outcomes into verified immutable artifacts and short Cell publication commands. Preparation, publication dispatch, certificates, policy pages, recovery records and readers exist, but the full serving path is not registered and the hard cutover remains incomplete. Multi-capability repository Cells are a separate proposal. + + +ACTIVE SERVING MODELNEW PACKED CATALOG MODEL / IMPLEMENTED BUILDING BLOCKSPer-object SQL authorityCanonical objects + edges + closure in Repository CellInline / chunks / external / archived packed blobsCompletePush publishes refs and stored responseCurrent gateway preparationSerialized native push phase within each gatewayVerified immutable body / archive uploadsSQL ingestion batches and graph certificationPrivate native preparationRetain exact input + native resultOwned staging / preparation attemptsPack + index + metadata artifactsVerify canonical identity and closureImmutable catalog metadataSorted directory runs + leveled indexSource roots identify physical packsCatalog binds directory + sourcesImmutable ref-state / outcome rootsPublication coordinatorBounded class / account dispatchRegister exact SDK command identityBind owner fence / attempt / baseTrusted certificates + ref policy pagesShort Cell publicationCheck current ACL / policy / generationPublish catalog + ref root + outcomeRetain generation / recovery factsCellule still supplies durability gatePinned catalog readersVerify catalog / source descriptorsResolve OID -> pack locationIndependent bounded reader facilitiesDo not infer access from pack membershipcertified rootpacks/publication/mod.rs explicitly says these commands are not registered on the legacy repository serving path.Required production workWire startup, HTTP, SSH and generated producers/readers; select the fresh schema;complete recovery, collection, backup and capacity qualification.Separate Cell capability proposalAdding KV, Queue, Workflow, Blob, Cron, Timer and Effects to the same Repository Cellneeds framework composition and autonomous-progress acceptance gates.What does not change in the intended designThe repository UUID, one fenced authoritative publisher, final policy checks and durable outcome acknowledgement remain. Immutable uploaded artifacts are preparation until an authorizedCell command publishes them.active serving pathnew implemented primitivesimmutable metadata / bytesintegration and qualification openPacked storage evolution and remaining integrationThese primitives exist in code; the replacement production request path and fresh-schema cutover remain open.INCOMPLETE PRODUCT INTEGRATION +
    +
    Read the explanation and source code

    The newer packs subsystem implements immutable native-pack metadata, canonical directory runs, leveled indexes, source roots, catalogs, ref-state roots and exact outcome artifacts. Staging and preparation retain attempts and input custody; trusted certificates bind verified work to the repository, owner fence and base. A bounded publication coordinator dispatches registered commands by class and account. Short final commands publish catalog/ref state and outcomes while rechecking authority and policy. Uploaded artifacts remain preparation until authorized publication selects them.

    +

    The publication module explicitly states that its commands are not registered on the legacy serving path. The implementation status lists production startup, HTTP, SSH, generated producers/readers and the fresh-schema hard cutover as open work, along with recovery, collection, backup and capacity qualification. The atlas therefore separates implemented primitives from a complete replacement serving system.

    +

    A separate repository capability proposal aims to compose SQL, KV, Queue, Workflow, Blob, Cron, Timer and Effects within one repository identity, fence and recovery boundary. Current RepositoryModule declares CatalogRole::Sql and has empty workflow/activity inventories. Autonomous per-repository work and composed capabilities remain acceptance goals. Framework support for a primitive does not mean Canopy has wired that primitive into its Repository Cells.

    +

    Hard cutover requirements

    +

    The new model shifts large inventories, ref snapshots and exact outcomes into immutable artifacts while retaining short authoritative SQL publication commands. Prepared certificates bind verified artifacts to the repository, owner fence and base state. Service-owned staging, checkpoints and recovery records must retain the original command inputs across uncertainty; rebuilding a similar-looking command is not exact recovery.

    +

    Completing this model requires compatible registration and recovery admission, production HTTP/SSH producers, all readers, a fresh-schema format selection, retained-root collection, isolated restore, maintenance and capacity qualification. The current implementation status, rather than the presence of individual types or local tests, determines whether each gate is closed. The active archived-pack optimization must not be presented as the completed immutable-catalog cutover.

    +

    Composed capabilities are a separate design

    +

    Installing additional primitives into a repository requires explicit capability metadata, registry validation, typed APIs, durable work inspection and runner/lifecycle integration under the same target and fence. A SQL table named queue is not sufficient. Current release/acquisition logic must account for every installed primitive's live work before moving or retiring the Cell.

    +

    Canopy's product-level blob/LFS store and check records do not instantiate Cellule Blob or Workflow capabilities. The proposal describes the desired composition; it must be rechecked against the exact framework revision selected for implementation. Historical dependency pins in that proposal are not the pin of the serving snapshot documented here.

    Read the complete design section

    Source map at the inspected revision

    +
    + \ No newline at end of file diff --git a/diagram/canopy-architecture/manifest.json b/diagram/canopy-architecture/manifest.json new file mode 100644 index 0000000..436456d --- /dev/null +++ b/diagram/canopy-architecture/manifest.json @@ -0,0 +1,111 @@ +[ + { + "slug": "01-system-overview", + "title": "Canopy system architecture", + "explanation": "Canopy is an embedded Cellule application. Public ingress and native Git run in Canopy, while each durable Directory or Repository Cell has one fenced owner. A receiving node may call a remote owner. The object store holds authority records, SQLite recovery roots and product body objects. Native Git caches are disposable.", + "sources": [ + "crates/canopy-server/src/server/mod.rs", + "crates/canopy-server/src/server/residency/mod.rs", + "crates/canopy-server/src/git_gateway/mod.rs", + "crates/canopy-server/src/lib.rs", + "crates/canopy-server/src/schema.sql" + ] + }, + { + "slug": "02-canopy-on-cellule", + "title": "How Canopy models its domain on Cellule", + "explanation": "CanopyApplication registers DirectoryModule and RepositoryModule as SQL Cell types. Directory uses a fixed shard; repository UUIDs derive entity partitions. Product wrappers hold typed SqlCell handles. Cellule owns execution and persistence mechanics, and Canopy owns protocols and authorization. Other framework primitives are available but are not composed into Canopy Repository Cells.", + "sources": [ + "crates/canopy-server/src/lib.rs", + "crates/canopy-server/src/directory/mod.rs", + "crates/canopy-server/Cargo.toml", + "docs/repository-cell-primitives.md" + ] + }, + { + "slug": "03-durable-command", + "title": "Cellule command execution and durability", + "explanation": "This is the object-store durability path assembled by Canopy. The Cell owner records domain state and the request outcome together, captures SQLite through LTX and publishes the exact root using a fenced conditional authority update. Only then does the caller receive Committed and a receipt. A Cellule follower-log mode exists, but this diagram does not claim Canopy enables it.", + "sources": [ + "crates/canopy-server/src/server/mod.rs", + "crates/canopy-server/src/lib.rs" + ] + }, + { + "slug": "04-request-routing", + "title": "Request routing and repository residency", + "explanation": "Directory state authenticates accounts and resolves ready repository names. RepositoryManager binds a stable target to either a local actor or a signed peer client. Cold transitions are admitted and supervised; remote-owner loss triggers acquisition on demand. Repository authorization is checked using current durable state. Response pins prevent eviction during streaming.", + "sources": [ + "crates/canopy-server/src/server/residency/mod.rs", + "crates/canopy-server/src/server/peer.rs", + "crates/canopy-server/src/repository_http/authorization.rs", + "crates/canopy-server/src/server/discovery.rs" + ] + }, + { + "slug": "05-git-push", + "title": "Git push from wire input to durable refs", + "explanation": "After outer authentication, GitGateway spools and hashes the encoded request and binds a logical push ID to the actor and digest. It can replay a completed push before decoding and native work. New attempts prepare a private ref snapshot, run receive-pack, ingest canonical objects, certify graph closure and stage the report and ref plan. CompletePush checks current policy and expected ref versions and commits accepted refs with the canonical response pointer. Cellule gates acknowledgement on durable publication.", + "sources": [ + "crates/canopy-server/src/git_gateway/mod.rs", + "crates/canopy-server/src/git_gateway/preflight.rs", + "crates/canopy-server/src/git_gateway/push.rs", + "crates/canopy-server/src/push/mod.rs", + "crates/canopy-server/src/refs.rs" + ] + }, + { + "slug": "06-read-and-lfs", + "title": "Fetch, browse and Git LFS data paths", + "explanation": "Fetch takes consistent refs, validates requested object reachability, hydrates selected verified objects and uses native upload-pack to stream the wire response. Browser APIs read through Cell metadata and verified body readers. LFS uploads first publish and verify immutable bytes, then commit authorized metadata; downloads verify manifest-bound parts. The active serving schema supports inline, chunked, external and packed storage records, distinct from the incomplete immutable catalog replacement.", + "sources": [ + "crates/canopy-server/src/git_gateway/fetch.rs", + "crates/canopy-server/src/git_gateway/hydration.rs", + "crates/canopy-server/src/git_gateway/discovery.rs", + "crates/canopy-server/src/git_read/mod.rs", + "crates/canopy-server/src/lfs/upload.rs", + "crates/canopy-server/src/lfs/read.rs", + "crates/canopy-server/src/schema.sql" + ] + }, + { + "slug": "07-owner-recovery", + "title": "Owner lifecycle and exact recovery", + "explanation": "Canopy locks its local runtime workspace, verifies storage capabilities, validates the selected release and enrolls a signed node lease before serving. New or cold Cells bootstrap, acquire idle authority or fence an expired owner and restore the authority-pinned SQLite root. Recovery preserves the request outcome ledger. Native caches rebuild afterward. Shutdown and maintenance supervise all admitted work; backup copies durable roots and referenced external bodies into an independent prefix.", + "sources": [ + "crates/canopy-server/src/server/mod.rs", + "crates/canopy-server/src/server/lifecycle.rs", + "crates/canopy-server/src/server/workspace/mod.rs", + "crates/canopy-server/src/deployment/recovery.rs", + "crates/canopy-server/src/deployment/backup/mod.rs", + "docs/operations.md" + ] + }, + { + "slug": "08-domain-and-policy", + "title": "Repository components and collaboration control", + "explanation": "The Repository Cell colocates Git refs and graph metadata with ACLs, visibility, issues, PRs, reviews, line threads, check attempts, branch rules and operation outcomes. Merge preparation may run native Git outside the authority boundary, but MergePull rechecks current exact revisions, policy and required evidence before atomically changing the branch and PR state. Directory and Repository state remain separate Cells.", + "sources": [ + "crates/canopy-server/src/schema.sql", + "crates/canopy-server/src/pulls/merge/command.rs", + "crates/canopy-server/src/pulls/candidates/mod.rs", + "crates/canopy-server/src/branch_rules/command.rs", + "crates/canopy-server/src/server/discovery.rs" + ] + }, + { + "slug": "09-packed-storage-evolution", + "title": "Packed storage evolution and remaining integration", + "explanation": "The current product still uses per-object SQL authority, with an active archived-pack optimization. The new subsystem moves canonical inventories, graph metadata, ref snapshots and exact outcomes into verified immutable artifacts and short Cell publication commands. Preparation, publication dispatch, certificates, policy pages, recovery records and readers exist, but the full serving path is not registered and the hard cutover remains incomplete. Multi-capability repository Cells are a separate proposal.", + "sources": [ + "crates/canopy-server/src/packs/publication/mod.rs", + "crates/canopy-server/src/packs/publication/coordinator.rs", + "crates/canopy-server/src/packs/catalog/mod.rs", + "crates/canopy-server/src/packs/directory/mod.rs", + "crates/canopy-server/src/packs/ref_state/mod.rs", + "docs/large-repository-implementation-status.md", + "docs/large-team-scalability.md", + "docs/repository-cell-primitives.md" + ] + } +] diff --git a/diagram/canopy-architecture/render_pngs.py b/diagram/canopy-architecture/render_pngs.py new file mode 100644 index 0000000..62af1e0 --- /dev/null +++ b/diagram/canopy-architecture/render_pngs.py @@ -0,0 +1,17 @@ +#!/usr/bin/env python3 +"""Render the SVG atlas at twice its viewBox resolution with librsvg.""" +from pathlib import Path +import shutil +import subprocess +import xml.etree.ElementTree as ET + +root = Path(__file__).resolve().parent +renderer = shutil.which('rsvg-convert') +if renderer is None: + raise SystemExit('rsvg-convert is required for this local PNG export helper.') +for svg in sorted(root.glob('*.svg')): + dims = list(map(float, ET.parse(svg).getroot().attrib['viewBox'].split())) + png = svg.with_name(svg.stem+'@2x.png') + subprocess.run([renderer, '-w', str(int(dims[2]*2)), '-h', str(int(dims[3]*2)), + '-o', str(png), str(svg)], check=True) + print(png.name) diff --git a/docs/README.md b/docs/README.md index 71320f8..91bfcdf 100644 --- a/docs/README.md +++ b/docs/README.md @@ -9,6 +9,7 @@ Use this page to choose a document by task. Canopy's hosting core supports stock | If you need to… | Read | What you will find | | --- | --- | --- | | Understand the server and try a local deployment | [Project README](../README.md) and [bounded deployment](../deploy/README.md) | Configuration, runtime commands, resource boundary and recovery procedures | +| Understand the system design | [Canopy design](../DESIGN.md) | Goals, invariants, Cellule integration, control flows, failures, resource ownership and tradeoffs | | Decide which Git operations work | [Git compatibility](git-compatibility.md) | Stock-client evidence, restrictions and provider qualification commands | | Implement an API or storage change | [Persisted contracts](contracts.md) | Identity, authorization, protocol, durability and HTTP behavior | | Find the right Rust crate | [Rust workspace](workspace.md) | Crate ownership, dependencies and build commands | @@ -28,6 +29,8 @@ Use this page to choose a document by task. Canopy's hosting core supports stock ## Understand where data lives +Read the root [Canopy design document](../DESIGN.md) for authority boundaries and request, push, read and recovery flows. Its [architecture gallery](../diagram/canopy-architecture/index.html) includes nine diagrams with standalone SVG/PNG exports and separates the active serving architecture from incomplete packed-storage integration. + The Directory Cell resolves names and accounts. Each repository has its own Repository Cell, which owns the durable Git and collaboration state. Native Git uses a disposable cache for wire protocols. Large Git blobs and Git LFS bodies live in immutable object-store objects referenced by the repository's SQLite state. ![Canopy gateway, Directory Cell, Repository Cell, object store and disposable Git cache](architecture.svg)