Skip to content

Repository files navigation

s32p: S3 to POSIX proxy and gateway

s32p-proxy is a Rust-based S3-compatible proxy + gateway manager built on Cloudflare Pingora.


⚠️ EARLY DEVELOPMENT WARNING ⚠️

This project is in an early development stage and is not ready for production use. All aspects of the implementation may change significantly in future versions. Use only for testing and development purposes.


It accepts S3 client requests, maps S3 identities (SigV4 access keys) to Unix users, and routes traffic to per-access-key worker processes (currently: VersityGW, or experimental alternative s32p-gateway) that expose a shared POSIX filesystem with kernel-enforced permissions.

A central design goal is to preserve Unix security semantics:

All filesystem access happens inside a process already running as the target Unix user.

The proxy itself does not perform filesystem I/O.


Goals

  • Provide an S3-compatible endpoint backed by a POSIX filesystem
  • Enforce per-user isolation using Unix UIDs/GIDs (kernel-level security)
  • Support standard S3 clients (AWS CLI, MinIO client / SDKs)
  • Start workers on demand, then reuse them (per access key + profile)
  • Prevent easy DoS via worker spawn storms
  • Keep proxy logic (auth/routing/management) separate from the storage worker

Repository layout

This repository is a Rust workspace with multiple crates:

  • s32p-proxy: the Pingora-based proxy binary
  • s32p-directory: shared Directory API + YAML/OpenBao backends + shared YAML file format
  • s32p-admin: management library (OpenBao write access, import/export, etc.)
  • s32p-ctl: CLI wrapper around s32p-admin (operator tooling)
  • s32p-support: shared support library (common config/types/errors/helpers used across crates)
  • s32p-gateway: the per-access-key worker gateway binary (launched by the proxy; it runs as alternative to VersityGW)

Outside the workspace:

  • tests-integration/: pixi-managed pytest harness. Spawns the real proxy + workers, drives them through a matrix of S3 clients (boto3 path-style + virtual-hosted, aws-cli, mount-s3 + rclone FUSE mounts) plus AWS-compat probes. Run with cd tests-integration && pixi run test. Layout + how to extend in tests-integration/README.md.

Architecture

    ┌──────────────┐
    │  S3 Clients  │  (aws cli, mcli, SDKs)
    └──────┬───────┘
           │  HTTP(S)
           ▼
┌───────────────────────────────────────────┐
│                 s32p-proxy                │
│            (Pingora HTTP proxy)           │
│                                           │
│  • parses Authorization → access key      │
│  • maps access key → unix user            │
│  • looks up user + ACLs via Directory     │
│    (YAML file or OpenBao)                 │
│  • classifies requests (query+path)       │
│  • routes by "request class" via YAML     │
│    (proxy to profile or local response)   │
│  • gates worker spawn with SigV4          │
│    verification (auth header or presign)  │
│  • reverse proxies to per-user worker     │
│  • optional response header rewriting     │
└───────────────────┬───────────────────────┘
                    │ internal HTTP (loopback) or
                    │ HTTP over Unix domain sockets (UDS)
                    ▼
      ┌────────────────────────────────┐
      │      Per-access-key workers    │
      │        (e.g. s32p-gateway      │
      │           or VersityGW)        │
      │                                │
      │  • runs as unix user           │
      │  • validates SigV4 again       │
      │  • uses a staged posix_root    │
      └───────────────┬────────────────┘
                      │ POSIX syscalls
                      ▼
┌───────────────────────────────────────────┐
│   Shared POSIX FS (real bucket data)      │
│                                           │
│  Each worker starts with a fresh temp     │
│  directory as its posix_root.             │
│  Symlinks to all accessible buckets       │ 
│  are created in that temp root before     │ 
│  the worker starts. The worker follows    │
│  those links.                             │
└───────────────────────────────────────────┘

Current Status

Proxy / request flow

  • Pingora proxy-mode HTTP server

    • Reverse proxies to locally spawned workers over loopback or Unix domain sockets (UDS)
    • Preserves SigV4-critical headers (notably the original Host)
    • Has a response_filter hook for response header rewriting
  • Request classification (crates/s32p-support/src/classifier.rs)

    • Shared by proxy + gateway (single source of truth for routing decisions)
    • Parses path + query parameters (and selected headers where needed, e.g. x-amz-copy-source)
    • Produces a high-level operation class key:
      • read (e.g. GetObject, HeadObject, ListObjectsV1/V2, ListBuckets, GetBucketLocation, HeadBucket, GetObjectAcl, GetBucketAcl)
      • write (e.g. PutObject, CopyObject, DeleteObject, DeleteObjects, RenameObject, PutObjectAcl, PutBucketAcl)
      • multipart (initiate/upload-part/list-parts/complete/abort + list uploads)
      • versioning (detected; read-side probes answered AWS-shape via aws_compat, write-side falls back to 501)
      • object_lock (same shape as versioning — feature-disabled responses via aws_compat)
      • bucket_admin (CreateBucket/DeleteBucket; detected and routed to NotImplemented — bucket lifecycle is operator-only via s32p-ctl)
      • session (S3 Express CreateSession)
      • other
  • Config-driven routing (etc/s32p-proxy.yaml)

    • Routes based on the classifier class keys above.
    • Each class maps to one of four actions:
      • proxy (selects a worker profile)
      • not_implemented (uniform 501 NotImplemented response, after SigV4 validation)
      • aws_compat (per-S3Op AWS-shaped feature-disabled responses — 200 with empty <VersioningConfiguration/> for GetBucketVersioning, 404 ObjectLockConfigurationNotFoundError for GetBucketObjectLockConfiguration, 400 InvalidRequest "Bucket is missing Object Lock Configuration" for PutObjectRetention / PutObjectLegalHold, etc. — falling back to 501 for ops AWS always implements like PutBucketVersioning)
      • create_session (local handling of S3 Express CreateSession)
    • If a specific class key is not configured, the proxy falls back to the other route.
    • Always-on pre-routing gate: any request carrying server-side-encryption headers (SSE-C / SSE-S3 / SSE-KMS / SSE-KMS-DSSE) is rejected with 400 InvalidRequest before routing — the gateway has no encryption path and silent plaintext storage would be a data-confidentiality bug. See s3_compatibility_analysis.md for the per-variant table.

Gateway (s32p-gateway)

Experimental alternative to versitygw. Implements a growing subset of the S3 REST API directly on top of a POSIX filesystem.

Implemented operations:

  • Read
    • GetObject
    • HeadObject
    • HeadBucket
    • ListBuckets
    • GetBucketLocation
    • ListObjectsV1 (GET /{bucket}, legacy form) and ListObjectsV2 (GET /{bucket}?list-type=2)
  • Write
    • PutObject (streaming upload). Append is supported via the S3 Express extension x-amz-write-offset-bytes: N: the request body is written at byte offset N of the existing object; N must equal the current size (no overwrite, no holes), otherwise 412 PreconditionFailed.
    • CopyObject (server-side copy; size-limited by configuration)
    • DeleteObject
    • DeleteObjects (POST /?delete)
    • RenameObject (PUT ?renameObject with x-amz-rename-source; same-bucket only). Idempotent: re-running an already-applied rename returns success.
  • Multipart
    • CreateMultipartUpload (POST ?uploads)
    • UploadPart (PUT ?partNumber=N&uploadId=...)
    • UploadPartCopy (PUT ?partNumber=N&uploadId=... with x-amz-copy-source; optional x-amz-copy-source-range and the four x-amz-copy-source-if-* headers). When the first part of an upload is a whole-source UploadPartCopy (no copy-source-range), the source's size is recorded as a total_size_hint on the upload and used to size the assembly file's Lustre stripe count.
    • ListParts (GET ?uploadId=...)
    • ListMultipartUploads (GET /bucket?uploads)
    • CompleteMultipartUpload (POST ?uploadId=...)
    • AbortMultipartUpload (DELETE ?uploadId=...)
  • Directory bucket (S3 Express) handshake
    • CreateSession (GET /{bucket}?session) — issues short-lived session credentials and a x-amz-s3session-token. Subsequent data-plane requests sign with service=s3express and pass the session token either as the x-amz-s3session-token header or, for presigned URLs, as the equivalent query parameter. The proxy validates both paths against an in-memory session store; workers re-validate on every request.
  • ACL
    • GetObjectAcl, GetBucketAcl (GET ?acl)
    • PutObjectAcl, PutBucketAcl (PUT ?acl) — accepted as a no-op when the requested ACL matches the current POSIX state; mismatches are rejected. The directory ACLs in s32p-ctl remain authoritative.
  • Object tagging
    • GetObjectTagging, PutObjectTagging, DeleteObjectTagging (GET/PUT/DELETE /{bucket}/{key}?tagging)
    • x-amz-tagging header on PutObject (tags-at-creation) and on CopyObject with x-amz-tagging-directive: COPY|REPLACE (default COPY mirrors source tags).
    • Storage model. Tags live in the user.s32p.tags xattr on the object file, encoded URL-form (team=a&stage=raw) — the same shape as the wire header. A POSIX-created file has no xattr → empty TagSet (still a valid S3 object). cp -a / rsync -X preserve tags across POSIX copies; plain cp drops them, which is equivalent to creating a fresh tagless object — acceptable degradation.
    • Limits enforced: ≤ 10 pairs, key 1..=128 chars, value 0..=256 chars, total payload ≤ 2 KB, AWS-documented charset ([A-Za-z0-9 +-=._:/@]), no duplicate keys.
    • Bucket-level tagging (?tagging on a bucket) is not implemented; the proxy continues to route it to the other worker (versitygw) per the default config.
  • Object metadata
    • x-amz-meta-* headers on PutObject (round-tripped on HeadObject / GetObject).
    • Content-Type on PutObject (round-tripped on HeadObject / GetObject).
    • CopyObject with x-amz-metadata-directive: COPY|REPLACE — default COPY mirrors source user metadata and Content-Type; REPLACE uses the request headers (clears when none are supplied).
    • Storage model. User metadata lives in user.s32p.meta (URL-form, lowercased keys); explicit Content-Type lives in user.s32p.content_type. The gateway also honors the freedesktop user.mime_type xattr read-only as a Content-Type fallback (GUI file managers and gio set populate this; the gateway never writes it). A POSIX-created file has no xattrs → empty metadata, Content-Type resolved by the fallback ladder below.
    • Content-Type resolution ladder (HEAD/GET): user.s32p.content_type → user.mime_type → extension map via mime_guess → application/octet-stream.
    • Limits enforced: user metadata payload ≤ 2 KB total, ≤ 32 pairs, key length 1..=128 chars (HTTP-token shape [A-Za-z0-9_-]), value ≤ 1024 chars, no duplicate keys. Content-Type ≤ 256 bytes and must parse as type/subtype[;params].
    • Not stored (deferred): the other five HTTP passthrough headers (Content-Encoding, Content-Language, Content-Disposition, Cache-Control, Expires), and metadata set on CreateMultipartUpload (multipart-completed objects carry no metadata even if the create-call supplied headers).

Notes / behavior:

  • Supports Range: bytes=... (returns 206 Partial Content; invalid ranges return 416 InvalidRange). A single range returns the bytes directly; multiple ranges return a multipart/byteranges body (RFC 7233), one part per range — capped at 50 ranges, and any unsatisfiable range fails the whole request with 416. A multi-range HEAD ignores the Range and returns the full object (200).
  • Rejects most query parameters for now, except those required for:
    • ?location, ?list-type=2, ?delete, ?acl, ?renameObject, ?session, and the multipart query parameters (?uploads, ?uploadId=..., ?partNumber=...)
  • SigV4 presigned URL query parameters (X-Amz-*) are supported and do not count as "effective" query parameters for routing/handling. A small fixed set of other keys is also treated as non-effective: x-id, content-type, cache-control, content-encoding, content-disposition, expires, x-amz-storage-class.
  • ListObjectsV2 supports Lustre Lazy Size on MDS (LSOM) when built with the Lustre feature.
  • When built with --features lustre, the gateway creates new files with Lustre striping via llapi_file_create().
    • Config: S32P_LUSTRE_MAX_STRIPE_COUNT (default: 4) caps the stripe count.
    • Serial uploads (PutObject, CopyObject, and any temp/staging files):
      • stripe_size = S32P_CHUNK_SIZE_MB × 1 MiB (env var is in MiB; default 4 → 4 MiB)
      • stripe_count = ceil(file_size / stripe_size), capped by S32P_LUSTRE_MAX_STRIPE_COUNT
    • Multipart uploads:
      • direct.bin: stripe_size = min(stripe_size_serial, part_size). stripe_count is sized from whatever total-size estimate is available — the upload's total_size_hint (set when the first part is a whole-source UploadPartCopy) or, in CompleteMultipartUpload's fast path, the now-known final size — via ceil(estimate / stripe_size) and capped by S32P_LUSTRE_MAX_STRIPE_COUNT. With no estimate, the count falls back to the configured max. Striping is fixed at file creation; the hint matters only the first time direct.bin is opened.
      • individual part files are striped like serial uploads.
  • ETag for final objects is generated from the inode number.
  • Object-level metadata is stored in user.* xattrs. Tagging, user metadata (x-amz-meta-*), and Content-Type each have a dedicated xattr (user.s32p.tags, user.s32p.meta, user.s32p.content_type). The POSIX interop contract still holds — absence of any auxiliary state yields a valid object (POSIX-created file = empty metadata, no error). cp -a / rsync -X preserve every category; plain cp drops them and the destination becomes a metadata-less-but-valid object. Not stored: the other five HTTP passthrough headers (Content-Encoding, Content-Language, Content-Disposition, Cache-Control, Expires) and any metadata set on CreateMultipartUpload — both deferred until a real client need surfaces. Same principle as PutObjectAcl being a no-op when the requested ACL matches current POSIX state.

Multipart upload (s32p-gateway) — server-side assembly algorithm

Multipart uploads are implemented in crates/s32p-gateway/src/multipart.rs.

The design goal is to avoid creating a full extra copy of the object on the server during completion. The gateway does this by writing parts into an assembly file that can be renamed into place as the final object.

On-disk layout (per bucket)

For a bucket with root <bucket_root>, the gateway reserves a hidden directory (default name: .s32p-mpu):

  • <bucket_root>/.s32p-mpu/uploads/<upload_id>/meta.json
    JSON metadata for the upload (bucket/key, state, and a map of uploaded parts).
  • <bucket_root>/.s32p-mpu/uploads/<upload_id>/lock
    A file used with flock(LOCK_EX) to serialize metadata updates and completion.
  • <bucket_root>/.s32p-mpu/uploads/<upload_id>/direct.bin
    The assembly file (random-access writes at part offsets).
  • <bucket_root>/.s32p-mpu/uploads/<upload_id>/parts/part-00001.bin (etc.)
    Fallback storage for parts that cannot safely be placed into direct.bin.

The gateway also prevents clients from reading/writing/deleting objects inside the reserved multipart directory by treating that first path segment as “reserved”.

Part upload placement

When a part arrives (PUT ?partNumber=N&uploadId=...), the gateway decides where to store it:

  1. It reads and updates meta.json under an exclusive lock (short critical section).

  2. It tries to determine a stable assumed part size:

    • If part #1 is uploaded and no size is known yet, its length becomes the assumed part size.
    • If no assumed size exists and the gateway sees two parts with the same size, it promotes that size to the assumed part size.
  3. If an assumed part size is known and the part is not larger than it, the gateway computes the direct placement offset:

    offset = (partNumber - 1) * assumed_part_size

    and writes the request body directly into direct.bin at that offset (random-access file write).

  4. Otherwise, it writes the part as an individual file under parts/ and records that in metadata.

Metadata records, per part:

  • size
  • an upload-time ETag (stable, client-visible; completion does not validate ETags)
  • timestamp
  • storage kind: { direct: off } or { file: name }

Completion: assembling without a full server-side copy

On CompleteMultipartUpload (POST ?uploadId=...), the gateway:

  1. Reads the XML body (requested part numbers).
  2. Takes an exclusive lock for the entire completion to prevent concurrent UploadPart and to make the final rename deterministic.
  3. Validates the request:
    • Parts must be contiguous starting at 1 (1..=lastPart).
    • Every requested part must exist in meta.json.

Then it chooses one of two assembly paths:

Fast path (rename direct.bin into place)

This path is taken when the upload matches the classic “fixed-size parts + last part shorter or equal” layout:

  • assumed_part_size is known
  • parts 1..last-1 are exactly assumed_part_size
  • any parts already stored as direct are at the expected offsets

Fast-path algorithm:

  1. Compute final object size:

    final_size = (lastPart - 1) * assumed_part_size + size(lastPart)

  2. Ensure all missing parts are present in direct.bin:

    • For parts stored as separate files, copy them into direct.bin at their final offsets (range copy).
  3. ftruncate(direct.bin, final_size)

  4. Rename direct.bin → <final_object_path>

If the rename fails with cross-device (EXDEV), the gateway performs a single copy of the already-assembled direct.bin into a temp file next to the destination and renames that temp file into place. Importantly, it still avoids “assemble → copy again” behavior.

✅ Why this avoids a full copy: in the common case (same filesystem), completion becomes a metadata operation (rename) after assembling into direct.bin. There is no “write whole object into a second file” step.

Fallback path (sequential staging file)

If the fixed-size/offset conditions don’t hold, the gateway assembles sequentially:

  1. Create a staging output file:
    • If the upload directory is on a different filesystem than the destination, the staging file is created next to the destination (to avoid an extra copy later).
  2. Copy parts in order into the staging file:
    • For direct parts: copy the required range out of direct.bin.
    • For file parts: copy the entire part file.
  3. Rename staging file → final object path.

Finally, on success the gateway marks the upload as completed, returns the completion XML + ETag, and removes the upload directory.

Practical highlights

  • The gateway’s “happy path” is optimized for large objects: parts are placed directly into their final offsets, and completion is typically a truncate + rename.
  • The implementation is careful to avoid data corruption:
    • Direct placement is only used when offsets can be computed safely (no overlap risk).
    • Metadata updates are protected by flock and written atomically (meta.json.tmp → rename).

Directory backends (users, buckets, ACLs)

A “Directory” provides:

  • user_by_access_key(access_key) → unix identity + secret key
  • buckets_for_access_key(access_key) → list of visible buckets with effective access

Backends:

  • YAML: load once at startup into HashMaps
  • OpenBao: AppRole login + KV v2 reads, using indices for efficient lookup

Notes:

  • Multiple access keys may map to the same Unix user (same username/uid/gid).
  • Workers are keyed by (access_key, worker_profile) (not just uid), because the staged bucket set is per access key.

Bucket ACL model

Buckets have an ACL list. Each ACL entry grants an access level to a principal:

  • access_key principal (direct grant)
  • group_name principal (POSIX group name grant)

Effective access for a caller is the max of all matching entries (e.g. read_write beats read_only).

Group membership is resolved from the OS at runtime (username → gids → group names).

Local responses (no proxying)

  • Shared S3 REST-XML helpers live in crates/s32p-support/src/s3resp.rs and crates/s32p-support/src/s3xml.rs (errors like AccessDenied, SignatureDoesNotMatch, NotImplemented, plus body builders); proxy-side glue (respond_bytes() and friends) is in crates/s32p-proxy/src/responses.rs.

SigV4 validation behavior

  • DoS mitigation via “spawn gating”

    • If a worker is not running, the proxy performs SigV4 verification without reading the body:
      • Standard SigV4 auth header (Authorization, x-amz-content-sha256)
      • SigV4 presigned URLs (query signature; typically UNSIGNED-PAYLOAD)
    • Only if the signature is valid will the proxy start the worker
    • Once a worker is already running, the proxy does not fully validate SigV4; it only extracts the access key for routing and forwards the request to the worker
  • Workers (VersityGW or s32p-gateway) still validate SigV4 again (cannot be disabled).

  • Clock-skew gate (header SigV4 only). Header-signed requests whose x-amz-date differs from the server clock by more than 15 minutes (in either direction) are rejected with RequestTimeTooSkewed (HTTP 403). This matches AWS S3's documented window and exists so a captured signed request can't be replayed indefinitely. Presigned URLs have their own X-Amz-Expires-based check (rejected with AccessDenied / "Request has expired") and don't use this window. The 15-minute value is a constant in s32p-support (HEADER_SIGV4_MAX_SKEW).

Worker lifecycle management (src/worker_manager.rs)

  • Workers are keyed by (access_key, worker_profile)
    • supports multiple access keys mapped to the same unix user but different bucket ACLs
    • allows routing different command classes to different worker profiles per access key
  • Workers are started via a configurable launcher (default: restricted-exec)
    • If running as root and pass_user_flag_if_root=true, the proxy passes --user <username>
    • When workers.launcher.landlock is enabled (default true), the proxy also passes --rw <posix_root>, then per accessible bucket either --rw <bucket.data_path> (for read_write grants) or --ro <bucket.data_path> (for read_only grants), plus --resolve-libs and --allow-nss. Landlock therefore enforces the same ACL access level the gateway already checks (see §Security Notes), at the kernel layer.
  • Upstream bind (per worker):
    • TCP loopback: 127.0.0.1:<port>
    • Unix domain socket: <uds_run_dir>/<uid>-<instance_id>/<profile>.sock
      • <uds_run_dir> defaults to /run/s32p (root) or $XDG_RUNTIME_DIR/s32p (non-root); override per profile via workers.profiles.<name>.upstream.uds_run_dir.
      • <instance_id> is a per-worker random suffix. It disambiguates concurrent proxy processes on the same host and disambiguates workers that share (uid, profile) within a single proxy — which happens when two access keys map to the same uid (a documented design point).
      • One socket file per worker profile under the per-worker run dir.
  • Readiness probing: connect loop until port is reachable
  • Idle shutdown after idle_timeout_secs
  • Sweeper removes dead/idle workers periodically (sweep_interval_secs)
  • Staged posix_root
    • On each worker start, the proxy creates a fresh temp directory under workers.posix_root
    • Creates symlinks for all buckets visible to the access key:
      • <temp>/<bucket_name> → <bucket.data_path>
    • Passes that temp directory as {{posix_root}} to the worker
    • When the worker stops, the temp directory is removed

Worker args / env templates are expanded per spawn by worker_manager.rs::render_template. Unknown tokens cause a hard error. Available tokens:

Token Value
{{username}} Unix username of the target user
{{uid}} numeric UID
{{gid}} numeric GID
{{access_key}} the SigV4 access key for this worker
{{secret_key}} matching secret key
{{posix_root}} the staged temp dir with bucket symlinks
{{bind_addr}} endpoint string (host:port for TCP, path for UDS)
{{bind_uds}} alias of {{bind_addr}} (use whichever reads better in your config)
{{region}} from server.region
{{virtual_hosted_suffixes}} comma-joined server.virtual_hosted_suffixes
{{log_level}} from server.log_level

Separately, {{install_bin_dir}} is a config-load-time placeholder resolved only inside workers.launcher.path and workers.profiles.*.exec (not in args/env). It expands to the directory containing the running s32p-proxy executable (e.g. target/debug in dev, /usr/local/bin after install).


Configuration

Primary configuration: etc/s32p-proxy.yaml

Key sections:

  • server.listen / server.public_scheme (both required)
  • server.region (required; used as the SigV4 region and exposed to workers via {{region}})
  • server.log_level (logging configuration, supports RUST_LOG format)
  • server.virtual_hosted_suffixes (virtual-hosted-style bucket detection)
  • server.tls_cert_path / server.tls_key_path (optional; enables HTTPS on the listen address)
  • server.shutdown_grace_period_secs (graceful-shutdown timeout, default 10s)
  • auth.* (directory backend selection and credentials)
  • workers.posix_root (base dir for per-worker temp roots)
  • workers.launcher.* (includes path, pass_user_flag_if_root, landlock)
  • workers.lifecycle.* (idle_timeout_secs, sweep_interval_secs)
  • workers.profiles.* (worker templates: exec, args, env, upstream)
  • routing.class_map.* (routes classifier classes to actions)

Virtual-hosted-style bucket support

The proxy automatically detects and supports both path-style and virtual-hosted-style bucket addressing:

  • Path-style: https://s3.example.com/bucket/key (bucket in path)
  • Virtual-hosted-style: https://bucket.s3.example.com/key (bucket in host)

Configure domain suffixes for virtual-hosted-style detection via server.virtual_hosted_suffixes:

server:
  virtual_hosted_suffixes:
    - "127.0.0.1.nip.io"
    - "s3.example.com"
    - "s3.localhost"

When a request's Host header ends with a configured suffix, the proxy extracts the bucket name from the host and the key from the path. Port numbers are handled correctly (e.g., bucket.suffix:9000).

Workers receive the suffixes via the S32P_VIRTUAL_HOSTED_SUFFIXES environment variable or {{virtual_hosted_suffixes}} template:

workers:
  profiles:
    s32p-gateway:
      env:
        S32P_VIRTUAL_HOSTED_SUFFIXES: "{{virtual_hosted_suffixes}}"

Directory backend selection

Choose a directory backend via auth.backend:

  • yaml — load a local directory file
  • openbao — use OpenBao (Vault-compatible API) with AppRole + KV v2 + indices

YAML backend config

auth:
  backend: "yaml"
  yaml:
    path: "/etc/s32p/directory.yaml"

OpenBao backend config (AppRole)

auth:
  backend: "openbao"
  openbao:
    address: "http://127.0.0.1:8200"
    approle_mount: "approle"
    role_id_file: "/etc/s32p/role_id"
    secret_id_file: "/etc/s32p/secret_id"
    kv_mount: "secret"
    prefix: "s32p"

Directory cache (auth.cache.*)

The proxy can hold a small in-process TTL cache in front of the directory backend. The cache decorates any backend (so YAML and OpenBao share the same code path) and only matters in practice for OpenBao — every uncached request would otherwise hit Vault.

auth:
  cache:
    enabled: true            # optional; default depends on backend (see below)
    user_ttl_secs: 30        # positive `user_by_access_key` results
    buckets_ttl_secs: 30     # positive `buckets_for_access_key` results
    negative_ttl_secs: 5     # cached "no such access key"
    max_entries: 4096        # soft cap per map; over cap, prune expired then evict soonest-expiring
  • Defaults: cache is on by default for the OpenBao backend and off by default for the YAML backend (YAML is an in-memory HashMap lookup; the cache adds no benefit, only lock overhead). Setting enabled: true or enabled: false overrides either default.
  • Errors are never cached — backend errors (network failures, KV decode errors) always pass through so the next attempt sees fresh state.
  • Snapshot vs. cache. ACL state is independently snapshotted into spawned workers via S32P_BUCKET_ACL (see §Security Notes); that snapshot freezes for the worker's lifetime. The cache only changes how fresh the next worker spawn's snapshot is — it does not propagate ACL changes to running workers. ACL changes therefore take effect at max(buckets_ttl_secs, idle_timeout_secs) worst-case.
  • Single-flight deduplication. Under a cold cache, N concurrent first-requests for the same access key all see the miss but only one actually hits the backend; the rest wait on a per-key Arc<Mutex<()>> and pick up the now-cached result. This bounds backend fan-out by the number of distinct access keys, not by request concurrency.
  • Observability. Hits log at trace and misses + evictions at debug under the s32p_directory::cache target. To see misses without spamming on hits:
    RUST_LOG=info,s32p_directory::cache=debug cargo run --bin s32p-proxy
    Each event carries the access_key, the affected map (users/buckets), and the TTL applied — useful for confirming the cache is actually warm in production and for tuning TTLs.

YAML directory file format

Example /etc/s32p/directory.yaml:

version: 1

users:
  - access_key: "alice_key_1"
    secret_key: "alice_secret"
    username: "alice"
    uid: 1001
    gid: 1001

  # Second access key mapping to the same unix user:
  - access_key: "alice_key_2"
    secret_key: "alice_secret_2"
    username: "alice"
    uid: 1001
    gid: 1001

buckets:
  - id: "bkt-alice-photos"
    name: "photos"
    data_path: "/srv/s3/alice/photos"
    acl:
      - principal: { type: "access_key", access_key: "alice_key_1" }
        access: "read_write"
      - principal: { type: "access_key", access_key: "alice_key_2" }
        access: "read_write"

  - id: "bkt-team"
    name: "team"
    data_path: "/srv/s3/team"
    acl:
      - principal: { type: "group_name", name: "s3-team" }
        access: "read_write"
      - principal: { type: "access_key", access_key: "alice_key_1" }
        access: "read_only"

OpenBao KV layout (conceptual)

Under auth.openbao.kv_mount + auth.openbao.prefix, the directory uses KV v2 documents:

  • prefix/users/<access_key> → UserDoc
  • prefix/buckets/<bucket_id> → BucketDoc
  • prefix/index/access_key/<access_key> → { bucket_ids: [...] }
  • prefix/index/group/<group_name> → { bucket_ids: [...] }

The indices make buckets_for_access_key() efficient:

  • read index for the access key
  • resolve group names for the unix user
  • read index for each group name
  • load bucket docs and evaluate ACLs (authoritative)

Administration CLI (s32p-ctl)

s32p-ctl manages the Directory state (users, buckets, ACLs) for both supported backends:

  • --backend openbao (default): OpenBao/Vault KV v2 + indices
  • --backend yaml: local directory.yaml file

All commands work with both backends. The backend is selected via --backend.

Backend selection

YAML backend

For YAML, you must provide the directory file path:

s32p-ctl --backend yaml --yaml-path /etc/s32p/directory.yaml <COMMAND...>

OpenBao backend

For OpenBao, provide --address (or VAULT_ADDR) plus authentication:

  • Setup typically uses a root/admin token (--token or VAULT_TOKEN)
  • Normal operations typically use AppRole (--role-id(-file) + --secret-id(-file))

Common OpenBao flags:

  • --address http://127.0.0.1:8200 (or VAULT_ADDR)
  • --kv-mount secret (default: secret)
  • --prefix s32p (default: s32p)
  • --approle-mount approle (default: approle)
  • --token ... (or VAULT_TOKEN)
  • --role-id ... / --role-id-file ...
  • --secret-id ... / --secret-id-file ...

Example:

s32p-ctl --backend openbao \
  --address http://127.0.0.1:8200 \
  --kv-mount secret \
  --prefix s32p \
  --approle-mount approle \
  --role-id-file /etc/s32p/admin_role_id \
  --secret-id-file /etc/s32p/admin_secret_id \
  <COMMAND...>

Setup

OpenBao setup

Setup is idempotent and verifying — re-runs are safe and abort hard rather than overwrite incompatible state. It creates/updates:

  • mounts KV v2 at --kv-mount (default: secret) if missing; if a mount already exists at that path, it is verified to be type=kv with options.version=2 and reused — any other engine type or KV v1 aborts setup
  • enables AppRole auth method at --approle-mount (default: approle)
  • creates policies + roles, scoped to <kv-mount>/data/<prefix>/* and <kv-mount>/metadata/<prefix>/*:
    • s32p-proxy (read-only)
    • s32p-admin (read-write)
  • reads the role IDs and generates a fresh secret ID per role; if the corresponding --*-id-file flags are given, the four IDs are written there with mode 0600 (Unix)

The bootstrap token (--token or VAULT_TOKEN) must be allowed to write sys/mounts/<kv-mount>, sys/auth/<approle-mount>, sys/policies/acl/* and auth/<approle-mount>/role/*. A root token covers all of these; restricted bootstrap tokens need the corresponding capabilities.

To keep the token out of your shell history and out of ps//proc/<pid>/cmdline, prompt for it without echo:

export VAULT_ADDR=http://127.0.0.1:8200
read -rs VAULT_TOKEN && export VAULT_TOKEN   # paste token, hit enter; no echo

s32p-ctl --backend openbao setup \
  --proxy-role-id-file   /etc/s32p/proxy_role_id \
  --proxy-secret-id-file /etc/s32p/proxy_secret_id \
  --admin-role-id-file   /etc/s32p/admin_role_id \
  --admin-secret-id-file /etc/s32p/admin_secret_id

unset VAULT_TOKEN   # token is no longer needed

After setup the bootstrap token is not needed for normal operation:

  • the proxy authenticates via the s32p-proxy AppRole (role_id_file + secret_id_file in etc/s32p-proxy.yaml under auth.openbao)
  • further s32p-ctl commands (user add, bucket add, import-yaml, …) authenticate via the s32p-admin AppRole (--role-id-file + --secret-id-file)

Re-running setup generates an additional secret_id for each role; existing secret IDs stay valid until destroyed via auth/<approle-mount>/role/<role>/secret-id/destroy. Rotate explicitly if you need the old ones revoked.

YAML setup

Creates an empty directory file skeleton:

s32p-ctl --backend yaml --yaml-path /etc/s32p/directory.yaml setup

Users

Commands:

  • user add
  • user rm
  • user ls

Add a user. Only --username is required; everything else is derived from it:

s32p-ctl ... user add --username alice
ok
user alice uid=1001 gid=1001 access_key=alice
secret_key=xXDk-bbkR7GweeyJM_mfMjkSJ0UG-BLK79PItgebzDA
  • --username is lower-cased (POSIX user names are lower-case by convention, and getpwnam_r matches exactly)
  • --access-key defaults to that lower-cased user name
  • --uid / --gid default to the passwd entry of that user name; without one, both must be given explicitly
  • --secret-key defaults to 32 CSPRNG bytes as URL-safe base64, printed only at creation — it cannot be recovered afterwards, only replaced by re-running user add

Any of them can still be given explicitly, e.g. for an account this host's passwd database doesn't know, or for an access key that differs from the user name. An explicit --access-key is stored verbatim, case included — SigV4 compares it byte-for-byte against what the client sends, and AWS-style keys are upper-case:

s32p-ctl ... user add \
  --access-key alice_key_1 \
  --secret-key alice_secret \
  --username alice \
  --uid 1001 \
  --gid 1001

Remove a user (optionally scrubs their access_key from bucket ACLs):

s32p-ctl ... user rm --access-key alice_key_1 --cleanup-acls true

List users:

s32p-ctl ... user ls

Buckets

Commands:

  • bucket add
  • bucket rm
  • bucket ls
  • bucket acl-set

Add a bucket:

s32p-ctl ... bucket add \
  --name photos \
  --data-path /srv/s3/alice/photos \
  --grant ak:alice_key_1:read_write

--grant is repeatable and supports:

  • ak:<ACCESS_KEY>:read_only|read_write
  • group:<GROUP_NAME>:read_only|read_write

Optional: provide a stable bucket id (otherwise a UUID is generated):

s32p-ctl ... bucket add \
  --bucket-id bkt-alice-photos \
  --name photos \
  --data-path /srv/s3/alice/photos \
  --grant ak:alice_key_1:read_write

Remove a bucket:

s32p-ctl ... bucket rm --bucket-id bkt-alice-photos

List buckets:

s32p-ctl ... bucket ls

Replace a bucket ACL:

s32p-ctl ... bucket acl-set \
  --bucket-id bkt-team \
  --grant group:s3-team:read_write \
  --grant ak:alice_key_1:read_only

Import / Export

Import a directory YAML into the selected backend:

s32p-ctl ... import-yaml --yaml /path/to/directory.yaml --replace true
  • openbao: --replace purges the directory subtrees under the configured prefix before import.
  • yaml: --replace overwrites the destination file; --replace false merges by users.access_key and buckets.id.

Export backend state to a directory YAML:

s32p-ctl ... export-yaml --yaml /path/to/directory.yaml

Importing VersityGW IAM JSON

Import users from a VersityGW IAM JSON file (the accessAccounts map). Works against both backends, always merges into the existing directory (existing users with the same access key are overwritten; buckets/ACLs are not touched since VG IAM has no bucket concept).

s32p-ctl ... import-versity-iam --json /path/to/iam.json

Mapping: VG access → access_key, secret → secret_key, userID → uid, groupID → gid. The username field is resolved at runtime on the host running s32p-ctl via getpwuid_r(userID). The VG role field and any other unknown fields (e.g. projectID) are ignored.

Filtering (mutually exclusive, both repeatable, matched against the access key):

  • --include <ACCESS_KEY> — import only the listed users.
  • --exclude <ACCESS_KEY> — import everything except the listed users.

What to do when getpwuid_r(userID) returns no entry on the local host:

  • --on-missing-user error (default) — hard error, stop the import.
  • --on-missing-user use-access-key — use the VG access key string as the username.
  • --on-missing-user skip — log a warning and skip that entry.

Examples:

# OpenBao backend, only two users, fall back to access key for unknown uids
s32p-ctl --backend openbao ... \
  import-versity-iam \
    --json iam.json \
    --include alice --include bob \
    --on-missing-user use-access-key

# YAML backend, exclude a service account
s32p-ctl --backend yaml --yaml-path /etc/s32p/directory.yaml \
  import-versity-iam --json iam.json --exclude svc-test

Building

# Clone the repository with submodules
git clone --recurse-submodules https://github.com/r5r3/s32p-proxy.git
cd s32p-proxy

# If you already cloned without submodules, initialize them:
git submodule update --init --recursive

# Build the workspace
cargo build

Running (development)

Due to the usage of io_uring, you need RHEL 9.3, or another Linux distribution with a compatible kernel. On RHEL, it is necessary to enable the io_uring kernel module:

sysctl -w kernel.io_uring_disabled=0

You can set the log level either via environment variable or in the config file:

Using environment variable (traditional):

RUST_LOG=s32p_proxy=debug,pingora=info,pingora_proxy=info cargo run --bin s32p-proxy

Using config file (recommended): The log level can be configured in etc/s32p-proxy.yaml:

server:
  log_level: "s32p_proxy=debug,s32p_gateway=debug,pingora=info,pingora_proxy=info"

The config file approach automatically forwards the log level to worker processes.

The proxy binds to the address configured under server.listen in etc/s32p-proxy.yaml (a required field — there is no compiled-in default). The shipped dev config listens on:

http://0.0.0.0:9000

Workers are launched on-demand and bind to loopback (127.0.0.1:<port>) or to a Unix socket under a per-uid (and per-proxy-instance) run dir, one socket per worker profile.


Testing with MinIO Client (mcli)

mcli alias set S32P http://localhost:9000 TESTACCESSKEY123 TESTSECRETKEY456
mcli ls --debug S32P

Notes:

  • For the first request when no worker exists, the proxy validates SigV4 (auth header or presigned URL) before spawning.
  • VersityGW validates SigV4 again.

Security Notes

  • Run the proxy as root if you want the launcher to actually switch users (restricted-exec --user).
  • If the proxy is not root, workers start as the proxy’s user (development convenience).
  • Workers should never run as root in production.
  • Keep worker listeners loopback-only or use Unix domain sockets (recommended for local-only traffic).

ACL notes:

  • read_only vs read_write is computed by the proxy at worker spawn (via directory.buckets_for_access_key) and snapshotted into the worker's S32P_BUCKET_ACL environment variable as a name:rw|ro list. The worker rejects writes against read_only buckets with AccessDenied on every request. Buckets not in the snapshot (or an absent snapshot, e.g. when paired with an older proxy) default to read_write, preserving prior behavior.
  • The snapshot is captured at spawn time and stays frozen for the worker's lifetime. ACL changes in the directory take effect at the next worker spawn — operators can force a refresh by waiting for idle_timeout_secs to expire or by restarting the proxy.
  • Actual filesystem enforcement is still done by the kernel permissions of the target user and the bucket data path; the worker-side ACL check sits on top of that.

Known limitations / TODO

  • Versioning is not implemented; read-side probes (GetBucketVersioning) answer with AWS's "Unversioned" 200 shape via aws_compat so client state-detection works, but write-side ops (PutBucketVersioning, ListObjectVersions, anything with ?versionId=) fall back to 501 NotImplemented.
  • Object Lock is not implemented; reads answer with the matching AWS feature-disabled shape (404 ObjectLockConfigurationNotFoundError for the bucket config, 404 NoSuchObjectLockConfiguration for per-object retention/legal-hold), writes reject with 400 InvalidRequest "Bucket is missing Object Lock Configuration".
  • Server-side encryption (SSE-C, SSE-S3, SSE-KMS, SSE-KMS-DSSE) is rejected at the proxy with 400 InvalidRequest — no encryption path exists, and silently storing plaintext would mislead the client.
  • Multipart notes / current constraints:
    • CompleteMultipartUpload currently requires contiguous part numbers starting at 1.
    • Completion does not validate client-provided part ETags (the gateway uses a stable, upload-time ETag per part).
  • More production hardening:
    • rate limiting / max concurrent starts
    • negative caching for repeated invalid requests
    • structured metrics
  • OpenBao provisioning tooling (managing bucket docs + indices) not included here

License

Apache 2.0


Acknowledgements

  • Cloudflare Pingora
  • AWS SDK for Rust and S3 documentation
  • MinIO client for testing
  • VersityGW

About

s32p-proxy is a Rust-based S3-compatible proxy + gateway manager built on Cloudflare Pingora.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages