s32p-proxy is a Rust-based S3-compatible proxy + gateway manager built on Cloudflare Pingora.
This project is in an early development stage and is not ready for production use. All aspects of the implementation may change significantly in future versions. Use only for testing and development purposes.
It accepts S3 client requests, maps S3 identities (SigV4 access keys) to Unix users, and routes traffic to per-access-key worker processes (currently: VersityGW, or experimental alternative s32p-gateway) that expose a shared POSIX filesystem with kernel-enforced permissions.
A central design goal is to preserve Unix security semantics:
All filesystem access happens inside a process already running as the target Unix user.
The proxy itself does not perform filesystem I/O.
- Provide an S3-compatible endpoint backed by a POSIX filesystem
- Enforce per-user isolation using Unix UIDs/GIDs (kernel-level security)
- Support standard S3 clients (AWS CLI, MinIO client / SDKs)
- Start workers on demand, then reuse them (per access key + profile)
- Prevent easy DoS via worker spawn storms
- Keep proxy logic (auth/routing/management) separate from the storage worker
This repository is a Rust workspace with multiple crates:
s32p-proxy: the Pingora-based proxy binarys32p-directory: shared Directory API + YAML/OpenBao backends + shared YAML file formats32p-admin: management library (OpenBao write access, import/export, etc.)s32p-ctl: CLI wrapper arounds32p-admin(operator tooling)s32p-support: shared support library (common config/types/errors/helpers used across crates)s32p-gateway: the per-access-key worker gateway binary (launched by the proxy; it runs as alternative to VersityGW)
Outside the workspace:
tests-integration/: pixi-managed pytest harness. Spawns the real proxy + workers, drives them through a matrix of S3 clients (boto3 path-style + virtual-hosted, aws-cli, mount-s3 + rclone FUSE mounts) plus AWS-compat probes. Run withcd tests-integration && pixi run test. Layout + how to extend intests-integration/README.md.
┌──────────────┐
│ S3 Clients │ (aws cli, mcli, SDKs)
└──────┬───────┘
│ HTTP(S)
▼
┌───────────────────────────────────────────┐
│ s32p-proxy │
│ (Pingora HTTP proxy) │
│ │
│ • parses Authorization → access key │
│ • maps access key → unix user │
│ • looks up user + ACLs via Directory │
│ (YAML file or OpenBao) │
│ • classifies requests (query+path) │
│ • routes by "request class" via YAML │
│ (proxy to profile or local response) │
│ • gates worker spawn with SigV4 │
│ verification (auth header or presign) │
│ • reverse proxies to per-user worker │
│ • optional response header rewriting │
└───────────────────┬───────────────────────┘
│ internal HTTP (loopback) or
│ HTTP over Unix domain sockets (UDS)
▼
┌────────────────────────────────┐
│ Per-access-key workers │
│ (e.g. s32p-gateway │
│ or VersityGW) │
│ │
│ • runs as unix user │
│ • validates SigV4 again │
│ • uses a staged posix_root │
└───────────────┬────────────────┘
│ POSIX syscalls
▼
┌───────────────────────────────────────────┐
│ Shared POSIX FS (real bucket data) │
│ │
│ Each worker starts with a fresh temp │
│ directory as its posix_root. │
│ Symlinks to all accessible buckets │
│ are created in that temp root before │
│ the worker starts. The worker follows │
│ those links. │
└───────────────────────────────────────────┘
-
Pingora proxy-mode HTTP server
- Reverse proxies to locally spawned workers over loopback or Unix domain sockets (UDS)
- Preserves SigV4-critical headers (notably the original
Host) - Has a
response_filterhook for response header rewriting
-
Request classification (
crates/s32p-support/src/classifier.rs)- Shared by proxy + gateway (single source of truth for routing decisions)
- Parses path + query parameters (and selected headers where needed, e.g.
x-amz-copy-source) - Produces a high-level operation class key:
read(e.g.GetObject,HeadObject,ListObjectsV1/V2,ListBuckets,GetBucketLocation,HeadBucket,GetObjectAcl,GetBucketAcl)write(e.g.PutObject,CopyObject,DeleteObject,DeleteObjects,RenameObject,PutObjectAcl,PutBucketAcl)multipart(initiate/upload-part/list-parts/complete/abort + list uploads)versioning(detected; read-side probes answered AWS-shape viaaws_compat, write-side falls back to 501)object_lock(same shape as versioning — feature-disabled responses viaaws_compat)bucket_admin(CreateBucket/DeleteBucket; detected and routed to NotImplemented — bucket lifecycle is operator-only vias32p-ctl)session(S3 ExpressCreateSession)other
-
Config-driven routing (
etc/s32p-proxy.yaml)- Routes based on the classifier class keys above.
- Each class maps to one of four actions:
proxy(selects a worker profile)not_implemented(uniform 501 NotImplemented response, after SigV4 validation)aws_compat(per-S3OpAWS-shaped feature-disabled responses —200with empty<VersioningConfiguration/>forGetBucketVersioning,404 ObjectLockConfigurationNotFoundErrorforGetBucketObjectLockConfiguration,400 InvalidRequest "Bucket is missing Object Lock Configuration"forPutObjectRetention/PutObjectLegalHold, etc. — falling back to 501 for ops AWS always implements likePutBucketVersioning)create_session(local handling of S3 ExpressCreateSession)
- If a specific class key is not configured, the proxy falls back to the
otherroute. - Always-on pre-routing gate: any request carrying server-side-encryption headers (SSE-C / SSE-S3 / SSE-KMS / SSE-KMS-DSSE) is rejected with
400 InvalidRequestbefore routing — the gateway has no encryption path and silent plaintext storage would be a data-confidentiality bug. Sees3_compatibility_analysis.mdfor the per-variant table.
Experimental alternative to versitygw. Implements a growing subset of the S3 REST API directly on top of a POSIX filesystem.
- Read
GetObjectHeadObjectHeadBucketListBucketsGetBucketLocationListObjectsV1(GET /{bucket}, legacy form) andListObjectsV2(GET /{bucket}?list-type=2)
- Write
PutObject(streaming upload). Append is supported via the S3 Express extensionx-amz-write-offset-bytes: N: the request body is written at byte offsetNof the existing object;Nmust equal the current size (no overwrite, no holes), otherwise412 PreconditionFailed.CopyObject(server-side copy; size-limited by configuration)DeleteObjectDeleteObjects(POST /?delete)RenameObject(PUT ?renameObjectwithx-amz-rename-source; same-bucket only). Idempotent: re-running an already-applied rename returns success.
- Multipart
CreateMultipartUpload(POST ?uploads)UploadPart(PUT ?partNumber=N&uploadId=...)UploadPartCopy(PUT ?partNumber=N&uploadId=...withx-amz-copy-source; optionalx-amz-copy-source-rangeand the fourx-amz-copy-source-if-*headers). When the first part of an upload is a whole-sourceUploadPartCopy(nocopy-source-range), the source's size is recorded as atotal_size_hinton the upload and used to size the assembly file's Lustre stripe count.ListParts(GET ?uploadId=...)ListMultipartUploads(GET /bucket?uploads)CompleteMultipartUpload(POST ?uploadId=...)AbortMultipartUpload(DELETE ?uploadId=...)
- Directory bucket (S3 Express) handshake
CreateSession(GET /{bucket}?session) — issues short-lived session credentials and ax-amz-s3session-token. Subsequent data-plane requests sign withservice=s3expressand pass the session token either as thex-amz-s3session-tokenheader or, for presigned URLs, as the equivalent query parameter. The proxy validates both paths against an in-memory session store; workers re-validate on every request.
- ACL
GetObjectAcl,GetBucketAcl(GET ?acl)PutObjectAcl,PutBucketAcl(PUT ?acl) — accepted as a no-op when the requested ACL matches the current POSIX state; mismatches are rejected. The directory ACLs ins32p-ctlremain authoritative.
- Object tagging
GetObjectTagging,PutObjectTagging,DeleteObjectTagging(GET/PUT/DELETE /{bucket}/{key}?tagging)x-amz-taggingheader onPutObject(tags-at-creation) and onCopyObjectwithx-amz-tagging-directive: COPY|REPLACE(defaultCOPYmirrors source tags).- Storage model. Tags live in the
user.s32p.tagsxattr on the object file, encoded URL-form (team=a&stage=raw) — the same shape as the wire header. A POSIX-created file has no xattr → emptyTagSet(still a valid S3 object).cp -a/rsync -Xpreserve tags across POSIX copies; plaincpdrops them, which is equivalent to creating a fresh tagless object — acceptable degradation. - Limits enforced: ≤ 10 pairs, key 1..=128 chars, value 0..=256 chars, total payload ≤ 2 KB, AWS-documented charset (
[A-Za-z0-9 +-=._:/@]), no duplicate keys. - Bucket-level tagging (
?taggingon a bucket) is not implemented; the proxy continues to route it to theotherworker (versitygw) per the default config.
- Object metadata
x-amz-meta-*headers onPutObject(round-tripped onHeadObject/GetObject).Content-TypeonPutObject(round-tripped onHeadObject/GetObject).CopyObjectwithx-amz-metadata-directive: COPY|REPLACE— defaultCOPYmirrors source user metadata and Content-Type;REPLACEuses the request headers (clears when none are supplied).- Storage model. User metadata lives in
user.s32p.meta(URL-form, lowercased keys); explicit Content-Type lives inuser.s32p.content_type. The gateway also honors the freedesktopuser.mime_typexattr read-only as a Content-Type fallback (GUI file managers andgio setpopulate this; the gateway never writes it). A POSIX-created file has no xattrs → empty metadata, Content-Type resolved by the fallback ladder below. - Content-Type resolution ladder (HEAD/GET):
user.s32p.content_type→user.mime_type→ extension map viamime_guess→application/octet-stream. - Limits enforced: user metadata payload ≤ 2 KB total, ≤ 32 pairs, key length 1..=128 chars (HTTP-token shape
[A-Za-z0-9_-]), value ≤ 1024 chars, no duplicate keys. Content-Type ≤ 256 bytes and must parse astype/subtype[;params]. - Not stored (deferred): the other five HTTP passthrough headers (
Content-Encoding,Content-Language,Content-Disposition,Cache-Control,Expires), and metadata set onCreateMultipartUpload(multipart-completed objects carry no metadata even if the create-call supplied headers).
- Supports
Range: bytes=...(returns206 Partial Content; invalid ranges return416 InvalidRange). A single range returns the bytes directly; multiple ranges return amultipart/byterangesbody (RFC 7233), one part per range — capped at 50 ranges, and any unsatisfiable range fails the whole request with416. A multi-rangeHEADignores theRangeand returns the full object (200). - Rejects most query parameters for now, except those required for:
?location,?list-type=2,?delete,?acl,?renameObject,?session, and the multipart query parameters (?uploads,?uploadId=...,?partNumber=...)
- SigV4 presigned URL query parameters (
X-Amz-*) are supported and do not count as "effective" query parameters for routing/handling. A small fixed set of other keys is also treated as non-effective:x-id,content-type,cache-control,content-encoding,content-disposition,expires,x-amz-storage-class. ListObjectsV2supports Lustre Lazy Size on MDS (LSOM) when built with the Lustre feature.- When built with
--features lustre, the gateway creates new files with Lustre striping viallapi_file_create().- Config:
S32P_LUSTRE_MAX_STRIPE_COUNT(default:4) caps the stripe count. - Serial uploads (
PutObject,CopyObject, and any temp/staging files):stripe_size = S32P_CHUNK_SIZE_MB × 1 MiB(env var is in MiB; default4→ 4 MiB)stripe_count = ceil(file_size / stripe_size), capped byS32P_LUSTRE_MAX_STRIPE_COUNT
- Multipart uploads:
direct.bin:stripe_size = min(stripe_size_serial, part_size).stripe_countis sized from whatever total-size estimate is available — the upload'stotal_size_hint(set when the first part is a whole-sourceUploadPartCopy) or, inCompleteMultipartUpload's fast path, the now-known final size — viaceil(estimate / stripe_size)and capped byS32P_LUSTRE_MAX_STRIPE_COUNT. With no estimate, the count falls back to the configured max. Striping is fixed at file creation; the hint matters only the first timedirect.binis opened.- individual part files are striped like serial uploads.
- Config:
ETagfor final objects is generated from the inode number.- Object-level metadata is stored in
user.*xattrs. Tagging, user metadata (x-amz-meta-*), andContent-Typeeach have a dedicated xattr (user.s32p.tags,user.s32p.meta,user.s32p.content_type). The POSIX interop contract still holds — absence of any auxiliary state yields a valid object (POSIX-created file = empty metadata, no error).cp -a/rsync -Xpreserve every category; plaincpdrops them and the destination becomes a metadata-less-but-valid object. Not stored: the other five HTTP passthrough headers (Content-Encoding,Content-Language,Content-Disposition,Cache-Control,Expires) and any metadata set onCreateMultipartUpload— both deferred until a real client need surfaces. Same principle asPutObjectAclbeing a no-op when the requested ACL matches current POSIX state.
Multipart uploads are implemented in crates/s32p-gateway/src/multipart.rs.
The design goal is to avoid creating a full extra copy of the object on the server during completion. The gateway does this by writing parts into an assembly file that can be renamed into place as the final object.
For a bucket with root <bucket_root>, the gateway reserves a hidden directory (default name: .s32p-mpu):
<bucket_root>/.s32p-mpu/uploads/<upload_id>/meta.json
JSON metadata for the upload (bucket/key, state, and a map of uploaded parts).<bucket_root>/.s32p-mpu/uploads/<upload_id>/lock
A file used withflock(LOCK_EX)to serialize metadata updates and completion.<bucket_root>/.s32p-mpu/uploads/<upload_id>/direct.bin
The assembly file (random-access writes at part offsets).<bucket_root>/.s32p-mpu/uploads/<upload_id>/parts/part-00001.bin(etc.)
Fallback storage for parts that cannot safely be placed intodirect.bin.
The gateway also prevents clients from reading/writing/deleting objects inside the reserved multipart directory by treating that first path segment as “reserved”.
When a part arrives (PUT ?partNumber=N&uploadId=...), the gateway decides where to store it:
-
It reads and updates
meta.jsonunder an exclusive lock (short critical section). -
It tries to determine a stable assumed part size:
- If part #1 is uploaded and no size is known yet, its length becomes the assumed part size.
- If no assumed size exists and the gateway sees two parts with the same size, it promotes that size to the assumed part size.
-
If an assumed part size is known and the part is not larger than it, the gateway computes the direct placement offset:
offset = (partNumber - 1) * assumed_part_sizeand writes the request body directly into
direct.binat that offset (random-access file write). -
Otherwise, it writes the part as an individual file under
parts/and records that in metadata.
Metadata records, per part:
- size
- an upload-time ETag (stable, client-visible; completion does not validate ETags)
- timestamp
- storage kind:
{ direct: off }or{ file: name }
On CompleteMultipartUpload (POST ?uploadId=...), the gateway:
- Reads the XML body (requested part numbers).
- Takes an exclusive lock for the entire completion to prevent concurrent
UploadPartand to make the final rename deterministic. - Validates the request:
- Parts must be contiguous starting at 1 (
1..=lastPart). - Every requested part must exist in
meta.json.
- Parts must be contiguous starting at 1 (
Then it chooses one of two assembly paths:
This path is taken when the upload matches the classic “fixed-size parts + last part shorter or equal” layout:
assumed_part_sizeis known- parts
1..last-1are exactlyassumed_part_size - any parts already stored as
directare at the expected offsets
Fast-path algorithm:
-
Compute final object size:
final_size = (lastPart - 1) * assumed_part_size + size(lastPart) -
Ensure all missing parts are present in
direct.bin:- For parts stored as separate files, copy them into
direct.binat their final offsets (range copy).
- For parts stored as separate files, copy them into
-
ftruncate(direct.bin, final_size) -
Rename
direct.bin→<final_object_path>
If the rename fails with cross-device (EXDEV), the gateway performs a single copy of the already-assembled direct.bin into a temp file next to the destination and renames that temp file into place. Importantly, it still avoids “assemble → copy again” behavior.
✅ Why this avoids a full copy: in the common case (same filesystem), completion becomes a metadata operation (rename) after assembling into direct.bin. There is no “write whole object into a second file” step.
If the fixed-size/offset conditions don’t hold, the gateway assembles sequentially:
- Create a staging output file:
- If the upload directory is on a different filesystem than the destination, the staging file is created next to the destination (to avoid an extra copy later).
- Copy parts in order into the staging file:
- For
directparts: copy the required range out ofdirect.bin. - For
fileparts: copy the entire part file.
- For
- Rename staging file → final object path.
Finally, on success the gateway marks the upload as completed, returns the completion XML + ETag, and removes the upload directory.
- The gateway’s “happy path” is optimized for large objects: parts are placed directly into their final offsets, and completion is typically a truncate + rename.
- The implementation is careful to avoid data corruption:
- Direct placement is only used when offsets can be computed safely (no overlap risk).
- Metadata updates are protected by
flockand written atomically (meta.json.tmp→ rename).
A “Directory” provides:
user_by_access_key(access_key)→ unix identity + secret keybuckets_for_access_key(access_key)→ list of visible buckets with effective access
Backends:
- YAML: load once at startup into HashMaps
- OpenBao: AppRole login + KV v2 reads, using indices for efficient lookup
Notes:
- Multiple access keys may map to the same Unix user (same
username/uid/gid). - Workers are keyed by (access_key, worker_profile) (not just uid), because the staged bucket set is per access key.
Buckets have an ACL list. Each ACL entry grants an access level to a principal:
access_keyprincipal (direct grant)group_nameprincipal (POSIX group name grant)
Effective access for a caller is the max of all matching entries (e.g. read_write beats read_only).
Group membership is resolved from the OS at runtime (username → gids → group names).
- Shared S3 REST-XML helpers live in
crates/s32p-support/src/s3resp.rsandcrates/s32p-support/src/s3xml.rs(errors likeAccessDenied,SignatureDoesNotMatch,NotImplemented, plus body builders); proxy-side glue (respond_bytes()and friends) is incrates/s32p-proxy/src/responses.rs.
-
DoS mitigation via “spawn gating”
- If a worker is not running, the proxy performs SigV4 verification without reading the body:
- Standard SigV4 auth header (
Authorization,x-amz-content-sha256) - SigV4 presigned URLs (query signature; typically
UNSIGNED-PAYLOAD)
- Standard SigV4 auth header (
- Only if the signature is valid will the proxy start the worker
- Once a worker is already running, the proxy does not fully validate SigV4; it only extracts the access key for routing and forwards the request to the worker
- If a worker is not running, the proxy performs SigV4 verification without reading the body:
-
Workers (VersityGW or s32p-gateway) still validate SigV4 again (cannot be disabled).
-
Clock-skew gate (header SigV4 only). Header-signed requests whose
x-amz-datediffers from the server clock by more than 15 minutes (in either direction) are rejected withRequestTimeTooSkewed(HTTP 403). This matches AWS S3's documented window and exists so a captured signed request can't be replayed indefinitely. Presigned URLs have their ownX-Amz-Expires-based check (rejected withAccessDenied/ "Request has expired") and don't use this window. The 15-minute value is a constant ins32p-support(HEADER_SIGV4_MAX_SKEW).
- Workers are keyed by (access_key, worker_profile)
- supports multiple access keys mapped to the same unix user but different bucket ACLs
- allows routing different command classes to different worker profiles per access key
- Workers are started via a configurable launcher (default:
restricted-exec)- If running as root and
pass_user_flag_if_root=true, the proxy passes--user <username> - When
workers.launcher.landlockis enabled (defaulttrue), the proxy also passes--rw <posix_root>, then per accessible bucket either--rw <bucket.data_path>(forread_writegrants) or--ro <bucket.data_path>(forread_onlygrants), plus--resolve-libsand--allow-nss. Landlock therefore enforces the same ACL access level the gateway already checks (see §Security Notes), at the kernel layer.
- If running as root and
- Upstream bind (per worker):
- TCP loopback:
127.0.0.1:<port> - Unix domain socket:
<uds_run_dir>/<uid>-<instance_id>/<profile>.sock<uds_run_dir>defaults to/run/s32p(root) or$XDG_RUNTIME_DIR/s32p(non-root); override per profile viaworkers.profiles.<name>.upstream.uds_run_dir.<instance_id>is a per-worker random suffix. It disambiguates concurrent proxy processes on the same host and disambiguates workers that share(uid, profile)within a single proxy — which happens when two access keys map to the same uid (a documented design point).- One socket file per worker profile under the per-worker run dir.
- TCP loopback:
- Readiness probing: connect loop until port is reachable
- Idle shutdown after
idle_timeout_secs - Sweeper removes dead/idle workers periodically (
sweep_interval_secs) - Staged posix_root
- On each worker start, the proxy creates a fresh temp directory under
workers.posix_root - Creates symlinks for all buckets visible to the access key:
<temp>/<bucket_name>→<bucket.data_path>
- Passes that temp directory as
{{posix_root}}to the worker - When the worker stops, the temp directory is removed
- On each worker start, the proxy creates a fresh temp directory under
Worker args / env templates are expanded per spawn by worker_manager.rs::render_template. Unknown tokens cause a hard error. Available tokens:
| Token | Value |
|---|---|
{{username}} |
Unix username of the target user |
{{uid}} |
numeric UID |
{{gid}} |
numeric GID |
{{access_key}} |
the SigV4 access key for this worker |
{{secret_key}} |
matching secret key |
{{posix_root}} |
the staged temp dir with bucket symlinks |
{{bind_addr}} |
endpoint string (host:port for TCP, path for UDS) |
{{bind_uds}} |
alias of {{bind_addr}} (use whichever reads better in your config) |
{{region}} |
from server.region |
{{virtual_hosted_suffixes}} |
comma-joined server.virtual_hosted_suffixes |
{{log_level}} |
from server.log_level |
Separately, {{install_bin_dir}} is a config-load-time placeholder resolved only inside workers.launcher.path and workers.profiles.*.exec (not in args/env). It expands to the directory containing the running s32p-proxy executable (e.g. target/debug in dev, /usr/local/bin after install).
Primary configuration: etc/s32p-proxy.yaml
Key sections:
server.listen/server.public_scheme(both required)server.region(required; used as the SigV4 region and exposed to workers via{{region}})server.log_level(logging configuration, supports RUST_LOG format)server.virtual_hosted_suffixes(virtual-hosted-style bucket detection)server.tls_cert_path/server.tls_key_path(optional; enables HTTPS on the listen address)server.shutdown_grace_period_secs(graceful-shutdown timeout, default 10s)auth.*(directory backend selection and credentials)workers.posix_root(base dir for per-worker temp roots)workers.launcher.*(includespath,pass_user_flag_if_root,landlock)workers.lifecycle.*(idle_timeout_secs,sweep_interval_secs)workers.profiles.*(worker templates:exec,args,env,upstream)routing.class_map.*(routes classifier classes to actions)
The proxy automatically detects and supports both path-style and virtual-hosted-style bucket addressing:
- Path-style:
https://s3.example.com/bucket/key(bucket in path) - Virtual-hosted-style:
https://bucket.s3.example.com/key(bucket in host)
Configure domain suffixes for virtual-hosted-style detection via server.virtual_hosted_suffixes:
server:
virtual_hosted_suffixes:
- "127.0.0.1.nip.io"
- "s3.example.com"
- "s3.localhost"When a request's Host header ends with a configured suffix, the proxy extracts the bucket name from the host and the key from the path. Port numbers are handled correctly (e.g., bucket.suffix:9000).
Workers receive the suffixes via the S32P_VIRTUAL_HOSTED_SUFFIXES environment variable or {{virtual_hosted_suffixes}} template:
workers:
profiles:
s32p-gateway:
env:
S32P_VIRTUAL_HOSTED_SUFFIXES: "{{virtual_hosted_suffixes}}"Choose a directory backend via auth.backend:
yaml— load a local directory fileopenbao— use OpenBao (Vault-compatible API) with AppRole + KV v2 + indices
auth:
backend: "yaml"
yaml:
path: "/etc/s32p/directory.yaml"auth:
backend: "openbao"
openbao:
address: "http://127.0.0.1:8200"
approle_mount: "approle"
role_id_file: "/etc/s32p/role_id"
secret_id_file: "/etc/s32p/secret_id"
kv_mount: "secret"
prefix: "s32p"The proxy can hold a small in-process TTL cache in front of the directory backend. The cache decorates any backend (so YAML and OpenBao share the same code path) and only matters in practice for OpenBao — every uncached request would otherwise hit Vault.
auth:
cache:
enabled: true # optional; default depends on backend (see below)
user_ttl_secs: 30 # positive `user_by_access_key` results
buckets_ttl_secs: 30 # positive `buckets_for_access_key` results
negative_ttl_secs: 5 # cached "no such access key"
max_entries: 4096 # soft cap per map; over cap, prune expired then evict soonest-expiring- Defaults: cache is on by default for the OpenBao backend and off by default for the YAML backend (YAML is an in-memory
HashMaplookup; the cache adds no benefit, only lock overhead). Settingenabled: trueorenabled: falseoverrides either default. - Errors are never cached — backend errors (network failures, KV decode errors) always pass through so the next attempt sees fresh state.
- Snapshot vs. cache. ACL state is independently snapshotted into spawned workers via
S32P_BUCKET_ACL(see §Security Notes); that snapshot freezes for the worker's lifetime. The cache only changes how fresh the next worker spawn's snapshot is — it does not propagate ACL changes to running workers. ACL changes therefore take effect atmax(buckets_ttl_secs, idle_timeout_secs)worst-case. - Single-flight deduplication. Under a cold cache, N concurrent first-requests for the same access key all see the miss but only one actually hits the backend; the rest wait on a per-key
Arc<Mutex<()>>and pick up the now-cached result. This bounds backend fan-out by the number of distinct access keys, not by request concurrency. - Observability. Hits log at
traceand misses + evictions atdebugunder thes32p_directory::cachetarget. To see misses without spamming on hits:Each event carries theRUST_LOG=info,s32p_directory::cache=debug cargo run --bin s32p-proxy
access_key, the affected map (users/buckets), and the TTL applied — useful for confirming the cache is actually warm in production and for tuning TTLs.
Example /etc/s32p/directory.yaml:
version: 1
users:
- access_key: "alice_key_1"
secret_key: "alice_secret"
username: "alice"
uid: 1001
gid: 1001
# Second access key mapping to the same unix user:
- access_key: "alice_key_2"
secret_key: "alice_secret_2"
username: "alice"
uid: 1001
gid: 1001
buckets:
- id: "bkt-alice-photos"
name: "photos"
data_path: "/srv/s3/alice/photos"
acl:
- principal: { type: "access_key", access_key: "alice_key_1" }
access: "read_write"
- principal: { type: "access_key", access_key: "alice_key_2" }
access: "read_write"
- id: "bkt-team"
name: "team"
data_path: "/srv/s3/team"
acl:
- principal: { type: "group_name", name: "s3-team" }
access: "read_write"
- principal: { type: "access_key", access_key: "alice_key_1" }
access: "read_only"Under auth.openbao.kv_mount + auth.openbao.prefix, the directory uses KV v2 documents:
prefix/users/<access_key>→UserDocprefix/buckets/<bucket_id>→BucketDocprefix/index/access_key/<access_key>→{ bucket_ids: [...] }prefix/index/group/<group_name>→{ bucket_ids: [...] }
The indices make buckets_for_access_key() efficient:
- read index for the access key
- resolve group names for the unix user
- read index for each group name
- load bucket docs and evaluate ACLs (authoritative)
s32p-ctl manages the Directory state (users, buckets, ACLs) for both supported backends:
--backend openbao(default): OpenBao/Vault KV v2 + indices--backend yaml: local directory.yaml file
All commands work with both backends. The backend is selected via --backend.
For YAML, you must provide the directory file path:
s32p-ctl --backend yaml --yaml-path /etc/s32p/directory.yaml <COMMAND...>For OpenBao, provide --address (or VAULT_ADDR) plus authentication:
- Setup typically uses a root/admin token (
--tokenorVAULT_TOKEN) - Normal operations typically use AppRole (
--role-id(-file)+--secret-id(-file))
Common OpenBao flags:
--address http://127.0.0.1:8200(orVAULT_ADDR)--kv-mount secret(default:secret)--prefix s32p(default:s32p)--approle-mount approle(default:approle)--token ...(orVAULT_TOKEN)--role-id .../--role-id-file ...--secret-id .../--secret-id-file ...
Example:
s32p-ctl --backend openbao \
--address http://127.0.0.1:8200 \
--kv-mount secret \
--prefix s32p \
--approle-mount approle \
--role-id-file /etc/s32p/admin_role_id \
--secret-id-file /etc/s32p/admin_secret_id \
<COMMAND...>Setup is idempotent and verifying — re-runs are safe and abort hard rather than overwrite incompatible state. It creates/updates:
- mounts KV v2 at
--kv-mount(default:secret) if missing; if a mount already exists at that path, it is verified to betype=kvwithoptions.version=2and reused — any other engine type or KV v1 aborts setup - enables AppRole auth method at
--approle-mount(default:approle) - creates policies + roles, scoped to
<kv-mount>/data/<prefix>/*and<kv-mount>/metadata/<prefix>/*:s32p-proxy(read-only)s32p-admin(read-write)
- reads the role IDs and generates a fresh secret ID per role; if the corresponding
--*-id-fileflags are given, the four IDs are written there with mode 0600 (Unix)
The bootstrap token (--token or VAULT_TOKEN) must be allowed to write sys/mounts/<kv-mount>, sys/auth/<approle-mount>, sys/policies/acl/* and auth/<approle-mount>/role/*. A root token covers all of these; restricted bootstrap tokens need the corresponding capabilities.
To keep the token out of your shell history and out of ps//proc/<pid>/cmdline, prompt for it without echo:
export VAULT_ADDR=http://127.0.0.1:8200
read -rs VAULT_TOKEN && export VAULT_TOKEN # paste token, hit enter; no echo
s32p-ctl --backend openbao setup \
--proxy-role-id-file /etc/s32p/proxy_role_id \
--proxy-secret-id-file /etc/s32p/proxy_secret_id \
--admin-role-id-file /etc/s32p/admin_role_id \
--admin-secret-id-file /etc/s32p/admin_secret_id
unset VAULT_TOKEN # token is no longer neededAfter setup the bootstrap token is not needed for normal operation:
- the proxy authenticates via the
s32p-proxyAppRole (role_id_file+secret_id_fileinetc/s32p-proxy.yamlunderauth.openbao) - further
s32p-ctlcommands (user add,bucket add,import-yaml, …) authenticate via thes32p-adminAppRole (--role-id-file+--secret-id-file)
Re-running setup generates an additional secret_id for each role; existing secret IDs stay valid until destroyed via auth/<approle-mount>/role/<role>/secret-id/destroy. Rotate explicitly if you need the old ones revoked.
Creates an empty directory file skeleton:
s32p-ctl --backend yaml --yaml-path /etc/s32p/directory.yaml setupCommands:
user adduser rmuser ls
Add a user. Only --username is required; everything else is derived from it:
s32p-ctl ... user add --username aliceok
user alice uid=1001 gid=1001 access_key=alice
secret_key=xXDk-bbkR7GweeyJM_mfMjkSJ0UG-BLK79PItgebzDA
--usernameis lower-cased (POSIX user names are lower-case by convention, andgetpwnam_rmatches exactly)--access-keydefaults to that lower-cased user name--uid/--giddefault to the passwd entry of that user name; without one, both must be given explicitly--secret-keydefaults to 32 CSPRNG bytes as URL-safe base64, printed only at creation — it cannot be recovered afterwards, only replaced by re-runninguser add
Any of them can still be given explicitly, e.g. for an account this host's
passwd database doesn't know, or for an access key that differs from the user
name. An explicit --access-key is stored verbatim, case included — SigV4
compares it byte-for-byte against what the client sends, and AWS-style keys are
upper-case:
s32p-ctl ... user add \
--access-key alice_key_1 \
--secret-key alice_secret \
--username alice \
--uid 1001 \
--gid 1001Remove a user (optionally scrubs their access_key from bucket ACLs):
s32p-ctl ... user rm --access-key alice_key_1 --cleanup-acls trueList users:
s32p-ctl ... user lsCommands:
bucket addbucket rmbucket lsbucket acl-set
Add a bucket:
s32p-ctl ... bucket add \
--name photos \
--data-path /srv/s3/alice/photos \
--grant ak:alice_key_1:read_write--grant is repeatable and supports:
ak:<ACCESS_KEY>:read_only|read_writegroup:<GROUP_NAME>:read_only|read_write
Optional: provide a stable bucket id (otherwise a UUID is generated):
s32p-ctl ... bucket add \
--bucket-id bkt-alice-photos \
--name photos \
--data-path /srv/s3/alice/photos \
--grant ak:alice_key_1:read_writeRemove a bucket:
s32p-ctl ... bucket rm --bucket-id bkt-alice-photosList buckets:
s32p-ctl ... bucket lsReplace a bucket ACL:
s32p-ctl ... bucket acl-set \
--bucket-id bkt-team \
--grant group:s3-team:read_write \
--grant ak:alice_key_1:read_onlyImport a directory YAML into the selected backend:
s32p-ctl ... import-yaml --yaml /path/to/directory.yaml --replace true- openbao:
--replacepurges the directory subtrees under the configured prefix before import. - yaml:
--replaceoverwrites the destination file;--replace falsemerges byusers.access_keyandbuckets.id.
Export backend state to a directory YAML:
s32p-ctl ... export-yaml --yaml /path/to/directory.yamlImport users from a VersityGW IAM JSON file (the accessAccounts map). Works
against both backends, always merges into the existing directory (existing
users with the same access key are overwritten; buckets/ACLs are not touched
since VG IAM has no bucket concept).
s32p-ctl ... import-versity-iam --json /path/to/iam.jsonMapping: VG access → access_key, secret → secret_key, userID → uid,
groupID → gid. The username field is resolved at runtime on the host
running s32p-ctl via getpwuid_r(userID). The VG role field and any other
unknown fields (e.g. projectID) are ignored.
Filtering (mutually exclusive, both repeatable, matched against the access key):
--include <ACCESS_KEY>— import only the listed users.--exclude <ACCESS_KEY>— import everything except the listed users.
What to do when getpwuid_r(userID) returns no entry on the local host:
--on-missing-user error(default) — hard error, stop the import.--on-missing-user use-access-key— use the VG access key string as the username.--on-missing-user skip— log a warning and skip that entry.
Examples:
# OpenBao backend, only two users, fall back to access key for unknown uids
s32p-ctl --backend openbao ... \
import-versity-iam \
--json iam.json \
--include alice --include bob \
--on-missing-user use-access-key
# YAML backend, exclude a service account
s32p-ctl --backend yaml --yaml-path /etc/s32p/directory.yaml \
import-versity-iam --json iam.json --exclude svc-test# Clone the repository with submodules
git clone --recurse-submodules https://github.com/r5r3/s32p-proxy.git
cd s32p-proxy
# If you already cloned without submodules, initialize them:
git submodule update --init --recursive
# Build the workspace
cargo buildDue to the usage of io_uring,
you need RHEL 9.3, or another Linux distribution with a compatible kernel.
On RHEL, it is necessary to enable the io_uring kernel module:
sysctl -w kernel.io_uring_disabled=0You can set the log level either via environment variable or in the config file:
Using environment variable (traditional):
RUST_LOG=s32p_proxy=debug,pingora=info,pingora_proxy=info cargo run --bin s32p-proxyUsing config file (recommended):
The log level can be configured in etc/s32p-proxy.yaml:
server:
log_level: "s32p_proxy=debug,s32p_gateway=debug,pingora=info,pingora_proxy=info"The config file approach automatically forwards the log level to worker processes.
The proxy binds to the address configured under server.listen in etc/s32p-proxy.yaml (a required field — there is no compiled-in default). The shipped dev config listens on:
http://0.0.0.0:9000
Workers are launched on-demand and bind to loopback (127.0.0.1:<port>) or to a Unix socket under a per-uid (and per-proxy-instance) run dir, one socket per worker profile.
mcli alias set S32P http://localhost:9000 TESTACCESSKEY123 TESTSECRETKEY456
mcli ls --debug S32PNotes:
- For the first request when no worker exists, the proxy validates SigV4 (auth header or presigned URL) before spawning.
- VersityGW validates SigV4 again.
- Run the proxy as root if you want the launcher to actually switch users (
restricted-exec --user). - If the proxy is not root, workers start as the proxy’s user (development convenience).
- Workers should never run as root in production.
- Keep worker listeners loopback-only or use Unix domain sockets (recommended for local-only traffic).
ACL notes:
read_onlyvsread_writeis computed by the proxy at worker spawn (viadirectory.buckets_for_access_key) and snapshotted into the worker'sS32P_BUCKET_ACLenvironment variable as aname:rw|rolist. The worker rejects writes againstread_onlybuckets withAccessDeniedon every request. Buckets not in the snapshot (or an absent snapshot, e.g. when paired with an older proxy) default toread_write, preserving prior behavior.- The snapshot is captured at spawn time and stays frozen for the worker's lifetime. ACL changes in the directory take effect at the next worker spawn — operators can force a refresh by waiting for
idle_timeout_secsto expire or by restarting the proxy. - Actual filesystem enforcement is still done by the kernel permissions of the target user and the bucket data path; the worker-side ACL check sits on top of that.
- Versioning is not implemented; read-side probes (
GetBucketVersioning) answer with AWS's "Unversioned" 200 shape viaaws_compatso client state-detection works, but write-side ops (PutBucketVersioning,ListObjectVersions, anything with?versionId=) fall back to501 NotImplemented. - Object Lock is not implemented; reads answer with the matching AWS feature-disabled shape (
404 ObjectLockConfigurationNotFoundErrorfor the bucket config,404 NoSuchObjectLockConfigurationfor per-object retention/legal-hold), writes reject with400 InvalidRequest "Bucket is missing Object Lock Configuration". - Server-side encryption (SSE-C, SSE-S3, SSE-KMS, SSE-KMS-DSSE) is rejected at the proxy with
400 InvalidRequest— no encryption path exists, and silently storing plaintext would mislead the client. - Multipart notes / current constraints:
CompleteMultipartUploadcurrently requires contiguous part numbers starting at 1.- Completion does not validate client-provided part ETags (the gateway uses a stable, upload-time ETag per part).
- More production hardening:
- rate limiting / max concurrent starts
- negative caching for repeated invalid requests
- structured metrics
- OpenBao provisioning tooling (managing bucket docs + indices) not included here
Apache 2.0
- Cloudflare Pingora
- AWS SDK for Rust and S3 documentation
- MinIO client for testing
- VersityGW