make all # fmt + lint + test + build
make test # cargo test --workspace
make lint # cargo clippy --workspace --all-targets -- -D warnings
make fmt # cargo fmt --all
make deny # cargo deny check (advisories + licenses)
make release # optimized gw-server binary (--locked)
make docker # build the container image
make run # cargo run -p gw-serverCI runs fmt/clippy/test (against live Postgres and Redis services, so the env-gated suites run), cargo deny, and the control-plane Go, web and Playwright gates on every push to main and every pull request. A v* tag cuts one
GitHub release (.github/workflows/release.yml): native runners build gw
for linux/darwin × amd64/arm64 and npm run build packs the web assets;
goreleaser builds the control-plane binaries, attaches the gw and
web-asset tarballs, and generates the changelog.
The tag is the version — the release build stamps it into the workspace before
compiling, so nothing needs bumping in Cargo.toml. Multi-arch container
images for both components go to ghcr on the same tag
(.github/workflows/docker.yml). Edition 2024; the workspace denies
unwrap/expect/undocumented unsafe outside tests.
Crates are strictly layered — lower layers never depend on higher ones:
server → {views, task} → handler → {dag, engines} → {models, state} → {protocol, config} → consts
| Crate | Role |
|---|---|
consts |
error codes, the Protocol enum |
models |
request/response types, typed params, usage, cost |
protocol |
OpenAI/Anthropic wire types + cross-protocol conversions |
config |
YAML config, provider presets, name indices |
state |
auth, account pool, health, cache; Store and Governance seams |
engines |
per-protocol engines behind the Transport seam, SSE, SigV4 |
dag |
the 4-layer request pipeline, nodes in declaration order, the admission token estimate |
handler |
online/offline orchestration, DLP/blocklist plugins |
task |
background tasks: quota reset, content purge, usage rollup, availability flush + alerts, alert dispatch |
views |
axum HTTP/WebSocket handlers, streaming, metrics |
server |
binary: wires config + state + transport, serves the router |
Every boundary to the outside world is a trait with a deterministic default, so the whole pipeline runs offline in tests:
| Trait | Default | Alternative |
|---|---|---|
Transport |
dispatch (mock in-process, HTTP for real URLs) | force mock / force HTTP |
Store |
in-memory | SQLite / Postgres (fleet) |
Governance |
in-memory counters | Redis |
Moderator |
allow-all | AWS Bedrock Guardrails (moderation:) |
TokenEncoder |
tiktoken cl100k BPE | heuristic fallback |
Unit tests live beside their code; integration tests are in crates/*/tests/.
Engine golden tests assert exact request wire shapes and response parsing
against recorded fixtures. crates/server/tests/e2e.rs boots the full router
in-process and exercises every surface offline. Tests that need real
infrastructure gate on an env var (GW_TEST_REDIS_URL, GW_TEST_PG_URL) and
no-op when it is unset; CI provisions both services so they run there. A release micro-benchmark lives in crates/server/tests/bench.rs:
cargo test --release -p gw-server --test bench -- --ignored --nocaptureIt is a manual diagnostic rather than a CI merge gate; measured HTTP-level numbers and the load-test recipe are in Performance.
scripts/live-matrix/ drives a running gateway against real vendors with real
keys — the check the mock cannot make. live.yaml declares one account per
vendor (keys come from the named env vars, never from the file) with prices
and token_rate weights chosen so every billing dimension is visible;
live_matrix.py runs, per provider group, non-streaming and streaming chat,
/v1/messages (native and cross-protocol), thinking (budget/adaptive/effort
dialects, signed replay through a tool loop), prompt cache (Anthropic
breakpoints, automatic prefix caching on OpenAI/DeepSeek/Qwen), the response
cache, embeddings, rerank, image and async video (submit, poll to done, one
settle row). For every call it recomputes the weighted total
and cost from the wire usage with the configured prices and compares them
to the newest ledger row — an oracle independent of the gateway's own
arithmetic — and asserts that reasoning content and cache reads/writes reached
the client.
export OPENAI_API_KEY=... ANTHROPIC_API_KEY=... # every api_key_env in live.yaml
GW_ADMIN_TOKEN=admin-live GW_CONFIG=scripts/live-matrix/live.yaml ./target/release/gw &
python3 scripts/live-matrix/live_matrix.py # or: ... anthropic bedrockLast full run (2026-09-17, Linux x86_64): 182/205, every ledger row matching the
oracle. The groups cover Anthropic, OpenAI (incl. gpt-realtime-mini through
/v1/realtime and sora-2 video), Gemini, DeepSeek, MiniMax (incl. Hailuo
video), Qwen/DashScope (incl. wan2.2-t2v-plus video), Qianfan, Moonshot,
SiliconFlow (incl. Wan2.2 video), OpenRouter, Cohere/Jina rerank, xAI Grok
(chat, Responses, image, video; the grok group serves the same models from
OpenRouter and Bedrock), Kling video, Brave search, Bedrock (InvokeModel,
Converse, Llama), a local Ollama through the generic OpenAI-compatible path,
OpenAI moderations/TTS/STT, the DashScope legacy wire and the Gemini Live
realtime dialect.
None of the 23 failures were gateway defects: MiniMax was out of credits, the
Kling and Brave keys were dead, that host has no Ollama and no
websocket-client for the two realtime cases, and one transient Bedrock 503 on
fable-5-1 latched the account unhealthy and cascaded into the nine later
aws-anthropic/aws-converse cases — bedrock bedrock-jp grok xai re-run
alone on a fresh gateway is 78/78. That cascade is worth knowing before reading
a report: a single upstream 503 can mark a whole protocol unserved for the rest
of the run.