Skip to content

Latest commit

 

History

History
119 lines (100 loc) · 6.02 KB

File metadata and controls

119 lines (100 loc) · 6.02 KB

Development

Build & check

make all         # fmt + lint + test + build
make test        # cargo test --workspace
make lint        # cargo clippy --workspace --all-targets -- -D warnings
make fmt         # cargo fmt --all
make deny        # cargo deny check (advisories + licenses)
make release     # optimized gw-server binary (--locked)
make docker      # build the container image
make run         # cargo run -p gw-server

CI runs fmt/clippy/test (against live Postgres and Redis services, so the env-gated suites run), cargo deny, and the control-plane Go, web and Playwright gates on every push to main and every pull request. A v* tag cuts one GitHub release (.github/workflows/release.yml): native runners build gw for linux/darwin × amd64/arm64 and npm run build packs the web assets; goreleaser builds the control-plane binaries, attaches the gw and web-asset tarballs, and generates the changelog. The tag is the version — the release build stamps it into the workspace before compiling, so nothing needs bumping in Cargo.toml. Multi-arch container images for both components go to ghcr on the same tag (.github/workflows/docker.yml). Edition 2024; the workspace denies unwrap/expect/undocumented unsafe outside tests.

Workspace layout

Crates are strictly layered — lower layers never depend on higher ones:

server → {views, task} → handler → {dag, engines} → {models, state} → {protocol, config} → consts
Crate Role
consts error codes, the Protocol enum
models request/response types, typed params, usage, cost
protocol OpenAI/Anthropic wire types + cross-protocol conversions
config YAML config, provider presets, name indices
state auth, account pool, health, cache; Store and Governance seams
engines per-protocol engines behind the Transport seam, SSE, SigV4
dag the 4-layer request pipeline, nodes in declaration order, the admission token estimate
handler online/offline orchestration, DLP/blocklist plugins
task background tasks: quota reset, content purge, usage rollup, availability flush + alerts, alert dispatch
views axum HTTP/WebSocket handlers, streaming, metrics
server binary: wires config + state + transport, serves the router

Seams

Every boundary to the outside world is a trait with a deterministic default, so the whole pipeline runs offline in tests:

Trait Default Alternative
Transport dispatch (mock in-process, HTTP for real URLs) force mock / force HTTP
Store in-memory SQLite / Postgres (fleet)
Governance in-memory counters Redis
Moderator allow-all AWS Bedrock Guardrails (moderation:)
TokenEncoder tiktoken cl100k BPE heuristic fallback

Testing

Unit tests live beside their code; integration tests are in crates/*/tests/. Engine golden tests assert exact request wire shapes and response parsing against recorded fixtures. crates/server/tests/e2e.rs boots the full router in-process and exercises every surface offline. Tests that need real infrastructure gate on an env var (GW_TEST_REDIS_URL, GW_TEST_PG_URL) and no-op when it is unset; CI provisions both services so they run there. A release micro-benchmark lives in crates/server/tests/bench.rs:

cargo test --release -p gw-server --test bench -- --ignored --nocapture

It is a manual diagnostic rather than a CI merge gate; measured HTTP-level numbers and the load-test recipe are in Performance.

Live vendor matrix

scripts/live-matrix/ drives a running gateway against real vendors with real keys — the check the mock cannot make. live.yaml declares one account per vendor (keys come from the named env vars, never from the file) with prices and token_rate weights chosen so every billing dimension is visible; live_matrix.py runs, per provider group, non-streaming and streaming chat, /v1/messages (native and cross-protocol), thinking (budget/adaptive/effort dialects, signed replay through a tool loop), prompt cache (Anthropic breakpoints, automatic prefix caching on OpenAI/DeepSeek/Qwen), the response cache, embeddings, rerank, image and async video (submit, poll to done, one settle row). For every call it recomputes the weighted total and cost from the wire usage with the configured prices and compares them to the newest ledger row — an oracle independent of the gateway's own arithmetic — and asserts that reasoning content and cache reads/writes reached the client.

export OPENAI_API_KEY=... ANTHROPIC_API_KEY=...          # every api_key_env in live.yaml
GW_ADMIN_TOKEN=admin-live GW_CONFIG=scripts/live-matrix/live.yaml ./target/release/gw &
python3 scripts/live-matrix/live_matrix.py               # or: ... anthropic bedrock

Last full run (2026-09-17, Linux x86_64): 182/205, every ledger row matching the oracle. The groups cover Anthropic, OpenAI (incl. gpt-realtime-mini through /v1/realtime and sora-2 video), Gemini, DeepSeek, MiniMax (incl. Hailuo video), Qwen/DashScope (incl. wan2.2-t2v-plus video), Qianfan, Moonshot, SiliconFlow (incl. Wan2.2 video), OpenRouter, Cohere/Jina rerank, xAI Grok (chat, Responses, image, video; the grok group serves the same models from OpenRouter and Bedrock), Kling video, Brave search, Bedrock (InvokeModel, Converse, Llama), a local Ollama through the generic OpenAI-compatible path, OpenAI moderations/TTS/STT, the DashScope legacy wire and the Gemini Live realtime dialect.

None of the 23 failures were gateway defects: MiniMax was out of credits, the Kling and Brave keys were dead, that host has no Ollama and no websocket-client for the two realtime cases, and one transient Bedrock 503 on fable-5-1 latched the account unhealthy and cascaded into the nine later aws-anthropic/aws-converse cases — bedrock bedrock-jp grok xai re-run alone on a fresh gateway is 78/78. That cascade is worth knowing before reading a report: a single upstream 503 can mark a whole protocol unserved for the rest of the run.