Repository navigation
feat(ai-rate-limiting): multi-window per-consumer quota limiting on different models #13225
SylvainVerdy
started this conversation in
Ideas
Replies: 1 comment
Finally, I implemented a custom Apisix plugin to do it myself. If you're interested, feel free to contact me. -- Custom APISIX plugin "multi-window-rate-limiter".
--
-- Multi-window rate limiting (hourly/daily/weekly) using FIXED WINDOW.
--
-- Storage: ngx.shared.DICT (shared dict "mwrl_counters") for
-- atomic increments via dict:incr(). Replaces the etcd v0.5 storage
-- which suffered from:
-- 1. Silent write failures (etcd_set only logged a warning)
-- 2. Race conditions (non-atomic read-modify-write under concurrency,
-- causing systematic undercounting of tokens)
-- The result was that the developer-portal (Loki) showed a quota
-- exceeded while the plugin (etcd) never blocked requests.
--
-- ngx.shared.DICT solves both problems:
-- - dict:incr() is atomic (no race condition)
-- - No network calls (no etcd failures)
-- - Auto-expiry via TTL (no manual cleanup)
-- Works with 1 APISIX replica (replicaCount: 1).
-- Counters are lost on pod restart, which is acceptable
-- (fixed window = periodic reset anyway).
--
-- Fixed window: quotas are aligned to the clock (on the hour, midnight,
-- Monday 00:00). At the end of each period, the counter resets to zero.
-- Example: hourly quota = 50 images. If the consumer uses them at 14:10,
-- they must wait until 15:00 for 50 fresh images.
--
-- Multi-modality support:
-- type = "text" -> counted in tokens (from response SSE, default)
-- type = "ocr" -> counted in images (from request body, pre-proxy)
-- type = "audio" -> counted in seconds (header X-Audio-Duration)
-- type = "video" -> counted in seconds (header X-Video-Duration)
-- type = "multimodal" -> tokens (primary) + images (secondary via image_quotas)
-- type = "embedding" -> counted in tokens (same mechanisms as text)
--
-- For ocr/audio/video types, usage is known BEFORE the proxy,
-- which allows rejecting without consuming GPU.
-- For text/multimodal/embedding, tokens are only known after the
-- response: access phase checks the accumulated quota, log phase records it.
--
-- Multimodal models (e.g., Qwen-27b):
-- Multi-dimensional quotas: tokens (primary, hourly/daily/weekly),
-- images (image_quotas), audio (audio_quotas), and video (video_quotas).
-- Banned if any dimension is exceeded.
--
-- Ban hierarchy (top-down cascade):
-- weekly > daily > hourly
-- If a higher window is BANNED, all lower windows are blocked.
--
-- Per-GROUP quotas (v0.8, optional):
-- conf.group_quotas = { <group> = { model_quotas = {...}, ... } }
-- conf.consumer_groups = { <uid> = <group> } (membership, GitOps)
-- A consumer's group is resolved via conf.consumer_groups[uid] with
-- priority (reliable, does not depend on ctx.consumer), falling back to
-- ctx.consumer.labels/group_id, then conf.default_group, and finally the
-- root model_quotas (implicit default group). Counters remain per consumer
-- (unchanged key): individual counters, inherited group limit. Without
-- group_quotas, behavior is identical to previous versions.
-- All configuration (group_quotas + consumer_groups) is versioned in GitOps
-- within the route's plugin-config. See select_quota_set().
--
-- Shared dict storage (one key per consumer/model/dimension/window):
-- Key: consumer|model|dimension|wtype|window_start
-- Value: total usage in the window (integer, incremented atomically)
-- TTL: remaining time in the window + 2min margin
--
-- Headers returned when banned:
-- X-RateLimit-Ban-Level: weekly|daily|hourly
-- X-RateLimit-Model: lightonai/LightOnOCR-2-1B
-- X-RateLimit-Model-Type: ocr
-- X-RateLimit-Unit: images
-- X-RateLimit-Window: weekly
-- X-RateLimit-Limit: 500
-- X-RateLimit-Remaining: 0
-- Retry-After: <seconds until window resets> |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Hello,
I’d like to define per-consumer token quotas per instance, directly on the route, without having to duplicate the config on every consumer.
Something like this:
Meaning: every consumer gets 1M tokens/hour on qwen and 5M tokens/hour on devstral, with counters isolated per consumer.
Why
Today the only way to achieve this is to attach ai-rate-limiting on every consumer object with the same instances[] block. With dozens of consumers, this means:
• duplicating the same config N times
• updating N objects every time a quota changes
• no single source of truth on the route
• the ApisixConsumer v2 CRD doesn’t even accept ai-rate-limiting, so Kubernetes users have to bypass it entirely
A rules[] field inside instances[], keyed on ${consumer_name} (or any APISIX variable), would let operators define the quota policy once on the route and have it apply uniformly to all authenticated consumers.
Environment
• APISIX 3.16.0
• Used together with ai-proxy-multi and openid-connect (Keycloak)
All reactions