Repository navigation
Conversation
…-call-type Fix/proxy only failure call type
…age-metadata fix(langsmith): populate usage_metadata in outputs for Cost column
…ssage-detection-performance Fix model repetition detection performance
fix: fix logging for response incomplete streaming + custom pricing on /v1/messages and /v1/responses
…idebar - Add 'Contributing to Guardrails' category with links to: - Generic Guardrail API (integrate without PR) - Adding a New Guardrail Integration tutorial - Adding Guardrail Support to Endpoints - Add 'Team Bring-Your-Own Guardrails' link for team BYOG workflow These docs existed but were only accessible from the 'LiteLLM AI Gateway' sidebar. Now they're also accessible when browsing the 'Guardrail Providers' section. Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
…ls-docs-143b docs: add Contributing to Guardrails section to Guardrail Providers sidebar
Adds `default_api_key_tpm_limit` and `default_api_key_rpm_limit` to `GenericLiteLLMParams` so operators can set per-deployment rate limit defaults in config.yaml. When a key has no model-specific tpm/rpm limit configured, the proxy falls back to these deployment defaults (Case 2 in spec). Key-level limits always take priority (Case 1). - Extends `get_key_model_tpm_limit` / `get_key_model_rpm_limit` with a `model_name` param and a priority-4 deployment-default fallback - Passes `model_name=requested_model` in the parallel request limiter so the fallback is triggered at enforcement time - Adds `"limit"` to `SensitiveDataMasker` non-sensitive overrides so `*_limit` fields are not masked in `/model/info` responses - Adds 17 unit tests covering both spec cases and the `/model/info` path Co-Authored-By: Claude (claude-sonnet-4-6) <noreply@anthropic.com>
…82.3 changelog Helicone (PRs BerriAI#19288, BerriAI#22603) and Langfuse (BerriAI#22390) were present in the v1.82.0-stable...v1.82.3-stable diff but omitted from the AI Integrations logging section. Also updates the AI Integrations diff summary count from 2 to 4. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Use min() across all matching deployments instead of first-wins when resolving default_api_key_tpm/rpm_limit for a model group, so load-balanced setups with different per-deployment limits always apply the most conservative value - Replace the global SensitiveDataMasker non_sensitive_overrides change with a targeted excluded_keys set at the remove_sensitive_info_from_deployment call site, avoiding unintended suppression of other fields - Update the v1 parallel request limiter to pass model_name to get_key_model_tpm/rpm_limit so deployment defaults apply there too - Add 4 tests covering multi-deployment min semantics Co-Authored-By: Claude (claude-sonnet-4-6) <noreply@anthropic.com>
Compute get_key_model_tpm/rpm_limit once before the guard condition instead of calling each function twice (once to check non-None, once to retrieve). Removes 2 extra llm_router.get_model_list() calls per request when deployment defaults are active. Co-Authored-By: Claude (claude-sonnet-4-6) <noreply@anthropic.com>
…ult limits async_log_success_event only updated the per-model cache counter when model_rpm_limit / model_tpm_limit were present in key metadata or model_max_budget was set. For the new deployment-default path (default_api_key_tpm_limit / default_api_key_rpm_limit), none of those conditions held, so current_tpm stayed at zero and tpm enforcement was never applied across multiple requests. Extend the guard condition to also trigger when the model group has a deployment-default tpm or rpm limit, and import the two helpers at module level. Co-Authored-By: Claude (claude-sonnet-4-6) <noreply@anthropic.com>
Replace bare _get_deployment_default_tpm/rpm_limit calls in the async_log_success_event condition with get_key_model_tpm/rpm_limit (model_name=model_group). The higher-level getters short-circuit on key/team metadata hits before ever reaching the router, so requests that don't use deployment defaults incur no extra router lookup. Remove the now-unused bare helper imports. Also fix invalid `int = None` type hints in test helper signatures to `Optional[int] = None`. Co-Authored-By: Claude (claude-sonnet-4-6) <noreply@anthropic.com>
…ures Full audit of 371 PRs in v1.82.0-stable...v1.82.3-stable range. Adds previously undocumented user-facing changes: - Key Highlights: Hashicorp Vault, Responses WebSocket, Org Admin RBAC, guardrail mode defaults - New Providers: Google Search API, Bedrock Mantle (7 total, was 5) - LLM API: Anthropic Files API, Mistral Voxtral transcription, WebRTC, Responses WebSocket, litellm.acount_tokens() public API, OpenRouter image edit, Vertex AI VIDEO token tracking, input_fidelity image edit, model cost aliases, per-request json schema validation, 15+ bug fixes - Management: RBAC expansion for Org Admins, Vector Store CRUD, MCP token auth + team scoping, BYOK key precedence, virtual key spend reset, batch expiry for teams, Admin Viewer audit log access, 12+ bug fixes - Guardrails: mode default list, tag-based modes, presidio fix, OTEL fix - Secret Managers: Hashicorp Vault (was "no changes") - Spend Tracking: new section — budget-linked reset fix, flex pricing, spend log cleanup, WebSearch dedup fix - Performance: 4 additional reliability fixes - Diff summary counts updated to reflect actual scope Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add MCP Gateway section (moved from Management per guide rule §11) - Rename Spend Tracking → Spend Tracking, Budgets and Rate Limiting - Fix Hashicorp Vault doc link: docs/secret → docs/secret_managers - Fix LLM API section: #### Bug Fixes → #### Bugs (matches guide) - Add Documentation Updates section (required by guide §11) - Update Diff Summary: correct section names, add MCP Gateway and Documentation Updates counts Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
BerriAI#23931) - Change chunk["id"] to chunk.get("id") for compatibility with MiniMax - ModelResponseStream auto-generates id when None is passed - Add regression test test_chunk_parser_without_id_field
Move pre-call checks (rate limits, guardrails, budget) to run BEFORE polling ID creation in the background streaming flow. This prevents the edge case where a rate-limited request receives a polling ID that immediately fails. Changes: - Add skip_pre_call_logic parameter to base_process_llm_request to allow skipping pre-call checks (avoiding double-counting of RPM/parallel requests) - Run common_processing_pre_call_logic before generating polling ID in the responses API endpoint. If rate limits/guardrails fail, return error immediately without creating a polling ID - Background streaming task passes skip_pre_call_logic=True to avoid re-running pre-call checks that were already done before polling ID creation - Add tests verifying skip_pre_call_logic parameter works correctly Fixes the edge case where polling_via_cache would return a polling ID for a request that immediately fails due to rate limiting.
- Guard logging_obj for None when skip_pre_call_logic=True: raise ValueError if litellm_logging_obj not in data, preventing AttributeError downstream - Add model=None to common_processing_pre_call_logic call in endpoints.py to match style of other call sites - Add test verifying rate-limited request never receives polling ID
…) directly Previously the test called common_processing_pre_call_logic in isolation, making generate_polling_id.assert_not_called() vacuously true. Now the test calls responses_api() end-to-end so it actually verifies that a rate-limited request never receives a polling ID. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Enable deployment_affinity, responses_api_deployment_check, and session_affinity to be configured per model group via router_settings.model_group_affinity_config, falling back to global settings for unconfigured groups. - Add model_group_affinity_config parameter to Router and DeploymentAffinityCheck - Add _get_effective_flags helper to resolve flags per model group - Update async_filter_deployments and async_pre_call_deployment_hook to use per-group config - Add 4 comprehensive tests covering per-group config, fallback, and override scenarios This allows fine-grained control of affinity behavior across model groups, e.g., enabling stickiness only for cross-provider deployments while leaving other groups free to load-balance. Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
…n on unknown affinity flags Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
merge main
…fault Aligns proxy default with litellm.AZURE_DEFAULT_API_VERSION (2025-02-01-preview) so Azure response_format + json_schema works without tools fallback. Made-with: Cursor
- Extract helper methods in langsmith._prepare_log_data to reduce from 51 to <50 statements - Extract helper methods in anthropic.transform_parsed_response to reduce from 57 to <50 statements - Fixes PLR0915 linter errors - All existing tests pass (10 langsmith tests, 126 anthropic tests) Made-with: Cursor
…17_2026 fix(fireworks): skip #transform=inline for base64 data URLs (BerriAI#23729)
…-activity-entity-breakdown fix(proxy): restore per-entity breakdown in aggregated daily activity endpoint
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Extract multiline `or` chain from LiteLLM_AuditLogs constructor to fix pydantic mypy plugin field-type misattribution, and add explicit Optional[bool] annotation to avoid variable name shadowing conflict.
…05_2026 Litellm oss staging 03 05 2026
…arch_week Litellm dev sameer 16 march week
New docs page covering the HA control plane architecture where each worker instance has its own DB, Redis, and master key. Includes a React component diagram, setup configs, SSO notes, and local testing instructions.
litellm ryan march 20
[Infra] Build UI for release
…activity endpoint" This reverts commit 9c3fab2.
…hboard routes - OldTeams: refresh table via fetchTeamsV2 after team create instead of appending - TeamDropdown: rewrite with useInfiniteTeams for paginated fetch, scroll-to-load, and debounced search - Update all TeamDropdown consumers to use the new self-fetching API - Dashboard layout: switch from Sidebar2 to SidebarProvider (leftnav) - Leftnav: add MIGRATED_PAGES routing for path-based navigation (api-reference) - Navbar: remove chat button Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Use Select.Option with font-medium alias + Text secondary ID to match OrganizationDropdown - Default page size to 20 - Add useInfiniteTeams mock to AddModelForm tests Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
[Fix] UI - Teams: Table refresh, infinite dropdown, leftnav migration
| ANTHROPIC_API_KEY: ${{ secrets.LITELLM_VIRTUAL_KEY }} | ||
| ANTHROPIC_BASE_URL: ${{ secrets.LITELLM_BASE_URL }} | ||
| - name: Check for potential duplicates | ||
| uses: wow-actions/potential-duplicates@v1 |
Check warning
Code scanning / CodeQL
Unpinned tag for a non-immutable Action in workflow Medium
| pip install pytest pytest-codspeed==4.3.0 | ||
| - name: Run benchmarks | ||
| uses: CodSpeedHQ/action@v4 |
Check warning
Code scanning / CodeQL
Unpinned tag for a non-immutable Action in workflow Medium
| runs-on: ubuntu-latest | ||
|
|
||
| steps: | ||
| - name: Checkout repository | ||
| uses: actions/checkout@v3 | ||
| with: | ||
| fetch-depth: 0 | ||
|
|
||
| - name: Create internal dev branch | ||
| env: | ||
| GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} | ||
| run: | | ||
| # Configure Git user | ||
| git config user.name "github-actions[bot]" | ||
| git config user.email "github-actions[bot]@users.noreply.github.com" | ||
| # Generate branch name with MM_DD_YYYY format | ||
| BRANCH_NAME="litellm_internal_dev_$(date +'%m_%d_%Y')" | ||
| echo "Creating branch: $BRANCH_NAME" | ||
| # Fetch all branches | ||
| git fetch --all | ||
| # Check if the branch already exists | ||
| if git show-ref --verify --quiet refs/remotes/origin/$BRANCH_NAME; then | ||
| echo "Branch $BRANCH_NAME already exists. Skipping creation." | ||
| else | ||
| echo "Creating new branch: $BRANCH_NAME" | ||
| # Create the new branch from main | ||
| git checkout -b $BRANCH_NAME origin/main | ||
| # Push the new branch | ||
| git push origin $BRANCH_NAME | ||
| echo "Successfully created and pushed branch: $BRANCH_NAME" | ||
| fi |
Check warning
Code scanning / CodeQL
Workflow does not contain permissions Medium
Show autofix suggestion
Hide autofix suggestion
Copilot Autofix
AI 7 months ago
In general, the problem is fixed by explicitly declaring a permissions block either at the top level of the workflow (applying to all jobs) or per job, and granting only the minimal required scopes. This workflow only needs to read from and write to repository contents to create and push new branches, so contents: write is sufficient. No other scopes (like issues, pull-requests, or packages) are required.
The best fix with no behavior change is to add a workflow-level permissions block just under the name: line. This will apply to both create-staging-branch and create-internal-dev-branch without repeating configuration, and will keep the ability to push branches (which requires contents: write). Concretely, in .github/workflows/create_daily_staging_branch.yml, after line 1 (name: Create Daily Staging Branch), insert:
permissions:
contents: writeNo additional imports or methods are needed, as this is purely a YAML configuration change for GitHub Actions.
| @@ -1,4 +1,6 @@ | ||
| name: Create Daily Staging Branch | ||
| permissions: | ||
| contents: write | ||
|
|
||
| on: | ||
| schedule: |
| runs-on: ubuntu-latest | ||
| timeout-minutes: 30 | ||
|
|
||
| services: | ||
| postgres: | ||
| image: postgres:15 | ||
| env: | ||
| POSTGRES_USER: llmproxy | ||
| POSTGRES_PASSWORD: dbpassword9090 | ||
| POSTGRES_DB: litellm | ||
| ports: | ||
| - 5432:5432 | ||
| options: >- | ||
| --health-cmd pg_isready | ||
| --health-interval 10s | ||
| --health-timeout 5s | ||
| --health-retries 5 | ||
| steps: | ||
| - uses: actions/checkout@v4 | ||
|
|
||
| - name: Set up Python | ||
| uses: actions/setup-python@v5 | ||
| with: | ||
| python-version: "3.12" | ||
|
|
||
| - name: Install Poetry | ||
| uses: snok/install-poetry@v1 | ||
|
|
||
| - name: Cache Poetry dependencies | ||
| uses: actions/cache@v4 | ||
| with: | ||
| path: | | ||
| ~/.cache/pypoetry | ||
| ~/.cache/pip | ||
| .venv | ||
| key: ${{ runner.os }}-poetry-e2e-batches-${{ hashFiles('poetry.lock') }} | ||
| restore-keys: | | ||
| ${{ runner.os }}-poetry-e2e-batches- | ||
| ${{ runner.os }}-poetry- | ||
| - name: Install dependencies | ||
| run: | | ||
| poetry config virtualenvs.in-project true | ||
| poetry install --with dev,proxy-dev --extras "proxy" | ||
| poetry run pip install psycopg2-binary uvicorn fastapi httpx tenacity | ||
| - name: Setup litellm-enterprise | ||
| run: | | ||
| poetry run pip install --force-reinstall --no-deps -e enterprise/ | ||
| - name: Generate Prisma client | ||
| run: | | ||
| poetry run prisma generate --schema litellm/proxy/schema.prisma | ||
| - name: Run Prisma migrations | ||
| env: | ||
| DATABASE_URL: postgresql://llmproxy:dbpassword9090@localhost:5432/litellm | ||
| run: | | ||
| cd litellm/proxy | ||
| poetry run prisma migrate deploy --schema schema.prisma | ||
| cd ../.. | ||
| - name: Run Azure Batch E2E Tests | ||
| env: | ||
| DATABASE_URL: postgresql://llmproxy:dbpassword9090@localhost:5432/litellm | ||
| USE_LOCAL_LITELLM: "true" | ||
| USE_MOCK_MODELS: "true" | ||
| USE_STATE_TRACKER: "true" | ||
| LITELLM_LOG: DEBUG | ||
| run: | | ||
| poetry run pytest tests/proxy_e2e_azure_batches_tests/test_proxy_e2e_azure_batches.py \ | ||
| -vv -s -k "test_e2e_managed_batch" \ | ||
| --tb=short \ | ||
| --maxfail=3 \ | ||
| --durations=10 |
Check warning
Code scanning / CodeQL
Workflow does not contain permissions Medium test
Show autofix suggestion
Hide autofix suggestion
Copilot Autofix
AI 7 months ago
To fix this, explicitly declare minimal GITHUB_TOKEN permissions in the workflow. Since the job only needs to read repository contents (for actions/checkout) and does not push code, manage releases, or modify issues/PRs, we can safely set contents: read. This should be done at the workflow (root) level so it applies to all jobs unless overridden.
Concretely, in .github/workflows/test-proxy-e2e-azure-batches.yml, add a top-level permissions: block after the on: section and before concurrency: (or before jobs:). The block should specify contents: read, which corresponds to GitHub’s suggested minimal starting point and is sufficient for actions/checkout@v4. No other steps in the shown snippet require additional scopes, so we avoid granting unnecessary write permissions. No code logic or job behavior changes; we only constrain the ambient token.
| @@ -5,6 +5,9 @@ | ||
| branches: [main] | ||
| workflow_dispatch: | ||
|
|
||
| permissions: | ||
| contents: read | ||
|
|
||
| concurrency: | ||
| group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }} | ||
| cancel-in-progress: true |
| python-version: "3.12" | ||
|
|
||
| - name: Install Poetry | ||
| uses: snok/install-poetry@v1 |
Check warning
Code scanning / CodeQL
Unpinned tag for a non-immutable Action in workflow Medium test
| // Exchange it for the JWT via the worker's /v3/login/exchange endpoint. | ||
| const params = new URLSearchParams(window.location.search); | ||
| const ssoCode = params.get("code"); | ||
| if (ssoCode) { |
Check failure
Code scanning / CodeQL
User-controlled bypass of security check High
| // Only redirect if the return URL is different from the current URL | ||
| // This prevents infinite redirect loops | ||
| if (normalizedReturnUrl !== normalizedCurrentUrl) { | ||
| window.location.replace(returnUrl); |
Check failure
Code scanning / CodeQL
Client-side cross-site scripting High
Show autofix suggestion
Hide autofix suggestion
Copilot Autofix
AI 7 months ago
To fix this, we should ensure that only safe, same-origin paths are used for client-side redirects and that the redirect is performed via a trusted base URL rather than the raw user-controlled value.
General approach:
- Treat values derived from
window.location.searchas untrusted. - Before redirecting, parse
returnUrland enforce:- Protocol is
http:orhttps:(and ideally matches current). - Origin matches
window.location.origin. - Or, more simply, only allow path-relative URLs (starting with
/) and never absolute external URLs.
- Protocol is
- Combine the validated path with the current origin when calling
window.location.replace.
Concrete change in this code:
- In
ui/litellm-dashboard/src/app/page.tsx, right before usingwindow.location.replace(returnUrl), derive a safe redirect URL:- If
returnUrlstarts with/, keep it (path-only redirect). - Else, try to
new URL(returnUrl, window.location.origin)and enforceurl.origin === window.location.origin; if not, skip redirect.
- If
- Call
window.location.replace(safeUrl)instead of the rawreturnUrl.
This change is local to the useEffect that performs the post-authentication redirect (lines 265–289) and does not require new imports or changes elsewhere. It preserves intended functionality (redirecting back within the same app) while blocking open redirects / potential client-side XSS via hostile URLs.
| @@ -283,7 +283,26 @@ | ||
| // Only redirect if the return URL is different from the current URL | ||
| // This prevents infinite redirect loops | ||
| if (normalizedReturnUrl !== normalizedCurrentUrl) { | ||
| window.location.replace(returnUrl); | ||
| // Enforce same-origin / safe redirects only | ||
| let safeRedirectUrl: string | null = null; | ||
| try { | ||
| if (returnUrl.startsWith("/")) { | ||
| // Path-relative URL within the same origin | ||
| safeRedirectUrl = returnUrl; | ||
| } else { | ||
| const parsed = new URL(returnUrl, window.location.origin); | ||
| if (parsed.origin === window.location.origin) { | ||
| safeRedirectUrl = parsed.href; | ||
| } | ||
| } | ||
| } catch { | ||
| // Malformed URLs are ignored | ||
| safeRedirectUrl = null; | ||
| } | ||
|
|
||
| if (safeRedirectUrl) { | ||
| window.location.replace(safeRedirectUrl); | ||
| } | ||
| } | ||
| } | ||
| }, [authLoading, token]); |
| // Only redirect if the return URL is different from the current URL | ||
| // This prevents infinite redirect loops | ||
| if (normalizedReturnUrl !== normalizedCurrentUrl) { | ||
| window.location.replace(returnUrl); |
Check warning
Code scanning / CodeQL
Client-side URL redirect Medium
Show autofix suggestion
Hide autofix suggestion
Copilot Autofix
AI 7 months ago
General fix: Ensure that any client-side redirect constructed from user-influenced data is constrained to safe destinations (typically same-origin, relative paths) or is chosen from a server-/code-defined allowlist. Never pass a raw query parameter or cookie value directly to window.location or window.location.replace without enforcing such constraints.
Best fix here without changing behavior: Keep using consumeReturnUrl() but, right before redirecting, normalize the URL into a safe, same-origin path. If returnUrl is absolute and same-origin, we can safely use its path/search/hash. If it is relative, we can construct a URL with window.location.origin as the base and again only use the path/search/hash. If the URL is invalid or cross-origin, we simply do not redirect. This both hardens the redirect and satisfies CodeQL by ensuring that the tainted string is not directly used as the redirect target.
Concretely in ui/litellm-dashboard/src/app/page.tsx:
-
In the
useEffectthat currently does:const returnUrl = consumeReturnUrl(); ... if (normalizedReturnUrl !== normalizedCurrentUrl) { window.location.replace(returnUrl); }
change it to:
- Parse
returnUrlusing theURLconstructor withwindow.location.originas a base (to handle relative URLs). - Check that
url.origin === window.location.origin. If not, skip redirect. - Build a safe redirect target from
url.pathname + url.search + url.hash. - Optionally compare normalized versions using this safe target instead of the raw
returnUrl. - Call
window.location.replace(safeTarget).
- Parse
No new methods or imports are required; we use the standard URL API and existing normalizeUrlForCompare. All changes stay within the shown useEffect in page.tsx. We do not need to modify returnUrlUtils.ts for this fix.
| @@ -278,12 +278,30 @@ | ||
| const returnUrl = consumeReturnUrl(); | ||
| if (returnUrl) { | ||
| const currentUrl = window.location.href; | ||
| const normalizedReturnUrl = normalizeUrlForCompare(returnUrl); | ||
| const normalizedCurrentUrl = normalizeUrlForCompare(currentUrl); | ||
| // Only redirect if the return URL is different from the current URL | ||
| // This prevents infinite redirect loops | ||
| if (normalizedReturnUrl !== normalizedCurrentUrl) { | ||
| window.location.replace(returnUrl); | ||
| try { | ||
| // Normalize the return URL against the current origin to handle relative URLs | ||
| const currentOrigin = new URL(currentUrl).origin; | ||
| const parsedReturnUrl = new URL(returnUrl, currentOrigin); | ||
|
|
||
| // Only allow redirects to the same origin to prevent open redirect attacks | ||
| if (parsedReturnUrl.origin !== currentOrigin) { | ||
| // Unsafe origin; do not redirect | ||
| return; | ||
| } | ||
|
|
||
| const safeReturnPath = | ||
| parsedReturnUrl.pathname + parsedReturnUrl.search + parsedReturnUrl.hash; | ||
|
|
||
| const normalizedReturnUrl = normalizeUrlForCompare(safeReturnPath); | ||
| const normalizedCurrentUrl = normalizeUrlForCompare(currentUrl); | ||
| // Only redirect if the return URL is different from the current URL | ||
| // This prevents infinite redirect loops | ||
| if (normalizedReturnUrl !== normalizedCurrentUrl) { | ||
| window.location.replace(safeReturnPath); | ||
| } | ||
| } catch { | ||
| // Malformed return URL; safely ignore and do not redirect | ||
| return; | ||
| } | ||
| } | ||
| }, [authLoading, token]); |
Summary
main1.82.6and the upstream responses/reasoning fixes landed before and around 2026-03-22Verification
git rev-list --left-right --count HEAD...upstream/main->0 0