Skip to content

fix(knowledge): give document processing per-tenant queue lanes - #7937

Merged
waleedlatif1 merged 1 commit into
stagingfrom
investigate/integ-failure-kb-ingestion
Sep 17, 2026
Merged

waleedlatif1 merged 1 commit into
stagingfrom
investigate/integ-failure-kb-ingestion

Conversation

@waleedlatif1

Copy link
Copy Markdown
Collaborator

Summary

  • Document processing ran through one Trigger.dev queue with a global concurrency limit and no concurrency key, so a single connector backfill could hold every slot and leave every other tenant's uploads queued behind it for hours.
  • Dispatch now names a lane — interactive for work a person is waiting on, backfill for connector-driven ingestion — and sets a concurrencyKey for the entity that owns the knowledge base, so the limit applies per tenant copy of the queue instead of fleet-wide.
  • Key on the owning entity, not the workspace: an organization-scoped knowledge base carries a null workspace id, so a workspace-only key would collapse every such tenant onto one bucket and reproduce the starvation. Owned scopes reuse resourceScopeKey so tenant identity keeps one spelling.
  • The lane is stamped on the payload, so quota and capacity continuations resume in the lane they were admitted against rather than promoting themselves out of the backfill ceiling on every retry.
  • An absent or unrecognized lane parses as backfill instead of throwing, so payloads written before the lanes existed — and payloads stamped by a newer version mid-rollout — cannot burn a run's retry budget.
  • A rejected backfill chunk is no longer absorbed in-process. Those documents are reported failed and reclaimed by the stuck-document sweep instead of outlasting the connector lease the dispatching worker is holding.

Throughput

Both lanes default to the limit the single shared queue carried, so nothing drains slower than before:

  • one tenant backfilling: same concurrency as today
  • one tenant backfilling while someone uploads: the upload starts immediately instead of queueing behind the backfill
  • two or more tenants: each gets its own allowance, where previously they split one

KB_CONFIG_BACKFILL_CONCURRENCY_LIMIT is a separate variable from KB_CONFIG_CONCURRENCY_LIMIT so backfill can be dialled down without slowing person-facing work. Worth noting for review: the fleet-wide ceiling is now the Trigger.dev environment concurrency limit rather than this queue's own limit, so aggregate ingestion can rise above what it could reach before. Lowering the backfill variable is the lever if that budget gets tight.

Deploy note

Interactive work deliberately stays on the existing queue name. A run naming a queue the running Trigger.dev worker has not registered parks in PENDING_VERSION, and the app deploys separately from the worker — so only the new backfill queue can be caught by that window, and stranded backfill work is exactly what the stuck-document sweep already recovers.

Type of Change

  • Bug fix

Testing

  • Full apps/sim suite: 53,005 passed, 1 pre-existing failure that needs DATABASE_URL
  • bun run type-check clean, bun run lint clean, all 46 audits pass, docs manifest in sync
  • 18 new tests covering lane routing, tenant-key derivation across all three billing scopes, continuation lane stickiness, and the backfill fallback guard; each verified to fail when its fix is reverted

Checklist

  • Code follows project style guidelines
  • Self-reviewed my changes
  • Tests added/updated and passing
  • No new warnings introduced
  • I confirm that I have read and agree to the terms outlined in the Contributor License Agreement (CLA)

Document processing ran through one Trigger.dev queue with a global
concurrency limit and no concurrency key, so a single connector backfill
could hold every slot and leave every other tenant's uploads queued
behind it for hours.

Dispatch now names a lane — interactive for work a person is waiting on,
backfill for connector-driven ingestion — and keys each run by the
entity that owns the knowledge base, so the limit applies per tenant
copy of the queue rather than fleet-wide.

- Key on the owning entity, not the workspace: an organization-scoped
  knowledge base carries a null workspace id and would otherwise share
  one bucket with every other such tenant. Owned scopes reuse
  resourceScopeKey so tenant identity keeps one spelling.
- Stamp the lane on the payload so quota and capacity continuations
  resume in the lane they were admitted against instead of promoting
  themselves out of the backfill ceiling.
- Parse an absent or unrecognized lane as backfill rather than throwing,
  so payloads written before the lanes existed and payloads stamped by a
  newer version mid-rollout cannot burn a run's retry budget.
- Keep interactive work on the pre-existing queue name; a queue the
  running worker has not registered parks its runs in PENDING_VERSION,
  and the app deploys separately from the worker.
- Stop absorbing a rejected backfill chunk in-process: those documents
  are reported failed and reclaimed by the stuck-document sweep instead
  of outlasting the connector lease they run under.
@vercel

vercel Bot commented Sep 17, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Sep 17, 2026 7:22pm UTC

Request Review

@greptile-apps

greptile-apps Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with lane routing, tenant isolation, continuation behavior, and rejected-backfill recovery handled consistently.

Summary

This PR separates knowledge-document processing into per-tenant interactive and connector-backfill lanes while preserving lane assignment through durable dispatch and continuation paths.

  • Adds independently configurable Trigger.dev queues with owner-derived concurrency keys.
  • Routes direct uploads and manual retries through the interactive lane and connector ingestion through backfill.
  • Preserves lane metadata across outbox, recovery, quota, and provider-capacity continuations.
  • Leaves rejected backfill chunks for the existing stuck-document recovery sweep rather than extending connector workers with in-process work.
  • Adds focused coverage for routing, billing scopes, continuation stickiness, compatibility fallback, and failed backfill enqueue behavior.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Direct upload or manual retry] -->|interactive| I[Interactive queue]
  B[Connector ingestion or stuck sweep] -->|backfill| F[Backfill queue]
  I --> K[Concurrency key derived from KB owner]
  F --> K
  K --> T[knowledge-process-document]
  T --> D{Quota or provider deferral?}
  D -->|No| P[Complete processing]
  D -->|Yes| C[Continuation preserves payload lane]
  C --> I
  C --> F
  F -->|Enqueue rejected| S[Return document to pending]
  S --> R[Later connector stuck-document sweep]
Loading

Reviews (1) · Last reviewed commit: "fix(knowledge): give document processing..."

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 25 files

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Re-trigger cubic

@waleedlatif1
waleedlatif1 merged commit 41488fd into staging Sep 17, 2026
34 checks passed
@waleedlatif1
waleedlatif1 deleted the investigate/integ-failure-kb-ingestion branch September 17, 2026 19:36

This branch was previously deployed

1 inactive deployment
Preview — 58dfb3d2 Deployed Sep 17, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant