Skip to content

perf: resolve LLMDescriptionScorer cache hits before the rate limiter - #361

Open
kayaal34 wants to merge 1 commit into
deeppavlov:devfrom
kayaal34:perf/llm-scorer-skip-limiter-on-cache-hits
Open

kayaal34 wants to merge 1 commit into
deeppavlov:devfrom
kayaal34:perf/llm-scorer-skip-limiter-on-cache-hits

Conversation

@kayaal34

Copy link
Copy Markdown

Closes #354

Summary

LLMDescriptionScorer._compute_similarities sent every utterance through aiometer.run_all(..., max_per_second=...), and the cache lookup only happened inside Generator.get_structured_output_async, which runs after the limiter. So a fully cached predict on N utterances took at least N / max_per_second seconds (10 s per 100 utterances at the default), just for cache lookups.

Changes

  • Generator.get_cached_structured_output(messages, output_model): a cache-only lookup that never calls the API. It uses the same key as get_structured_output_* and returns None when caching is disabled.
  • LLMDescriptionScorer (async path): looks up the cache for all utterances first and sends only the misses to aiometer.run_all, so max_per_second now limits real API calls only. Results stay in input order. The sync path (max_concurrent=None) is unchanged.
  • Test fixtures: the mock generators return None from the new method (cache miss), so existing tests behave as before.

Testing

  • New tests: with a mixed cache, only the miss reaches get_structured_output_async, and cached rows score correctly. With a fully cached batch, the API is never awaited.
  • pytest tests/modules/scoring/test_description_llm.py tests/generation: 72 passed
  • ruff check, ruff format, and mypy on the changed files pass

This is scoped to #354. The error handling from #351 is a separate change.

🤖 Generated with Claude Code

Every utterance was handed to `aiometer.run_all(..., max_per_second=...)`,
and the cache lookup only happened inside `Generator.get_structured_output_async`,
i.e. after the limiter. A fully cached `predict` on N utterances therefore took
at least N / `max_per_second` seconds of pure cache lookups.

Add `Generator.get_cached_structured_output` (cache-only lookup, no API call)
and look up all utterances up front; only cache misses go through aiometer, so
`max_per_second` now gates real API calls only.

Closes deeppavlov#354

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 19, 2026 11:26

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The implementation preserves ordering, leaves synchronous behavior unchanged, and includes coverage for partial and complete cache hits.

Review effort: Lite
Findings: None

What changed in this PR

Optimizes asynchronous LLM description scoring by resolving cached results before rate limiting, avoiding unnecessary delays for cached utterances.

Changes:

  • Added a cache-only structured-output lookup to Generator.
  • Processes only cache misses through aiometer, preserving input order.
  • Updated fixtures and added mixed-cache and fully-cached tests.
File Description
src/​autointent/​generation/​_generator.py Adds cache-only lookup support.
src/​autointent/​modules/​scoring/​_description/​llm_encoder.py Bypasses rate limiting for cache hits.
tests/​_fixtures/​mock_generator.py Configures cache-miss behavior in mocks.
tests/​modules/​scoring/​test_description_llm.py Tests mixed and fully cached batches.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

LLMDescriptionScorer runs cache hits through the max_per_second limiter — 10 s per 100 cached utterances

2 participants