Conversation
Every utterance was handed to `aiometer.run_all(..., max_per_second=...)`, and the cache lookup only happened inside `Generator.get_structured_output_async`, i.e. after the limiter. A fully cached `predict` on N utterances therefore took at least N / `max_per_second` seconds of pure cache lookups. Add `Generator.get_cached_structured_output` (cache-only lookup, no API call) and look up all utterances up front; only cache misses go through aiometer, so `max_per_second` now gates real API calls only. Closes deeppavlov#354 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The implementation preserves ordering, leaves synchronous behavior unchanged, and includes coverage for partial and complete cache hits.
Review effort: Lite
Findings: None
What changed in this PR
Optimizes asynchronous LLM description scoring by resolving cached results before rate limiting, avoiding unnecessary delays for cached utterances.
Changes:
- Added a cache-only structured-output lookup to
Generator. - Processes only cache misses through
aiometer, preserving input order. - Updated fixtures and added mixed-cache and fully-cached tests.
| File | Description |
|---|---|
src/autointent/generation/_generator.py |
Adds cache-only lookup support. |
src/autointent/modules/scoring/_description/llm_encoder.py |
Bypasses rate limiting for cache hits. |
tests/_fixtures/mock_generator.py |
Configures cache-miss behavior in mocks. |
tests/modules/scoring/test_description_llm.py |
Tests mixed and fully cached batches. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #354
Summary
LLMDescriptionScorer._compute_similaritiessent every utterance throughaiometer.run_all(..., max_per_second=...), and the cache lookup only happened insideGenerator.get_structured_output_async, which runs after the limiter. So a fully cachedpredicton N utterances took at least N /max_per_secondseconds (10 s per 100 utterances at the default), just for cache lookups.Changes
Generator.get_cached_structured_output(messages, output_model): a cache-only lookup that never calls the API. It uses the same key asget_structured_output_*and returnsNonewhen caching is disabled.LLMDescriptionScorer(async path): looks up the cache for all utterances first and sends only the misses toaiometer.run_all, somax_per_secondnow limits real API calls only. Results stay in input order. The sync path (max_concurrent=None) is unchanged.Nonefrom the new method (cache miss), so existing tests behave as before.Testing
get_structured_output_async, and cached rows score correctly. With a fully cached batch, the API is never awaited.pytest tests/modules/scoring/test_description_llm.py tests/generation: 72 passedruff check,ruff format, andmypyon the changed files passThis is scoped to #354. The error handling from #351 is a separate change.
🤖 Generated with Claude Code