What happens
LLMDescriptionScorer._compute_similarities hands every utterance to aiometer.run_all(..., max_at_once=max_concurrent, max_per_second=max_per_second) (llm_encoder.py:247-252); the cache lookup happens inside Generator.get_structured_output_async (_generator.py:266-270), i.e. after the limiter. A fully cached predict on N utterances therefore takes ≥ N / max_per_second seconds — 10 s per 100 utterances at the default max_per_second=10 — doing nothing but dictionary/disk lookups.
- warm
pipeline.predict on 100 test utterances: 10.07 s (banking77), 10.12 s (hwu64) — vs 0.002 s for description_typesafe, which looks the cache up first and only sends misses through aiometer (typesafe.py:290-300);
- HPO: each of the 30 scoring trials re-scores
train_1 + validation (all cache hits after trial 1), so the pipeline fit is limiter-bound whether cold or warm — 758 s cold vs 729 s warm on banking77, 649 s vs 618 s on hwu64.
Proposed
Check the cache for all utterances up front and submit only the misses to aiometer.run_all. Needs a cache-only lookup on Generator (e.g. get_cached(messages, output_model) / aget_cached) or access to self.cache; max_per_second then gates real API calls only. Cuts the API arm's HPO wall-clock by roughly an order of magnitude.
Follow-up from #350.
What happens
LLMDescriptionScorer._compute_similaritieshands every utterance toaiometer.run_all(..., max_at_once=max_concurrent, max_per_second=max_per_second)(llm_encoder.py:247-252); the cache lookup happens insideGenerator.get_structured_output_async(_generator.py:266-270), i.e. after the limiter. A fully cachedpredicton N utterances therefore takes ≥ N /max_per_secondseconds — 10 s per 100 utterances at the defaultmax_per_second=10— doing nothing but dictionary/disk lookups.Measured (Darinochka/AutoIntent-experiments#43, gpt-6 via OpenRouter)
pipeline.predicton 100 test utterances: 10.07 s (banking77), 10.12 s (hwu64) — vs 0.002 s fordescription_typesafe, which looks the cache up first and only sends misses through aiometer (typesafe.py:290-300);train_1+validation(all cache hits after trial 1), so the pipeline fit is limiter-bound whether cold or warm — 758 s cold vs 729 s warm on banking77, 649 s vs 618 s on hwu64.Proposed
Check the cache for all utterances up front and submit only the misses to
aiometer.run_all. Needs a cache-only lookup onGenerator(e.g.get_cached(messages, output_model)/aget_cached) or access toself.cache;max_per_secondthen gates real API calls only. Cuts the API arm's HPO wall-clock by roughly an order of magnitude.Follow-up from #350.