Skip to content

LLMDescriptionScorer runs cache hits through the max_per_second limiter — 10 s per 100 cached utterances #354

Description

@voorhs

What happens

LLMDescriptionScorer._compute_similarities hands every utterance to aiometer.run_all(..., max_at_once=max_concurrent, max_per_second=max_per_second) (llm_encoder.py:247-252); the cache lookup happens inside Generator.get_structured_output_async (_generator.py:266-270), i.e. after the limiter. A fully cached predict on N utterances therefore takes ≥ N / max_per_second seconds — 10 s per 100 utterances at the default max_per_second=10 — doing nothing but dictionary/disk lookups.

Measured (Darinochka/AutoIntent-experiments#43, gpt-6 via OpenRouter)

  • warm pipeline.predict on 100 test utterances: 10.07 s (banking77), 10.12 s (hwu64) — vs 0.002 s for description_typesafe, which looks the cache up first and only sends misses through aiometer (typesafe.py:290-300);
  • HPO: each of the 30 scoring trials re-scores train_1 + validation (all cache hits after trial 1), so the pipeline fit is limiter-bound whether cold or warm — 758 s cold vs 729 s warm on banking77, 649 s vs 618 s on hwu64.

Proposed

Check the cache for all utterances up front and submit only the misses to aiometer.run_all. Needs a cache-only lookup on Generator (e.g. get_cached(messages, output_model) / aget_cached) or access to self.cache; max_per_second then gates real API calls only. Cuts the API arm's HPO wall-clock by roughly an order of magnitude.

Follow-up from #350.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions