Skip to content

LocalHFBackend silently answers a blank prompt when SimpleContext is given an empty action #1705

Description

@planetf1

What goes wrong

#1632 makes the OpenAI, LiteLLM, and Ollama backends raise ValueError when the assembled conversation has no user content. That happens when a caller records turns on a SimpleContext and then generates with an empty action (#1597). LocalHFBackend has no such check. The same call sends the model one empty user turn, raises nothing, and returns an answer to a prompt the caller never wrote.

With ibm-granite/granite-4.2-3b the recorded turn is dropped and the model reasons about a message that doesn't exist:

import asyncio

from mellea.backends.huggingface import LocalHFBackend
from mellea.core import CBlock
from mellea.stdlib.context import SimpleContext


async def main():
    backend = LocalHFBackend(model_id="ibm-granite/granite-4.2-3b")
    ctx = SimpleContext().add(CBlock("Summarise this report: revenue rose 12% in Q3."))
    mot, _ = await backend.generate_from_context(CBlock(""), ctx)
    print(repr(await mot.avalue()))


asyncio.run(main())

Output:

'Okay, the user sent a message that\'s just the word "hello". I need to respond appropriately'

Expected: the same ValueError the other backends raise.

Mechanism

Both HF chat paths, standard (mellea/backends/huggingface.py:1782) and KV-cache (:1587), build messages through to_chat() in mellea/backends/utils.py:74. It takes ctx.view_for_generation(), which is empty for SimpleContext by design, appends the action, and hands the result to apply_chat_template without checking for user content. SimpleContext inherits _is_chat_context = True (mellea/core/base.py:2066), so it takes this path.

Capturing the conversation passed to apply_chat_template shows [{'role': 'user', 'content': ''}] on both HuggingFaceTB/SmolLM2-135M-Instruct and ibm-granite/granite-4.2-3b. A non-empty prompt through the same capture shows the expected user turn.

Suggested fix

Call has_user_content() in to_chat() straight after the action is appended and raise the same ValueError as the other backends. One call covers both HF paths. HF already rejects images and audio via _check_no_multimodal_blocks, so only the text and documents branches apply.

has_user_content() lives in mellea/helpers/openai_compatible_helpers.py. Once a non-OpenAI-compatible backend calls it, a backend-neutral helpers module is a better home.

Test: mirror test/backends/test_simple_context_guard.py, with SimpleContext + CBlock("") raising before generation, plus a documents pass-through case.

Scope and limits

  • Out of scope: WatsonxAIBackend (deprecated) and _generate_from_intrinsic, which doesn't append the action.
  • The token-burning stall from feat: add Granite 4.2 model defaults #1587 didn't reproduce here. Granite 4.2-3b answered the blank prompt in 74 tokens against a 1024-token cap (seed 0, one run). On HF the cost is a silently wrong answer.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions