You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
#1632 makes the OpenAI, LiteLLM, and Ollama backends raise ValueError when the assembled conversation has no user content. That happens when a caller records turns on a SimpleContext and then generates with an empty action (#1597). LocalHFBackend has no such check. The same call sends the model one empty user turn, raises nothing, and returns an answer to a prompt the caller never wrote.
With ibm-granite/granite-4.2-3b the recorded turn is dropped and the model reasons about a message that doesn't exist:
importasynciofrommellea.backends.huggingfaceimportLocalHFBackendfrommellea.coreimportCBlockfrommellea.stdlib.contextimportSimpleContextasyncdefmain():
backend=LocalHFBackend(model_id="ibm-granite/granite-4.2-3b")
ctx=SimpleContext().add(CBlock("Summarise this report: revenue rose 12% in Q3."))
mot, _=awaitbackend.generate_from_context(CBlock(""), ctx)
print(repr(awaitmot.avalue()))
asyncio.run(main())
Output:
'Okay, the user sent a message that\'s just the word "hello". I need to respond appropriately'
Expected: the same ValueError the other backends raise.
Mechanism
Both HF chat paths, standard (mellea/backends/huggingface.py:1782) and KV-cache (:1587), build messages through to_chat() in mellea/backends/utils.py:74. It takes ctx.view_for_generation(), which is empty for SimpleContext by design, appends the action, and hands the result to apply_chat_template without checking for user content. SimpleContext inherits _is_chat_context = True (mellea/core/base.py:2066), so it takes this path.
Capturing the conversation passed to apply_chat_template shows [{'role': 'user', 'content': ''}] on both HuggingFaceTB/SmolLM2-135M-Instruct and ibm-granite/granite-4.2-3b. A non-empty prompt through the same capture shows the expected user turn.
Suggested fix
Call has_user_content() in to_chat() straight after the action is appended and raise the same ValueError as the other backends. One call covers both HF paths. HF already rejects images and audio via _check_no_multimodal_blocks, so only the text and documents branches apply.
has_user_content() lives in mellea/helpers/openai_compatible_helpers.py. Once a non-OpenAI-compatible backend calls it, a backend-neutral helpers module is a better home.
Test: mirror test/backends/test_simple_context_guard.py, with SimpleContext + CBlock("") raising before generation, plus a documents pass-through case.
Scope and limits
Out of scope: WatsonxAIBackend (deprecated) and _generate_from_intrinsic, which doesn't append the action.
The token-burning stall from feat: add Granite 4.2 model defaults #1587 didn't reproduce here. Granite 4.2-3b answered the blank prompt in 74 tokens against a 1024-token cap (seed 0, one run). On HF the cost is a silently wrong answer.
What goes wrong
#1632 makes the OpenAI, LiteLLM, and Ollama backends raise
ValueErrorwhen the assembled conversation has no user content. That happens when a caller records turns on aSimpleContextand then generates with an empty action (#1597).LocalHFBackendhas no such check. The same call sends the model one empty user turn, raises nothing, and returns an answer to a prompt the caller never wrote.With
ibm-granite/granite-4.2-3bthe recorded turn is dropped and the model reasons about a message that doesn't exist:Output:
Expected: the same
ValueErrorthe other backends raise.Mechanism
Both HF chat paths, standard (
mellea/backends/huggingface.py:1782) and KV-cache (:1587), build messages throughto_chat()inmellea/backends/utils.py:74. It takesctx.view_for_generation(), which is empty forSimpleContextby design, appends the action, and hands the result toapply_chat_templatewithout checking for user content.SimpleContextinherits_is_chat_context = True(mellea/core/base.py:2066), so it takes this path.Capturing the conversation passed to
apply_chat_templateshows[{'role': 'user', 'content': ''}]on bothHuggingFaceTB/SmolLM2-135M-Instructandibm-granite/granite-4.2-3b. A non-empty prompt through the same capture shows the expected user turn.Suggested fix
Call
has_user_content()into_chat()straight after the action is appended and raise the sameValueErroras the other backends. One call covers both HF paths. HF already rejects images and audio via_check_no_multimodal_blocks, so only the text and documents branches apply.has_user_content()lives inmellea/helpers/openai_compatible_helpers.py. Once a non-OpenAI-compatible backend calls it, a backend-neutral helpers module is a better home.Test: mirror
test/backends/test_simple_context_guard.py, withSimpleContext+CBlock("")raising before generation, plus a documents pass-through case.Scope and limits
WatsonxAIBackend(deprecated) and_generate_from_intrinsic, which doesn't append the action.