Conversation
LocalHFBackend parses tool calls with parse_tools(), which only understood JSON. From 4.2 the Granite chat template instructs the model to use an XML function format, so those calls were dropped silently and the markup reached the caller as response text. parse_tools() now extracts XML calls, returning values as raw strings, and blanks them out before the JSON pass so a JSON-looking parameter value is no longer parsed as a second call. to_tool_calls() JSON-decodes string values for object/array parameters and logs a warning when <tool_call> markup is present but nothing parses. Fixes generative-computing#1689 Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
…ol args Schemas brought in through MelleaTool.from_langchain() or from_smolagents() can carry an anyOf branch whose type is a list, such as ["array", "null"]. Collecting those into a set raised TypeError and failed the whole generation. Normalise each branch's type the same way as the top-level one. Refs generative-computing#1689 Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
Code review found three ways the XML pass could produce the wrong call: - XML markup inside a JSON call's string argument was parsed as a second call and blanked out of the argument. JSON tool-call spans are now located first and XML matches inside them are skipped. - A call missing `</function>` ran on into the next call and took its parameters. The body can no longer cross another `<function=` or a `<tool_call>` tag. - A `<function=NAME>` mentioned in text before the real call started a match there, so the wrong tool was called. Matches must now open with `<tool_call>`, as the template requires; the closing tag stays optional for truncated output. to_tool_calls() now warns whenever there are more `<tool_call>` blocks than parsed calls, so a dropped call is logged even when others parse. Refs generative-computing#1689 Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
Granite's chat template renders a Python None as the text "None", and XML tool calls carry every value as text. So "None" or "null" for an optional parameter now becomes None before validation, instead of reaching the tool as a string. Required parameters are left as sent. Validation of an explicit None for an optional parameter is fixed separately in the tool-schema work for generative-computing#1693. Refs generative-computing#1689 Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
_try_load_granite_tokenizer lived in test_huggingface_filter_options.py, which skips without torch and imports the HF backend at module level. test_utils.py and test_huggingface_thinking.py imported it from there. Move it, with the model id, into test/backends/_granite_tokenizer.py, which needs only transformers and skips cleanly without it. Refs generative-computing#1689 Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
…eview
A second three-reviewer round found:
- to_tool_calls() text-decoded every parsed call, so a JSON call with
an optional string argument of "None" reached the tool as None.
Decoding now applies only to XML-format calls, which are the ones
that carry every value as text. "None"/"null" also becomes None only
where the parameter can take it: optional, and either nullable in the
schema or with no default other than None. `title: str = "Untitled"`
keeps the text.
- An XML value quoting a whole `<tool_call>{json}</tool_call>` made the
outer call fail to match, and the quoted JSON then ran as a call. A
call body must now be parameter blocks separated by whitespace, so a
value can hold any markup except `</parameter>`.
- The dropped-call warning compared a raw tag count with the number of
parsed calls, so it fired on tags quoted in a JSON argument and
missed a drop hidden by another call. It now counts `<tool_call>`
tags that open no parsed call.
Refs generative-computing#1689
Assisted-by: Claude Code
Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
3 of 8 tasks
Comment and docstring changes only; the code is unchanged. Refs generative-computing#1689 Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
jakelorocco
reviewed
Oct 1, 2026
Lenient validate_tool_arguments returned every argument raw when any one failed, so a single bad value, including the explicit None an XML call decodes for an optional parameter, left ints and bools as strings. Keep the failing arguments as given and validate the rest. Addresses review on generative-computing#1692. Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pull Request
Issue
Fixes #1689
Description
Before: tool calling didn't work with Granite 4.2 on the local Hugging Face backend (
LocalHFBackend). The model asked for a tool, Mellea never ran it, and the raw<tool_call>markup came back as the answer text. Nothing was logged. Ollama and OpenAI-compatible servers were fine, because they parse tool calls themselves.After: Granite 4.2 tool calls on the local HF backend run the tool, with arguments converted to the tool's parameter types. A malformed or truncated tool call is logged as a warning instead of vanishing.
Models that write tool calls as JSON (Granite 4.0/4.1, Mistral and others) behave as before, with two fixes: tool-call markup quoted inside an argument no longer runs as a second call, and one invalid argument no longer stops the others being converted.
How it works
Granite 4.2's chat template asks for an XML format,
<tool_call><function=NAME><parameter=KEY>VALUE</parameter></function></tool_call>, andparse_tools()only understood JSON.parse_tools()now reads the XML format. A call must open with<tool_call>and contain only parameter blocks, so a value can hold any text and one broken call can't swallow the next.to_tool_calls()decodes list and dict values (the template writes them as JSON), and turnsNone/nullintoNonewhere the parameter can take it. Numbers and booleans go through the existing validation.<tool_call>that doesn't produce a call.Related work and limitations
[1, 2], not['1', '2']) and validation accepts an explicitNone. Without it, an explicitNonefor an optional parameter is kept but logged as a validation warning. I tested the two branches merged together.</parameter>can't be parsed; it's logged.<think>block too (see feat: implement better hugging face output parsing #1604). Validation trims whitespace from string arguments.Testing
Written test-first, and every new test failed before its fix. Two rounds of three-reviewer code review shaped the later commits. The tests cover the template's shapes, quoted markup in both directions, broken and truncated calls, the
Nonerules, the warnings, the HF backend's post-processing, and a round trip through the real granite-4.2-3b template (tokenizer only).uv run pytest test/ -m "not qualitative": 4725 passed, 20 skipped, 0 failed.ruff format,ruff checkandmypyare clean.Attribution
Adding a new component, requirement, sampling strategy, or tool?
If your PR adds or modifies one of the types below, check the matching box. A checklist of type-specific review items will be posted as a comment.
NOTE: Please ensure you have an issue that has been acknowledged by a core contributor and routed you to open a pull request against this repository. Otherwise, please open an issue before continuing with this pull request.