Skip to content

bug(tools): list[int] tool arguments arrive as strings, and list/dict element types never reach the model #1693

Description

@planetf1

What users see

A tool that takes a list of numbers gets strings instead. Take total(nums: list[int]). The model calls it correctly with [1, 2], Mellea runs it with nums=['1', '2'], and sum(nums) raises a TypeError. The error comes from inside the user's tool, so it looks like a bug in their own code.

This happens on every backend (HF, Ollama, OpenAI, LiteLLM and WatsonX) and for any list of numbers, including optional ones: list[int], list[float], list[int] | None.

Other parameter types fail more quietly:

  • list[bool], nested lists such as list[list[int]], and lists of Pydantic models fail Mellea's argument validation on every call. Mellea logs a warning and hands the model's arguments to the tool unchecked.
  • The model is never told what goes inside a list or a dict. list[int], list[Point] and dict[str, int] all reach it as a bare "array" or "object". For a list of models, it has no way of knowing which fields to send.
  • Constraints such as Field(ge=0, le=10) or Field(min_length=1) never reach the model either.

Reproduction

from pydantic import BaseModel
from mellea.backends.tools import MelleaTool, validate_tool_arguments

class Point(BaseModel):
    x: int
    y: int

def total(nums: list[int]) -> int:
    """Sum some numbers.

    Args:
        nums: the numbers
    """
    return sum(nums)

def count(pts: list[Point]) -> int:
    """Count points.

    Args:
        pts: the points
    """
    return len(pts)

t = MelleaTool.from_callable(total)
print(t.as_json_tool["function"]["parameters"]["properties"]["nums"])
# {'type': 'array', 'description': 'the numbers'}
# (Pydantic's own schema for this field has "items": {"type": "integer"})

args = validate_tool_arguments(t, {"nums": [1, 2]})
print(args)
# {'nums': ['1', '2']}

print(validate_tool_arguments(MelleaTool.from_callable(count), {"pts": [{"x": 1, "y": 2}]}))
# warns "pts.0: Input should be a valid string", then returns the original args unvalidated
# {'pts': [{'x': 1, 'y': 2}]}

t.run(**args)
# TypeError: unsupported operand type(s) for +: 'int' and 'str'

Why it happens

Mellea builds each tool's JSON Schema from the function signature in convert_function_to_ollama_tool(). Pydantic produces a complete schema, but for any parameter that isn't a nested model, Mellea rebuilds the property from four keys: description, type, enum and default (mellea/backends/tools.py:1486-1496). Everything else is thrown away, including items, additionalProperties, minimum/maximum and minItems.

That stripped schema is then used twice.

  1. It's the tool definition every backend sends to the model (as_json_tool, convert_tools_to_json). That's where element types and constraints drop out of the model's view.
  2. validate_tool_arguments() builds its validator from the same schema. With no items, a list's element type falls back to string (tools.py:604, tools.py:637), and the validator sets coerce_numbers_to_str=True (tools.py:732). So [1, 2] passes validation and comes out as ['1', '2']. Booleans, nested lists and model dicts can't be coerced to strings, so they fail validation instead, and on failure the validator returns the original arguments with a warning (tools.py:796, tools.py:809).

The validator runs on every backend: from the HF path (mellea/backends/utils.py:175), from Ollama (mellea/backends/ollama.py:1393), and from extract_model_tool_requests() (mellea/helpers/openai_compatible_helpers.py:207), which the OpenAI, LiteLLM and WatsonX backends share.

Every case measured

Each row is a single-parameter tool: what the model is sent, what Pydantic produced, and what validate_tool_arguments() returns for a correct argument.

Parameter Sent to the model Pydantic's schema also has Correct input → validated output
list[int] {"type": "array"} items: {type: integer} [1, 2] → ['1', '2']
list[float] {"type": "array"} items: {type: number} [1.5] → ['1.5']
list[int] | None {"type": "array"} items on the array branch [1, 2] → ['1', '2']
list[bool] {"type": "array"} items: {type: boolean} fails validation, passed through unchecked
list[list[int]] {"type": "array"} nested items fails validation, passed through unchecked
list[Point] {"type": "array"} items: {$ref: Point} fails validation, passed through unchecked
dict[str, int] {"type": "object"} additionalProperties: {type: integer} {"a": "1"} passed through as-is
Annotated[int, Field(ge=0, le=10)] {"type": "integer"} minimum: 0, maximum: 10 50 accepted
Annotated[list[int], Field(min_length=1)] {"type": "array"} items, minItems: 1 same as list[int]

The dict and bounded-int results would stay the same after a fix: the validator doesn't read additionalProperties, minimum or maximum today. For those types the damage is what the model is told.

Why the tests don't catch it

The only list-parameter validation tests use list_tool(items: list[str]) (test/backends/test_tool_validation_integration.py:63). Strings are exactly what the fallback element type produces, so they pass. Nothing tests a list[int], list[float] or list[bool] parameter.

Notes for a fix

  • The property model already allows these keys: it has an items field (tools.py:982) and extra="allow" (tools.py:979). The final _recursively_inline_refs pass (tools.py:1501) already walks items (tools.py:1184), so copying items onto the rebuilt property should also get list[Model] refs inlined. For list[int] | None, items lives on the array branch of the anyOf and has to be taken from there.
  • Carrying additionalProperties and the numeric and length constraints through as well would improve what the model is told, but won't change validation (see above).
  • The change alters the schema every backend sends for any tool with these parameter types, so it wants its own tests: list[int], list[float], list[bool], list[Point] and list[int] | None through both schema generation and validation.

Related

Verified on Python 3.12 against main at f15fd89. Line numbers are for that commit.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area/toolsTool framework, Bash/Python tools, tool call lifecyclebugSomething isn't workingp1High: important bugs (workaround exists) or high-value core features. Do soon, not on fire.triage/acceptedIssues which should be fixed (post-triage)

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions