What users see
A tool that takes a list of numbers gets strings instead. Take total(nums: list[int]). The model calls it correctly with [1, 2], Mellea runs it with nums=['1', '2'], and sum(nums) raises a TypeError. The error comes from inside the user's tool, so it looks like a bug in their own code.
This happens on every backend (HF, Ollama, OpenAI, LiteLLM and WatsonX) and for any list of numbers, including optional ones: list[int], list[float], list[int] | None.
Other parameter types fail more quietly:
list[bool], nested lists such as list[list[int]], and lists of Pydantic models fail Mellea's argument validation on every call. Mellea logs a warning and hands the model's arguments to the tool unchecked.
- The model is never told what goes inside a list or a dict.
list[int], list[Point] and dict[str, int] all reach it as a bare "array" or "object". For a list of models, it has no way of knowing which fields to send.
- Constraints such as
Field(ge=0, le=10) or Field(min_length=1) never reach the model either.
Reproduction
from pydantic import BaseModel
from mellea.backends.tools import MelleaTool, validate_tool_arguments
class Point(BaseModel):
x: int
y: int
def total(nums: list[int]) -> int:
"""Sum some numbers.
Args:
nums: the numbers
"""
return sum(nums)
def count(pts: list[Point]) -> int:
"""Count points.
Args:
pts: the points
"""
return len(pts)
t = MelleaTool.from_callable(total)
print(t.as_json_tool["function"]["parameters"]["properties"]["nums"])
# {'type': 'array', 'description': 'the numbers'}
# (Pydantic's own schema for this field has "items": {"type": "integer"})
args = validate_tool_arguments(t, {"nums": [1, 2]})
print(args)
# {'nums': ['1', '2']}
print(validate_tool_arguments(MelleaTool.from_callable(count), {"pts": [{"x": 1, "y": 2}]}))
# warns "pts.0: Input should be a valid string", then returns the original args unvalidated
# {'pts': [{'x': 1, 'y': 2}]}
t.run(**args)
# TypeError: unsupported operand type(s) for +: 'int' and 'str'
Why it happens
Mellea builds each tool's JSON Schema from the function signature in convert_function_to_ollama_tool(). Pydantic produces a complete schema, but for any parameter that isn't a nested model, Mellea rebuilds the property from four keys: description, type, enum and default (mellea/backends/tools.py:1486-1496). Everything else is thrown away, including items, additionalProperties, minimum/maximum and minItems.
That stripped schema is then used twice.
- It's the tool definition every backend sends to the model (
as_json_tool, convert_tools_to_json). That's where element types and constraints drop out of the model's view.
validate_tool_arguments() builds its validator from the same schema. With no items, a list's element type falls back to string (tools.py:604, tools.py:637), and the validator sets coerce_numbers_to_str=True (tools.py:732). So [1, 2] passes validation and comes out as ['1', '2']. Booleans, nested lists and model dicts can't be coerced to strings, so they fail validation instead, and on failure the validator returns the original arguments with a warning (tools.py:796, tools.py:809).
The validator runs on every backend: from the HF path (mellea/backends/utils.py:175), from Ollama (mellea/backends/ollama.py:1393), and from extract_model_tool_requests() (mellea/helpers/openai_compatible_helpers.py:207), which the OpenAI, LiteLLM and WatsonX backends share.
Every case measured
Each row is a single-parameter tool: what the model is sent, what Pydantic produced, and what validate_tool_arguments() returns for a correct argument.
| Parameter |
Sent to the model |
Pydantic's schema also has |
Correct input → validated output |
list[int] |
{"type": "array"} |
items: {type: integer} |
[1, 2] → ['1', '2'] |
list[float] |
{"type": "array"} |
items: {type: number} |
[1.5] → ['1.5'] |
list[int] | None |
{"type": "array"} |
items on the array branch |
[1, 2] → ['1', '2'] |
list[bool] |
{"type": "array"} |
items: {type: boolean} |
fails validation, passed through unchecked |
list[list[int]] |
{"type": "array"} |
nested items |
fails validation, passed through unchecked |
list[Point] |
{"type": "array"} |
items: {$ref: Point} |
fails validation, passed through unchecked |
dict[str, int] |
{"type": "object"} |
additionalProperties: {type: integer} |
{"a": "1"} passed through as-is |
Annotated[int, Field(ge=0, le=10)] |
{"type": "integer"} |
minimum: 0, maximum: 10 |
50 accepted |
Annotated[list[int], Field(min_length=1)] |
{"type": "array"} |
items, minItems: 1 |
same as list[int] |
The dict and bounded-int results would stay the same after a fix: the validator doesn't read additionalProperties, minimum or maximum today. For those types the damage is what the model is told.
Why the tests don't catch it
The only list-parameter validation tests use list_tool(items: list[str]) (test/backends/test_tool_validation_integration.py:63). Strings are exactly what the fallback element type produces, so they pass. Nothing tests a list[int], list[float] or list[bool] parameter.
Notes for a fix
- The property model already allows these keys: it has an
items field (tools.py:982) and extra="allow" (tools.py:979). The final _recursively_inline_refs pass (tools.py:1501) already walks items (tools.py:1184), so copying items onto the rebuilt property should also get list[Model] refs inlined. For list[int] | None, items lives on the array branch of the anyOf and has to be taken from there.
- Carrying
additionalProperties and the numeric and length constraints through as well would improve what the model is told, but won't change validation (see above).
- The change alters the schema every backend sends for any tool with these parameter types, so it wants its own tests:
list[int], list[float], list[bool], list[Point] and list[int] | None through both schema generation and validation.
Related
Verified on Python 3.12 against main at f15fd89. Line numbers are for that commit.
What users see
A tool that takes a list of numbers gets strings instead. Take
total(nums: list[int]). The model calls it correctly with[1, 2], Mellea runs it withnums=['1', '2'], andsum(nums)raises aTypeError. The error comes from inside the user's tool, so it looks like a bug in their own code.This happens on every backend (HF, Ollama, OpenAI, LiteLLM and WatsonX) and for any list of numbers, including optional ones:
list[int],list[float],list[int] | None.Other parameter types fail more quietly:
list[bool], nested lists such aslist[list[int]], and lists of Pydantic models fail Mellea's argument validation on every call. Mellea logs a warning and hands the model's arguments to the tool unchecked.list[int],list[Point]anddict[str, int]all reach it as a bare "array" or "object". For a list of models, it has no way of knowing which fields to send.Field(ge=0, le=10)orField(min_length=1)never reach the model either.Reproduction
Why it happens
Mellea builds each tool's JSON Schema from the function signature in
convert_function_to_ollama_tool(). Pydantic produces a complete schema, but for any parameter that isn't a nested model, Mellea rebuilds the property from four keys:description,type,enumanddefault(mellea/backends/tools.py:1486-1496). Everything else is thrown away, includingitems,additionalProperties,minimum/maximumandminItems.That stripped schema is then used twice.
as_json_tool,convert_tools_to_json). That's where element types and constraints drop out of the model's view.validate_tool_arguments()builds its validator from the same schema. With noitems, a list's element type falls back tostring(tools.py:604,tools.py:637), and the validator setscoerce_numbers_to_str=True(tools.py:732). So[1, 2]passes validation and comes out as['1', '2']. Booleans, nested lists and model dicts can't be coerced to strings, so they fail validation instead, and on failure the validator returns the original arguments with a warning (tools.py:796,tools.py:809).The validator runs on every backend: from the HF path (
mellea/backends/utils.py:175), from Ollama (mellea/backends/ollama.py:1393), and fromextract_model_tool_requests()(mellea/helpers/openai_compatible_helpers.py:207), which the OpenAI, LiteLLM and WatsonX backends share.Every case measured
Each row is a single-parameter tool: what the model is sent, what Pydantic produced, and what
validate_tool_arguments()returns for a correct argument.list[int]{"type": "array"}items: {type: integer}[1, 2]→['1', '2']list[float]{"type": "array"}items: {type: number}[1.5]→['1.5']list[int] | None{"type": "array"}itemson the array branch[1, 2]→['1', '2']list[bool]{"type": "array"}items: {type: boolean}list[list[int]]{"type": "array"}itemslist[Point]{"type": "array"}items: {$ref: Point}dict[str, int]{"type": "object"}additionalProperties: {type: integer}{"a": "1"}passed through as-isAnnotated[int, Field(ge=0, le=10)]{"type": "integer"}minimum: 0, maximum: 1050acceptedAnnotated[list[int], Field(min_length=1)]{"type": "array"}items,minItems: 1list[int]The
dictand bounded-intresults would stay the same after a fix: the validator doesn't readadditionalProperties,minimumormaximumtoday. For those types the damage is what the model is told.Why the tests don't catch it
The only list-parameter validation tests use
list_tool(items: list[str])(test/backends/test_tool_validation_integration.py:63). Strings are exactly what the fallback element type produces, so they pass. Nothing tests alist[int],list[float]orlist[bool]parameter.Notes for a fix
itemsfield (tools.py:982) andextra="allow"(tools.py:979). The final_recursively_inline_refspass (tools.py:1501) already walksitems(tools.py:1184), so copyingitemsonto the rebuilt property should also getlist[Model]refs inlined. Forlist[int] | None,itemslives on the array branch of theanyOfand has to be taken from there.additionalPropertiesand the numeric and length constraints through as well would improve what the model is told, but won't change validation (see above).list[int],list[float],list[bool],list[Point]andlist[int] | Nonethrough both schema generation and validation.Related
*args/**kwargs. It doesn't cover any of the cases above.Verified on Python 3.12 against
mainat f15fd89. Line numbers are for that commit.