You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
On the Ollama backend, which start_session() uses by default, a tool parameter typed as a Pydantic model reaches the model with none of its fields. The model is told only that the parameter is an object (or, for Model | None, nothing but its description), so it has to guess the field names and nesting.
When it guesses wrong, the call fails validation and lenient mode hands the guessed arguments to the tool anyway. The failure surfaces when the tool runs (a TypeError for unexpected keyword arguments, or a dict missing the fields the tool reads), so it looks like a bug in the user's code. Whether a given tool works depends on the model and the prompt, not on anything the user controls. The same tool works through OpenAIBackend on the same Ollama server. A reproducer and fix checks are at the end.
Measured impact
Local Ollama 0.34.4, ollama Python client 0.6.1, prompt "Book a double room for Alice for 3 nights.", 5 samples per cell at temperature 0.7, counting calls whose booking argument validates as a Booking.
Controlled A/B, raw /api/chat requests identical except for the tool schema:
Model
Schema as the client sends it
Full schema
granite4.1:3b
0/5
5/5
granite4.2:3b
0/5
5/5
The failed calls used guest and duration instead of guest_name and nights.
Through mellea (OllamaModelBackend vs OpenAIBackend on Ollama's /v1 endpoint, same models):
Model
Parameter
Native backend
OpenAIBackend → /v1
granite4.1:3b
Booking
0/5 (arguments sent flat, without the booking wrapper)
5/5
granite4.2:3b
Booking
5/5
5/5
granite4.1:3b
Booking | None
5/5
5/5
granite4.2:3b
Booking | None
5/5
5/5
So the native backend works only when the model happens to guess the field names, which these names make fairly easy. We'd expect names that can't be inferred from the request to fail more often, but that wasn't measured.
Cause
Mellea builds the full schema. The Ollama backend passes it to the ollama client as a dict, and the client re-validates it into its own Tool model. That model's Property declares only type, items, description and enum and ignores everything else, so properties, required and anyOf are dropped before the request (0.6.1, and unchanged in 0.6.3). The Ollama server itself accepts all three (api.ToolProperty in 0.34.4).
What survives, per parameter shape:
list[Model]: intact, because the client leaves items untyped.
Model: {"type": "object", "description": ...}, no fields.
Model | None, unions and discriminated unions: only description, since the type lives in anyOf.
dict[str, int] value types and Field constraints (minimum, maxLength, ...): dropped, but the Ollama server doesn't support those either, so they can't be fixed on this path.
fix: preserve nested tool properties ollama/ollama-python#725 (open, no review since 2026-09-05) adds properties only. required and anyOf would still be dropped, so Model | None and unions stay broken after it merges. A dependency bump alone won't fix this.
Mitigation
Today, with no code change: use OpenAIBackend against Ollama's OpenAI-compatible endpoint, which sends the schema untouched. That's 5/5 in every configuration above.
Document it. Add a note to docs/docs/integrations/ollama.md and the tools docs saying that Pydantic-model tool parameters need the /v1 route on Ollama. The page already documents that route.
Warn at runtime. Have OllamaModelBackend log a warning, once per tool, when a parameter carries properties, required or anyOf that the client will drop, pointing at the workaround.
Fix upstream, then bump. Get required and anyOf added alongside feat: citation requirement #725's properties (a small change, three fields plus tests), then raise mellea's ollama floor to the first release carrying it and add a regression test on the outgoing request.
Change the default. Route the default Ollama path through the OpenAI-compatible endpoint. This fixes it without upstream, but changes behaviour for every Ollama user and needs its own assessment of what the native backend does that /v1 doesn't.
A patch inside the native backend is possible: pre-building the client's Tool objects with model_construct gets the full schema onto the wire. It relies on Pydantic's fallback serialization and logs a serialization warning on every request, so it isn't recommended.
Reproducing and testing a fix
Offline check (no model needed). This is what the ollama client turns mellea's schema into, and it's the assertion a regression test should make:
fromtypingimportLiteralfrompydanticimportBaseModelfromollamaimportToolfrommellea.backends.toolsimportMelleaToolclassBooking(BaseModel):
guest_name: strnights: introom_type: Literal["single", "double", "suite"]
defbook(booking: Booking, backup: Booking|None=None, extras: list[Booking] |None=None) ->str:
"""Book a hotel room. Args: booking: the booking details backup: a fallback booking extras: additional bookings """return"ok"spec=MelleaTool.from_callable(book).as_json_toolsent=Tool.model_validate(spec).model_dump(exclude_none=True)["function"]["parameters"]["properties"]
fornamein ("booking", "backup", "extras"):
print(name, "->", sent[name])
A fix passes when booking keeps its properties and required, and backup keeps its anyOf branches, as extras already does.
Live check (needs ollama serve and ollama pull granite4.1:3b). This runs the same tool through both mellea backends against the same server:
fromtypingimportLiteralfrompydanticimportBaseModel, ValidationErrorfrommelleaimportMelleaSessionfrommellea.backendsimportModelOptionfrommellea.backends.ollamaimportOllamaModelBackendfrommellea.backends.openaiimportOpenAIBackendfrommellea.backends.toolsimportMelleaToolfrommellea.stdlib.componentsimportMessagefrommellea.stdlib.contextimportChatContextMODEL="granite4.1:3b"classBooking(BaseModel):
guest_name: strnights: introom_type: Literal["single", "double", "suite"]
defbook(booking: Booking) ->str:
"""Book a hotel room. Args: booking: the booking details """return"ok"tool=MelleaTool.from_callable(book)
backends= {
"native": lambda: OllamaModelBackend(model_id=MODEL),
"openai /v1": lambda: OpenAIBackend(model_id=MODEL, base_url="http://localhost:11434/v1", api_key="ollama"),
}
forlabel, makeinbackends.items():
ok=0forseedinrange(5):
m=MelleaSession(make(), ctx=ChatContext())
out=m.act(
Message("user", "Book a double room for Alice for 3 nights."),
model_options={ModelOption.TOOLS: [tool], ModelOption.SEED: seed, ModelOption.TEMPERATURE: 0.7},
tool_calls=True,
)
call=next((cforcinout.tool_callsor [] ifc.name=="book"), None)
try:
Booking.model_validate(call.args["booking"])
ok+=1except (AttributeError, KeyError, TypeError, ValidationError):
print(f" {label} seed {seed}: tool would receive {call.argsifcallelseNone}")
print(f"{label}: {ok}/5 valid")
Today it prints native: 0/5 valid, with each failure showing the tool would receive {'guest_name': 'Alice', 'nights': 3, 'room_type': 'double'} (no booking wrapper), and openai /v1: 5/5 valid. A fix passes when native matches openai /v1. Sampling makes this a smoke test, so keep the offline check as the regression gate.
Related
Found while working on #1693 / #1694, which make list, dict and constraint information reach the model on every backend. On Ollama that fix works for list[...] element types, but this issue keeps model-typed parameters from benefiting.
What users see
On the Ollama backend, which
start_session()uses by default, a tool parameter typed as a Pydantic model reaches the model with none of its fields. The model is told only that the parameter is anobject(or, forModel | None, nothing but its description), so it has to guess the field names and nesting.When it guesses wrong, the call fails validation and lenient mode hands the guessed arguments to the tool anyway. The failure surfaces when the tool runs (a
TypeErrorfor unexpected keyword arguments, or a dict missing the fields the tool reads), so it looks like a bug in the user's code. Whether a given tool works depends on the model and the prompt, not on anything the user controls. The same tool works throughOpenAIBackendon the same Ollama server. A reproducer and fix checks are at the end.Measured impact
Local Ollama 0.34.4,
ollamaPython client 0.6.1, prompt "Book a double room for Alice for 3 nights.", 5 samples per cell at temperature 0.7, counting calls whosebookingargument validates as aBooking.Controlled A/B, raw
/api/chatrequests identical except for the tool schema:granite4.1:3bgranite4.2:3bThe failed calls used
guestanddurationinstead ofguest_nameandnights.Through mellea (
OllamaModelBackendvsOpenAIBackendon Ollama's/v1endpoint, same models):OpenAIBackend→/v1granite4.1:3bBookingbookingwrapper)granite4.2:3bBookinggranite4.1:3bBooking | Nonegranite4.2:3bBooking | NoneSo the native backend works only when the model happens to guess the field names, which these names make fairly easy. We'd expect names that can't be inferred from the request to fail more often, but that wasn't measured.
Cause
Mellea builds the full schema. The Ollama backend passes it to the
ollamaclient as a dict, and the client re-validates it into its ownToolmodel. That model'sPropertydeclares onlytype,items,descriptionandenumand ignores everything else, soproperties,requiredandanyOfare dropped before the request (0.6.1, and unchanged in 0.6.3). The Ollama server itself accepts all three (api.ToolPropertyin 0.34.4).What survives, per parameter shape:
list[Model]: intact, because the client leavesitemsuntyped.Model:{"type": "object", "description": ...}, no fields.Model | None, unions and discriminated unions: onlydescription, since the type lives inanyOf.dict[str, int]value types andFieldconstraints (minimum,maxLength, ...): dropped, but the Ollama server doesn't support those either, so they can't be fixed on this path.Upstream
properties.propertiesonly.requiredandanyOfwould still be dropped, soModel | Noneand unions stay broken after it merges. A dependency bump alone won't fix this.Mitigation
Today, with no code change: use
OpenAIBackendagainst Ollama's OpenAI-compatible endpoint, which sends the schema untouched. That's 5/5 in every configuration above.Options for mellea, roughly in order of cost:
docs/docs/integrations/ollama.mdand the tools docs saying that Pydantic-model tool parameters need the/v1route on Ollama. The page already documents that route.OllamaModelBackendlog a warning, once per tool, when a parameter carriesproperties,requiredoranyOfthat the client will drop, pointing at the workaround.requiredandanyOfadded alongside feat: citation requirement #725'sproperties(a small change, three fields plus tests), then raise mellea'sollamafloor to the first release carrying it and add a regression test on the outgoing request./v1doesn't.A patch inside the native backend is possible: pre-building the client's
Toolobjects withmodel_constructgets the full schema onto the wire. It relies on Pydantic's fallback serialization and logs a serialization warning on every request, so it isn't recommended.Reproducing and testing a fix
Offline check (no model needed). This is what the
ollamaclient turns mellea's schema into, and it's the assertion a regression test should make:Today it prints:
A fix passes when
bookingkeeps itspropertiesandrequired, andbackupkeeps itsanyOfbranches, asextrasalready does.Live check (needs
ollama serveandollama pull granite4.1:3b). This runs the same tool through both mellea backends against the same server:Today it prints
native: 0/5 valid, with each failure showing the tool would receive{'guest_name': 'Alice', 'nights': 3, 'room_type': 'double'}(nobookingwrapper), andopenai /v1: 5/5 valid. A fix passes whennativematchesopenai /v1. Sampling makes this a smoke test, so keep the offline check as the regression gate.Related
Found while working on #1693 / #1694, which make list, dict and constraint information reach the model on every backend. On Ollama that fix works for
list[...]element types, but this issue keeps model-typed parameters from benefiting.