Skip to content

bug(ollama): Pydantic-model tool parameters reach the model without their fields #1696

Description

@planetf1

What users see

On the Ollama backend, which start_session() uses by default, a tool parameter typed as a Pydantic model reaches the model with none of its fields. The model is told only that the parameter is an object (or, for Model | None, nothing but its description), so it has to guess the field names and nesting.

When it guesses wrong, the call fails validation and lenient mode hands the guessed arguments to the tool anyway. The failure surfaces when the tool runs (a TypeError for unexpected keyword arguments, or a dict missing the fields the tool reads), so it looks like a bug in the user's code. Whether a given tool works depends on the model and the prompt, not on anything the user controls. The same tool works through OpenAIBackend on the same Ollama server. A reproducer and fix checks are at the end.

Measured impact

Local Ollama 0.34.4, ollama Python client 0.6.1, prompt "Book a double room for Alice for 3 nights.", 5 samples per cell at temperature 0.7, counting calls whose booking argument validates as a Booking.

Controlled A/B, raw /api/chat requests identical except for the tool schema:

Model Schema as the client sends it Full schema
granite4.1:3b 0/5 5/5
granite4.2:3b 0/5 5/5

The failed calls used guest and duration instead of guest_name and nights.

Through mellea (OllamaModelBackend vs OpenAIBackend on Ollama's /v1 endpoint, same models):

Model Parameter Native backend OpenAIBackend → /v1
granite4.1:3b Booking 0/5 (arguments sent flat, without the booking wrapper) 5/5
granite4.2:3b Booking 5/5 5/5
granite4.1:3b Booking | None 5/5 5/5
granite4.2:3b Booking | None 5/5 5/5

So the native backend works only when the model happens to guess the field names, which these names make fairly easy. We'd expect names that can't be inferred from the request to fail more often, but that wasn't measured.

Cause

Mellea builds the full schema. The Ollama backend passes it to the ollama client as a dict, and the client re-validates it into its own Tool model. That model's Property declares only type, items, description and enum and ignores everything else, so properties, required and anyOf are dropped before the request (0.6.1, and unchanged in 0.6.3). The Ollama server itself accepts all three (api.ToolProperty in 0.34.4).

What survives, per parameter shape:

  • list[Model]: intact, because the client leaves items untyped.
  • Model: {"type": "object", "description": ...}, no fields.
  • Model | None, unions and discriminated unions: only description, since the type lives in anyOf.
  • dict[str, int] value types and Field constraints (minimum, maxLength, ...): dropped, but the Ollama server doesn't support those either, so they can't be fixed on this path.

Upstream

Mitigation

Today, with no code change: use OpenAIBackend against Ollama's OpenAI-compatible endpoint, which sends the schema untouched. That's 5/5 in every configuration above.

from mellea.backends.openai import OpenAIBackend

backend = OpenAIBackend(model_id="granite4.1:3b", base_url="http://localhost:11434/v1", api_key="ollama")

Options for mellea, roughly in order of cost:

  1. Document it. Add a note to docs/docs/integrations/ollama.md and the tools docs saying that Pydantic-model tool parameters need the /v1 route on Ollama. The page already documents that route.
  2. Warn at runtime. Have OllamaModelBackend log a warning, once per tool, when a parameter carries properties, required or anyOf that the client will drop, pointing at the workaround.
  3. Fix upstream, then bump. Get required and anyOf added alongside feat: citation requirement #725's properties (a small change, three fields plus tests), then raise mellea's ollama floor to the first release carrying it and add a regression test on the outgoing request.
  4. Change the default. Route the default Ollama path through the OpenAI-compatible endpoint. This fixes it without upstream, but changes behaviour for every Ollama user and needs its own assessment of what the native backend does that /v1 doesn't.

A patch inside the native backend is possible: pre-building the client's Tool objects with model_construct gets the full schema onto the wire. It relies on Pydantic's fallback serialization and logs a serialization warning on every request, so it isn't recommended.

Reproducing and testing a fix

Offline check (no model needed). This is what the ollama client turns mellea's schema into, and it's the assertion a regression test should make:

from typing import Literal
from pydantic import BaseModel
from ollama import Tool
from mellea.backends.tools import MelleaTool

class Booking(BaseModel):
    guest_name: str
    nights: int
    room_type: Literal["single", "double", "suite"]

def book(booking: Booking, backup: Booking | None = None, extras: list[Booking] | None = None) -> str:
    """Book a hotel room.

    Args:
        booking: the booking details
        backup: a fallback booking
        extras: additional bookings
    """
    return "ok"

spec = MelleaTool.from_callable(book).as_json_tool
sent = Tool.model_validate(spec).model_dump(exclude_none=True)["function"]["parameters"]["properties"]
for name in ("booking", "backup", "extras"):
    print(name, "->", sent[name])

Today it prints:

booking -> {'type': 'object', 'description': 'the booking details'}
backup -> {'description': 'a fallback booking'}
extras -> {'type': 'array', 'items': {'properties': {'guest_name': ..., 'nights': ..., 'room_type': ...}, ...}}

A fix passes when booking keeps its properties and required, and backup keeps its anyOf branches, as extras already does.

Live check (needs ollama serve and ollama pull granite4.1:3b). This runs the same tool through both mellea backends against the same server:

from typing import Literal
from pydantic import BaseModel, ValidationError
from mellea import MelleaSession
from mellea.backends import ModelOption
from mellea.backends.ollama import OllamaModelBackend
from mellea.backends.openai import OpenAIBackend
from mellea.backends.tools import MelleaTool
from mellea.stdlib.components import Message
from mellea.stdlib.context import ChatContext

MODEL = "granite4.1:3b"

class Booking(BaseModel):
    guest_name: str
    nights: int
    room_type: Literal["single", "double", "suite"]

def book(booking: Booking) -> str:
    """Book a hotel room.

    Args:
        booking: the booking details
    """
    return "ok"

tool = MelleaTool.from_callable(book)
backends = {
    "native": lambda: OllamaModelBackend(model_id=MODEL),
    "openai /v1": lambda: OpenAIBackend(model_id=MODEL, base_url="http://localhost:11434/v1", api_key="ollama"),
}
for label, make in backends.items():
    ok = 0
    for seed in range(5):
        m = MelleaSession(make(), ctx=ChatContext())
        out = m.act(
            Message("user", "Book a double room for Alice for 3 nights."),
            model_options={ModelOption.TOOLS: [tool], ModelOption.SEED: seed, ModelOption.TEMPERATURE: 0.7},
            tool_calls=True,
        )
        call = next((c for c in out.tool_calls or [] if c.name == "book"), None)
        try:
            Booking.model_validate(call.args["booking"])
            ok += 1
        except (AttributeError, KeyError, TypeError, ValidationError):
            print(f"  {label} seed {seed}: tool would receive {call.args if call else None}")
    print(f"{label}: {ok}/5 valid")

Today it prints native: 0/5 valid, with each failure showing the tool would receive {'guest_name': 'Alice', 'nights': 3, 'room_type': 'double'} (no booking wrapper), and openai /v1: 5/5 valid. A fix passes when native matches openai /v1. Sampling makes this a smoke test, so keep the offline check as the regression gate.

Related

Found while working on #1693 / #1694, which make list, dict and constraint information reach the model on every backend. On Ollama that fix works for list[...] element types, but this issue keeps model-typed parameters from benefiting.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/backendsProvider-specific work: Ollama, HF, LiteLLM, OpenAI, Bedrock, vLLMarea/toolsTool framework, Bash/Python tools, tool call lifecyclebugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions