Repository navigation
Conversation
Add support for the OpenAI Responses API alongside the existing /v1/chat/completions endpoint. Both routes are registered on the same server and share the same user-supplied serve() function. New endpoint behaviour: - Accepts ResponseRequest (input, instructions, tools, conversation, previous_response_id, max_output_tokens, temperature, etc.) - Converts string or message-array input to the internal ChatMessage format; maps the `developer` role to `system` - Returns a typed Response object with an output[] array, output_text convenience field, and ResponseUsage (including reasoning_tokens) - Streaming emits semantic SSE events: response.created, response.in_progress, response.output_text.delta, response.output_text.done, response.function_call_arguments.done, response.completed (or response.failed on error) - Tool calls produce ResponseFunctionCall output items - background=true returns 400 (not yet supported) New models in cli/serve/models.py: ResponseRequest, ResponseInputItem, InputContent, ResponseTool, WebSearchTool, FileSearchTool, MCPTool, ContextManagementConfig, PromptCacheOptions, Response, ResponseError New helpers in mellea/helpers/openai_compatible_helpers.py: ResponseUsage, ResponseOutputItem, ResponseOutputMessage, ResponseFunctionCall, OutputTextContent, build_response_usage(), build_response_output_items() Assisted-by: IBM Bob Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Update m-serve.md to mention both endpoints in the intro, add /v1/responses to the exposed routes list, and add a "Responses API" section with curl, OpenAI SDK, and streaming examples. Fix the API Endpoints section in docs/examples/m_serve/README.md, which incorrectly listed POST /generate (not a real route). Replace it with the actual registered routes: /v1/chat/completions and /v1/responses. Assisted-by: IBM Bob Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Add client_responses.py, client_responses_streaming.py, and client_responses_tool_calling.py alongside the existing Chat Completions clients in simple/, streaming/, and tool-calling/. Each uses the same server program as its sibling client — no server changes are needed to use /v1/responses. client_responses.py: minimal smoke test using client.responses.create() and output_text. client_responses_streaming.py: toggleable streaming/non-streaming using client.responses.stream() and response.output_text.delta events. client_responses_tool_calling.py: three scenarios (weather, stock price, tool call + follow-up) reading function_call items from the output[] array. Update docs/examples/m_serve/README.md to list the new files and note that all Responses API clients share the same server program. Assisted-by: IBM Bob Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
…s API Assisted-by: IBM Bob Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Fitting the new responses paradigm, we can continue a conversation w/o passing everything back and forth. At this time, there is no store that persists a server restart. There is a time-to-live (TTL) option as a simplistic way to not hog memory forever. Store, compaction, cancel, etc... would likely be future enhancements. Assisted-by: IBM Bob Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
…letion IDs Replace uuid.uuid4().hex[:N] with secrets.choice over ascii_letters+digits for both response_id (resp_...) and completion_id (chatcmpl-...) generation. The hex approach produced only lowercase a-f plus 0-9, which does not match the mixed-case alphanumeric format used by the OpenAI API. The new IDs use the full base-62 alphabet (0-9a-zA-Z), matching the OpenAI spec. Also fix test_previous_response_id_echoed, which was passing an unknown ID and expecting it to be echoed back. The correct behaviour (404 for unknown previous_response_id) meant the endpoint returned a JSONResponse, causing an AttributeError. The test now does a proper two-step round-trip: store a response first, then reference it via previous_response_id. Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com> Assisted-by: IBM Bob
Add expires_at: int | None to the Response model. Set to created_at + TTL when store=True, None when store=False. Updates README to document in-memory-only behaviour and loss on restart. Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com> Assisted-by: IBM Bob
web_search, file_search, mcp, and code_interpreter were accepted by the schema but had no implementation, causing an internal AssertionError deep in the backend. Now _build_model_options_from_response_request raises ValueError for any non-function tool type, which the existing handler converts to a 400 with an actionable message. Remove the unused WebSearchTool, FileSearchTool, and MCPTool model classes — the type literals on ResponseTool are sufficient for parsing before rejection. Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com> Assisted-by: IBM Bob
stream_response_chunks ignored the request's include list and always emitted usage in response.completed. Pass request.include through from make_responses_endpoint and gate usage on "usage" in include. Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com> Assisted-by: IBM Bob
Response.output was typed as Union[ResponseOutputMessage, ResponseFunctionCall, ResponseOutputItem] but Pydantic was matching items against the base class first, stripping the content and role fields from serialized message items. Reorder the Union so ResponseOutputMessage and ResponseFunctionCall are listed before the base ResponseOutputItem, ensuring subclass fields are preserved on serialization. Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
stream_response_chunks() had two gaps:
1. response.completed emitted only {id, status, usage}; SDK clients
expect the full Response object. Now builds a complete Response and
serializes it as {"type": "response.completed", "response": ...}.
2. Tool call arguments were only emitted as .done with no preceding
.delta. Now emits response.function_call_arguments.delta (full
arguments as one chunk) before .done, which is spec-conformant when
the backend doesn't stream arguments incrementally.
Also adds the missing SDK envelope fields to every event:
- type field matching the event name
- sequence_number counter
- response.output_item.added before content events
- response.content_part.added before text delta/done events
Passes store, ttl, and previous_response_id into the generator so
the completed envelope has the correct expires_at and
previous_response_id fields.
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Supports structured output same as chat/completions. Follows OpenAI API spec. Does not support json_object mode yet (same as completions). Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Consistent with chat/completions implementation. Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
|
marking as draft because some fixes are coming |
Fix streaming events for OpenAI SDK spec. Fix tool call index. Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
planetf1
left a comment
There was a problem hiding this comment.
I ran this against the real OpenAI SDK (openai 2.24.0, with TestClient(app) as the SDK's http_client, so no network needed). The plain non-streaming text path works well, and /v1/chat/completions is unaffected. Beyond that, a few places in the request and stream formats differ from the Responses API, and most of #1635's SDK criteria fail as a result.
Blocking: SDK tool shapes, structured output via text.format, the stream event sequence, streamed turns not being stored, stream not reaching serve(), and instructions carrying over with previous_response_id. Should fix: reasoning-token usage, the server-side tools example, and SDK-level integration tests. The rest are non-blocking suggestions and small nits, all inline.
| """ | ||
|
|
||
| type: Literal["function", "web_search", "file_search", "mcp", "code_interpreter"] | ||
| function: FunctionDefinition | None = None |
There was a problem hiding this comment.
The OpenAI SDK sends Responses-style tools, but ResponseTool only accepts the Chat Completions shape. Reproduced with openai 2.24.0, using OpenAI(http_client=TestClient(app)) and a stub serve():
- Tools. The SDK's function tool is flat:
{"type": "function", "name", "parameters", "strict"}(openai/types/responses/function_tool_param.py). Only the nestedfunction: {...}form is modelled here, sonameandparametersare dropped andserve()receives[{"type": "function"}]. With the server-side tools example,_resolve_toolsthen hitstool_def["function"]["name"], raisesKeyError, and the client gets a 500. tool_choice. The SDK form is{"type": "function", "name": ...}. It's passed through unchanged, and the example server readstool_choice["function"]["name"], so it's ignored.- Tool results.
ResponseInputItemrequiresrole, so afunction_call_outputitem gets a 400. An SDK client can't send a tool result back.
This is the "SDK works without modifications" and "function tool calling works end-to-end" criteria in #1635. Changing the wire shape after people have built against it would be a breaking change. Could we accept the SDK shapes at the edge and normalize them into what serve() gets today? The nested form could stay as an alias so the current examples keep working.
| parallel_tool_calls: bool | None = True | ||
| context_management: ContextManagementConfig | None = None | ||
| prompt_cache_options: PromptCacheOptions | None = None | ||
| format: ResponseFormat | None = None |
There was a problem hiding this comment.
The SDK sends structured output as text={"format": {"type": "json_schema", "name", "schema", "strict"}}, flat, and Responses.create() has no format argument. Reproduced with openai 2.24.0:
client_responses_format.py:65andREADME.md:383fail withTypeError: Responses.create() got an unexpected keyword argument 'format'.- A real
text.formatrequest is accepted but ignored:serve()getsformat=None, andtextends up inmodel_optionsas a raw key.
Could this read text.format in the flat shape (keeping top-level format as an alias if you like) and keep text out of model_options?
| output_index = 0 | ||
| content_index = 0 | ||
| # Generate a message item ID for the output_text events | ||
| item_id = f"msg_{uuid.uuid4().hex[:24]}" |
There was a problem hiding this comment.
client.responses.stream() fails with IndexError: list index out of range on the first text delta. That's the call the streaming client example makes (client_responses_streaming.py:30). Reproduced with openai 2.24.0.
The SDK's stream helper builds its snapshot from response.output_item.added and response.content_part.added, then indexes snapshot.output[event.output_index] on each output_text.delta (openai/lib/streaming/responses/_responses.py:330-354). This generator never sends those two events, so the list is still empty when the first delta arrives.
The tool-call events have a related problem. The function_call_arguments.delta and .done payloads (L296-305) have no type or sequence_number. The SDK reads the data body, not the event: line, so they decode as ResponseAudioDeltaEvent with type=None and never match response.function_call_arguments.done. response.failed (L329) is the same. The stream tests check the event: line, which is why they pass.
The fix is the SDK's per-item sequence: output_item.added, content_part.added, the deltas, output_text.done, content_part.done, output_item.done, plus an output_item.added/.done pair around each function call, with type and sequence_number on every event. When there's no text, skip the message item, otherwise it shares output_index 0 with the first function call.
|
|
||
| storing = request.store is not False | ||
|
|
||
| if request.stream: |
There was a problem hiding this comment.
On the streaming path nothing is written to _response_store. The if request.stream: branch returns before the store write at L542, and stream_response_chunks only uses store to set expires_at.
So a streamed response says it's retained until expires_at, but GET /v1/responses/{id} returns 404, and so does a follow-up turn using previous_response_id. Reproduced with openai 2.24.0: after a streamed turn the entry isn't in the store, and the chained create() raises NotFoundError.
Agent clients usually stream and chain with previous_response_id, so this breaks multi-turn for them. Could the generator store the response, with history built from accumulated_text, just before it emits response.completed? Passing in a small store callback from app.py would keep the store logic in one place.
| "input", | ||
| "instructions", | ||
| "model", | ||
| "stream", |
There was a problem hiding this comment.
stream is excluded here and never mapped to ModelOption.STREAM, whereas the chat path maps it (L208). So serve() can't tell the client asked to stream, and the "one per streamed token" behaviour in the stream_response_chunks docstring (streaming.py:197) never happens.
With the streaming example server, ModelOption.STREAM is missing, so it takes the await_result=True branch (m_serve_example_streaming.py:26-41). The whole generation finishes, and then the client gets it as a single delta. Reproduced with openai 2.24.0: stream=True reached serve() with model_options={'temperature': 1.0}.
Mapping it the way the chat path does should be enough, as the generator already handles an uncomputed thunk through astream().
| ] | ||
|
|
||
| from typing import Any, Literal | ||
| from typing import Any, Literal, Union |
| Using `curl`: | ||
|
|
||
| ```bash | ||
| curl http://localhost:8000/v1/responses \ |
There was a problem hiding this comment.
The new snippets here and at L134 use port 8000, but the CLI default is 8080 (commands.py:20). The 8000 at L76 predates this PR.
|
|
||
| import pytest | ||
|
|
||
| import cli.serve.app as app_module |
There was a problem hiding this comment.
This file tests cli.serve.app, so it probably belongs in test/cli/ with the other serve tests.
| if output.value: | ||
| items.append( | ||
| ResponseOutputMessage( | ||
| id=f"msg_{uuid.uuid4().hex[:24]}", |
There was a problem hiding this comment.
This mints a fresh message id, so it won't match the item_id the stream used (streaming.py:253). Reusing one id would let a client tie the deltas to the final item.
| ), | ||
| host: str = typer.Option("0.0.0.0", help="Host to bind to"), | ||
| port: int = typer.Option(8080, help="Port to bind to"), | ||
| response_ttl: int = typer.Option( |
There was a problem hiding this comment.
--response-ttl has no minimum, so 0 or a negative value is accepted.
Pull Request
Issue
Fixes #1635
Description
Adds responses endpoint consistent with the OpenAI API and our chat/completions endpoint.
This brings the updated way of interacting which can include server-side storing of context and server-side tools.
The server-side storage is very basic in-memory with time-to-live setting. Storage that lasts across restarts and is more management options could be added later.
Testing
Attribution
Adding a new component, requirement, sampling strategy, or tool?
If your PR adds or modifies one of the types below, check the matching box. A checklist of type-specific review items will be posted as a comment.
NOTE: Please ensure you have an issue that has been acknowledged by a core contributor and routed you to open a pull request against this repository. Otherwise, please open an issue before continuing with this pull request.