Skip to content

feat: add OpenAI Responses API endpoint (/v1/responses) to m serve - #1708

Open
markstur wants to merge 23 commits into
generative-computing:mainfrom
markstur:responses
Open

markstur wants to merge 23 commits into
generative-computing:mainfrom
markstur:responses

Conversation

@markstur

@markstur markstur commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Pull Request

Issue

Fixes #1635

Description

Adds responses endpoint consistent with the OpenAI API and our chat/completions endpoint.
This brings the updated way of interacting which can include server-side storing of context and server-side tools.
The server-side storage is very basic in-memory with time-to-live setting. Storage that lasts across restarts and is more management options could be added later.

Testing

  • Tests added to the respective file if code was changed
  • New code has 100% coverage if code was added
  • Ensure existing tests and github automation passes (a maintainer will kick off the github automation when the rest of the PR is populated)

Attribution

  • AI coding assistants used

Adding a new component, requirement, sampling strategy, or tool?

If your PR adds or modifies one of the types below, check the matching box. A checklist of type-specific review items will be posted as a comment.

  • Component
  • Requirement
  • Sampling Strategy
  • Tool

NOTE: Please ensure you have an issue that has been acknowledged by a core contributor and routed you to open a pull request against this repository. Otherwise, please open an issue before continuing with this pull request.

Add support for the OpenAI Responses API alongside the existing
/v1/chat/completions endpoint. Both routes are registered on the same
server and share the same user-supplied serve() function.

New endpoint behaviour:
- Accepts ResponseRequest (input, instructions, tools, conversation,
  previous_response_id, max_output_tokens, temperature, etc.)
- Converts string or message-array input to the internal ChatMessage
  format; maps the `developer` role to `system`
- Returns a typed Response object with an output[] array, output_text
  convenience field, and ResponseUsage (including reasoning_tokens)
- Streaming emits semantic SSE events: response.created,
  response.in_progress, response.output_text.delta,
  response.output_text.done, response.function_call_arguments.done,
  response.completed (or response.failed on error)
- Tool calls produce ResponseFunctionCall output items
- background=true returns 400 (not yet supported)

New models in cli/serve/models.py:
  ResponseRequest, ResponseInputItem, InputContent, ResponseTool,
  WebSearchTool, FileSearchTool, MCPTool, ContextManagementConfig,
  PromptCacheOptions, Response, ResponseError

New helpers in mellea/helpers/openai_compatible_helpers.py:
  ResponseUsage, ResponseOutputItem, ResponseOutputMessage,
  ResponseFunctionCall, OutputTextContent, build_response_usage(),
  build_response_output_items()

Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Update m-serve.md to mention both endpoints in the intro, add
/v1/responses to the exposed routes list, and add a "Responses API"
section with curl, OpenAI SDK, and streaming examples.

Fix the API Endpoints section in docs/examples/m_serve/README.md,
which incorrectly listed POST /generate (not a real route). Replace
it with the actual registered routes: /v1/chat/completions and
/v1/responses.

Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Add client_responses.py, client_responses_streaming.py, and
client_responses_tool_calling.py alongside the existing Chat
Completions clients in simple/, streaming/, and tool-calling/.
Each uses the same server program as its sibling client — no server
changes are needed to use /v1/responses.

client_responses.py: minimal smoke test using client.responses.create()
and output_text.

client_responses_streaming.py: toggleable streaming/non-streaming using
client.responses.stream() and response.output_text.delta events.

client_responses_tool_calling.py: three scenarios (weather, stock price,
tool call + follow-up) reading function_call items from the output[]
array.

Update docs/examples/m_serve/README.md to list the new files and note
that all Responses API clients share the same server program.

Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
…s API

Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Fitting the new responses paradigm, we can continue a conversation w/o passing everything back and forth.

At this time, there is no store that persists a server restart.
There is a time-to-live (TTL) option as a simplistic way to not hog memory forever.
Store, compaction, cancel, etc... would likely be future enhancements.

Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
…letion IDs

Replace uuid.uuid4().hex[:N] with secrets.choice over ascii_letters+digits
for both response_id (resp_...) and completion_id (chatcmpl-...) generation.

The hex approach produced only lowercase a-f plus 0-9, which does not match
the mixed-case alphanumeric format used by the OpenAI API. The new IDs use
the full base-62 alphabet (0-9a-zA-Z), matching the OpenAI spec.

Also fix test_previous_response_id_echoed, which was passing an unknown ID
and expecting it to be echoed back. The correct behaviour (404 for unknown
previous_response_id) meant the endpoint returned a JSONResponse, causing
an AttributeError. The test now does a proper two-step round-trip: store a
response first, then reference it via previous_response_id.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
Add expires_at: int | None to the Response model. Set to created_at +
TTL when store=True, None when store=False. Updates README to document
in-memory-only behaviour and loss on restart.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
web_search, file_search, mcp, and code_interpreter were accepted by the
schema but had no implementation, causing an internal AssertionError deep
in the backend. Now _build_model_options_from_response_request raises
ValueError for any non-function tool type, which the existing handler
converts to a 400 with an actionable message.

Remove the unused WebSearchTool, FileSearchTool, and MCPTool model
classes — the type literals on ResponseTool are sufficient for parsing
before rejection.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
stream_response_chunks ignored the request's include list and always
emitted usage in response.completed. Pass request.include through from
make_responses_endpoint and gate usage on "usage" in include.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
Response.output was typed as Union[ResponseOutputMessage,
ResponseFunctionCall, ResponseOutputItem] but Pydantic was matching
items against the base class first, stripping the content and role
fields from serialized message items.

Reorder the Union so ResponseOutputMessage and ResponseFunctionCall
are listed before the base ResponseOutputItem, ensuring subclass
fields are preserved on serialization.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
stream_response_chunks() had two gaps:

1. response.completed emitted only {id, status, usage}; SDK clients
   expect the full Response object. Now builds a complete Response and
   serializes it as {"type": "response.completed", "response": ...}.

2. Tool call arguments were only emitted as .done with no preceding
   .delta. Now emits response.function_call_arguments.delta (full
   arguments as one chunk) before .done, which is spec-conformant when
   the backend doesn't stream arguments incrementally.

Also adds the missing SDK envelope fields to every event:
- type field matching the event name
- sequence_number counter
- response.output_item.added before content events
- response.content_part.added before text delta/done events

Passes store, ttl, and previous_response_id into the generator so
the completed envelope has the correct expires_at and
previous_response_id fields.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Supports structured output same as chat/completions.
Follows OpenAI API spec.
Does not support json_object mode yet (same as completions).

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Consistent with chat/completions implementation.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
@markstur
markstur requested a review from a team as a code owner October 5, 2026 15:59
@markstur
markstur marked this pull request as draft October 5, 2026 15:59
@markstur

markstur commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

marking as draft because some fixes are coming

Fix streaming events for OpenAI SDK spec.
Fix tool call index.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
@markstur markstur changed the title Responses feat: add OpenAI Responses API endpoint (/v1/responses) to m serve Oct 5, 2026
@github-actions github-actions Bot added the enhancement New feature or request label Oct 5, 2026
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>

@planetf1 planetf1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I ran this against the real OpenAI SDK (openai 2.24.0, with TestClient(app) as the SDK's http_client, so no network needed). The plain non-streaming text path works well, and /v1/chat/completions is unaffected. Beyond that, a few places in the request and stream formats differ from the Responses API, and most of #1635's SDK criteria fail as a result.

Blocking: SDK tool shapes, structured output via text.format, the stream event sequence, streamed turns not being stored, stream not reaching serve(), and instructions carrying over with previous_response_id. Should fix: reasoning-token usage, the server-side tools example, and SDK-level integration tests. The rest are non-blocking suggestions and small nits, all inline.

Comment thread cli/serve/models.py
"""

type: Literal["function", "web_search", "file_search", "mcp", "code_interpreter"]
function: FunctionDefinition | None = None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The OpenAI SDK sends Responses-style tools, but ResponseTool only accepts the Chat Completions shape. Reproduced with openai 2.24.0, using OpenAI(http_client=TestClient(app)) and a stub serve():

  • Tools. The SDK's function tool is flat: {"type": "function", "name", "parameters", "strict"} (openai/types/responses/function_tool_param.py). Only the nested function: {...} form is modelled here, so name and parameters are dropped and serve() receives [{"type": "function"}]. With the server-side tools example, _resolve_tools then hits tool_def["function"]["name"], raises KeyError, and the client gets a 500.
  • tool_choice. The SDK form is {"type": "function", "name": ...}. It's passed through unchanged, and the example server reads tool_choice["function"]["name"], so it's ignored.
  • Tool results. ResponseInputItem requires role, so a function_call_output item gets a 400. An SDK client can't send a tool result back.

This is the "SDK works without modifications" and "function tool calling works end-to-end" criteria in #1635. Changing the wire shape after people have built against it would be a breaking change. Could we accept the SDK shapes at the edge and normalize them into what serve() gets today? The nested form could stay as an alias so the current examples keep working.

Comment thread cli/serve/models.py
parallel_tool_calls: bool | None = True
context_management: ContextManagementConfig | None = None
prompt_cache_options: PromptCacheOptions | None = None
format: ResponseFormat | None = None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The SDK sends structured output as text={"format": {"type": "json_schema", "name", "schema", "strict"}}, flat, and Responses.create() has no format argument. Reproduced with openai 2.24.0:

  • client_responses_format.py:65 and README.md:383 fail with TypeError: Responses.create() got an unexpected keyword argument 'format'.
  • A real text.format request is accepted but ignored: serve() gets format=None, and text ends up in model_options as a raw key.

Could this read text.format in the flat shape (keeping top-level format as an alias if you like) and keep text out of model_options?

Comment thread cli/serve/streaming.py
output_index = 0
content_index = 0
# Generate a message item ID for the output_text events
item_id = f"msg_{uuid.uuid4().hex[:24]}"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

client.responses.stream() fails with IndexError: list index out of range on the first text delta. That's the call the streaming client example makes (client_responses_streaming.py:30). Reproduced with openai 2.24.0.

The SDK's stream helper builds its snapshot from response.output_item.added and response.content_part.added, then indexes snapshot.output[event.output_index] on each output_text.delta (openai/lib/streaming/responses/_responses.py:330-354). This generator never sends those two events, so the list is still empty when the first delta arrives.

The tool-call events have a related problem. The function_call_arguments.delta and .done payloads (L296-305) have no type or sequence_number. The SDK reads the data body, not the event: line, so they decode as ResponseAudioDeltaEvent with type=None and never match response.function_call_arguments.done. response.failed (L329) is the same. The stream tests check the event: line, which is why they pass.

The fix is the SDK's per-item sequence: output_item.added, content_part.added, the deltas, output_text.done, content_part.done, output_item.done, plus an output_item.added/.done pair around each function call, with type and sequence_number on every event. When there's no text, skip the message item, otherwise it shares output_index 0 with the first function call.

Comment thread cli/serve/app.py

storing = request.store is not False

if request.stream:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On the streaming path nothing is written to _response_store. The if request.stream: branch returns before the store write at L542, and stream_response_chunks only uses store to set expires_at.

So a streamed response says it's retained until expires_at, but GET /v1/responses/{id} returns 404, and so does a follow-up turn using previous_response_id. Reproduced with openai 2.24.0: after a streamed turn the entry isn't in the store, and the chained create() raises NotFoundError.

Agent clients usually stream and chain with previous_response_id, so this breaks multi-turn for them. Could the generator store the response, with history built from accumulated_text, just before it emits response.completed? Passing in a small store callback from app.py would keep the store logic in one place.

Comment thread cli/serve/app.py
"input",
"instructions",
"model",
"stream",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

stream is excluded here and never mapped to ModelOption.STREAM, whereas the chat path maps it (L208). So serve() can't tell the client asked to stream, and the "one per streamed token" behaviour in the stream_response_chunks docstring (streaming.py:197) never happens.

With the streaming example server, ModelOption.STREAM is missing, so it takes the await_result=True branch (m_serve_example_streaming.py:26-41). The whole generation finishes, and then the client gets it as a single delta. Reproduced with openai 2.24.0: stream=True reached serve() with model_options={'temperature': 1.0}.

Mapping it the way the chat path does should be enough, as the generator already handles an uncomputed thunk through astream().

Comment thread cli/serve/models.py
]

from typing import Any, Literal
from typing import Any, Literal, Union

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Union looks unused here.

Using `curl`:

```bash
curl http://localhost:8000/v1/responses \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new snippets here and at L134 use port 8000, but the CLI default is 8080 (commands.py:20). The 8000 at L76 predates this PR.


import pytest

import cli.serve.app as app_module

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This file tests cli.serve.app, so it probably belongs in test/cli/ with the other serve tests.

if output.value:
items.append(
ResponseOutputMessage(
id=f"msg_{uuid.uuid4().hex[:24]}",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This mints a fresh message id, so it won't match the item_id the stream used (streaming.py:253). Reusing one id would let a client tie the deltas to the final item.

Comment thread cli/serve/commands.py
),
host: str = typer.Option("0.0.0.0", help="Host to bind to"),
port: int = typer.Option(8080, help="Port to bind to"),
response_ttl: int = typer.Option(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

--response-ttl has no minimum, so 0 or a negative value is accepted.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat: add OpenAI Responses API endpoint (/v1/responses) to m serve

2 participants