Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,7 +74,8 @@ requirements R236–R249. Each row names one package owner, not a new class to s
The existing components are the migration starting points on V2; the linked issues
implement the V3 behavior. In particular, native AgentEngine execution, separate
Working/Strategic Plans, typed provider responses and typed runtime events are
future work, not capabilities delivered by this ownership change.
separate implementation work, not capabilities established by ownership names
alone. The structured provider foundation is described in [providers](providers.md).

| Requirement / subsystem | Package owner and existing components | Responsibility and authority boundary | V3 implementation |
| --- | --- | --- | --- |
Expand Down Expand Up @@ -139,7 +140,8 @@ V3 architecture is already implemented.
The inbound chat contract is
`AgentRuntime.chat(session, message, *, clear_tasks=False, profile=None,
memory_context=None, on_event=None, approve=None) -> str`.
Providers keep `LLMProvider.chat/chat_stream`. Small feature-owned protocols describe
Providers expose structured `LLMProvider.chat_response` and `get_capabilities`,
with checked `chat` tuple compatibility and existing text-only `chat_stream`. Small feature-owned protocols describe
memory/vector, task, learning, scheduling, history and interaction storage; event and
approval boundaries are callables. Concrete adapters are supplied at composition.
Foreground, durable, MCP and subagent actions share the executor and produce the same
Expand Down
1 change: 1 addition & 0 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,7 @@ This index covers the current product guides, developer references, and dated en
- [Reports index](reports.md) explains the dates, scope, and limits of recorded audits, validation runs, and benchmarks.
- [Modular refactor validation](modular-refactor-validation.md), [production audit](production-audit-2026-10-04.md), [shipping audit](production-shipping-audit-2026-10-04.md), and [self-learning audit](self-learning-audit-2026-10-03.md) preserve historical evidence.
- [V3 architecture ownership plan](superpowers/plans/2026-10-07-v3-issue-227.md) records the scoped implementation and verification for issue #227.
- [V3 provider contract plan](superpowers/plans/2026-10-08-v3-issue-228.md) records the structured response and compatibility migration for issue #228.
- Machine-readable audit receipts and benchmark inputs remain beside their corresponding reports under `docs/` and `docs/benchmarks/`.

The repository's [English overview](../README.md) links here. Technical guides are maintained in English; README translations provide localized entry points.
42 changes: 42 additions & 0 deletions docs/providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,48 @@ The normal provider setting has separate simple and complex model slots. The age

The registry marks a subset of providers for automatic selection and records preferred fallback providers. A fallback is not guaranteed: it still needs credentials, an available compatible model, and a supported transport. Alternate sub-agent providers use their own configured environment credentials and do not inherit a generic main-provider key. See [sub-agents](subagents.md) before using cross-provider assignments.

## Structured response contract (V3 issue #228)

`LLMProvider.chat_response(messages, model=None)` returns `ModelResponse`:
`text`, ordered `ToolCall` objects, normalized `finish_reason`, usage and explicit
provider/model/response metadata. `ToolCall` preserves a native call ID, name and
JSON-object arguments. Contract validation raises `ProviderContractError`, a
`ValueError` subclass; fallback does not replay a received malformed response.
All processing after a successful SDK return is marked as the received-response
phase, so parsing or accounting failures cannot trigger another generation. IDs omitted by a transport are generated within the response
scope and listed in `metadata["synthesized_call_ids"]`. Malformed arguments,
duplicate IDs/JSON keys and non-JSON values are rejected; calls are never rendered
as assistant text or executed by the provider boundary.

Finish reasons distinguish `final`, `tool_request`, `length`, `provider_error`,
`cancelled`, `blocked` and `unknown`. The original status remains in metadata.
Absent or unfamiliar status does not prove completion. The existing `chat()`
tuple API delegates to this contract and refuses tool requests, incomplete,
blocked, cancelled or failed responses rather than silently losing their meaning.
Ordinary text retains its existing behavior. Direct Ollama `chat()` transport
failures now raise exceptions instead of returning an error-string tuple. Legacy-only providers are bridged
with unknown finish status and an explicit `legacy` marker.

`get_capabilities(model=None)` reports flags for `native_tools`, `strict_schemas`,
`parallel_calls`, `text_streaming`, `streaming_tool_calls` and `reasoning_controls`.
These describe usable features of the shipped adapter API, not every feature an
underlying model might support. Text streaming reflects the existing implementation;
other flags remain false until the corresponding request paths are implemented.
Fallback capabilities are the conservative intersection of the configured,
model-mapped candidates. Structured responses retain the responding provider's
metadata. A successful response rejected by text conversion does not cause another
provider request.

The runtime's `_get_model_response` preserves this object under the existing
bounded-call, usage and context scopes; `_get_llm_response` is the checked text
compatibility boundary. Existing stream methods remain text-only. Tool request
serialization, native history round trips and structured stream events belong to
[#229](https://github.com/EvanProgramming/OpenKyrozen/issues/229) and
[#230](https://github.com/EvanProgramming/OpenKyrozen/issues/230). This foundation
also relates to [provider adapter issue #225](https://github.com/EvanProgramming/OpenKyrozen/issues/225).
Offline fixtures and installed SDK types validate the contract; they do not establish
live provider interoperability.

## Optional dependencies

`pip install -e '.[web]'` installs FastAPI and Uvicorn. Provider extras are `.[claude]`, `.[gemini]`, `.[perplexity]`, `.[bedrock]`, and `.[vertex]`; `.[cloud]` groups the cloud integrations. `.[browser]` installs Playwright; browser execution also needs a browser installation. `.[all]` installs the supported optional Python integrations. The official installer has its own pinned set of dependencies; see [installation](installation.md).
Expand Down
63 changes: 63 additions & 0 deletions docs/superpowers/plans/2026-10-08-v3-issue-228.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
# V3 issue #228: structured provider response contract

Issue: https://github.com/EvanProgramming/OpenKyrozen/issues/228 (R001, R002, R004, R007)
Related: https://github.com/EvanProgramming/OpenKyrozen/issues/225
Base: current main 25588796b4ed187dbde84b763042512c0e598405
Branch: Evan/v3-228-provider-contracts

## Task 1 — SDK-independent contract (RED → GREEN)

Define ToolCall (nonempty id/name, JSON object arguments), ModelResponse (text,
ordered calls, usage, finish reason, explicit metadata), ProviderCapabilities and
FinishReason in providers/models.py. Normalize final/tool/length/error/cancelled/
blocked/unknown outcomes. Never infer final from text or an absent finish status.
Preserve native IDs; synthesize response-scoped IDs only for transports omitting
IDs and identify them in metadata. Reject malformed/non-object/non-JSON arguments.
Tests must first fail for lost calls and finish reasons before adding code.

## Task 2 — adapter and compatibility migration

Add LLMProvider.chat_response() and get_capabilities(); keep chat() and text
streaming. Base chat_response bridges legacy-only providers with UNKNOWN status.
Native adapters expose full non-stream responses and chat delegates through the
shared checked compatibility conversion. That conversion refuses calls, tool
requests, length stops, errors, cancellation and blocked output. Preserve ordinary
text behavior and existing accounting exactly once, even for rejected responses.

Implement each current transport: OpenAI chat/Azure, OpenAI Responses, Anthropic,
Google/Vertex, Bedrock, Ollama, Perplexity. Do not add tool request serialization.
Capabilities represent implemented adapter request features, not guessed endpoint
or model capabilities: native tools, strict schemas, parallel/streaming tools and
reasoning controls remain false until those paths exist; text streaming reflects
the actual override. Fallback uses model-mapped conservative intersections and
preserves the responding provider metadata; conversion failure is not a reason
to replay a successful structured response through another provider.

## Task 3 — shared runtime boundary

Add _get_model_response() to the existing providers/calls boundary and bind it
on AgentRuntime. Reuse bounded calls, scoped usage, context and token accounting.
_get_llm_response retains its string API through checked conversion. Legacy duck
providers remain supported, without treating dynamically invented mock attributes
as an explicit structured API. Current stream handling remains text-only; native
streaming is explicitly unsupported here and owned by #230. Do not invoke tools.

## Task 4 — verification and delivery

Cover all transports with text-only, mixed/tool-only, multiple-call and malformed
argument fixtures; validate IDs, usage, finish status, metadata and inheritance.
Use real installed SDK response types where available to check fixture shape.
Test legacy callers/providers, fallback identity/model mapping/capabilities,
lossy conversions, runtime structured receipt, no tool execution and cost-once.
Run focused tests on supported Python 3.12/3.13, full make test, make check, lint,
docs-check and offline agent/subagent acceptance. Obtain fresh independent review
and fix material findings. Sign/verify commits with GPG, create/attach one PR,
resolve valid review feedback and wait for final-head CI. Do not merge.

## Limits

No paid provider calls, new dependencies, raw SDK objects/credentials in metadata,
or changes to native request serialization (#229), registry (#232), structured
streaming (#230), or AgentEngine execution (#238). Live interoperability remains
unverified. Record this limitation in the PR rather than claiming fixture tests
are live acceptance.
1 change: 1 addition & 0 deletions openkyrozen/agent/runtime.py
Original file line number Diff line number Diff line change
Expand Up @@ -229,6 +229,7 @@ class AgentRuntime:
_bounded_provider_call,
_bounded_provider_stream,
_get_llm_response,
_get_model_response,
)

# agent.planner
Expand Down
2 changes: 2 additions & 0 deletions openkyrozen/providers/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -14,3 +14,5 @@
from openkyrozen.providers.fallback import (FallbackProvider)
from openkyrozen.providers.factory import (_PROVIDER_CLASSES, get_provider, get_fallback_provider, detect_provider, save_provider_config)
from openkyrozen.security.credentials import (_get_encryption_key, _get_fernet, encrypt_api_key, decrypt_api_key, save_provider_config_encrypted)

from openkyrozen.providers.models import ModelResponse, ToolCall, ProviderCapabilities, FinishReason, ProviderContractError
37 changes: 23 additions & 14 deletions openkyrozen/providers/anthropic.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@
import time
from typing import Any, Iterator
from openkyrozen.providers.base import LLMProvider
from openkyrozen.providers.models import received_response, ModelResponse, model_response
from openkyrozen.providers.config import ProviderConfig
from openkyrozen.providers.retry import _retry_with_backoff

Expand Down Expand Up @@ -39,6 +40,9 @@ def _prepare_messages(self, messages):
return system_prompts, claude_messages

def chat(self, messages: list[dict[str, str]], model: str | None = None) -> tuple[str, dict | None]:
return self.chat_response(messages, model).as_legacy_tuple()

def chat_response(self, messages: list[dict[str, str]], model: str | None = None) -> ModelResponse:
model = model or self.config.model_simple
system_prompts, claude_messages = self._prepare_messages(messages)
started = time.monotonic()
Expand All @@ -55,20 +59,25 @@ def _call():
return self._client.messages.create(**kwargs)

response = _retry_with_backoff(_call)
text = ""
for block in response.content:
if hasattr(block, "text"):
text += block.text
usage = getattr(response, "usage", None)
usage_dict = None
if usage is not None:
usage_dict = {
"prompt_tokens": getattr(usage, "input_tokens", 0) or 0,
"completion_tokens": getattr(usage, "output_tokens", 0) or 0,
}
usage_ledger._track_cost(self.config.provider, usage_dict, model=getattr(response, "model", None) or model,
latency_ms=round((time.monotonic() - started) * 1000))
return text.strip(), usage_dict
with received_response():
usage = getattr(response, "usage", None)
usage_dict = None
if usage is not None:
usage_dict = {
"prompt_tokens": getattr(usage, "input_tokens", 0) or 0,
"completion_tokens": getattr(usage, "output_tokens", 0) or 0,
}
usage_ledger._track_cost(self.config.provider, usage_dict, model=getattr(response, "model", None) or model,
latency_ms=round((time.monotonic() - started) * 1000))
text = ""
for block in response.content:
if hasattr(block, "text"):
text += block.text
calls = [(getattr(block, "id", None), getattr(block, "name", None), getattr(block, "input", None))
for block in response.content if getattr(block, "type", None) == "tool_use"]
return model_response(provider=self.name, model=model, actual_model=getattr(response, "model", None), text=text, usage=usage_dict,
calls=calls, response_id=getattr(response, "id", None),
raw_finish_reason=getattr(response, "stop_reason", None))

def chat_stream(self, messages: list[dict[str, str]], model: str | None = None) -> Iterator[str]:
model = model or self.config.model_simple
Expand Down
23 changes: 23 additions & 0 deletions openkyrozen/providers/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@

from typing import Iterator
from abc import ABC, abstractmethod
from inspect import getattr_static
from openkyrozen.providers.models import ModelResponse, ProviderCapabilities, ProviderContractError
from openkyrozen.providers.config import ProviderConfig

class LLMProvider(ABC):
Expand All @@ -15,6 +17,16 @@ def chat(self, messages: list[dict[str, str]], model: str | None = None) -> tupl
"""Send messages to the LLM. Returns (content, usage_dict_or_None)."""
...

def chat_response(self, messages: list[dict], model: str | None = None) -> ModelResponse:
"""Bridge providers implementing only the existing text contract."""
text, usage = self.chat(messages, model)
return ModelResponse(text=text, usage=usage,
metadata={"provider": self.name, "model": model or self.config.model_simple, "legacy": True})

def get_capabilities(self, model: str | None = None) -> ProviderCapabilities:
"""Only advertise features usable through the shipped adapter API."""
return ProviderCapabilities(text_streaming=type(self).chat_stream is not LLMProvider.chat_stream)

def chat_stream(self, messages: list[dict[str, str]], model: str | None = None) -> Iterator[str]:
"""Stream response tokens. Default: fall back to non-streaming chat()."""
text, _ = self.chat(messages, model)
Expand All @@ -23,3 +35,14 @@ def chat_stream(self, messages: list[dict[str, str]], model: str | None = None)
@property
def name(self) -> str:
return self.config.provider


def get_model_response(provider, messages, model=None) -> ModelResponse:
"""Use an explicitly supplied structured API, or bridge a legacy duck provider."""
if callable(getattr_static(provider, "chat_response", None)):
response = provider.chat_response(messages, model)
if not isinstance(response, ModelResponse):
raise ProviderContractError("chat_response must return ModelResponse")
return response
text, usage = provider.chat(messages, model)
return ModelResponse(text=text, usage=usage, metadata={"legacy": True})
21 changes: 15 additions & 6 deletions openkyrozen/providers/bedrock.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@
import time
from typing import Any, Iterator
from openkyrozen.providers.base import LLMProvider
from openkyrozen.providers.models import received_response, ModelResponse, model_response
from openkyrozen.providers.config import ProviderConfig
from openkyrozen.providers.retry import _retry_with_backoff

Expand Down Expand Up @@ -49,6 +50,9 @@ def _usage(data: dict[str, Any] | None) -> dict[str, int | None] | None:
}

def chat(self, messages: list[dict[str, str]], model: str | None = None) -> tuple[str, dict | None]:
return self.chat_response(messages, model).as_legacy_tuple()

def chat_response(self, messages: list[dict[str, str]], model: str | None = None) -> ModelResponse:
model = model or self.config.model_simple
conversation, system = self._request(messages)
started = time.monotonic()
Expand All @@ -60,12 +64,17 @@ def chat(self, messages: list[dict[str, str]], model: str | None = None) -> tupl
if system:
kwargs["system"] = system
response = _retry_with_backoff(lambda: self._client.converse(**kwargs))
content = response.get("output", {}).get("message", {}).get("content", [])
text = "".join(str(item.get("text", "")) for item in content if isinstance(item, dict))
usage = self._usage(response)
usage_ledger._track_cost(self.config.provider, usage, model=model,
latency_ms=round((time.monotonic() - started) * 1000))
return text.strip(), usage
with received_response():
usage = self._usage(response)
usage_ledger._track_cost(self.config.provider, usage, model=model,
latency_ms=round((time.monotonic() - started) * 1000))
content = response.get("output", {}).get("message", {}).get("content", [])
text = "".join(str(item.get("text", "")) for item in content if isinstance(item, dict))
calls = [(item["toolUse"].get("toolUseId"), item["toolUse"].get("name"), item["toolUse"].get("input"))
for item in content if isinstance(item, dict) and "toolUse" in item]
return model_response(provider=self.name, model=model, text=text, usage=usage,
calls=calls, response_id=response.get("ResponseMetadata", {}).get("RequestId"),
raw_finish_reason=response.get("stopReason"))

def chat_stream(self, messages: list[dict[str, str]], model: str | None = None) -> Iterator[str]:
model = model or self.config.model_simple
Expand Down
Loading
Loading