Skip to content

feat(tools): add AIMLAPITools for image, video, speech and transcription - #3

Open
Lookoff-AIMLAPI wants to merge 31 commits into
mainfrom
feat/aimlapi-tools
Open

Lookoff-AIMLAPI wants to merge 31 commits into
mainfrom
feat/aimlapi-tools

Conversation

@Lookoff-AIMLAPI

Copy link
Copy Markdown
Member

Provider toolkit for AI/ML API, built on the GeminiTools pattern the Agno maintainers pointed at after agno-agi#9898.

What it adds

agno.tools.models.aimlapi.AIMLAPITools with four tools, each behind an enable_* flag and configured by its own model and options:

tool endpoint default model
generate_image POST /v1/images/generations openai/gpt-image-2
generate_video POST /v2/video/generations + poll GET ?generation_id= bytedance/seedance-2-5
generate_speech POST /v1/tts openai/tts-1
transcribe_audio POST /v1/stt/create (multipart file or URL) + poll GET /v1/stt/{id} deepgram/nova-3

Generated assets are downloaded from the public CDN link (no key sent) and returned as Image / Video / Audio artifacts on a ToolResult; the transcript comes back as text.

Attribution: the same AIMLAPI_HEADERS as agno.models.aimlapi, attached only when the base URL host is exactly api.aimlapi.com.

Files

  • libs/agno/agno/tools/models/aimlapi.py
  • libs/agno/tests/unit/tools/models/test_aimlapi.py — 15 tests against an httpx mock transport
  • cookbook/91_tools/models/aimlapi_tools.py

Verification

  • Unit: pytest libs/agno/tests/unit/tools/models/ — all green; ruff + mypy clean with the repo config.
  • Live (2026-09-16, real key): all four tools direct, then image → speech → transcription through an Agent(model=AIMLAPI(...)); transcript of the generated speech came back exact. Video (seedance-2-5, 480p, 4s) took ~400s, hence the 900s default timeout.

Note for the cookbook: send_media_to_model=False on the Agent, because the chat model does not take audio/video back as input (same as models_lab_tools.py).

A provider toolkit for AI/ML API in the shape of GeminiTools: one class,
one tool per capability, each behind an enable_* flag with its own model
and options, so an agent can be handed only the media tools it needs.

- generate_image: POST /v1/images/generations, downloads the asset
- generate_video: POST /v2/video/generations, polls until completed
- generate_speech: POST /v1/tts
- transcribe_audio: POST /v1/stt/create (local file as multipart or a
  URL), polls GET /v1/stt/{id}

The attribution headers shared with agno.models.aimlapi are sent only
when the base URL host is api.aimlapi.com. Unit tests run against an
httpx mock transport that answers with the gateway's documented shapes;
all four tools were also exercised live through an Agent.
xizhuomengcontin and others added 28 commits September 16, 2026 14:43
…cstrings (agno-agi#10192)

## What

Two `base_url` docstrings name a URL the class does not use. One line
each.

## Perplexity

`libs/agno/agno/models/perplexity/perplexity.py`

| | |
| --- | --- |
| docstring (`:37`) | `https://api.perplexity.ai/chat/completions` |
| field (`:48`) | `https://api.perplexity.ai/` |

The documented value is a full endpoint, not a base URL. Since
`Perplexity` extends `OpenAILike`, the path is appended by the client,
so a reader who copies the documented string into `base_url=` gets
requests aimed at `.../chat/completions/chat/completions`.

## Vercel v0

`libs/agno/agno/models/vercel/v0.py`

| | |
| --- | --- |
| docstring (`:19`) | `https://v0.dev/chat/settings/keys` |
| field (`:27`) | `https://api.v0.dev/v1/` |

The documented value is the browser page where a user goes to *obtain*
an API key — a different host (`v0.dev` vs `api.v0.dev`) and not an API
at all. It looks like the signup link landed in the wrong docstring row;
the `api_key` attribute is the one that relates to that page.

## How these were found, and what else was checked

Mechanical comparison of every `base_url ... Defaults to <url>`
docstring in `libs/agno/agno/models/**` against the `base_url: str =
"<url>"` field in the same file. **Thirty-odd providers, exactly these
two disagree** — aimlapi, cometapi, dashscope, deepinfra, deepseek,
fireworks, inception, minimax, moonshot, n1n, nebius, neosantara,
openrouter, ramp, requesty, sambanova, siliconflow, synthorai, together,
tokenlab, trustedrouter, tuning_engines, xai, xiaomi and the rest all
match.

So this is two isolated typos rather than a pattern, and there is no
third case waiting behind them.

## Scope

Docstrings only. No field, no default, no behaviour. Nothing else in
either file is touched.

---------

Co-authored-by: Sannya Singal <32308435+sannya-singal@users.noreply.github.com>
## Summary

`TextReader.read()` and `async_read()` unconditionally decode the result
of `file.read()`. For a `StringIO` or a file opened in text mode, that
result is already a string. The resulting `AttributeError` is caught and
the reader returns no documents.

```python
from io import StringIO
from agno.knowledge.reader.text_reader import TextReader

documents = TextReader(chunk=False).read(StringIO("Agent context"))
# Before: []; after: one document containing "Agent context".
```

Decode only byte results in both paths, following the behavior recently
added to `MarkdownReader` in agno-agi#10168. Binary streams retain the
configured encoding (UTF-8 by default). The regression tests cover text
and binary streams, sync and async reads, chunking on/off, rewinding a
partially consumed input, and leaving caller-owned streams open.

## Type of change

- [x] Bug fix

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [x] Self-review completed (automated diff review by Codex)
- [ ] Documentation updated (existing file-object interface is
unchanged)
- [ ] Examples and guides updated (not applicable)
- [x] Tested in a clean virtual environment
- [x] Tests added/updated

### Duplicate and AI-Generated PR Check

- [x] Searched existing open pull requests and found no PR addressing
this issue
- [x] This PR was entirely AI-generated

## Additional Notes

Validation on Windows with Python 3.12:

- Before the fix: the TextReader suite has 8 failures and 31 passes; all
failures are the new text-stream cases.
- After the fix: 49 tests pass across `test_text_reader.py` and
`test_markdown_reader.py`.
- Repository format and validation scripts pass, including Ruff, mypy,
and the cookbook pattern check. An unrelated formatter change was
excluded, and validation was rerun on the final tree.
- `git diff --check` passes. No model calls or external services are
required by these tests.

Linux validation also passed on Python 3.10 and 3.12: [fork CI
run](https://github.com/RaycarlLei/agno/actions/runs/35069275441). Each
job checked out the exact PR commit,
`ac1d30869c35b34619f03de37cc5dd67c5321203`, passed all 49 related tests,
and passed contribution formatting and the repository validation script
(Ruff, mypy, cookbook checks).

The upstream jobs did not execute: GitHub reports an account billing
lock in the [PR Lint
run](https://github.com/agno-agi/agno/actions/runs/35044944144) and
Validation job annotations. Upstream CI remains blocked; the fork run
provides independent verification of this commit.

Implementation, diff review, and validation were performed by Codex.

Co-authored-by: Sannya Singal <32308435+sannya-singal@users.noreply.github.com>
…oolbox_demo (agno-agi#10184)

Updates urllib3 and requests in the MCP toolbox demo requirements to
address reported advisories.

Evidence:
- cookbook/91_tools/mcp/mcp_toolbox_demo/requirements.txt referenced
urllib3@2.5.0 and requests@2.32.4
- osv-scanner reported PYSEC-2026-1994, PYSEC-2026-1996,
PYSEC-2026-1998, PYSEC-2026-141 (urllib3 2.5.0) and GHSA-gc5v-m9x4-r6x2,
PYSEC-2026-2275 (requests 2.32.4) before the update
- updated versions: urllib3 2.7.0, requests 2.33.0

Validation:
- osv-scanner no longer reports any advisory for urllib3 or requests
after the update

Scope: cookbook/91_tools/mcp/mcp_toolbox_demo/requirements.txt only.

---------

Co-authored-by: katsugtgz <katsugtgz@users.noreply.github.com>
Co-authored-by: Sannya Singal <32308435+sannya-singal@users.noreply.github.com>
Co-authored-by: sannya-singal <sannyasingal@gmail.com>
## Summary

`test_tool_hook_receives_messages` fails intermittently in the release
validation suite (`test-agents-2` on agno-agi#10211), on the first attempt and
on the rerun:

```
FAILED libs/agno/tests/integration/agent/test_tool_hooks.py::test_tool_hook_receives_messages - IndexError: list index out of range
```

The agent had no instruction to use the `modulo` tool. For `"Compute 10
mod 3"` the model sometimes answers `1` directly, so `response.tools` is
empty and `response.tools[0]` raises before the hook assertions run. The
other tool tests in this file already pass an instruction (`"Always use
the mul tool to compute products."`); this test now does the same.

Measured with the real default model (`OpenAIResponses(id="gpt-5.4")`),
one `pytest` process per run:

| | Runs | Passed | Failed |
|---|---|---|---|
| Without the instruction | 30 | 28 | 2 (`IndexError`, same as CI) |
| With the instruction | 90 | 90 | 0 |

## Type of change

- [ ] Bug fix
- [ ] New feature
- [ ] Breaking change
- [ ] Improvement
- [ ] Model update
- [x] Other: flaky integration test

---

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [x] Self-review completed
- [ ] Documentation updated (comments, docstrings)
- [ ] Examples and guides: Relevant cookbook examples have been included
or updated (if applicable)
- [ ] Tested in clean environment
- [x] Tests added/updated (if applicable)

### Duplicate and AI-Generated PR Check

- [x] I have searched existing [open pull requests](../../pulls) and
confirmed that no other PR already addresses this issue
- [ ] If a similar PR exists, I have explained below why this PR is a
better approach
- [ ] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

---

## Additional Notes

- The two `divide` tests in the same file also run without an
instruction. They passed 30/30 without it and 90/90 in the runs above,
so they are unchanged.
- `./scripts/validate.sh`: ruff passes. mypy reports errors only in
`libs/agno/agno/` source files, none of which this PR touches.
…on path (agno-agi#9948)

## Summary

On the synchronous tool-execution path, a tool whose return value is
falsy but meaningful (`0`, `0.0`, `False`, `[]`, `{}`) is sent to the
model as an empty tool message. `Model.run_function_call` stringified
the result only when it was truthy:

```python
function_call_output = str(function_execution_result.result) if function_execution_result.result else ""
```

while `Model.arun_function_calls` already used
`str(function_call.result)`. So `agent.run()` and `agent.arun()` gave
the model different tool results for the same tool, and the model could
not tell "zero items" apart from "the tool produced no output". The
empty string also ends up in `RunOutput.tools[i].result`, so it is what
gets persisted to the session.

This PR drops the truthiness check so the sync branch matches the async
one. See the note below on `None`.

Issue number: agno-agi#9947

## Type of change

- [x] Bug fix
- [ ] New feature
- [ ] Breaking change
- [ ] Improvement
- [ ] Model update
- [ ] Other:

---

## Changes

- `libs/agno/agno/models/base.py`: one-line change in
`run_function_call`, `str(function_execution_result.result)`
unconditionally, matching `arun_function_calls`.
- `libs/agno/tests/unit/models/test_tool_result_falsy_values.py`: new
regression test running a tool that returns `0`, `0.0`, `False`, `[]`,
`{}`, `1`, `"text"` through both `run_function_call` and
`arun_function_calls` with a stub `Model` (no network), asserting the
tool message content on each path and that the two paths agree.

**Note on `None`:** with this change a tool returning `None` produces
`"None"` on the sync path, which is what the async path has always
produced (verified: both paths now give `'None'`). Previously sync gave
`""`. I kept the two paths identical rather than special-casing `None`.
If you would rather have `""` for `None`, the change is a one-liner (`if
result is not None else ""`) and I'm happy to apply it to both paths.

## Testing

Commands actually run, in a `uv` venv (`uv pip install -e
"libs/agno[dev]"` plus provider SDKs), on `main` @ `56eae14f`:

- Reproduction (Agent level, stub model that requests one tool call for
a `count_items() -> int` tool returning `0`):
  - Before: `run()` tool message content `''`, `arun()` `'0'`.
  - After: both `'0'`.
- New test file: 18 passed with the fix. Sabotage check with the fix
reverted and the test kept: 8 failed (all sync-path cases for `0`,
`0.0`, `False`, `[]`, `{}` and the sync/async equality cases), 10
passed. Restoring the fix: 18 passed.
- `pytest libs/agno/tests/unit/models libs/agno/tests/unit/agent
libs/agno/tests/unit/team libs/agno/tests/unit/tools/test_functions.py
libs/agno/tests/unit/tools/test_decorator.py
libs/agno/tests/unit/tools/test_toolkit.py` → 3121 passed, 12 skipped.
- `pytest libs/agno/tests/unit/workflow` → 560 passed, 7 skipped.
- `ruff format --check` and `ruff check` on the two changed files →
clean.
- `mypy --config-file libs/agno/pyproject.toml
libs/agno/agno/models/base.py` → 0 errors.

I did not run the entire `libs/agno/tests/unit` suite: some tool modules
need optional extras not installed here. CI will cover the rest.

---

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (ruff format/check and mypy with the
repo's `pyproject.toml` config, as `./scripts/validate.sh` does)
- [x] Self-review completed
- [ ] Documentation updated (comments, docstrings) — not applicable, no
public API change
- [ ] Examples and guides: Relevant cookbook examples have been included
or updated (if applicable) — not applicable
- [x] Tested in clean environment
- [x] Tests added/updated (if applicable)

### Duplicate and AI-Generated PR Check

- [x] I have searched existing [open pull requests](../../pulls) and
confirmed that no other PR already addresses this issue
- [x] If a similar PR exists, I have explained below why this PR is a
better approach
- [x] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

---

## Additional Notes

**Related open PR:** agno-agi#6633 (for agno-agi#6361) refactors both branches into a
shared `_format_non_generator_result` helper and, in passing, changes
this check to `is not None`. That PR is much broader (also touches event
handling for generators and workflow events), has been conflicting since
June, and does not call out the falsy-result behaviour or test it. This
PR is the minimal fix for the specific bug in agno-agi#9947 with a targeted
regression test; if agno-agi#6633 lands first this becomes redundant and can be
closed.

**AI disclosure:** not entirely AI-generated. I found and reproduced the
bug, decided on the fix, and ran every command listed above; an AI
assistant helped draft the test file and this description, and I
reviewed and understand every line.

---------

Co-authored-by: Harsh <74086017+harshsinha03@users.noreply.github.com>
Co-authored-by: Harsh Sinha <sinha03harsh@gmail.com>
…acy (agno-agi#9722)

## Summary
All verified at HEAD:

- **CONTRIBUTING.md** referenced four paths that don't exist in-repo:
`tools/toolkit/toolkit.py` → `tools/toolkit.py`, `cookbook/tools/` →
`cookbook/91_tools/`, `cookbook/models/` → `cookbook/90_models/`,
`07_knowledge/vector_db/` → `07_knowledge/05_integrations/vector_dbs/`
(each adjacent line already pointed at the real location).
- **libs/agno_infra/README.md** claimed "Mozilla Public License 2.0",
but the LICENSE file and `pyproject.toml` are Apache License 2.0 (as are
all sibling libs).
- **db layer**: 92 docstrings said `deserialize (Optional[bool]):
Whether to serialize ...` — inverted wording; now `Whether to
deserialize`. Mechanical change across the 11 db backends'
session/memory/eval methods.
- **github tool**: `get_pull_requests` docstring said `limit` "Defaults
to 20"; the signature default is 50.
- **gmail tools**: `search_threads`/`list_drafts` docstrings documented
`next_page_token`; the parameter is `page_token`.
- **bigquery tool**: `describe_table` docstring documented `table_name`;
the parameter is `table_id`.

---------

Co-authored-by: simpleqt <simpleqt@users.noreply.github.com>
Co-authored-by: Harsh Sinha <sinha03harsh@gmail.com>
Co-authored-by: Harsh <74086017+harshsinha03@users.noreply.github.com>
## Summary

Workflow step content checks used truthiness, so valid step results such
as a zero count or False were omitted when building downstream context.

Treat content as present unless it is None or blank text (`content is
not None and str(content).strip()`, the check `Parallel` already used
for its combined output) in:

- `StepInput.get_all_previous_content`
- `StepInput.get_step_content` for Parallel sub-steps and steps nested
inside them
- `Parallel._build_aggregated_content`
- `Step._get_deepest_content_from_step_output`, whose return type now
matches the non-string content it returns

Add regressions for zero, False, empty containers, blank text, absent
content, ordering and the Parallel paths.

No existing issue is linked; this behavior was reproduced from the
current main branch.

## Type of change

- [x] Bug fix
- [ ] New feature
- [ ] Breaking change
- [ ] Improvement
- [ ] Model update
- [ ] Other

## Validation

New tests: 7 failed and 4 passed against the unchanged base; all 11
passed with the patch. Including existing StepInput serialization tests:
15 passed. No external services or model calls.

```text
python -m pytest libs/agno/tests/unit/workflow/test_previous_content_values.py libs/agno/tests/unit/workflow/test_step_input_serialization.py -q
```

Changed-file Ruff format/check and source-file mypy pass. On Windows,
the underlying checks from scripts/format.sh and scripts/validate.sh
were run directly: repository-wide Ruff lint/import checks and the
cookbook pattern check pass; mypy passes for all 1055 agno and 21
agnoctl source files. Repository-wide format checking reports one
unchanged baseline file,
libs/agno/tests/unit/vectordb/test_elasticsearch.py (also reproduced on
pristine main); changed files pass. Bash scripts themselves and the
complete test suite were not run. Python 3.12 on Windows; pytest 9.1.1,
pytest-asyncio 1.4.0, Ruff 0.15.20, mypy 2.1.0.

## Checklist

- [x] Code complies with style guidelines (changed files checked)
- [ ] Ran full format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [ ] Contributor self-review completed
- [ ] Documentation updated (not part of this patch)
- [ ] Examples and guides updated (not part of this patch)
- [ ] Tested in a clean, dependency-consistent environment
- [x] Tests added/updated

## Duplicate and AI-Generated PR Check

- [x] Searched existing open PR titles and relevant issue/PR keywords on
2026-09-16; no matching fix was found. Recheck immediately before
submission.
- [ ] If a similar PR exists, explain why this PR is a better approach
(related but different fixes discussed below)
- [x] This patch and PR draft were entirely AI-generated

## Additional Notes

Empty strings now retain the step heading; None continues to be omitted.
This semantic choice needs maintainer review. Related PR agno-agi#9948 fixes
model tool-result formatting in models/base.py, not this workflow
helper. Other workflow content accessors are outside this patch.

An AI coding assistant discovered the behavior, drafted the patch/tests
and executed the reported local checks. The contributor approved
submission of this patch. The reported implementation and validation
work was performed by the assistant; no independent human implementation
claim is made.

---------

Co-authored-by: Harsh <74086017+harshsinha03@users.noreply.github.com>
Co-authored-by: Harsh Sinha <sinha03harsh@gmail.com>
…agi#9995)

## Summary

`CSVReader.async_read()` creates a document for each page. With
`RowChunking(skip_header=True)`, the first data row on every
continuation page is currently removed as though it were another header.
For a CSV containing one header and 1,002 data records, synchronous
reading returns 1,002 records while asynchronous reading returns 1,001.

Use the existing `start_row` metadata in `RowChunking` to skip only the
original header and retain logical row numbers across pages. Add
regression coverage for both header modes, a header-only first page,
continuation pages, and single-page input. Standalone documents and
Excel sheets retain their existing behavior.

Fixes agno-agi#9994.

## Type of change

- [x] Bug fix
- [ ] New feature
- [ ] Breaking change
- [ ] Improvement
- [ ] Model update
- [ ] Other:

---

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [x] Self-review completed
- [ ] Documentation updated (comments, docstrings)
- [ ] Examples and guides: Relevant cookbook examples have been included
or updated (if applicable)
- [x] Tested in clean environment
- [x] Tests added/updated (if applicable)

### Duplicate and AI-Generated PR Check

- [x] I have searched existing [open pull requests](../../pulls) and
confirmed that no other PR already addresses this issue
- [ ] If a similar PR exists, I have explained below why this PR is a
better approach
- [x] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

---

## Additional Notes

The implementation and regression test were generated by Codex. Review
covered the complete diff, public reader behavior, affected callers and
real local knowledge ingestion; two independent automated reviews also
completed without findings. No new API or cookbook pattern is
introduced, so documentation/example updates are not applicable.

Verified against main `d1a388446e1b44b20498c772e91303588e2734cf` on
2026-09-06 with Python 3.12.10 on macOS arm64:

- The regression failed before the fix and passed afterward. CSV,
field-labeled CSV, chunk-ID and Excel checks: 153 passed.
- Public reader reproductions pass for Path, string path and BytesIO,
including the default page size and concurrent reuse.
- Offline `Knowledge.ainsert()` with a local embedder and temporary
LanceDB stored all 2,002 records after the fix; unmodified main stored
2,000. Ordered content and row numbers were checked.
- Standalone row cleaning and IDs, per-sheet XLS/XLSX header handling,
custom async chunking, encoding, delimiters, multiline cells and page
metadata were checked.
- `./scripts/format.sh` succeeded. Its unrelated existing formatting
changes were excluded from this contribution. `./scripts/validate.sh`
passed with the documented development dependencies.
- Adding the optional LanceDB SDK exposes 11 mypy errors in unchanged
`lance_db.py`; all 11 reproduce against unmodified main with that SDK
installed.
- `./scripts/test.sh`: 19,453 passed, 202 skipped, 307 warnings; exit
code 0. On macOS, this required a process-local file descriptor limit of
4096; the initial default limit of 256 caused cascading `Too many open
files` errors. Skipped service/provider checks, including unavailable
PostgreSQL, remain unverified. Warnings include dependency deprecations,
test-key checks, Pydantic mock types and unrelated mock/tool coroutine
warnings; no warnings were suppressed. Linux CI remains to be run for
this contribution.

Related work checked on 2026-09-06: agno-agi#8025 fixed newline joining and is
already merged; agno-agi#9972 concerns text-stream decoding, agno-agi#9460 concerns
Excel memory use, and agno-agi#7125 concerns default strategy construction. None
addresses continuation-page header removal.

Classification: bug fix. Applying the repository’s `bug` label returned
an `AddLabelsToLabelable` permission error for this account. Please
apply it to this PR and the linked issue.

---------

Co-authored-by: Harsh <74086017+harshsinha03@users.noreply.github.com>
Co-authored-by: Harsh Sinha <sinha03harsh@gmail.com>
## Summary

Security follow-up to agno-agi#10210 (CodingTools `run_shell` interpreter-RCE
hardening). I audited the sibling toolkits that pair a command allowlist
or "restricted" mode with shell/interpreter execution, for the same
bypass class and the same default-on exposure.

Audit results:

- **`ShellTools` (`shell.py`)** — no change. Docstring is already honest
(warns it is an RCE sink, points to `requires_confirmation_tools`); no
false boundary. agno-agi#8854 already made the deliberate call to keep hardening
at docs + confirmation and leave the default on.
- **`LocalFileSystemTools` (`local_file_system.py`)** — no change.
`restrict_to_base_dir=True` by default and the path check
(`safe_join_relative_path`) is real and enforced (blocks absolute paths,
`..`, symlink escape). Restriction claimed and enforced.
- **`PythonTools` (`python.py`)** — **fixed (this PR).** This is the
agno-agi#10210-class problem: `run_python_code` calls `exec()` and every tool is
on by default. `safe_globals`/`safe_locals` are named like a sandbox but
restrict nothing, and `restrict_to_base_dir` only guards *file-path
arguments* — executed code can read `/etc/passwd` or dump `os.environ`
regardless. It had no honest class-level warning, unlike its siblings.
- Confirmed clean: `DaytonaTools` (remote sandbox), `CodeMode` (already
documents "not a sandbox"), `Workspace` (confirm-by-default, honest
docs).

### Changes

Mirror agno-agi#10210's honest-docs approach (no false-boundary blocklist around
`exec`, no default change):

- Added a class `.. warning::` and honest `__init__` arg docs to
`PythonTools` stating plainly that there is no sandbox, that
`safe_globals`/`safe_locals` and `restrict_to_base_dir` are not security
boundaries, and pointing to the real, already-supported mitigations
(`requires_confirmation_tools`, `exclude_tools`) and a real sandbox /
`DaytonaTools` for untrusted input.
- Strengthened the runtime warning to match.
- Added tests: one pins the documented limitation (code escapes
`base_dir` even with `restrict_to_base_dir=True`), two prove the
recommended mitigations actually work (`requires_confirmation_tools`
marks the exec tools; `exclude_tools` drops them while keeping benign
helpers).

Docs-only behavior change; no runtime behavior or defaults changed.

## Type of change

- [x] Bug fix
- [ ] New feature
- [ ] Breaking change
- [x] Improvement
- [ ] Model update
- [ ] Other:

---

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [x] Self-review completed
- [x] Documentation updated (comments, docstrings)
- [ ] Examples and guides: Relevant cookbook examples have been included
or updated (if applicable)
- [x] Tested in clean environment
- [x] Tests added/updated (if applicable)

### Duplicate and AI-Generated PR Check

- [x] I have searched existing open pull requests and confirmed that no
other PR already addresses this issue
- [ ] If a similar PR exists, I have explained below why this PR is a
better approach
- [x] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

---

## Additional Notes

Security considerations: this closes an expectations gap, not an
enforcement gap. There is no safe blocklist for `exec`, so the honest
fix is accurate docs plus the existing HITL/exclusion mechanisms and a
real sandbox for untrusted input.

Note on validation: `python.py` is mypy-clean in isolation and
`./scripts/format.sh` passes. A full `./scripts/validate.sh` run
requires `./scripts/dev_setup.sh` first (fresh worktree has no `.venv`);
the pre-existing mypy errors observed were all in the unrelated
`os/interfaces/a2a/router.py`.

Open coherence question (not addressed here): `PythonTools` stays
on-by-default. Moving both `PythonTools` and `ShellTools` to
confirm-by-default (the `Workspace` model) would be a larger, consistent
breaking change and can be a follow-up if desired.
…0210)

## Summary

`CodingTools.run_shell` had an arbitrary-code-execution hole in its
"restricted" mode. The mode blocks shell metacharacters (`;`, `|`, `&&`,
redirects, substitution) and validates the first token against a command
allowlist, but the allowlist includes interpreters (`python`, `python3`,
`pip`). An interpreter runs code that the allowlist and path checks
never see, so:

```
python3 -c "print(__import__('os').environ)"
```

executes past the restriction and dumps the full process environment
(live API keys, DB password, JWT secret). Arbitrary file reads via
`open()` work the same way. `restrict_to_base_dir=True` is the default
and `run_shell` was enabled by default, so a caller got a full RCE
against the host out of the box while the API advertised containment
("file and shell operations cannot escape base_dir").

A string filter around a general-purpose interpreter cannot be a
security boundary. This PR stops the toolkit from claiming otherwise and
closes the default-exposed path.

### Changes

- **`enable_run_shell` now defaults to `False`.** A shell is never
mounted unless the caller opts in. This is a behavior change (see
below).
- **Inline code-execution flags are blocked in restricted mode.** For
interpreter commands (`python*`), the flags `-c`, `-m`, `-e`, and stdin
(`-`) are rejected. This kills the reported one-liner. It is harm
reduction, not a boundary: an interpreter can still escape by running a
written script file, and the code and docs say so.
- **Honest documentation.** The class docstring, `restrict_to_base_dir`
/ `enable_run_shell` arg docs, the runtime warning, the cookbook README,
and `01_basic_usage.py` now state plainly that `restrict_to_base_dir` is
not a security sandbox and that `run_shell` must not be exposed to
untrusted input. For untrusted input, use a real sandbox (separate
process/container, scrubbed env, no network, read-only mount).

`restrict_to_base_dir=False` behavior is unchanged (fully unrestricted,
explicit opt-in).

## Type of change

- [x] Bug fix
- [x] Breaking change

The breaking part: agents that relied on `run_shell` being present by
default must now pass `enable_run_shell=True`. `all=True` still enables
it.

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [x] Self-review completed
- [x] Documentation updated (comments, docstrings)
- [x] Examples and guides updated (cookbook README +
`01_basic_usage.py`)
- [x] Tested in clean environment
- [x] Tests added/updated

New/updated tests in `libs/agno/tests/unit/tools/test_coding_tools.py`:
- `test_run_shell_blocks_inline_interpreter_code` — `python3 -c`, `-m`,
and versioned interpreter basenames are blocked; a plain `python3
script.py` still runs.
- `test_run_shell_inline_code_allowed_when_unrestricted` —
`restrict_to_base_dir=False` still allows it.
- `test_optional_tools_disabled_by_default` /
`test_instructions_default_no_opt_in_tools` — assert the new opt-in
default.

54 passed.

## Additional Notes

Security context: this is the RCE and secret-exfiltration path behind
the demo-environment key compromise. This PR removes the SDK-level
footgun. Removing/gating the tool in the agent-builder deployment and
rotating the compromised secrets are separate operational actions, not
part of this change.

I did not attempt to make restricted mode airtight (e.g. blocking
script-file execution or `git`/`pip` escapes), because extending the
blocklist produces false confidence without a real boundary. The honest
fix is opt-in + clear docs + a sandbox for untrusted use.

---------

Co-authored-by: Harsh <74086017+harshsinha03@users.noreply.github.com>
)

## Summary

A truncated or corrupt `.xml.gz` sitemap can raise `EOFError` or
`zlib.error` from `gzip.decompress()`. `SitemapReader` currently catches
only `OSError`, so these failures escape both `read()` and
`async_read()` and abort discovery.

Handle the other two invalid-gzip exceptions in the existing decode
fallback. An unreadable candidate then allows discovery to try the next
sitemap; an unreadable index child allows healthy sibling pages to load
with `discovery_incomplete=True`, preserving the existing signal that
missing pages must not be pruned.

The 16 regression cases exercise both public read paths, four malformed
gzip payloads, and both discovery scenarios. HTTP is mocked. Valid gzip
and plain XML continue to be covered by the existing suite.

## Type of change

- [x] Bug fix

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [x] Self-review completed (automated diff review by Codex)
- [ ] Documentation updated (no public interface changes)
- [ ] Examples and guides updated (not applicable)
- [x] Tested in clean environment
- [x] Tests added/updated

### Duplicate and AI-Generated PR Check

- [x] Searched existing open pull requests and confirmed no other PR
addresses these decompression exceptions
- [x] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

Related agno-agi#9925 hardens XML parsing after decompression. This fix is
complementary: it handles invalid gzip data before XML parsing and does
not change the XML parser or add dependencies.

## Additional Notes

- Before the fix: 12 new cases fail with `EOFError` or `zlib.error`; the
4 cases for the already-handled `OSError` pass.
- After the fix on Windows/Python 3.12: all 85 tests in
`test_sitemap_reader.py` and `test_page_fetcher.py` pass.
- Repository validation passes: Ruff, mypy (1,055 Agno and 21 agnoctl
files), and cookbook pattern checks. The format script passed; an
unrelated baseline formatting change was excluded.
- [Fresh Linux
CI](https://github.com/RaycarlLei/agno/actions/runs/35094290792) passes
on Python 3.10 and 3.12: each job verified commit
`f354987cb3216e8f6950f4424a76a35d5b815127`, passed all 85 related tests,
checked contribution formatting, and ran the repository validation
script.

Implementation, diff review, and validation were performed by Codex.

---------

Co-authored-by: Harsh <74086017+harshsinha03@users.noreply.github.com>
Co-authored-by: Harsh Sinha <sinha03harsh@gmail.com>
## Problem

The shared agnoctl client config helpers read and write
JSON/TOML-adjacent text through Python's platform-default encoding.
These files are user configuration files, so relying on the process
locale can make behavior differ across machines.

Closes agno-agi#10039

## Before / after

Before:
- `read_json_lenient()` and `read_json_strict()` used `Path.read_text()`
without an encoding
- `atomic_write_text()` opened the temporary file without an encoding
before replacing the target config

After:
- all three paths explicitly use UTF-8 text encoding
- the existing atomic replace and file-mode behavior stays unchanged
- a focused adapter test covers non-ASCII JSON content through the
helpers

## Verification

- `PYTHONPATH=D:\fable 5
files\GitHub-Regular-Session-2026-09-08\agno\libs\agnoctl python -m
pytest
libs/agnoctl/tests/test_adapters.py::test_json_helpers_use_utf8_text`

Notes:
- `uv run --project libs/agnoctl pytest ...` is currently blocked
locally because the project advertises Python `>=3.9,<4` while the dev
dependency `mypy==2.1.0` requires Python `>=3.10`.
- I also ran the full `libs/agnoctl/tests/test_adapters.py` file via
`PYTHONPATH`; 48 tests passed, and the remaining 9 failures are existing
Windows permission-mode assertions around chmod/0600 behavior rather
than this encoding change.

---------

Co-authored-by: Harsh <74086017+harshsinha03@users.noreply.github.com>
Co-authored-by: Harsh Sinha <sinha03harsh@gmail.com>
…perator injection (agno-agi#10220)

## Summary

Follow-up to agno-agi#10210. That PR hardened `CodingTools.run_shell` but left a
chaining gap: a bare `&`, a newline, and other separators still slipped
a second command past the first-token allowlist because commands ran
under `shell=True`.

Rather than keep patching a metacharacter blocklist (an endless game
against shell quoting), this fixes the root cause: **restricted mode no
longer invokes a shell.**

### What changed

- **`run_shell` runs without a shell in restricted mode.** The command
is tokenized with `shlex` and executed with `shell=False`. Chaining
(`&&`, `;`, `&`, newline), redirection, command substitution (`$(...)`,
backticks), and globbing are therefore inert by construction: they are
passed to the program as literal arguments, not interpreted. When
`restrict_to_base_dir=False`, the raw string still runs through the
system shell (full power, opt-in, supervised only).
- **`_check_command` is now a small policy check** over the tokenized
command: reject standalone control-operator tokens as unsupported (a
clear error instead of a confusing literal run), then the existing
allowlist, interpreter `-c`/`-m` block, and path-escape checks. Quoted
operators like `git commit -m "A & B"` pass untouched.
- **Deleted** the hand-rolled quote-aware operator scanner and the
operator pattern tables from the earlier iteration of this PR. Security
no longer depends on emulating shell quoting.

Verified end to end: `echo "$(touch owned.txt)"` does not create the
file, `echo hi; touch chained.txt` does not run the second command,
`python3 -c "..."` is still blocked, `echo 'A & B'` runs and prints `A &
B`, a spaced `|` is rejected with a clear message, and unrestricted mode
still chains normally.

Incorporates the finding from agno-agi#9472 (coderdailyone), who spotted the
bare-`&` gap, and addresses the review note that quoted/escaped
operators must not be rejected.

### Behavior change to note

In restricted mode, shell features stop working: pipes, redirection,
command chaining, globbing (`ls *.py`), and shell builtins. These were
already blocked or unsafe under the old metacharacter check, so this is
not a new capability loss for chaining/redirection, but **glob expansion
in restricted mode is a real change** (use the `ls`/`find`/`grep` tools,
or `restrict_to_base_dir=False` for a full shell).

## Type of change

- [x] Bug fix
- [x] Improvement

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [x] Self-review completed
- [x] Documentation updated (class/arg/method docstrings)
- [x] Tested in clean environment
- [x] Tests added/updated

56 tests pass. New tests: control operators rejected, operators proven
inert via side effects that never happen, quoted operators allowed.

## Additional Notes

Restricted mode remains harm reduction, not a security sandbox (see
agno-agi#10210): an allowlisted interpreter can still read files or reach the
network. Running without a shell removes the entire operator-injection
class and the fragile parsing it required.

---------

Co-authored-by: Harsh <74086017+harshsinha03@users.noreply.github.com>
## Description

`ChunkingStrategy.clean_text` collapsed runs of each whitespace type
independently — or so it intended. The second substitution used
`re.sub(r"\s+", " ", ...)` which matches **every** whitespace character
(including the newlines and tabs the first rule just preserved),
flattening the document to single spaces. The subsequent `\t+`, `\r+`,
`\f+`, `\v+` rules were dead code.

Both `FixedSizeChunking` and `RecursiveChunking` call `clean_text`, so
every chunked document with meaningful formatting (code, markdown) lost
its newlines and indentation before chunking.

## Fix

Replace `\s+` with ` +` so only space runs collapse, preserving newlines
and tabs:

```python
# Before (destroys structure):
cleaned = re.sub(r"\s+", " ", text)  # \n and \t → " "

# After (preserves structure):
cleaned = re.sub(r" +", " ", text)   # only spaces collapse
```

## Verification

```python
raw = "para1\n\n\npara2\n\tindented code\n   spaced"
cleaned = clean_text(raw)
# Before: "para1 para2 indented code spaced"   (all whitespace → spaces)
# After:  "para1\npara2\n\tindented code spaced" (structure preserved)
```

Fixes agno-agi#9984

---------

Co-authored-by: icn5381 <255778606+icn5381@users.noreply.github.com>
Co-authored-by: Harsh <74086017+harshsinha03@users.noreply.github.com>
Co-authored-by: Harsh Sinha <sinha03harsh@gmail.com>
# Changelog

## New Features:

- **Azure OpenAI Responses**: Added `AzureOpenAIResponses` to use the
Responses API with Azure OpenAI deployments. See
[cookbook](https://github.com/agno-agi/agno/blob/main/cookbook/90_models/azure/openai/responses.py).
- **Elasticsearch**: Added `Elasticsearch` vector database with vector,
keyword and hybrid search. See
[cookbook](https://github.com/agno-agi/agno/blob/main/cookbook/07_knowledge/09_archive/vector_dbs/elasticsearch_db.py).
- **DocumentationMarkdown**: Added a `transform` for
`Knowledge.sync_pages` that converts Mintlify and Fumadocs components
into plain Markdown. See
[cookbook](https://github.com/agno-agi/agno/blob/main/cookbook/05_agent_os/27_public_pages/documentation_markdown.py).

## Improvements:

- **MCPConfig**: Added `root_host`, `path` and `path_aliases` to serve
MCP on a dedicated hostname or a custom endpoint path.
- **MCP Server Card**: `/mcp/server-card` now returns pretty-printed
JSON.

## Bug Fixes:

- **MCP Tools**: Preserve `AudioContent` from MCP tool results as audio
artifacts.
- **Slack**: Deduplicate event retries by `event_id` instead of dropping
them.
- **Remote Agents**: `RemoteAgent.role` and `RemoteTeam.role` are now
properties that return the role instead of a bound method; use `.role`,
not `.role()`.
- **Readers**: `TextReader`, `MarkdownReader` and
`FieldLabeledCSVReader` now accept text streams.
- **PPTXReader**: Read text inside grouped shapes.
- **JSONReader**: `async_read()` now uses the chunking strategy's async
`achunk()`.
- **Chunking**: Preserve newlines and tabs when cleaning text before
chunking.
- **CSVReader**: `async_read()` keeps every data row across pages with
`RowChunking(skip_header=True)`.
- **SitemapReader**: Keep filenames like `myindex.html` intact and
continue discovery past an invalid `.xml.gz` sitemap.
- **Tools & Workflows**: Keep falsy values like `0`, `False` and `[]` in
sync tool results and `Parallel` workflow step outputs.
- **agnoctl**: `agno connect` now reads and writes client configs as
UTF-8 (agnoctl 0.2.1).
- **DynamoDB**: Serialize booleans as native `BOOL` values.

## Breaking Changes:

- **CodingTools**: `run_shell` is now opt-in (`enable_run_shell=True`),
and restricted mode runs commands without a shell.
- **PublicSurface**: With `authorization=True` and
`PublicSurface(mcp=True)`, MCP now accepts only localhost by default;
add your domain to `MCPConfig(allowed_hosts=[...])`.
…0229)

Tool docstrings that did not match their signatures, and async tool
methods whose docstrings dropped the tool's parameter descriptions:

- `libs/agno/agno/tools/workflow.py`: `async_run_workflow` documented
`input_data`/`additional_data`, but the signature takes `input:
RunWorkflowInput`; it now uses the sync twin's `input` description.
- `libs/agno/agno/tools/calcom.py`: three methods documented
`user_timezone` / `event_type_id` entries — those are toolkit config
attributes, not parameters.
- `libs/agno/agno/tools/neo4j.py`: the docstring referenced the
pre-rename API (`host`/`port`, bare `list_labels`-style flags); aligned
with the actual `__init__` parameters (`uri`, `enable_*` flags, `all`).
- `libs/agno/agno/tools/zoom.py`: `__init__` documented a `name`
parameter; the toolkit name is fixed to `zoom_tool`, and passing `name=`
raises `TypeError`.
- `libs/agno/agno/tools/adanos.py`, `minimax.py`, `studio_runner.py`,
`superserve.py`: each async method registers under the same tool name as
its sync twin, and `agent.arun` builds the tool definition from the
async method's docstring. Those docstrings were one-liners (e.g. `Async
variant of run_python_code.`), so async runs sent the model no parameter
descriptions. The 27 async methods now carry their sync twin's
docstring.
- `libs/agno/agno/tools/minimax.py`: `resolution` said MiniMax H3
supports only 2K; the video generation API accepts `768P` or `2K` for
`MiniMax-H3`.

Docstring-only.

- [x] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

---------

Co-authored-by: Harsh <74086017+harshsinha03@users.noreply.github.com>
Co-authored-by: Harsh Sinha <sinha03harsh@gmail.com>
## Summary

Adds a runnable observability cookbook for [Confident
AI](https://www.confident-ai.com/) using its OpenTelemetry-native
`confident-trace` SDK, which detects Agno automatically.

- Add `cookbook/observability/confident_ai.py`: calls `init()` once at
startup, runs a Hacker News agent (OpenAIResponses, `gpt-5.6-luna`)
twice, wraps the second run in `trace_context` to attach tags, metadata,
and a user ID, and calls `shutdown()` in a `finally` block to flush
spans.
- List the example in `cookbook/observability/README.md`.
- Record the test run in `cookbook/observability/TEST_LOG.md`.

Companion docs: agno-agi/docs#767 (Mintlify) and agno-agi/agno-docs#141
(in-house). The in-house examples page links to this cookbook on `main`.

## Type of change

- [ ] Bug fix
- [ ] New feature
- [ ] Breaking change
- [x] Improvement
- [ ] Model update
- [ ] Other:

---

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [x] Self-review completed
- [x] Documentation updated (comments, docstrings)
- [x] Examples and guides: Relevant cookbook examples have been included
or updated (if applicable)
- [x] Tested in clean environment
- [ ] Tests added/updated (if applicable)

### Duplicate and AI-Generated PR Check

- [x] I have searched existing [open pull
requests](https://github.com/agno-agi/agno/pulls) and confirmed that no
other PR already addresses this issue
- [ ] If a similar PR exists, I have explained below why this PR is a
better approach
- [x] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

---

## Additional Notes

- Verified against `confident-trace==0.1.3`. Its OpenAI integration
patches the Responses API, so `OpenAIResponses` model spans are captured
alongside the Agno agent and tool spans.
- Ran end to end with `CONFIDENT_API_KEY` and `OPENAI_API_KEY` set. Both
runs completed and span batches were accepted by the collector with HTTP
200.
- `CONFIDENT_OTEL_ENDPOINT` must match the project's region. A US
project key sent to the EU endpoint returns 401 on export, so only set
it for EU projects. The module docstring calls this out.
- `cookbook/scripts/check_cookbook_pattern.py` passes for the new file.
The two existing violations it reports in `mlflow_via_autolog.py` are
pre-existing and untouched.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
…agno-agi#10228)

Streaming parser docstrings in six model classes did not match their
signatures:

- `libs/agno/agno/models/meta/llama.py:443`
(`_parse_provider_response_delta`): documented `response_delta`,
parameter is `response`
- `libs/agno/agno/models/openai/responses.py:1291`
(`_parse_provider_response_delta`): documented `response`, parameters
are `stream_event` / `assistant_message` / `tool_use`; Returns said
`ModelResponse`, the method returns `Tuple[ModelResponse, Dict[str,
Any]]`
- `libs/agno/agno/models/openai/responses.py:1398` (`_get_metrics`):
documented `response`, parameter is `response_usage`
- `libs/agno/agno/models/cohere/chat.py:367`
(`_parse_provider_response_delta`): `tool_use` was not documented;
Returns said `ModelResponse`, the method returns `Tuple[ModelResponse,
Dict[str, Any]]`
- `libs/agno/agno/models/ollama/chat.py:407`
(`_parse_provider_response_delta`): Returns said
`Iterator[ModelResponse]`, the method returns `ModelResponse`
- `libs/agno/agno/models/anthropic/claude.py:1127`
(`_parse_provider_response_delta`): summary named
`ModelProviderResponse`, which does not exist; Returns described an
iterator
- `libs/agno/agno/models/groq/groq.py:530`
(`_parse_provider_response_delta`): Returns described an iterator

The new lines reuse the existing wording for the same parameters in
`aws/bedrock.py`, `models/base.py`, `openai/chat.py` and
`cerebras/cerebras.py`.

Docstring-only change.

- [x] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

---------

Co-authored-by: Harsh <74086017+harshsinha03@users.noreply.github.com>
Co-authored-by: Harsh Sinha <sinha03harsh@gmail.com>
…gi#10115)

## Summary

Adds MMR (Maximal Marginal Relevance) as a retrieval strategy, along
with the knowledge level pipeline it needs to work.

`Knowledge.search()` returned the vector db's results directly, so
anything that reorders results had to be implemented per adapter.
Reordering also needs more candidates than the caller asked for: a
document can only be surfaced if it was retrieved in the first place.

**Knowledge level reranking.** `Knowledge` takes an optional `reranker`
that runs after the vector db returns candidates, with a widened fetch
to give it something to choose between:
- `reranker`: applied to results before they are returned
- `candidate_multiplier`: candidates fetched per requested result
(default 5)
- `max_candidates`: ceiling on the widened fetch (default 100)
One implementation covers every adapter, including those that never
implemented reranking. A reranker configured on the vector db still runs
first; this runs on its output, and `Knowledge` warns when both are set.
**MMRReranker.** Selects results that are relevant to the query but
unlike each other, so a search returns several distinct answers instead
of one answer repeated. `lambda_mult` trades relevance against diversity
(1.0 is pure relevance, 0.0 pure
  difference).      

Selection tracks each candidate's similarity to the nearest already
selected document instead of recomputing it every iteration.

**Vector db fixes.** MMR needs an embedding on every candidate and an
embedder for the query, and neither was returned everywhere:
- Cassandra, Chroma, Couchbase, Elasticsearch, OpenSearch, Pinecone and
Upstash returned embeddings without the embedder that produced them.
They now attach it, as the other adapters already did.
- Qdrant named vector searches return a mapping of dense and sparse
vectors, which was assigned to `Document.embedding` whole. Scoring
raised `TypeError`, which `Knowledge` caught as a transient reranker
failure, so MMR silently did nothing.
- Pinecone omits vectors unless asked. `PineconeDb` takes
`return_vectors`, off by default so ordinary searches keep their current
response size.
`Reranker.arerank` now runs the sync implementation in a worker thread,
since a reranker that calls a provider would otherwise block the event
loop. Page backed knowledge routes through the reranker instead of
returning ahead of it.

## Type of change

- [x] Bug fix
- [x] New feature
- [ ] Breaking change
- [x] Improvement
- [ ] Model update
- [ ] Other:

---

## Checklist

- [ ] Code complies with style guidelines
- [ ] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [ ] Self-review completed
- [ ] Documentation updated (comments, docstrings)
- [ ] Examples and guides: Relevant cookbook examples have been included
or updated (if applicable)
- [ ] Tested in clean environment
- [ ] Tests added/updated (if applicable)

### Duplicate and AI-Generated PR Check

- [ ] I have searched existing [open pull requests](../../pulls) and
confirmed that no other PR already addresses this issue
- [ ] If a similar PR exists, I have explained below why this PR is a
better approach
- [ ] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

---

## Additional Notes

Add any important context (deployment instructions, screenshots,
security considerations, etc.)

---------
…sion rules (agno-agi#10179)

## Summary

The HTTP and WebSocket doors for workflow runs share the durable core
(validation, enqueue, register, prepare, tail), but the WebSocket door
had drifted on three admission rules that run before it. Found while
checking queue behaviour over WebSocket against HTTP after agno-agi#10169.

- **Version-pinned submissions were queued without their pin.** HTTP
refuses to queue a submission that pins a workflow version, since the
worker resolves the registry instance and a ticket cannot carry the pin;
pinned runs go in-process with the version stamped on the run. The
WebSocket gate had no such check: a `start-workflow` with a `version`
was accepted onto the queue with an empty kwargs payload, and the worker
executed whatever version was current. Verified live before the change.
The WebSocket gate now excludes pinned submissions like HTTP and they
fall through to the existing in-process path, which already stamps the
pin.
- **No session ownership check on the WebSocket write path.** HTTP calls
`assert_session_writable` before any background work because the runs
table has no ownership predicate, so an unguarded write is replayed into
the owner's history as their own turn. The WebSocket door pinned the
caller's identity to the token but let the client choose any session id,
with no check on the start path, the continue path or the shared
prepare. The start-workflow handler now applies the same guard with the
same effective identity as HTTP (the caller's resolved user id, else the
workflow's own default) and answers a refusal with an error frame. The
guard keeps its allowances: admins, unowned sessions, sessions that do
not exist yet, remote databases, and no database.
- **Session id defaulting differed.** HTTP mints a new session for a
submission that names none; WebSocket fell back to the workflow's own
`session_id` first, so every client omitting the field on a workflow
configured with one pooled into a single session. Under per-session
queueing they would all line up behind each other. WebSocket now mints
like HTTP.

One commit per change. Continue doors already matched and are untouched.

## Type of change

- [x] Bug fix
- [ ] New feature
- [ ] Breaking change
- [ ] Improvement
- [ ] Model update
- [ ] Other:

---

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [x] Self-review completed
- [x] Documentation updated (comments, docstrings)
- [ ] Examples and guides: Relevant cookbook examples have been included
or updated (if applicable)
- [x] Tested in clean environment
- [x] Tests added/updated (if applicable)

### Duplicate and AI-Generated PR Check

- [x] I have searched existing [open pull requests](../../pulls) and
confirmed that no other PR already addresses this issue
- [ ] If a similar PR exists, I have explained below why this PR is a
better approach
- [x] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

---

## Additional Notes

Tests in `tests/unit/os/test_ws_workflow_submission_parity.py` drive
`handle_workflow_via_websocket` through a durable deployment with a
stubbed queue worker and an in-memory event stream. Each of the three
behaviour tests failed before its change:

- a pinned submission is not queued, runs in-process, and carries the
pin on the run
- a caller scoped to one user is refused into a session owned by
another, with nothing queued or executed; the owner is admitted
- two submissions without a `session_id` on a workflow configured with
one get two fresh sessions

Behaviour change to be aware of on the third item: a WebSocket client
that relied on omitting `session_id` to land in the workflow's
configured session now gets a fresh session per submission, which is
what HTTP has always done. Pass `session_id` explicitly to keep a
conversation.

Independent of agno-agi#10169 (per-run tail pumps); both touch the WebSocket
handler in different regions and either merge order works.
## Summary

Fixes agno-agi#10011.

When MCPTools receives a raw ClientSession, discovery currently
registers only the first tools/list page. A tool on the second page is
missing, and include/exclude validation can incorrectly reject its name.

Collect every ListToolsResult page with the SDK's PaginatedRequestParams
before filtering or registering functions. Continue until next_cursor is
None, preserving opaque cursor values, including empty and repeated
tokens. Bound automatic discovery to 250 pages, matching the FastMCP
Client's default page budget; reaching the budget raises instead of
silently registering a partial catalog.

Copy the initial tool list so cached result objects are not extended in
place. A later-page failure or cancellation leaves the existing function
registry unchanged. The plain-list path used by FastMCP Client remains a
single call.

## Type of change

- [x] Bug fix
- [ ] New feature
- [ ] Breaking change
- [ ] Improvement
- [ ] Model update
- [ ] Other:

---

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [ ] Self-review completed
- [x] Documentation updated (comments, docstrings)
- [ ] Examples and guides: Relevant cookbook examples have been included
or updated (if applicable)
- [ ] Tested in clean environment
- [x] Tests added/updated (if applicable)

### Duplicate and AI-Generated PR Check

- [x] I have searched existing open pull requests and confirmed that no
other PR already addresses this issue
- [ ] If a similar PR exists, I have explained below why this PR is a
better approach
- [x] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

---

## Additional Notes

Prepared and verified with Codex assistance in an existing isolated
development environment. Human self-review is not asserted by the
unchecked checklist item.

Validation on Python 3.12.13, mcp 2.1.1 and fastmcp 4.0.3:

- Unmodified baseline: 104 existing MCP tests passed.
- Final regression tests against unmodified main: 13 failed, 2 passed.
- After the patch: 119 tests passed across test_mcp.py and
test_mcp_pagination.py.
- A real in-process FastMCP server uses list_page_size=1. Through its
raw ClientSession, MCPTools now discovers both tools and successfully
invokes the second-page entrypoint.
- Cases cover opaque/empty/repeated cursors, empty pages with
continuation, later-page filters, non-terminating listings, completion
on the last allowed page, later-page errors/cancellation, preservation
of the original response list, and complete list responses.
- Repository formatting and validation passed (Ruff, mypy and cookbook
pattern checks); git diff --check passed.

The [MCP pagination
specification](https://modelcontextprotocol.io/specification/2026-07-28/server/utilities/pagination)
requires opaque cursor handling and termination based on the absence of
a continuation token. Repeated token values alone do not establish that
the server stopped advancing.

This changes list collection before registration. agno-agi#10007 separately
handles ownership and removal of stale registered functions; no changes
from that PR are included here. Connection lifecycle, per-request
timeouts and tool entrypoint behavior are unchanged. Full repository
integration-suite execution and external LLM evaluation are not claimed.
## Summary

`ag-ui-protocol` 1.0.0 was released on 2026-09-17. agno allows any
version from 0.1.15 up, so CI and new installs now get 1.0.0, and `main`
has been failing since.

What fails on `main` with 1.0.0:

- Two tests in `test_agui_app.py` and one in
`test_validation_error_body.py`. The third was hidden because fail-fast
cancelled its CI shard.
- The mypy step of `style-check-agno`, with two errors in
`agui/resume.py`.

One of these is a real bug. In 1.0 the content of a tool result message
(`ToolMessage.content`) can be a list of content parts instead of a
string. The AG-UI resume code still treated it as a string. When a
paused run was answered with a list:

- a confirmation ended in `RUN_ERROR` and the tool never ran
- a frontend tool result reached the model as raw objects, the run could
not be saved, and it stayed `PAUSED`

Older versions reject list content before agno sees it, so this only
happens on 1.0.

## Changes

- `agui/resume.py`: turn the tool result into text once, before it is
used. A string is kept as is. For a list, the text parts are joined and
any other parts are dropped with a warning. It checks the part's `type`
string instead of importing the 1.0 classes, because those do not exist
on 0.1.x.
- `test_agui_hitl.py`: new tests for answers sent as content parts. One
goes through the real `/agui` route with SQLite and checks the run is
saved as `COMPLETED`.
- `test_agui_app.py` and `test_validation_error_body.py`: three tests
assumed 0.x shapes. They now work on both. The binary-part test skips on
1.0, because 1.0 removed that part.

Behaviour on 0.1.15 to 0.1.22 is unchanged. The version range in
`pyproject.toml` is unchanged.

## Testing

- The new tests fail on 1.0.0 without the fix and pass with it. They
skip on 0.1.x, which cannot send list content.
- The AG-UI test files pass on 1.0.0, 0.1.22 and 0.1.15.
- Full unit suite with CI's command on 1.0.0: 20,499 passed, 0 failed,
236 skipped. I had no Postgres service locally, so those suites were
among the skips.
- `ruff check` and `mypy` are clean on Python 3.10 with 1.0.0 installed.
`format.sh` and `validate.sh` pass.
- I ran the AG-UI cookbook examples against a real model using the
official `@ag-ui/client` 1.0.0. They work on 1.0.0 and on 0.1.22.
`agent_with_media` was run with an OpenAI model because I did not have a
valid Gemini key.

## Not changed here

These come from 1.0 itself and can be follow-ups:

- A legacy `binary` content part is now rejected with 422 by the SDK.
- The new `file` source on media parts is accepted and skipped without a
log line.

## Type of change

- [x] Bug fix
- [ ] New feature
- [ ] Breaking change
- [ ] Improvement
- [ ] Model update
- [ ] Other:

---

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [x] Self-review completed
- [x] Documentation updated (comments, docstrings)
- [ ] Examples and guides: Relevant cookbook examples have been included
or updated (if applicable)
- [x] Tested in clean environment
- [x] Tests added/updated (if applicable)

### Duplicate and AI-Generated PR Check

- [x] I have searched existing [open pull
requests](https://github.com/agno-agi/agno/pulls) and confirmed that no
other PR already addresses this issue
- [ ] If a similar PR exists, I have explained below why this PR is a
better approach
- [ ] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

---

## Additional Notes

Reference: the "Migrating to 1.0" page on docs.ag-ui.com (Python
section).

agno-agi#10102 and agno-agi#10125 also edit `test_agui_app.py` and `resume.py`, so they
will need a small rebase after this.
## Summary

Adds **Y-API** (`https://y-api.bestvirtualgoods.com`) as an
OpenAI-compatible model provider,
following the "Adding a new Model Provider" section of
`CONTRIBUTING.md`:

- `libs/agno/agno/models/yapi/yapi.py` — `YAPI(OpenAILike)`, with the
usual `id` / `name` /
`provider` / `api_key` / `base_url` attributes and a `YAPI_API_KEY`
fallback in
  `_get_client_params()`.
- `libs/agno/agno/models/utils.py` — one row in the `_PROVIDERS` table
(`"yapi": ("agno.models.yapi", "YAPI", "YAPI", "yapi")`), inserted after
`xiaomi` to keep the
  table's alphabetical order.
- `libs/agno/tests/unit/models/yapi/test_yapi.py` — unit tests
(defaults, env-var auth, client
params, `"yapi:<model>"` string resolution,
`to_dict()`/`get_model_from_dict()` round trip).
- `cookbook/90_models/yapi/` — `README.md`, `basic.py`, `tool_use.py`,
`TEST_LOG.md`, showing
  both the class and the string syntax.

Motivation: the gateway is already listed in models.dev, LiteLLM, Fabric
and a few other
provider registries, and users have asked how to drive it from Agno.
Model ids are
org-prefixed on the wire (`deepseek/deepseek-v4-flash`, `z-ai/glm-5.3`,
`openai/gpt-5.6-sol`, …),
so the default is set to a real id rather than a placeholder.

(If applicable, issue number: Closes agno-agi#10327)

## Type of change

- [ ] Bug fix
- [x] New feature
- [ ] Breaking change
- [ ] Improvement
- [ ] Model update
- [ ] Other:

---

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [x] Self-review completed
- [x] Documentation updated (comments, docstrings)
- [x] Examples and guides: Relevant cookbook examples have been included
or updated (if applicable)
- [x] Tested in clean environment
- [x] Tests added/updated (if applicable)

### Duplicate and AI-Generated PR Check

- [x] I have searched existing [open pull
requests](https://github.com/agno-agi/agno/pulls) and confirmed that no
other PR already addresses this issue
- [ ] If a similar PR exists, I have explained below why this PR is a
better approach
- [x] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

---

## Additional Notes

**AI disclosure**: the diff was drafted with AI assistance. I reviewed
every line, ran the
verification below myself, and can explain any of it — but per the
contributing guide I'd
rather say so up front than have you guess.

**Verification** (in a clean venv with `pip install -e libs/agno`, using
the tool versions
pinned in `libs/agno/pyproject.toml`: ruff 0.15.20, mypy 2.1.0):

| Check | Result |
| --- | --- |
| `pytest libs/agno/tests/unit/models/yapi
libs/agno/tests/unit/models/test_provider_resolution.py -q` | **153
passed, 40 skipped** |
| `ruff format --check` / `ruff check` on the touched paths | clean |
| `ruff check --select I` (import sort) | clean |
| `python3 cookbook/scripts/check_cookbook_pattern.py --base-dir
cookbook/90_models/yapi --recursive` | 2 files, 0 violations |
| `ruff check libs/agno cookbook` | **20 errors before, 20 after,
identical set** — all pre-existing |
| `ruff format --check libs/agno cookbook` | 1 pre-existing file would
be reformatted (`tests/unit/vectordb/test_elasticsearch.py`) both before
and after; 2497 files already formatted |
| `mypy libs/agno --config-file libs/agno/pyproject.toml` | **18 errors
before, 18 after, identical set** — all pre-existing
(`knowledge/loaders/azure_blob.py`, `context/web/parallel_mcp.py`, …),
none in `models/yapi` |
| `pytest libs/agno/tests/unit/models -q` | 14 collection errors —
identical with and without this change, all `ImportError` for optional
SDKs (`anthropic`, `google-genai`, `litellm`, `ollama`) that aren't
installed here |

I also drove the class against the live endpoint to make sure it isn't
just structurally
correct:

```
$ python -c "..."   # YAPI_API_KEY set, real requests
base_url: https://api.y-api.bestvirtualgoods.com/v1 | provider: YAPI | name: YAPI
1) 类调用 -> 'ok'
2) 字符串语法 -> YAPI z-ai/glm-5.3
3) Agent.run -> 'ok'
```

**One caveat worth knowing about, documented in the cookbook README**:
the four
`openai/gpt-5.6-*` / `openai/gpt-6-astra` ids reject requests that carry
`tools` unless
`reasoning_effort` is set explicitly (the upstream answers
`Function tools with reasoning_effort are not supported`). That's why
the default id is
`deepseek/deepseek-v4-flash`, which supports tool calls without extra
parameters. Happy to
move that note into the class docstring instead if you'd prefer it
there.

---------

Co-authored-by: jiweiyeah <jiweiyeah@users.noreply.github.com>
Co-authored-by: Kaustubh <shuklakaustubh84@gmail.com>
Fixes agno-agi#10318.

> **AI disclosure:** this PR was authored with AI assistance and
reviewed by the author; per the contributing guidelines it is held to
the same quality bar.

## Summary

One duplicated word in
`libs/agno/agno/db/migrations/V3_MIGRATION_GUIDE.md` (line 200): "a sign
that **that that** session was not migrated" → "a sign that **that the**
session was not migrated". The sentence explains when
`db.cleanup_legacy_runs_column()` refuses to drop the legacy `runs`
column; the duplication reads as a typo and doesn't change the meaning.

Linked issue: agno-agi#10318 (created on the triage bot's request).

## Type of change

- [x] Bug fix

## Checklist

- [x] Code complies with style guidelines
- [ ] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`) — not applicable: the change is one word in a
Markdown file, no code or formatting surface touched
- [x] Self-review completed
- [x] Documentation updated (this PR is the documentation change)
- [ ] Examples and guides: Relevant cookbook examples have been included
or updated — N/A
- [ ] Tested in clean environment — N/A: nothing executable changed
- [ ] Tests added/updated — N/A: nothing executable changed

### Duplicate and AI-Generated PR Check

- [x] I have searched existing [open pull
requests](https://github.com/agno-agi/agno/pulls) and confirmed that no
other PR already addresses this issue (searched open PR titles for
"migration guide" and the duplicated phrase — no matches)
- [ ] If a similar PR exists, I have explained below why this PR is a
better approach — no similar PR exists
- [x] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

## Additional Notes

Verification performed: `grep -c "that that"` on the file returns 0
after the change; `git diff` touches only line 200 of the guide.

Co-authored-by: Kaustubh <shuklakaustubh84@gmail.com>
…ences (agno-agi#10160)

## Summary

Replaces 12 dead or moved URLs in cookbook knowledge sources and docs
references. Every replacement target was verified to return HTTP 200
(checked with a real browser UA, following redirects) and every old URL
confirmed dead:

| Old | New | Note |
|---|---|---|
| `docs.agno.com/introduction.md` (404) | `docs.agno.com/introduction`
(200) | Three Knowledge sources ingested nothing because of this —
`TEST_LOG.md` in this repo already records the ingestion failure
(`httpx.HTTPStatusError: 307 → broken target`) |
|
`raw.githubusercontent.com/run-llama/llama_index/main/docs/docs/examples/.../paul_graham_essay.txt`
(404) | `.../main/docs/examples/...` (200) | llama_index dropped the
nested `docs/docs/` on main |
| `mcp.deepwiki.com/sse` (410 Gone) | `mcp.deepwiki.com/mcp` (streamable
HTTP) | DeepWiki retired its SSE endpoint |
| `mcp.pipedream.com/app/<app>` (404, ×6) | `pipedream.com/apps/<app>`
(200) | Pipedream moved the per-app MCP pages |
| `docs.agno.com/tools/mcp` (404) | `docs.agno.com/tools/mcp/overview`
(200) | Per docs.agno.com/llms.txt |

Left untouched: URLs inside `TEST_LOG.md` historical records (they
document what a past run actually did), API host roots that only answer
POST, and template placeholders.

## Testing

- Scripted sweep over all 375 unique external URLs in `cookbook/` (HEAD,
GET fallback, browser UA); the table above lists every genuinely dead
one with a live replacement.
- Each replacement target re-verified returning 200 individually.
- No code paths changed — URL strings only.

Disclosure: PR prepared with AI assistance (link sweep + verification);
all replacements manually reviewed.


Fixes agno-agi#10163

Co-authored-by: Sannya Singal <32308435+sannya-singal@users.noreply.github.com>
Co-authored-by: Kaustubh <shuklakaustubh84@gmail.com>
## Summary

Adds `cookbook/11_memory/integrations/inspeximus_integration.py`,
alongside the dakera, mem0, memori and zep examples, and lists it in
that directory's README.

inspeximus is a zero-dependency memory file with a correction channel: a
later write to the same key retires the earlier value, so a fact the
user corrected does not come back on a later recall, and `revert(key)`
makes the previous value current again without naming it. The example
shows both, and feeds the current values to an Agent through
`dependencies` and `add_dependencies_to_context`, the same way
`zep_integration.py` does.

The memory half needs no API key, so the behaviour can be checked before
a model is wired to it.

One detail worth flagging, because it is the reason for the first three
lines of the example: the file deletes its store on start. Re-running it
against the store it left behind would restate a value that `revert()`
had already retired, and inspeximus refuses that on purpose. Without the
delete, the example fails on its second run.

Closes agno-agi#10151

## Type of change

- [ ] Bug fix
- [ ] New feature
- [ ] Breaking change
- [ ] Improvement
- [ ] Model update
- [x] Other: cookbook example

---

## Checklist

- [x] Code complies with style guidelines
- [x] Ran format/validation scripts (`./scripts/format.sh` and
`./scripts/validate.sh`)
- [x] Self-review completed
- [x] Documentation updated (comments, docstrings)
- [x] Examples and guides: Relevant cookbook examples have been included
or updated (if applicable)
- [x] Tested in clean environment
- [ ] Tests added/updated (if applicable)

On the format and validation boxes, precisely what I ran: `ruff format
--check` and `ruff check` on the added file, which reports "1 file
already formatted" and "All checks passed", plus `ruff check --select I
--fix`. I did not run the `mypy` targets in `validate.sh`, which cover
`libs/agno` and `libs/agnoctl` and are untouched by this change.

On the tests box: there is no separate test. The two asserts in the
example are its proof, and I ran the file three times in the same
directory to confirm they hold on every run rather than only the first.

### Duplicate and AI-Generated PR Check

- [x] I have searched existing [open pull requests](../../pulls) and
confirmed that no other PR already addresses this issue
- [ ] If a similar PR exists, I have explained below why this PR is a
better approach
- [x] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

## Additional Notes

Written with AI assistance and reviewed by a human before opening. I
maintain inspeximus, so treat the framing as interested and these checks
as the part to re-run:

- with agno 3.0.9 installed, every `Agent` parameter the file uses
exists: `model`, `instructions`, `dependencies`,
`add_dependencies_to_context`, `markdown`
- the file executes to completion and the agent receives
`dependencies={'memory': 'The staging database is db-7.internal'}`, the
corrected value only
- both asserts hold against inspeximus 2.27.4 from PyPI, in a clean
virtualenv, on three consecutive runs

Co-authored-by: DanceNitra <DanceNitra@users.noreply.github.com>
Co-authored-by: Kaustubh <shuklakaustubh84@gmail.com>
…ection (agno-agi#10342)

## Summary

Fixes agno-agi#10277.

`ReasoningManager._detect_model_type_uncached` failed to classify a
`reasoning_model` served through an `OpenAILike` provider (e.g.
`DashScope`) when its ID is a DeepSeek thinking-mode model such as
`deepseek-v4-pro`:

- `is_deepseek_reasoning_model()` requires the model class to be exactly
`DeepSeek`, so `DashScope` never matches.
- `is_openai_reasoning_model()` only accepted `deepseek-r1`,
`minimax-m2`, and `minimax-m3` in its `OpenAILike` branch.

As a result the reasoning stage was skipped for
`DashScope(id="deepseek-v4-pro")` even though the same ID is treated as
native reasoning by `is_deepseek_reasoning_model()` when the class is
`DeepSeek`, and by the OpenRouter fallback substrings (`deepseek-v4`).

This PR extends the `OpenAILike` branch of `is_openai_reasoning_model()`
with the DeepSeek thinking-mode ID families that
`is_deepseek_reasoning_model()` already recognizes: `deepseek-reasoner`,
`deepseek-v3.1*`, `deepseek-v3.2*`, and `deepseek-v4*`.

## Type of change

- [x] Bug fix

---

## Checklist

- [x] Code complies with style guidelines (`ruff check` / `ruff format
--check` clean on the touched files; repo-wide `ruff check` reports the
same pre-existing findings on `main` and on this branch)
- [x] Ran format/validation scripts: format check run locally on changed
files
- [x] Self-review completed
- [x] Documentation updated: docstring-level only (no user-facing docs
needed)
- [ ] Examples and guides: not applicable
- [x] Tested in clean environment
- [x] Tests added/updated: 7 new cases in
`libs/agno/tests/unit/reasoning/test_reasoning_checkers.py`

### Duplicate and AI-Generated PR Check

- [x] I have searched existing [open pull
requests](https://github.com/agno-agi/agno/pulls) and confirmed that no
other PR already addresses this issue
- [ ] If a similar PR exists, I have explained below why this PR is a
better approach
- [x] Check if this PR was entirely AI-generated (by Copilot, Claude
Code, Cursor, etc.)

**AI disclosure:** this PR was authored with AI assistance (operated by
@yetuge) and reviewed against the repository's code and tests.

## Additional Notes

Verification performed locally (Python 3.12, `PYTHONPATH=libs/agno`, no
network):

- Reproduced on `origin/main`:
`is_openai_reasoning_model(DashScope(id="deepseek-v4-pro"))` returned
`False`.
- After the change it returns `True`; the full checker suite passes: `86
passed` (79 existing + 7 new).
- Negative case covered: `deepseek-chat` (non-thinking) stays `False`,
so the dispatch order in `ReasoningManager` is unchanged for
non-reasoning IDs.

Co-authored-by: Kaustubh <shuklakaustubh84@gmail.com>
No-go items:
- importing the toolkit no longer needs `openai`: the aimlapi model package
  loads the chat model lazily, so its attribution constants come without it
- async variants for every tool (agenerate_image, agenerate_video,
  agenerate_speech, atranscribe_audio) on httpx.AsyncClient with
  asyncio.sleep, so a video job no longer blocks the loop under arun
- generated media carries `format` next to `mime_type`

Should-fix items:
- transcribe_audio reads local files through Toolkit._check_path inside a
  base_dir (default cwd), so a prompt cannot make the agent upload an
  arbitrary file
- an empty transcript is a transcript, not a failure
- a job in any status outside the in-progress set ends the poll
- transient poll errors (429, 5xx, transport) are retried three times
  before the paid job is abandoned
- string-shaped `error` fields are read as such
- a base URL ending in /v1 (the chat model's form) is accepted
- application/octet-stream assets are typed from the URL

Nits: `timeout` reaches the base Toolkit, speech_format is validated up
front, one poll helper serves both jobs, transcript length is what gets
logged, the cookbook names files by their media type and lists run
results in TEST_LOG.md.
…y sends

Deepgram-backed jobs end as "error", but AssemblyAI-backed ones end as
"failed": aai/universal and aai/slam-1 both reported it live on
2026-09-23, with the provider's message in the usual error object. Only
"error" was read as a failure, so those jobs surfaced as "job ended with
status 'failed'" and dropped the message that says what to do.

On the open question from the review: "waiting" never appears. Six
transcription models were submitted and polled to a terminal state on
2026-09-23 and the gateway only ever reported queued, generating,
completed and failed. The Nova-3 docs example still tests for "waiting",
so it joins the in-progress set anyway — if it ever does appear it costs
one more poll, where treating it as terminal would report a running job
as failed.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.