Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
31 commits
Select commit Hold shift + click to select a range
6499617
feat(tools): add AIMLAPITools for image, video, speech and transcription
Lookoff-AIMLAPI Sep 16, 2026
a7af81b
fix: docs correct the Perplexity and v0 base_url defaults in their do…
xizhuomengcontin Sep 16, 2026
bca8485
[fix] Accept text streams in TextReader (#10203)
RaycarlLei Sep 16, 2026
d079b77
fix: dependency bump urllib3 to 2.7.0 and requests to 2.33.0 in mcp_t…
katsugtgz Sep 16, 2026
9fabf65
test: fix flaky test_tool_hook_receives_messages (#10212)
harshsinha03 Sep 16, 2026
ccc22c8
[fix] keep falsy tool results (0, False, []) on the sync tool executi…
BlueX888 Sep 16, 2026
93aae98
fix: CONTRIBUTING paths, agno_infra license note, and docstring accur…
simpleqt Sep 16, 2026
cf4fc53
fix: retain falsy values in previous workflow content (#10208)
guanzhuohe273-hub Sep 16, 2026
57c07d4
fix: Preserve CSV data rows when skipping headers across pages (#9995)
baba9811 Sep 16, 2026
0d0db98
fix: document PythonTools has no code sandbox (#10218)
ysolanky Sep 16, 2026
ed64fab
fix: harden CodingTools run_shell against interpreter RCE (#10210)
ysolanky Sep 16, 2026
b315e17
[fix] Continue sitemap discovery after invalid gzip data (#10213)
RaycarlLei Sep 16, 2026
e2114d2
fix: use UTF-8 for agnoctl client config text (#10038)
Ghraven Sep 16, 2026
f6233f4
fix: run CodingTools restricted shell without a shell to neutralize o…
ysolanky Sep 16, 2026
a7615d9
fix: Preserve newlines and tabs in clean_text (#9985)
icn5381 Sep 16, 2026
cbc10df
chore: Release v3.0.10 (#10211)
harshsinha03 Sep 16, 2026
097bae6
chore: remove phantom Args entries in the tool docstrings (#10229)
simpleqt Sep 17, 2026
38039bf
cookbook: add Confident AI observability example (#10257)
kausmeows Sep 17, 2026
a71c4d0
fix: correct streaming docstring parameter names in two model classes…
simpleqt Sep 17, 2026
cc64676
feat: add knowledge level retrieval pipeline and MMR reranker (#10115)
sannya-singal Sep 18, 2026
f1cb6a3
fix: bring WebSocket workflow submissions in line with the HTTP admis…
ysolanky Sep 18, 2026
9467136
[fix] Paginate MCP tool discovery for ClientSession (#10012)
sicauzxl Sep 18, 2026
85b6d1d
fix: support ag-ui-protocol 1.0 in the AG-UI interface (#10283)
Himanshu040604 Sep 18, 2026
d173851
[feat] Add Y-API as an OpenAI-compatible model provider (#10326)
jiweiyeah Sep 20, 2026
0073e1d
[fix] Correct duplicated word in V3 migration guide (#10315)
yetuge Sep 20, 2026
d2c325a
[fix] replace dead links in cookbook knowledge sources and docs refer…
Theater-ahyeon Sep 20, 2026
da8aaaf
[cookbook] Add inspeximus memory integration example (#10146)
DanceNitra Sep 20, 2026
dc86cf2
fix: recognize DeepSeek thinking-mode IDs in OpenAILike reasoning det…
yetuge Sep 20, 2026
90cb710
Merge remote-tracking branch 'upstream/main' into feat/aimlapi-tools
Lookoff-AIMLAPI Sep 21, 2026
99492cb
fix(tools): address the AIMLAPITools review
Lookoff-AIMLAPI Sep 21, 2026
639d2e7
fix(tools): read the transcription failure status the gateway actuall…
Lookoff-AIMLAPI Sep 23, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/pull_request_template.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ Describe key changes, mention related issues or motivation for the changes.

### Duplicate and AI-Generated PR Check

- [ ] I have searched existing [open pull requests](../../pulls) and confirmed that no other PR already addresses this issue
- [ ] I have searched existing [open pull requests](https://github.com/agno-agi/agno/pulls) and confirmed that no other PR already addresses this issue
- [ ] If a similar PR exists, I have explained below why this PR is a better approach
- [ ] Check if this PR was entirely AI-generated (by Copilot, Claude Code, Cursor, etc.)

Expand Down
2 changes: 1 addition & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -148,4 +148,4 @@ Message us on [Discord](https://discord.gg/4MtYHHrgA8) if you have any questions

## 📝 License

This project is licensed under the terms of the [Apache-2.0 license](/LICENSE)
This project is licensed under the terms of the [Apache-2.0 license](LICENSE)
Original file line number Diff line number Diff line change
Expand Up @@ -41,5 +41,5 @@
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
knowledge.insert(name="Agno Docs", url="https://docs.agno.com/introduction.md")
knowledge.insert(name="Agno Docs", url="https://docs.agno.com/introduction")
agent.print_response("What are Agno's key features?")
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
"""
Knowledge-Level Reranking
=========================
A reranker set on Knowledge runs after the vector db returns results, rather than
inside the vector db itself. Two differences follow from that:

1. It works with any vector db, so the same reranker moves between backends.
2. It widens the fetch, so the reranker chooses from a real pool rather than only
reordering what the vector db already returned. candidate_multiplier (capped by
max_candidates) is set on the reranker itself.

The widened fetch is what makes ordering strategies possible: a reranker can only
surface a document that was retrieved in the first place.

See also: 03_reranking.py for vector db level reranking.
"""

import asyncio

from agno.agent import Agent
from agno.knowledge.knowledge import Knowledge
from agno.knowledge.reranker.cohere import CohereReranker
from agno.models.openai import OpenAIResponses
from agno.vectordb.qdrant import Qdrant

# ---------------------------------------------------------------------------
# Setup
# ---------------------------------------------------------------------------

qdrant_url = "http://localhost:6333"

knowledge = Knowledge(
vector_db=Qdrant(collection="knowledge_reranking_demo", url=qdrant_url),
reranker=CohereReranker(
# Candidates fetched per requested result, so Cohere can rescue a document that
# plain search ranked outside max_results. Costs that many times the API calls,
# so lower it to 1 to only reorder what the vector db already returned.
candidate_multiplier=3,
# Ceiling on the widened fetch, once the multiplier is above 1.
max_candidates=100,
),
)

agent = Agent(
model=OpenAIResponses(id="gpt-5.6-luna"),
knowledge=knowledge,
search_knowledge=True,
markdown=True,
)


async def main():
await knowledge.ainsert(
url="https://agno-public.s3.amazonaws.com/recipes/ThaiRecipes.pdf"
)

# Retrieves 25 candidates, reranks them, returns the top 5.
results = await knowledge.asearch("What are some Thai curry dishes?", max_results=5)
print("Reranked results:")
for document in results:
print(f" {document.name}")

await agent.aprint_response("What are some Thai curry dishes?", stream=True)


if __name__ == "__main__":
asyncio.run(main())
96 changes: 96 additions & 0 deletions cookbook/07_knowledge/02_building_blocks/08_mmr_diverse_results.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,96 @@
"""
MMR: Diverse, Non-Redundant Results
===================================
Vector search returns the closest matches to a query, which are often near-duplicates
of each other: five chunks that all say the same thing. MMR (Maximal Marginal
Relevance) picks documents one at a time, discounting each candidate by how similar it
already is to what has been selected.

lambda_mult controls the tradeoff:
- 1.0 ranks by relevance alone (equivalent to plain vector search)
- 0.5 balances relevance against difference
- 0.0 ranks by difference alone

MMR needs a pool larger than the number of results requested, which is what the
reranker provides: candidate_multiplier widens the fetch, MMR selects from it, and
max_results are returned.

MMR reads the embedding on each search result. Not every vector db returns one:
Milvus, MongoDB, Redis and Valkey do not, so MMR raises there rather
than silently returning unreranked results.

Take the returned order as the result: reranking_score holds the MMR score at the
moment each document was picked, which is not descending, so re-sorting by it discards
the diversity ordering.

Set a reranker in one place: with one on both Knowledge and the vector db, only the
one on Knowledge is applied and the vector db's is ignored.

See also: 07_knowledge_level_reranking.py for how the widened fetch works.
"""

import asyncio

from agno.agent import Agent
from agno.knowledge.knowledge import Knowledge
from agno.knowledge.reranker.mmr import MMRReranker
from agno.models.openai import OpenAIResponses
from agno.vectordb.qdrant import Qdrant

# ---------------------------------------------------------------------------
# Setup
# ---------------------------------------------------------------------------

qdrant_url = "http://localhost:6333"

knowledge = Knowledge(
vector_db=Qdrant(collection="mmr_demo", url=qdrant_url),
reranker=MMRReranker(
# Relevance against diversity: 1.0 is relevance alone, 0.0 difference alone.
lambda_mult=0.5,
# Candidates fetched per requested result, so MMR has a pool to choose from.
candidate_multiplier=5,
# Ceiling on that widened fetch, whatever max_results is asked for.
max_candidates=100,
),
)

agent = Agent(
model=OpenAIResponses(id="gpt-5.6-luna"),
knowledge=knowledge,
markdown=True,
)


def show(results, candidates: int) -> None:
"""Print a snippet per result: every chunk shares the source file name."""
print(f"Selected {len(results)} of {candidates} candidates:\n")
for document in results:
snippet = " ".join(document.content.split())[:100]
print(f" - {snippet}...")
print()


async def main():
await knowledge.ainsert(
url="https://agno-public.s3.amazonaws.com/recipes/ThaiRecipes.pdf"
)

query = "What are some Thai curry dishes?"

# Same query without MMR, to compare against.
plain = Knowledge(vector_db=knowledge.vector_db)
candidates = len(await plain.asearch(query, max_results=25))

print("\nWithout MMR")
show(await plain.asearch(query, max_results=5), candidates)

# Retrieves 25 candidates, selects 5 that are relevant but unlike each other.
print("With MMR")
show(await knowledge.asearch(query, max_results=5), candidates)

await agent.aprint_response("What are some Thai curry dishes?", stream=True)


if __name__ == "__main__":
asyncio.run(main())
104 changes: 104 additions & 0 deletions cookbook/07_knowledge/02_building_blocks/09_mmr_with_pgvector.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,104 @@
"""
MMR with PgVector
=================
The same diversity selection as 08_mmr_diverse_results.py, against PgVector.

MMR compares candidates to each other, so it needs the embedding of every search
result. PgVector returns embeddings on search, so MMR works against it directly.

Setup:
./cookbook/scripts/run_pgvector.sh

See also: 08_mmr_diverse_results.py for what lambda_mult controls.
"""

import asyncio

from agno.agent import Agent
from agno.knowledge.embedder.openai import OpenAIEmbedder
from agno.knowledge.knowledge import Knowledge
from agno.knowledge.reranker.mmr import MMRReranker
from agno.models.openai import OpenAIResponses
from agno.vectordb.pgvector import PgVector
from agno.vectordb.search import SearchType

# ---------------------------------------------------------------------------
# Setup
# ---------------------------------------------------------------------------

db_url = "postgresql+psycopg://ai:ai@localhost:5532/ai"

knowledge = Knowledge(
vector_db=PgVector(
table_name="mmr_demo",
db_url=db_url,
search_type=SearchType.hybrid,
embedder=OpenAIEmbedder(id="text-embedding-3-small"),
),
# Runs after PgVector returns candidates.
reranker=MMRReranker(
# Relevance against diversity: 1.0 is relevance alone, 0.0 difference alone.
lambda_mult=0.5,
# Candidates fetched per requested result, so MMR has a pool to choose from.
candidate_multiplier=5,
# Ceiling on that widened fetch, whatever max_results is asked for.
max_candidates=100,
),
)

# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------

agent = Agent(
model=OpenAIResponses(id="gpt-5.6-luna"),
knowledge=knowledge,
search_knowledge=True,
instructions=[
"Always search your knowledge base before answering.",
"Include sources in your response.",
],
markdown=True,
)

# ---------------------------------------------------------------------------
# Run Demo
# ---------------------------------------------------------------------------


def show(results, candidates: int) -> None:
"""Print a snippet per result: every chunk shares the source file name."""
print(f"Selected {len(results)} of {candidates} candidates:\n")
for document in results:
snippet = " ".join(document.content.split())[:100]
print(f" - {snippet}...")
print()


if __name__ == "__main__":

async def main():
await knowledge.ainsert(
url="https://agno-public.s3.amazonaws.com/recipes/ThaiRecipes.pdf"
)

print("\n" + "=" * 60)
print("PgVector hybrid search + MMR")
print("=" * 60 + "\n")

query = "What are some Thai curry dishes?"

# Same query without MMR, to compare against.
plain = Knowledge(vector_db=knowledge.vector_db)
candidates = len(await plain.asearch(query, max_results=25))

print("Without MMR")
show(await plain.asearch(query, max_results=5), candidates)

# Retrieves 25 candidates, selects 5 that are relevant but unlike each other.
print("With MMR")
show(await knowledge.asearch(query, max_results=5), candidates)

await agent.aprint_response("What are some Thai curry dishes?", stream=True)

asyncio.run(main())
Loading
Loading