Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -183,7 +183,7 @@ For the exploration workflow and local-memory evaluation, see [LIBERO exploratio

### Interactive CLI mode

Add `--interactive` (`-i`) to steer the agent live from your terminal. At the `you>` prompt, the built-in task is pre-filled — press Enter to use it or replace it with your own — then type any message while it runs to steer the agent at the next turn (`/help` lists commands; `/quit` or Ctrl-D ends). Requires an interactive terminal (TTY).
With `claude_code` or `codex`, add `--interactive` (`-i`) to steer the agent live from your terminal. The built-in task is pre-filled at the `you>` prompt. Press Enter to submit it as-is. To change or replace the task, edit the input before pressing Enter. While the agent runs, type follow-up messages to steer it (`/help` lists commands; `/quit` or Ctrl-D ends). Requires a TTY. The `api` planner uses a native CLI that runs the preset task first and accepts follow-ups between runs; use `/exit` to close it.

```bash
rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
Expand Down
2 changes: 1 addition & 1 deletion README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -184,7 +184,7 @@ rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \

### 交互模式

加上 `--interactive`(`-i`)即可在终端里实时引导智能体。在 `you>` 提示符处,内置任务已预填——按 Enter 直接使用,或替换为你自己的任务;智能体运行时,随时输入消息即可在下一轮引导它(`/help` 查看命令,`/quit` 或 Ctrl-D 结束)。需要交互式终端(TTY)。
使用 `claude_code` 或 `codex` 时,加上 `--interactive`(`-i`)即可在终端里实时引导智能体。内置任务会预填在 `you>` 提示符处。直接按 Enter 即可提交该任务;如需修改或替换任务,请先编辑输入内容,再按 Enter 提交。任务运行期间,可以继续输入消息引导智能体(`/help` 查看命令,`/quit` 或 Ctrl-D 结束)。需要真实终端(TTY)。`api` planner 使用原生 CLI,先运行预设任务,再在每轮运行完成后接收输入,使用 `/exit` 退出。

```bash
rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
Expand Down
4 changes: 2 additions & 2 deletions docs/source-en/rst_source/development/architecture.rst
Original file line number Diff line number Diff line change
Expand Up @@ -34,8 +34,8 @@ to add a new primitive.

**A swappable planner.** The planner is the LLM agent runtime that drives the
tool-calling loop. One ``--planner`` flag switches it while the tools and
prompts stay put. Three are built in: ``api`` is RPent's own tool-calling loop
(built on pydantic-ai, the default, provider-agnostic across model APIs);
prompts stay put. Three are built in: ``api`` uses the native Pydantic AI loop
with Harness sliding-window history trimming (the default, provider-agnostic);
``claude_code`` reuses the Claude Agent SDK runtime; ``codex`` reuses the Codex
SDK runtime. Because all three face the exact same tools, they can be compared
head-to-head on the same physical benchmark. See
Expand Down
5 changes: 5 additions & 0 deletions docs/source-en/rst_source/quickstart.rst
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,11 @@ streams agent reasoning, camera views, and the action timeline; submit another
task after the current one finishes. Use ``--dashboard-language zh-cn`` for the
Chinese UI.

.. _quickstart-interactive:

For terminal interaction, add ``--interactive`` (``-i``) to enter follow-up
instructions. It cannot be combined with ``--dashboard``.

Key CLI options
---------------

Expand Down
31 changes: 17 additions & 14 deletions docs/source-en/rst_source/usage/configure_planner.rst
Original file line number Diff line number Diff line change
Expand Up @@ -20,11 +20,9 @@ loop is orchestrated, and which model SDK is used.
- What it is
- When to pick it
* - ``api``
- Provider-agnostic tool-calling loop built on
`pydantic-ai <https://ai.pydantic.dev/>`_. It currently supports
the Anthropic Messages API, the OpenAI Responses API, and
OpenAI-compatible Chat Completions APIs. It handles prompt caching
and history-image pruning.
- A tool-calling loop built on `Pydantic AI <https://ai.pydantic.dev/>`_
that supports multiple model APIs. For long conversations, it sends
fewer older messages to the model.
- You want the tightest control over model calls, the widest
provider coverage, or the cheapest per-turn spend.
* - ``claude_code``
Expand All @@ -50,8 +48,8 @@ loop is orchestrated, and which model SDK is used.
The ``api`` planner (direct model API)
---------------------------------------

``--planner api`` is the default. It uses Pydantic AI to implement the
tool-calling loop and requires a provider prefix in ``--model``. The
``--planner api`` is the default. It uses the native Pydantic AI tool-calling
runtime and requires a provider prefix in ``--model``. The
project currently installs the Anthropic and OpenAI integrations, so it
can directly use the Anthropic Messages API, the OpenAI Responses API,
and OpenAI-compatible Chat Completions APIs.
Expand Down Expand Up @@ -79,12 +77,16 @@ needed):
Relevant ``api`` planner knobs:

- ``--max-tokens`` — cap each LLM reply (default ``8192``).
- ``--max-turns`` — cap the number of tool-calling turns (default
``100``).
- ``--max-turns`` — cap model requests across the whole conversation,
including retries and follow-ups (default ``100``). Reaching the cap is
a normal stop, not a planner error or a claim of task success. Exploration
can continue with the next session and merge memory when otherwise eligible.
- ``--no-images`` — never send image bytes; this is required for
text-only models. The agent then reasons from textual state alone,
so task performance may not be satisfactory.

For ``--interactive`` usage, see :ref:`Terminal interaction <quickstart-interactive>`.

.. _planner-claude-code:

The ``claude_code`` planner
Expand Down Expand Up @@ -262,17 +264,17 @@ schemas, or your context length.
Add a custom planner
--------------------

If none of the three planners fit — say you want to plug in an
If none of the built-in planners fit — say you want to plug in an
in-house planner, a research prototype, or a different agent SDK —
implement the ``rpent.planner.base.Planner`` protocol and add a
construction branch to ``rpent.planner.base.build_planner``:
subclass ``rpent.planner.base.Planner``, implement its abstract ``solve()``
method, and update ``rpent.planner.base.build_planner`` to create the new backend:

.. code-block:: python

# rpent/planner/my_planner.py
from rpent.planner.base import PlannerResult
from rpent.planner.base import Planner, PlannerResult

class MyPlanner:
class MyPlanner(Planner):
def solve(
self,
*,
Expand All @@ -281,6 +283,7 @@ construction branch to ``rpent.planner.base.build_planner``:
toolkit,
max_turns,
input_queue=None,
dashboard_interaction=None,
):
tool_specs = toolkit.get_tools_spec()
# Call the model with system_prompt, user_message, and tool_specs.
Expand Down
2 changes: 1 addition & 1 deletion docs/source-zh/rst_source/development/architecture.rst
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@

**可替换的 planner。** planner 就是驱动工具调用循环的 LLM agent 运行时。
一个 ``--planner`` 参数就能切换它,而工具和提示词保持不变。内置三种:
``api`` 是 RPent 自研的工具调用循环(基于 pydantic-ai,为默认值,
``api`` 使用 Pydantic AI 原生运行时与 Harness 滑动窗口历史裁剪(为默认值,
不绑定具体模型提供商);``claude_code`` 复用 Claude Agent SDK 运行时;
``codex`` 复用 Codex SDK 运行时。由于三者面对完全相同的工具,
可以在同一套物理基准上正面对比。配置方法见 :doc:`../usage/configure_planner`。
Expand Down
5 changes: 5 additions & 0 deletions docs/source-zh/rst_source/quickstart.rst
Original file line number Diff line number Diff line change
Expand Up @@ -64,6 +64,11 @@ Session 配置全部来自命令行,打开地址后直接进入实时监控;
推理过程、相机画面和动作时间线;任务结束后可以继续提交下一任务。使用
``--dashboard-language zh-cn`` 可切换到中文界面。

.. _quickstart-interactive:

也可以添加 ``--interactive``(``-i``),在终端输入后续指令。
该选项不能与 ``--dashboard`` 同时使用。

关键 CLI 选项
-------------

Expand Down
22 changes: 13 additions & 9 deletions docs/source-zh/rst_source/usage/configure_planner.rst
Original file line number Diff line number Diff line change
Expand Up @@ -20,8 +20,7 @@ SDK。
- 什么时候选它
* - ``api``
- 基于 `Pydantic AI <https://pydantic.dev/docs/ai/>`_ 实现的工具调用循环,
不绑定特定模型提供商。当前支持 Anthropic Messages API、OpenAI Responses
API 和 OpenAI 兼容的 Chat Completions API,内置 prompt 缓存和历史图片剪枝。
支持多种模型 API。对话过长时,会减少发送给模型的早期记录。
- 需要精细控制模型调用、支持更多模型提供商,或降低单轮调用成本。
* - ``claude_code``
- `Claude Agent SDK
Expand All @@ -44,7 +43,7 @@ SDK。
``api`` planner(直接调用模型 API)
-------------------------------------

``--planner api`` 是默认选项。它使用 Pydantic AI 实现工具调用循环,并要求
``--planner api`` 是默认选项。它使用 Pydantic AI 原生工具调用运行时,并要求
``--model`` 带有模型提供商前缀。当前项目安装的依赖包含 Anthropic 和 OpenAI
集成,因此可以直接使用 Anthropic Messages API、OpenAI Responses API,
以及 OpenAI 兼容的 Chat Completions API。
Expand All @@ -71,10 +70,14 @@ SDK。
``api`` planner 的相关调节参数:

- ``--max-tokens`` —— 单次 LLM 回复的 token 上限(默认 ``8192``)。
- ``--max-turns`` —— 工具调用轮数上限(默认 ``100``)。
- ``--max-turns`` —— 整段对话的模型请求次数上限,包含重试和后续输入
(默认 ``100``)。达到上限时正常停止,不记为规划器错误,也不代表任务成功。
探索模式仍可继续下一会话,并在满足其他条件时合并记忆。
- ``--no-images`` —— 不向模型发送图片字节;纯文本模型必须加此参数。此时
智能体只依赖文本状态推理,任务表现可能不够理想。

``--interactive`` 的用法见 :ref:`终端交互 <quickstart-interactive>`。

.. _planner-claude-code:

``claude_code`` planner
Expand Down Expand Up @@ -237,16 +240,16 @@ Dashboard 提供同一项检查:启动页的 **测试连接** 按钮会针对
接入自定义 planner
------------------

如果三种内置 planner 都不合适,例如需要接入内部 planner、研究原型或其他
agent SDK,可以实现 ``rpent.planner.base.Planner`` 协议,并在
``rpent.planner.base.build_planner`` 中增加对应的构造分支:
如果内置 planner 都不合适,例如需要接入内部 planner、研究原型或其他
agent SDK,可以继承 ``rpent.planner.base.Planner``,实现抽象方法 ``solve()``,
并修改 ``rpent.planner.base.build_planner``,使其能够创建新后端的实例:

.. code-block:: python

# rpent/planner/my_planner.py
from rpent.planner.base import PlannerResult
from rpent.planner.base import Planner, PlannerResult

class MyPlanner:
class MyPlanner(Planner):
def solve(
self,
*,
Expand All @@ -255,6 +258,7 @@ agent SDK,可以实现 ``rpent.planner.base.Planner`` 协议,并在
toolkit,
max_turns,
input_queue=None,
dashboard_interaction=None,
):
tool_specs = toolkit.get_tools_spec()
# 使用 system_prompt、user_message 和 tool_specs 调用模型。
Expand Down
3 changes: 2 additions & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,8 @@ classifiers = [
]

dependencies = [
"pydantic-ai-slim[anthropic,openai]>=2.1",
"pydantic-ai-slim[anthropic,openai,cli]>=2.43",
"pydantic-ai-harness>=0.31",
"pydantic>=2",
"fastapi>=0.110",
"uvicorn>=0.27",
Expand Down
14 changes: 10 additions & 4 deletions rpent/cli/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -306,6 +306,7 @@ def _start_continuation_session(
claude_code_max_budget_usd=args.claude_code_max_budget_usd,
dashboard_events=dashboard_events,
no_images=args.no_images,
interactive=args.interactive,
)
system_prompt = prompt_bundle.render(
"system",
Expand Down Expand Up @@ -368,6 +369,7 @@ def main() -> int:
)
if sys.stdin is None or not sys.stdin.isatty():
parser.error("This robot requires a TTY for operator confirmation.")
native_cli = args.interactive and args.planner == "api"
Comment thread
wilburx813 marked this conversation as resolved.
if args.base_url and args.planner in BASE_URL_ENV_BY_PLANNER:
parser.error(
"--base-url applies to the 'api' planner only; "
Expand Down Expand Up @@ -452,6 +454,7 @@ def main() -> int:
claude_code_max_budget_usd=args.claude_code_max_budget_usd,
dashboard_events=dashboard_events,
no_images=args.no_images,
interactive=args.interactive,
)
prompt_bundle = robot_spec.prompts
prompt_vars = {**prompt_vars, "output_dir": output_dir}
Expand All @@ -468,10 +471,13 @@ def main() -> int:
if human_interactive_exploration:
from rpent.tools.human_in_the_loop import HumanInTheLoopInput

operator_input = HumanInTheLoopInput(interactive=args.interactive)
# The native CLI reads between runs, leaving the TTY available to tools.
operator_input = HumanInTheLoopInput(
interactive=args.interactive and not native_cli
)
input_queue: "queue.Queue[str | None] | None" = None
await_first_prompt: "Callable[[], str | None] | None" = None
if args.interactive:
if args.interactive and not native_cli:
input_queue = queue.Queue()
# Pre-fill the first prompt with the rendered default task (editable
# preset);
Expand Down Expand Up @@ -575,7 +581,7 @@ def main() -> int:
config=run_config,
)
memory_manager = toolkit.memory
if operator_input is not None and args.interactive:
if operator_input is not None and input_queue is not None:

def accept_verdict(verdict: str, active_toolkit=toolkit) -> bool:
if not active_toolkit.request_direct_verdict(verdict):
Expand Down Expand Up @@ -618,7 +624,7 @@ def accept_verdict(verdict: str, active_toolkit=toolkit) -> bool:
if solved and callable(write_recipe):
recipe_path = write_recipe(recipe_tag) or recipe_path
finally:
if operator_input is not None and args.interactive:
if operator_input is not None and input_queue is not None:
operator_input.bind_verdict(None)
try:
if robot_spec.finalize_run is not None:
Expand Down
25 changes: 25 additions & 0 deletions rpent/dashboard/interaction.py
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,31 @@ def as_dict(self) -> dict[str, Any]:
}


class PlannerSessionDriver(Protocol):
"""Backend operations used by the Dashboard planner control.

The controller drains physical toolkit work before requesting an interrupt.
Drivers own their SDK resources, event consumers, and cleanup.
"""

async def submit(self, message: DashboardMessage) -> int:
"""Submit input and return the number of new completion events expected.

Steering an active turn may return zero. With deferred acknowledgement,
the driver reports when the message starts or is discarded through
``DashboardPlannerControl``; returning only confirms acceptance.
"""
...

async def interrupt(self) -> int:
"""Interrupt execution and return completions to remove from the count.

Count only work that will not report completion through the normal event
path, such as discarded queued input. Do not count those events twice.
"""
...


class DashboardInteractionPort(Protocol):
"""Planner-facing access to one Dashboard interaction Session."""

Expand Down
22 changes: 10 additions & 12 deletions rpent/dashboard/planner_control.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,9 +18,8 @@

import asyncio
from collections.abc import Callable
from typing import Any

from rpent.dashboard.interaction import DashboardInteractionPort
from rpent.dashboard.interaction import DashboardInteractionPort, PlannerSessionDriver


class DashboardPlannerControl:
Expand All @@ -34,12 +33,14 @@ def __init__(
emit_user: Callable[[str], None],
emit_initial_user: Callable[[], None],
defer_message_ack: bool = False,
submit_while_busy: bool = False,
) -> None:
self._interaction = interaction
self._cancel_active_and_wait = cancel_active_and_wait
self._emit_user = emit_user
self._emit_initial_user = emit_initial_user
self._defer_message_ack = defer_message_ack
self._submit_while_busy = submit_while_busy
self._lock = asyncio.Lock()
self._outstanding_completions = 0

Expand All @@ -52,7 +53,7 @@ async def start(self) -> None:
)
self._emit_initial_user()

async def run(self, driver: Any) -> None:
async def run(self, driver: PlannerSessionDriver) -> None:
"""Forward Dashboard commands until the interaction ends."""
version = self._interaction.interaction_version
while self._interaction.planner_activity != "ended":
Expand All @@ -62,7 +63,7 @@ async def run(self, driver: Any) -> None:
version,
)

async def complete(self, driver: Any) -> None:
async def complete(self, driver: PlannerSessionDriver) -> None:
"""Record one completed backend request and flush queued input."""
async with self._lock:
if self._interaction.planner_activity == "ended":
Expand All @@ -73,7 +74,7 @@ async def complete(self, driver: Any) -> None:
)
await self._flush(driver)

async def tool_completed(self, driver: Any) -> None:
async def tool_completed(self, driver: PlannerSessionDriver) -> None:
"""Flush input queued while the backend was running a tool."""
async with self._lock:
await self._flush(driver)
Expand All @@ -96,7 +97,7 @@ async def cancel_active_toolkit(self) -> None:
"""Cancel and drain the active toolkit operation off the event loop."""
await asyncio.to_thread(self._cancel_active_and_wait)

async def _process(self, driver: Any) -> None:
async def _process(self, driver: PlannerSessionDriver) -> None:
async with self._lock:
if self._interaction.planner_activity == "ended":
return
Expand Down Expand Up @@ -127,17 +128,14 @@ async def _process(self, driver: Any) -> None:
)
await self._flush(driver)
return
if self._interaction.planner_activity == "idle":
if self._interaction.planner_activity == "idle" or self._submit_while_busy:
await self._flush(driver)

async def _flush(self, driver: Any) -> None:
async def _flush(self, driver: PlannerSessionDriver) -> None:
message = self._interaction.claim_next_pending_message()
while message is not None and not self._interaction.task_replacement_requested:
try:
if self._defer_message_ack:
added_completions = await driver.submit_dashboard_message(message)
else:
added_completions = await driver.submit(message.text)
added_completions = await driver.submit(message)
except Exception as exc:
self._interaction.mark_message_failed(
message.message_id,
Expand Down
Loading
Loading