A general-purpose AI agent with sandboxed code execution, sub-agent parallelism, and multi-provider LLM support.
English | 中文
Get started in 30 seconds:
uv tool install box-agent # or: pip install box-agent (Python 3.10+)
box-agent setup # interactive config wizard
box-agent # start chattingOr run a one-shot task:
box-agent --task "Analyze sales.csv — show top 10 products by revenue with a bar chart"Most agent frameworks are either too simple (no sandbox, no tools) or too complex (massive dependencies, rigid architecture). Box Agent hits the sweet spot:
| Feature | Box Agent | Open Interpreter | Aider |
|---|---|---|---|
| Sandboxed code execution | Jupyter kernel in isolated venv | Runs in host Python | N/A |
| Sub-agent parallelism | Multiple sub-agents run concurrently | No | No |
| Multi-provider LLM | Anthropic, OpenAI, DeepSeek, SiliconFlow, any API | OpenAI + a few others | OpenAI + Anthropic |
| MCP tool integration | Native | No | No |
| ACP protocol (embed in apps) | Full support | No | No |
| Standalone binary | PyInstaller runtime, no Python needed | No | No |
| Context compression | Staged automatic compaction + LLM summary | Manual | Git-based |
Delegate isolated work through a flat task contract with optional tools, Skills,
files, write scope, and hard step/tool-call budgets. Omitted tools resolve only
to trusted local readers; explicit capabilities still pass a fail-closed runtime
policy. Passing known local text paths in files selects the bounded,
completeness-checked batch fast path automatically when read_file is the only
resolved tool; requesting additional tools keeps the general child loop. The parent remains
responsible for conflict handling, the final deliverable, and verification.
You: "Analyze data1.csv, data2.csv, and data3.csv separately, then give me a combined summary"
┌─ Sub-Agent 1 ──────┐ ┌─ Sub-Agent 2 ──────┐ ┌─ Sub-Agent 3 ──────┐
│ Read data1.csv │ │ Read data2.csv │ │ Read data3.csv │
│ Run statistics │ │ Run statistics │ │ Run statistics │
│ Generate charts │ │ Generate charts │ │ Generate charts │
│ → Summary: ... │ │ → Summary: ... │ │ → Summary: ... │
└─────────────────────┘ └─────────────────────┘ └─────────────────────┘
↓ parallel ↓
┌─ Parent Agent ──────────┐
│ Combines 3 summaries │
│ Produces final report │
└─────────────────────────┘
Child policy is derived by the runtime: process tools, external side effects, and unknown MCP tools fail closed; path writes require an exact scope. See the sub-agent delegation contract for schemas, limits, compatibility behavior, and host diagnostics.
Python runs in an isolated Jupyter kernel with pre-installed data science packages (pandas, numpy, matplotlib, scikit-learn, openpyxl, xlrd). Generated files (charts, CSVs, PDFs) are automatically detected and surfaced as structured artifacts.
One config, any provider:
# Anthropic
api_base: "https://api.anthropic.com"
provider: "anthropic"
model: "claude-sonnet-4-20250514"
# DeepSeek
api_base: "https://api.deepseek.com"
provider: "openai"
model: "deepseek-chat"
# Any OpenAI-compatible endpoint
api_base: "https://your-api.example.com/v1"
provider: "openai"
model: "your-model"- Oversized tool results: Individual results are persisted immediately when needed; fresh parallel results also share a 50k-character pre-request budget. The model receives a stable preview while full text remains on disk. Read results are exempt and stay bounded by Read's own line/character controls.
- Usage-aware auto-summary: The next request is estimated from the latest real API usage plus subsequent messages. When it reaches the model-derived safety threshold, older history is summarized into a
usermessage while bounded recent messages and todo, plan, and skill state are restored. - Tool-call arguments: Write/edit arguments remain verbatim until a whole-history summary replaces their turn; they are not independently compacted.
- Legacy safety guard: Internal history placeholders from older or externally supplied sessions are rejected if a model tries to reuse them as executable file/code arguments; Box-Agent requests one clean regeneration instead of writing the placeholder to disk.
- MCP Tools: Connect to any MCP server — web search, knowledge graphs, databases
- Claude Skills: 12 built-in runtime, Office, and artifact skills, plus on-demand marketplace skill sources
- ACP Protocol: Embed Box Agent in Electron apps, Zed Editor, or any ACP-compatible host via JSON-RPC over stdio
- Standalone Runtime: PyInstaller binary bundles Python + all dependencies. No external Python needed — download and run
- Cross-session Memory: Persistent memory lets the agent retain key information across conversations
- Safety Layer: Dangerous command detection, workspace scope control, auto-backup before file modifications. Interactive permission negotiation for out-of-workspace access (CLI prompts user, ACP sends reverse RPC to host)
- Planning Snapshots: Structured plan tool for rendering objective, scope, steps, verification, and risks in host UIs
- Task Tracking: Built-in todo tool for multi-step task decomposition and progress tracking
The agent creates a webpage and opens it in the browser.
The agent uses a skill to create a professional document.
The agent searches the web and summarizes results.
Requires Python 3.10+. If your system Python is older (e.g. 3.9), use
uv tool install— it manages Python automatically.
uv handles Python version management for you — no need to upgrade your system Python:
# Install uv (if not already)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Install box-agent (auto-downloads Python 3.10+ if needed)
uv tool install box-agent
box-agent setup # interactive config wizard
box-agent # start chatting
# Upgrade later
uv tool upgrade box-agentIf you already have Python 3.10+:
pip install box-agent
box-agent setup
box-agentgit clone https://github.com/Raccoon-Office/Box-Agent.git
cd Box-Agent
uv sync
uv run python -m box_agent.cliIf you are joining the project as a collaborator, start here before changing code:
git clone https://github.com/Raccoon-Office/Box-Agent.git
cd Box-Agent
git submodule update --init --recursive # needed for bundled skills
uv sync
uv run python -m box_agent.cli --help
uv run pytest tests/test_core.py -qRead these files first:
AGENTS.md— repo-local engineering rules and verification expectations.CONTRIBUTING.md— contribution flow, PR checklist, and commit style.docs/REVIEW_GUIDE.md— maintainer review order, blockers, and proof requirements.docs/DEVELOPMENT_GUIDE.md— deeper architecture and development notes.docs/INTEGRATION.md— ACP/runtime integration details for host apps.
Project map:
| Area | Where to start |
|---|---|
| Agent execution loop | box_agent/core.py, box_agent/agent.py, box_agent/events.py |
| CLI and config | box_agent/cli.py, box_agent/config.py, box_agent/config/ |
| LLM providers | box_agent/llm/ |
| Built-in tools | box_agent/tools/ |
| ACP server/runtime embedding | box_agent/acp/, box_agent/build_runtime_cli.py |
| Skills | box_agent/skills/, box_agent/tools/skill_loader.py |
| Tests | tests/test_<area>.py |
Common development loop:
# Run the smallest relevant test while iterating
uv run pytest tests/test_bash_tool.py -q
# Run the broader suite before handing off
uv run pytest tests/ -q
# Catch whitespace/patch formatting issues
git diff --checkUse focused tests for the area you touched: tools in tests/test_*_tool.py,
LLM behavior in tests/test_llm*.py / tests/test_error_messages.py, ACP in
tests/test_acp*.py, memory in tests/test_memory*.py, and runtime packaging
in tests/test_build_runtime.py / tests/test_cli_runtime.py. Tests that need
real provider credentials are skipped unless the required API keys are present.
When a change affects the standalone runtime used by a host app, source changes are not enough: rebuild the runtime, install it into the host, restart the running ACP process, then probe the installed runtime. For local packaging:
uv run box-agent-build-runtimeBuild a versioned runtime and install the resulting archive into the usual officev3 checkout in one command:
uv run box-agent-build-runtime --version 0.8.82 --install-officev3Pass an explicit checkout path after --install-officev3, or set
BOX_AGENT_OFFICEV3_DIR, when officev3 is stored elsewhere.
After running box-agent setup, your config lives at ~/.box-agent/config/config.yaml:
api_key: "your-api-key"
api_base: "https://api.anthropic.com"
model: "claude-sonnet-4-20250514"
provider: "anthropic" # "anthropic" or "openai"
max_steps: 300
max_parallel_tools: 8
parallel_tool_timeout_seconds: 900
provider_stale_seconds: 300
sub_agent_token_limit: 50000
sub_agent_batch_synthesis_timeout_seconds: 600 # 0 disables the extra batch synthesis cap
goal_autopilot_enabled: true
goal_autopilot_max_turns: 3
goal_autopilot_max_seconds: 14400
goal_autopilot_no_progress_turns: 2Tool limits are omitted by default so runtime upgrades can supply updated
defaults from box_agent/config.py. Add only deliberate overrides under
tool_limits:; inspect the current effective values with
box-agent config --json.
box-agent config # show current config summary
box-agent config --get model # print one config value
box-agent config --set max_steps 300
box-agent config --set goal_autopilot_max_turns 5
box-agent config --set tool_limits.external_skill.max_tool_calls 160
box-agent config --set tool_limits.external_skill.max_delegated_tool_calls 512
box-agent config --set tool_limits.completion.deadline_seconds 1800
box-agent config --set tools.bash_default_timeout_seconds 300
box-agent config --json # machine-readable config summary
box-agent config --edit # open in editor
box-agent doctor # check environment & API connectivity
box-agent doctor --json # machine-readable health checkHosts may opt into a dedicated absolute BOX_AGENT_HOME before starting the runtime. This routes Box-Agent-owned configuration and state away from the default user profile without changing the operating-system HOME; unset behavior stays compatible. Explicit profiles require their own config/config.yaml and never fall back to legacy configuration or login-token environment variables.
For endpoints that reject reasoning_effort: "none", see
reasoning configuration for the explicit
reasoning_effort_when_disabled: low compatibility option and child Agent inheritance.
See isolated runtime profiles and the no-background-startup example configuration. This is path/configuration isolation, not an OS sandbox or permission to call a model.
# Interactive mode
box-agent
box-agent --workspace /path/to/project
box-agent --no-sandbox # disable Jupyter sandbox
# Non-interactive (CI/CD, scripts)
box-agent --task "analyze data.csv and create a report"
box-agent --task "analyze data.csv" --json # append execution summary JSON
box-agent --task "local file task" --no-verify-api # skip startup API probe
box-agent --task "create a PPT" --force-plan-start # publish a plan before work
box-agent --task "create a PPT" --no-completion-gate
box-agent --goal "Ship CLI parity" --task "finish tests"
box-agent --goal "Ship CLI parity" --task "finish tests" --no-goal-autopilot
box-agent --deep-think --task "review this repo" # enable thinking mode when supported
# Subcommands
box-agent setup # config wizard
box-agent config # show/edit config
box-agent doctor # health check
box-agent log # open log directory
box-agent trace-viewer # open the offline Agent Trace diagnostics page
box-agent goal status # show persistent workspace goal
box-agent goal complete --evidence "tests passed"
box-agent install-browser # install Chromium for Playwright MCP (~200MB)
box-agent install-node # install managed Node.js runtime for skills (macOS)Run box-agent trace-viewer to open the packaged, offline developer viewer. Open ~/.box-agent/log/sessions/ for a newest-first overview of every trace, then select one run to inspect per-turn metrics, LLM/tool waterfalls, raw events, and the complete system → user → assistant/tool → final-response chain. You can still open one .jsonl file directly.
Both ACP and CLI create session traces automatically. Each CLI invocation writes
cli-<unique-id>.jsonl; interactive user turns have distinct turn IDs in that
file. /clear resets conversation history without deleting the diagnostic trace,
and Goal autopilot continuations stay within the originating user turn. Native
CLI traces include input, LLM/tool activity, final output, stop reason and turn
duration; no test-harness writer is required. Startup probes and idle CLI commands
are outside the turn trace. Set BOX_AGENT_SESSION_TRACE_DIR to override the
directory, or BOX_AGENT_SESSION_TRACE_ENABLED=0 to disable tracing. CLI trace
initialization or write failures do not fail the task. These files contain sensitive task data; the same
retention policy below applies to both entrypoints.
Choose Compare sources to compare equivalent runs from any number of sources. The selected root uses source directories directly, without an extra data-directory level:
comparison-traces/
baseline/*.jsonl
current/*.jsonl
candidate-x/*.jsonl
The viewer groups traces by normalized turn.input first and may use a compatible normalized filename when input is missing or non-conflicting. It keeps repeated runs per source selectable, shows each source side by side, and calculates duration, call, token, and error deltas against the selected reference source.
If an embedded browser does not expose the native file picker, run the loopback-only service and enter the trace directory path in the page:
uv run python -m box_agent.trace_viewer.server --port 8766The offline page reads files in the browser. In ledger mode, the service reads only top-level .jsonl files from the directory you enter. In comparison mode, it reads top-level .jsonl files from each immediate non-symlink source directory and goes no deeper. It checks metadata once per second and refreshes the active view when files are added or changed; trace bodies are transferred over 127.0.0.1 only when that metadata changes. The service rejects requests whose Host or Origin is not its exact loopback authority, preventing a rebinding site from reading local traces. Neither mode makes external network requests. Chromium and Edge can keep following appended records after you grant a file handle; drag/drop and ordinary file inputs load a snapshot. Session traces may contain prompts, tool arguments, outputs, and business data—handle them as sensitive diagnostic artifacts.
Box-Agent ships with a disabled @playwright/mcp entry. To enable browser tools locally:
box-agent install-browser # downloads Chromium and flips the entry to enabledRequires Node.js ≥ 18 on PATH. Chromium lands in ~/.box-agent/browsers/ (shared by CLI and ACP runtime) and mcpServers.playwright.disabled in ~/.box-agent/config/mcp.json is set to false.
ACP embedders: no env-var plumbing required — box-agent-acp defaults PLAYWRIGHT_BROWSERS_PATH to the same ~/.box-agent/browsers/ path. To point at a different cache, export PLAYWRIGHT_BROWSERS_PATH=<your path> before spawning box-agent-acp (our setdefault won't override it).
In-session commands: /help, /clear, /clear_all, /history, /stats, /sandbox_status, /log, /goal, /memory review, /exit
CLI Skill subprocesses receive the original user requests, including constraints
such as "do not generate images", through the same source-binding contract as
ACP. Interactive requests accumulate until /clear or /clear_all; assistant
output and automatic goal-continuation prompts are not added as source facts.
ACP session traces keep their existing ~/.box-agent/log/sessions/<session-id>.jsonl
name and box-agent-session-trace/v1 record format. Retention removes only whole,
inactive session files: files older than 7 days are eligible, and the directory
has a soft 512 MiB cap. The current append target, files modified within 24
hours, and the newest two sessions are protected. Cleanup runs best-effort at
most once every 6 hours; cleanup failures never interrupt agent execution.
Operators can override the defaults with BOX_AGENT_SESSION_TRACE_RETENTION_DAYS,
BOX_AGENT_SESSION_TRACE_MAX_TOTAL_BYTES, and
BOX_AGENT_SESSION_TRACE_CLEANUP_INTERVAL_SECONDS, or disable cleanup with
BOX_AGENT_SESSION_TRACE_RETENTION_ENABLED=0.
Use /goal <objective> or --goal "<objective>" to keep a durable workspace objective attached to later turns. The CLI persists it under ~/.box-agent/goals/; later turns include that goal until you run /goal pause, /goal resume, /goal block <reason>, /goal complete <evidence>, or /goal clear. Scripted runs can manage it with box-agent goal ....
In non-interactive --task mode and ACP sessions, active goals also use bounded autopilot: when a turn ends naturally but the goal is still active, Box-Agent automatically continues in the same session until the model marks the goal complete, marks it blocked, the user cancels, goal_autopilot_max_turns / goal_autopilot_max_seconds is reached, or goal_autopilot_no_progress_turns consecutive automatic continuations make no recorded goal progress. Use --no-goal-autopilot for one CLI run, or set goal_autopilot_enabled: false in config.
Box Agent supports the Agent Communication Protocol for embedding in editors and apps.
Zed Editor — add to settings.json:
{
"agent_servers": {
"box-agent": {
"command": "/path/to/box-agent-acp"
}
}
}Standalone Runtime — for Electron apps and other hosts:
# Download pre-built binary (latest release; omit the tag to always get the newest)
gh release download --repo Raccoon-Office/Box-Agent --pattern "box-agent-runtime-*.tar.gz"
# Or build from source (current platform)
uv run box-agent-build-runtime
# Build macOS Intel/x64 runtime from Apple Silicon
# Requires a separate x86_64 venv because PyInstaller cannot bundle arm64 wheels into an x64 binary.
# One-time setup:
# arch -x86_64 /bin/bash -c 'curl -LsSf https://astral.sh/uv/install.sh | INSTALLER_NO_MODIFY_PATH=1 UV_INSTALL_DIR="$HOME/.local/bin-x64" sh'
# UV_PROJECT_ENVIRONMENT=.venv-x64 arch -x86_64 ~/.local/bin-x64/uv sync
# Reuse the existing x64 Python without depending on the x64 uv install path.
# Isolate intermediate outputs from ARM builds; add --version X.Y.Z if needed.
.venv-x64/bin/python -m box_agent.build_runtime_cli --target darwin-x64 --output dist/runtime-darwin-x64The runtime communicates via JSON-RPC over stdio. stdout = protocol only, stderr = diagnostics. macOS runtime archives contain ACP and its internal dependencies. Stable tool Python/Node runtimes are host-managed and must match the target architecture; they are not bundled into this ACP-only archive.
uv run pytest tests/ -v # all tests
uv run pytest tests/test_core.py -v # core + context compression
uv run pytest --cov # with coverageSSL Certificate Error: pip install --upgrade certifi or set verify=False for testing.
Module Not Found: Make sure you're in the project directory: cd Box-Agent && uv run python -m box_agent.cli
Issues and PRs welcome! See Contributing Guide.
If this project helps you, give it a ⭐!


