A production-grade autonomous AI agent system — built from scratch, 161 tests, 0 failures.
Status note (2026-10): this system previously ran on an Alibaba Cloud ECS instance under systemd, serving ~100K requests with 0 recorded incidents. That instance has been released, so there is no live demo URL at present. Everything below is reproducible locally — see Quick Start.
Most AI Agent tutorials stop at "call an API, write a prompt." This project goes further — it's a complete autonomous agent system with 5 self-improvement engines across ~55 modules, automated security hardening, and production deployment. Built from zero by a 2026 graduate over one semester. 161 tests, 0 failures. Was deployed 24/7 on Alibaba Cloud ECS.
Key numbers:
| Metric | Result |
|---|---|
| Security | 14/14 penetration tests passed, 10/10 b3 benchmark attacks blocked |
| Code Repair | 90% fix rate on real-world Python bugs |
| Load Test | 1000/1000 requests, P95=150ms, 50 concurrent |
| Test Suite | 161 passed, 0 failed — zero regressions |
| Autonomy | 5 self-improvement engines (evolution, debate, bootstrap, self-play, meta-agent) |
| Deployment | Alibaba Cloud ECS + systemd daemon (instance since released) |
An autonomous agent system that doesn't just answer questions — it improves itself:
- Detects its own capability gaps → generates missing tools at runtime (Bootstrap Engine)
- Evolves tool implementations → optimizes code, rolls back on regression (Evolution Engine)
- Debates with itself → Primary proposes, Challenger critiques, Arbitrator synthesizes (Debate Engine)
- Self-play training → generates curriculum, improves via feedback loop (SelfPlay Engine)
- Meta-agent oversight → observes, decides, acts on system health (MetaAgent)
These 5 are the self-improvement engines; the wider system has ~55 modules under agent/
(memory, cost tracking, governance, sandbox, reliability, and others).
scripts/benchmark_engines.py measures the three that are directly comparable —
baseline vs Debate vs Matrix.
┌──────────────────────────────────────────────────────────┐
│ AutoPilot Loop │
│ │
│ CLASSIFY → EXECUTE → VERIFY → REFLECT → IMPROVE → RETRY
│ │ │ │
│ ▼ ▼ │
│ Matrix Router ┌───────┴───────┐ │
│ (role + model) │ 4 Strategies │ │
│ ├── Debate │ │
│ ├── Evolution │ │
│ ├── Bootstrap │ │
│ └── Meta Observe │ │
│ └───────┬───────┘ │
│ ▼ │
│ Evaluation Gate │
│ (I+F+U scoring) │
└──────────────────────────────────────────────────────────┘
Agent State Machine: IDLE → PLANNING → TOOL_CALL → REFLECT → LEARN → DONE
Detailed architecture: see ARCHITECTURE.md
I audited my own code against OWASP Top 10 for LLM Applications and found 12 critical/high vulnerabilities — then fixed every one:
| # | Vulnerability | Fix |
|---|---|---|
| 1 | Sandbox timeout bypassable (thread join) | multiprocessing.Process + terminate()/kill() |
| 2 | Path traversal via .. / Unicode / case |
Path.resolve() + case-insensitive match |
| 3 | API Key auth silently disabled when unset | production mode refuses startup without key |
| 4 | Token entropy 64-bit (truncation) | HMAC-SHA256, 64-char hex, random salt |
| 5 | CORS allow_origins=["*"] with credentials |
Explicit origin whitelist |
| 6 | Token brute-force (no rate limit) | 5 attempts/min/IP, counter resets on success |
| 7 | Prompt injection unprotected | 30+ pattern guard (CN + EN), all endpoints |
| 8 | Identity creator not tracked | Audit trail: who created which identity |
| 9 | Audit log redaction incomplete | Regex: API keys, JWTs, Bearer tokens |
| 10 | Coarse-grained permissions (4 roles) | Resource-level PermissionGrant with fnmatch |
| 11 | Tenant ID spoofable (Header) | Cryptographic tenant binding |
| 12 | Tool risk misclassified | 4-level risk model with CISO approval gate |
Result: 14/14 automated penetration tests pass. Run uv run python scripts/pentest.py to verify.
Prompt Injection ████████████████████ 100% blocked
Path Traversal ████████████████████ 100% blocked
Token Brute Force ████████████████████ 100% blocked
Bootstrap Safety ████████████████████ 100% blocked
Evolution Safety ████████████████████ 100% blocked
Stress Test (50 concurrent):
Total: 1000 | OK: 1000 (100%)
Avg: 87ms | P50: 68ms | P95: 150ms | P99: 300ms
Code Repair (10 real bugs):
Fix Rate: 90% (9/10)
Detect Rate: 70% (7/10)
Self-Correction: 30% (feedback-driven retry)
| Engine | Avg Score | Latency |
|---|---|---|
| Baseline (DeepSeek V4 only) | 8.9/10 | ~13s |
| Debate (process-centric) | 8.3/10 | ~119s |
| Matrix (multi-model route) | 8.9/10 | ~30s |
Insight: DeepSeek V4 is already strong. Debate doesn't help simple tasks but fixes 1/5 baseline errors on hard code-debugging tasks. Key: selective use, not blind debate for every request.
| Layer | Technology | Why |
|---|---|---|
| LLM | DeepSeek V4 (primary), Qwen2.5:7b (reviewer) | OpenAI-compatible, model-agnostic |
| Framework | FastAPI + AsyncIO + Uvicorn | Standard Python AI serving stack |
| Vector DB | ChromaDB + all-MiniLM-L6-v2 | Lightweight RAG, easy to deploy |
| Security | Process sandbox + AST scan + HMAC auth | Defense in depth |
| Deployment | Docker + Alibaba Cloud ECS + systemd | production-grade config |
| Monitoring | Prometheus + CLEAR 5D panel | Cost, Latency, Efficacy, Assurance, Reliability |
git clone https://github.com/aidless/ai-agent-playground.git
cd ai-agent-playground
cp .env.example .env
# Edit .env: add DEEPSEEK_API_KEY
uv sync
uv run uvicorn agent.server:app --host 0.0.0.0 --port 8000Docker: docker-compose up -d | Production: ./deploy.sh setup && nano .env && ./deploy.sh start
161 tests, 1 skipped. The suite runs fully offline — no API calls are made.
git clone https://github.com/aidless/ai-agent-playground.git
cd ai-agent-playground
# --group dev is required: pytest lives in the dev dependency group
uv sync --group dev
# agent/server.py validates the key's format (must start with `sk-`) at import
# time; the value is never used to reach the network during tests.
export DEEPSEEK_API_KEY="sk-local-test-dummy"
uv run --group dev pytest tests/ -q
# 161 passed, 1 skippedThe skipped test is the locust load test (tests/load_test.py); add it with
uv add --dev locust.
If pytest resolves to a global install instead of the venv, prefix with
uv run --group dev as above — plain uv run pytest picks up whatever
pytest is on PATH.
Note for contributors: prometheus-client, sse-starlette, and tenacity
are imported by the code and the tests but were missing from dependencies;
pytest and pytest-asyncio were missing entirely (now in [dependency-groups].dev).
| Endpoint | Method | Description |
|---|---|---|
/health |
GET | System health — 10 subsystems |
/chat/completions |
POST | OpenAI-compatible chat (drop-in replacement) |
/autopilot/solve |
POST | Full 9-engine autonomous loop |
/super/debate |
POST | Multi-model debate |
/super/evolve |
POST | Tool evolution (optimize → rollback) |
/super/meta/experiment |
POST | Sandboxed self-evolution |
/eval/gate |
POST | 3D quality evaluation (I+F+U) |
/eval/ab |
POST | A/B testing |
/matrix/solve |
POST | Multi-agent routing |
/security/intrusion |
GET | Intrusion detection status |
agent/ 43 Python files (production engine)
├── async_core.py Streaming agent + state machine
├── autopilot.py Full 9-engine autonomous coordinator
├── debate.py Multi-model debate
├── evolution.py Tool optimization + template learning
├── bootstrap.py Runtime tool generation + AST safety
├── reflect_action.py Failure detection + auto-degradation
├── self_play.py Curriculum learning self-improvement
├── sandbox_meta.py Safe self-modification sandbox
├── matrix.py Multi-model routing
├── eval_gate.py 3D quality evaluation gate
├── intrusion.py Anomaly detection (5 types)
├── identity.py RBAC + session tokens + rate limiting
└── server.py 30+ REST endpoints
ai_agent_playground/ 23 files (framework layer, HuggingFace-inspired)
tests/ 161 test cases (0 failures)
scripts/ 22 benchmark & deployment scripts
blog/ Technical blog (CN + EN)
This project demonstrates:
- Real production engineering — not a tutorial clone, but a system designed, built, deployed, and maintained by one person
- Security mindset — found and fixed 12 vulnerabilities, automated penetration testing
- Self-directed learning — read 20+ AI research papers, applied insights to implementation
- System-level thinking — architecture, observability, deployment, not just model calling
I'm open to AI Application Developer / AI Engineer roles.
GitHub: @aidless
Liu Zewen (刘泽文) — B.Eng. Software Engineering 2026, Qilu Institute of Technology