Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions .agents/plugins/marketplace.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
{
"name": "planman",
"interface": {
"displayName": "Planman"
},
"plugins": [
{
"name": "planman",
"source": {
"source": "local",
"path": "./plugins/planman"
},
"policy": {
"installation": "AVAILABLE",
"authentication": "ON_INSTALL"
},
"category": "Developer Tools"
}
]
}
4 changes: 2 additions & 2 deletions .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,13 +4,13 @@
"name": "rusdyn"
},
"metadata": {
"description": "Evaluates Claude Code plans using OpenAI Codex CLI",
"description": "Evaluates AI coding-agent plans using another agent",
"version": "0.4.9"
},
"plugins": [
{
"name": "planman",
"description": "Evaluates Claude Code plans using OpenAI Codex CLI",
"description": "Evaluates AI coding-agent plans using another agent",
"version": "0.4.9",
"source": "./",
"author": {
Expand Down
4 changes: 2 additions & 2 deletions .claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,10 +1,10 @@
{
"name": "planman",
"version": "0.4.9",
"description": "Evaluates Claude Code plans using OpenAI Codex CLI and rejects low-scoring plans with feedback",
"description": "Evaluates AI coding-agent plans using another agent and rejects low-scoring plans with feedback",
"author": { "name": "dyn" },
"license": "MIT",
"keywords": ["plan", "evaluation", "quality-gate", "codex"],
"keywords": ["plan", "evaluation", "quality-gate", "codex", "claude"],
"homepage": "https://github.com/RusDyn/planman",
"repository": "https://github.com/RusDyn/planman"
}
24 changes: 24 additions & 0 deletions .codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
{
"name": "planman",
"version": "0.4.9",
"description": "Evaluates AI coding-agent plans using another agent",
"author": { "name": "dyn" },
"license": "MIT",
"keywords": ["plan", "evaluation", "quality-gate", "codex", "claude"],
"homepage": "https://github.com/RusDyn/planman",
"repository": "https://github.com/RusDyn/planman",
"interface": {
"displayName": "Planman",
"shortDescription": "Evaluates coding-agent plans before execution",
"longDescription": "Planman evaluates implementation plans from Codex or Claude Code with another local coding-agent CLI and blocks weak plans with actionable feedback.",
"developerName": "dyn",
"category": "Developer Tools",
"capabilities": ["Read", "Interactive"],
"defaultPrompt": [
"Evaluate my implementation plan",
"Stress-test this coding plan",
"Show planman status"
],
"brandColor": "#2563EB"
}
}
12 changes: 6 additions & 6 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -8,16 +8,16 @@ build/
.env

# Claude Code local state
.claude/planman.log
.claude/plans/
.claude/settings.local.json
.claude/

# Playwright MCP auto-generated logs
.playwright-mcp/

# Benchmark data (large/ephemeral)
benchmark/exercises/
benchmark/results/
benchmark/swebench/results/
logs/
planman-*.json
submission
benchmark/
test_*.log
.omc/
sb-cli-reports/
91 changes: 61 additions & 30 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,16 +7,16 @@
<p align="center">
<a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-blue.svg" alt="License: MIT"></a>
<img src="https://img.shields.io/badge/python-3.8%2B-blue.svg" alt="Python 3.8+">
<img src="https://img.shields.io/badge/claude--code-plugin-blueviolet.svg" alt="Claude Code Plugin">
<img src="https://img.shields.io/badge/claude--code%20%2B%20codex-plugin-blueviolet.svg" alt="Claude Code and Codex Plugin">
</p>

A quality gate for AI-generated plans. This [Claude Code](https://docs.anthropic.com/en/docs/claude-code) plugin sends implementation plans to [OpenAI Codex CLI](https://github.com/openai/codex) for independent scoring, automatically rejecting low-quality plans with actionable feedback — so Claude iterates before you review.
A quality gate for AI-generated plans. This plugin evaluates implementation plans from Claude Code or Codex with another coding agent, automatically rejecting low-quality plans with actionable feedback before you review.

> *"Every plan deserves a second opinion."*

When Claude presents an implementation plan, planman intercepts it, sends it to [OpenAI Codex CLI](https://github.com/openai/codex) for scoring, and rejects low-scoring plans with actionable feedback. Claude revises and re-presents. After a configurable number of rounds, you decide.
When an agent presents an implementation plan, planman intercepts it, sends it to the configured evaluator for scoring, and rejects low-scoring plans with actionable feedback. By default, Claude Code plans are reviewed by Codex and Codex plans are reviewed by Claude Code. After a configurable number of rounds, you decide.

**No API keys required** — uses your ChatGPT subscription via the `codex` CLI.
**No API keys required for Codex review** — uses your ChatGPT subscription via the `codex` CLI. Claude review uses your local Claude Code CLI authentication.

## How It Works

Expand All @@ -26,7 +26,7 @@ Planman uses **two hooks** in a plan-mode-only architecture:
PostToolUse(Write) — records plan file path when Claude writes to .claude/plans/
│
▼
PreToolUse(ExitPlanMode) — evaluates plan via codex when Claude exits plan mode
PreToolUse(ExitPlanMode) — evaluates plan via configured evaluator when Claude exits plan mode
│
├── Round 1: mandatory review — plan always gets scored feedback
│ │
Expand All @@ -41,6 +41,7 @@ PreToolUse(ExitPlanMode) — evaluates plan via codex when Claude exits plan mod
```

**Deterministic:** Files in `.claude/plans/` are always treated as plans — no LLM-based plan detection.
For Codex-hosted sessions, planman uses Codex lifecycle hooks and evaluates only clearly plan-like content; if no plan is found, it fails open.

## Quick Start

Expand All @@ -54,9 +55,14 @@ PreToolUse(ExitPlanMode) — evaluates plan via codex when Claude exits plan mod
/plugin marketplace add RusDyn/planman
/plugin install planman@planman
```
4. Restart Claude Code
4. Or add planman to Codex:
```bash
codex plugin marketplace add RusDyn/planman
codex plugin add planman@planman
```
5. Restart the host agent

That's it. The next time Claude exits plan mode, planman evaluates the plan and blocks with feedback if the score is below threshold (default 7/10). Run `/planman:init` to customize settings.
That's it. The next time the host agent presents a detected plan, planman evaluates it and blocks with feedback if the score is below threshold (default 7/10). Run `/planman:init` to customize settings.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Make the quick-start customization step host-specific.

After presenting both Claude Code and Codex install paths, this line sends everyone to /planman:init, but Codex users do not have that command. Point Codex users here to .planman.jsonc or PLANMAN_* instead.

Suggested wording
-That's it. The next time the host agent presents a detected plan, planman evaluates it and blocks with feedback if the score is below threshold (default 7/10). Run `/planman:init` to customize settings.
+That's it. The next time the host agent presents a detected plan, planman evaluates it and blocks with feedback if the score is below threshold (default 7/10). In Claude Code, run `/planman:init` to customize settings. In Codex, create `.planman.jsonc` or set `PLANMAN_*` environment variables.
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
That's it. The next time the host agent presents a detected plan, planman evaluates it and blocks with feedback if the score is below threshold (default 7/10). Run `/planman:init` to customize settings.
That's it. The next time the host agent presents a detected plan, planman evaluates it and blocks with feedback if the score is below threshold (default 7/10). In Claude Code, run `/planman:init` to customize settings. In Codex, create `.planman.jsonc` or set `PLANMAN_*` environment variables.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@README.md` at line 65, Update the README sentence so the customization step
is host-specific: mention that users of hosts with the /planman:init command
(e.g., Claude Code) can run /planman:init, while Codex users should instead
configure via .planman.jsonc or with environment variables prefixed PLANMAN_*;
ensure both options are presented clearly and the wording replaces the current
single recommendation to run /planman:init.


> *"You're four steps from better plans."*

Expand All @@ -77,6 +83,8 @@ That's it. The next time Claude exits plan mode, planman evaluates the plan and

### From GitHub

Claude Code:

```bash
# Add the marketplace
/plugin marketplace add RusDyn/planman
Expand All @@ -85,8 +93,17 @@ That's it. The next time Claude exits plan mode, planman evaluates the plan and
/plugin install planman@planman
```

Codex:

```bash
codex plugin marketplace add RusDyn/planman
codex plugin add planman@planman
```

### Local Development

Claude Code:

```bash
# Add the local directory as a marketplace
/plugin marketplace add /path/to/planman
Expand All @@ -95,28 +112,38 @@ That's it. The next time Claude exits plan mode, planman evaluates the plan and
/plugin install planman@planman
```

Then restart Claude Code.
Codex:

```bash
codex plugin marketplace add /path/to/planman
codex plugin add planman@planman
```

Then restart the host agent.

## Configuration

Settings are loaded from env vars (highest priority) or `.claude/planman.jsonc`:
Settings are loaded from env vars (highest priority), `.planman.jsonc`, or legacy `.claude/planman.jsonc`:

| Setting | Env Var | Default | Description |
|---------|---------|---------|-------------|
| `threshold` | `PLANMAN_THRESHOLD` | `7` | Minimum score (0-10) to pass; 0 = pass all |
| `max_rounds` | `PLANMAN_MAX_ROUNDS` | `3` | Evaluation rounds before you decide (1-100) |
| `model` | `PLANMAN_MODEL` | *(codex default)* | Override Codex model (`-m` flag) |
| `fail_open` | `PLANMAN_FAIL_OPEN` | `true` | Pass through if Codex fails |
| `model` | `PLANMAN_MODEL` | *(evaluator default)* | Override evaluator model |
| `evaluator` | `PLANMAN_EVALUATOR` | `auto` | `auto`, `codex`, or `claude`; auto picks the other agent |
| `codex_bin` | `PLANMAN_CODEX_BIN` | `codex` | Codex CLI binary/path |
| `claude_bin` | `PLANMAN_CLAUDE_BIN` | `claude` | Claude Code CLI binary/path |
| `fail_open` | `PLANMAN_FAIL_OPEN` | `true` | Pass through if evaluator fails |
| `enabled` | `PLANMAN_ENABLED` | `true` | Master switch |
| `custom_rubric` | `PLANMAN_RUBRIC` | *(built-in)* | Custom evaluation rubric |
| `verbose` | `PLANMAN_VERBOSE` | `false` | Debug output to stderr |
| `source_verify` | `PLANMAN_SOURCE_VERIFY` | `true` | Codex verifies plan against actual source files |
| `source_verify` | `PLANMAN_SOURCE_VERIFY` | `true` | Evaluator verifies plan against actual source files |
| `stress_test` | `PLANMAN_STRESS_TEST` | `false` | Stress-test rounds (`false`/`true`/number N) |
| `context` | `PLANMAN_CONTEXT` | *(empty)* | Project context injected into evaluation prompt |

`stress_test` accepts `false` (off), `true` (1 round), or a number N (N stress-test rounds). Stress-test rounds skip Codex and auto-reject with the stress-test prompt. Codex evaluation begins at round N+1.
`stress_test` accepts `false` (off), `true` (1 round), or a number N (N stress-test rounds). Stress-test rounds skip the external evaluator and auto-reject with the stress-test prompt. External evaluation begins at round N+1.

Run `/planman:init` to generate `.claude/planman.jsonc` with all settings and inline documentation.
Run `/planman:init` in Claude Code to generate `.planman.jsonc` with all settings and inline documentation. In Codex, create the same file manually or use `PLANMAN_*` environment variables.

## Scoring Rubric

Expand All @@ -140,7 +167,7 @@ Override the built-in rubric for domain-specific evaluation:
export PLANMAN_RUBRIC="Score the plan focusing on security implications, test coverage, and backwards compatibility. Be strict about migration safety."
```

Or in `.claude/planman.jsonc`:
Or in `.planman.jsonc`:

```json
{
Expand All @@ -150,13 +177,13 @@ Or in `.claude/planman.jsonc`:

## Commands

All commands use the `planman:` namespace prefix:
Claude Code slash commands use the `planman:` namespace prefix. Codex uses the installed Stop hook plus `.planman.jsonc` or `PLANMAN_*` environment variables.

| Command | Description |
|---------|-------------|
| `/planman:status` | Show status, codex version, effective config |
| `/planman:status` | Show status, evaluator versions, effective config |
| `/planman:help` | Full usage guide |
| `/planman:init` | Create `.claude/planman.jsonc` with all defaults |
| `/planman:init` | Create `.planman.jsonc` with all defaults |
| `/planman:clear` | Clear session state (reset evaluation rounds) |

## Multi-Round Behavior
Expand All @@ -168,10 +195,11 @@ All commands use the `planman:` namespace prefix:

## Zero Friction Design

- **No API keys** — uses ChatGPT subscription via `codex` CLI
- **Cross-agent by default** — Claude plans use Codex review; Codex plans use Claude review
- **No API keys for Codex review** — uses ChatGPT subscription via `codex` CLI
- **No pip dependencies** — stdlib only (Python 3.8+)
- **Fail-open by default** — Codex errors never block your workflow
- **Auto-detect codex** — if `codex` isn't installed, hook silently passes through
- **Fail-open by default** — evaluator errors never block your workflow
- **Auto-detect evaluator** — if the resolved evaluator CLI isn't installed, hook silently passes through
- **Plan-mode only** — deterministic detection via `.claude/plans/` path

## State Files
Expand All @@ -183,34 +211,37 @@ Session state is stored in the system temp directory (run `python3 -c "import te

## Plugin Structure

- `.claude-plugin/marketplace.json` — marketplace registry (used by `/plugin marketplace add`)
- `.claude-plugin/plugin.json` — plugin definition (hooks, commands, schemas)
- `hooks/hooks.json` — two hooks: PostToolUse(Write) + PreToolUse(ExitPlanMode)
- `.claude-plugin/marketplace.json` — Claude Code marketplace registry
- `.claude-plugin/plugin.json` — Claude Code plugin definition
- `.codex-plugin/plugin.json` — Codex plugin definition
- `.agents/plugins/marketplace.json` — Codex marketplace registry
- `hooks/hooks.json` — Claude Code hooks: PostToolUse(Write) + PreToolUse(ExitPlanMode)
- `hooks.json` — Codex Stop hook
- `scripts/` — hook implementation (Python, stdlib only)
- `post_tool_hook.py` — records plan file path
- `pre_exit_plan_hook.py` — evaluates plan via codex
- `pre_exit_plan_hook.py` — evaluates plan via configured evaluator
- `hook_utils.py` — shared evaluation logic
- `evaluator.py` — codex subprocess wrapper
- `evaluator.py` — evaluator subprocess wrappers
- `state.py` — multi-round session state
- `config.py` — configuration loader
- `clear_state.py` — session cleanup utility
- `schemas/` — JSON output schema for codex structured output
- `commands/` — slash commands (`/planman:status`, `/planman:help`, `/planman:init`, `/planman:clear`)
- `schemas/` — JSON output schema for structured evaluator output
- `commands/` — Claude Code slash commands (`/planman:status`, `/planman:help`, `/planman:init`, `/planman:clear`)

## Uninstalling

```
/plugin uninstall planman@planman
```
Comment on lines 233 to 235

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Add a language tag to this fenced block.

This trips markdownlint (MD040) as written; bash would be enough here.

🧰 Tools
🪛 markdownlint-cli2 (0.22.1)

[warning] 233-233: Fenced code blocks should have a language specified

(MD040, fenced-code-language)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@README.md` around lines 233 - 235, The fenced code block containing the
command "/plugin uninstall planman@planman" in README.md lacks a language tag;
add a language identifier (e.g., "bash") immediately after the opening triple
backticks so the block becomes fenced as a bash snippet to satisfy markdownlint
MD040 and improve highlighting for the command.


This removes planman's hooks and commands. Your `.claude/planman.jsonc` config file is preserved — delete it manually if no longer needed.
This removes planman's hooks and commands. Your `.planman.jsonc` or legacy `.claude/planman.jsonc` config file is preserved — delete it manually if no longer needed.

## Troubleshooting

### "Nothing happens" when Claude presents a plan

1. **Check planman is installed**: Run `/planman:status` — it should show status and config
2. **Enable verbose mode**: Set `PLANMAN_VERBOSE=true` in your env or `.claude/planman.jsonc`
2. **Enable verbose mode**: Set `PLANMAN_VERBOSE=true` in your env or `.planman.jsonc`
3. **Check threshold**: A threshold of `0` disables evaluation. Set `PLANMAN_THRESHOLD=1` for testing

### I set `PLANMAN_VERBOSE=true` but see no output
Expand Down
2 changes: 1 addition & 1 deletion commands/clear.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
description: Clear planman session state (reset evaluation rounds)
---
Clear all planman session state files by running:
`python3 ${CLAUDE_PLUGIN_ROOT}/scripts/clear_state.py`
`python3 -c "import os, subprocess; root = os.environ.get('CLAUDE_PLUGIN_ROOT') or os.environ.get('CODEX_PLUGIN_ROOT') or os.getcwd(); subprocess.run(['python3', os.path.join(root, 'scripts', 'clear_state.py')])"`

## What Gets Cleared

Expand Down
Loading