diff --git a/.github/ISSUE_TEMPLATE/first-use.yml b/.github/ISSUE_TEMPLATE/first-use.yml
new file mode 100644
index 0000000..b5e4042
--- /dev/null
+++ b/.github/ISSUE_TEMPLATE/first-use.yml
@@ -0,0 +1,53 @@
+name: First-use feedback
+description: Tell us what you tried, where you got stuck, and whether the review helped.
+title: "[First use] "
+body:
+ - type: markdown
+ attributes:
+ value: |
+ Setup failures and unsupported export formats are useful feedback.
+ Share only a minimal redacted example. Please omit credentials,
+ private prompts, customer data and full production traces.
+ - type: dropdown
+ id: workflow
+ attributes:
+ label: Which workflow did you try?
+ options:
+ - Browser demonstration
+ - Released wheel and recorded comparison
+ - My own EvalArc evaluations
+ - AgentCore or another trace export
+ - Repeated judgments
+ - Candidate execution
+ - Something else
+ validations:
+ required: true
+ - type: textarea
+ id: goal
+ attributes:
+ label: What decision were you trying to make?
+ placeholder: Describe the task or change you wanted to review.
+ validations:
+ required: true
+ - type: textarea
+ id: result
+ attributes:
+ label: What happened?
+ description: Include the first confusing step or failure, and the result you expected.
+ validations:
+ required: true
+ - type: input
+ id: environment
+ attributes:
+ label: Version and environment
+ placeholder: EvalArc version, OS/Python or browser; export format if relevant.
+ - type: textarea
+ id: reproduction
+ attributes:
+ label: Minimal reproduction or redacted diagnostic
+ description: Commands or a small example are enough. Leave blank if you cannot share it.
+ - type: textarea
+ id: value
+ attributes:
+ label: Did the review help?
+ description: Did it reveal a problem, change a decision, or leave you unsure? Would you use it on another record?
diff --git a/CHANGELOG.md b/CHANGELOG.md
index d1b9b2e..a02bf24 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -1,5 +1,11 @@
# Changelog
+## Website and first-use documentation — 2026-09-16
+
+- Lead the English/Chinese README and evidence lab with the recorded regression, with a 30-second annotated UI walkthrough and a direct path to local review.
+- Document a source-free install from the released 0.12.1 wheel, recomputation of the comparison and suite acceptance checks, and the bounded AgentCore export workflow.
+- Add a first-use feedback form and replace historical version stacks in the HF card and outreach descriptions. This is a presentation and onboarding update; the Python version and released evidence remain unchanged.
+
## 0.12.1 — 2026-09-16
- Request attachment disposition for judge, suite and research ZIP links with `download=true`. Hugging Face's inline CDN redirects prevented the new ZIP link from triggering a download inside the Hub iframe; the attachment query was verified in the live browser.
diff --git a/README.md b/README.md
index 3b307b2..65b8ee4 100644
--- a/README.md
+++ b/README.md
@@ -1,172 +1,114 @@
-
+
-
- Open environments and evaluations for AI agents.
- Python 3.11+ · Linux host · No runtime dependencies · MIT · Research preview
-
+
Find the agent regression behind a better score.
+Review changed checks, follow the recorded actions, and hand off evidence someone else can verify.
-EvalArc is a research preview for **auditable agent evaluations**. It starts by
-checking whether a grader can distinguish correct work from plausible defects:
-run known-good and deliberately flawed submissions, inspect the evidence, and
-record exactly what was evaluated.
-
-**The score rose from 90% to 93.75%. A previously passing check now fails.**
-The [interactive evidence lab](https://huggingface.co/spaces/glayguo/evalarc)
-lets you compare revisions side by side, switch between correct and faulty
-implementations, and step through the tool call that changed the state. It replays the
-committed Docker audits without a model API or installation.
-Share the exact case and trace step with **Copy evidence link**, return from
-details to the case list, and retry failed sections independently. New links include
-SHA-256 of the loaded audit bytes: changed evidence is flagged before restoring a
-view, while legacy links disclose that the original audit identity is unknown.
-The fingerprint identifies content, not its author.
-[Explorer guide](docs/explorer.md).
-
-[](https://huggingface.co/spaces/glayguo/evalarc)
-
-**v0.8: verify the whole handoff.** Download the suite evidence ZIP from the
-lab, then run `evalarc verify suite-evidence --json` to recompute its original
-TOML, plan, five attempts, custom gates and JUnit. Add `--require-accepted`
-for CI acceptance. Consistency, configured acceptance and full resolution are
-reported separately. No candidate execution is required.
-[Offline verification and limits](docs/verification.md).
-
-**Three working task packs** share an evidence format and
-configurable candidate commands:
-
-| Task | Interaction | Host verification | Declared faults | Caught by one case |
-| --- | --- | --- | ---: | ---: |
-| `durable-kv` | Run a coding agent's completed service | Responses, transactions, restart durability | 8 | 3 |
-| `support-routing` | Drive a policy through simulated ticket tools | Routing, exact notes, closure, unrelated state, protocol | 7 | 2 |
-| `robot-evidence-review` | Report on attributed recording data | Coordinate and clock transforms, missing observations, source attribution | 6 | 1 |
-
-Every pack detects every declared fault: 21 faults, 21 detected. Six of the 21 are detected by a single case each, so the
-suite would lose coverage if that case were removed or stopped detecting its
-fault. A fresh audit would then lower the mutation score. Detection margins
-identify these dependencies before a change, alongside the current score.
-[How this relates to hack-verifiable environments](docs/methodology.md#relation-to-hack-verifiable-environments).
-
-v0.6 adds **Python and JavaScript workspace templates for both tasks**.
-Use `init --language javascript` for a starter or `--reference` for a scripted
-control, and `audit --language javascript` to check the same 15 fault models
-with independent Node.js implementations. JavaScript requires Node.js 22+;
-Docker runs explicitly select `--image node:22-slim`. See the
-[multilanguage guide](docs/languages.md).
-
-v0.5 adds `evalarc suite`: declare tasks, candidates, repeats, budgets, and
-acceptance gates in TOML. Preview the plan, execute all jobs, and inspect HTML,
-JSON, and JUnit results. Scores remain task-specific. See the
-[suite and CI guide](docs/suites.md).
-The [new acceptance-gate showcase](https://glayguo-evalarc.static.hf.space/#suite)
-uses one frozen faulty policy in two jobs: both score 93.75% with no resolved
-attempts. A permissive rule accepts the partial result; requiring every notes
-check rejects it. The full three-job Docker suite, five attempts, TOML and JUnit
-remain inspectable. Configured acceptance is separate from task resolution.
-
-[](https://glayguo-evalarc.static.hf.space/#suite)
-
-Prefer tables or Python? The [Hugging Face casebook](https://huggingface.co/datasets/glayguo/evalarc-casebook)
-separates 251 audit cases, six repeated attempts and three suite jobs into
-filterable configurations, with unchanged source JSON and provenance.
-Start with `suite_jobs` to compare `gate_accepted` and `fully_resolved`.
-These are scripted public-development records, not a held-out model benchmark.
-[Data guide and reproduction](docs/casebook.md).
-
-`evalarc repeat` freezes one candidate, runs fresh attempts on fixed
-cases, and reports every outcome with per-check pass rates. Runs now record
-JSONL progress, enforce a total case budget, and save bounded process diagnostics.
-See the [repeatability guide](docs/reliability.md).
-The [repeatability showcase](https://glayguo-evalarc.static.hf.space/#repeat)
-preserves three Docker attempts of each scripted control: the reference resolves
-3/3, while the duplicate-write policy resolves 0/3 despite its 93.75% mean score.
-Open every attempt's full evidence and per-check counts. No variation was observed;
-this is not a model reliability estimate.
-
-[](https://glayguo-evalarc.static.hf.space/#repeat)
-
-The workflow includes `evalarc doctor`, individual HTML reports, and `evalarc compare` for
-check regressions that a higher average score can hide. Every run preserves
-earlier outputs. See the [run-and-compare guide](docs/workflow.md).
-
-The support pack records tool calls and state changes, including retries after
-ambiguous write outcomes. Python and JavaScript scripted policies use the same
-host verifier. Browser environments, LLM-provider adapters, and RL training
-integrations remain planned. No frontier-model benchmark result is claimed.
-
-**Received a report? Verify it without running the candidate.**
-`evalarc verify path/to/report --json` checks evaluation, repetition or comparison
-evidence and fingerprints every input. Use `--require-resolved` when your handoff
-also requires all checks to pass. [Verification and limits](docs/verification.md).
-
-
-[](https://noteflowai.github.io/evalarc/#coverage)
-
-**Inspect coverage before trusting a perfect score.** The website now lists all three task packs and links each single-case dependency directly to its recorded checks and seeds. Offline audit reports provide the same disclosures without scripts or remote assets. These are new views of the original records, not new model runs.
-
-## Judge Stability — 0.12.0
-
-**Same trace. Same verdict?** Compare repeated saved judgments on one fixed
-recording. Keep score variation, pass/reject disagreement and incomplete
-assessments separate; inspect every value and verify preserved inputs offline.
-All-reject agreement remains rejection. [Interactive controls](https://noteflowai.github.io/evalarc/judge-stability/index.html)
-· [Local import guide](docs/judge-stability.md). The five controls are synthetic;
-this diagnostic does not rerun an agent or judge or establish calibration.
-
-[](https://noteflowai.github.io/evalarc/judge-stability/index.html)
-
-## Trace Workbench — 0.11.0
-
-Import saved AgentCore Evaluate responses, versioned golden cases and Skills Anywhere delivery receipts. Inspect zero scores, skipped judges, missing results and missed skills separately. Compare matching datasets/rubrics and verify preserved input bytes offline. [Try the authored controls](https://noteflowai.github.io/evalarc/trace-workbench/index.html) · [Actual local MCP delivery](https://noteflowai.github.io/evalarc/trace-mcp/index.html) · [Input contract](docs/trace-workbench.md). No live AWS evaluation is claimed.
+**90% → 93.75%. Two checks improve. One previously passing check fails.**
+A tool commits a note but returns an error. Retrying with a new key writes the
+note again. EvalArc exposes that regression instead of letting the higher
+average score settle the review.
-## New in 0.9.0: research you can inspect
+
+
+
+
-[Explore all 27 real GPU skill trials](https://noteflowai.github.io/evalarc/skill-impact/index.html) and [the research pilots](docs/research-pilots.md). Robot Reel's [captured-scene editor](https://noteflowai.github.io/robot-reel/scene-lab/) and [official LIBERO-Plus replay](https://noteflowai.github.io/robot-reel/libero-plus/) connect real source records with portable skill delivery and independent grading. Every failed attempt stays visible; no skill efficacy, full-benchmark or real-hardware result is implied.
+The demonstration replays saved Docker runs of scripted controls. No installation,
+account or model key is needed to explore it. **Research preview** · MIT ·
+Python 3.11+ · Linux for local workflows · no third-party Python runtime dependencies.
+## Start with one review
-## Run an audit
+1. **See the regression.** [Compare the two revisions](https://noteflowai.github.io/evalarc/#regression),
+ then inspect `retry-after-commit` in the [case explorer](https://noteflowai.github.io/evalarc/#explorer).
+2. **Check the decision.** [Compare the acceptance gates](https://noteflowai.github.io/evalarc/#suite):
+ the same 93.75% score passes a permissive rule and fails the strict notes rule.
+3. **Recompute it locally.** The [first-review walkthrough](docs/first-review.md)
+ installs the published wheel, downloads the records and rebuilds the comparison.
+ Verification exits 0 for consistency; comparison exits 1 for the regression.
-Clone the source, then install in an isolated Python environment:
+Install the released reviewer in a fresh virtual environment:
```bash
-git clone https://github.com/noteflowai/evalarc.git
-cd evalarc
python3 -m venv .venv
-source .venv/bin/activate
-python -m pip install -e .
+. .venv/bin/activate
+python -m pip install "https://github.com/noteflowai/evalarc/releases/download/v0.12.1/evalarc-0.12.1-py3-none-any.whl#sha256=115f3d8d452dee2b5d3aed12f736880aca69aaf3ea42937ef4f8f222c0b2291b"
+evalarc --version
+```
+
+The offline review needs no Docker, Node, GPU or model API. Follow the
+[download and comparison commands](docs/first-review.md#3-recompute-the-recorded-regression)
+to produce your first HTML report without cloning the source.
+
+## Bring your own work
+
+| What you need to review | Use EvalArc to | Start here |
+| --- | --- | --- |
+| A changed agent implementation | Compare matching evaluations and inspect regressed checks | [Run and compare](docs/workflow.md) |
+| Saved AgentCore Evaluate results and spans | Inspect valid zero scores, skipped judgments, missing results and skill delivery | [Export-to-review walkthrough](docs/agentcore-first-review.md) |
+| Repeated judgments on one fixed recording | Separate score variation, verdict disagreement and incomplete assessments | [Judge Stability](docs/judge-stability.md) |
+| A report received from another developer | Recompute summaries, configured gates and JUnit from original inputs | [Offline verification](docs/verification.md) |
+| A grader or candidate you want to execute | Run a reference and deliberate faults against a task contract | [Run an audit](#run-an-audit) |
+
+Trace import accepts a [bounded export format](docs/trace-workbench.md), not arbitrary
+cloud exports. Its scored controls are synthetic; the separate MCP example records
+actual local delivery with no evaluator scores. No live AgentCore evaluation is claimed.
+
+Trying your own records? [Tell us where the first review helped or got stuck](https://github.com/noteflowai/evalarc/issues/new?template=first-use.yml).
+A minimal redacted example is enough; a failed setup is useful feedback too.
+
+## What the recorded evidence covers
+
+| Task | Interaction | Declared faults | Detected by only one case |
+| --- | --- | ---: | ---: |
+| `durable-kv` | Coding artifact: responses, transactions and restart durability | 8 | 3 |
+| `support-routing` | Simulated ticket tools: routing, exact notes, closure and unrelated state | 7 | 2 |
+| `robot-evidence-review` | Attributed recordings: coordinates, clocks and missing observations | 6 | 1 |
+
+The saved audits detect **21/21 declared faults across three task packs**. Six
+faults depend on one detecting case each. Removing a sole detector lowers a fresh
+audit's mutation score; these margins expose that dependency before the change.
+They do not establish coverage of unseen faults.
+[Inspect coverage](https://noteflowai.github.io/evalarc/#coverage) ·
+[Methodology](docs/methodology.md) · [251 audit case records](docs/casebook.md).
+
+For deeper exploration: [repeated attempts](docs/reliability.md),
+[TOML suites and CI](docs/suites.md), [Python/JavaScript candidates](docs/languages.md),
+[recorded GPU research pilots](docs/research-pilots.md),
+[architecture](docs/architecture.md) and [papers](docs/research.zh-CN.md).
+Feature history lives in the [changelog](CHANGELOG.md).
+
+## Run an audit
+
+To execute the built-in Python reference and eight deliberate coding faults,
+install the wheel above, then use Docker:
+
+```bash
docker pull python:3.12-slim
evalarc audit --seeds 17 41 97 --output runs/audit
```
-The command evaluates the reference and eight negative controls, writes
-`runs/audit/audit.json`, and creates a standalone `runs/audit/index.html` report.
-Exit code `0` means the reference passed and every declared defect was detected
-in its intended dimension. Exit code `1` means an audit or candidate failed;
-`2` indicates a usage/configuration error or an invalid run caused by an
-environment failure.
+Open `runs/audit/index.html`. Exit 0 means the reference passed and all declared
+faults were detected in their intended dimensions; 1 means an audit/candidate
+failed, and 2 means invalid input or an environment failure.
-For the bundled, trusted controls, a faster CPU-only demo is:
+For the bundled trusted controls, this shorter CPU-only run uses the host:
```bash
-evalarc audit --backend local --trust-local --output runs/local-audit
evalarc audit --task support-routing --backend local --trust-local --output runs/support-audit
```
-Local execution has the host user's privileges. Use Docker for candidate
-isolation and read the [execution boundaries](SECURITY.md).
-If your Docker setup requires a wrapper, set `EVALARC_DOCKER` to that command
-or pass `--docker-command`.
+Local execution has your user privileges. Use Docker for candidate isolation;
+see [execution boundaries](SECURITY.md) and [readiness checks](docs/workflow.md).
+Use a fresh output path for another run. To develop EvalArc itself, see
+[Development](#development).
## Coding task
@@ -286,7 +228,13 @@ failures must be established before this becomes a research benchmark.
## Development
+Clone the source for development and for the `examples/` commands in this README:
+
```bash
+git clone https://github.com/noteflowai/evalarc.git
+cd evalarc
+python3 -m venv .venv
+. .venv/bin/activate
python -m pip install -e ".[dev]"
pytest -q
ruff check .
@@ -297,4 +245,4 @@ python -m build
See [CONTRIBUTING.md](CONTRIBUTING.md), [SECURITY.md](SECURITY.md), and
[LICENSE](LICENSE). The [task-author guide](docs/task-authoring.md) explains
the current built-in extension points. CI includes Python checks, the Node
-policy, and Docker audits for both packs.
+policies, and Docker audits for all three task packs.
diff --git a/README.zh-CN.md b/README.zh-CN.md
index 86dad3a..8fa1caf 100644
--- a/README.zh-CN.md
+++ b/README.zh-CN.md
@@ -1,146 +1,101 @@
-
+
# EvalArc
-**面向 AI 智能体的开放任务环境与可审计评测。**
+**找出更高评分背后的 Agent 回归。**
-Python 3.11+,Linux 主机,零运行时第三方依赖,MIT 许可证。
+查看退步的检查,定位已记录的工具动作,交付他人可以复核的证据。
-[在线交互演示](https://huggingface.co/spaces/glayguo/evalarc) ·
-[网页镜像](https://noteflowai.github.io/evalarc/) ·
-[可筛选证据数据集](https://huggingface.co/datasets/glayguo/evalarc-casebook) ·
-[版本下载](https://github.com/noteflowai/evalarc/releases) ·
-[English](README.md) · [中文调研与论文分析](docs/research.zh-CN.md) ·
-[架构设计](docs/architecture.md) · [方法说明](docs/methodology.md) · [开发路线](docs/roadmap.md)
+[**立即查看失败案例 →**](https://noteflowai.github.io/evalarc/#regression) ·
+[首次本地复核](docs/first-review.zh-CN.md) ·
+[Hugging Face 演示](https://huggingface.co/spaces/glayguo/evalarc) · [English](README.md)
-EvalArc 关注智能体实际完成的结果,以及支撑评分结论的证据。
-通过正确实现和刻意带有缺陷的实现进行对照,
-检查它能发现哪些问题,并保存可复查的报告。
+**90% → 93.75%。两项检查改善,一项原本通过的检查却失败了。**
+工具已写入备注,却返回错误;策略更换幂等键后重试,多写了一次。
+EvalArc 把这次退步展开给你看,帮助判断分数提高是否满足发布要求。
-**分数从 90% 升到 93.75%,原本通过的检查却失败了。**
-[交互证据实验室](https://huggingface.co/spaces/glayguo/evalarc)
-可以并排比较两个版本,查看一处退步、两处改进,再逐步检查工具调用和状态变化。
-页面读取仓库保存的 Docker 审计记录,无需安装,也不调用模型。
-使用 **Copy evidence link** 分享具体案例和步骤;打开详情后可返回原案例,
-某一分区加载失败时可单独重试。新链接携带实际加载审计文件的 SHA-256,
-记录变化时会先提示并停止恢复;旧链接明确说明未保存原始审计身份。
-指纹标识内容,不认证作者。[交互使用说明](docs/explorer.md)。
+
+
+
+
-[](https://huggingface.co/spaces/glayguo/evalarc)
+演示回放的是脚本对照在 Docker 中运行后保存的记录,无需安装、账号或模型密钥。
+研究预览 · MIT · Python 3.11+ · 本地流程使用 Linux · Python 包无第三方运行时依赖。
-**v0.8:完整验收证据可离线复核。** 从首页下载套件 ZIP,解压后运行
-`evalarc verify suite-evidence --json`,复核原始 TOML、执行计划、五次尝试、
-验收规则与 JUnit。使用 `--require-accepted` 接入 CI 验收;证据一致、规则接受、
-任务完全完成分别报告。整个检查不执行候选程序。
-[离线验证流程与边界](docs/verification.md)。
+## 从一次复核开始
-**已实现 coding、业务工具与记录复核三个场景。**
+1. **看退步。** [对照两个版本](https://noteflowai.github.io/evalarc/#regression),
+ 再在[案例浏览器](https://noteflowai.github.io/evalarc/#explorer)查看 `retry-after-commit`。
+2. **看验收。** [比较两条规则](https://noteflowai.github.io/evalarc/#suite):
+ 相同的 93.75% 分数,宽松规则接受,严格备注规则拒绝。
+3. **本地复算。** [首次复核指南](docs/first-review.zh-CN.md)使用发布的 wheel 和原始记录重建报告;
+ 复核退出 0 表示一致,对照退出 1 表示发现退步。
-| 任务 | 交互方式 | 验证内容 | 声明缺陷 | 仅单个用例检出 |
-| --- | --- | --- | ---: | ---: |
-| `durable-kv` | 执行代码智能体交付的服务 | 读写、事务、CAS、持久化与异常恢复 | 8 | 3 |
-| `support-routing` | 策略通过工具操作模拟工单 | 路由、精确备注、条件关闭、无关数据保护与协议完成 | 7 | 2 |
-| `robot-evidence-review` | 基于有出处的记录数据出报告 | 坐标与时钟换算、缺失观测、来源归属 | 6 | 1 |
+在新的虚拟环境安装已发布的复核工具:
-三个任务包都检出了全部声明缺陷:21 个缺陷,21 个检出。其中六个各自只靠一个用例检出;删除该用例,或使其无法再检出对应缺陷,就会失去这部分覆盖,重新审计的变异分数也会下降。检出余量在改动之前就能指出这些依赖,与当前分数一起报告。[与 hack-verifiable environments 的关系](docs/methodology.md#relation-to-hack-verifiable-environments)。
-
-v0.6 为两个任务都提供 **Python 和 JavaScript 工作区模板**:
-`init --language javascript` 生成起步代码,添加 `--reference` 生成脚本对照;
-`audit --language javascript` 使用独立的 Node.js 实现检查相同的 21 类故障,实测检出余量与 Python 完全一致。
-JavaScript 需要 Node.js 22+,Docker 模式显式指定 `--image node:22-slim`。
-详见[多语言接入指南](docs/languages.zh-CN.md)。
-
-v0.5 新增 `evalarc suite`:用 TOML 声明任务、候选、轮次、预算及验收门槛,
-先预览执行计划,再批量运行并输出 HTML、JSON 和 JUnit。各任务单独评分。
-详见[套件与 CI 指南](docs/suites.zh-CN.md)。
-
-[新增验收规则展示](https://glayguo-evalarc.static.hf.space/#suite)让同一份缺陷策略分别按两套规则验收:
-都是 93.75%、0/2 轮完全通过,宽松规则允许部分进展,要求备注检查全部通过的规则则拒绝。
-完整三项评测作业的 Docker 记录、五轮尝试、TOML 和 JUnit 均可检查;“规则接受”与“任务完全完成”
-分别展示。
-
-[](https://glayguo-evalarc.static.hf.space/#suite)
-
-[Hugging Face Casebook](https://huggingface.co/datasets/glayguo/evalarc-casebook)
-把 251 条审计用例、6 次重复尝试和 3 项验收作业分别整理为可筛选的表,
-保留未经改写的原始 JSON 和版本指纹。可先选择 `suite_jobs`,
-对照 `gate_accepted` 与 `fully_resolved`,或用 Python 读取。
-这些是公开开发任务中的脚本对照,不是隐藏模型测试集。详见[数据说明](docs/casebook.md)。
-
-`evalarc repeat` 固定一份候选快照,在相同场景上重新启动多轮评测,
-保存每轮证据并显示逐项通过率与结果波动。同时补齐场景总时间预算、JSONL 进度
-和受限进程诊断。详见[重复评测指南](docs/reliability.zh-CN.md)。
-
-[重复评测展示](https://glayguo-evalarc.static.hf.space/#repeat)保留了两种脚本策略各三次
-Docker 评测:参考策略 3/3 轮完全通过,重复写入策略虽然平均分为 93.75%,却 0/3 轮
-完全通过。可以查看每轮原始证据和逐项计数。这些观察中未出现检查结果波动,
-也不能据此估计模型在未见任务上的可靠性。
-
-[](https://glayguo-evalarc.static.hf.space/#repeat)
-
-现有流程包含环境预检查、单次评测 HTML 报告和逐项回归比较,即使总分上升也能指出
-退步的检查;输出保护会保留之前的运行证据。详见[使用流程](docs/workflow.zh-CN.md)。
-
-两个场景使用共同的报告元数据,各自定义评分规则。工单环境记录工具调用和
-真实状态变化,支持写入前失败、写入后响应失败及幂等重试。
-已提供 Python 和 JavaScript 策略;浏览器环境、真实模型 API 适配与 RL 集成
-仍在路线图中。具体边界见[架构设计](docs/architecture.md)。
-
-当前版本尚未经过前沿模型、真实人类工时或强化学习收益标定。
-
-**收到报告后,先独立复核。** `evalarc verify 报告路径 --json` 无需执行候选程序,
-即可重算单次评测、重复运行和前后对照的汇总,并记录每份输入的指纹。
-需要全部任务通过时加 `--require-resolved`。
-[使用说明与校验范围](docs/verification.md)。
-
-
-[](https://noteflowai.github.io/evalarc/#coverage)
+```bash
+python3 -m venv .venv
+. .venv/bin/activate
+python -m pip install "https://github.com/noteflowai/evalarc/releases/download/v0.12.1/evalarc-0.12.1-py3-none-any.whl#sha256=115f3d8d452dee2b5d3aed12f736880aca69aaf3ea42937ef4f8f222c0b2291b"
+evalarc --version
+```
-**在信任满分前,先检查覆盖薄弱点。** 网页现展示全部三个任务包,可从仅靠一个用例检出的缺陷直接定位到原始检查与种子记录。离线审计报告也提供相同的证据展开入口,不依赖脚本或远程资源;新增界面沿用原始数据,不冒充新模型运行。
+离线复核不需要 Docker、Node、GPU 或模型 API。继续按
+[下载与对照步骤](docs/first-review.zh-CN.md#3-复算这次退步)生成 HTML 报告,无需克隆源码。
-## 0.11.0:运行记录评估工作台
+## 用于自己的工作
-导入已保存的 AgentCore Evaluate 结果、版本化黄金案例和 Skills Anywhere 加载回执,分别查看有效零分、评估跳过、缺少结果及漏调用技能。支持相同测试集与评分规则下的对比,以及原始输入的离线复核。[交互示例](https://noteflowai.github.io/evalarc/trace-workbench/index.html) · [真实本地 MCP 加载](https://noteflowai.github.io/evalarc/trace-mcp/index.html) · [数据契约](docs/trace-workbench.md)。示例明确区分合成评分与真实加载记录,未运行云端评估。
+| 已有材料 | 可以检查什么 | 入口 |
+| --- | --- | --- |
+| 变更前后的 Agent 评测 | 匹配条件下哪些检查退步 | [运行与对照](docs/workflow.md) |
+| AgentCore Evaluate 结果与 spans | 有效零分、跳过、缺失以及技能交付 | [导出到审阅](docs/agentcore-first-review.md) |
+| 同一记录上的多次评判 | 分数变化、通过/拒绝翻转及未评判情况 | [中文指南](docs/judge-stability.zh-CN.md) |
+| 他人交付的报告 | 原始输入与汇总、验收规则、JUnit 是否一致 | [离线复核](docs/verification.md) |
+| 评分器或待执行候选 | 正确实现和刻意缺陷是否被评分器区分 | [运行审计](#运行审计) |
-## v0.12.0:同一记录,多次评判
+运行记录导入有明确的[受限格式约定](docs/trace-workbench.md),不直接接受任意云端导出。
+带分数的示例为合成数据,独立 MCP 示例是真实本地加载但没有评委分数;未完成实时 AgentCore 评测。
-新增 `trace-stability`,固定执行记录与评估规则,分别检查分数变化、通过/拒绝翻转、
-跳过与缺失,并保存原始输入供离线复核。三次都拒绝仍然是拒绝,不把一致性当作
-任务成功。[在线交互示例](https://noteflowai.github.io/evalarc/judge-stability/index.html)
-· [中文使用指南](docs/judge-stability.zh-CN.md)。五个展示案例均为人工构造;
-此功能不重新运行 Agent 或裁判,也不代替独立人工校准。
+尝试自己的记录后,欢迎[反馈首次使用的卡点或发现](https://github.com/noteflowai/evalarc/issues/new?template=first-use.yml)。
+最小脱敏样例即可;安装失败同样有价值。
-[](https://noteflowai.github.io/evalarc/judge-stability/index.html)
+## 原始证据覆盖什么
-## 0.9.0:有原始证据的研究场景
+| 任务 | 交互场景 | 声明缺陷 | 仅一个用例检出 |
+| --- | --- | ---: | ---: |
+| `durable-kv` | 代码交付物:响应、事务、重启持久化 | 8 | 3 |
+| `support-routing` | 模拟工单:路由、精确备注、关闭与无关状态保护 | 7 | 2 |
+| `robot-evidence-review` | 有出处的记录:坐标、时钟与缺失观察 | 6 | 1 |
-[查看 27 次真实 GPU 技能评测](https://noteflowai.github.io/evalarc/skill-impact/index.html),并阅读[完整方法与限制](docs/research-pilots.md)。新增[实景 Blender 编辑](https://noteflowai.github.io/robot-reel/scene-lab/)与[官方 LIBERO-Plus 子集回放](https://noteflowai.github.io/robot-reel/libero-plus/),把原始记录、技能交付与独立验收连接起来。失败尝试全部保留;不宣称技能提分、完整基准成绩或真机效果。
+三个任务包保存的审计检出 **21/21 种声明缺陷**。其中六种各依赖一个检测用例;
+移除唯一检测用例后,重新审计的变异分数会下降。覆盖余量用于提前暴露这种依赖,
+不代表覆盖未知缺陷。[检查覆盖](https://noteflowai.github.io/evalarc/#coverage) ·
+[方法说明](docs/methodology.md) · [251 条审计记录](docs/casebook.md)。
+更多工作流:[重复运行](docs/reliability.md)、[TOML 套件与 CI](docs/suites.md)、
+[Python/JavaScript](docs/languages.md)、[GPU 研究记录](docs/research-pilots.md)、
+[架构](docs/architecture.md)和[论文分析](docs/research.zh-CN.md)。历史更新见 [CHANGELOG](CHANGELOG.md)。
-## 直接运行
+## 运行审计
-克隆仓库后,在独立 Python 环境中安装:
+安装上面的 wheel 后,用 Docker 执行内置 Python 参考实现和八种刻意带错的代码实现:
```bash
-git clone https://github.com/noteflowai/evalarc.git
-cd evalarc
-python3 -m venv .venv
-source .venv/bin/activate
-python -m pip install -e .
docker pull python:3.12-slim
evalarc audit --seeds 17 41 97 --output runs/audit
```
-输出 `runs/audit/audit.json` 和可独立打开的 `runs/audit/index.html`。
-对仓库自带的可信对照代码,可以运行更快的本机演示:
+打开 `runs/audit/index.html`。退出 0 表示参考实现通过且所有声明缺陷被检出;
+1 表示审计或候选失败;2 表示参数或环境导致结果无效。
+
+对包内可信对照,可以在 CPU 上直接运行:
```bash
-evalarc audit --backend local --trust-local --output runs/local-audit
evalarc audit --task support-routing --backend local --trust-local --output runs/support-audit
```
-本机模式具有当前用户的文件和网络权限。Docker 模式的边界见
-[SECURITY.md](SECURITY.md)。无模型调用,无 API 费用,无 GPU 依赖。
+本机模式具有当前用户权限;候选隔离使用 Docker,边界见 [SECURITY.md](SECURITY.md)。
+重复运行时使用新的输出路径。
## 已实现
diff --git a/docs/agentcore-first-review.md b/docs/agentcore-first-review.md
new file mode 100644
index 0000000..371224d
--- /dev/null
+++ b/docs/agentcore-first-review.md
@@ -0,0 +1,98 @@
+# From an AgentCore export to a local review
+
+This walkthrough connects EvalArc's existing `trace-import` command to your
+review process. It starts with published controls, then explains which pieces
+to replace with actual records. It does not deploy a cloud agent or collect
+AWS data for you.
+
+**Validation scope:** the first half is reproducible with the released 0.12.1
+wheel. The five scored controls are synthetic API-shaped inputs. A separate
+record contains actual local MCP delivery and no evaluator scores. A live
+AgentCore or Strands evaluation has not been validated by these examples.
+
+## Inspect a complete input before collecting data
+
+After [installing the wheel](first-review.md#2-install-the-reviewer), run
+these commands in a fresh directory:
+
+```bash
+curl --fail --location \
+ https://github.com/noteflowai/evalarc/releases/download/v0.12.1/trace-workbench-evidence.zip \
+ --output trace-workbench-evidence.zip
+python -m zipfile -e trace-workbench-evidence.zip .
+evalarc trace-import trace-workbench/input.json \
+ --baseline trace-workbench/baseline-input.json \
+ --output my-trace-review
+evalarc trace-verify my-trace-review
+```
+
+Both commands exit **0** for a valid import and consistent evidence. That does
+not mean every case was accepted. Open `my-trace-review/index.html` to see the
+accepted, zero-scored, skipped, missing-result and missing-skill controls.
+The ZIP SHA-256 is
+`2e8ea5ba062167e8de5420400e60837a1aae68e465b5cfdf5337686d6fe2be3f`.
+
+The copied `my-trace-review/input.json` and `baseline-input.json` preserve the
+original bytes; `review.json` contains the derived review. The report and
+filters work offline.
+
+## Replace authored inputs with a recorded export
+
+Use `trace-workbench/input.json` as a structural example. Create your own
+wrapper rather than relabeling the demonstration as a real run:
+
+1. Record the golden set, ordered case IDs, intended goals and required skills.
+ Set `provenance.kind` to `recorded` and describe how the data was collected
+ and what the collection omits.
+2. Copy each case's native **AgentCore Evaluate API** `evaluationResults` into
+ `evaluation_response`. Keep the service's original fields and matching
+ session/trace/span identifiers.
+3. Supply a flat array of corresponding OpenTelemetry spans. Extract spans
+ from any transport envelope first; arbitrary OTLP envelopes and CloudWatch
+ query wrappers are not accepted directly.
+4. Explicitly declare the evaluator's target level, revision, rating scale
+ and acceptance rule. Keep model/configuration identity and hashes where
+ required. An evaluator name does not establish its scale or version.
+5. If using Skills Anywhere, preserve actual `open_skill` receipts and
+ annotate the matching tool spans. Only assert complete skill observation
+ when your collector actually observed all loads.
+
+The AWS workshop CLI's `run.results[].sessionScores` is a different envelope.
+Renaming that field is not a conversion. Strands traces alone also do not
+provide the required golden cases and judgments. See the
+[full input contract](trace-workbench.md#prepare-a-real-export) before adapting
+an exporter.
+
+## Review, compare, then gate
+
+```bash
+evalarc trace-import recorded-current.json --output recorded-review
+evalarc trace-verify recorded-review
+```
+
+For a baseline comparison, add `--baseline recorded-baseline.json`. Both
+inputs must use the same complete golden set and evaluator definitions;
+synthetic inputs cannot be compared with recorded inputs.
+
+Add `--require-accepted` to **trace-import** when a CI step should enforce the
+rules. Its exit codes distinguish **0** (every case accepted), **1** (complete
+assessments with rejected gates), and **2** (incomplete assessments or invalid
+input). A well-formed input still produces its report when a gate rejects or
+a judgment is missing. `trace-verify` checks consistency; it is not the gate.
+
+If importing fails, start with the named case/evaluator in the diagnostic.
+Check linkage, duplicate targets, explicit scale bounds and expected-result
+coverage. A missing or skipped evaluation must not be filled with a fabricated
+zero to make the file pass validation.
+
+## Bring back one useful finding
+
+Use the [first-use form](https://github.com/noteflowai/evalarc/issues/new?template=first-use.yml)
+to report the exporter/envelope you used, the command and the first point of
+friction. A small redacted fixture helps establish whether an adapter is
+useful. Do not attach production traces merely to demonstrate interest.
+
+Sources and collection boundaries are documented in the
+[Trace Workbench guide](trace-workbench.md#scope-and-sources). EvalArc reviews
+imported judgments; it does not independently establish their truth or the
+completeness of the supplied execution record.
diff --git a/docs/assets/banner.svg b/docs/assets/banner.svg
index a52f3ae..e3a2969 100644
--- a/docs/assets/banner.svg
+++ b/docs/assets/banner.svg
@@ -1,14 +1,14 @@
-