diff --git a/.github/ISSUE_TEMPLATE/first-use.yml b/.github/ISSUE_TEMPLATE/first-use.yml new file mode 100644 index 0000000..b5e4042 --- /dev/null +++ b/.github/ISSUE_TEMPLATE/first-use.yml @@ -0,0 +1,53 @@ +name: First-use feedback +description: Tell us what you tried, where you got stuck, and whether the review helped. +title: "[First use] " +body: + - type: markdown + attributes: + value: | + Setup failures and unsupported export formats are useful feedback. + Share only a minimal redacted example. Please omit credentials, + private prompts, customer data and full production traces. + - type: dropdown + id: workflow + attributes: + label: Which workflow did you try? + options: + - Browser demonstration + - Released wheel and recorded comparison + - My own EvalArc evaluations + - AgentCore or another trace export + - Repeated judgments + - Candidate execution + - Something else + validations: + required: true + - type: textarea + id: goal + attributes: + label: What decision were you trying to make? + placeholder: Describe the task or change you wanted to review. + validations: + required: true + - type: textarea + id: result + attributes: + label: What happened? + description: Include the first confusing step or failure, and the result you expected. + validations: + required: true + - type: input + id: environment + attributes: + label: Version and environment + placeholder: EvalArc version, OS/Python or browser; export format if relevant. + - type: textarea + id: reproduction + attributes: + label: Minimal reproduction or redacted diagnostic + description: Commands or a small example are enough. Leave blank if you cannot share it. + - type: textarea + id: value + attributes: + label: Did the review help? + description: Did it reveal a problem, change a decision, or leave you unsure? Would you use it on another record? diff --git a/CHANGELOG.md b/CHANGELOG.md index d1b9b2e..a02bf24 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,11 @@ # Changelog +## Website and first-use documentation — 2026-09-16 + +- Lead the English/Chinese README and evidence lab with the recorded regression, with a 30-second annotated UI walkthrough and a direct path to local review. +- Document a source-free install from the released 0.12.1 wheel, recomputation of the comparison and suite acceptance checks, and the bounded AgentCore export workflow. +- Add a first-use feedback form and replace historical version stacks in the HF card and outreach descriptions. This is a presentation and onboarding update; the Python version and released evidence remain unchanged. + ## 0.12.1 — 2026-09-16 - Request attachment disposition for judge, suite and research ZIP links with `download=true`. Hugging Face's inline CDN redirects prevented the new ZIP link from triggering a download inside the Hub iframe; the attachment query was verified in the live browser. diff --git a/README.md b/README.md index 3b307b2..65b8ee4 100644 --- a/README.md +++ b/README.md @@ -1,172 +1,114 @@ -

EvalArc — Run agents. Measure outcomes.

+

EvalArc — Higher score. New failure.

-

- Open environments and evaluations for AI agents.
- Python 3.11+ · Linux host · No runtime dependencies · MIT · Research preview -

+

Find the agent regression behind a better score.
+Review changed checks, follow the recorded actions, and hand off evidence someone else can verify.

- Interactive evidence lab · - Web demo · - Filterable casebook · - Releases · - 简体中文 · - Research & papers · - Architecture · - Methodology · - Roadmap + Try the recorded failure → · + First local review · + Hugging Face · + 简体中文

-EvalArc is a research preview for **auditable agent evaluations**. It starts by -checking whether a grader can distinguish correct work from plausible defects: -run known-good and deliberately flawed submissions, inspect the evidence, and -record exactly what was evaluated. - -**The score rose from 90% to 93.75%. A previously passing check now fails.** -The [interactive evidence lab](https://huggingface.co/spaces/glayguo/evalarc) -lets you compare revisions side by side, switch between correct and faulty -implementations, and step through the tool call that changed the state. It replays the -committed Docker audits without a model API or installation. -Share the exact case and trace step with **Copy evidence link**, return from -details to the case list, and retry failed sections independently. New links include -SHA-256 of the loaded audit bytes: changed evidence is flagged before restoring a -view, while legacy links disclose that the original audit identity is unknown. -The fingerprint identifies content, not its author. -[Explorer guide](docs/explorer.md). - -[![EvalArc v0.3: score rises from 90% to 93.75% while a check regresses](docs/assets/regression-lab.png)](https://huggingface.co/spaces/glayguo/evalarc) - -**v0.8: verify the whole handoff.** Download the suite evidence ZIP from the -lab, then run `evalarc verify suite-evidence --json` to recompute its original -TOML, plan, five attempts, custom gates and JUnit. Add `--require-accepted` -for CI acceptance. Consistency, configured acceptance and full resolution are -reported separately. No candidate execution is required. -[Offline verification and limits](docs/verification.md). - -**Three working task packs** share an evidence format and -configurable candidate commands: - -| Task | Interaction | Host verification | Declared faults | Caught by one case | -| --- | --- | --- | ---: | ---: | -| `durable-kv` | Run a coding agent's completed service | Responses, transactions, restart durability | 8 | 3 | -| `support-routing` | Drive a policy through simulated ticket tools | Routing, exact notes, closure, unrelated state, protocol | 7 | 2 | -| `robot-evidence-review` | Report on attributed recording data | Coordinate and clock transforms, missing observations, source attribution | 6 | 1 | - -Every pack detects every declared fault: 21 faults, 21 detected. Six of the 21 are detected by a single case each, so the -suite would lose coverage if that case were removed or stopped detecting its -fault. A fresh audit would then lower the mutation score. Detection margins -identify these dependencies before a change, alongside the current score. -[How this relates to hack-verifiable environments](docs/methodology.md#relation-to-hack-verifiable-environments). - -v0.6 adds **Python and JavaScript workspace templates for both tasks**. -Use `init --language javascript` for a starter or `--reference` for a scripted -control, and `audit --language javascript` to check the same 15 fault models -with independent Node.js implementations. JavaScript requires Node.js 22+; -Docker runs explicitly select `--image node:22-slim`. See the -[multilanguage guide](docs/languages.md). - -v0.5 adds `evalarc suite`: declare tasks, candidates, repeats, budgets, and -acceptance gates in TOML. Preview the plan, execute all jobs, and inspect HTML, -JSON, and JUnit results. Scores remain task-specific. See the -[suite and CI guide](docs/suites.md). -The [new acceptance-gate showcase](https://glayguo-evalarc.static.hf.space/#suite) -uses one frozen faulty policy in two jobs: both score 93.75% with no resolved -attempts. A permissive rule accepts the partial result; requiring every notes -check rejects it. The full three-job Docker suite, five attempts, TOML and JUnit -remain inspectable. Configured acceptance is separate from task resolution. - -[![EvalArc v0.5: the same score meets one gate and fails another](docs/assets/suite-lab.png)](https://glayguo-evalarc.static.hf.space/#suite) - -Prefer tables or Python? The [Hugging Face casebook](https://huggingface.co/datasets/glayguo/evalarc-casebook) -separates 251 audit cases, six repeated attempts and three suite jobs into -filterable configurations, with unchanged source JSON and provenance. -Start with `suite_jobs` to compare `gate_accepted` and `fully_resolved`. -These are scripted public-development records, not a held-out model benchmark. -[Data guide and reproduction](docs/casebook.md). - -`evalarc repeat` freezes one candidate, runs fresh attempts on fixed -cases, and reports every outcome with per-check pass rates. Runs now record -JSONL progress, enforce a total case budget, and save bounded process diagnostics. -See the [repeatability guide](docs/reliability.md). -The [repeatability showcase](https://glayguo-evalarc.static.hf.space/#repeat) -preserves three Docker attempts of each scripted control: the reference resolves -3/3, while the duplicate-write policy resolves 0/3 despite its 93.75% mean score. -Open every attempt's full evidence and per-check counts. No variation was observed; -this is not a model reliability estimate. - -[![EvalArc v0.4: three 93.75% attempts, zero fully resolved runs](docs/assets/repeat-lab.png)](https://glayguo-evalarc.static.hf.space/#repeat) - -The workflow includes `evalarc doctor`, individual HTML reports, and `evalarc compare` for -check regressions that a higher average score can hide. Every run preserves -earlier outputs. See the [run-and-compare guide](docs/workflow.md). - -The support pack records tool calls and state changes, including retries after -ambiguous write outcomes. Python and JavaScript scripted policies use the same -host verifier. Browser environments, LLM-provider adapters, and RL training -integrations remain planned. No frontier-model benchmark result is claimed. - -**Received a report? Verify it without running the candidate.** -`evalarc verify path/to/report --json` checks evaluation, repetition or comparison -evidence and fingerprints every input. Use `--require-resolved` when your handoff -also requires all checks to pass. [Verification and limits](docs/verification.md). - - -[![Three task packs with direct evidence links for six single-case dependencies.](docs/coverage-review.png)](https://noteflowai.github.io/evalarc/#coverage) - -**Inspect coverage before trusting a perfect score.** The website now lists all three task packs and links each single-case dependency directly to its recorded checks and seeds. Offline audit reports provide the same disclosures without scripts or remote assets. These are new views of the original records, not new model runs. - -## Judge Stability — 0.12.0 - -**Same trace. Same verdict?** Compare repeated saved judgments on one fixed -recording. Keep score variation, pass/reject disagreement and incomplete -assessments separate; inspect every value and verify preserved inputs offline. -All-reject agreement remains rejection. [Interactive controls](https://noteflowai.github.io/evalarc/judge-stability/index.html) -· [Local import guide](docs/judge-stability.md). The five controls are synthetic; -this diagnostic does not rerun an agent or judge or establish calibration. - -[![Five synthetic controls separate score changes, gate flips and unavailable judgments](docs/assets/judge-stability.png)](https://noteflowai.github.io/evalarc/judge-stability/index.html) - -## Trace Workbench — 0.11.0 - -Import saved AgentCore Evaluate responses, versioned golden cases and Skills Anywhere delivery receipts. Inspect zero scores, skipped judges, missing results and missed skills separately. Compare matching datasets/rubrics and verify preserved input bytes offline. [Try the authored controls](https://noteflowai.github.io/evalarc/trace-workbench/index.html) · [Actual local MCP delivery](https://noteflowai.github.io/evalarc/trace-mcp/index.html) · [Input contract](docs/trace-workbench.md). No live AWS evaluation is claimed. +**90% → 93.75%. Two checks improve. One previously passing check fails.** +A tool commits a note but returns an error. Retrying with a new key writes the +note again. EvalArc exposes that regression instead of letting the higher +average score settle the review. -## New in 0.9.0: research you can inspect + + + Recorded walkthrough: the score rises, a retry duplicates a note, and the strict acceptance gate rejects the policy. + -[Explore all 27 real GPU skill trials](https://noteflowai.github.io/evalarc/skill-impact/index.html) and [the research pilots](docs/research-pilots.md). Robot Reel's [captured-scene editor](https://noteflowai.github.io/robot-reel/scene-lab/) and [official LIBERO-Plus replay](https://noteflowai.github.io/robot-reel/libero-plus/) connect real source records with portable skill delivery and independent grading. Every failed attempt stays visible; no skill efficacy, full-benchmark or real-hardware result is implied. +The demonstration replays saved Docker runs of scripted controls. No installation, +account or model key is needed to explore it. **Research preview** · MIT · +Python 3.11+ · Linux for local workflows · no third-party Python runtime dependencies. +## Start with one review -## Run an audit +1. **See the regression.** [Compare the two revisions](https://noteflowai.github.io/evalarc/#regression), + then inspect `retry-after-commit` in the [case explorer](https://noteflowai.github.io/evalarc/#explorer). +2. **Check the decision.** [Compare the acceptance gates](https://noteflowai.github.io/evalarc/#suite): + the same 93.75% score passes a permissive rule and fails the strict notes rule. +3. **Recompute it locally.** The [first-review walkthrough](docs/first-review.md) + installs the published wheel, downloads the records and rebuilds the comparison. + Verification exits 0 for consistency; comparison exits 1 for the regression. -Clone the source, then install in an isolated Python environment: +Install the released reviewer in a fresh virtual environment: ```bash -git clone https://github.com/noteflowai/evalarc.git -cd evalarc python3 -m venv .venv -source .venv/bin/activate -python -m pip install -e . +. .venv/bin/activate +python -m pip install "https://github.com/noteflowai/evalarc/releases/download/v0.12.1/evalarc-0.12.1-py3-none-any.whl#sha256=115f3d8d452dee2b5d3aed12f736880aca69aaf3ea42937ef4f8f222c0b2291b" +evalarc --version +``` + +The offline review needs no Docker, Node, GPU or model API. Follow the +[download and comparison commands](docs/first-review.md#3-recompute-the-recorded-regression) +to produce your first HTML report without cloning the source. + +## Bring your own work + +| What you need to review | Use EvalArc to | Start here | +| --- | --- | --- | +| A changed agent implementation | Compare matching evaluations and inspect regressed checks | [Run and compare](docs/workflow.md) | +| Saved AgentCore Evaluate results and spans | Inspect valid zero scores, skipped judgments, missing results and skill delivery | [Export-to-review walkthrough](docs/agentcore-first-review.md) | +| Repeated judgments on one fixed recording | Separate score variation, verdict disagreement and incomplete assessments | [Judge Stability](docs/judge-stability.md) | +| A report received from another developer | Recompute summaries, configured gates and JUnit from original inputs | [Offline verification](docs/verification.md) | +| A grader or candidate you want to execute | Run a reference and deliberate faults against a task contract | [Run an audit](#run-an-audit) | + +Trace import accepts a [bounded export format](docs/trace-workbench.md), not arbitrary +cloud exports. Its scored controls are synthetic; the separate MCP example records +actual local delivery with no evaluator scores. No live AgentCore evaluation is claimed. + +Trying your own records? [Tell us where the first review helped or got stuck](https://github.com/noteflowai/evalarc/issues/new?template=first-use.yml). +A minimal redacted example is enough; a failed setup is useful feedback too. + +## What the recorded evidence covers + +| Task | Interaction | Declared faults | Detected by only one case | +| --- | --- | ---: | ---: | +| `durable-kv` | Coding artifact: responses, transactions and restart durability | 8 | 3 | +| `support-routing` | Simulated ticket tools: routing, exact notes, closure and unrelated state | 7 | 2 | +| `robot-evidence-review` | Attributed recordings: coordinates, clocks and missing observations | 6 | 1 | + +The saved audits detect **21/21 declared faults across three task packs**. Six +faults depend on one detecting case each. Removing a sole detector lowers a fresh +audit's mutation score; these margins expose that dependency before the change. +They do not establish coverage of unseen faults. +[Inspect coverage](https://noteflowai.github.io/evalarc/#coverage) · +[Methodology](docs/methodology.md) · [251 audit case records](docs/casebook.md). + +For deeper exploration: [repeated attempts](docs/reliability.md), +[TOML suites and CI](docs/suites.md), [Python/JavaScript candidates](docs/languages.md), +[recorded GPU research pilots](docs/research-pilots.md), +[architecture](docs/architecture.md) and [papers](docs/research.zh-CN.md). +Feature history lives in the [changelog](CHANGELOG.md). + +## Run an audit + +To execute the built-in Python reference and eight deliberate coding faults, +install the wheel above, then use Docker: + +```bash docker pull python:3.12-slim evalarc audit --seeds 17 41 97 --output runs/audit ``` -The command evaluates the reference and eight negative controls, writes -`runs/audit/audit.json`, and creates a standalone `runs/audit/index.html` report. -Exit code `0` means the reference passed and every declared defect was detected -in its intended dimension. Exit code `1` means an audit or candidate failed; -`2` indicates a usage/configuration error or an invalid run caused by an -environment failure. +Open `runs/audit/index.html`. Exit 0 means the reference passed and all declared +faults were detected in their intended dimensions; 1 means an audit/candidate +failed, and 2 means invalid input or an environment failure. -For the bundled, trusted controls, a faster CPU-only demo is: +For the bundled trusted controls, this shorter CPU-only run uses the host: ```bash -evalarc audit --backend local --trust-local --output runs/local-audit evalarc audit --task support-routing --backend local --trust-local --output runs/support-audit ``` -Local execution has the host user's privileges. Use Docker for candidate -isolation and read the [execution boundaries](SECURITY.md). -If your Docker setup requires a wrapper, set `EVALARC_DOCKER` to that command -or pass `--docker-command`. +Local execution has your user privileges. Use Docker for candidate isolation; +see [execution boundaries](SECURITY.md) and [readiness checks](docs/workflow.md). +Use a fresh output path for another run. To develop EvalArc itself, see +[Development](#development). ## Coding task @@ -286,7 +228,13 @@ failures must be established before this becomes a research benchmark. ## Development +Clone the source for development and for the `examples/` commands in this README: + ```bash +git clone https://github.com/noteflowai/evalarc.git +cd evalarc +python3 -m venv .venv +. .venv/bin/activate python -m pip install -e ".[dev]" pytest -q ruff check . @@ -297,4 +245,4 @@ python -m build See [CONTRIBUTING.md](CONTRIBUTING.md), [SECURITY.md](SECURITY.md), and [LICENSE](LICENSE). The [task-author guide](docs/task-authoring.md) explains the current built-in extension points. CI includes Python checks, the Node -policy, and Docker audits for both packs. +policies, and Docker audits for all three task packs. diff --git a/README.zh-CN.md b/README.zh-CN.md index 86dad3a..8fa1caf 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -1,146 +1,101 @@ -

EvalArc — Run agents. Measure outcomes.

+

EvalArc — 分数提高,检查却退步了。

# EvalArc -**面向 AI 智能体的开放任务环境与可审计评测。** +**找出更高评分背后的 Agent 回归。** -Python 3.11+,Linux 主机,零运行时第三方依赖,MIT 许可证。 +查看退步的检查,定位已记录的工具动作,交付他人可以复核的证据。 -[在线交互演示](https://huggingface.co/spaces/glayguo/evalarc) · -[网页镜像](https://noteflowai.github.io/evalarc/) · -[可筛选证据数据集](https://huggingface.co/datasets/glayguo/evalarc-casebook) · -[版本下载](https://github.com/noteflowai/evalarc/releases) · -[English](README.md) · [中文调研与论文分析](docs/research.zh-CN.md) · -[架构设计](docs/architecture.md) · [方法说明](docs/methodology.md) · [开发路线](docs/roadmap.md) +[**立即查看失败案例 →**](https://noteflowai.github.io/evalarc/#regression) · +[首次本地复核](docs/first-review.zh-CN.md) · +[Hugging Face 演示](https://huggingface.co/spaces/glayguo/evalarc) · [English](README.md) -EvalArc 关注智能体实际完成的结果,以及支撑评分结论的证据。 -通过正确实现和刻意带有缺陷的实现进行对照, -检查它能发现哪些问题,并保存可复查的报告。 +**90% → 93.75%。两项检查改善,一项原本通过的检查却失败了。** +工具已写入备注,却返回错误;策略更换幂等键后重试,多写了一次。 +EvalArc 把这次退步展开给你看,帮助判断分数提高是否满足发布要求。 -**分数从 90% 升到 93.75%,原本通过的检查却失败了。** -[交互证据实验室](https://huggingface.co/spaces/glayguo/evalarc) -可以并排比较两个版本,查看一处退步、两处改进,再逐步检查工具调用和状态变化。 -页面读取仓库保存的 Docker 审计记录,无需安装,也不调用模型。 -使用 **Copy evidence link** 分享具体案例和步骤;打开详情后可返回原案例, -某一分区加载失败时可单独重试。新链接携带实际加载审计文件的 SHA-256, -记录变化时会先提示并停止恢复;旧链接明确说明未保存原始审计身份。 -指纹标识内容,不认证作者。[交互使用说明](docs/explorer.md)。 + + + 已记录案例的操作演示:分数提高,重试导致重复备注,严格验收规则拒绝结果。 + -[![EvalArc v0.3:分数上升,一项检查却退步](docs/assets/regression-lab.png)](https://huggingface.co/spaces/glayguo/evalarc) +演示回放的是脚本对照在 Docker 中运行后保存的记录,无需安装、账号或模型密钥。 +研究预览 · MIT · Python 3.11+ · 本地流程使用 Linux · Python 包无第三方运行时依赖。 -**v0.8:完整验收证据可离线复核。** 从首页下载套件 ZIP,解压后运行 -`evalarc verify suite-evidence --json`,复核原始 TOML、执行计划、五次尝试、 -验收规则与 JUnit。使用 `--require-accepted` 接入 CI 验收;证据一致、规则接受、 -任务完全完成分别报告。整个检查不执行候选程序。 -[离线验证流程与边界](docs/verification.md)。 +## 从一次复核开始 -**已实现 coding、业务工具与记录复核三个场景。** +1. **看退步。** [对照两个版本](https://noteflowai.github.io/evalarc/#regression), + 再在[案例浏览器](https://noteflowai.github.io/evalarc/#explorer)查看 `retry-after-commit`。 +2. **看验收。** [比较两条规则](https://noteflowai.github.io/evalarc/#suite): + 相同的 93.75% 分数,宽松规则接受,严格备注规则拒绝。 +3. **本地复算。** [首次复核指南](docs/first-review.zh-CN.md)使用发布的 wheel 和原始记录重建报告; + 复核退出 0 表示一致,对照退出 1 表示发现退步。 -| 任务 | 交互方式 | 验证内容 | 声明缺陷 | 仅单个用例检出 | -| --- | --- | --- | ---: | ---: | -| `durable-kv` | 执行代码智能体交付的服务 | 读写、事务、CAS、持久化与异常恢复 | 8 | 3 | -| `support-routing` | 策略通过工具操作模拟工单 | 路由、精确备注、条件关闭、无关数据保护与协议完成 | 7 | 2 | -| `robot-evidence-review` | 基于有出处的记录数据出报告 | 坐标与时钟换算、缺失观测、来源归属 | 6 | 1 | +在新的虚拟环境安装已发布的复核工具: -三个任务包都检出了全部声明缺陷:21 个缺陷,21 个检出。其中六个各自只靠一个用例检出;删除该用例,或使其无法再检出对应缺陷,就会失去这部分覆盖,重新审计的变异分数也会下降。检出余量在改动之前就能指出这些依赖,与当前分数一起报告。[与 hack-verifiable environments 的关系](docs/methodology.md#relation-to-hack-verifiable-environments)。 - -v0.6 为两个任务都提供 **Python 和 JavaScript 工作区模板**: -`init --language javascript` 生成起步代码,添加 `--reference` 生成脚本对照; -`audit --language javascript` 使用独立的 Node.js 实现检查相同的 21 类故障,实测检出余量与 Python 完全一致。 -JavaScript 需要 Node.js 22+,Docker 模式显式指定 `--image node:22-slim`。 -详见[多语言接入指南](docs/languages.zh-CN.md)。 - -v0.5 新增 `evalarc suite`:用 TOML 声明任务、候选、轮次、预算及验收门槛, -先预览执行计划,再批量运行并输出 HTML、JSON 和 JUnit。各任务单独评分。 -详见[套件与 CI 指南](docs/suites.zh-CN.md)。 - -[新增验收规则展示](https://glayguo-evalarc.static.hf.space/#suite)让同一份缺陷策略分别按两套规则验收: -都是 93.75%、0/2 轮完全通过,宽松规则允许部分进展,要求备注检查全部通过的规则则拒绝。 -完整三项评测作业的 Docker 记录、五轮尝试、TOML 和 JUnit 均可检查;“规则接受”与“任务完全完成” -分别展示。 - -[![EvalArc v0.5:相同分数,不同验收结果](docs/assets/suite-lab.png)](https://glayguo-evalarc.static.hf.space/#suite) - -[Hugging Face Casebook](https://huggingface.co/datasets/glayguo/evalarc-casebook) -把 251 条审计用例、6 次重复尝试和 3 项验收作业分别整理为可筛选的表, -保留未经改写的原始 JSON 和版本指纹。可先选择 `suite_jobs`, -对照 `gate_accepted` 与 `fully_resolved`,或用 Python 读取。 -这些是公开开发任务中的脚本对照,不是隐藏模型测试集。详见[数据说明](docs/casebook.md)。 - -`evalarc repeat` 固定一份候选快照,在相同场景上重新启动多轮评测, -保存每轮证据并显示逐项通过率与结果波动。同时补齐场景总时间预算、JSONL 进度 -和受限进程诊断。详见[重复评测指南](docs/reliability.zh-CN.md)。 - -[重复评测展示](https://glayguo-evalarc.static.hf.space/#repeat)保留了两种脚本策略各三次 -Docker 评测:参考策略 3/3 轮完全通过,重复写入策略虽然平均分为 93.75%,却 0/3 轮 -完全通过。可以查看每轮原始证据和逐项计数。这些观察中未出现检查结果波动, -也不能据此估计模型在未见任务上的可靠性。 - -[![EvalArc v0.4:三轮平均分 93.75%,但没有一轮完全通过](docs/assets/repeat-lab.png)](https://glayguo-evalarc.static.hf.space/#repeat) - -现有流程包含环境预检查、单次评测 HTML 报告和逐项回归比较,即使总分上升也能指出 -退步的检查;输出保护会保留之前的运行证据。详见[使用流程](docs/workflow.zh-CN.md)。 - -两个场景使用共同的报告元数据,各自定义评分规则。工单环境记录工具调用和 -真实状态变化,支持写入前失败、写入后响应失败及幂等重试。 -已提供 Python 和 JavaScript 策略;浏览器环境、真实模型 API 适配与 RL 集成 -仍在路线图中。具体边界见[架构设计](docs/architecture.md)。 - -当前版本尚未经过前沿模型、真实人类工时或强化学习收益标定。 - -**收到报告后,先独立复核。** `evalarc verify 报告路径 --json` 无需执行候选程序, -即可重算单次评测、重复运行和前后对照的汇总,并记录每份输入的指纹。 -需要全部任务通过时加 `--require-resolved`。 -[使用说明与校验范围](docs/verification.md)。 - - -[![三个任务包的覆盖薄弱点与原始证据入口](docs/coverage-review.png)](https://noteflowai.github.io/evalarc/#coverage) +```bash +python3 -m venv .venv +. .venv/bin/activate +python -m pip install "https://github.com/noteflowai/evalarc/releases/download/v0.12.1/evalarc-0.12.1-py3-none-any.whl#sha256=115f3d8d452dee2b5d3aed12f736880aca69aaf3ea42937ef4f8f222c0b2291b" +evalarc --version +``` -**在信任满分前,先检查覆盖薄弱点。** 网页现展示全部三个任务包,可从仅靠一个用例检出的缺陷直接定位到原始检查与种子记录。离线审计报告也提供相同的证据展开入口,不依赖脚本或远程资源;新增界面沿用原始数据,不冒充新模型运行。 +离线复核不需要 Docker、Node、GPU 或模型 API。继续按 +[下载与对照步骤](docs/first-review.zh-CN.md#3-复算这次退步)生成 HTML 报告,无需克隆源码。 -## 0.11.0:运行记录评估工作台 +## 用于自己的工作 -导入已保存的 AgentCore Evaluate 结果、版本化黄金案例和 Skills Anywhere 加载回执,分别查看有效零分、评估跳过、缺少结果及漏调用技能。支持相同测试集与评分规则下的对比,以及原始输入的离线复核。[交互示例](https://noteflowai.github.io/evalarc/trace-workbench/index.html) · [真实本地 MCP 加载](https://noteflowai.github.io/evalarc/trace-mcp/index.html) · [数据契约](docs/trace-workbench.md)。示例明确区分合成评分与真实加载记录,未运行云端评估。 +| 已有材料 | 可以检查什么 | 入口 | +| --- | --- | --- | +| 变更前后的 Agent 评测 | 匹配条件下哪些检查退步 | [运行与对照](docs/workflow.md) | +| AgentCore Evaluate 结果与 spans | 有效零分、跳过、缺失以及技能交付 | [导出到审阅](docs/agentcore-first-review.md) | +| 同一记录上的多次评判 | 分数变化、通过/拒绝翻转及未评判情况 | [中文指南](docs/judge-stability.zh-CN.md) | +| 他人交付的报告 | 原始输入与汇总、验收规则、JUnit 是否一致 | [离线复核](docs/verification.md) | +| 评分器或待执行候选 | 正确实现和刻意缺陷是否被评分器区分 | [运行审计](#运行审计) | -## v0.12.0:同一记录,多次评判 +运行记录导入有明确的[受限格式约定](docs/trace-workbench.md),不直接接受任意云端导出。 +带分数的示例为合成数据,独立 MCP 示例是真实本地加载但没有评委分数;未完成实时 AgentCore 评测。 -新增 `trace-stability`,固定执行记录与评估规则,分别检查分数变化、通过/拒绝翻转、 -跳过与缺失,并保存原始输入供离线复核。三次都拒绝仍然是拒绝,不把一致性当作 -任务成功。[在线交互示例](https://noteflowai.github.io/evalarc/judge-stability/index.html) -· [中文使用指南](docs/judge-stability.zh-CN.md)。五个展示案例均为人工构造; -此功能不重新运行 Agent 或裁判,也不代替独立人工校准。 +尝试自己的记录后,欢迎[反馈首次使用的卡点或发现](https://github.com/noteflowai/evalarc/issues/new?template=first-use.yml)。 +最小脱敏样例即可;安装失败同样有价值。 -[![五个人工构造的案例,分别展示分数变化、结论翻转和未评判结果](docs/assets/judge-stability.png)](https://noteflowai.github.io/evalarc/judge-stability/index.html) +## 原始证据覆盖什么 -## 0.9.0:有原始证据的研究场景 +| 任务 | 交互场景 | 声明缺陷 | 仅一个用例检出 | +| --- | --- | ---: | ---: | +| `durable-kv` | 代码交付物:响应、事务、重启持久化 | 8 | 3 | +| `support-routing` | 模拟工单:路由、精确备注、关闭与无关状态保护 | 7 | 2 | +| `robot-evidence-review` | 有出处的记录:坐标、时钟与缺失观察 | 6 | 1 | -[查看 27 次真实 GPU 技能评测](https://noteflowai.github.io/evalarc/skill-impact/index.html),并阅读[完整方法与限制](docs/research-pilots.md)。新增[实景 Blender 编辑](https://noteflowai.github.io/robot-reel/scene-lab/)与[官方 LIBERO-Plus 子集回放](https://noteflowai.github.io/robot-reel/libero-plus/),把原始记录、技能交付与独立验收连接起来。失败尝试全部保留;不宣称技能提分、完整基准成绩或真机效果。 +三个任务包保存的审计检出 **21/21 种声明缺陷**。其中六种各依赖一个检测用例; +移除唯一检测用例后,重新审计的变异分数会下降。覆盖余量用于提前暴露这种依赖, +不代表覆盖未知缺陷。[检查覆盖](https://noteflowai.github.io/evalarc/#coverage) · +[方法说明](docs/methodology.md) · [251 条审计记录](docs/casebook.md)。 +更多工作流:[重复运行](docs/reliability.md)、[TOML 套件与 CI](docs/suites.md)、 +[Python/JavaScript](docs/languages.md)、[GPU 研究记录](docs/research-pilots.md)、 +[架构](docs/architecture.md)和[论文分析](docs/research.zh-CN.md)。历史更新见 [CHANGELOG](CHANGELOG.md)。 -## 直接运行 +## 运行审计 -克隆仓库后,在独立 Python 环境中安装: +安装上面的 wheel 后,用 Docker 执行内置 Python 参考实现和八种刻意带错的代码实现: ```bash -git clone https://github.com/noteflowai/evalarc.git -cd evalarc -python3 -m venv .venv -source .venv/bin/activate -python -m pip install -e . docker pull python:3.12-slim evalarc audit --seeds 17 41 97 --output runs/audit ``` -输出 `runs/audit/audit.json` 和可独立打开的 `runs/audit/index.html`。 -对仓库自带的可信对照代码,可以运行更快的本机演示: +打开 `runs/audit/index.html`。退出 0 表示参考实现通过且所有声明缺陷被检出; +1 表示审计或候选失败;2 表示参数或环境导致结果无效。 + +对包内可信对照,可以在 CPU 上直接运行: ```bash -evalarc audit --backend local --trust-local --output runs/local-audit evalarc audit --task support-routing --backend local --trust-local --output runs/support-audit ``` -本机模式具有当前用户的文件和网络权限。Docker 模式的边界见 -[SECURITY.md](SECURITY.md)。无模型调用,无 API 费用,无 GPU 依赖。 +本机模式具有当前用户权限;候选隔离使用 Docker,边界见 [SECURITY.md](SECURITY.md)。 +重复运行时使用新的输出路径。 ## 已实现 diff --git a/docs/agentcore-first-review.md b/docs/agentcore-first-review.md new file mode 100644 index 0000000..371224d --- /dev/null +++ b/docs/agentcore-first-review.md @@ -0,0 +1,98 @@ +# From an AgentCore export to a local review + +This walkthrough connects EvalArc's existing `trace-import` command to your +review process. It starts with published controls, then explains which pieces +to replace with actual records. It does not deploy a cloud agent or collect +AWS data for you. + +**Validation scope:** the first half is reproducible with the released 0.12.1 +wheel. The five scored controls are synthetic API-shaped inputs. A separate +record contains actual local MCP delivery and no evaluator scores. A live +AgentCore or Strands evaluation has not been validated by these examples. + +## Inspect a complete input before collecting data + +After [installing the wheel](first-review.md#2-install-the-reviewer), run +these commands in a fresh directory: + +```bash +curl --fail --location \ + https://github.com/noteflowai/evalarc/releases/download/v0.12.1/trace-workbench-evidence.zip \ + --output trace-workbench-evidence.zip +python -m zipfile -e trace-workbench-evidence.zip . +evalarc trace-import trace-workbench/input.json \ + --baseline trace-workbench/baseline-input.json \ + --output my-trace-review +evalarc trace-verify my-trace-review +``` + +Both commands exit **0** for a valid import and consistent evidence. That does +not mean every case was accepted. Open `my-trace-review/index.html` to see the +accepted, zero-scored, skipped, missing-result and missing-skill controls. +The ZIP SHA-256 is +`2e8ea5ba062167e8de5420400e60837a1aae68e465b5cfdf5337686d6fe2be3f`. + +The copied `my-trace-review/input.json` and `baseline-input.json` preserve the +original bytes; `review.json` contains the derived review. The report and +filters work offline. + +## Replace authored inputs with a recorded export + +Use `trace-workbench/input.json` as a structural example. Create your own +wrapper rather than relabeling the demonstration as a real run: + +1. Record the golden set, ordered case IDs, intended goals and required skills. + Set `provenance.kind` to `recorded` and describe how the data was collected + and what the collection omits. +2. Copy each case's native **AgentCore Evaluate API** `evaluationResults` into + `evaluation_response`. Keep the service's original fields and matching + session/trace/span identifiers. +3. Supply a flat array of corresponding OpenTelemetry spans. Extract spans + from any transport envelope first; arbitrary OTLP envelopes and CloudWatch + query wrappers are not accepted directly. +4. Explicitly declare the evaluator's target level, revision, rating scale + and acceptance rule. Keep model/configuration identity and hashes where + required. An evaluator name does not establish its scale or version. +5. If using Skills Anywhere, preserve actual `open_skill` receipts and + annotate the matching tool spans. Only assert complete skill observation + when your collector actually observed all loads. + +The AWS workshop CLI's `run.results[].sessionScores` is a different envelope. +Renaming that field is not a conversion. Strands traces alone also do not +provide the required golden cases and judgments. See the +[full input contract](trace-workbench.md#prepare-a-real-export) before adapting +an exporter. + +## Review, compare, then gate + +```bash +evalarc trace-import recorded-current.json --output recorded-review +evalarc trace-verify recorded-review +``` + +For a baseline comparison, add `--baseline recorded-baseline.json`. Both +inputs must use the same complete golden set and evaluator definitions; +synthetic inputs cannot be compared with recorded inputs. + +Add `--require-accepted` to **trace-import** when a CI step should enforce the +rules. Its exit codes distinguish **0** (every case accepted), **1** (complete +assessments with rejected gates), and **2** (incomplete assessments or invalid +input). A well-formed input still produces its report when a gate rejects or +a judgment is missing. `trace-verify` checks consistency; it is not the gate. + +If importing fails, start with the named case/evaluator in the diagnostic. +Check linkage, duplicate targets, explicit scale bounds and expected-result +coverage. A missing or skipped evaluation must not be filled with a fabricated +zero to make the file pass validation. + +## Bring back one useful finding + +Use the [first-use form](https://github.com/noteflowai/evalarc/issues/new?template=first-use.yml) +to report the exporter/envelope you used, the command and the first point of +friction. A small redacted fixture helps establish whether an adapter is +useful. Do not attach production traces merely to demonstrate interest. + +Sources and collection boundaries are documented in the +[Trace Workbench guide](trace-workbench.md#scope-and-sources). EvalArc reviews +imported judgments; it does not independently establish their truth or the +completeness of the supplied execution record. diff --git a/docs/assets/banner.svg b/docs/assets/banner.svg index a52f3ae..e3a2969 100644 --- a/docs/assets/banner.svg +++ b/docs/assets/banner.svg @@ -1,14 +1,14 @@ - + EvalArc - Run agents. Measure outcomes. Open environments and evaluations for AI agents. Research preview. - + Higher score. New failure. Review agent regressions with reproducible evidence. Research preview. + - EXECUTABLE EVALUATION / RESEARCH PREVIEW - EvalArc - Run agents. Measure outcomes. - STATEFUL TASKS / NEGATIVE CONTROLS / REPRODUCIBLE EVIDENCE + AGENT REGRESSION REVIEW / RESEARCH PREVIEW + EvalArc + Higher score. New failure. + COMPARE CHECKS / FOLLOW ACTIONS / VERIFY EVIDENCE diff --git a/docs/assets/first-review-media.json b/docs/assets/first-review-media.json new file mode 100644 index 0000000..0e0eaeb --- /dev/null +++ b/docs/assets/first-review-media.json @@ -0,0 +1,47 @@ +{ + "kind": "annotated screenshots of actual UI states", + "duration_seconds": 30, + "evidence": { + "comparison/baseline.json": "4ec043dd1de0b0d852737cf6e8f2ebe04d517d0c945335e686773e00a8f970b5", + "comparison/current.json": "40d0726cd8378b237c514d69b4fda9e4f00550dc3012f4c276a360f4deed25ad", + "support/audit.json": "596393f262a4b0d411ffd776bb83b80ef37b36663ee9bd29b447cf472ae20c95", + "suite/suite.json": "532788bc3655814616f1ae89c527d8eb3c39c61e299cb47fd1910b5fd4f80a50" + }, + "frames": [ + { + "title": "A higher score. A new regression.", + "detail": "90% → 93.75% · Two closure checks improve; the notes check regresses." + }, + { + "title": "The tool returned an error. The note was written.", + "detail": "Recorded step 3 · The state changed before the transient error returned." + }, + { + "title": "A new key. A duplicate note.", + "detail": "Recorded step 4 · Retrying the write with a different key adds the note again." + }, + { + "title": "Keep the score. Check the acceptance rule.", + "detail": "Same 93.75% policy · The strict notes gate rejects what the permissive gate accepts." + } + ], + "files": { + "first-review.png": { + "bytes": 156318, + "sha256": "2d5e6f83fb4d9ed5b430e9ec965c41e09c385838e21718cc200d44aea62c4dc9" + }, + "first-review.gif": { + "bytes": 242960, + "sha256": "57d7abb3f1233475113319b5e7bf92928d6fd004a45357150ec8ac597cc5f5af" + }, + "first-review.mp4": { + "bytes": 195851, + "sha256": "257a3fcf0ffc1515d79efebff52461bab279eea7729fcba2da367f10dae61909" + }, + "first-review.vtt": { + "bytes": 600, + "sha256": "7c6508c00c0802484f907e7bdfd3d018a39442938dc98ef582c20eb1b5f63b71" + } + }, + "scope": "Saved scripted Docker controls. No new agent execution or customer evidence." +} diff --git a/docs/assets/first-review.gif b/docs/assets/first-review.gif new file mode 100644 index 0000000..57a28b4 Binary files /dev/null and b/docs/assets/first-review.gif differ diff --git a/docs/assets/first-review.mp4 b/docs/assets/first-review.mp4 new file mode 100644 index 0000000..d1d1e05 Binary files /dev/null and b/docs/assets/first-review.mp4 differ diff --git a/docs/assets/first-review.png b/docs/assets/first-review.png new file mode 100644 index 0000000..36121e0 Binary files /dev/null and b/docs/assets/first-review.png differ diff --git a/docs/assets/first-review.vtt b/docs/assets/first-review.vtt new file mode 100644 index 0000000..b426a74 --- /dev/null +++ b/docs/assets/first-review.vtt @@ -0,0 +1,17 @@ +WEBVTT + +00:00:00.000 --> 00:00:07.500 +A higher score. A new regression. +90% → 93.75% · Two closure checks improve; the notes check regresses. + +00:00:07.500 --> 00:00:15.000 +The tool returned an error. The note was written. +Recorded step 3 · The state changed before the transient error returned. + +00:00:15.000 --> 00:00:22.500 +A new key. A duplicate note. +Recorded step 4 · Retrying the write with a different key adds the note again. + +00:00:22.500 --> 00:00:30.000 +Keep the score. Check the acceptance rule. +Same 93.75% policy · The strict notes gate rejects what the permissive gate accepts. diff --git a/docs/first-review-media.md b/docs/first-review-media.md new file mode 100644 index 0000000..9b59a27 --- /dev/null +++ b/docs/first-review-media.md @@ -0,0 +1,43 @@ +# First-review walkthrough + +The 30-second animation and video show four **actual interface states** with +explanatory captions. They are an annotated walkthrough of saved scripted +Docker controls, not a real-time recording of an agent run. + +| Time | View | Finding | +| --- | --- | --- | +| 0–7.5 s | Revision comparison | The score rises from 90% to 93.75%. Two closure checks improve; the notes check regresses. | +| 7.5–15 s | `retry-after-commit`, step 3 | The note is committed before the tool returns `temporarily_unavailable`. | +| 15–22.5 s | The same case, step 4 | A retry with a different idempotency key adds the note again. | +| 22.5–30 s | Suite acceptance rules | The same 93.75% policy passes a permissive gate and fails the strict notes gate. | + +The MP4 has controls, optional captions and no audio or autoplay. The README +uses a GIF and selects a static image for reduced-motion preferences. + +## Reproduce the media + +Use the repository's Python development environment, installed npm +dependencies and Playwright Chromium. The capture also requires an FFmpeg +executable with libx264 and GIF encoding support. + +```bash +python scripts/build_site.py --output dist/media-source +python -m http.server 8767 --bind 127.0.0.1 --directory dist/media-source +``` + +In a second terminal, from the same repository: + +```bash +SITE_URL=http://127.0.0.1:8767/ FFMPEG=ffmpeg node scripts/capture_first_review.cjs +``` + +The script asserts the displayed scores, regression count, committed error, +duplicate note and gate decisions before taking screenshots. It adds captions +in a separate composition and writes `docs/assets/first-review.*`. +`first-review-media.json` records media hashes and the original evidence hashes. +Font and encoder differences may change media bytes. The original evidence +is never rewritten. + +Rebuild the site into a fresh directory to include regenerated media. Read the +[first-review guide](first-review.md) for the actual offline verification +steps and their limits. diff --git a/docs/first-review.md b/docs/first-review.md new file mode 100644 index 0000000..908a797 --- /dev/null +++ b/docs/first-review.md @@ -0,0 +1,94 @@ +# Your first EvalArc review + +Start with a recorded failure, check it on your machine, then use the same +workflow with your own evidence. The steps below use the published **0.12.1** +wheel and records. You need Python 3.11+ on Linux and `curl` for downloads. +The review does not need a source checkout, Docker, Node, a GPU or a model key. + +## 1. See the problem + +[Open the revision comparison](https://noteflowai.github.io/evalarc/#regression). +The score rises from **90% to 93.75%**, but the notes check regresses. Two +closure checks improve. Select `retry-after-commit` to see the duplicate note. + +These are saved Docker runs of scripted controls. They illustrate a grader +and a workflow, not a customer's incident or a model ranking. + +## 2. Install the reviewer + +Use a new folder so that existing downloads and outputs stay separate: + +```bash +mkdir evalarc-first-review +cd evalarc-first-review +python3 -m venv .venv +. .venv/bin/activate +python -m pip install "https://github.com/noteflowai/evalarc/releases/download/v0.12.1/evalarc-0.12.1-py3-none-any.whl#sha256=115f3d8d452dee2b5d3aed12f736880aca69aaf3ea42937ef4f8f222c0b2291b" +evalarc --version +``` + +Expected: `EvalArc 0.12.1`. This installs the released wheel, which has no +third-party runtime dependencies. The URL includes its SHA-256. It does not +require a package named `evalarc` to exist on PyPI. + +## 3. Recompute the recorded regression + +```bash +curl --fail --location \ + https://github.com/noteflowai/evalarc/releases/download/v0.12.1/evalarc-evidence-explorer.zip \ + --output evalarc-evidence-explorer.zip +python -m zipfile -e evalarc-evidence-explorer.zip . +evalarc verify evalarc-evidence-explorer/comparison --json +evalarc compare evalarc-evidence-explorer/comparison/baseline.json \ + evalarc-evidence-explorer/comparison/current.json \ + --output comparison-review +``` + +The verification exits **0**: the records are consistent. The comparison exits +**1**: one check regressed, even though the score improved. Open +`comparison-review/index.html` directly in your browser to inspect the result. +It works offline. Use a new output folder when repeating the comparison. + +The released ZIP has SHA-256 +`4ac78211c718dc930a0091227e0d4d9baf0543e05222792d82e05cb44362a073`; +all release checksums are in `SHA256SUMS` on the +[release page](https://github.com/noteflowai/evalarc/releases/tag/v0.12.1). + +## 4. Check the acceptance decision + +```bash +evalarc verify evalarc-evidence-explorer/suite --json +evalarc verify evalarc-evidence-explorer/suite --json --require-accepted +``` + +The first command exits **0** for consistency; the second exits **1** because +the strict notes gate rejects the policy. Across the recorded suite, **2/3 +jobs are accepted and 1/3 is fully resolved**. A configured gate accepting +partial progress does not mean every task check passed. + +Open `evalarc-evidence-explorer/suite/index.html` to inspect the rules. Keep +`suite/junit.xml` with the original suite if you hand the report to someone +else. [Verification checks and limits](verification.md) explain what is +recomputed; this is not a fresh execution of the candidate or authentication +of the producer. + +## 5. Use your own evidence + +| What you have | Next step | +| --- | --- | +| Two EvalArc evaluations of a candidate | Replace the two JSON inputs to `evalarc compare`; keep matching tasks, graders, cases and runtime conditions. [Workflow](workflow.md) | +| Saved AgentCore Evaluate results and spans | Prepare the bounded input wrapper, import locally, and inspect missing judgments as well as rejected gates. [Export-to-review walkthrough](agentcore-first-review.md) | +| A coding or tool-using candidate to execute | Run a task audit or evaluate its workspace with the Docker backend. [Run an audit](../README.md#run-an-audit) | +| Repeated judgments on one fixed trace | Check score variation and verdict disagreement separately. [Judge Stability](judge-stability.md) | + +## Share a useful first-use report + +If you try your own records, a small +[first-use report](https://github.com/noteflowai/evalarc/issues/new?template=first-use.yml) +helps improve the workflow: what you were reviewing, where you got stuck, and +whether the result changed a decision. A failed setup is useful feedback too. +Only include a minimal redacted example; private prompts, customer data and +credentials are not needed. + +The first-use form is an invitation for independent feedback. It is not a +claim that independent users have completed this workflow. diff --git a/docs/first-review.zh-CN.md b/docs/first-review.zh-CN.md new file mode 100644 index 0000000..bea8bd2 --- /dev/null +++ b/docs/first-review.zh-CN.md @@ -0,0 +1,81 @@ +# 第一次用 EvalArc 复核 + +先看一个已记录的失败,在本机复算,再换成自己的证据。下列流程使用已发布的 +**0.12.1 wheel 和记录**,需要 Linux、Python 3.11+ 和用于下载的 `curl`。 +审阅不需要克隆源码、Docker、Node、GPU 或模型密钥。 + +## 1. 先看问题 + +[打开版本对照](https://noteflowai.github.io/evalarc/#regression):分数从 +**90% 升到 93.75%**,两项关闭检查改善,但备注检查退步。 +选择 `retry-after-commit`,查看多写的那条备注。 + +这是脚本对照在 Docker 中运行后保存的记录,用于展示评分器与复核流程, +不是客户事故或模型排名。 + +## 2. 安装复核工具 + +在新文件夹中运行,避免混用已有下载和输出: + +```bash +mkdir evalarc-first-review +cd evalarc-first-review +python3 -m venv .venv +. .venv/bin/activate +python -m pip install "https://github.com/noteflowai/evalarc/releases/download/v0.12.1/evalarc-0.12.1-py3-none-any.whl#sha256=115f3d8d452dee2b5d3aed12f736880aca69aaf3ea42937ef4f8f222c0b2291b" +evalarc --version +``` + +预期输出 `EvalArc 0.12.1`。安装地址固定了发布版本及 SHA-256, +包本身没有第三方运行时依赖;不依赖 PyPI 存在同名包。 + +## 3. 复算这次退步 + +```bash +curl --fail --location \ + https://github.com/noteflowai/evalarc/releases/download/v0.12.1/evalarc-evidence-explorer.zip \ + --output evalarc-evidence-explorer.zip +python -m zipfile -e evalarc-evidence-explorer.zip . +evalarc verify evalarc-evidence-explorer/comparison --json +evalarc compare evalarc-evidence-explorer/comparison/baseline.json \ + evalarc-evidence-explorer/comparison/current.json \ + --output comparison-review +``` + +复核命令退出 **0**,表示记录内部一致;对照命令退出 **1**,表示一项检查退步, +即使总分提高。用浏览器直接打开 `comparison-review/index.html`,可以离线查看。 +重复运行时换一个输出文件夹。 + +此发布 ZIP 的 SHA-256 为 +`4ac78211c718dc930a0091227e0d4d9baf0543e05222792d82e05cb44362a073`。 +完整校验清单见[发布页](https://github.com/noteflowai/evalarc/releases/tag/v0.12.1) +的 `SHA256SUMS`。 + +## 4. 检查验收结论 + +```bash +evalarc verify evalarc-evidence-explorer/suite --json +evalarc verify evalarc-evidence-explorer/suite --json --require-accepted +``` + +第一条命令退出 **0**,表示证据一致;第二条退出 **1**,因为严格备注规则拒绝了 +该策略。整个套件中 **2/3 作业被规则接受,1/3 完全完成**。宽松规则允许部分进展, +不代表所有任务检查通过。 + +打开 `evalarc-evidence-explorer/suite/index.html` 查看规则。交付时保留原始套件 +及 `suite/junit.xml`。此流程复算记录,不重新执行候选程序,也不认证生产者身份; +详见[复核范围](verification.md)。 + +## 5. 换成自己的证据 + +| 已有材料 | 下一步 | +| --- | --- | +| 两份 EvalArc 评测 | 替换 `compare` 的输入 JSON;任务、评分器、案例与运行条件必须匹配。[操作指南](workflow.md) | +| AgentCore Evaluate 结果与运行 spans | 按受限输入约定整理后离线导入,区分拒绝、零分与缺少评判。[导出到审阅](agentcore-first-review.md) | +| 待执行的代码或工具策略 | 使用 Docker 后端执行任务审计或候选评测。[运行审计](../README.md#run-an-audit) | +| 同一记录上的多次评判 | 分别检查分数变化与结论翻转。[中文指南](judge-stability.zh-CN.md) | + +欢迎提交[首次使用反馈](https://github.com/noteflowai/evalarc/issues/new?template=first-use.yml): +你原本要检查什么、在哪一步卡住、结果是否帮助做出决定。安装失败同样有价值。 +只需最小脱敏样例,不必提供客户数据、私有提示词或凭证。该入口用于收集反馈, +不代表已经获得独立用户验证。 diff --git a/docs/outreach/first-use-plan.md b/docs/outreach/first-use-plan.md new file mode 100644 index 0000000..e37aad1 --- /dev/null +++ b/docs/outreach/first-use-plan.md @@ -0,0 +1,45 @@ +# First-use validation plan · 2026-09-16 + +Focus on Agent developers reviewing a change before accepting it. The primary +example is the recorded score increase from 90% to 93.75% with a regressed +notes check. Keep the three-step path visible: inspect, review locally, try +one's own evidence. + +## Ready for independent use + +- English and Chinese first-review guides use the public 0.12.1 wheel and + original release files. They do not require a source checkout or Docker. +- Maintainer verification in a fresh virtual environment outside the checkout + reproduced the regression, suite decisions and trace imports. Restricting + PATH to that environment left no Docker or Node available to those commands. +- The 30-second media uses actual UI screenshots with original evidence + fingerprints. It does not depict a new agent run. +- The first-use issue form asks for the intended decision, first failure, + environment and practical value, using a minimal redacted example. +- Current editorial drafts are tailored to their channels. HelloGitHub's + description is 153 characters, within its current 32–256 requirement. + +These are maintainer checks and prepared invitations, not independent adoption. +The scored AgentCore-shaped controls remain synthetic; live service collection +has not been validated. + +## Next seven days + +1. Keep the existing HF introduction and editorial submissions aligned with + the deployed guide. Read edits back; pending submissions are not listings. +2. Seek three independent first-use attempts with a real review question. + Record setup failures and unsupported formats alongside successful imports. +3. Prioritize fixes that unblock those attempts. Add adapters only for an + actual export format, with a redacted fixture and explicit semantics. +4. Publish a reproducible finding when a user can share one; never relabel + synthetic controls as customer evidence or manufacture usage reports. +5. Review at the end of the week: did anyone use a second record, change a + decision, report a concrete blocker, or contribute an example? + +Three participants is a validation target, not a promised growth outcome. +Maintenance checks, CI clones and release-download verification are logged +separately from independent use. Do not calculate conversion by mixing +GitHub's traffic window with cumulative Stars or attachment downloads. + +Use aggregate channel analytics only where available. No trace upload or +client tracking was added in this update. diff --git a/docs/outreach/hellogithub-submission.md b/docs/outreach/hellogithub-submission.md index 0bef7a7..53fafdb 100644 --- a/docs/outreach/hellogithub-submission.md +++ b/docs/outreach/hellogithub-submission.md @@ -8,35 +8,36 @@ https://github.com/noteflowai/evalarc ### 项目标题 -分享具体失败证据,检查智能体评分器的盲点 +找出更高评分背后的 Agent 回归 ### 项目描述 -EvalArc 是面向 AI 智能体的开源评测与验收工具。支持 Python/JavaScript 候选、TOML 套件、JUnit 和离线报告复核,保留每轮原始记录。交互实验室展示高分仍违反关键业务约束的案例;0.8.0 可下载完整套件证据,并离线复算原始配置、执行计划、验收规则和 JUnit;浏览器可分享到具体调用步骤。提供可筛选 HF Casebook 和中英文文档,现有演示为脚本对照。 +EvalArc 是开源的 Agent 评测复核工具。交互案例展示分数从 90% 升到 93.75%,重试却多写了一条备注。可对照失败检查、定位工具动作,并下载证据在本机复算。适合检查 Agent 变更和验收规则;演示免安装,本地复核无需 Docker、GPU 或模型密钥。当前为研究预览,示例来自脚本对照。 ### 亮点 -- **把讨论定位到同一步**: Copy evidence link 保存任务、对照实现、种子、案例与 trace 步骤。接收者打开同一观察点,键盘可从详情返回案例列表;受限剪贴板有手动复制入口。 -- **失败可恢复**:suite、重复尝试、前后比较和任务包可以分别重试;一个任务包不可用时仍可检查另一个。页面经过 320/390/1440 像素与键盘路径验证。 -- **高分不代表完成**:工单策略得 93.75% 却重复写备注;代码实现得 92.5% 却混淆 JSON true 与 1。可检查全部 15 种声明缺陷,以及三项 suite 作业、六条重复尝试和 167 条审计用例的数据表。 -- **可复核交付**:`evalarc verify` 不执行候选代码,重算单次、重复、对照和完整套件报告,包括 TOML、plan、五次尝试与 JUnit;`--require-accepted` 检查配置的验收规则,`--require-resolved` 另行要求任务完整通过。Python/JavaScript 参考实现共享任务约定。 +- 用一个具体失败解释分数、验收与任务完成的区别。 +- 发布 wheel 与原始记录,可按中英文指南复算结果。 +- 可以比较自己的 EvalArc 评测,或按受限格式导入保存的运行记录。 -本账号为维护者,项目与 AI 结对开发,采用 MIT 许可,处于研究预览阶段。现有展示来自已保存的公开开发任务和脚本 Docker 对照,不提供真实大模型排名或 RL 收益结论。离线一致性校验不等于重新运行评分器,也不认证报告作者;原始候选路径与计时仍是报告元数据。 +### 首次体验 -### 示例代码 +[查看失败案例](https://noteflowai.github.io/evalarc/#regression) · +[首次本地复核](https://github.com/noteflowai/evalarc/blob/main/docs/first-review.zh-CN.md) · +[Hugging Face](https://huggingface.co/spaces/glayguo/evalarc) -安装发布的 Python wheel,从首页下载 suite-evidence.zip 并解压后,无需 Docker 或 Node 即可核验;默认退出 0 表示证据一致,添加 --require-accepted 退出 1 表示严格备注规则拒绝该结果: +按指南安装并解压记录后: ```sh -evalarc verify suite-evidence --json - -# Require every configured suite gate: -evalarc verify suite-evidence --json --require-accepted +evalarc verify evalarc-evidence-explorer/comparison --json +# 退出 0:记录一致。 +evalarc compare evalarc-evidence-explorer/comparison/baseline.json \ + evalarc-evidence-explorer/comparison/current.json --output comparison-review +# 退出 1:一项检查退步,即使总分提高。 ``` ### 截图或演示视频 -在线体验:https://huggingface.co/spaces/glayguo/evalarc -版本:https://github.com/noteflowai/evalarc/releases/tag/v0.8.0 +![分数提高,记录中的检查却退步](https://raw.githubusercontent.com/noteflowai/evalarc/main/docs/assets/first-review.gif) -![分享具体失败证据,检查智能体评分器的盲点](https://github.com/noteflowai/evalarc/raw/main/docs/assets/suite-lab.png) +本账号为维护者,项目与 AI 结对开发,MIT 许可。此案例是脚本 Docker 对照,不是客户效果或模型排名;离线复核不认证生产者或重新执行评分器。云端导入样例采用合成评分,未完成实时 AgentCore 评测。 diff --git a/docs/outreach/huggingface-introduction.md b/docs/outreach/huggingface-introduction.md index c62b144..e55f54e 100644 --- a/docs/outreach/huggingface-introduction.md +++ b/docs/outreach/huggingface-introduction.md @@ -1,9 +1,29 @@ -**Download the evidence. Check every gate.** EvalArc 0.8.0 verifies a whole suite from its original TOML, plan, every repetition and attempt, custom acceptance rules and JUnit. It runs offline without a candidate process, Docker, Node or the original candidate directory. +**The score went up. A previously passing check failed.** -Download **suite-evidence.zip** from the lab and unzip it. `evalarc verify suite-evidence --json` checks all 12 original input files and exits 0 for consistency. Adding `--require-accepted` exits 1: the strict notes gate rejects the faulty policy. Two of three jobs are accepted; only one is fully resolved. Those are separate outcomes. +In EvalArc's recorded support example, two checks improve and the score rises +from 90% to 93.75%. A retry also duplicates a note. The changed check remains +visible instead of being hidden by the average. -[Open the evidence lab](https://huggingface.co/spaces/glayguo/evalarc) · [Release 0.8.0](https://github.com/noteflowai/evalarc/releases/tag/v0.8.0) · [Verification workflow and limits](https://github.com/noteflowai/evalarc/blob/main/docs/verification.md) +[**Inspect the regression →**](https://glayguo-evalarc.static.hf.space/#regression) -The public Casebook still provides 167 audit cases, six repeated attempts and three suite jobs. Historical source bytes and scoring rules are unchanged. These are scripted development controls; offline record consistency does not authenticate the producer or independently rerun the grader. Original candidate paths and durations remain reported metadata. EvalArc remains a research preview. +Follow the recorded action, compare the acceptance rules, then +[recompute the report on your machine](https://github.com/noteflowai/evalarc/blob/main/docs/first-review.md). +The walkthrough uses the published 0.12.1 wheel and downloadable records; +no source checkout, Docker, GPU or model key is needed for that review. +Verification exits 0 for consistent evidence; comparison exits 1 for the +regressed check. -Maintainer update to the existing introduction, developed with AI assistance. Independent community project; no upstream or Hugging Face endorsement is implied. +Have your own records? Compare matching EvalArc evaluations, or use the +[bounded AgentCore export walkthrough](https://github.com/noteflowai/evalarc/blob/main/docs/agentcore-first-review.md). +[First-use feedback](https://github.com/noteflowai/evalarc/issues/new?template=first-use.yml) +about a failed setup or useful finding is welcome. Please use a minimal +redacted example, not a production trace dump. + +The featured case is a scripted Docker control, not a customer incident or +model benchmark. The scored trace-import controls are synthetic; the separate +MCP example records actual local delivery without evaluator scores. No live +AgentCore evaluation or independent adoption is claimed. + +Maintainer update to this existing introduction, developed with AI assistance. +EvalArc is an independent MIT research preview. Offline consistency does not +authenticate the producer or rerun the candidate. diff --git a/docs/outreach/launch.md b/docs/outreach/launch.md index 4217b93..423f55d 100644 --- a/docs/outreach/launch.md +++ b/docs/outreach/launch.md @@ -2,71 +2,38 @@ ## Customer description · 中文 -0.7.1 把证据讨论定位到具体案例和步骤:接收者打开同一观察点,键盘可从详情返回列表;某个分区加载失败时单独重试,其余记录继续可用。已验证桌面、手机和 320 像素窄屏,便于客户现场演示与远程复核。 - -新版 0.7 增加独立交付复核:收到单次、重复或前后对照报告后,可在不执行候选 -程序、不连接 Docker 的情况下重算已记录结果,并列出输入文件指纹。 -“报告内部一致”与“任务全部通过”分别判断,适合团队交接、客户复核和 CI 证据检查。 -这是对已记录检查的验证,不替代重新运行评分器,也不认证报告发布者。 - - -EvalArc 是面向 AI 智能体的开源评测与评分器审计工具,帮助团队检查“高分” -背后是否仍存在关键交付缺陷。当前提供代码服务和业务工具两个任务环境, -通过正确实现与 15 种已声明缺陷进行对照,验证事务、异常恢复、工具重试、 -幂等写入及最终业务状态。在线演示可逐步查看调用、失败检查和状态变化, -并下载带有运行配置与指纹的原始证据。v0.5 新增 TOML 评测套件、自定义验收规则 -和 JUnit 导出:同样是 93.75%、没有完全通过的尝试,宽松规则可允许部分进展, -要求备注检查全部通过的规则则拒绝。页面分别展示任务完成情况和规则决策, -并保留重复评测、版本回归与全部运行证据,适合 PAI 智能体创新场景的评测设计、 -验收验证与技术交流。HF Casebook 进一步提供可筛选、可用 Python 读取的公开 -证据表,分别保留 167 条用例、6 次重复尝试和 3 项验收作业及其原始记录。 -支持生成 Python/JavaScript 候选,并通过独立实现验证同一任务约定, -可在统一套件中为不同语言明确指定运行镜像。 -当前为研究预览版,演示使用脚本对照,尚未给出真实大模型 -性能或强化学习收益结论。 +EvalArc 帮助 Agent 开发团队找出更高评分背后的关键退步:对照变更前后的检查,定位已记录的工具动作,并交付可离线复核的验收证据。公开演示中,分数从 90% 提高到 93.75%,重试却多写了一条备注;用户可以直接查看原因,再用发布的 Python wheel 在本机复算。适合 PAI 智能体场景的评测设计、发布复核与技术交流。当前为 MIT 研究预览,展示案例来自脚本对照,不代表客户生产效果或模型排名。 ## English -EvalArc audits the graders behind AI-agent evaluations. Its two executable task -packs compare known-good implementations with 15 declared faults, then preserve -the checks, state changes and reproduction metadata. In the evidence lab, a -support policy scores 93.75% while duplicating a write; a coding artifact scores -92.5% while violating compare-and-swap. Explore why each fails acceptance, -compare the reference, and download the full audit. MIT licensed, CPU-friendly, -and currently a research preview with scripted controls, not model rankings. -Version 0.3 adds runtime readiness checks, standalone evaluation reports, and -matched revision comparisons. Its new showcase exposes a regressed note check -even as two closure improvements raise the score from 90% to 93.75%. -Version 0.4 preserves every attempt of a frozen candidate, with explicit -assessed/invalid denominators, case deadlines and progress events. The browser -exposes six recorded Docker attempts: the reference resolves 3/3; the faulty -policy resolves 0/3 despite three 93.75% scores. No check variation was observed, -and repeated public cases do not establish general model reliability. -Version 0.5 coordinates multi-task TOML suites with explicit per-job gates and -JUnit exports. The new recorded showcase keeps scores and full resolution -separate from configurable acceptance: the same faulty policy meets a -permissive threshold but fails a gate requiring every notes check. Three jobs, -five attempts and the complete configuration remain inspectable. No hosted -CI importer or model-provider integration is claimed. +EvalArc helps agent developers find regressions hidden by a better average +score. Compare changed checks, follow recorded tool actions, and hand off +evidence another developer can verify offline. In the public example, the +score rises from 90% to 93.75% while a retry duplicates a note. Explore the +failure without installing anything, then recompute the comparison using the +published Python wheel. MIT research preview; the featured case is a scripted +control, not a customer incident or a model ranking. -The Hugging Face Casebook makes 167 audit cases, six repeated attempts and three -suite jobs filterable and readable from Python as separate development -configurations. Each row retains a pointer and hash for its original source -record. This is a tabular view of existing evidence, not additional model trials. +## One useful invitation -Version 0.6 supplies Python and JavaScript starters and independent references -for both tasks. The Node audits exercise the same declared faults, and a -mixed-language suite retains three resolved jobs. Shared task contracts do not -establish a language ranking. +Have two agent revisions you need to review? Try one comparison and tell us +where the first review helped or got stuck. A minimal redacted example is +enough. Installation failures and unsupported export formats are useful +feedback too. -## Entry points +Use the [first-use form](https://github.com/noteflowai/evalarc/issues/new?template=first-use.yml). +An invitation is not evidence of independent adoption. -- Source: https://github.com/noteflowai/evalarc -- Demo: https://huggingface.co/spaces/glayguo/evalarc -- Web mirror: https://noteflowai.github.io/evalarc/ -- Releases: https://github.com/noteflowai/evalarc/releases -- Data: https://huggingface.co/datasets/glayguo/evalarc-casebook +## Entry points -For each community, write a description appropriate to its rules and audience. -Disclose maintainer affiliation. Do not present pending editorial submissions -as an endorsement or listing. +- [Recorded regression](https://noteflowai.github.io/evalarc/#regression) +- [First local review](https://github.com/noteflowai/evalarc/blob/main/docs/first-review.md) +- [首次本地复核](https://github.com/noteflowai/evalarc/blob/main/docs/first-review.zh-CN.md) +- [Hugging Face Space](https://huggingface.co/spaces/glayguo/evalarc) +- [Source](https://github.com/noteflowai/evalarc) +- [Release files](https://github.com/noteflowai/evalarc/releases/tag/v0.12.1) + +Use one relevant finding for each audience. Maintain the existing weekly, +HelloGitHub and HF introductions instead of appending a release chronology. +Disclose maintainer affiliation, follow channel rules, and distinguish pending +editorial review from acceptance. diff --git a/docs/outreach/weekly-submission.md b/docs/outreach/weekly-submission.md index 755c96a..70c7c15 100644 --- a/docs/outreach/weekly-submission.md +++ b/docs/outreach/weekly-submission.md @@ -1,24 +1,18 @@ -维护者自荐:**evalarc**。EvalArc 是面向 AI 智能体的开源评测与验收工具。支持 Python/JavaScript 候选、TOML 套件、JUnit 和离线报告复核,保留每轮原始记录。交互实验室展示高分仍违反关键业务约束的案例;0.8.0 可下载完整套件证据,并离线复算原始配置、执行计划、验收规则和 JUnit;浏览器可分享到具体调用步骤。提供可筛选 HF Casebook 和中英文文档,现有演示为脚本对照。 +维护者自荐:**EvalArc——分数提高,Agent 就更可靠了吗?** -在线体验:https://huggingface.co/spaces/glayguo/evalarc +一个已保存的对照案例中,分数从 90% 提高到 93.75%,两项检查改善,但重试多写了一条备注。EvalArc 可以并排查看退步的检查、逐步定位工具动作,并下载原始证据在本机复算。 -- **把讨论定位到同一步**: Copy evidence link 保存任务、对照实现、种子、案例与 trace 步骤。接收者打开同一观察点,键盘可从详情返回案例列表;受限剪贴板有手动复制入口。 -- **失败可恢复**:suite、重复尝试、前后比较和任务包可以分别重试;一个任务包不可用时仍可检查另一个。页面经过 320/390/1440 像素与键盘路径验证。 -- **高分不代表完成**:工单策略得 93.75% 却重复写备注;代码实现得 92.5% 却混淆 JSON true 与 1。可检查全部 15 种声明缺陷,以及三项 suite 作业、六条重复尝试和 167 条审计用例的数据表。 -- **可复核交付**:`evalarc verify` 不执行候选代码,重算单次、重复、对照和完整套件报告,包括 TOML、plan、五次尝试与 JUnit;`--require-accepted` 检查配置的验收规则,`--require-resolved` 另行要求任务完整通过。Python/JavaScript 参考实现共享任务约定。 +**先看一个问题,再决定是否适合自己的项目:** -```sh -evalarc verify suite-evidence --json +1. [打开免安装演示](https://noteflowai.github.io/evalarc/#regression),查看 `retry-after-commit`。 +2. 检查相同分数为什么通过宽松规则、却被严格备注规则拒绝。 +3. 按[首次复核指南](https://github.com/noteflowai/evalarc/blob/main/docs/first-review.zh-CN.md)安装发布的 wheel,重建对照报告;无需 Docker、GPU 或模型密钥。 -# Require every configured suite gate: -evalarc verify suite-evidence --json --require-accepted -``` - -安装发布的 Python wheel,从首页下载 suite-evidence.zip 并解压后,无需 Docker 或 Node 即可核验;默认退出 0 表示证据一致,添加 --require-accepted 退出 1 表示严格备注规则拒绝该结果: +适合需要复查 Agent 变更与发布验收的开发者。已有自己的记录时,可比较 EvalArc 评测,或按照受限格式导入保存的 AgentCore Evaluate 结果。欢迎反馈首次使用的卡点,不需要提供私有数据。 项目:https://github.com/noteflowai/evalarc -版本:https://github.com/noteflowai/evalarc/releases/tag/v0.8.0 +Hugging Face:https://huggingface.co/spaces/glayguo/evalarc -![分享具体失败证据,检查智能体评分器的盲点](https://github.com/noteflowai/evalarc/raw/main/docs/assets/suite-lab.png) +![分数提高,记录中的检查却退步](https://raw.githubusercontent.com/noteflowai/evalarc/main/docs/assets/first-review.gif) -本账号为维护者,项目与 AI 结对开发,采用 MIT 许可,处于研究预览阶段。现有展示来自已保存的公开开发任务和脚本 Docker 对照,不提供真实大模型排名或 RL 收益结论。离线一致性校验不等于重新运行评分器,也不认证报告作者;原始候选路径与计时仍是报告元数据。 +本账号为维护者,项目与 AI 结对开发,MIT 许可,处于研究预览。此案例是已保存的脚本 Docker 对照,不是客户事故或模型排名;离线复核不认证生产者或重新运行评分器。云端导入示例采用合成评分,未完成实时 AgentCore 评测。 diff --git a/huggingface/README.md b/huggingface/README.md index 4a6b77e..b564ea2 100644 --- a/huggingface/README.md +++ b/huggingface/README.md @@ -7,7 +7,7 @@ sdk: static app_file: index.html pinned: false license: mit -short_description: Inspect scores, acceptance gates and every agent attempt. +short_description: Find agent regressions behind a better score. Inspect the evidence. tags: - agent-evaluation - tool-use @@ -16,99 +16,70 @@ tags: - developer-tools --- -# EvalArc — Look past the score. - -**New in v0.12: same trace, same verdict?** Inspect three saved judgments for -each of five synthetic controls. Separate changing scores, flipped acceptance -decisions and unavailable results; download the complete evidence for offline -verification. All-reject agreement remains rejection. No model or AWS evaluation -was run for these controls. [Judge Stability](https://glayguo-evalarc.static.hf.space/judge-stability/index.html) -· [Import guide](https://github.com/noteflowai/evalarc/blob/main/docs/judge-stability.md). - -**Share the exact evidence.** Select a case and trace step, then **Copy evidence -link**. Recipients reopen the same recorded observation. Keyboard users can -inspect a case and return to its list; failed sections can be retried separately. -If one task pack fails to load, the other remains inspectable. - -**Python and JavaScript candidates.** Generate starters or independent -references for all three task packs, audit the declared faults, and combine runtimes -in one suite. The task contracts and graders are unchanged. See the -[language guide](https://github.com/noteflowai/evalarc/blob/main/docs/languages.md) -and recorded mixed-language Docker suite; this does not establish a language ranking. - -**The score rose from 90% to 93.75%. A previously passing check now fails.** - -**New in v0.5: same score, different gate.** Two support jobs use the same -frozen defective policy and score 93.75%, with 0/2 resolved attempts each. -A deliberately permissive gate accepts the partial result; requiring every -notes check rejects it. Inspect the three-job Docker suite, all five attempts, -the original TOML, and JUnit output distinguishing a failed gate from an -environment error. Gate acceptance remains separate from full task resolution. -A hosted CI importer was not exercised. - -**Explore the data as tables:** the -[EvalArc Casebook](https://huggingface.co/datasets/glayguo/evalarc-casebook) -offers three separate configurations for 251 audit cases, six repeated attempts -and three suite jobs. Filter the results or load the JSONL in Python; original -source records and fingerprints accompany every row. - -**Every v0.4 attempt remains visible.** Switch between three -recorded Docker attempts of the reference and three of the duplicate-write -control. The reference resolves 3/3 attempts; the faulty control resolves 0/3 -despite a mean score of 93.75%. Open every attempt, inspect per-check -denominators, and download the full summary and progress JSONL. -No check variation was observed in either scripted control. - -The **v0.3 comparison** remains available: compare two recorded support -policies side by side. Two closure -checks improve, while a retry introduces a duplicate note. Inspect all three -changed cases, then open the standalone comparison and individual reports. -`evalarc compare` returns exit code 1 for the regression despite the higher score. - -Explore saved evidence from two executable task packs for AI-agent evaluation. -Switch between known-good references and 15 declared faulty implementations, -inspect failed checks, and step through tool calls and state changes. - -- **Support tools:** a note is committed, its response fails, and a retry with a - new idempotency key duplicates it. The recorded partial score is 0.9375; - full resolution fails. -- **Coding artifacts:** treating JSON `true` and `1` as equal breaks - compare-and-swap. The recorded partial score is 0.925; full resolution fails. -- **Reproduction:** full JSON, grader/candidate/case fingerprints, seeds and - container image IDs accompany the reports. - -The browser replays the committed Docker audits and evaluation records. It does not execute arbitrary -submissions or call a model. The Python CLI has no third-party runtime -dependencies; the bundled trusted controls can run on a CPU. - -[Source & quickstart](https://github.com/noteflowai/evalarc) · -[中文说明](https://github.com/noteflowai/evalarc/blob/main/README.zh-CN.md) · +# EvalArc — Find the regression behind the score. + +**90% → 93.75%. Two checks improve. One previously passing check fails.** +A tool committed a note but returned an error. Retrying with a new key wrote it +again. Inspect the recorded regression, follow the action, and check the +acceptance rule before trusting the higher score. + +[**Try the revision comparison →**](https://glayguo-evalarc.static.hf.space/#regression) +· [First local review](https://github.com/noteflowai/evalarc/blob/main/docs/first-review.md) +· [中文](https://github.com/noteflowai/evalarc/blob/main/README.zh-CN.md) + +No account, installation or model key is needed to explore the saved records. +The featured comparison contains scripted Docker controls, not customer data +or a model leaderboard. + +## Start with one review + +1. Compare the two revisions and select `retry-after-commit`. +2. Step through the action that duplicates the note in the case explorer. +3. Compare the permissive gate with the strict notes gate. The score stays + 93.75%; acceptance changes with the declared rule. +4. [Install the published wheel and recompute the report](https://github.com/noteflowai/evalarc/blob/main/docs/first-review.md). + The local review needs Python 3.11+, with no Docker, Node, GPU or model call. + +## Bring your own evidence + +- **Matching EvalArc evaluations:** compare changed checks and retain original inputs. +- **Saved AgentCore Evaluate results and spans:** use the + [export-to-review walkthrough](https://github.com/noteflowai/evalarc/blob/main/docs/agentcore-first-review.md). + The adapter accepts a bounded input wrapper, not arbitrary cloud exports. +- **Repeated judgments on a fixed trace:** inspect + [score variation and verdict disagreement](https://glayguo-evalarc.static.hf.space/judge-stability/index.html) + separately from missing judgments. The five displayed controls are synthetic. +- **A received suite:** download `suite-evidence.zip` and verify the preserved + configuration, plan, attempts, gates and JUnit offline. + +Trying your own records? [Share a first-use finding or setup problem](https://github.com/noteflowai/evalarc/issues/new?template=first-use.yml). +A minimal redacted example is enough. + +## Evidence and scope + +The saved audits detect **21 declared faults across three task packs**; six +faults each depend on one detecting case. The +[Casebook](https://huggingface.co/datasets/glayguo/evalarc-casebook) contains +251 audit case records, six repeated attempts and three suite jobs in separate +configurations. These public-development controls do not establish coverage +of unseen faults or general model performance. + +The recorded suite has two accepted jobs and one fully resolved job out of +three. Consistency, acceptance and full resolution are separate outcomes. +Offline verification recomputes supplied records; it does not authenticate +the producer or rerun the candidate. + +The Trace Workbench's five scored controls are synthetic. Its separate MCP +example contains real local delivery and no evaluator scores; no live +AgentCore evaluation is claimed. The +[GPU pilot](https://glayguo-evalarc.static.hf.space/skill-impact/index.html) +contains 27 recorded model trials with every failure retained; it does not +establish skill efficacy or a model ranking. + +MIT · Research preview. `manifest.json` identifies the deployed source commit +and SHA-256 of published files. Maintainer publication does not imply +Hugging Face endorsement. + +[Source](https://github.com/noteflowai/evalarc) · [Methodology](https://github.com/noteflowai/evalarc/blob/main/docs/methodology.md) · -[Security boundaries](https://github.com/noteflowai/evalarc/blob/main/SECURITY.md) - -## Scope - -Research preview 0.8.0. These are scripted controls and public development -tasks, not held-out frontier-model results. Detection applies only to the -declared faults. No arbitrary reward-hack resistance, human time horizon, -hardware-agent validation or RL improvement is established. Repeated fixed -cases do not establish reliability on unseen tasks or a model success rate. - -The source and evidence are MIT licensed. `manifest.json` identifies the source -commit and SHA-256 of each published file. Publication is performed by the -maintainer and does not imply endorsement by Hugging Face. - -### Verify a handoff offline - -With EvalArc 0.8+, download **suite-evidence.zip** from the lab and unzip it. -Run `evalarc verify suite-evidence --json` to check the original configuration, -plan, all five attempts, custom gates and JUnit without executing a candidate. -The 12 original input files are unchanged: two of three jobs are accepted, one -is fully resolved. Default verification exits 0 for consistency; -`--require-accepted` exits 1 because the strict notes gate rejects the result. -`--require-resolved` separately requires full resolution. -[Workflow and limits](https://github.com/noteflowai/evalarc/blob/main/docs/verification.md). - -## Recorded GPU pilot - -[Inspect 27 actual model trials](https://noteflowai.github.io/evalarc/skill-impact/): no skill, direct delivery and real MCP, with independent grading and every failure retained. [Composition and public-session handoff records](https://noteflowai.github.io/evalarc/research/) are separate small pilots. No accuracy or memory efficacy gain is claimed. +[Changelog](https://github.com/noteflowai/evalarc/blob/main/CHANGELOG.md) diff --git a/scripts/build_site.py b/scripts/build_site.py index ab53f41..c73a364 100644 --- a/scripts/build_site.py +++ b/scripts/build_site.py @@ -362,6 +362,12 @@ def build(destination: Path) -> dict: ) else: shutil.copyfile(path, destination / path.name) + for suffix in ("png", "gif", "mp4", "vtt"): + name = f"first-review.{suffix}" + shutil.copyfile(ROOT / "docs" / "assets" / name, destination / name) + shutil.copyfile( + ROOT / "docs/assets/first-review-media.json", destination / "first-review-media.json" + ) for name, directory, *_ in specifications: target = destination / name target.mkdir() diff --git a/scripts/capture_first_review.cjs b/scripts/capture_first_review.cjs new file mode 100644 index 0000000..2fbe461 --- /dev/null +++ b/scripts/capture_first_review.cjs @@ -0,0 +1,95 @@ +// Capture actual UI states; annotate them without changing any displayed evidence. +// Run after build_site.py, with SITE_URL pointing to a local HTTP server. +// FFMPEG must name an ffmpeg executable with libx264 and GIF encoding support. +const {chromium} = require("playwright"); +const assert = require("node:assert/strict"); +const fs = require("node:fs"); +const os = require("node:os"); +const path = require("node:path"); +const crypto = require("node:crypto"); +const {execFileSync} = require("node:child_process"); + +async function main() { + const base = process.env.SITE_URL || "http://127.0.0.1:8767/"; + const output = path.resolve("docs/assets"); + const temporary = fs.mkdtempSync(path.join(os.tmpdir(), "evalarc-tour-")); + const browser = await chromium.launch({headless:true}); + const frames = []; + try { + const page = await browser.newPage({viewport:{width:1280,height:1100}}); + await page.emulateMedia({reducedMotion:"reduce"}); + await page.goto(base, {waitUntil:"networkidle"}); + await page.locator("#suite-workspace").waitFor({state:"visible"}); + assert.equal(await page.locator("#before-score").innerText(), "90%"); + assert.equal(await page.locator("#after-score").innerText(), "93.75%"); + assert.equal(await page.locator("#regression-count").innerText(), "1 REGRESSED CHECK"); + async function capture(selector, title, detail) { + const image = await page.locator(selector).screenshot(); + frames.push({title,detail,image}); + } + await capture("#comparison-workspace", "A higher score. A new regression.", + "90% → 93.75% · Two closure checks improve; the notes check regresses."); + await page.locator("#next-step").click(); + await page.locator("#next-step").click(); + assert.match(await page.locator("#trace-action").innerText(), /temporarily_unavailable/); + assert.match(await page.locator("#trace-changes").innerText(), /Reviewed request/); + await capture("#trace-panel", "The tool returned an error. The note was written.", + "Recorded step 3 · The state changed before the transient error returned."); + await page.locator("#next-step").click(); + const change = JSON.parse(await page.locator("#trace-changes").innerText()); + assert.equal(Object.values(change)[0].after.notes.length, 3); + await capture("#trace-panel", "A new key. A duplicate note.", + "Recorded step 4 · Retrying the write with a different key adds the note again."); + assert.equal(await page.locator("#suite-partial-verdict").innerText(), "GATE ACCEPTED"); + assert.equal(await page.locator("#suite-protected-verdict").innerText(), "GATE REJECTED"); + await capture(".suite-comparison", "Keep the score. Check the acceptance rule.", + "Same 93.75% policy · The strict notes gate rejects what the permissive gate accepts."); + const canvas = await browser.newPage({viewport:{width:1280,height:960},deviceScaleFactor:1}); + for (const [index,frame] of frames.entries()) { + await canvas.setContent(` + +
EVALARC / FIRST REVIEW / ${index+1} OF 4

+ Captured evidence view`); + await canvas.locator("h1").evaluate((node,text) => {node.textContent=text}, frame.title); + await canvas.locator("p").evaluate((node,text) => {node.textContent=text}, frame.detail); + await canvas.locator("img").evaluate((node,src) => {node.src=src}, "data:image/png;base64,"+frame.image.toString("base64")); + await canvas.locator("img").evaluate(node => node.decode()); + await canvas.screenshot({path:path.join(temporary,`frame-${index}.png`)}); + } + fs.copyFileSync(path.join(temporary,"frame-0.png"),path.join(output,"first-review.png")); + const sequence = frames.map((_,i) => `file '${path.join(temporary,`frame-${i}.png`)}'\nduration 7.5`).join("\n"); + const list = path.join(temporary,"frames.txt"); + fs.writeFileSync(list,sequence+`\nfile '${path.join(temporary,"frame-3.png")}'\n`); + const ffmpeg = process.env.FFMPEG || "ffmpeg"; + execFileSync(ffmpeg,["-v","error","-y","-f","concat","-safe","0","-i",list,"-t","30", + "-vf","fps=10,scale=1000:-2:flags=lanczos","-c:v","libx264","-crf","22","-pix_fmt","yuv420p", + "-movflags","+faststart",path.join(output,"first-review.mp4")],{stdio:"inherit"}); + execFileSync(ffmpeg,["-v","error","-y","-f","concat","-safe","0","-i",list,"-t","30", + "-filter_complex","fps=2,scale=960:-2:flags=lanczos,split[a][b];[a]palettegen[p];[b][p]paletteuse", + "-loop","0",path.join(output,"first-review.gif")],{stdio:"inherit"}); + const cues = ["00:00:00.000 --> 00:00:07.500","00:00:07.500 --> 00:00:15.000", + "00:00:15.000 --> 00:00:22.500","00:00:22.500 --> 00:00:30.000"]; + fs.writeFileSync(path.join(output,"first-review.vtt"),"WEBVTT\n\n"+frames.map((f,i) => + `${cues[i]}\n${f.title}\n${f.detail}\n`).join("\n")); + const manifest = await (await page.request.get(new URL("manifest.json",base).href)).json(); + const evidence = ["comparison/baseline.json","comparison/current.json","support/audit.json","suite/suite.json"]; + const files = Object.fromEntries(["png","gif","mp4","vtt"].map(ext => { + const name="first-review."+ext, bytes=fs.readFileSync(path.join(output,name)); + return [name,{bytes:bytes.length,sha256:crypto.createHash("sha256").update(bytes).digest("hex")}]; + })); + fs.writeFileSync(path.join(output,"first-review-media.json"),JSON.stringify({ + kind:"annotated screenshots of actual UI states",duration_seconds:30, + evidence:Object.fromEntries(evidence.map(name => [name,manifest.files[name]])), + frames:frames.map(({title,detail}) => ({title,detail})),files, + scope:"Saved scripted Docker controls. No new agent execution or customer evidence.", + },null,2)+"\n"); + console.log(JSON.stringify({frames:frames.length,duration_seconds:30,files},null,2)); + } finally { + await browser.close(); + fs.rmSync(temporary,{recursive:true,force:true}); + } +} +main().catch(error => {console.error(error);process.exitCode=1}); diff --git a/scripts/check_site.cjs b/scripts/check_site.cjs index e38d234..7949787 100644 --- a/scripts/check_site.cjs +++ b/scripts/check_site.cjs @@ -10,7 +10,7 @@ async function main() { let server; let base = process.env.SITE_URL; if (!base) { - const mime = {".html":"text/html", ".js":"text/javascript", ".json":"application/json", ".css":"text/css", ".svg":"image/svg+xml", ".zip":"application/zip"}; + const mime = {".html":"text/html", ".js":"text/javascript", ".json":"application/json", ".css":"text/css", ".svg":"image/svg+xml", ".zip":"application/zip", ".png":"image/png", ".gif":"image/gif", ".mp4":"video/mp4", ".vtt":"text/vtt"}; server = http.createServer((req, res) => { const route = decodeURIComponent(new URL(req.url, "http://localhost").pathname); // HF static hosting does not resolve subdirectory indexes. Keep that @@ -45,6 +45,29 @@ async function main() { await app.locator("#comparison-workspace").waitFor({state:"visible"}); await app.locator("#repeat-workspace").waitFor({state:"visible"}); await app.locator("#suite-workspace").waitFor({state:"visible"}); + await app.getByRole("link", {name:"Find the regression", exact:false}).click(); + assert.equal(await app.locator("body").evaluate(() => location.hash), "#regression"); + const comparisonHeading = await app.locator("#regression h2").boundingBox(); + assert(comparisonHeading && comparisonHeading.y >= 0 && comparisonHeading.y < 1000, + "The primary action must expose the recorded regression"); + if (width === 1440) { + await app.locator(".tour summary").click(); + const video = app.getByLabel("Recorded regression walkthrough"); + assert.equal(await video.getAttribute("autoplay"), null); + assert.equal(await video.getAttribute("preload"), "none"); + const playback = await video.evaluate(async element => { + await element.play(); + await new Promise((resolve, reject) => { + const timer = setTimeout(() => reject(new Error("Walkthrough playback stalled")), 15000); + element.addEventListener("timeupdate", () => {clearTimeout(timer);resolve()}, {once:true}); + }); + element.pause(); + return {duration:element.duration,time:element.currentTime,width:element.videoWidth}; + }); + assert(Math.abs(playback.duration - 30) < 0.2 && playback.time > 0 && playback.width > 0, + "The walkthrough must decode and play, not just show its poster"); + await app.locator(".tour summary").click(); + } assert.equal(await app.locator(".coverage-card").count(),3); assert.equal(await app.locator(".coverage-card li").count(),6); await page.waitForLoadState("networkidle"); diff --git a/site/index.html b/site/index.html index b8ed368..ffe65ee 100644 --- a/site/index.html +++ b/site/index.html @@ -3,7 +3,7 @@ - + EvalArc | Look past the score. @@ -22,10 +22,10 @@

OPEN SOURCE   /   RESEARCH PREVIEW __EVALARC_VERSION__

Look past
the score.

A tool returned an error. The write already happened.
A retry added the same note twice.

-

EvalArc tests whether an agent's grader catches the difference between plausible progress and a completed task.

- Inspect the acceptance gate ↘ - Every repeated attempt ↓ - Compare revision regressions ↓ +

Review an agent change before you accept it. Compare failed checks, follow the recorded actions, and keep the evidence behind your decision.

+ Find the regression ↘ +

Interactive recorded example · no account, install or model key

+ Run your first local review ↗
RECORDED CONTROL / 01NOT RESOLVED
@@ -36,79 +36,30 @@

Look past
the score.

Scripted negative control · public seed 17
No model call or live customer data.

-

Generate Python or JavaScript candidates and audit independent controls against the same task contracts. Try both runtimes ↗. The recorded showcases below retain their original versions and fingerprints.

-

NEW / RECORDED MODEL EVIDENCE

A skill loaded. Did the task pass?

27 real Qwen3-8B trials compare no skill, direct loading and MCP delivery on attributed robot recordings. Inspect every tool receipt, candidate, independent score and ATIF trajectory. Three engineering profiles, including negative results.

L40S recordings; one public development task. No skill accuracy gain or general model ranking is claimed.

-

NEW / BRING YOUR AGENT RECORDS

Zero, skipped, or missing?

A zero can be a valid judgment. A skipped evaluator needs context. A required skill may never have loaded. Inspect each against a versioned golden case, with recording identity and explicit acceptance rules.

The five controls are synthetic. The separate MCP record contains a real local load and no evaluator scores. Review your own saved AgentCore Evaluate export offline with evalarc trace-import.

-
-

NEW / FIXED TRACE, REPEATED JUDGMENTS

-

Same trace. Same verdict?

-

Keep the execution fixed. Inspect three saved judgments per target: a changing score, a flipped acceptance decision, and unavailable assessments are different findings.

-
- __JUDGE_PREVIEW__
Five synthetic controls · configured pass threshold 0.8 · no model or AWS call
ControlJudge run 1Judge run 2Judge run 3
-
- -

Agreement is not accuracy: three rejections still reject the case. Agent reruns remain separate in the execution repeatability report.

+ +
+ 01 / SEE THE CHANGE90% → 93.75%. One regression.Compare the checks behind the improved score. + 02 / FOLLOW THE RETRYThe note was already written.Step through the recorded action and changed state. + 03 / CHECK IT LOCALLYKeep the evidence and the rules.Install the wheel and rebuild the report. No Docker needed.
+
+ Watch the 30-second walkthrough + +

Four annotated views of saved scripted Docker controls. No audio or new agent run. Transcript and recording method ↗

+
__TASK_PACK_COUNT__task packs
__FAULT_COUNT__declared faults caught
JSONdownloadable evidence
CPUreproduce without a model API
- -
-

THREE TASK PACKS / DECLARED FAULTS

-

A perfect score.
How much coverage remains?

-

A fault detected by one case loses its coverage if that case is removed or weakened. Open each dependency to inspect the recorded checks, status and seed. Margins count distinct case IDs, not repeated runs.

-
__AUDIT_COVERAGE__
-

These summaries derive from the original scripted audit records. They are not new agent runs or a guarantee against unseen faults.

-
-
-

NEW IN 0.5 / DECLARE YOUR ACCEPTANCE RULES

Same score. Different gate.

One frozen policy.
Two explicit acceptance rules.

-

Both support jobs score 93.75% and resolve 0/2 attempts. A deliberately permissive gate accepts the result; requiring every notes check to pass rejects it. The task outcome stays the same.

-

Loading the recorded suite…

- - -

Recorded Docker suite: three jobs, five attempts, 31 case executions. JUnit distinguishes a rejected gate from an environment error; a hosted CI importer was not exercised. This is scripted development evidence.

- -

Prefer a filterable table or Python? Open the Hugging Face Casebook ↗ and compare gate_accepted with fully_resolved. The original records accompany every row.

-
-
-

NEW IN 0.4 / REPEAT ONE FROZEN CANDIDATE

One run is not the whole record.

Same cases. Fresh state.
Every attempt stays visible.

-

Three Docker attempts per scripted control. The reference passes every time; the duplicate-write policy keeps its 93.75% score and fails acceptance every time. Switch controls and open any attempt's full evidence.

-
- - -
-

Loading the recorded attempts…

- - -

These v0.4 recordings show fixed public cases, with no observed check variation in either control. They are not stochastic model trials or an estimate of reliability on unseen tasks. The v0.3 comparison below retains its original records and grading fingerprint.

- -
+
-

NEW IN 0.3 / COMPARE TWO REVISIONS

Better score. New failure.

Two closure checks improve.
A previously passing note check fails.

+

COMPARE TWO RECORDED REVISIONS

Better score. New failure.

Two closure checks improve.
A previously passing note check fails.

These recorded policies use the same task, seeds, grader and runtime. Select a changed case to compare the observed outcomes.

Loading the recorded comparison…

@@ -126,10 +77,10 @@

A perfect score.
How much coverage remains?

After

Final ticket state

-
Reproduce this comparison without running candidates
evalarc compare examples/comparison/baseline.json \
-  examples/comparison/current.json \
+        
Reproduce this comparison without running candidates
evalarc compare evalarc-evidence-explorer/comparison/baseline.json \
+  evalarc-evidence-explorer/comparison/current.json \
   --output runs/comparison-001
-# Exit 1: a check regressed, despite the higher score.

Use a fresh output directory. The CLI validates report consistency and matching execution conditions before comparison. It does not authenticate the producer of the reports.

+# Exit 1: a check regressed, despite the higher score.

Download the records using the first-review guide and use a fresh output directory. The CLI validates report consistency and matching execution conditions before comparison. It does not authenticate the producer of the reports.

@@ -177,8 +128,76 @@

A perfect score.
How much coverage remains?

+
+

DECLARE YOUR ACCEPTANCE RULES

Same score. Different gate.

One frozen policy.
Two explicit acceptance rules.

+

Both support jobs score 93.75% and resolve 0/2 attempts. A deliberately permissive gate accepts the result; requiring every notes check to pass rejects it. The task outcome stays the same.

+

Loading the recorded suite…

+ + +

Recorded Docker suite: three jobs, five attempts, 31 case executions. JUnit distinguishes a rejected gate from an environment error; a hosted CI importer was not exercised. This is scripted development evidence.

+ +

Prefer a filterable table or Python? Open the Hugging Face Casebook ↗ and compare gate_accepted with fully_resolved. The original records accompany every row.

+
+
+

REPEAT ONE FROZEN CANDIDATE

One run is not the whole record.

Same cases. Fresh state.
Every attempt stays visible.

+

Three Docker attempts per scripted control. The reference passes every time; the duplicate-write policy keeps its 93.75% score and fails acceptance every time. Switch controls and open any attempt's full evidence.

+
+ + +
+

Loading the recorded attempts…

+ + +

These v0.4 recordings show fixed public cases, with no observed check variation in either control. They are not stochastic model trials or an estimate of reliability on unseen tasks. The v0.3 comparison below retains its original records and grading fingerprint.

+ +
+
+

THREE TASK PACKS / DECLARED FAULTS

+

A perfect score.
How much coverage remains?

+

A fault detected by one case loses its coverage if that case is removed or weakened. Open each dependency to inspect the recorded checks, status and seed. Margins count distinct case IDs, not repeated runs.

+
__AUDIT_COVERAGE__
+

These summaries derive from the original scripted audit records. They are not new agent runs or a guarantee against unseen faults.

+
+
+

YOUR NEXT REVIEW

Bring a change
you need to trust.

Start with two EvalArc evaluations, or prepare a saved AgentCore export using the documented input contract. Review locally, then share a minimal finding.

Export-to-review walkthrough ↗

Bounded import format; no automatic cloud collection. The scored trace controls are synthetic. Live AgentCore evaluation has not been validated by this demo.

+

FIRST-USE FEEDBACK

Where did the review help?

Tell us what you were checking, where you got stuck, and whether the result changed a decision. A failed setup is useful feedback too.

Share a first-use report ↗

Use a minimal redacted example. No production traces, private prompts or credentials are needed.

+
+

SAVED TRACE REVIEW

Zero, skipped, or missing?

A zero can be a valid judgment. A skipped evaluator needs context. A required skill may never have loaded. Inspect each against a versioned golden case, with recording identity and explicit acceptance rules.

The five controls are synthetic. The separate MCP record contains a real local load and no evaluator scores. Review your own saved AgentCore Evaluate export offline with evalarc trace-import.

+
+

FIXED TRACE, REPEATED JUDGMENTS

+

Same trace. Same verdict?

+

Keep the execution fixed. Inspect three saved judgments per target: a changing score, a flipped acceptance decision, and unavailable assessments are different findings.

+
+ __JUDGE_PREVIEW__
Five synthetic controls · configured pass threshold 0.8 · no model or AWS call
ControlJudge run 1Judge run 2Judge run 3
+
+ +

Agreement is not accuracy: three rejections still reject the case. Agent reruns remain separate in the execution repeatability report.

+
+

RECORDED MODEL EVIDENCE

A skill loaded. Did the task pass?

27 real Qwen3-8B trials compare no skill, direct loading and MCP delivery on attributed robot recordings. Inspect every tool receipt, candidate, independent score and ATIF trajectory. Three engineering profiles, including negative results.

L40S recordings; one public development task. No skill accuracy gain or general model ranking is claimed.

+
-

HAND OFF EVIDENCE

Download the evidence.
Check every gate.

Recompute the original suite configuration, plan, five attempts, acceptance decisions and JUnit without executing a candidate.

Download suite evidence (ZIP) · Offline verification guide ↗

+

HAND OFF EVIDENCE

Download the evidence.
Check every gate.

Recompute the original suite configuration, plan, five attempts, acceptance decisions and JUnit without executing a candidate.

Download suite evidence (ZIP) · Install and verify locally ↗

EVALARC 0.8+ / READ-ONLY
# Unzip the downloaded archive first.
 evalarc verify suite-evidence --json
 
@@ -198,8 +217,8 @@ 

A perfect score.
How much coverage remains?

--backend local --trust-local \ --output runs/support

Local mode runs with your user privileges. Use the documented Docker backend for candidate isolation.

-

What this evidence establishes

The bundled references pass and the 15 declared faulty implementations fail their intended checks. These are public development tasks and scripted policies. They do not establish frontier-model performance, coverage of arbitrary reward hacks, a human time horizon, or RL training gains. The browser replays saved reports; it does not execute submissions.

+

What this evidence establishes

The saved audits detect __FAULT_COUNT__ declared faults across __TASK_PACK_COUNT__ task packs. These are public development tasks and scripted policies. They do not establish frontier-model performance, coverage of arbitrary reward hacks, a human time horizon, or RL training gains. The browser replays saved reports; it does not execute submissions.

- + diff --git a/site/style.css b/site/style.css index 9b7be65..fb58ae4 100644 --- a/site/style.css +++ b/site/style.css @@ -43,3 +43,16 @@ button:disabled{opacity:.5;cursor:default} .judge-preview caption{text-align:left;color:var(--muted,#a8b5ab);margin-bottom:14px} .judge-preview th,.judge-preview td{text-align:left;padding:12px;border-bottom:1px solid #46534b} .judge-preview .unknown{color:#f0d48b;border-color:#f0d48b} + +/* Clear first-use choices, with readable targets in the embedded mobile view. */ +.start-path{display:grid;grid-template-columns:repeat(3,minmax(0,1fr));gap:16px;margin:0 0 30px} +.start-path>a,.feedback-card{min-width:0;padding:24px;border:1px solid var(--line);border-radius:8px;background:var(--panel)} +.start-path>a{display:flex;flex-direction:column;gap:14px;text-decoration:none} +.start-path>a:hover{border-color:var(--green)} +.start-path strong{font-size:18px;color:var(--ink)} +.start-path>a>span:last-child{font-size:14px;line-height:1.6;color:var(--muted)} +.feedback-card a{display:inline-flex;min-height:44px;align-items:center;color:var(--green)} +.tour{border:1px solid var(--line);border-radius:8px;padding:0 20px;margin:0 0 30px} +.tour summary{min-height:44px;color:var(--green)} +.tour video{display:block;width:100%;height:auto;max-height:80vh;margin:10px 0} +@media(max-width:750px){.start-path{grid-template-columns:minmax(0,1fr)}.start-path>a{padding:20px}}