Skip to content

Commit 1c95704

Browse files
author
Sim Pi Agent
committed
Pi Babysit: address PR #8064 feedback
1 parent 71a3717 commit 1c95704

1 file changed

Lines changed: 1 addition & 9 deletions

File tree

  • apps/sim/content/library/reproducible-ai-coding-agent-benchmark

‎apps/sim/content/library/reproducible-ai-coding-agent-benchmark/index.mdx‎

Lines changed: 1 addition & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -204,15 +204,7 @@ Cost records must identify what was measured. If a vendor exposes model token ch
204204

205205
The Sim Reproducible Coding-Agent Benchmark must provide public fixtures, raw results, patches, logs, evaluators, and analysis scripts so researchers can inspect or rerun every score.
206206

207-
The final publication should link to these artifacts:
208-
209-
- Benchmark repository: [BENCHMARK_REPOSITORY_URL]
210-
- Task manifest: [TASK_MANIFEST_URL]
211-
- Raw results in CSV format: [RAW_RESULTS_CSV_URL]
212-
- Raw results in JSONL format: [RAW_RESULTS_JSONL_URL]
213-
- Agent patches and logs: [RUN_ARTIFACTS_URL]
214-
- Scoring and analysis script: [ANALYSIS_SCRIPT_URL]
215-
- Archived benchmark release: [VERSIONED_RELEASE_URL]
207+
These artifacts are not yet available because the public benchmark run is pending. After the run, this section will link to the benchmark repository, task manifest, CSV and JSONL results, agent patches and logs, scoring and analysis script, and versioned release. No download links will be published before the corresponding artifacts are public and auditable.
216208

217209
The machine-readable result schema should include:
218210

0 commit comments

Comments
 (0)