Skip to content

docs: publish five-language Compass vs Graphify evaluation - #354

Open
forhappy wants to merge 1 commit into
mainfrom
codex/main-vs-graphify-evaluation
Open

forhappy wants to merge 1 commit into
mainfrom
codex/main-vs-graphify-evaluation

Conversation

@forhappy

@forhappy forhappy commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Summary

Publish the October 2 comparison of Compass main b7b7fa81 (0.4.1) and Graphify 0.9.74 on Cobra, Flask, Gson, Zod and the Axum package. The report separates source-verified graph facts, query recall, interface capabilities, indexing/query cost, memory measurement and rebuild consistency across Compass JSON/SQLite and Graphify base/native-Leiden configurations.

The evidence includes all 120 builds, 1,020 query workflows, source oracles and audits, 2,412 stdout/stderr captures deduplicated into 389 streams, package identities, archived method snapshots and publication hashes. Local paths use documented placeholders; the original archive remains unchanged. Graphs, stores, caches, binaries and environments remain external.

Start with REPORT.md. It preserves competitor wins, Compass failures and publication omissions. The selected development corpora, busy shared host and post-capture directional audit limit the conclusions; this does not establish population precision or a controlled performance ranking.

Motivation

Provide reviewable evidence for the engineering team's tooling choice, including the underlying answers and failed checks rather than only aggregate scores. Link the capture from the agent-query guide and PERFORMANCE.md without replacing the qualification baseline.

Verification

python3 -m unittest benchmarks.agent_query.tests.test_runner benchmarks.agent_query.tests.test_edge_audit benchmarks.performance.tests.test_process
  45 tests passed
sh scripts/check_product_boundary.sh
  passed
git diff --cached --check
  passed
Publication validation
  37 published file hashes; all normalized JSON records equal original observations
  2,412 original capture hashes and 389 normalized content hashes verified
  archived Python snapshot syntax and changed local Markdown links checked
  local personal paths and credential prefixes absent from the published bundle

The measurement used a successful pinned Rust release build with cargo build -p compass-cli --bin compass --release --locked and a dedicated external Cargo target. Original completion verification checked all 120 frozen graphs, 20 query graphs, 1,020 graph-pinned observations and corpus/tool identities. The full Rust baseline and eight-repository promotable performance qualification were not run for this documentation/evidence publication; no runtime source changed, and diagnostic timings are not promoted to a release baseline.

Compatibility and documentation

No CLI, graph, storage, integration or runtime-dependency change. PERFORMANCE.md and the agent-query guide link the dated report. The evidence is an archival comparison bundle with documented path normalization and separate original/published hashes; archived scripts are text snapshots, not installed commands.

Checklist

  • The change is focused and excludes unrelated formatting or generated files
  • Tests cover changed behavior, or this pull request changes documentation only
  • User-facing commands, flags, limits, and examples are documented
  • Compatibility or migration effects are described
  • No credentials, private source code, or sensitive report details are included
  • I agree to license my contribution under MIT OR Apache-2.0
  • I followed the Compass code of conduct

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant