Empirical local runtime evaluation: pure Rust, pure Zig, and a Rust control plane calling a Zig semantic-data kernel. This is not production CSL. Go is out of scope. Semantic equality comes first; no architecture recommendation has been made.
./ci/setup-local.sh
source .venv/bin/activate
./ci/validate-schemas.sh
./ci/validate-oracle.sh
./ci/smoke.sh
./ci/check-candidates.sh
./ci/abi-diagnostics.shRust/Cargo must be installed separately if unavailable. Scripts use the repository virtual environment when present. Missing toolchains are SKIP/NOT AVAILABLE, never candidate PASS. See TOOLCHAINS.md for the tested versions and the optional macOS 27 SDK workaround used for this run.
python harness/benchctl.py env
python harness/benchctl.py prepare --preset SMOKE
python harness/benchctl.py validate
python harness/benchctl.py build rust
python harness/benchctl.py build zig
python harness/benchctl.py build hybrid
python harness/benchctl.py conformance --candidate rust --candidate zig --candidate hybrid
python harness/benchctl.py run W1 --candidate rust --corpus SMOKE --repeat 10
python harness/benchctl.py compare W2 --corpus SMOKE --repeat 10
python harness/benchctl.py compare W1 --corpus SMOKE --repeat 10 --profile
python harness/benchctl.py compare W2 --corpus SMOKE --repeat 10 --profile
python harness/benchctl.py conformance --candidate rust --candidate zig --profile
python harness/benchctl.py w10 --repeat 10
python harness/benchctl.py w11
python harness/benchctl.py report gate1--corpus accepts SMOKE/S/M or an explicit fixture path. prepare --preset S
generates 100,000 entities/1,000,000 edges; M generates 1,000,000/10,000,000.
Generation streams arrays with an explicit seed; loading still materializes JSON
and large runs need adequate RAM. M is never generated by smoke CI.
All candidates implement JSON-only stdout for info, load --fixture PATH,
query --fixture PATH --query PATH, and
workload --id W1|W2 --corpus PATH --params '{"query": QueryIR}'.
The harness invokes binaries as processes; Python is only the oracle/orchestrator.
Hybrid also supports echo "CSL Rust + Zig", an ABI test rather than a CSL query.
After building and passing conformance, run the paired campaign on a quiet native toolchain/SDK environment:
python harness/campaign.py run --preset S --repeat 10 --seed 20260928 \
--output results/s-scale-nativeThe runner prepares schema-validated oracle references in a separate process,
alternates candidate order, records machine conditions for each invocation, and
rejects unsuitable controlled runs. Output directories must be new. For an
explicitly uncontrolled local scale check, add --exploratory; this classification
is retained regardless of whether correctness passes. See
ADR-0003 for phase scope, thresholds and limitations.
STATUS.md records actual completion and limits. The E0 contract is in
oracle/SEMANTICS.md, and ownership/W10 details are in
ABI.md. Raw evidence is under results/.
W1 loads/validates entities, looks up known/missing IDs, and scans by kind.
W2 builds outgoing/incoming indexes and runs one-hop, depth 2/4/8, mixed-relation,
cycle-heavy and high-fanout workloads. Both pure candidates now use typed records, hash-based entity and adjacency indexes,
hash sets and sorted traversal seeds/results. Representation pass 1 and remaining
allocator/query differences are documented in ADR-0002.
Original untuned records are retained in results/bootstrap-20260927/.
--profile adds separate native load, index, query and result nanosecond
measurements for the pure candidates, saved in results/phases/. The ordinary
workload result contract is unchanged. Profile envelopes must pass timing/schema
checks and full oracle equality. Hybrid retains its whole-fixture JSON FFI batch
and is excluded from profiled comparisons.
These remain cold-process smoke measurements. Load includes reading/parsing; index includes semantic validation; query includes evidence scanning and sorting; result includes encoding and hashing. Process elapsed also includes startup, parameter decoding, output and cleanup. RSS is whole-process, with different allocation lifetimes in Rust and Zig. This is not controlled S/M throughput and does not justify a language winner.
Every accepted W1/W2 iteration matches the full oracle result, including digest and snapshot. W10 hashes every output. RSS uses wait4 per-child peak memory where available; bytes/entity includes the entire process, not just entity storage. Missing metrics are null. Git commit is null for an unborn repository; a source hash and dirty flag identify these measurements. No native persistence, mutation, or cancellation API is claimed; JSON reload equality is covered.
The manifest uses repository-relative paths and covers evaluation sources,
lockfiles and the checked-in SMOKE fixture. It excludes itself, generated results,
other corpora, build caches, .DS_Store files, and the unrelated pre-existing web
page. Verify with ./ci/verify-manifest.sh; regenerate after source edits with
python3 ci/manifest.py --write. Evidence files can be regenerated without
invalidating the source manifest. No huge generated corpora need to be committed.