Repository navigation
CUDA: software-pipelined internal rounds, free-addend products, no early exit - #104
Conversation
…rly exit Kernel: - Internal rounds are software-pipelined: elements 1..11 are summed before the S-box result exists, element 0's output is produced first, and the next round's S-box is issued immediately so the other 11 products retire under it. The last round is peeled. - The unreduced 96-bit row sum (plus the round constant for element 0) rides in the 64-bit accumulators of the partial products, replacing a 4-op carry chain per element. - x^7 uses a depth-3 chain (x2; x3 and x4 in parallel; x7 = x4*x3). - No thread stops early. Up to eight candidate indices per launch are recorded with an atomic slot counter. Host: - Every recorded candidate is recomputed on the CPU; the lowest valid one is returned. A launch always evaluates its whole rectangle, so hash_count is the dispatched count and the resume-after-rejection path is gone. - The engine contract is restated accordingly.
Isolated on a Vast RTX 4090 (driver 570, ptxas 12.8 JIT, kernel-only harness, alternating with controls): the pipelined ordering alone is +1.5% over v4.2.0 with no local memory, the depth-3 S-box is neutral, and the mad.wide accumulator form of the diagonal product is -5% with a 16 B spill on this ptxas. The row sum goes back to an explicit 4-op carry add per element.
n13
left a comment
There was a problem hiding this comment.
Reviewer model: GPT 5.6 Sol
Verdict (advisory): Approve
No blocking findings. The no-early-exit result protocol is bounded on both sides, preserves exact dispatched-hash accounting, CPU-verifies every retained candidate, and selects the lowest valid retained nonce. The pipelined internal rounds preserve all 22 rounds, including the final zero round constant, and the benchmark record is internally consistent.
Non-blocking cleanup: crates/engine-cuda/src/kernels/mining.cu:218 still contains the unused mul128_add_wide experiment and its now-misleading “add costs nothing” comment even though commit 7730cb5 dropped that path after measuring the spill/regression. crates/engine-cuda/src/lib.rs:19 also retains two obsolete lines describing the old two-word result buffer. Removing both would keep the landed source aligned with the implementation.
Validation: inspected the complete c1cf0a3...e5d9a68 diff and affected host/kernel call paths; git diff --check, cargo fmt --all -- --check, cargo test --locked -p engine-cuda, strict focused Clippy, and benchmark JSON parsing passed. CUDA-dependent tests self-skipped locally because this arm64 macOS host has no NVIDIA runtime; the PR records three successful ten-test GPU runs on each of two RTX 4090 hosts. All exact-head GitHub checks are green.
Summary
Third round of CUDA kernel work, after #100, from comparing our kernel with the
quantus-cuda-minersource tree.Kernel
Host
hash_countis the dispatched count, exactly, and the resume-after-rejection path from CUDA: cheaper Goldilocks arithmetic, no third squeeze on the GPU, arch-targeted NVRTC #100 is gone. More than eight candidates in one launch (only at difficulties far below real mining) logs a warning; the extras are skipped.CudaEnginecontract is restated for the new record.Tried and dropped in this round: the row sum riding in
mad.wide.u32accumulators (the source tree's free-addend product) measured −5% on ptxas 12.8 and −3% on 13.2 with a 16 B spill on our kernel.Benchmark
Same-host protocol: released v4.2.0 first, then this binary, three alternating 30 s runs each, GPU only, default 32M batch. Ten
engine-cudaGPU tests passed three times on each host before timing. Record:docs/benchmarks/2026-09-11-vast-pr104.json.mining_main: 64 registers, 0 B local on both. The ptxas version (13.2 driver JIT vs NVRTC 12.8 cubin) made no measurable difference to v4.2.0, this kernel, or the rejected variant.Tests
cuda_found_batch_counts_every_dispatched_nonce: difficulty 1, 4M batch; the returned candidate is CPU-valid andhash_count == 4,000,000.cuda_search_rejects_prefix_equal_candidate_and_returns_lowest_valid: target equal to a golden hash over an 8-nonce range so every candidate is recorded; the golden nonce is rejected and the engine returns exactly the lowest CPU-valid nonce.