Skip to content

CUDA: software-pipelined internal rounds, free-addend products, no early exit - #104

Merged
n13 merged 3 commits into
mainfrom
n13/cuda-pipelined-internal
Sep 12, 2026
Merged

n13 merged 3 commits into
mainfrom
n13/cuda-pipelined-internal

Conversation

@n13

@n13 n13 commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Third round of CUDA kernel work, after #100, from comparing our kernel with the quantus-cuda-miner source tree.

Kernel

  • Software-pipelined internal rounds. Elements 1..11 are summed before the S-box result exists, element 0's output is produced first, and the next round's S-box is issued immediately so the other 11 diagonal products retire underneath it. The last round is peeled. This removed the last local-memory spill.
  • Depth-3 x^7. x2, then x3 = x2·x and x4 = x2² in parallel, x7 = x4·x3. Measured neutral; kept because it shortens the serial path at no cost.
  • No early exit. Candidate threads record their index in one of eight slots via an atomic counter and keep hashing. A launch always evaluates its whole rectangle.

Host

  • Every recorded candidate is recomputed on the CPU and the lowest valid one is returned. hash_count is the dispatched count, exactly, and the resume-after-rejection path from CUDA: cheaper Goldilocks arithmetic, no third squeeze on the GPU, arch-targeted NVRTC #100 is gone. More than eight candidates in one launch (only at difficulties far below real mining) logs a warning; the extras are skipped.
  • The CudaEngine contract is restated for the new record.

Tried and dropped in this round: the row sum riding in mad.wide.u32 accumulators (the source tree's free-addend product) measured −5% on ptxas 12.8 and −3% on 13.2 with a 16 B spill on our kernel.

Benchmark

Same-host protocol: released v4.2.0 first, then this binary, three alternating 30 s runs each, GPU only, default 32M batch. Ten engine-cuda GPU tests passed three times on each host before timing. Record: docs/benchmarks/2026-09-11-vast-pr104.json.

GPU Driver (ptxas) Power cap v4.2.0 This PR Gain
RTX 4090, Texas 570.86.16 (12.8) 450 W 862-870 MH/s 880-882 MH/s +2.2%
RTX 4090, Washington 595.58.03 (13.2) 420 W 876-886 MH/s 896-902 MH/s +2.1%

mining_main: 64 registers, 0 B local on both. The ptxas version (13.2 driver JIT vs NVRTC 12.8 cubin) made no measurable difference to v4.2.0, this kernel, or the rejected variant.

Tests

  • cuda_found_batch_counts_every_dispatched_nonce: difficulty 1, 4M batch; the returned candidate is CPU-valid and hash_count == 4,000,000.
  • cuda_search_rejects_prefix_equal_candidate_and_returns_lowest_valid: target equal to a golden hash over an 8-nonce range so every candidate is recorded; the golden nonce is rejected and the engine returns exactly the lowest CPU-valid nonce.
  • Golden vectors, target boundaries, nonce carries, bit-exact reducer models, CPU-verified solution: unchanged.

n13 added 3 commits September 11, 2026 15:47
…rly exit

Kernel:
- Internal rounds are software-pipelined: elements 1..11 are summed before
  the S-box result exists, element 0's output is produced first, and the next
  round's S-box is issued immediately so the other 11 products retire under
  it. The last round is peeled.
- The unreduced 96-bit row sum (plus the round constant for element 0) rides
  in the 64-bit accumulators of the partial products, replacing a 4-op carry
  chain per element.
- x^7 uses a depth-3 chain (x2; x3 and x4 in parallel; x7 = x4*x3).
- No thread stops early. Up to eight candidate indices per launch are
  recorded with an atomic slot counter.

Host:
- Every recorded candidate is recomputed on the CPU; the lowest valid one is
  returned. A launch always evaluates its whole rectangle, so hash_count is
  the dispatched count and the resume-after-rejection path is gone.
- The engine contract is restated accordingly.
Isolated on a Vast RTX 4090 (driver 570, ptxas 12.8 JIT, kernel-only
harness, alternating with controls): the pipelined ordering alone is
+1.5% over v4.2.0 with no local memory, the depth-3 S-box is neutral,
and the mad.wide accumulator form of the diagonal product is -5% with a
16 B spill on this ptxas. The row sum goes back to an explicit 4-op
carry add per element.
@n13
n13 marked this pull request as ready for review September 11, 2026 08:37
@n13 n13 added the bot-review label Sep 11, 2026

@n13 n13 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer model: GPT 5.6 Sol

Verdict (advisory): Approve

No blocking findings. The no-early-exit result protocol is bounded on both sides, preserves exact dispatched-hash accounting, CPU-verifies every retained candidate, and selects the lowest valid retained nonce. The pipelined internal rounds preserve all 22 rounds, including the final zero round constant, and the benchmark record is internally consistent.

Non-blocking cleanup: crates/engine-cuda/src/kernels/mining.cu:218 still contains the unused mul128_add_wide experiment and its now-misleading “add costs nothing” comment even though commit 7730cb5 dropped that path after measuring the spill/regression. crates/engine-cuda/src/lib.rs:19 also retains two obsolete lines describing the old two-word result buffer. Removing both would keep the landed source aligned with the implementation.

Validation: inspected the complete c1cf0a3...e5d9a68 diff and affected host/kernel call paths; git diff --check, cargo fmt --all -- --check, cargo test --locked -p engine-cuda, strict focused Clippy, and benchmark JSON parsing passed. CUDA-dependent tests self-skipped locally because this arm64 macOS host has no NVIDIA runtime; the PR records three successful ten-test GPU runs on each of two RTX 4090 hosts. All exact-head GitHub checks are green.

@n13 n13 removed the bot-review label Sep 11, 2026
@n13
n13 merged commit 0ba8ddf into main Sep 12, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant