Skip to content

CUDA: low-64 nonce specialization, compact permutation loop, tiny result record - #97

Merged
n13 merged 2 commits into
mainfrom
n13/cuda-40xx-2x
Sep 10, 2026
Merged

n13 merged 2 commits into
mainfrom
n13/cuda-40xx-2x

Conversation

@n13

@n13 n13 commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

Summary

Per-nonce work and hot-code footprint of the native CUDA engine (--cuda-gpu) are reduced, and the host stops spinning a CPU core per GPU.

Kernel

  • The host precomputes the Poseidon2 state up to the first external-round S-box for the 448 fixed nonce bits (pow_core::mining_prestate_low64). The kernel adds only the sparse linear contribution of the two byte-swapped low-64 nonce words per nonce, so the general 16-limb nonce increment and the repeated first linear layer are gone from the hot loop.
  • The two hot permutations run as a non-unrolled two-pass loop instead of two inlined copies.
  • The result record is a claim flag plus the 32-bit logical index. The host reconstructs the nonce, recomputes the full hash on the CPU, and enforces strict hash < target before anything is reported.
  • No per-iteration global result polling inside the nonce loop.

Host

  • Blocking CUDA synchronization (CU_CTX_SCHED_BLOCKING_SYNC) instead of busy-waiting.
  • Batches are split at the low-64 carry so a launch never wraps the low nonce word (unit test low64_batches_stop_before_carry).
  • --gpu-batch-size now defaults to 32M nonces with --cuda-gpu (Vulkan keeps 1M). With blocking sync, 1M-nonce launches lose throughput to host wake-up latency; 32M is about 50 ms per launch on a 4090.
  • Each CUDA search logs the GPU duty cycle since the previous search (GPU busy N% since previous search) so idle time on a live node is measurable.

Measurements (RTX 4090, 300 W cap, driver 580.173.02, NVRTC 12.8)

Configuration MH/s
main (v4.1.0), 67M-nonce launches, alternating samples 585-592
This branch, 67M-nonce launches 609-611
Both at 4M-nonce launches tied (~600)

Steady-state CPU use of the GPU worker thread drops from ~100% of a core to near zero.

Testing

  • The kernel and host changes were run through the engine-cuda real-GPU tests on the 4090 host above (golden vectors, target boundaries at hash-1/hash/hash+1, nonce carries at bits 32/64/128/224, Goldilocks reduction edge cases, CPU-verified solution).
  • cargo fmt, cargo clippy --workspace --all-features, and cargo test pass locally; CUDA tests self-skip without a GPU.
  • The CLI default change and duty-cycle log were not re-run on a GPU; they do not touch the kernel or launch path beyond a timer.

Not included

The static analysis notes under docs/research/ stay untracked; whether to publish them is a separate decision.

… result record

Kernel:
- Host precomputes the Poseidon2 state up to the first external-round S-box
  for the 448 fixed nonce bits; the kernel adds only the sparse linear
  contribution of the two byte-swapped low-64 words per nonce.
- The two hot permutations run as a non-unrolled two-pass loop instead of
  two inlined copies, shrinking hot code and spill traffic.
- The result record is a claim flag plus the 32-bit logical index; the host
  reconstructs the nonce and re-verifies the full hash before reporting.
- No per-iteration global result polling inside the nonce loop.

Host:
- Blocking CUDA synchronization instead of spinning a CPU core per GPU.
- Batches split at the low-64 carry so a launch never wraps the low word.
- CUDA-specific default batch size of 32M nonces (Vulkan keeps 1M); short
  launches lose throughput to wake-up latency with blocking sync.
- Log the GPU duty cycle between searches so idle time on a live node is
  visible.
@n13 n13 added the bot-review label Sep 10, 2026

@n13 n13 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verdict: changes requested

[P2] Recompute hash_count after removing result polling

The kernel no longer checks results[0] inside the per-thread loop, so after one thread claims a solution the other threads continue through their assigned nonces. The host still calculates completed work as total_threads * (winning_iteration + 1), which assumes the old lockstep early-exit behavior. With the new 32M default and 31 nonces per thread, a winner at j=0 can report about 1.05M hashes even though nearly the full 32M dispatch executes. That incorrect count feeds gpu_hashes_total, completion-rate logs, and MiningResult.hash_count.

Please make the returned count match the new execution model (for example, report the dispatched batch under scheduled-work semantics, or add exact accounting if exact completed hashes are required) and cover a multi-iteration found batch in a regression test.

Without in-loop result polling, threads finish their assigned nonces after
another thread claims a solution, so the old lockstep estimate undercounted
found batches. Drop the kernel-entry early exit as well so the launch always
executes the full rectangle, and report the dispatched count for found and
exhausted batches alike. Add a multi-iteration found-batch regression test.
@n13

n13 commented Sep 10, 2026

Copy link
Copy Markdown
Contributor Author

Fixed in 857c057: found batches now report the dispatched count. The kernel-entry early exit is gone too, so the full launch rectangle always executes and the dispatched count is exact rather than an upper bound. Regression test cuda_found_batch_counts_every_dispatched_nonce covers a 4M-nonce batch with 4 nonces per thread.

@n13 n13 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verdict: approve

Re-reviewed current head 857c057. The prior hash_count issue is resolved under scheduled-work semantics: found and exhausted batches now report the dispatched nonce count consistently, the kernel-entry early exit no longer makes that scheduled rectangle conditional, and cuda_found_batch_counts_every_dispatched_nonce covers the multi-iteration case.

I found no remaining blocking issues. All current GitHub checks pass. Local validation also passed: git diff --check and cargo test -p pow-core -p engine-cuda -p miner-cli --locked.

@n13
n13 merged commit 92f7771 into main Sep 10, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant