Repository navigation
CUDA: low-64 nonce specialization, compact permutation loop, tiny result record - #97
Conversation
… result record Kernel: - Host precomputes the Poseidon2 state up to the first external-round S-box for the 448 fixed nonce bits; the kernel adds only the sparse linear contribution of the two byte-swapped low-64 words per nonce. - The two hot permutations run as a non-unrolled two-pass loop instead of two inlined copies, shrinking hot code and spill traffic. - The result record is a claim flag plus the 32-bit logical index; the host reconstructs the nonce and re-verifies the full hash before reporting. - No per-iteration global result polling inside the nonce loop. Host: - Blocking CUDA synchronization instead of spinning a CPU core per GPU. - Batches split at the low-64 carry so a launch never wraps the low word. - CUDA-specific default batch size of 32M nonces (Vulkan keeps 1M); short launches lose throughput to wake-up latency with blocking sync. - Log the GPU duty cycle between searches so idle time on a live node is visible.
n13
left a comment
There was a problem hiding this comment.
Verdict: changes requested
[P2] Recompute hash_count after removing result polling
The kernel no longer checks results[0] inside the per-thread loop, so after one thread claims a solution the other threads continue through their assigned nonces. The host still calculates completed work as total_threads * (winning_iteration + 1), which assumes the old lockstep early-exit behavior. With the new 32M default and 31 nonces per thread, a winner at j=0 can report about 1.05M hashes even though nearly the full 32M dispatch executes. That incorrect count feeds gpu_hashes_total, completion-rate logs, and MiningResult.hash_count.
Please make the returned count match the new execution model (for example, report the dispatched batch under scheduled-work semantics, or add exact accounting if exact completed hashes are required) and cover a multi-iteration found batch in a regression test.
Without in-loop result polling, threads finish their assigned nonces after another thread claims a solution, so the old lockstep estimate undercounted found batches. Drop the kernel-entry early exit as well so the launch always executes the full rectangle, and report the dispatched count for found and exhausted batches alike. Add a multi-iteration found-batch regression test.
|
Fixed in 857c057: found batches now report the dispatched count. The kernel-entry early exit is gone too, so the full launch rectangle always executes and the dispatched count is exact rather than an upper bound. Regression test |
n13
left a comment
There was a problem hiding this comment.
Verdict: approve
Re-reviewed current head 857c057. The prior hash_count issue is resolved under scheduled-work semantics: found and exhausted batches now report the dispatched nonce count consistently, the kernel-entry early exit no longer makes that scheduled rectangle conditional, and cuda_found_batch_counts_every_dispatched_nonce covers the multi-iteration case.
I found no remaining blocking issues. All current GitHub checks pass. Local validation also passed: git diff --check and cargo test -p pow-core -p engine-cuda -p miner-cli --locked.
Summary
Per-nonce work and hot-code footprint of the native CUDA engine (
--cuda-gpu) are reduced, and the host stops spinning a CPU core per GPU.Kernel
pow_core::mining_prestate_low64). The kernel adds only the sparse linear contribution of the two byte-swapped low-64 nonce words per nonce, so the general 16-limb nonce increment and the repeated first linear layer are gone from the hot loop.hash < targetbefore anything is reported.Host
CU_CTX_SCHED_BLOCKING_SYNC) instead of busy-waiting.low64_batches_stop_before_carry).--gpu-batch-sizenow defaults to 32M nonces with--cuda-gpu(Vulkan keeps 1M). With blocking sync, 1M-nonce launches lose throughput to host wake-up latency; 32M is about 50 ms per launch on a 4090.GPU busy N% since previous search) so idle time on a live node is measurable.Measurements (RTX 4090, 300 W cap, driver 580.173.02, NVRTC 12.8)
main(v4.1.0), 67M-nonce launches, alternating samplesSteady-state CPU use of the GPU worker thread drops from ~100% of a core to near zero.
Testing
engine-cudareal-GPU tests on the 4090 host above (golden vectors, target boundaries athash-1/hash/hash+1, nonce carries at bits 32/64/128/224, Goldilocks reduction edge cases, CPU-verified solution).cargo fmt,cargo clippy --workspace --all-features, andcargo testpass locally; CUDA tests self-skip without a GPU.Not included
The static analysis notes under
docs/research/stay untracked; whether to publish them is a separate decision.