Repository navigation
Optimize CPU mining and Apple GPU hashing - #96
Merged
n13 merged 5 commits intoSep 10, 2026
Merged
Conversation
Reduce Apple u64 kernel overhead by combining reduction carry/borrow
corrections, using explicit u32 middle-limb carries, indexing round
constants directly, and delaying full result construction until needed.
Keep the existing hash, target comparison, and nonce coverage semantics.
Add offline arithmetic edge/reference checks, exhaustive small-range
coverage checks, and a trusted-kernel CPU parity/ABBA comparison harness.
Measured about 7.8% faster than v4.1.0 on one M4 Pro at identical dispatch
parameters; this is not a claim about other GPUs or live mining rewards.
Validation: stable full-workspace build/tests (58 passed), all-target
Clippy with warnings denied, rustfmt, Taplo, GPU reference examples,
and an offline CPU benchmark smoke test passed. Independent review of
the modular arithmetic and measurements on other Apple GPUs is welcome.
Code: crates/engine-gpu/src/kernels/mining_u64_apple.wgsl
Validation: crates/engine-gpu/examples/{arithmetic_edges,dispatch_coverage,kernel_compare}.rs
Replace carry/borrow comparisons with signed radix folding and explicit matrix doubling shifts. Bounds guarantee one final EPS correction. Extend arithmetic checks with independent four-limb edge cases and optional candidate source input. Validated full workspace build, 58 tests, all-target Clippy, rustfmt, Taplo, 4416 arithmetic cases, CPU parity and range coverage. Offline M4 Pro comparisons gain 2.3-3.2% over the previous commit and 9.8-10.8% over v4.1.0; no mainnet or cross-device performance claim. Code and validation: crates/engine-gpu/src/kernels/mining_u64_apple.wgsl; crates/engine-gpu/examples/arithmetic_edges.rs. PR: Quantus-Network#96
Add a job-local MiningHasher in pow-core and use it in FastCpuEngine. Refresh cached state on high-half nonce changes, return full hashes for candidates, and preserve strict target comparison, search order, counts and cancellation cadence. Keep hash_from_nonce independent for reference verification. Add cache/boundary/golden-vector/cancellation regressions and an offline CPU reference comparison example. M4 Pro reference comparisons improve 2.57-2.59x; full CLI one-worker ABBA mean improves 2.55x. This affects CPU workers only, not GPU throughput or proven mainnet rewards. Validation: full workspace locked build and 62 tests; all-target Clippy, rustfmt, Taplo; 100 GPU/CPU parity tasks passed. Code: crates/pow-core/src/lib.rs and crates/engine-cpu. PR: Quantus-Network#96
Move shared Apple kernel inputs into the Metal constant address space through WGSL uniform bindings. Pack words into vec4 arrays without changing byte order and pad dispatch config to 16 bytes. Allocate matching input usages in the host and adapt raw-kernel validation tools; other kernels retain storage inputs. M4 Pro same-parameter ABBA comparisons gain 0.94-1.52% over the previous kernel. Final comparisons against v4.1.0 gain 12.22% and 12.00% cumulatively. These are offline GPU kernel results, not CPU gains or mainnet revenue claims. Validated all three GPU kernel component/end-to-end suites, 100 CPU parity tasks, strict target boundaries, range coverage, full workspace build and 62 tests, Clippy, rustfmt and Taplo. Code: crates/engine-gpu. PR: Quantus-Network#96
Replace signed 64-bit intermediate reduction sums with explicit u32 limb carries/borrows, and fold matrix accumulators in radix 2^32. Preserve lazy residues, strict target comparison, rounds and dispatch. Document correction bounds in the shader. Expand the independent u128 arithmetic example to 65,920 cases, including accumulator carry edges. M4 Pro offline ABBA: final kernel +18.386% over the previous commit; +30.919–32.357% versus v4.1.0. GPU-only CLI averages 35.9475 to 42.9250 MH/s (+19.41%). Results are hardware-specific, not earnings. Validation: 62 workspace tests, locked builds, all-target Clippy, fmt, Taplo; 65,920 arithmetic cases, all three GPU kernel suites, 100 parity tasks and 30 dispatch coverage cases passed. Mainnet mining stays off. PR: Quantus-Network#96 Code: crates/engine-gpu/src/kernels/mining_u64_apple.wgsl
n13
approved these changes
Sep 10, 2026
n13
left a comment
Contributor
There was a problem hiding this comment.
Verdict: APPROVE
No blocking findings at 29c10ee.
I reviewed the CPU MiningHasher cache boundary and strict target comparison, the Metal uniform-buffer ABI and alignment changes, the u32 Goldilocks reduction/carry logic, round-constant indexing, early hash rejection, and the new validation helpers.
Local validation on the repository-pinned Rust 1.93.0 toolchain:
- cargo fmt --all -- --check
- taplo fmt --check
- cargo build --workspace --locked
- cargo clippy --workspace --all-targets -- -D warnings
- cargo test --workspace --locked: 62 passed
- GPU component/end-to-end suite: u32, generic u64, and Apple u64 all passed on Metal
- arithmetic_edges: 65,920 cases passed
- dispatch_coverage: 15 minimum-hash cases and 15 exhausted ranges passed
Non-blocking caveat: gpu_cpu_parity verified 100/100 jobs against CPU, then the process aborted during wgpu thread-local destruction. That helper is unchanged by this PR, so I did not treat the pre-existing teardown issue as a blocker.
Contributor
|
Very nice! |
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Optimize CPU nonce scanning and the Apple u64 GPU kernel. On one M4 Pro (48 GB), the CPU reference comparison measured 2.57–2.59× throughput, and a one-worker CLI comparison against the previous build measured 2.55×. The Apple GPU kernel separately measured about 31–32% higher throughput than v4.1.0. These are separate offline results, not additive whole-machine gains or mainnet earnings claims.
What changed
CPU (
pow-core,engine-cpu): add a job-localMiningHasherthat caches the header/high-nonce midstate and refreshes it when the high 256 bits change. Reject hashes whose high half exceeds the target before the final squeeze; compute the complete hash for potential solutions and preserve strict comparison. The independenthash_from_noncereference remains unchanged. Range order, candidate bytes, cancellation cadence, and counts are retained.Apple GPU (
engine-gpu):Dispatch counts, round counts, constants, nonce ranges, difficulty rules, CUDA kernel, and generic shader remain unchanged. Apple input binding types change; the other kernels retain storage inputs and ignore dispatch padding.
For reduction, let B=2^32 and P=B²−B+1. For limbs w0..w3, the input is congruent to
(w0−w2−w3) + (w1+w2)B. Setlow=w0−w2−w3andhigh=w1+w2+floor(low/B). Assemble their low 32-bit limbs and addfloor(high/B)*(B−1). Since −1≤high≤2B−2, this correction is −1, 0, or 1 times EPS. A positive correction has result≤B²−B−1, so cannot overflow; a negative correction has a high limb of B−1, so cannot underflow. The implementation obtains the low borrow from two u32 subtractions and the high correction from carry minus borrow. For accumulator lo+c·2^64, subtract c from the low limb, add c−borrow to the high limb, and fold one high carry; c−borrow is nonnegative even when c=0. The overflow case has result≤B²−B−1 before adding EPS. Lazy residues remain permitted and subsequent canonicalization is unchanged.Validation
Locally on macOS / M4 Pro with Rust stable 1.95.0:
Each Apple GPU performance run uses 256 threads per workgroup and 1,000,000 nonces per batch. After warmup, baseline/candidate/candidate/baseline is interleaved 100 times, totaling 200 million nonces per side. Timing covers submission through readback at a hard target.
The final u32 carry/borrow changes also measured 36.6983→43.4457 MH/s (+18.386%) against the previous uniform-input shader. A separate complete CLI comparison against commit 560f65b used zero CPU workers, one GPU, 1,000,000 nonces per batch, and two ABBA rounds of five-second runs: mean 35.9475→42.9250 MH/s (+19.41%). These incremental comparisons have a different baseline from the official-kernel table.
CPU comparison uses 20,000 nonces per batch, a hard target, a warmup round, then 12 ABBA rounds (480,000 nonces per side) with changing headers. The baseline calls the independent full reference hash; the candidate runs the CPU engine's range search, including cache setup. Three confirmations measured 0.2674→0.6910, 0.2655→0.6882, and 0.2695→0.6924 MH/s.
The complete CLI was also compared in two ABBA rounds with one CPU worker, zero GPU workers and five seconds per run: mean 0.2650→0.6768 MH/s (+155.4%). This gain does not apply when CPU workers are disabled.
Reproduce from this checkout (GPU examples require suitable hardware; arithmetic/comparison require SHADER_INT64):
cargo +stable run --release --locked -p engine-cpu --example compare_reference git show 220712c8d73c539474c9c8b955972b2a8f559dd6:crates/engine-gpu/src/kernels/mining_u64_apple.wgsl > /tmp/quantus-baseline.wgsl cargo +stable build --release --locked -p engine-gpu --examples ./target/release/examples/arithmetic_edges ./target/release/examples/gpu_cpu_parity 100 ./target/release/examples/dispatch_coverage ./target/release/examples/kernel_compare /tmp/quantus-baseline.wgsl crates/engine-gpu/src/kernels/mining_u64_apple.wgslRisks and mitigations
The Apple input ABI now uses uniform bindings 1–4 and 16-byte dispatch config; external raw-shader callers must use matching resources. Word order is unchanged and other shaders retain storage bindings.
Arithmetic changes are consensus-sensitive; CPU reference, strict target boundaries, and exhaustive small-range checks cover the changed behavior, but independent review is welcome. The comparison example disables shader runtime checks like production and must only receive trusted compatible shaders.
Performance is validated only on this M4 Pro and may vary with GPU/compiler, target, thermals, and workload. No sustained power or mainnet mining improvement is claimed.
Follow-ups
Independent CPU/Apple GPU measurements and review of cache invalidation and limb-reduction bounds would be valuable before merging.