Skip to content

Optimize CPU mining and Apple GPU hashing - #96

Merged
n13 merged 5 commits into
Quantus-Network:mainfrom
userInner:perf/apple-goldilocks-reduction
Sep 10, 2026
Merged

n13 merged 5 commits into
Quantus-Network:mainfrom
userInner:perf/apple-goldilocks-reduction

Conversation

@userInner

@userInner userInner commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Optimize CPU nonce scanning and the Apple u64 GPU kernel. On one M4 Pro (48 GB), the CPU reference comparison measured 2.57–2.59× throughput, and a one-worker CLI comparison against the previous build measured 2.55×. The Apple GPU kernel separately measured about 31–32% higher throughput than v4.1.0. These are separate offline results, not additive whole-machine gains or mainnet earnings claims.

What changed

CPU (pow-core, engine-cpu): add a job-local MiningHasher that caches the header/high-nonce midstate and refreshes it when the high 256 bits change. Reject hashes whose high half exceeds the target before the final squeeze; compute the complete hash for potential solutions and preserve strict comparison. The independent hash_from_nonce reference remains unchanged. Range order, candidate bytes, cancellation cadence, and counts are retained.

Apple GPU (engine-gpu):

  • Use uniform inputs for the Apple kernel so Metal can access shared inputs through the constant address space. Pack scalar words in vec4 arrays to retain byte order; pad dispatch config to 16 bytes. The runtime allocates matching uniform/storage inputs, and benchmark/end-to-end helpers support both layouts.
  • Fold Goldilocks reduction directly from four 32-bit limbs, using u32 carries and borrows instead of signed 64-bit intermediate sums. Fold matrix accumulator carries with the same radix-2^32 approach.
  • Accumulate multiplication/squaring middle limbs in u32 with explicit carry bits.
  • Express two external-matrix doublings as shifts with explicit carry extraction.
  • Pass an external round-constant index instead of a 12-element array.
  • Reject on the first hash word before materializing the full result. The original kernel already skips the final permutation for rejected high halves; this retains that optimization and strict full-hash comparison.
  • Add offline arithmetic, dispatch coverage, and CPU-parity/ABBA comparison examples.

Dispatch counts, round counts, constants, nonce ranges, difficulty rules, CUDA kernel, and generic shader remain unchanged. Apple input binding types change; the other kernels retain storage inputs and ignore dispatch padding.

For reduction, let B=2^32 and P=B²−B+1. For limbs w0..w3, the input is congruent to (w0−w2−w3) + (w1+w2)B. Set low=w0−w2−w3 and high=w1+w2+floor(low/B). Assemble their low 32-bit limbs and add floor(high/B)*(B−1). Since −1≤high≤2B−2, this correction is −1, 0, or 1 times EPS. A positive correction has result≤B²−B−1, so cannot overflow; a negative correction has a high limb of B−1, so cannot underflow. The implementation obtains the low borrow from two u32 subtractions and the high correction from carry minus borrow. For accumulator lo+c·2^64, subtract c from the low limb, add c−borrow to the high limb, and fold one high carry; c−borrow is nonnegative even when c=0. The overflow case has result≤B²−B−1 before adding EPS. Lazy residues remain permitted and subsequent canonicalization is unchanged.

Validation

Locally on macOS / M4 Pro with Rust stable 1.95.0:

  • New CPU regression coverage: cold/hot cache strict boundaries, random and non-monotonic nonces, high-half carry, maximum nonce, first qualifying candidate and count, and cancellation cadence.
  • Full-workspace locked build and tests: 62 tests passed, zero failures.
  • All-target full-workspace Clippy with warnings denied, rustfmt, and Taplo passed.
  • 65,920 arithmetic cases passed against Rust u128: 65,536 seeded random pairs, 64 original edge pairs, 256 independent four-limb edge combinations, and 64 accumulator edge cases; checks cover wide multiplication, modular multiplication, square, arbitrary 128-bit reduction, and accumulator folding against an independent u128 reference.
  • Full GPU component and end-to-end suites passed for u32, generic u64, and Apple u64 kernels.
  • GPU/CPU hash parity: 100 tasks passed.
  • Both kernels passed 64 deterministic samples at hash−1, hash, and hash+1.
  • Dispatch coverage: 15 exhaustive minimum-hash cases and 15 exhausted ranges passed, including tails and high-256-bit nonce carry.

Each Apple GPU performance run uses 256 threads per workgroup and 1,000,000 nonces per batch. After warmup, baseline/candidate/candidate/baseline is interleaved 100 times, totaling 200 million nonces per side. Timing covers submission through readback at a hard target.

Run Official MH/s Candidate MH/s Improvement
Final source 1 32.7751 43.3800 32.357%
Final source 2 30.4389 39.8502 30.919%

The final u32 carry/borrow changes also measured 36.6983→43.4457 MH/s (+18.386%) against the previous uniform-input shader. A separate complete CLI comparison against commit 560f65b used zero CPU workers, one GPU, 1,000,000 nonces per batch, and two ABBA rounds of five-second runs: mean 35.9475→42.9250 MH/s (+19.41%). These incremental comparisons have a different baseline from the official-kernel table.

CPU comparison uses 20,000 nonces per batch, a hard target, a warmup round, then 12 ABBA rounds (480,000 nonces per side) with changing headers. The baseline calls the independent full reference hash; the candidate runs the CPU engine's range search, including cache setup. Three confirmations measured 0.2674→0.6910, 0.2655→0.6882, and 0.2695→0.6924 MH/s.

The complete CLI was also compared in two ABBA rounds with one CPU worker, zero GPU workers and five seconds per run: mean 0.2650→0.6768 MH/s (+155.4%). This gain does not apply when CPU workers are disabled.

Reproduce from this checkout (GPU examples require suitable hardware; arithmetic/comparison require SHADER_INT64):

cargo +stable run --release --locked -p engine-cpu --example compare_reference
git show 220712c8d73c539474c9c8b955972b2a8f559dd6:crates/engine-gpu/src/kernels/mining_u64_apple.wgsl > /tmp/quantus-baseline.wgsl
cargo +stable build --release --locked -p engine-gpu --examples
./target/release/examples/arithmetic_edges
./target/release/examples/gpu_cpu_parity 100
./target/release/examples/dispatch_coverage
./target/release/examples/kernel_compare /tmp/quantus-baseline.wgsl crates/engine-gpu/src/kernels/mining_u64_apple.wgsl

Risks and mitigations

The Apple input ABI now uses uniform bindings 1–4 and 16-byte dispatch config; external raw-shader callers must use matching resources. Word order is unchanged and other shaders retain storage bindings.

Arithmetic changes are consensus-sensitive; CPU reference, strict target boundaries, and exhaustive small-range checks cover the changed behavior, but independent review is welcome. The comparison example disables shader runtime checks like production and must only receive trusted compatible shaders.

Performance is validated only on this M4 Pro and may vary with GPU/compiler, target, thermals, and workload. No sustained power or mainnet mining improvement is claimed.

Follow-ups

Independent CPU/Apple GPU measurements and review of cache invalidation and limb-reduction bounds would be valuable before merging.

Reduce Apple u64 kernel overhead by combining reduction carry/borrow
corrections, using explicit u32 middle-limb carries, indexing round
constants directly, and delaying full result construction until needed.
Keep the existing hash, target comparison, and nonce coverage semantics.

Add offline arithmetic edge/reference checks, exhaustive small-range
coverage checks, and a trusted-kernel CPU parity/ABBA comparison harness.
Measured about 7.8% faster than v4.1.0 on one M4 Pro at identical dispatch
parameters; this is not a claim about other GPUs or live mining rewards.

Validation: stable full-workspace build/tests (58 passed), all-target
Clippy with warnings denied, rustfmt, Taplo, GPU reference examples,
and an offline CPU benchmark smoke test passed. Independent review of
the modular arithmetic and measurements on other Apple GPUs is welcome.

Code: crates/engine-gpu/src/kernels/mining_u64_apple.wgsl
Validation: crates/engine-gpu/examples/{arithmetic_edges,dispatch_coverage,kernel_compare}.rs
Replace carry/borrow comparisons with signed radix folding and explicit matrix doubling shifts. Bounds guarantee one final EPS correction. Extend arithmetic checks with independent four-limb edge cases and optional candidate source input.

Validated full workspace build, 58 tests, all-target Clippy, rustfmt, Taplo, 4416 arithmetic cases, CPU parity and range coverage. Offline M4 Pro comparisons gain 2.3-3.2% over the previous commit and 9.8-10.8% over v4.1.0; no mainnet or cross-device performance claim.

Code and validation: crates/engine-gpu/src/kernels/mining_u64_apple.wgsl; crates/engine-gpu/examples/arithmetic_edges.rs. PR: Quantus-Network#96
Add a job-local MiningHasher in pow-core and use it in FastCpuEngine. Refresh cached state on high-half nonce changes, return full hashes for candidates, and preserve strict target comparison, search order, counts and cancellation cadence. Keep hash_from_nonce independent for reference verification.

Add cache/boundary/golden-vector/cancellation regressions and an offline CPU reference comparison example. M4 Pro reference comparisons improve 2.57-2.59x; full CLI one-worker ABBA mean improves 2.55x. This affects CPU workers only, not GPU throughput or proven mainnet rewards.

Validation: full workspace locked build and 62 tests; all-target Clippy, rustfmt, Taplo; 100 GPU/CPU parity tasks passed. Code: crates/pow-core/src/lib.rs and crates/engine-cpu. PR: Quantus-Network#96
@userInner userInner changed the title Optimize Apple Goldilocks reduction and hash rejection Optimize CPU mining and Apple GPU hashing Sep 9, 2026
Move shared Apple kernel inputs into the Metal constant address space through WGSL uniform bindings. Pack words into vec4 arrays without changing byte order and pad dispatch config to 16 bytes. Allocate matching input usages in the host and adapt raw-kernel validation tools; other kernels retain storage inputs.

M4 Pro same-parameter ABBA comparisons gain 0.94-1.52% over the previous kernel. Final comparisons against v4.1.0 gain 12.22% and 12.00% cumulatively. These are offline GPU kernel results, not CPU gains or mainnet revenue claims.

Validated all three GPU kernel component/end-to-end suites, 100 CPU parity tasks, strict target boundaries, range coverage, full workspace build and 62 tests, Clippy, rustfmt and Taplo. Code: crates/engine-gpu. PR: Quantus-Network#96
Replace signed 64-bit intermediate reduction sums with explicit u32
limb carries/borrows, and fold matrix accumulators in radix 2^32.
Preserve lazy residues, strict target comparison, rounds and dispatch.
Document correction bounds in the shader. Expand the independent u128
arithmetic example to 65,920 cases, including accumulator carry edges.

M4 Pro offline ABBA: final kernel +18.386% over the previous commit;
+30.919–32.357% versus v4.1.0. GPU-only CLI averages 35.9475 to
42.9250 MH/s (+19.41%). Results are hardware-specific, not earnings.

Validation: 62 workspace tests, locked builds, all-target Clippy, fmt,
Taplo; 65,920 arithmetic cases, all three GPU kernel suites, 100 parity
tasks and 30 dispatch coverage cases passed. Mainnet mining stays off.

PR: Quantus-Network#96
Code: crates/engine-gpu/src/kernels/mining_u64_apple.wgsl

@n13 n13 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verdict: APPROVE

No blocking findings at 29c10ee.

I reviewed the CPU MiningHasher cache boundary and strict target comparison, the Metal uniform-buffer ABI and alignment changes, the u32 Goldilocks reduction/carry logic, round-constant indexing, early hash rejection, and the new validation helpers.

Local validation on the repository-pinned Rust 1.93.0 toolchain:

  • cargo fmt --all -- --check
  • taplo fmt --check
  • cargo build --workspace --locked
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo test --workspace --locked: 62 passed
  • GPU component/end-to-end suite: u32, generic u64, and Apple u64 all passed on Metal
  • arithmetic_edges: 65,920 cases passed
  • dispatch_coverage: 15 minimum-hash cases and 15 exhausted ranges passed

Non-blocking caveat: gpu_cpu_parity verified 100/100 jobs against CPU, then the process aborted during wgpu thread-local destruction. That helper is unchanged by this PR, so I did not treat the pre-existing teardown issue as a blocker.

@n13

n13 commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

Very nice!

@n13
n13 merged commit 66e9692 into Quantus-Network:main Sep 10, 2026
6 checks passed
@n13 n13 mentioned this pull request Sep 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants