Skip to content

feat(fp16): accumulation-hardness PoUW for A100 (header-bound ZK consensus certificate + sm_80 miner pipeline) - #1

Open
jonpry wants to merge 8 commits into
fp8-schemefrom
fp16-scheme
Open

jonpry wants to merge 8 commits into
fp8-schemefrom
fp16-scheme

Conversation

@jonpry

@jonpry jonpry commented Oct 3, 2026 •

Copy link
Copy Markdown

FP16 accumulation-hardness proof-of-useful-work (A100)

An alternative Pearl PoUW instantiation whose unit of work is FP16 matrix multiplication on NVIDIA A100 (GA100, sm_80) tensor cores. It extends the FP8 certificate scheme rather than replacing it: commitments, seed chain, state window, ticket, target, and the ZK split are shared, and FP8 remains available unchanged.

Core idea. Hardness comes from the nonlinearity of the device's own accumulation, not a coarse rounding grid. The A100 tensor core truncates each product onto a per-group alignment grid before summing (groups of 8, 24-bit window, round-toward-zero per group). That per-product truncation does not commute with the reduction, so it cannot be expressed as a matrix multiplication — the same obstruction that makes the integer transcript scheme hard, applied for free on every MAC. This lets recovered products carry FP16-level accuracy instead of FP8-residual accuracy. BF16 is shown to be too close to linear and is not admitted.

The consensus path (read first)

The FP16 (A100) consensus certificate is the header-bound recursive ZK certificate (wire.CertificateV5 → verify_fp16_zk_cert_ffi → verify_wrapped_proof_with_headers): a constant-size wrapped plonky2 proof whose public inputs — opening keys, noise seeds, jackpot key, operand roots, geometry — are re-derived from the block header + authenticated ancestor and pinned by equality before the proof is verified and difficulty checked, against an embedded per-profile verifier cache (no circuit build at verify time). The earlier plaintext certificate is retired as a consensus path; Fp16PlainProof now serves only as the prover's operand-opener bundle (the PlainFp16 wire name is historical). Everything is additive: V1–V4 (FP8 and earlier) are byte-for-byte untouched.

What's included

  • Whitepaper & design (docs/fp16_scheme/) — LaTeX + built PDF in the style of the FP8 spec, a STARK feasibility/soundness note, and the ZK-binding design.
  • Scheme core + verifier (zk-pow/src/api/fp16/) — FP16 (E5M10) dtype, the bit-exact A100 accumulation model, seed-derived low-rank noise (B-then-A keyed-BLAKE3 chain), FP16 noisy quantization, the "unpredictable accumulation steps" jackpot policy (breakpoint density ≥ 0.30, certified-work ratio ≥ 1.2), and a keyed Merkle commitment over raw FP16 rows.
  • FP16 STARK (zk-pow/src/circuit/fp16/) — a matmul_a100 STARK, policy STARK, forked BLAKE3 egress AIR, noise/noisy-quant/row-scale/xor-fold STARKs with CTL-bound decode/shift/width columns, a tight per-group census, limb-range soundness splits, and a batched driver producing a constant-size, verifiable, tamper-rejecting recursive proof. All datapath branches (normal/zero/subnormal) are constrained; AIRs are degree ≤ 3.
  • Bit-exact sm_80 miner pipeline (miner/pearl-gemm/) — raw-CUDA GA100 kernels for the full datapath (mma.sync.m16n8k16.f32.f16 GEMM, noise lines, noisy-quant, policy, commitment), the torch-free tile pipeline, the full-matrix lottery search with atomic hit latch, the host seed-chain/search driver, and the sm_80 jackpot hit-signal reset — each with a bit-exact CPU/Python reference and tests.
  • V5 consensus certificate + node verification — the CertificateV5 wire format and node-side verifyCertificateV5, the Go FFI bindings, base-aware tile opening, and the Python certificate assembly (operand-commitment opener + proof builder).
  • Header-bound ZK consensus certificate — makes the header-bound recursive proof the V5 consensus path (retiring plaintext), with the embedded verifier cache, the zk/zk_cert API, a wrapper-legal consensus envelope, and an end-to-end prove/verify fixture.
  • Integration — gateway submit-direct V5 path, the vLLM miner FP16 layer + mining dispatch, miner-base scheme admission/dispatch glue, and the node mining/solve switches that arm V5.
  • Silicon validation harness (docs/fp16_scheme/validation/) — a reproducible A100 HMMA capture / calibration / rounding-mode probe and an attack-cost benchmark, run directly on real tensor cores (no software model in the loop), plus an on-GPU winning-tile → ZK-consensus example and test.

Validation

  • api::fp16 + circuit::fp16: full suite green, including the end-to-end batched + wrapped ZK proof (verify + tamper-reject) and the subnormal-branch constraints.
  • circuit::fp8: unchanged (no regression).
  • Go (zkpow): CertificateV5 routes + verifies through VerifyCertificate (accept when the embedded cache covers the fixture / tamper-reject / commitment-mismatch / wire round-trip / version-aware size cap / activation gate).
  • Miner kernels: bit-exact vs the accumulation oracle across shapes/scales + carry-in + determinism, validated on real GA100 silicon.
  • Silicon harness: model-vs-device bit-exactness confirmed on real sm_80 HMMA; attack-cost benchmark measures hardness above the whitepaper floor. (The harness runs on GPU hardware; CI compares the model to the committed capture vectors.)

Gating / activation

V5 is deliberately inert for consensus and does not activate on any network as-is: IsCertVersionAllowed(V5) is false, Fp16ForkHeight is unset on every network, and CPU solve errors on V5. Decode/route/verify are wired but blocked at header-sanity; activation requires flipping the allow-list and setting a fork height together (a deliberate hard-fork decision).

Remaining before activation (documented, out of scope here)

  • Embed the full per-profile FP16 verifier cache. The per-shape wrapper loads its circuit from an embedded cache keyed by degree profile; the dev/CI build embeds only a sample (smallest-profile) bootstrap blob, so fp16_cache.bin must be built over the full consensus-legal envelope (or the universal wrapper below must land) before V5 is allowed. Fail-closed until then: a legal geometry whose profile is absent is rejected.
  • Universal (single-circuit) FRI wrapper — removes the per-shape cache; blocked on a shared starky limitation (efficiency, not soundness).
  • Consensus activation of V5 — the coupled allow-list + fork-height gate above.

Note on the hardness assumption

The scheme's security rests on the software accumulation model being bit-exact to A100 silicon. The committed validation harness captures from real HMMA independently of the model and confirms the match; this independence is exercised on GPU hardware, not in CI (CI compares the model to the committed capture vectors).

@jonpry
jonpry changed the base branch from master to fp8-scheme October 3, 2026 05:11
@jonpry
jonpry force-pushed the fp16-scheme branch 3 times, most recently from 39756f8 to 81a230a Compare October 4, 2026 21:13
@jonpry jonpry changed the title feat(fp16): accumulation-hardness PoUW for A100 (whitepaper, plaintext scheme, ZK proof, V5 certificate, SM80 miner kernel) feat(fp16): accumulation-hardness PoUW for A100 (header-bound ZK consensus certificate + sm_80 miner pipeline) Oct 4, 2026
jonpry added 8 commits October 4, 2026 14:42
Whitepaper, STARK-feasibility and ZK-binding design notes for the A100
FP16 accumulation-hardness proof-of-useful-work scheme.
zk-pow API for the A100 scheme: dtype, params, noise, quantization,
policy, operand commitment, accumulation and the plaintext proof/verify
datapath, with A100 reference dot-product vectors.
Plonky2 circuits proving the scheme datapath: blake3 commitment, noise,
noisy-quant FMA, row-scale, policy and xor-fold STARKs, the A100 matmul
STARK, cross-table lookups, driver and recursive wrapper.
pearl-gemm sm_80 kernels and host code for the full miner datapath:
FP16 GEMM, noise lines, noisy-quant, policy and commitment, the
end-to-end tile pipeline, the full-matrix lottery search kernel with hit
latch, the host seed-chain/tile-search driver and the sm_80 jackpot
hit-signal reset, each with bit-exact CPU references and tests.
V5 certificate wire format and node-side verification, the Go FFI
bindings for the FP16 scheme, base-aware tile opening and the
py-pearl-mining certificate assembly (operand-commitment opener and
proof builder).
Gateway submit-direct certificate path, the vLLM miner FP16 layer and
mining dispatch, miner-base scheme admission/dispatch glue and the node
mining/solve switches that arm V5.
Make the header-bound recursive ZK proof the consensus path and retire
the plaintext certificate: embedded verifier cache, zk/zk_cert API,
tightened wrapper-legal consensus envelope and the e2e prove/verify
fixture.
Reproducible A100 HMMA capture/calibration/attack-cost harness, the
on-GPU winning-tile to ZK-consensus validation example and test, and the
whitepaper appendix referencing the harness.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant