Conversation
jonpry
force-pushed
the
fp16-scheme
branch
3 times, most recently
from
October 4, 2026 21:13
39756f8 to
81a230a
Compare
Whitepaper, STARK-feasibility and ZK-binding design notes for the A100 FP16 accumulation-hardness proof-of-useful-work scheme.
zk-pow API for the A100 scheme: dtype, params, noise, quantization, policy, operand commitment, accumulation and the plaintext proof/verify datapath, with A100 reference dot-product vectors.
Plonky2 circuits proving the scheme datapath: blake3 commitment, noise, noisy-quant FMA, row-scale, policy and xor-fold STARKs, the A100 matmul STARK, cross-table lookups, driver and recursive wrapper.
pearl-gemm sm_80 kernels and host code for the full miner datapath: FP16 GEMM, noise lines, noisy-quant, policy and commitment, the end-to-end tile pipeline, the full-matrix lottery search kernel with hit latch, the host seed-chain/tile-search driver and the sm_80 jackpot hit-signal reset, each with bit-exact CPU references and tests.
V5 certificate wire format and node-side verification, the Go FFI bindings for the FP16 scheme, base-aware tile opening and the py-pearl-mining certificate assembly (operand-commitment opener and proof builder).
Gateway submit-direct certificate path, the vLLM miner FP16 layer and mining dispatch, miner-base scheme admission/dispatch glue and the node mining/solve switches that arm V5.
Make the header-bound recursive ZK proof the consensus path and retire the plaintext certificate: embedded verifier cache, zk/zk_cert API, tightened wrapper-legal consensus envelope and the e2e prove/verify fixture.
Reproducible A100 HMMA capture/calibration/attack-cost harness, the on-GPU winning-tile to ZK-consensus validation example and test, and the whitepaper appendix referencing the harness.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
FP16 accumulation-hardness proof-of-useful-work (A100)
An alternative Pearl PoUW instantiation whose unit of work is FP16 matrix multiplication on NVIDIA A100 (GA100,
sm_80) tensor cores. It extends the FP8 certificate scheme rather than replacing it: commitments, seed chain, state window, ticket, target, and the ZK split are shared, and FP8 remains available unchanged.Core idea. Hardness comes from the nonlinearity of the device's own accumulation, not a coarse rounding grid. The A100 tensor core truncates each product onto a per-group alignment grid before summing (groups of 8, 24-bit window, round-toward-zero per group). That per-product truncation does not commute with the reduction, so it cannot be expressed as a matrix multiplication — the same obstruction that makes the integer transcript scheme hard, applied for free on every MAC. This lets recovered products carry FP16-level accuracy instead of FP8-residual accuracy. BF16 is shown to be too close to linear and is not admitted.
The consensus path (read first)
The FP16 (A100) consensus certificate is the header-bound recursive ZK certificate (
wire.CertificateV5→verify_fp16_zk_cert_ffi→verify_wrapped_proof_with_headers): a constant-size wrapped plonky2 proof whose public inputs — opening keys, noise seeds, jackpot key, operand roots, geometry — are re-derived from the block header + authenticated ancestor and pinned by equality before the proof is verified and difficulty checked, against an embedded per-profile verifier cache (no circuit build at verify time). The earlier plaintext certificate is retired as a consensus path;Fp16PlainProofnow serves only as the prover's operand-opener bundle (thePlainFp16wire name is historical). Everything is additive: V1–V4 (FP8 and earlier) are byte-for-byte untouched.What's included
docs/fp16_scheme/) — LaTeX + built PDF in the style of the FP8 spec, a STARK feasibility/soundness note, and the ZK-binding design.zk-pow/src/api/fp16/) — FP16 (E5M10) dtype, the bit-exact A100 accumulation model, seed-derived low-rank noise (B-then-A keyed-BLAKE3 chain), FP16 noisy quantization, the "unpredictable accumulation steps" jackpot policy (breakpoint density ≥ 0.30, certified-work ratio ≥ 1.2), and a keyed Merkle commitment over raw FP16 rows.zk-pow/src/circuit/fp16/) — amatmul_a100STARK, policy STARK, forked BLAKE3 egress AIR, noise/noisy-quant/row-scale/xor-fold STARKs with CTL-bound decode/shift/width columns, a tight per-group census, limb-range soundness splits, and a batched driver producing a constant-size, verifiable, tamper-rejecting recursive proof. All datapath branches (normal/zero/subnormal) are constrained; AIRs are degree ≤ 3.sm_80miner pipeline (miner/pearl-gemm/) — raw-CUDA GA100 kernels for the full datapath (mma.sync.m16n8k16.f32.f16GEMM, noise lines, noisy-quant, policy, commitment), the torch-free tile pipeline, the full-matrix lottery search with atomic hit latch, the host seed-chain/search driver, and thesm_80jackpot hit-signal reset — each with a bit-exact CPU/Python reference and tests.CertificateV5wire format and node-sideverifyCertificateV5, the Go FFI bindings, base-aware tile opening, and the Python certificate assembly (operand-commitment opener + proof builder).zk/zk_certAPI, a wrapper-legal consensus envelope, and an end-to-end prove/verify fixture.docs/fp16_scheme/validation/) — a reproducible A100 HMMA capture / calibration / rounding-mode probe and an attack-cost benchmark, run directly on real tensor cores (no software model in the loop), plus an on-GPU winning-tile → ZK-consensus example and test.Validation
api::fp16+circuit::fp16: full suite green, including the end-to-end batched + wrapped ZK proof (verify + tamper-reject) and the subnormal-branch constraints.circuit::fp8: unchanged (no regression).zkpow):CertificateV5routes + verifies throughVerifyCertificate(accept when the embedded cache covers the fixture / tamper-reject / commitment-mismatch / wire round-trip / version-aware size cap / activation gate).sm_80HMMA; attack-cost benchmark measures hardness above the whitepaper floor. (The harness runs on GPU hardware; CI compares the model to the committed capture vectors.)Gating / activation
V5 is deliberately inert for consensus and does not activate on any network as-is:
IsCertVersionAllowed(V5)isfalse,Fp16ForkHeightis unset on every network, and CPU solve errors on V5. Decode/route/verify are wired but blocked at header-sanity; activation requires flipping the allow-list and setting a fork height together (a deliberate hard-fork decision).Remaining before activation (documented, out of scope here)
fp16_cache.binmust be built over the full consensus-legal envelope (or the universal wrapper below must land) before V5 is allowed. Fail-closed until then: a legal geometry whose profile is absent is rejected.Note on the hardness assumption
The scheme's security rests on the software accumulation model being bit-exact to A100 silicon. The committed validation harness captures from real HMMA independently of the model and confirms the match; this independence is exercised on GPU hardware, not in CI (CI compares the model to the committed capture vectors).