Repository navigation
Conversation
The prototype in sha256.h was guarded by WOLFSSL_HAVE_LMS together with WOLFSSL_LMS_FULL_HASH, and nothing in the tree ever defines the latter. The definition, meanwhile, is only compiled in the arm that selects a software transform. On a port that supplies its own Update and Final, LMS therefore declared and called a function that was never built. Add WOLFSSL_HAVE_SHA256_HASH_BLOCK to sha256.h, listing in one place the ports where the helper is absent, and use it for both the prototype and the definition. The list is negative so that naming a port that does have the transform only costs the fast path, while a build combining two ports falls back instead of guessing. A compile time check in sha256.c catches the list drifting away from the arms that select XTRANSFORM. Move every in-tree caller onto the new macro in the same commit: the wolfCrypt and API test suites and the MC/DC hash fault injector all gated on WOLFSSL_LMS_FULL_HASH too, so a build without raw hash access kept calling a function that was not compiled. WOLFSSL_LMS_FULL_HASH is now referenced nowhere and leaves the known macro list. LMS now falls back to the full hash API when the helper is missing.
Two independent changes to the hot path. The SHA2 parameter sets ran every F, H and PRF through the streaming hash API: a free, a context copy, two updates and a final, all to compute one 64 byte compression. The free was redundant, because wc_Sha256Copy releases the destination itself, and the rest can be replaced by building the block that follows the pre-computed PK.seed midstate and compressing it once. The four helpers now share one dispatcher, which takes wc_Sha256HashBlock where it is available and the streaming API otherwise. Measured on a Xeon at 2.1 GHz, SHA2-128f signing goes from 28.1 ms to 22.7 ms. The SHAKE parameter sets spend about 93 percent of a signature in the four-way AVX2 Keccak permutation, while the eight-way AVX512 permutation that ML-DSA, ML-KEM and FrodoKEM already use sat unused. Measured per lane, eight ways beat four by two to one or better. Add an eight-way path for the WOTS+ chains, dispatched at run time on capable CPUs. SHAKE-128s signing goes from 309 ms to 194 ms, SHAKE-256s from 481 ms to 332 ms and SHAKE-128f from 15.6 ms to 11.2 ms. FORS and the verify side still use the four-way path. The eight-way helpers loop over the lanes rather than unrolling like their four-way counterparts. Setting up the state is a fraction of a percent of a signature, so the unrolling buys nothing worth the risk of an index error. Public keys and signatures are unchanged. Checked byte for byte against the previous code for all twelve parameter sets, with the eight-way path active and with WOLFSSL_SLHDSA_FULL_HASH forcing the streaming API. The PRF output that now takes the wc_Sha256HashBlock path is secret, and on x86-64 CPUs without MOVBE that function byte-reverses the digest through a stack buffer it never cleared. Wipe it.
The eight-way AVX512 Keccak path covered the WOTS+ chains only. Extend it to the FORS subtrees, which is where the remaining batched hashing sits. A FORS level narrower than eight nodes still finishes on the four-way path, so the last level of each subtree is unchanged. Separately, the batched helpers each allocated a Keccak state, used it for one group of hashes and freed it again. A WOTS+ public key needs one state, not one per group of four or eight, and a FORS subtree likewise. Hoist the state and the unchanging head of it into the callers that own the whole computation and pass them down. A SHAKE-128s signature under WOLFSSL_SMALL_STACK went from 143755 allocations to 11982, and the same change takes the buffers out of the callees' stack frames. Signing, measured on a Xeon at 2.1 GHz against the state before this work, taking each parameter set from one benchmark run: SHAKE-128s 344 ms -> 190 ms SHA2-128s 575 ms -> 484 ms SHAKE-192s 535 ms -> 315 ms SHA2-192s 1076 ms -> 816 ms SHAKE-256s 461 ms -> 274 ms SHA2-256s 907 ms -> 736 ms SHAKE-128f 15.5 ms -> 10.8 ms SHA2-128f 28.2 ms -> 22.0 ms SHAKE-192f 24.4 ms -> 17.1 ms SHA2-192f 45.8 ms -> 37.6 ms SHAKE-256f 51.5 ms -> 30.7 ms SHA2-256f 95.4 ms -> 76.5 ms SHA2 verification gains about a quarter from the hash rework. SHAKE verification is unchanged because it does not use the batched helpers. Public keys and signatures are unchanged. Checked byte for byte against the previous code for all twelve parameter sets, on the eight-way path and with NO_AVX512_SUPPORT forcing the four-way one.
Completing the WOTS+ chains of a signature was the last batched hashing
still on the four-way AVX2 path, so verification saw none of the earlier
work. Chains are gathered in order of the message value they start at,
which means a group of eight has to be brought up in steps: each chain
joins the ones already running when its own start is reached, and the
whole group then grows to w-1 together. Written as a loop over the group
rather than the unrolled ladder the four-way code uses. The last few
chains, which do not fill a group, finish one at a time.
The buffers a WOTS+ public key needs were still allocated per public
key, and a subtree builds thousands of them. Hand them to the subtree
instead: xmss_sign and key generation own one set and pass it down. A
SHAKE-128s signature under WOLFSSL_SMALL_STACK now makes 1081
allocations where it made 143755 before this series, and 225 on the
build without assembly. Peak stack is unchanged, since the buffers were
live across the same calls either way.
Measured on a Xeon at 2.1 GHz against the state before this series,
each parameter set from one benchmark run:
sign verify
SHAKE-128s 344 -> 190 ms 0.370 -> 0.269 ms
SHAKE-192s 535 -> 315 ms 0.523 -> 0.366 ms
SHAKE-256s 461 -> 276 ms 0.743 -> 0.519 ms
SHAKE-128f 15.5 -> 10.1 ms 1.039 -> 0.702 ms
SHAKE-192f 24.4 -> 15.6 ms 1.433 -> 0.959 ms
SHAKE-256f 51.5 -> 30.4 ms 1.467 -> 0.969 ms
SHA2-128s 575 -> 469 ms 0.573 -> 0.479 ms
SHA2-256f 95.4 -> 75.8 ms 2.456 -> 1.995 ms
FORS verification and the WOTS+ signing chains stay on the four-way
path. They are about a tenth and a five hundredth of their operations
respectively, so batching them further is not worth the code.
Public keys and signatures are unchanged, checked byte for byte against
the previous code for all twelve parameter sets on both paths.
Verification is the one path here where a defect means a false accept
rather than a visible failure, so the check for that goes in the test
suite rather than beside it. test_wc_slhdsa_sign_vfy now hands each
signature it makes to a helper that requires SIG_VERIFY_E from nine
single bit flips spread evenly over the signature, a changed message, a
changed context and an absent context, then re-checks the restored
signature so a rejection cannot be an artefact of leftover damage.
Reusing the keys and signatures sign_vfy already makes keeps the cost to
a few verifies per parameter set. Confirmed to bite by reverting the
root comparison to always accept.
The eight-way Keccak permutation is emitted by the SHA-3 assembly only for ML-KEM and ML-DSA builds, so an SLH-DSA build with the Intel assembly and neither of those failed to link. Add WOLFSSL_HAVE_SLHDSA to the x8 guards, matching the x4 ones, and close the guard after the permutation so the seed helpers that follow stay as they were. Two entries of the WOLFSSL_HAVE_SHA256_HASH_BLOCK port list named macros a port only sets inside its own .c file, so they read as undefined in the header and excluded nothing. Test the gate macros instead, and record the constraint. Keygen allocated the shared WOTS+ buffers with XMALLOC while xmss_sign used WC_DECLARE_VAR. Consolidate on the latter: heap-free builds work again, and a partial allocation no longer leaks. Guard the chain value buffer so a small memory build without the batched path stops carrying one it never reads. Clear the Keccak state and node buffers that hold FORS secret keys before releasing them, and clear the SHA-256 block helper's working buffers unconditionally rather than per call site. Also: reject a message that cannot fit the block with its padding, return early from a zero step chain, assert the len parity the buffer overshoot depends on, and correct the parameter documentation the refactor left behind. Added two CI configs covering the build breaks: SLH-DSA with the assembly but without ML-KEM/ML-DSA, and LMS/SLH-DSA without raw hash access.
Only the generic wc_Sha256Copy() frees its destination. The KCAPI and AF_ALG ports overwrite it wholesale and allocate a fresh handle, and the generic one returns before its free when a crypto callback takes the copy. Dropping the wc_Sha256Free()/wc_Sha512Free() ahead of the copy therefore stranded a hash handle on every inner hash on those builds. Restore it in all four places and say in the comment why it is needed. wc_InitSha256() takes the crypto callback default device id, not INVALID_DEVID, so initialising the consumer objects with INVALID_DEVID while the midstate objects kept the default made SLHDSA_SHA256_RAW_OK() always true. The direct block path then restored a midstate the callback had never written. Initialise both the same way again and have the guard check the midstate as well, since that is the object whose state words are read. The memory zeroization check registered the block buffer after the hash had run and only when the caller asked for a wipe, so it could not catch an early return and did not match the now unconditional clear. Register it while the data is live and drop the flag. Wipe the WOTS+ chain values at the size of the buffer rather than a parameter-set expression that no longer relates to it, and the FORS nodes at the size of theirs. The obvious length for the nodes, m * n, is wrong there: the Merkle loop halves m once per level, so on the success path it names the width of the level it stopped on rather than the number of entries written, and all but the last few leaf values would survive the free. Also: assert the lane and rate bounds the eight-way helpers assume, taking the rate bound from the widest absorb, H over a tree index, rather than the narrower WOTS+ one, assert the SHA-256 layout the direct path depends on, zero the shared buffer struct at both owners, rename a test that is no longer LMS specific, drop a configure flag that is an alias of one already passed, follow the generated assembly's condition naming, and record the verification stack growth in the ChangeLog.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Please describe the scope of the fix or feature addition.
Fixes zd#
Testing
How did you test?
Checklist