Skip to content

Slhdsa optimization - #52

Closed
Frauschi wants to merge 6 commits into
masterfrom
slhdsa_optimization
Closed

Frauschi wants to merge 6 commits into
masterfrom
slhdsa_optimization

Conversation

@Frauschi

Copy link
Copy Markdown
Owner

Description

Please describe the scope of the fix or feature addition.

Fixes zd#

Testing

How did you test?

Checklist

  • added tests
  • updated/added doxygen
  • updated appropriate READMEs
  • Updated manual and documentation

The prototype in sha256.h was guarded by WOLFSSL_HAVE_LMS together with
WOLFSSL_LMS_FULL_HASH, and nothing in the tree ever defines the latter.
The definition, meanwhile, is only compiled in the arm that selects a
software transform. On a port that supplies its own Update and Final,
LMS therefore declared and called a function that was never built.

Add WOLFSSL_HAVE_SHA256_HASH_BLOCK to sha256.h, listing in one place the
ports where the helper is absent, and use it for both the prototype and
the definition. The list is negative so that naming a port that does
have the transform only costs the fast path, while a build combining two
ports falls back instead of guessing. A compile time check in sha256.c
catches the list drifting away from the arms that select XTRANSFORM.

Move every in-tree caller onto the new macro in the same commit: the
wolfCrypt and API test suites and the MC/DC hash fault injector all
gated on WOLFSSL_LMS_FULL_HASH too, so a build without raw hash access
kept calling a function that was not compiled. WOLFSSL_LMS_FULL_HASH is
now referenced nowhere and leaves the known macro list.

LMS now falls back to the full hash API when the helper is missing.
Two independent changes to the hot path.

The SHA2 parameter sets ran every F, H and PRF through the streaming
hash API: a free, a context copy, two updates and a final, all to
compute one 64 byte compression. The free was redundant, because
wc_Sha256Copy releases the destination itself, and the rest can be
replaced by building the block that follows the pre-computed PK.seed
midstate and compressing it once. The four helpers now share one
dispatcher, which takes wc_Sha256HashBlock where it is available and the
streaming API otherwise. Measured on a Xeon at 2.1 GHz, SHA2-128f
signing goes from 28.1 ms to 22.7 ms.

The SHAKE parameter sets spend about 93 percent of a signature in the
four-way AVX2 Keccak permutation, while the eight-way AVX512 permutation
that ML-DSA, ML-KEM and FrodoKEM already use sat unused. Measured
per lane, eight ways beat four by two to one or better. Add an eight-way
path for the WOTS+ chains, dispatched at run time on capable CPUs.
SHAKE-128s signing goes from 309 ms to 194 ms, SHAKE-256s from 481 ms
to 332 ms and SHAKE-128f from 15.6 ms to 11.2 ms. FORS and the verify
side still use the four-way path.

The eight-way helpers loop over the lanes rather than unrolling like
their four-way counterparts. Setting up the state is a fraction of a
percent of a signature, so the unrolling buys nothing worth the risk of
an index error.

Public keys and signatures are unchanged. Checked byte for byte against
the previous code for all twelve parameter sets, with the eight-way path
active and with WOLFSSL_SLHDSA_FULL_HASH forcing the streaming API.

The PRF output that now takes the wc_Sha256HashBlock path is secret, and
on x86-64 CPUs without MOVBE that function byte-reverses the digest
through a stack buffer it never cleared. Wipe it.
The eight-way AVX512 Keccak path covered the WOTS+ chains only. Extend it
to the FORS subtrees, which is where the remaining batched hashing sits.
A FORS level narrower than eight nodes still finishes on the four-way
path, so the last level of each subtree is unchanged.

Separately, the batched helpers each allocated a Keccak state, used it
for one group of hashes and freed it again. A WOTS+ public key needs one
state, not one per group of four or eight, and a FORS subtree likewise.
Hoist the state and the unchanging head of it into the callers that own
the whole computation and pass them down. A SHAKE-128s signature under
WOLFSSL_SMALL_STACK went from 143755 allocations to 11982, and the same
change takes the buffers out of the callees' stack frames.

Signing, measured on a Xeon at 2.1 GHz against the state before this
work, taking each parameter set from one benchmark run:

  SHAKE-128s  344 ms -> 190 ms    SHA2-128s  575 ms -> 484 ms
  SHAKE-192s  535 ms -> 315 ms    SHA2-192s 1076 ms -> 816 ms
  SHAKE-256s  461 ms -> 274 ms    SHA2-256s  907 ms -> 736 ms
  SHAKE-128f 15.5 ms -> 10.8 ms   SHA2-128f 28.2 ms -> 22.0 ms
  SHAKE-192f 24.4 ms -> 17.1 ms   SHA2-192f 45.8 ms -> 37.6 ms
  SHAKE-256f 51.5 ms -> 30.7 ms   SHA2-256f 95.4 ms -> 76.5 ms

SHA2 verification gains about a quarter from the hash rework. SHAKE
verification is unchanged because it does not use the batched helpers.

Public keys and signatures are unchanged. Checked byte for byte against
the previous code for all twelve parameter sets, on the eight-way path
and with NO_AVX512_SUPPORT forcing the four-way one.
Completing the WOTS+ chains of a signature was the last batched hashing
still on the four-way AVX2 path, so verification saw none of the earlier
work. Chains are gathered in order of the message value they start at,
which means a group of eight has to be brought up in steps: each chain
joins the ones already running when its own start is reached, and the
whole group then grows to w-1 together. Written as a loop over the group
rather than the unrolled ladder the four-way code uses. The last few
chains, which do not fill a group, finish one at a time.

The buffers a WOTS+ public key needs were still allocated per public
key, and a subtree builds thousands of them. Hand them to the subtree
instead: xmss_sign and key generation own one set and pass it down. A
SHAKE-128s signature under WOLFSSL_SMALL_STACK now makes 1081
allocations where it made 143755 before this series, and 225 on the
build without assembly. Peak stack is unchanged, since the buffers were
live across the same calls either way.

Measured on a Xeon at 2.1 GHz against the state before this series,
each parameter set from one benchmark run:

               sign              verify
  SHAKE-128s   344 -> 190 ms     0.370 -> 0.269 ms
  SHAKE-192s   535 -> 315 ms     0.523 -> 0.366 ms
  SHAKE-256s   461 -> 276 ms     0.743 -> 0.519 ms
  SHAKE-128f  15.5 -> 10.1 ms    1.039 -> 0.702 ms
  SHAKE-192f  24.4 -> 15.6 ms    1.433 -> 0.959 ms
  SHAKE-256f  51.5 -> 30.4 ms    1.467 -> 0.969 ms
  SHA2-128s    575 -> 469 ms     0.573 -> 0.479 ms
  SHA2-256f   95.4 -> 75.8 ms    2.456 -> 1.995 ms

FORS verification and the WOTS+ signing chains stay on the four-way
path. They are about a tenth and a five hundredth of their operations
respectively, so batching them further is not worth the code.

Public keys and signatures are unchanged, checked byte for byte against
the previous code for all twelve parameter sets on both paths.

Verification is the one path here where a defect means a false accept
rather than a visible failure, so the check for that goes in the test
suite rather than beside it. test_wc_slhdsa_sign_vfy now hands each
signature it makes to a helper that requires SIG_VERIFY_E from nine
single bit flips spread evenly over the signature, a changed message, a
changed context and an absent context, then re-checks the restored
signature so a rejection cannot be an artefact of leftover damage.
Reusing the keys and signatures sign_vfy already makes keeps the cost to
a few verifies per parameter set. Confirmed to bite by reverting the
root comparison to always accept.
The eight-way Keccak permutation is emitted by the SHA-3 assembly only for
ML-KEM and ML-DSA builds, so an SLH-DSA build with the Intel assembly and
neither of those failed to link. Add WOLFSSL_HAVE_SLHDSA to the x8 guards,
matching the x4 ones, and close the guard after the permutation so the seed
helpers that follow stay as they were.

Two entries of the WOLFSSL_HAVE_SHA256_HASH_BLOCK port list named macros a
port only sets inside its own .c file, so they read as undefined in the
header and excluded nothing. Test the gate macros instead, and record the
constraint.

Keygen allocated the shared WOTS+ buffers with XMALLOC while xmss_sign used
WC_DECLARE_VAR. Consolidate on the latter: heap-free builds work again, and
a partial allocation no longer leaks. Guard the chain value buffer so a
small memory build without the batched path stops carrying one it never
reads.

Clear the Keccak state and node buffers that hold FORS secret keys before
releasing them, and clear the SHA-256 block helper's working buffers
unconditionally rather than per call site.

Also: reject a message that cannot fit the block with its padding, return
early from a zero step chain, assert the len parity the buffer overshoot
depends on, and correct the parameter documentation the refactor left
behind.

Added two CI configs covering the build breaks: SLH-DSA with the assembly
but without ML-KEM/ML-DSA, and LMS/SLH-DSA without raw hash access.
Only the generic wc_Sha256Copy() frees its destination. The KCAPI and AF_ALG
ports overwrite it wholesale and allocate a fresh handle, and the generic one
returns before its free when a crypto callback takes the copy. Dropping the
wc_Sha256Free()/wc_Sha512Free() ahead of the copy therefore stranded a hash
handle on every inner hash on those builds. Restore it in all four places and
say in the comment why it is needed.

wc_InitSha256() takes the crypto callback default device id, not
INVALID_DEVID, so initialising the consumer objects with INVALID_DEVID while
the midstate objects kept the default made SLHDSA_SHA256_RAW_OK() always
true. The direct block path then restored a midstate the callback had never
written. Initialise both the same way again and have the guard check the
midstate as well, since that is the object whose state words are read.

The memory zeroization check registered the block buffer after the hash had
run and only when the caller asked for a wipe, so it could not catch an early
return and did not match the now unconditional clear. Register it while the
data is live and drop the flag.

Wipe the WOTS+ chain values at the size of the buffer rather than a
parameter-set expression that no longer relates to it, and the FORS nodes at
the size of theirs. The obvious length for the nodes, m * n, is wrong there:
the Merkle loop halves m once per level, so on the success path it names the
width of the level it stopped on rather than the number of entries written,
and all but the last few leaf values would survive the free.

Also: assert the lane and rate bounds the eight-way helpers assume, taking
the rate bound from the widest absorb, H over a tree index, rather than the
narrower WOTS+ one, assert the SHA-256 layout the direct path depends on,
zero the shared buffer struct at both owners, rename a test that is no longer
LMS specific, drop a configure flag that is an alias of one already passed,
follow the generated assembly's condition naming, and record the verification
stack growth in the ChangeLog.
@Frauschi Frauschi closed this Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant