Repository navigation
Conversation
Frauschi
commented
Sep 7, 2026
Frauschi
force-pushed
the
zephyr_rw612
branch
from
September 7, 2026 13:27
f1dde98 to
853c5a4
Compare
Frauschi
force-pushed
the
zephyr_rw612
branch
from
September 7, 2026 15:09
853c5a4 to
98d85e3
Compare
…x-M3 WOLFSSL_ARM_ARCH_7M selects the UMAAL-free variants in sp_cortexm.c and in the thumb2-* assembly for Poly1305, Curve25519 and ML-KEM. It was derived from __ARM_ARCH_7M__, which names the Cortex-M3 and nothing else. UMAAL is not an ARMv7-M property though, it belongs to the DSP extension, which on ARMv8-M is optional. An ARMv8-M part built without it, such as the Cortex-M33 in the NXP RW612, has no UMAAL and no __ARM_ARCH_7M__ either, so it took the UMAAL path and would not assemble. WOLFSSL_SP_NO_UMAAL is not a way out: it is derived from this symbol and covers only a subset of the sites. Toolchains define __ARM_FEATURE_DSP exactly when the extension is present, so key on its absence on an M-profile core. Verified against gcc for M0, M3, M4, M7, M23, M33 and M33+nodsp: it is set only where UMAAL is genuinely missing. The __ARM_ARCH_7M__ arm stays, so no configuration that worked before changes, and the __ARM_ARCH_PROFILE test matches the one already in cpuid.c.
The WOLFSSL_ZEPHYR block includes <zephyr/kernel.h>, <zephyr/sys/printk.h>, <zephyr/sys/util.h> and <stdlib.h>, and declares z_realloc. Assembly sources reach settings.h through libwolfssl_sources_asm.h, so all of that was being handed to the assembler, which stops at the first C declaration it meets: stddef.h:160: Error: no such instruction: `typedef long int ptrdiff_t' That makes every wolfSSL .S file unbuildable on Zephyr. It is not specific to one architecture or one algorithm - reproduced on x86_64 with chacha_asm.S, sha256_asm.S, sha512_asm.S and aes_asm.S, and on a Cortex-M33 with thumb2-aes-asm.S, which fails identically: stddef.h:160: Error: bad instruction `typedef int ptrdiff_t' Zephyr ports that use the *_c.c inline-assembly variants never noticed, because those are compiled as C. Guard the block with __ASSEMBLER__, matching what settings.h already does for the STM32MP13 header. Assembly still gets the feature macros from user_settings.h; it just no longer gets the C declarations. No effect on any C translation unit - __ASSEMBLER__ is only defined when the compiler driver is assembling.
…s them
The nine div-word helpers bind a variable to rax with the bare asm keyword:
register sp_digit r asm("rax");
asm is a GNU extension, not an ISO C keyword, so it is unavailable under
-std=c17 and the file will not compile:
sp_x86_64.c:597: error: expected '=', ',', ';', 'asm' or '__attribute__'
before 'asm'
Zephyr compiles with -std=c17, which is how this surfaced, but any strict-ISO
build hits it. __asm__ is the alternate spelling GCC and Clang keep available
regardless of -std, and it is what the very next line of each of these
functions already uses for the statement asm - so this only makes the two
consistent.
Frauschi
force-pushed
the
zephyr_rw612
branch
from
September 8, 2026 08:55
98d85e3 to
9793e7f
Compare
Frauschi
commented
Sep 8, 2026
CPUID's AVX and AVX-512 bits say the silicon has the unit. Executing one of those instructions also requires the OS to have turned on extended state: CR4.OSXSAVE, plus the matching XCR0 components - SSE and AVX for VEX encodings, and opmask, ZMM_Hi256 and Hi16_ZMM on top of those for EVEX. An OS that does not context-switch those registers leaves them clear, and the instruction then raises #UD no matter what CPUID advertises. cpuid_set_flags() tested the feature bits alone, so on such a system wolfSSL set CPUID_AVX1/AVX2/VAES/AVX512* and dispatched to implementations the CPU would refuse. Zephyr is one such system - its x86 context switch is fxsave/fxrstor and it never sets CR4.OSXSAVE - and the result is an immediate crash the first time SHA-256 or AES-GCM picks a vector path: <err> os: Invalid opcode >>> ZEPHYR FATAL ERROR 0: CPU exception on CPU 0 Gate each family on OSXSAVE plus the XCR0 components it actually needs, the sequence Intel documents for this. The AVX-512 dispatch sites are the reason the masks have to differ: AesEcbEncryptBlocks(), AES_GCM_init_*, chacha_avx512_beneficial() and poly1305_use_avx512() all branch on IS_INTEL_AVX512 && IS_INTEL_VAES without consulting AVX1 or AVX2, so gating only the 256-bit flags would leave the 512-bit paths reachable on exactly the systems this is meant to protect. Keeping them on one gate would also be wrong in the other direction, since 0x06 does not cover the ZMM state. VAES is VEX or EVEX encoded, so it is gated with the AVX state; the AVX-512 dispatch sites test it alongside CPUID_AVX512, which carries the stricter mask. Where an OS does enable XSAVE, which is every mainstream one, nothing changes. XGETBV is itself only legal once OSXSAVE is set, so the two tests have to stay in that order.
The ARMv8 AES and SHA-256 ports emit aese, aesmc and the sha256 crypto
instructions with no .arch or .arch_extension directive of their own, so they
assemble only when the command line already names a CPU that has the
extension. Built against plain ARMv8-A, which is what -mcpu=cortex-a53 gives
you, the assembler refuses them:
armv8-aes-asm.S: Error: selected processor does not support
`aese v0.16b,v1.16b'
Zephyr is one caller that lands there. It derives -mcpu from the board's
Kconfig and has no symbol for the crypto extension, so there is nothing for a
build system layered on top to key on, and naming a CPU in the module would
retarget the library on every board that is not that CPU. Anything else
compiling for the base architecture meets the same wall.
A file-scope .arch_extension crypto applies to the rest of the translation
unit, so one directive per file does what would otherwise take sixty inline
copies. All three outputs carry it - bare in the GNU assembler file, wrapped
in __asm__() in the C file, and a comment in the armasm64 file, which
assembles these instructions unconditionally. Each sits inside the existing
WOLFSSL_ARMASM_NO_HW_CRYPTO guard, so a build that has opted out of the
hardware paths never asks for an extension it will not use.
Regenerated from the scripts repository, which carries the matching change.
Verified on qemu_cortex_a53 under Zephyr: wolfcrypt_test passes 43 of 43 with
2723 aese instructions in the image, where before neither file assembled.
The SHA-3, SHA-512 and FrodoKEM ARM64 ports guard their .arch_extension sha3
with __APPLE__, on the assumption that every other assembler learns about the
extension from the command line. That does not hold. Assembling for plain
ARMv8-A with the GNU assembler:
armv8-sha3-asm.S: Error: selected processor does not support
`eor3 v31.16b,v0.16b,v5.16b,v10.16b'
Emit the directive unconditionally. It is a no-op where the extension is
already enabled, so the Apple path is unchanged.
Regenerated from the scripts repository, which carries the matching change.
L_mlkem_aarch64_consts sat outside the WOLFSSL_HAVE_MLKEM conditional while
every other constant in the file - zetas, zetas_inv, q - sat inside it. With
ML-KEM off the C output then declares a static const nothing references, and a
-Werror build stops:
armv8-mlkem-asm_c.c:37: error: 'L_mlkem_aarch64_consts' defined but not used
[-Werror=unused-const-variable=]
Zephyr compiles wolfSSL with -Werror, so this made armv8-mlkem-asm_c.c
unbuildable on any aarch64 target that does not enable ML-KEM - which is every
one of them by default.
The cause is in the generator: Kyber_ARM64_Neon_ASM#initialize emitted the
constant, but write() is what opens the guard, and it runs later. Fixed there by
moving the emission into a define_consts() called from write(), mirroring the
existing define_q(); these files are the regenerated result. All three outputs
change by the same three-line move and nothing else.
A WOLFSSL_CUSTOM_CURVES build fails at runtime on its own largest enabled curve. Enabling Brainpool with sizes 256, 384 and 512 gives ecc_test_curve_size 64 failed! (WC_KEY_SIZE_E, -234) and it is not a bp512 quirk: a bp256-only build fails on bp256, a bp256+bp384 build on bp384. Every time it is the largest curve in the set. The ceiling is off by exactly one bit. Under custom curves ECC_KEY_MAX_BITS() takes the "dp->size * 8 + 1" variant, for orders a bit larger than their prime, but MAX_ECC_BITS_NEEDED is the plain curve size. MP_BITS_CNT() rounds up to whole digits, so that single bit puts the key one digit past the ceiling and MP_BITS_OVER_MAX() rejects it. Carry the same bit in the ceiling via MAX_ECC_BITS_EXTRA, and compare against it in the guard that rejects a too-small user MAX_ECC_BITS. Second, MAX_ECC_BITS_USE derived from MAX_ECC_BITS_NEEDED, so the overridable MAX_ECC_BITS never reached the runtime guard - defining it larger changed the key struct but not the check, and there was no way to size the working values for a curve that is not compiled in. Derive it from MAX_ECC_BITS instead; ecc.h already refuses a value below what the enabled curves need, so this only ever widens, and it is what an arbitrary curve passed to wc_ecc_set_custom_curve() needs. Both are no-ops for a build without custom curves.
The module hard-coded ECC_USER_CURVES with SECP256R1 and nothing else, so an application needing P-384, P-521 or a Brainpool curve had to patch user_settings.h. Each curve is now its own Kconfig option under WOLFSSL_ECC, with P-256 the default so existing configurations are unaffected. Two dependencies are encoded here rather than left for the next person to hit, because both fail in a way that does not point at the cause. Brainpool needs custom-curve support. Its curves are not prime-field NIST curves, and ecc.c refuses to build them without WOLFSSL_CUSTOM_CURVES. Custom curves in turn cannot coexist with the per-curve SP math this module selects, because SP's fast paths are written per curve and cannot represent an arbitrary one. Asking for Brainpool therefore moves the whole build onto the generic SP variant, which costs size and costs speed on the common curves. That is a real trade-off, and it is why Brainpool is opt-in. Brainpool also selects P-521, which is worth spelling out because it looks like overreach. On the generic variant an operation needs digit headroom above the modulus, so the enabled-curve ceiling has to be strictly larger than the largest Brainpool curve reachable; when it is not, that curve fails at runtime with WC_KEY_SIZE_E. Measured on native_sim: brainpoolP256r1 fails with a 256-bit ceiling, brainpoolP384r1 with a 384-bit one and brainpoolP512r1 with a 512-bit one - each time the largest curve in the set. The ceiling is MAX_ECC_BITS_NEEDED, which ecc.c derives from the enabled HAVE_ECCxxx set alone, so no other knob raises it. Brainpool stops at 512 bits, which makes P-521 the only size that lifts the ceiling without bringing a new Brainpool curve along with it. P-256 keeps its inverted sense: it is the one curve wolfCrypt enables by default, so it is turned off by defining NO_ECC256 rather than by omission. Verified on frdm_rw612. With P-384 and P-521 selected, all three curves generate a key and round-trip a signature on device; with the defaults the image is unchanged at P-256 only and 20 KB smaller, since neither the extra curves nor the generic math are pulled in.
The curve Kconfigs decide which curves wolfCrypt compiles and which SP implementation backs them, and getting that pairing wrong fails at runtime with WC_KEY_SIZE_E rather than at build time. Two wolfssl_test scenarios pin the combinations that matter: the NIST curves on their per-curve SP paths with no custom curve in sight, and the Brainpool set that forces the generic backend and depends on the curve ceiling P-521 provides.
…s used The build-profile options only write the module's user_settings.h, which is never read when the application supplies its own settings file. So CONFIG_WOLFSSL_ECC_384=y alongside a settings file was accepted and silently ignored, and the same went for every other feature option. They now depend on not having one, and disappear from the menu instead. This needs a tracking bool because WOLFSSL_SETTINGS_FILE is a string symbol and Kconfig evaluates a string in a logical context as always-false, which would have made every one of these dependencies quietly unsatisfiable. FIPS is gated on the master switch rather than on each version. WOLFCRYPT_FIPS writes HAVE_FIPS into user_settings.h, so it belongs under the rule; putting the dependency on the choice members alone would leave the version prompt visible with nothing selectable and its default unreachable, while CMake still compiled the FIPS bundle. Kconfig.tls-generic is left alone here. Only two of its 53 symbols reach any code, so gating the rest would make dead options look conditional; the commit after this one deletes the file and rehomes those two with the dependency. wolfssl_tls_sock's no-malloc scenario was setting five of those options next to a settings file, so they now name symbols that are not visible. Drop them and say why they are gone. Options that drive the build rather than the wolfCrypt configuration, meaning the implementation choice, the assembly port and the install path, are untouched and stay selectable either way.
Kconfig.tls-generic is a copy of Zephyr's mbedTLS Kconfig.tls-generic with the symbol prefix renamed, which is why a wolfSSL module carries options named after mbedTLS internals - ECP is their elliptic-curve-over-prime-field module and DP their domain parameters, neither of which means anything here. Renaming the prefix was all that happened: nothing was wired up behind it. Of the 53 config symbols in the file, exactly two reach any code, WOLFSSL_TLS_VERSION_1_2 and _1_3. Checked every symbol against the whole workspace - all .c, .h, .cmake, CMakeLists.txt, Kconfig and .conf files under zephyr/ and modules/crypto/wolfssl. The other 51 appear only in that file. Setting CONFIG_WOLFSSL_ECP_DP_SECP384R1_ENABLED=y has always done precisely nothing, and unlike the options the previous commit gated, these do nothing whether or not a settings file is present. Three of them are worse than inert. WOLFSSL_TLS_VERSION_1_0, _1_1 and _1_3 select WOLFSSL_ALLOW_TLSV10_ENABLED, WOLFSSL_NO_OLD_TLS_DISABLED and WOLFSSL_TLS13_ENABLED, none of which is defined anywhere. Kconfig accepts a select of an undefined symbol silently, so the reader sees a mechanism that does not exist. user_settings.h even tested CONFIG_WOLFSSL_TLS13_ENABLED alongside the real symbol; that half of the condition could never be true. Delete the file and move the two live symbols into the module's own Kconfig, next to WOLFSSL_DTLS where the other protocol options are, keeping the settings-file dependency. The dead selects go with it, as do the four dead options the two TLS samples were setting and the stale entry in .wolfssl_known_macro_extras. The ECC curve options that replaced the WOLFSSL_ECP_DP_* ones - and that do reach user_settings.h - went in earlier in this branch.
Four options guarded the assembly - WOLFCRYPT_ARMASM, its THUMB2 companion, WOLFCRYPT_INTELASM, and nothing at all for the single-precision math - and between them they could not express a working configuration on any target. WOLFCRYPT_ARMASM_THUMB2 selected the thumb2-* sources but nothing anywhere defined WOLFSSL_ARMASM_THUMB2, so every caller still took the ARMv8 path and the link failed on AES_*_AARCH32. It also needs WOLFSSL_ARMASM_NO_HW_CRYPTO, since Cortex-M has no AES or SHA extension and the Thumb2 code is the whole of it. The single-precision assembly had no option at all: WOLFSSL_SP_ARM_CORTEX_M_ASM sat commented out inside an #if 0 while sp_cortexm.c was already in the source list. WOLFCRYPT_INTELASM was worse: it named the 32-bit source pair while user_settings.h declared WOLFSSL_X86_64_BUILD, so aes_asm.S compiled to an empty object behind its WOLFSSL_X86_BUILD guard and aes_gcm_x86_asm.S emitted 32-bit code into a 64-bit build. It could not have produced a working image. Only one combination is ever right for a given board, so the four collapse into WOLFCRYPT_ASM and the module derives the rest from the CPU Zephyr reports: ARMv7-M and ARMv8-M mainline take Thumb2 plus the Cortex-M math, AArch64 the ARMv8 and ARM64 ones, other 32-bit ARM the ARMv8-32 and ARM32 ones, and x86_64 the Intel symmetric set with its correct 64-bit filenames. ARMv6-M and ARMv8-M baseline deliberately get no symmetric assembly. The Thumb2 port uses UBFX and LDRD, neither of which those cores have, so keying it on CONFIG_CPU_CORTEX_M would hand an M0 or M23 board code that cannot assemble. They still get the Thumb math, which is the only thing wolfCrypt ships for that profile - the same split the SP half already made. Each ARM port sets two single-precision macros, not one. WOLFSSL_SP_<cpu>_ASM compiles sp_<cpu>.c, the per-size backend for the RSA, DH and ECC sizes it covers; the unsuffixed WOLFSSL_SP_<cpu> enables the inline-asm word primitives inside sp_int.c, which is what the generic WOLFSSL_SP_MATH_ALL path uses for every other size and curve. sp_int.c's own header block names them as separate features and configure.ac sets them from separate options, neither implying the other. Defining only the _ASM half - which is what this module did - left Brainpool, custom curves and any uncovered RSA or DH size running with no assembly at all. The x86_64 set carries fe_x25519_asm.S but not the AES pair. curve25519.c takes the Intel path whenever USE_INTEL_SPEEDUP is defined, so omitting the former is a link error the moment an application enables the curve; the latter define nothing but _aesni entry points whose callers sit behind WOLFSSL_AESNI, which this module does not set, so they were bytes with no reachable caller. 32-bit x86 is not offered. Only the AES sources have an implementation at that width; sha256_asm.S, chacha_asm.S and poly1305_asm.S are all guarded to WOLFSSL_X86_64_BUILD and would contribute nothing. The x86_64 single-precision assembly is not offered either, and that one is worth recording. sp_x86_64_asm.S is AVX throughout - over fifteen hundred VEX instructions - and sp_x86_64.c calls into it unconditionally, with no CPUID dispatch and no scalar counterpart. Zephyr's x86 context switch is fxsave/fxrstor and it never sets CR4.OSXSAVE, so YMM state is neither saved nor enabled and those instructions fault. The symmetric sources are fine because they pick their implementation from CPUID at run time and now fall back correctly. Both set(TOOLCHAIN_C_FLAGS ...) calls go as well. They could never have had an effect: Zephyr applies that variable globally at zephyr/CMakeLists.txt:425 and only add_subdirectory's modules at line 788, so the assignment ran after the flags were consumed, and a plain set() in a subdirectory scope never reached the parent either. Confirmed by forcing the assignment to run unconditionally and finding neither -mcpu=cortex-a53+crypto nor -mstrict-align anywhere in the resulting compile database. HAVE___UINT128_T is defined wherever the compiler has the type. Every 64-bit single-precision backend works in 128-bit intermediates; an autoconf build learns the type is available from a configure probe, but with user settings nobody sets it and sp_int.c fails on an undeclared sp_int_word. AArch64 needs this as much as x86_64 does - qemu_cortex_a53 stops there before it ever reaches the assembler. One wolfssl_test scenario per port now covers the option in CI rather than by hand: qemu_x86_64 for Intel, qemu_cortex_a53 for AArch64, and mps2/an521/cpu0, a Cortex-M33, for Thumb2. The aarch64 one earned its keep immediately by catching an unused-constant error in the generated ML-KEM assembly, fixed earlier in this branch. Measured on frdm_rw612 with the benchmark, ops/sec: ECDSA P-256 sign 16 to 136, verify 10 to 102, RSA-2048 public 86 to 194. Thumb2 is worth 1.5x on AES-CBC and 1.3x on ChaCha20, and costs about 5% on SHA-256. wolfcrypt_test passes on frdm_rw612 and on all three QEMU targets with the option set.
Frauschi
force-pushed
the
zephyr_rw612
branch
from
September 8, 2026 15:12
9793e7f to
ee24352
Compare
Frauschi
commented
Sep 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Generic wolfSSL Zephyr-module work extracted from the NXP RW612 EdgeLock port, plus the one
settings.hchange that port needs. Everything here stands on its own and is independent of the crypto callback key store API (wolfSSL#11336); this branch ismaster+ 5 commits and carries no key store code.What is in it
settings: detect any M-profile core without UMAAL, not just the Cortex-M3WOLFSSL_ARM_ARCH_7Mwas derived from__ARM_ARCH_7M__, which names the Cortex-M3 and nothing else. UMAAL is not an ARMv7-M property though - it belongs to the DSP extension, optional on ARMv8-M. A Cortex-M33 built without DSP took the UMAAL path and would not assemble. The detection now keys on__ARM_FEATURE_DSPbeing absent on an M-profile core; the old__ARM_ARCH_7M__arm stays, so nothing that built before changes.zephyr: select ECC curves through KconfigThe module hard-coded
ECC_USER_CURVESwith SECP256R1 only, so needing P-384, P-521 or a Brainpool curve meant patchinguser_settings.h. Each curve is now its own Kconfig option, P-256 still the default. The menu covers P-256, P-384, the 512-bit size, P-521 and the Brainpool family - deliberately not the sizes below 256, the Koblitz curves or brainpoolP320r1, none of which have a TLS or wolfPSA use case worth the generic-SP cost.Brainpool selects P-521, which looks like overreach and is not: on the generic SP variant the enabled-curve ceiling must be strictly larger than the largest reachable Brainpool curve, and Brainpool stops at 512, so P-521 is the only size that lifts the ceiling without bringing a new Brainpool curve with it. Measured, not assumed - see the commit message.
zephyr: cover the ECC curve options with TwisterGetting the curve/SP pairing wrong fails at runtime with
WC_KEY_SIZE_E, not at build time. Twowolfssl_testscenarios pin the NIST-only and Brainpool combinations.zephyr: make the feature Kconfigs mean nothing when a settings file is usedThe build-profile options only write the module's
user_settings.h, which is never read when the application supplies its own settings file - soCONFIG_WOLFSSL_ECC_384=ynext to a settings file was accepted and silently ignored. They now depend on not having one and disappear from the menu.Kconfig.tls-genericis deliberately left alone: 47 of its 53 symbols are read by no code at all, so gating them would only make dead options look conditional. That file needs its own cleanup.zephyr: replace the assembly options with a single arch-derived oneFour options guarded the assembly and between them could not express a working configuration on any target.
WOLFCRYPT_ARMASM_THUMB2selected thethumb2-*sources but nothing definedWOLFSSL_ARMASM_THUMB2, so the link failed onAES_*_AARCH32. The single-precision assembly had no option at all.WOLFCRYPT_INTELASMnamed the 32-bit source pair whileuser_settings.hdeclaredWOLFSSL_X86_64_BUILD, soaes_asm.Scompiled to an empty object andaes_gcm_x86_asm.Semitted 32-bit code into a 64-bit build.They collapse into one
WOLFCRYPT_ASMand the module derives the port from the CPU Zephyr reports: Cortex-M, AArch64, other 32-bit ARM, and x86_64 with its correct 64-bit filenames.32-bit x86 is not offered - only the AES sources exist at that width. The x86_64 single-precision assembly is not offered either:
sp_x86_64_asm.Sis AVX throughout with no CPUID dispatch and no scalar counterpart, and Zephyr never enables YMM state.Both
set(TOOLCHAIN_C_FLAGS ...)calls go too. They could never have had an effect - Zephyr consumes that variable atzephyr/CMakeLists.txt:425and only adds modules at line 788, and a plainset()in a subdirectory scope never reaches the parent. Confirmed by forcing the assignment to run unconditionally and finding neither flag in the resulting compile database.Three library-wide fixes this needed
Enabling the assembly surfaced three defects in wolfSSL proper. Each is independent of Zephyr and is a separate commit at the front of the branch.
settings: keep the Zephyr platform block away from the assembler- theWOLFSSL_ZEPHYRblock includes<zephyr/kernel.h>and<stdlib.h>and declaresz_realloc, all of which reach.Sfiles throughlibwolfssl_sources_asm.h. Every wolfSSL.Sfile was unbuildable on Zephyr, on every architecture - reproduced on x86_64 and, identically, on a Cortex-M33 withthumb2-aes-asm.S. Guarded with__ASSEMBLER__, matching whatsettings.halready does for STM32MP13.sp_x86_64: spell the register bindings __asm__- nineregister sp_digit r asm("rax")declarations use the GNUasmkeyword, unavailable under-std=c17. The statement asm on the very next line of each function already uses__asm__.cpuid: only report AVX when the OS has enabled the YMM state-cpuid_set_flags()tested only the CPUID AVX feature bits, neverOSXSAVEorXGETBV. An OS that does not context-switch YMM leavesCR4.OSXSAVEclear, so those instructions raise#UDno matter what CPUID says. wolfSSL dispatched to AVX anyway and crashed on the first AVX SHA-256 or AES-GCM path. Now gated onOSXSAVEplusXGETBV(0) & 0x6, the sequence Intel documents. No change on any OS that enables XSAVE.Testing
qemu_x86(100%), including the Brainpool curve scenario and both no-malloc settings-file scenarios.qemu_x86_64withCONFIG_WOLFCRYPT_ASM=y: fullwolfcrypt_testrun, 43 passed / 0 failed,Exiting main with return code: 0, all seven Intel asm objects linked in. Before the cpuid fix the same image died withInvalid opcode/ZEPHYR FATAL ERROR 0: CPU exceptioninsidesp_256_get_point_33_4.frdm_rw612builds and links withCONFIG_WOLFCRYPT_ASM=y; 32-bitqemu_x86builds with the option correctly not offered.wolfcrypt_testpasses on device.