Skip to content

openssl compat: cut EVP, EVP_PKEY and PEM overhead, and fix defects found - #11689

Open
SparkiDev wants to merge 1 commit into
wolfSSL:masterfrom
SparkiDev:openssl_compat_api_perf_fixes_1
Open

SparkiDev wants to merge 1 commit into
wolfSSL:masterfrom
SparkiDev:openssl_compat_api_perf_fixes_1

Conversation

@SparkiDev

@SparkiDev SparkiDev commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Description

Performance:

  • Size WOLFSSL_SHA3_CTX's holder in pointers, not bytes. It reserved sizeof(wc_Sha3) pointers for a 440 byte object, so WOLFSSL_Hasher goes 3520 -> 448 bytes and WOLFSSL_EVP_MD_CTX 3552 -> 944. A static assert now catches the same mistake.
  • Resolve digests and ciphers by table pointer before comparing names, and return the table's own pointer from the accessors. EVP_sha256() 47.4 -> 1.2ns, EVP_MD_type 64.7 -> 4.3ns, EVP_CIPHER_nid 101.2 -> 8.1ns.
  • Copy only the fields past the hash union in EVP_MD_CTX_copy_ex, and leave the cipher union alone in EVP_CIPHER_CTX_init.
  • Set up an EVP_PKEY's RNG, an RSA key's blinding RNG and an EVP_PKEY's DER encoding on first use rather than at creation. Seeding a DRBG cost more than everything else about making a key put together. EVP_PKEY_new + free 13366 -> 34ns, RSA_new + free 13748 -> 230ns, EVP_PKEY_set1_RSA 13994 -> 45ns. Certificate parsing roughly halves as a result: d2i_X509 45978 -> 17198ns, X509_dup 45945 -> 17116ns, both now faster than OpenSSL. set1_DSA still encodes immediately, as WOLFSSL_DSA is not reference counted and the pkey only borrows the caller's key.
  • Keep the decoded public key on the X509 and hand out a reference, as OpenSSL does. X509_get_pubkey + free 27692 -> 9.8ns, against 9.6ns.
  • Give public key operations the shared global RNG instead of seeding one per call, where the instance locks itself or the build is single threaded. WOLFSSL_PK_LOCAL_RNG forces an RNG per call. EVP_DigestSign with ECDSA P-256 32291 -> 17008ns, against 18852ns.
  • Encode base64 without a call per output character, and hand whole blocks to Base64_Encode() rather than one 48 byte line at a time. EVP_EncodeUpdate 6446 -> 1855ns per 4KB. Skip the call into Base64_SkipNewline() for characters needing no skipping: EVP_DecodeUpdate 37059 -> 16194ns. What remains there is the constant time character decode, which is deliberate; wc_PemToDer() already picks the table decoder by PEM type.
  • Search cipher_tbl before the alias table in EVP_get_cipherbyname(), since over half the aliases only spell a name already there in another case, and compare without the platform's strcasecmp: 330 -> 67ns. The comparison is ASCII only, so it no longer follows the locale and works on platforms whose XSTRCASECMP is strcmp.
  • Find the PEM footer in a memory BIO's buffer and read up to it, rather than reading a byte at a time through the whole input. Loading a certificate from a file that carries a text dump alongside it 100311 -> 23301ns, against 33006ns.

Fixes:

  • AES-CCM through EVP now behaves as OpenSSL's does: the payload length comes from the first Update() carrying payload, AAD before the length is refused, and Final() has nothing left to do. Of 376 observables compared against OpenSSL - the tag and IV length matrices, the defaults, length inference over 48 sizes, and an encrypt/decrypt matrix with the tag then corrupted - 375 agree. The one that differs, EVP_CIPHER_iv_length() reporting CCM_NONCE_MIN_SZ where OpenSSL reports 12, is unchanged by this work.
  • EVP_get_cipherbyname(NULL) and EVP_get_digestbyname(NULL) dereferenced the name; OpenSSL returns NULL.
  • EVP_EncodeUpdate() wrote the block it was holding and then wrote over it, so a caller feeding data in pieces got wrong bytes, including bytes never written.
  • EVP_PKEY_keygen() on DH read the parameters after replacing the key object they came from, which is the same object when a caller passes its own key back in.
  • EVP_PKEY_set1_EC_KEY() with an empty EC_KEY now succeeds. It reported failure only as a side effect of encoding eagerly; OpenSSL returns 1 and the empty key is caught when the encoding is asked for.
  • Remove three GMULT macros naming functions that do not exist.

Portability:

  • Add XMEMCHR, with wc_memchr() for builds naming their own string functions through STRING_USER that have no memchr among them.

Tests:

  • Known answer vectors for the base64 encoders, generated independently of wolfSSL. testwolfcrypt passed with the encoder deliberately corrupted, as it only checked return codes.
  • EVP_EncodeUpdate() over every chunking of ten sizes straddling the 48 byte block, and PEM_read_bio_X509() reading several certificates from one BIO while leaving nothing behind after one.
  • EVP_get_cipherbyname() over canonical names, alternative spellings, every casing of each, and NULL.
  • One RSA private key operation per freshly decoded key: sharing a key lets the first operation provide the blinding RNG and hides a missing one in the rest. Plus EVP_PKEY_set1_DSA() followed by the caller releasing its key.
  • The reference counting behind the cached X509 public key: outliving the certificate and the caller, replacement, and duplication.
  • wc_memchr(), guarded by USE_WOLF_MEMCHR so it runs where that function is built.

Testing

Tested on ARM64 with SM algorithms.

Note: wolfsm fixes required:
wolfSSL/wolfsm#37

@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

MemBrowse Memory Report

gcc-arm-cortex-m0plus

  • FLASH: .text +132 B (+0.2%, 69,323 B / 262,144 B, total: 26% used)

gcc-arm-cortex-m3

  • FLASH: .text +144 B (+0.1%, 129,209 B / 262,144 B, total: 49% used)

gcc-arm-cortex-m4

  • FLASH: .text +448 B (+0.2%, 208,300 B / 262,144 B, total: 79% used)

gcc-arm-cortex-m4-baremetal

  • FLASH: .text +128 B (+0.2%, 71,715 B / 262,144 B, total: 27% used)

gcc-arm-cortex-m4-crypto-only

  • FLASH: .text +448 B (+0.2%, 181,341 B / 262,144 B, total: 69% used)

gcc-arm-cortex-m4-dtls13

  • FLASH: .text +128 B (+0.1%, 192,580 B / 1,048,576 B, total: 18% used)

gcc-arm-cortex-m4-min-ecc

  • FLASH: .text +128 B (+0.2%, 66,501 B / 262,144 B, total: 25% used)

gcc-arm-cortex-m4-openssl-compat

  • FLASH: .rodata +56 B, .text -576 B (-0.1%, 791,892 B / 1,048,576 B, total: 76% used)
  • RAM: .bss +16 B (+0.0%, 133,820 B / 262,144 B, total: 51% used)

gcc-arm-cortex-m4-pkcs7

  • FLASH: .text +448 B (+0.2%, 222,238 B / 262,144 B, total: 85% used)

gcc-arm-cortex-m4-pq

  • FLASH: .text +448 B (+0.1%, 309,104 B / 1,048,576 B, total: 29% used)

gcc-arm-cortex-m4-rsa-only

  • FLASH: .text +448 B (+0.1%, 338,928 B / 1,048,576 B, total: 32% used)

gcc-arm-cortex-m4-sp-math

  • FLASH: .text +128 B (+0.2%, 66,501 B / 262,144 B, total: 25% used)

gcc-arm-cortex-m4-tls12

  • FLASH: .text +128 B (+0.1%, 129,981 B / 262,144 B, total: 50% used)

gcc-arm-cortex-m4-tls13

  • FLASH: .text +448 B (+0.2%, 247,518 B / 262,144 B, total: 94% used)

gcc-arm-cortex-m7

  • FLASH: .text +448 B (+0.2%, 208,300 B / 262,144 B, total: 79% used)

gcc-arm-cortex-m7-pq

  • FLASH: .text +448 B (+0.1%, 310,000 B / 1,048,576 B, total: 30% used)

gcc-arm-cortex-m7-tls13

  • FLASH: .text +512 B (+0.2%, 247,582 B / 262,144 B, total: 94% used)

linuxkm-pie

  • Data: __patchable_function_entries +24 B (+0.1%, 28,688 B)

linuxkm-standard

  • Data: __patchable_function_entries +64 B (+0.1%, 51,664 B)

stm32-sim-stm32h753

  • FLASH: .text +512 B (+0.3%, 195,476 B / 2,097,152 B, total: 9% used)

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Portability failures, unsafe failure paths, RNG lifetime defects, and mutex initialization races remain unresolved.

10 open findings
What changed in this PR

Optimizes OpenSSL compatibility APIs while correcting EVP, PEM, CCM, Base64, key-lifecycle, and X509 caching behavior.

Changes:

  • Defers costly RNG and DER initialization and caches decoded X509 public keys.
  • Optimizes cipher/digest lookup, Base64 processing, and PEM BIO parsing.
  • Expands compatibility and regression coverage.
File Description
wolfssl/​wolfcrypt/​types.h Adds XMEMCHR portability support.
wolfssl/​wolfcrypt/​coding.h Declares streaming Base64 decoding.
wolfssl/​ssl.h Tracks deferred PKEY state.
wolfssl/​openssl/​sha3.h Corrects SHA3 holder sizing.
wolfssl/​openssl/​rsa.h Tracks deferred RSA RNG state.
wolfssl/​openssl/​evp.h Adds CCM message state.
wolfssl/​internal.h Adds cached X509 key support.
wolfcrypt/​src/​wc_port.c Implements wc_memchr.
wolfcrypt/​src/​evp_pk.c Integrates deferred DER generation.
wolfcrypt/​src/​coding.c Optimizes Base64 encoding and decoding.
wolfcrypt/​src/​evp.c Implements EVP optimizations and compatibility fixes.
wolfcrypt/​src/​aes.c Removes invalid GMULT macros.
tests/​api/​test_port.c Tests wc_memchr.
tests/​api/​test_ossl_x509_pk.h Registers X509 cache tests.
tests/​api/​test_ossl_x509_pk.c Tests cached-key reference handling.
tests/​api/​test_ossl_x509_io.h Registers PEM BIO tests.
tests/​api/​test_ossl_x509_io.c Tests multi-certificate BIO parsing.
tests/​api/​test_ossl_rsa.h Registers lazy-RNG tests.
tests/​api/​test_ossl_rsa.c Tests deferred RSA blinding RNGs.
tests/​api/​test_evp.h Registers chunked Base64 tests.
tests/​api/​test_evp.c Tests chunk-independent encoding.
tests/​api/​test_evp_pkey.h Registers deferred-DER tests.
tests/​api/​test_evp_pkey.c Tests lazy PKEY encoding and ownership.
tests/​api/​test_evp_digest.h Registers digest regression tests.
tests/​api/​test_evp_digest.c Tests digest context lookup and copying.
tests/​api/​test_evp_cipher.h Registers cipher and CCM tests.
tests/​api/​test_evp_cipher.c Tests CCM semantics and cipher lookup.
tests/​api/​test_coding.h Registers Base64 vectors.
tests/​api/​test_coding.c Adds Base64 known-answer tests.
src/​x509.c Caches keys and optimizes PEM parsing.
src/​ssl.c Enables shared RNG use where safe.
src/​ssl_p7p12.c Ensures DER before PKCS operations.
src/​ssl_load.c Ensures DER before loading keys.
src/​ssl_crypto.c Adds SHA3 size validation.
src/​pk.c Ensures DER before key serialization.
src/​pk_rsa.c Defers RSA blinding RNG creation.
src/​internal.c Releases cached X509 keys.

🧠 Review effort: Balanced


Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.

Comment thread src/pk_rsa.c Outdated
Comment thread src/ssl_crypto.c Outdated
Comment thread src/x509.c Outdated
Comment thread wolfcrypt/src/evp.c
Comment thread wolfcrypt/src/evp.c
Comment thread src/x509.c Outdated
Comment thread tests/api/test_evp_pkey.c Outdated
Comment thread tests/api/test_evp_cipher.c Outdated
Comment thread tests/api/test_evp_cipher.c Outdated
Comment thread tests/api/test_evp_cipher.c Outdated
@SparkiDev
SparkiDev force-pushed the openssl_compat_api_perf_fixes_1 branch from c923bc3 to 7602e88 Compare October 8, 2026 08:47
…ound

Performance:

 - Size WOLFSSL_SHA3_CTX's holder in pointers, not bytes. It reserved
   sizeof(wc_Sha3) pointers for a 440 byte object, so WOLFSSL_Hasher goes
   3520 -> 448 bytes and WOLFSSL_EVP_MD_CTX 3552 -> 944. A static assert now
   catches the same mistake.
 - Resolve digests and ciphers by table pointer before comparing names, and
   return the table's own pointer from the accessors. EVP_sha256()
   47.4 -> 1.2ns, EVP_MD_type 64.7 -> 4.3ns, EVP_CIPHER_nid 101.2 -> 8.1ns.
 - Copy only the fields past the hash union in EVP_MD_CTX_copy_ex, and leave
   the cipher union alone in EVP_CIPHER_CTX_init.
 - Set up an EVP_PKEY's RNG, an RSA key's blinding RNG and an EVP_PKEY's DER
   encoding on first use rather than at creation. Seeding a DRBG cost more
   than everything else about making a key put together.
   EVP_PKEY_new + free 13366 -> 34ns, RSA_new + free 13748 -> 230ns,
   EVP_PKEY_set1_RSA 13994 -> 45ns. Certificate parsing roughly halves as a
   result: d2i_X509 45978 -> 17198ns, X509_dup 45945 -> 17116ns, both now
   faster than OpenSSL. set1_DSA still encodes immediately, as WOLFSSL_DSA is
   not reference counted and the pkey only borrows the caller's key.
 - Keep the decoded public key on the X509 and hand out a reference, as
   OpenSSL does. X509_get_pubkey + free 27692 -> 9.8ns, against 9.6ns.
 - Give public key operations the shared global RNG instead of seeding one per
   call, where the instance locks itself or the build is single threaded.
   WOLFSSL_PK_LOCAL_RNG forces an RNG per call. EVP_DigestSign with ECDSA
   P-256 32291 -> 17008ns, against 18852ns.
 - Encode base64 without a call per output character, and hand whole blocks to
   Base64_Encode() rather than one 48 byte line at a time. EVP_EncodeUpdate
   6446 -> 1855ns per 4KB. Skip the call into Base64_SkipNewline() for
   characters needing no skipping: EVP_DecodeUpdate 37059 -> 16194ns. What
   remains there is the constant time character decode, which is deliberate;
   wc_PemToDer() already picks the table decoder by PEM type.
 - Search cipher_tbl before the alias table in EVP_get_cipherbyname(), since
   over half the aliases only spell a name already there in another case, and
   compare without the platform's strcasecmp: 330 -> 67ns. The comparison is
   ASCII only, so it no longer follows the locale and works on platforms whose
   XSTRCASECMP is strcmp.
 - Find the PEM footer in a memory BIO's buffer and read up to it, rather than
   reading a byte at a time through the whole input. Loading a certificate
   from a file that carries a text dump alongside it 100311 -> 23301ns,
   against 33006ns.

Fixes:

 - AES-CCM through EVP now behaves as OpenSSL's does: the payload length comes
   from the first Update() carrying payload, AAD before the length is refused,
   and Final() has nothing left to do. Of 376 observables compared against
   OpenSSL - the tag and IV length matrices, the defaults, length inference
   over 48 sizes, and an encrypt/decrypt matrix with the tag then corrupted -
   375 agree. The one that differs, EVP_CIPHER_iv_length() reporting
   CCM_NONCE_MIN_SZ where OpenSSL reports 12, is unchanged by this work.
 - EVP_get_cipherbyname(NULL) and EVP_get_digestbyname(NULL) dereferenced the
   name; OpenSSL returns NULL.
 - EVP_EncodeUpdate() wrote the block it was holding and then wrote over it,
   so a caller feeding data in pieces got wrong bytes, including bytes never
   written.
 - EVP_PKEY_keygen() on DH read the parameters after replacing the key object
   they came from, which is the same object when a caller passes its own key
   back in.
 - EVP_PKEY_set1_EC_KEY() with an empty EC_KEY now succeeds. It reported
   failure only as a side effect of encoding eagerly; OpenSSL returns 1 and
   the empty key is caught when the encoding is asked for.
 - Remove three GMULT macros naming functions that do not exist.

Portability:

 - Add XMEMCHR, with wc_memchr() for builds naming their own string functions
   through STRING_USER that have no memchr among them.

Tests:

 - Known answer vectors for the base64 encoders, generated independently of
   wolfSSL. testwolfcrypt passed with the encoder deliberately corrupted, as
   it only checked return codes.
 - EVP_EncodeUpdate() over every chunking of ten sizes straddling the 48 byte
   block, and PEM_read_bio_X509() reading several certificates from one BIO
   while leaving nothing behind after one.
 - EVP_get_cipherbyname() over canonical names, alternative spellings, every
   casing of each, and NULL.
 - One RSA private key operation per freshly decoded key: sharing a key lets
   the first operation provide the blinding RNG and hides a missing one in the
   rest. Plus EVP_PKEY_set1_DSA() followed by the caller releasing its key.
 - The reference counting behind the cached X509 public key: outliving the
   certificate and the caller, replacement, and duplication.
 - wc_memchr(), guarded by USE_WOLF_MEMCHR so it runs where that function is
   built.
@SparkiDev
SparkiDev force-pushed the openssl_compat_api_perf_fixes_1 branch from 7602e88 to 0591bab Compare October 8, 2026 09:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants