SharkSSL™ Embedded SSL/TLS Stack
SharkSSL crypto acceleration for 32-bit Windows

This plugin speeds up encryption, authentication, and X25519 key exchange in SharkSSL's Transport Layer Security (TLS) connections. It supports AES-GCM through AES instructions (AES-NI), and ChaCha20-Poly1305 through Advanced Vector Extensions 2 (AVX2). X25519 also uses AVX2, with ordinary 32-bit integer assembly as its fallback. SharkSSL checks processor and operating-system support before using vector acceleration. The cipher implementations fall back to C when unavailable.

Use these files for an MSVC Win32/x86 application. They are additions to an existing SharkSSL build, not a standalone TLS library. A 32-bit application can run on 64-bit Windows, but it still needs these 32-bit objects. The objects in X64_MSVC have a different calling convention and cannot replace them. This directory does not contain a VAES implementation.

Start with the standard build

Keep SharkSSL's shipped algorithm settings and add all seven files:

File Purpose
SharkSslCrypto_X86.asm AES-GCM encryption and authentication using AES-NI and carry-less multiplication
SharkSslCrypto_X86_AVX2.asm ChaCha20 encryption using AVX2
SharkSslCrypto_X86_Poly1305.asm Poly1305 authentication using AVX2
SharkSslCrypto_X86_CPUID.c Checks processor features and operating-system vector support
SharkSslX25519_X86.asm Integer X25519 fallback
SharkSslX25519_X86_AVX2.asm X25519 field arithmetic and scalar multiplication using AVX2
SharkSslX25519_X86_Adapter.c Connects the existing generic X25519 hook to this backend

In Visual Studio:

  1. Select Win32, enable masm under Build Customizations, and add the seven files to the project. Use Microsoft Macro Assembler as the item type for .asm files; compile the two .c files as C. Exclude these files from x64 builds. The adapter also needs SharkSSL's src and inc directories on its include path, in addition to your normal target include directories.
  2. Under C/C++ > Preprocessor > Preprocessor Definitions, enable the five switches below. Apply them to every configuration that builds SharkSSL, including the X25519 adapter.
  3. Under Microsoft Macro Assembler > Advanced, set Use Safe Exception Handlers to Yes (/safeseh) for all five assembler files. Keep the linker's SafeSEH setting enabled. These kernels install no exception handlers, and /safeseh supplies the object metadata required by a SafeSEH link.
  4. Leave MASM preprocessor definitions empty for the standard context layout. Rebuild and check that all seven objects appear in the link command.
/* Enable both cipher families and the optional X25519 backend. */
#define SHARKSSL_OPTIMIZED_GHASH_ASM 1
#define SHARKSSL_OPTIMIZED_GCM_ASM 1
#define SHARKSSL_OPTIMIZED_CHACHA_ASM 1
#define SHARKSSL_OPTIMIZED_POLY1305_ASM 1
#define SHARKSSL_X25519_ASM_HOOK 1

Enter these in Visual Studio as NAME=1, without #define. They are integer switches: 1 enables an optimization and 0 disables it. The four cipher switches default to 0 in configuration macros. Leave SHARKSSL_OPTIMIZED_GCM_VAES_ASM=0 for this target. The X25519 hook is undefined by default; undefined or 0 selects the existing C implementation.

The standard recipe keeps AES-128, AES-256, AES-GCM, ChaCha20 and Poly1305 enabled; AES-192 remains disabled and SHARKSSL_NOPACK=0. The optimization switches do not enable an algorithm that you disabled elsewhere. Keep the surrounding C code compiled for your oldest supported processor; globally enabling /arch:AVX or /arch:AVX2 can introduce instructions outside the guarded kernels.

For command-line assembly, run these commands from this directory in an x86 Native Tools Command Prompt. Use ml.exe, not ml64.exe:

rem Standard AES context layout; SafeSEH-compatible 32-bit objects.
ml /nologo /c /coff /safeseh SharkSslCrypto_X86.asm
ml /nologo /c /coff /safeseh SharkSslCrypto_X86_AVX2.asm
ml /nologo /c /coff /safeseh SharkSslCrypto_X86_Poly1305.asm
ml /nologo /c /coff /safeseh SharkSslX25519_X86.asm
ml /nologo /c /coff /safeseh SharkSslX25519_X86_AVX2.asm
cl /nologo /c /TC SharkSslCrypto_X86_CPUID.c
rem Standalone SharkSSL includes; BAS builds use their own TargConfig.h.
cl /nologo /c /TC /DSHARKSSL_X25519_ASM_HOOK=1 /I..\.. /I..\..\..\inc /I..\..\..\inc\arch\Windows SharkSslX25519_X86_Adapter.c

Add the resulting seven .obj files to your application link, and compile SharkSSL itself with the five C definitions above. This step does not replace your normal SharkSSL source files or their configuration.

Choose non-default settings when you need them

Include fewer acceleration components

Use fewer components to reduce code size or compare implementations. Set unused optimization switches explicitly to 0 when changing an existing project.

Build choice Switches set to 1 Files to include
All acceleration All five All seven files
AES-GCM only SHARKSSL_OPTIMIZED_GHASH_ASM, SHARKSSL_OPTIMIZED_GCM_ASM SharkSslCrypto_X86.asm, SharkSslCrypto_X86_CPUID.c
ChaCha20-Poly1305 only SHARKSSL_OPTIMIZED_CHACHA_ASM, SHARKSSL_OPTIMIZED_POLY1305_ASM SharkSslCrypto_X86_AVX2.asm, SharkSslCrypto_X86_Poly1305.asm, SharkSslCrypto_X86_CPUID.c
ChaCha20 only SHARKSSL_OPTIMIZED_CHACHA_ASM SharkSslCrypto_X86_AVX2.asm, SharkSslCrypto_X86_CPUID.c
Poly1305 only SHARKSSL_OPTIMIZED_POLY1305_ASM SharkSslCrypto_X86_Poly1305.asm, SharkSslCrypto_X86_CPUID.c
X25519 only SHARKSSL_X25519_ASM_HOOK SharkSslX25519_X86.asm, SharkSslX25519_X86_AVX2.asm, SharkSslX25519_X86_Adapter.c, SharkSslCrypto_X86_CPUID.c
Portable C only None None of the seven plugin files

Combine rows as needed; include the shared CPUID file only once. Enable the GCM and GHASH switches together. ChaCha20 and Poly1305 can be enabled independently. The other component continues to use C. These settings select local implementations; cipher-suite availability is configured separately.

Reduce the largest supported AES key size

Disabling larger AES keys can reduce the AES context size. The assembler must then use the same context layout as C. With AES-128 enabled and SHARKSSL_NOPACK=0, use:

Largest enabled key C changes from the defaults MASM definition for SharkSslCrypto_X86.asm
AES-256 None None; the default is SHARKSSL_GCM_M0=244
AES-192 SHARKSSL_USE_AES_256=0, SHARKSSL_USE_AES_192=1 SHARKSSL_GCM_M0=212
AES-128 SHARKSSL_USE_AES_256=0, keep AES-192 disabled SHARKSSL_GCM_M0=180

SHARKSSL_USE_AES_128, SHARKSSL_USE_AES_192, SHARKSSL_USE_AES_256, and SHARKSSL_NOPACK are integer C switches accepting 0 or 1. Their defaults are 1, 0, 1, and 0, respectively. Apply public context configuration consistently to SharkSSL and its callers.

SHARKSSL_GCM_M0 is a MASM byte offset, not a C enable switch. Set it in Microsoft Macro Assembler > General > Preprocessor Definitions, or pass it to ml. For example:

rem Also compile C with AES-256 and AES-192 disabled, and NOPACK=0.
ml /nologo /c /coff /safeseh /DSHARKSSL_GCM_M0=180 SharkSslCrypto_X86.asm

With SHARKSSL_NOPACK=1, the AES context retains the largest key schedule. Use the default 244 offset even if AES-256 is disabled. The ChaCha20, Poly1305, and X25519 files need no AES layout definition. Compile-time context checks and layout-specific linker symbols help reject incompatible C/assembler settings.

AES-192 is not used by SharkSSL's TLS suites. Its configuration is useful for applications using the separate crypto API. Disabling AES-256 removes its TLS suites and can change which peers you can communicate with.

Configure X25519

The x86 adapter requires these existing C settings:

Setting Required value Shipped default
SHARKSSL_USE_ECC 1 1
SHARKSSL_ECC_USE_CURVE25519 1 1
SHARKSSL_X25519_DEDICATED 1 1
SHARKSSL_BIGINT_WORDSIZE 32 bits 32
SHARKSSL_X25519_ASM Undefined or 0; the existing Cortex-M4 switch Undefined

Leave SHARKSSL_OPTIMIZED_BIGINT_ASM=0. This directory does not supply the generic BigInt assembly interface. X25519 uses SHARKSSL_X25519_ASM_HOOK=1 and the same single optional call in SharkSslBigInt_X25519_mult as x64. No additional common C changes or header definitions are required. Select only one platform adapter per build. The adapter checks the compiler, target, algorithms, and word size.

Enabling the hook selects AVX2 when the processor and OS support it; otherwise it selects the integer assembly backend. The adapter caches that decision using atomic accesses and reuses the shared CPUID helper. No additional application switch is needed. Disable the hook to compare against C.

The AVX2 backend uses ten alternating 26/25-bit limbs and computes four field products in parallel. The integer fallback uses eight 32-bit limbs and ordinary MUL, ADC, and SBB instructions, requiring no AVX, BMI2, or ADX. Public API clamping, opaque private-key format, and zero-secret rejection retain their existing behavior. The internal raw kernels expect the scalar format supplied by SharkSSL; applications should continue using the public API.

SHARKSSL_X25519_TEST is a presence-based MASM test-only definition exposing field helper symbols in either X25519 assembler file. Omit it in application builds. It does not enable the C hook and is not a runtime feature option.

Upgrade an existing x86 build

For an existing X25519 hook build, add SharkSslX25519_X86_AVX2.asm and rebuild the adapter. Include SharkSslCrypto_X86_CPUID.c if it is not already present for cipher acceleration. Keep the integer X25519 object for fallback. There are no new common C changes, header changes, or application definitions.

Poly1305 has moved out of SharkSslCrypto_X86_AVX2.asm. Add SharkSslCrypto_X86_Poly1305.asm to projects using Poly1305 acceleration and rebuild all plugin objects. Do not link the old combined AVX2 object alongside the new Poly1305 object: both would define the same symbol.

The optimized GHASH table now uses scaled powers in a reciprocal polynomial representation. Rebuild the complete AES-GCM assembly object and construct new contexts normally; do not mix old and new GHASH helpers or serialized internal contexts. The public context size and APIs have not changed.

A legacy GHASH-only build (SHARKSSL_OPTIMIZED_GHASH_ASM=1 and SHARKSSL_OPTIMIZED_GCM_ASM=0) also uses this assembly file, but has no runtime feature gate. It requires SSSE3 and PCLMULQDQ on the deployment CPU. Prefer the paired GCM/GHASH switches for automatic C fallback.

Understand cipher selection and fallback

The client offers supported cipher suites, and the server chooses a compatible suite. Each endpoint then chooses its own local implementation. Installing this plugin does not force the server to choose a particular cipher.

AES-GCM and ChaCha20-Poly1305 use different algorithms. Poly1305 is the authentication component of ChaCha20-Poly1305, so AES-NI or VAES does not replace it. On this 32-bit target, SharkSSL selects:

Negotiated cipher Implementation used locally
AES-GCM AES-NI when SSSE3, AES-NI and PCLMULQDQ are present; otherwise C
ChaCha20-Poly1305 AVX2 for eligible chunks when the CPU and OS support it; otherwise C
X25519 key exchange With the hook enabled: AVX2 when supported, otherwise integer assembly. With the hook disabled: C

The AVX2 check requires AVX, OSXSAVE, enabled XMM/YMM state in XCR0, and AVX2. It checks OS support before executing XGETBV. The AES-GCM check does not require AVX. The cipher dispatcher and X25519 adapter cache their feature results using atomic accesses.

The x86 dispatcher uses ChaCha20 AVX2 for every nonempty call when available. Its four-block kernel handles 256-byte batches; a register-resident single-block path handles short messages and remaining bytes. Poly1305 AVX2 still starts at 256 bytes, processes multiples of 64 bytes, and leaves remaining bytes to C. These are internal processing sizes, not minimum HTTP request sizes.

The same Win32 executable works without AES-NI or AVX2, provided the rest of the application supports that processor and Windows version. Fallback keeps the negotiated cipher unchanged. It does not convert a Windows binary into an ARM, Linux, or 64-bit application; those targets need their own build.

Applications continue using SharkSSL's normal TLS APIs. The internal assembler entry points have specific context and buffer requirements and should not be called directly by application code.

Performance and benefit

On an Intel Core i7-1165G7, optimized MSVC 14.51 Win32 builds running under WOW64 achieved about 2.7x the C throughput for internal X25519 multiplication using AVX2. Short ChaCha20 calls (16-255 bytes) were 1.3-1.6x faster than C processing of the same message sizes. This reduces CPU time spent on key exchange and small encrypted messages while retaining a 32-bit application build.

These operation benchmarks require AVX2 and OS support; they do not measure complete TLS or server throughput. Results vary with compiler settings, workload, processor, and fallback implementation.