|
SharkSSL™ Embedded SSL/TLS Stack
|
This plugin speeds up the encryption, authentication, and X25519 key exchange used by SharkSSL's Transport Layer Security (TLS) connections. It supports AES-GCM through AES instructions (AES-NI) and optional vector AES (VAES), ChaCha20-Poly1305 through Advanced Vector Extensions 2 (AVX2), and X25519 through BMI2/ADX integer instructions. With the standard settings below, SharkSSL checks processor and operating-system support and uses C when acceleration is unavailable.
Use these files for an MSVC x64 application. They are additions to an existing SharkSSL build, not a standalone TLS library. Win32/x86 uses the separate `X86_MSVC` implementation. ARM64, ARM64EC, and System V calling conventions are not supported by these objects.
Keep SharkSSL's shipped algorithm settings and add all eight files:
| File | Purpose |
|---|---|
SharkSslCrypto_X64.asm | AES-GCM encryption and GHASH authentication using AES-NI and carry-less multiplication |
SharkSslCrypto_X64_VAES.asm | Optional 256-bit VAES/VPCLMULQDQ AES-GCM kernel |
SharkSslCrypto_X64_AVX2.asm | ChaCha20 encryption using AVX2 |
SharkSslCrypto_X64_Poly1305.asm | Poly1305 authentication using AVX2 |
SharkSslCrypto_X64_CPUID.c | Processor and operating-system checks shared by AES-NI, ChaCha20, and Poly1305 |
SharkSslCrypto_X64_VAES_CPUID.c | Processor and operating-system checks for the VAES kernel |
SharkSslX25519_X64.asm | X25519 field arithmetic, scalar multiplication, and BMI2/ADX detection |
SharkSslX25519_X64_Adapter.c | Connects the common X25519 hook to the assembly and preserves C fallback |
In Visual Studio:
.asm files; compile the three .c files as C. Exclude these files from Win32 builds.src and inc directories so it finds SharkSslBigInt.h and SharkSSL.h. Use the same target configuration as the rest of SharkSSL.ml64.exe and its unwind metadata; the Win32 /safeseh recipe does not apply.Enter these in Visual Studio as NAME=1, without #define. They are integer switches: 1 enables an optimization and 0 disables it. The five SHARKSSL_OPTIMIZED_* switches default to 0 in configuration macros. The hook is undefined by default; undefined or 0 disables it without requiring an adapter or assembly object.
The standard recipe keeps AES-128, AES-256, AES-GCM, ChaCha20, Poly1305, ECC, and dedicated Curve25519 enabled. AES-192 remains disabled and SHARKSSL_NOPACK=0. Optimization switches do not enable algorithms disabled elsewhere. Keep the surrounding C code compiled for your oldest supported x64 processor; globally enabling /arch:AVX, /arch:AVX2, or newer instruction sets can introduce instructions outside the guarded kernels.
For command-line assembly, run these commands from this directory in an x64 Native Tools Command Prompt. Use ml64.exe, not ml.exe:
Add the resulting eight .obj files to the application link, and compile SharkSSL itself with the six definitions above. Keep the adapter's algorithm and context configuration consistent with that build. MASM does not read SharkSSL_cfg.h; C enable switches belong in the C compiler settings.
Use fewer components to reduce code size or compare implementations. Set unused optimization switches explicitly to 0 when changing an existing project. Include a shared CPUID file only once.
| Build choice | Switches set to 1 | Files to include |
|---|---|---|
| All acceleration | All six | All eight files |
| AES-GCM with AES-NI | SHARKSSL_OPTIMIZED_GHASH_ASM, SHARKSSL_OPTIMIZED_GCM_ASM | SharkSslCrypto_X64.asm, SharkSslCrypto_X64_CPUID.c |
| AES-GCM with VAES and AES-NI fallback | The two AES-NI switches plus SHARKSSL_OPTIMIZED_GCM_VAES_ASM | AES-NI files plus SharkSslCrypto_X64_VAES.asm, SharkSslCrypto_X64_VAES_CPUID.c |
| ChaCha20-Poly1305 | SHARKSSL_OPTIMIZED_CHACHA_ASM, SHARKSSL_OPTIMIZED_POLY1305_ASM | SharkSslCrypto_X64_AVX2.asm, SharkSslCrypto_X64_Poly1305.asm, SharkSslCrypto_X64_CPUID.c |
| ChaCha20 only | SHARKSSL_OPTIMIZED_CHACHA_ASM | SharkSslCrypto_X64_AVX2.asm, SharkSslCrypto_X64_CPUID.c |
| Poly1305 only | SHARKSSL_OPTIMIZED_POLY1305_ASM | SharkSslCrypto_X64_Poly1305.asm, SharkSslCrypto_X64_CPUID.c |
| X25519 only | SHARKSSL_X25519_ASM_HOOK | SharkSslX25519_X64.asm, SharkSslX25519_X64_Adapter.c |
| Portable C only | None | None of the eight plugin files |
Combine rows as needed. Enable the GCM and GHASH switches together. GCM assembly without GHASH assembly is rejected on Microsoft x64. VAES requires both of those switches and both AES-NI files; it is an additional backend, not a replacement for the fallback files. Enabling VAES alone does not select its dispatcher.
A legacy GHASH-only setting (SHARKSSL_OPTIMIZED_GHASH_ASM=1 with GCM assembly and VAES disabled) uses SharkSslCrypto_X64.asm for GHASH. That path does not perform the combined kernel's runtime feature check. It requires a deployment CPU with SSSE3 and PCLMULQDQ; use the paired GCM/GHASH settings for automatic fallback. Its MASM context offset must still match C.
These settings select local implementations; cipher-suite availability is configured separately. The AES-GCM kernels do not accelerate every other AES mode or replace the general AES key-schedule implementation.
Disabling larger AES keys can reduce the AES context size. Both AES-GCM assembler files must then use the same context layout as C. With AES-128 enabled and SHARKSSL_NOPACK=0, use:
| Largest enabled key | C changes from the defaults | MASM definition for both AES-GCM files |
|---|---|---|
| AES-256 | None | None; the default is SHARKSSL_GCM_M0=244 |
| AES-192 | SHARKSSL_USE_AES_256=0, SHARKSSL_USE_AES_192=1 | SHARKSSL_GCM_M0=212 |
| AES-128 | SHARKSSL_USE_AES_256=0, keep AES-192 disabled | SHARKSSL_GCM_M0=180 |
SHARKSSL_USE_AES_128, SHARKSSL_USE_AES_192, SHARKSSL_USE_AES_256, and SHARKSSL_NOPACK are integer C switches accepting 0 or 1. Their defaults are 1, 0, 1, and 0, respectively. Apply public context configuration consistently to SharkSSL and its callers.
SHARKSSL_GCM_M0 is a MASM byte offset, not a C enable switch. Set it in Microsoft Macro Assembler > General > Preprocessor Definitions, or pass it to ml64. For example:
With SHARKSSL_NOPACK=1, the AES context retains the largest key schedule. Use the default 244 offset even if AES-256 is disabled. SHARKSSL_AES_NR is derived internally as SHARKSSL_GCM_M0-4; do not configure it separately. The ChaCha20, Poly1305, and X25519 files need no AES layout definition. Compile-time context checks and layout-specific linker symbols help reject incompatible C/assembler settings in the combined GCM configuration.
AES-192 is not used by SharkSSL's TLS suites. Its configuration is useful for applications using the separate crypto API. Disabling AES-256 removes its TLS suites and can change which peers you can communicate with.
The x64 adapter requires the following existing C settings:
| Setting | Required value | Shipped default |
|---|---|---|
SHARKSSL_USE_ECC | 1 | 1 |
SHARKSSL_ECC_USE_CURVE25519 | 1 | 1 |
SHARKSSL_X25519_DEDICATED | 1 | 1 |
SHARKSSL_BIGINT_WORDSIZE | 32 bits | 32 |
SHARKSSL_X25519_ASM | Undefined or 0; this selects the existing Cortex-M4 backend | Undefined |
Leave SHARKSSL_OPTIMIZED_BIGINT_ASM=0 for this x64 build. It controls the existing generic BigInt assembly interface; this directory does not implement that interface. It does not enable the X25519 hook.
The common C integration is the single optional SharkSslX25519_tryAsm call in SharkSslBigInt_X25519_mult. The adapter handles representation conversion, CPU checks, and its scratch cleanup. A CPU without BMI2/ADX continues through the existing C implementation. Use SHARKSSL_X25519_ASM_HOOK to select this integration; no additional platform-specific C definitions are required.
The optional MASM definition SHARKSSL_X25519_TEST exports internal arithmetic entry points for the regression harness. It is an IFDEF switch: defining it, even as 0, enables those test exports. Leave it undefined in application builds. Applications should use the normal SharkSSL APIs, not the raw assembly entries.
The client offers supported cipher suites, and the server chooses a compatible suite. Each endpoint then chooses its own local implementation. Installing this plugin does not force a particular cipher or key-exchange group. X25519 is key exchange, independent of whether TLS uses AES-GCM or ChaCha20-Poly1305 for records.
With the standard switches enabled, SharkSSL selects:
| Operation | Implementation used locally |
|---|---|
| AES-GCM | VAES/VPCLMULQDQ when available; otherwise AES-NI/PCLMULQDQ; otherwise C |
| ChaCha20-Poly1305 | AVX2 when available; C for unsupported CPUs and Poly1305 tails |
| X25519 | BMI2/ADX when both are available; otherwise the existing dedicated C implementation |
AES-NI GCM requires SSSE3, AES-NI, and PCLMULQDQ. It does not require AVX. The AVX2 check requires AVX, OSXSAVE, enabled XMM/YMM state in XCR0, and AVX2. The VAES check additionally requires VAES, VPCLMULQDQ, and the AES-NI GCM features used by its tails. These probes check OS support before executing XGETBV. The VAES kernel uses 256-bit vectors and does not require AVX-512 state.
Symmetric feature results are cached by the shared C dispatcher using atomic accesses. The X25519 adapter checks BMI2/ADX for each call; it uses integer instructions and baseline x64 SSE2, so it needs no AVX OS-state check.
The x64 dispatcher offers every nonempty ChaCha20 call to AVX2 when available. The kernel uses eight-block batches for bulk data and a register-resident single-block path for short messages and remaining bytes. Poly1305 AVX2 still starts at 256 bytes and leaves remaining bytes to C. These are internal processing sizes, not minimum HTTP request sizes. The AES-GCM dispatcher sends complete 16-byte blocks to its selected kernel and retains C processing for remaining bytes.
The same x64 executable works without AES-NI, VAES, AVX2, or BMI2/ADX, provided the rest of the application supports that processor and Windows version and the standard runtime-dispatched settings are used. Fallback preserves the negotiated algorithms. It does not make these objects suitable for Win32, ARM, or Linux.
On an Intel Core i7-1165G7 with optimized MSVC 14.51 builds, X25519 shared-secret calculation took 25.15 us with assembly versus 68.77 us with C, about 2.7x the throughput. Short ChaCha20 calls (16-511 bytes) were 1.3-1.5x faster than C processing of the same message sizes. These savings leave more CPU time for application work and reduce the cryptographic cost of key exchange and small messages.
These figures describe individual operations in the benchmark builds, not whole-server throughput. Gains require the relevant CPU features and vary with compiler settings, workload, and API integration.