SharkSSL™ Embedded SSL/TLS Stack
SharkSSL Cortex-M Assembly (GCC)

These Thumb-2 sources target little-endian Cortex-M3, Cortex-M4, and Cortex-M7. Compile .S files with arm-none-eabi-gcc, not directly with as, so assembly and C see the same SharkSSL feature definitions. Use -mthumb and the appropriate -mcpu=cortex-m3, -mcpu=cortex-m4, or -mcpu=cortex-m7.

Big-integer assembly

Select exactly one of SharkSslBigInt_M3.S (Cortex-M3) or SharkSslBigInt_M4.S (Cortex-M4/M7 with DSP instructions). They export the same symbols and must not be linked together. Keep SharkSslBigInt.c in the normal SharkSSL build and set:

#define SHARKSSL_OPTIMIZED_BIGINT_ASM 1
#define SHARKSSL_BIGINT_WORDSIZE 32

X25519

SharkSslX25519_M4.S supplies the dedicated 8x32 field arithmetic and ladder for Cortex-M4/M7 with DSP instructions. It is separate from the generic BigInt assembly choice. Keep SharkSslBigInt.c in the build and set:

#define SHARKSSL_ECC_USE_CURVE25519 1
#define SHARKSSL_X25519_DEDICATED 1
#define SHARKSSL_X25519_ASM 1

The X25519 assembly backend is not available for Cortex-M3. The dedicated C or generic C path remains available when this assembly is not selected.

ChaCha20, Poly1305, and AES-GCM

SharkSslCrypto_M3.S contains the Thumb-2 crypto routines. Keep SharkSslCrypto.c in the build. Enable only the implementations needed:

#define SHARKSSL_USE_CHACHA20 1
#define SHARKSSL_OPTIMIZED_CHACHA_ASM 1
#define SHARKSSL_USE_POLY1305 1
#define SHARKSSL_OPTIMIZED_POLY1305_ASM 1

The same file also contains optional GHASH and combined GCM routines, guarded by SHARKSSL_OPTIMIZED_GHASH_ASM and SHARKSSL_OPTIMIZED_GCM_ASM. Those options require the corresponding AES-GCM configuration. There is no separate AES block-encryption assembly backend in this directory.

Target configuration

Define B_LITTLE_ENDIAN for these backends. Set SHARKSSL_UNALIGNED_ACCESS only if the target, memory regions, and integration permit such accesses. Choose cipher suites, code-size options, and a secure RNG in the active SharkSSL configuration.

Performance and benefit

These hand-written Thumb-2 routines use processor-specific arithmetic and register scheduling to target lower cryptographic processing cost on small MCUs. The potential benefit is more CPU time available for application tasks without changing SharkSSL's public APIs. Performance depends on the selected routines, compiler settings, and target MCU. Measure with the intended compiler and Cortex-M processor; ARM64 and x86 results do not apply here.