Paper 2026/2340

Optimizing Montgomery Arithmetic for RSA on Cortex-M0+ and Cortex-M3

Minoo Sim, Hansung University
Minwoo Lee, Hansung University
Seungwon Lee, Hansung University
SuBeen Cho, Hansung University
Jiwon Bang, Hansung University
Hwajeong Seo, Hansung University
Abstract

Cryptographic tokens that retain RSA credentials for compatibility with existing authentication systems require efficient private-key operations on small processors. Cortex-M0+ (M0+) lacks native widening multiplication, whereas Cortex-M3 (M3) provides it with operand-dependent latency. We optimize Montgomery multiplication and squaring using 15-bit limbs and integrate the kernels into RSA-2048 and RSA-3072 private operations based on the Chinese remainder theorem (CRT). On M0+, a wrapped sum and accumulated block quotients enable exact carry recovery. The squaring schedule integrates square and reduction columns without storing the complete square. On M3, we retain the established half-diagonal representation and double each buffer limb when it is first incorporated into Montgomery reduction, eliminating a separate buffer traversal. We establish accumulator bounds for exact Montgomery digit and carry recovery in both schedules. At operand widths of 1024 and 1536 bits, the M0+ representation and register mapping reduce Montgomery multiplication and squaring cycles by 12.60% to 13.11% against blocked two-word controls. The M3 schedule reduces Montgomery squaring cycles by 2.12% to 3.08% with half-diagonal initialization and triangular accumulation held fixed. Kernel replacement in otherwise unchanged RSA implementations reduces warm CRT cycles by 12.59% to 12.80% on M0+ and by 1.69% to 2.39% on M3. Our M3 Montgomery multiplication also uses 31.47% fewer cycles than a reproduced number-theoretic transform (NTT) implementation for 2048-bit operands with matched integers and canonical outputs.

Metadata
Available format(s)
PDF
Category
Implementation
Publication info
Preprint.
Keywords
RSAMontgomery arithmeticCortex-MEfficient implementation
Contact author(s)
minjoos9797 @ gmail com
minunejip @ gmail com
dkajdfhd1 @ gmail com
chosubin1208 @ gmail com
bang050922 @ gmail com
hwajeong84 @ gmail com
History
2026-10-06: approved
2026-10-05: received
See all versions
Short URL
https://ia.cr/2026/2340
License
No rights reserved
CC0

BibTeX

@misc{cryptoeprint:2026/2340,
      author = {Minoo Sim and Minwoo Lee and Seungwon Lee and SuBeen Cho and Jiwon Bang and Hwajeong Seo},
      title = {Optimizing Montgomery Arithmetic for {RSA} on Cortex-M0+ and Cortex-M3},
      howpublished = {Cryptology {ePrint} Archive, Paper 2026/2340},
      year = {2026},
      url = {https://eprint.iacr.org/2026/2340}
}
Note: In order to protect the privacy of readers, eprint.iacr.org does not use cookies or embedded third party content.