Paper 2026/344
Area-Efficient LUT-Based Multipliers for AMD Versal FPGAs
Abstract
AMD Versal FPGAs introduce a new CLB micro-architecture featuring the LOOKAHEAD8 carry structure in place of the legacy CARRY4/8 chains, on which existing area-efficient LUT-based multiplier designs map inefficiently. This paper proposes a LUT-based integer multiplier architecture tailored to the Versal fabric. By jointly exploiting radix-4 modified Booth recoding and the new Versal LUT micro-architecture, only $\mathtt{\sim}n^{2}/4$ LUTs are required to generate the partial-product bit heap for an $n$-bit multiplication. A new heuristic for compressor-tree synthesis further improves the area--delay product by 8--20% over state-of-the-art Versal heuristics. Overall, the proposed multipliers achieve up to 40% LUT reduction relative to AMD LogiCORE IP multipliers at comparable critical-path delay. An open-source Python RTL generator with configurable operand widths and pipeline depths is provided for scalable deployment.
Metadata
- Available format(s)
-
PDF
- Category
- Implementation
- Publication info
- Published elsewhere. Minor revision. 33rd IEEE International Symposium on Computer Arithmetic
- Keywords
- LUT-based multiplierradix-4 Booth recodingcompressor tree synthesisFPGAVersal
- Contact author(s)
-
zetao miao @ kuleuven be
xander pottier @ kuleuven be
jonas bertels @ kuleuven be
wouter legiest @ kuleuven be
ingrid verbauwhede @ kuleuven be - History
- 2026-05-21: revised
- 2026-02-20: received
- See all versions
- Short URL
- https://ia.cr/2026/344
- License
-
CC BY-NC-ND
BibTeX
@misc{cryptoeprint:2026/344,
author = {Zetao Miao and Xander Pottier and Jonas Bertels and Wouter Legiest and Ingrid Verbauwhede},
title = {Area-Efficient {LUT}-Based Multipliers for {AMD} Versal {FPGAs}},
howpublished = {Cryptology {ePrint} Archive, Paper 2026/344},
year = {2026},
url = {https://eprint.iacr.org/2026/344}
}