WiP: Towards a Secure SECP256K1 for Crypto Wallets: Hardware Architecture and Implementation

arXiv:2411.03910 · cs.CR, eess.SP · Submitted 2024-11-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "WiP: Towards a Secure SECP256K1 for Crypto Wallets: Hardware Architecture and Implementation".

Jane: The paper was written by Joel Poncha Lemayian, Ghyslain Gagnon, Kaiwen Zhang and Pascal Giard from Department of Electrical Engineering, École de technologie supérieure (ÉTS), Montréal, Canada and Department of Software Engineering and IT, École de technologie supérieure (ÉTS), Montréal, Canada.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Paper discussion segment 1: Tom: We're looking at "WiP: Towards a Secure SECP256K1 for Crypto Wallets: Hardware Architecture and Implementation" from a group at École de technologie supérieure. Jane, these authors are coming at this from a heavy electrical engineering background, which makes sense given the focus on hardware.

Jane: It really does, Tom. They're tackling SECP256K1, which is basically the math engine that powers Bitcoin and Ethereum to prove you own your digital assets. If someone steals your private key through a side-channel attack—which is like listening to the electrical hum of a chip to guess its secrets—you lose everything.

Tom: Right, so they aren't just writing software; they are designing actual physical circuit layouts. Lu, when you see hardware-level security like this, what does your mind jump to?

Lu: I see the potential for a massive shift in how we trust decentralized systems. If we can bake this level of protection directly into the silicon, we're creating a foundation for autonomous agents that can hold value without needing a human to guard them constantly. It’s about making the hardware itself an incorruptible witness to truth.

Meng: That sounds great in theory, Lu, but I wonder how much extra cost this adds to a standard chip. The paper mentions optimization and resource utilization, which tells me they are worried about the physical footprint on the silicon die. If it's too big or too expensive, manufacturers won't use it in small consumer wallets.

Jane: That’s exactly why their approach is so focused on "resource efficiency," as they put it. They want to make sure a tiny, portable device can run these heavy math operations without draining the battery or needing a massive processor.

Lalam: Building that reliability into the hardware could fundamentally change how society views digital sovereignty. If people feel their assets are physically protected by the laws of mathematics and silicon, they will trust digital ecosystems much more deeply. This creates a cultural shift from trusting institutions to trusting verifiable physical hardware.

Tom: It’s a high-stakes game, for sure. Let's talk about what they actually did to solve these security leaks in the next part.

Paper discussion segment 2: Jane: Now that we know the stakes, let's look at how they actually built this thing in "WiP: Towards a Secure SECP256K1 for Crypto Wallets: Hardware Architecture and Implementation." They're using something called the Montgomery Ladder algorithm to perform point multiplication.

Tom: Yeah, and they realized that standard ways of doing this math have "branches" in the code. If a chip does one thing when a bit is one and another when it's zero an attacker can just measure the power spikes to read your private key. It’s like watching someone type by listening to the different sounds of each key hitting the board.

Jane: To fix that, they used "complete addition formulas" and projective coordinates. Instead of having different paths for different math operations, they made every step look almost identical to an outside observer.

Meng: I'm looking at their implementation details here, and the way they use a temporary register—that Rt register mentioned in Algorithm three—is clever. It keeps the power consumption steady because you're always performing similar loads and stores regardless of whether the bit is a zero or a one. It’s a very grounded way to hide those signal variations.

Lu: It reminds me of how we try to mask noise in neural network training to prevent adversarial attacks! They are essentially creating "mathematical camouflage" at the hardware level. If you can't see the difference between an addition and a doubling, you can't steal the key.

Tom: And they aren't just doing it for show; they actually measured how much space this takes up on an FPGA chip.

Lalam: The elegance of this design lies in its ability to provide high-level security without requiring massive amounts of extra memory or processing power. It suggests a future where security isn't a heavy layer we add on top, but something that is naturally inherent to the way hardware functions.

Jane: They even used a Binary Inversion Algorithm at the end to turn those complicated coordinates back into something a wallet can actually use. But how much better is this than what's already out there?

Paper discussion segment 3: Tom: This is where the numbers come in, and they are pretty impressive. The authors compared their work against several other implementations in Table II, and they saw a massive reduction in LUT usage—that's Look-Up Tables, the basic building blocks of these chips.

Jane: They achieved an average reduction of forty-five percent in LUT usage compared to similar works. That’s huge because it means you can fit this security module into much smaller, cheaper, and more power-efficient hardware.

Meng: I noticed they didn't use any DSP blocks or RAM for their core implementation, which is a massive win for area efficiency. Most of the other papers in the table, like Asif et al., are using hundreds or even thousands of DSPs. Using only LUTs and registers makes this much more scalable for mass-produced consumer electronics.

Lu: Even though they saved space, they still managed to hit a frequency of two hundred fifty MHz on the Xilinx ZCU104! That's incredibly fast for something so lean. Imagine a world where every single microchip has this level of built-in cryptographic resilience without any noticeable performance hit.

Tom: It really does change the math on what you can fit into a tiny device. If you don't need to waste space on massive memory blocks, you have more room for other features or just a smaller device overall.

Jane: They also addressed the latency and throughput, showing that their design is competitive even when it's being highly efficient. It’s not just about being small; it’s about being fast enough to be useful in real-time transactions.

Lalam: This efficiency could lead to a democratization of security, where high-level protection isn't a luxury for expensive hardware but a standard feature for everyone. It bridges the gap between high-end cryptographic research and the everyday devices we carry in our pockets.

Tom: They've really hit a sweet spot here. Let’s wrap this up and see what the big picture looks like.

Conclusion: Jane: We've covered a lot of ground today, from how SECP256K1 math works to how this specific hardware architecture uses complete addition formulas and temporary registers to stop side-channel attacks. This paper, "WiP: Towards a Secure SECP256K1 for Crypto Wallets: Hardware Architecture and Implementation," really shows how you can balance security, speed, and size.

Tom: It’s a solid piece of work that moves the needle for hardware wallets. They've proven you don't need massive amounts of power or space to build something incredibly secure.

Lu: I think this is a blueprint for the next generation of secure silicon! It makes me wonder how much more we can hide in the physical properties of a chip.

Meng: From my side, it’s a practical win for anyone designing consumer hardware. This is exactly the kind of optimization that makes new tech viable for the real world.

Lalam: Ultimately, this research strengthens the foundation of digital trust by making security an invisible, efficient part of our physical world. It's a step toward a more secure and seamless digital culture.

Tom: Thanks for joining us on the show! We'll see you next time when we tackle another fascinating paper from the arXiv. Goodbye everyone!

Jane: Bye!

Lu: See ya!

Meng: Take care!

Lalam: Goodbye!--- END OF SCRIPT ------thought

Department of Electrical Engineering, École de technologie supérieure (ÉTS), Montréal, Canada · Department of Software Engineering and IT, École de technologie supérieure (ÉTS), Montréal, Canada

cs.CR, eess.SP

Submitted: 2024-11-06

Updated: 2024-11-06

Comments: Presented at HASP 2024 @ MICRO 2024 https://haspworkshop.org/2024/program.html

Journal ref: Int. Workshop on Hardware and Archit. Support for Security and Privacy (HASP) (2024)

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 71/100

The gist: "The SECP256K1 elliptic curve algorithm is fundamental in cryptocurrency wallets for generating secure public keys from private keys, thereby ensuring the protection and ownership of blockchain-based

Terminology

Summary

"The SECP256K1 elliptic curve algorithm is fundamental in cryptocurrency wallets for generating secure public keys from private keys, thereby ensuring the protection and ownership of blockchain-based digital assets. However, the literature highlights several successful side-channel attacks on hardware wallets that exploit SECP256K1 to extract private keys. This work proposes a novel hardware architecture for SECP256K1, optimized for side-channel attack resistance and efficient resource utilization. The architecture incorporates complete addition formulas, temporary registers, and parallel processing techniques, making elliptic curve point addition and doubling operations indistinguishable. Implementation results demonstrate an average reduction of 45% in LUT usage compared to similar works, emphasizing the design’s resource efficiency."

"This work proposes a hardware architecture and implementation of a SECP256K1 module for crypto wallets that are secure against SCA attacks while utilizing minimum resources hence adhering to the industry standard for small, compact, and portable wallets. The module employs temporary registers and parallel processing to prevent variations during ECPM. Moreover, the complete addition formulas are used to prevent timing variability. The proposed architecture comprises two main parts, SECP256K1 ECPM and BIA [Binary Inversion Algorithm]. ECPM is accomplished by employing ECPA and the ECPD in the Montgomery Ladder algorithm. This work uses the equations in Alg. 2 to design ECPA. The Montgomery Ladder algorithm employing the ECPM architecture is shown in Alg. 3, which includes a temporary register file Rt added to prevent variability and maintain uniformity in both branches when processing ki = 1 and ki = 0. It also utilizes two ECPA modules that run in parallel."

"The output of the Montgomery Ladder R̂ must be converted to affine coordinates by calculating the modular multiplicative inverse of the value in the z-axis and multiplying it with the x- and y-axis values. This work employs the binary inversion algorithm (BIA) to calculate r. The architecture also reuses registers R1 and Rt shown by the rugged red region."

"Implementation results indicate that the proposed architecture requires, on average, 45% less LUTs compared to analogous implementations in literature. Our implementation stands out for its minimal LUT count, surpassed only by [14]. However, [14] requires significantly more DSPs and RAM blocks, highlighting a critical efficiency trade-off. Unlike other designs such as [16] and [18]—which also avoid DSPs but demand higher LUTs—our approach efficiently manages operations with fewer resources. The lack of RAM and minimal register usage in our design further underscores its suitability for resource-constrained, low-power applications like crypto wallets."

"A Xilinx ZCU104 and an Artix-7 field programmable gate array (FPGA) boards were used to implement the proposed SECP256K1 architecture. Moreover, Vivado 2022.2 was used for simulation and a reference software implementation was used to verify the output." "Implementation results: LUT: 21 (Zynq-US), 24 (Artix-7); DSP: 0; Area RAM (kbits): 0; Registers: 13,881 (Zynq-US), 13,385 (Artix-7); Frequency: 250 MHz (Zynq-US), 90 MHz (Artix-7); Latency: 7.58 ms / 1,895 kCC (Zynq-US), 21 ms / 1,895 kCC (Artix-7); Throughput: 34 kbps (Zynq-US), 12 kbps (Artix-7)." "The proposed hardware architecture for the SECP256K1 algorithm enhances resistance against SCA attacks by maintaining uniformity in register operations during the private key processing. Moreover, the architecture is designed for crypto wallet application, where it utilizes minimum resources and adheres to the industry standard for small, portable crypto wallets."

Improvements for AI systems

To improve AI systems—specifically those integrated into hardware security modules (HSMs), edge computing devices, and IoT-based cryptographic processors—using the findings of this paper, I propose the following specific improvements:

I. Integration of Side-Channel Resistant Hardware Primitives into AI Acceleration Architectures

The paper demonstrates that using complete addition formulas in projective coordinates and a modified Montgomery Ladder (using temporary registers to ensure uniform execution paths) significantly reduces side-channel leakage while maintaining low resource footprints.

  1. Improvement: Implement the proposed Uniform Execution Path logic within the hardware-level cryptographic accelerators used by AI edge devices for secure model weight loading and identity verification.

  2. What the improved AI system can do: The AI system will be able to perform high-speed, secure inference on untrusted hardware (like mobile phones or IoT sensors) without leaking its private authentication keys through power consumption or timing variations during the SECP256K1 handshake/signature process.

II. Resource-Optimized Cryptographic Co-processors for TinyML Deployment

The paper achieves a 45% reduction in Look-Up Table (LUT) usage compared to existing literature, making it ideal for extremely resource-constrained environments.

  1. Improvement: Incorporate this specific low-area SECP256K1 hardware architecture into the SoC (System on Chip) design of AI microcontrollers used in TinyML applications.

  2. What the improved AI system can do: It allows ultra-low-power, battery-operated AI sensors (e.g., smart medical implants or remote environmental monitors) to perform blockchain-based data integrity verification and secure identity management locally, without needing a high-power processor or external security chip, thereby extending device battery life significantly.

III. Secure Federated Learning via Hardware-Rooted Identity

Federated learning requires devices to prove their identity and sign model updates without exposing the private keys used for authentication.

  1. Improvement: Use the proposed architecture as a Hardware Root of Trust (RoT) specifically optimized for the SECP256K1 curve to sign local gradient updates in a Federated Learning protocol.

  2. What the improved AI system can do: It ensures that an attacker cannot perform a Sybil attack or Model Poisoning attack by intercepting power signatures of the device during the signing process; it allows for highly secure, decentralized model training across millions of edge devices where each device's cryptographic identity is physically protected by the proposed hardware logic.

Abstract

The SECP256K1 elliptic curve algorithm is fundamental in cryptocurrency wallets for generating secure public keys from private keys, thereby ensuring the protection and ownership of blockchain-based digital assets. However, the literature highlights several successful side-channel attacks on hardware wallets that exploit SECP256K1 to extract private keys. This work proposes a novel hardware architecture for SECP256K1, optimized for side-channel attack resistance and efficient resource utilization. The architecture incorporates complete addition formulas, temporary registers, and parallel processing techniques, making elliptic curve point addition and doubling operations indistinguishable. Implementation results demonstrate an average reduction of 45% in LUT usage compared to similar works, emphasizing the design's resource efficiency.

Related papers