Accelerating HQC for Post-Quantum TLS 1.3 on x86 IoT Gateways

summary

Video file (mp4)

The gist

The gist: This work optimizes HQC for x86 processors using AVX2 and AVX-512 with GFNI instructions and integrates these implementations into TLS 1.3 to minimize computational costs while evaluating

In short

The work optimized HQC for TLS 1.3 using AVX2 and AVX-512 instructions with GFNI extensions on x86 processors. These optimizations significantly reduced computational costs for key generation, encapsulation, and decapsulation compared to previous methods. Evaluations on constrained networks showed that while implementation speed improved, communication sizes of HQC determined the total handshake time.

Key concepts

HQC
Hybrid Key Encapsulation is a method used in Post-Quantum cryptography to establish secure connections for TLS 1.3. It combines traditional key exchange methods with new quantum-resistant algorithms to ensure security against future quantum computers.
AVX2 and AVX-512
These are instruction set extensions for x86 processors that allow CPUs to perform many mathematical operations simultaneously on large sets of data at once. AVX-512 provides even greater parallelism, which the authors used to speed up complex cryptographic tasks like polynomial multiplication and decoding.
GFNI
GFNI stands for Galois Field Native Instructions. These are specialized instructions that allow the CPU to perform arithmetic directly within finite fields, which is crucial for efficiently handling the complex mathematical operations required by HQC algorithms.

Terminology used across episodes

This episode discusses

The paper

Accelerating HQC for Post-Quantum TLS 1.3 on x86 IoT Gateways · Read on arXiv

Jihoon Jang, Hyunju Park, Jebin Kim, Seokhie Hong, Suhri Kim

School of Cyber Security, Korea University · Sungshin Women’s University

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Accelerating HQC for Post-Quantum TLS 1.3 on x86 IoT Gateways".

Elias: The gist:

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So we’re looking at a paper today called "Accelerating HQC for Post-Quantum TLS one point three on xeighty-six IoT Gateways <ref:2610.11107#pg1,Accelerating HQC for Post-Quantum TLS 1.3 on x86 IoT Gateways>." It’s about making post-quantum cryptography work faster on those small devices we use every day.

Elias: Yeah, it tackles the problem that HQC, which is this specific post-quantum key exchange method, can be slow or use too much CPU when you try to run it in a TLS one point three handshake on hardware like xeighty-six processors <ref:2610.11107#pg1>.

Nadia: Exactly, and the authors are showing how they can use AVXtwo and AVX-five hundred twelve instructions alongside Galois Field New Instructions GFNI to speed things up significantly <ref:2610.11107#pg1,Galois Field New Instructions GFNI>.

Priya: For someone just listening, it means that when you use this for secure connections on a device, the initial setup time should be much better than what we’re used to right now.

Nadia: The paper explains that HQC has these big computational costs related to polynomial multiplication and decoding, and they are attacking those parts directly.

Elias: Specifically, they extend branch-free Toom-Cook or Karatsuba multiplication to AVX-five hundred twelve which helps with the key generation part of the process <ref:2610.11107#pg1,extend branch-free Toom-Cook>.

Priya: And the data shows that this optimization gives them gains in how fast they can generate keys and handle those initial setup steps.

Nadia: Right, and they also tackled decoding, which is a big part of running these schemes, by using GFNI arithmetic for both AVXtwo and AVX-five hundred twelve across all five phases of the Reed–Solomon decoding <ref:2610.11107#pg1>.

Elias: They also improved the Reed–Muller decoding by changing how it finds the maximum in those peaks, switching to a horizontal vector reduction instead of a binary search, which cuts down the full decoder cost by about eighteen percent to twenty percent <ref:2610.11107#pg2,the full decoder cost by>.

Priya: So what this means for us on constrained devices is that error correction during decapsulation should happen much more efficiently when using these optimized methods.

Nadia: The results they show are pretty big here, especially when you look at the AVX-five hundred twelve implementation, which reportedly reduces key generation by twenty-four point nine to twenty-nine point five percent compared to prior work across those three HQC parameter sets <ref:2610.11107#pg1,reduces key generation by 24.9>.

Elias: And that same AVX-five hundred twelve implementation for encapsulation and decapsulation shows reductions between ten point four and eleven point one percent for encapsulation, and between twenty point six and twenty-six point one percent for decapsulation relative to the work by Cabral et al across those three parameter sets in the paper "Accelerating HQC for Post-Quantum TLS one point three on xeighty-six IoT Gateways <ref:2610.11107#pg1,Accelerating HQC for Post-Quantum TLS 1.3 on x86 IoT Gateways>."

Priya: That’s a solid number because it shows how much faster the actual cryptographic math is running compared to what we’ve seen before, which is crucial when you are trying to save processing power on an IoT gateway.

Title and authors: Nadia: But there's a catch, because this paper also looks at real-world performance on constrained links, and they found that even with these speedups, the communication size of HQC-five still matters a lot for the total handshake time <ref:2610.11107#pg1>.

Elias: They test this on a path with one Mbit/s bandwidth and an added round trip time of fifty milliseconds, and they found the HQC-five handshake took two hundred forty-three milliseconds compared to eighty-seven milliseconds for MLKEM-one thousand twenty-four.

Priya: So even though the math itself is faster, if the ciphertext size is huge, you still hit a bottleneck because of how much data has to travel over that constrained network.

Nadia: Exactly, and they point out that this suggests that for these kinds of links, it’s not just about how fast the CPU crunches the numbers; it’s about how big the public-key and ciphertext sizes are determining the total handshake cost.

Elias: They also looked at TLS one point three integration using liboqs and oqs-provider, showing that on a local loopback setup, their AVX-five hundred twelve HQC-five implementation brought the handshake latency down from three point two two milliseconds to two point nine two milliseconds compared to Cabral et al <ref:2610.11107#pg3>.

Priya: That drop from three point two two down to two point nine two on a loopback is pretty nice, showing that when you control the network environment, those implementation optimizations do provide a tangible reduction in delay for the protocol itself <ref:2610.11107#pg2>.

Nadia: So we’ve seen how they optimized the core math and how it translates into faster handshake times under ideal conditions, but we need to look at what these numbers really mean for real-world deployment.

Elias: They also found that specific component improvements show very high reductions in cycle counts, like the FAFFT implementation for HQC-five reducing the transform cost by seventeen point seven percent over Robert and Veron in page two of this paper <ref:2610.11107#pg2>.

Priya: It’s interesting that these component gains are so high because they are targeting specific bottlenecks, which tells us exactly where we need to focus our optimization efforts for future cryptography on these platforms.

Nadia: They also showed that the AVXtwo implementation for decapsulation reduced the cycle count by eighty-three point two to eighty-seven point one percent compared to Chen et al in page two of this paper, and the AVX-five hundred twelve version was even better at reducing that cost by eighty-eight point nine to ninety-one percent over Cabral et al in page two of this paper <ref:2610.11107#pg2,cycle count by 83.2>.

Elias: And they also accelerated the Keccak-based functions, like SHA3-five hundred twelve oracles, using AVX-five hundred twelve which cut those cycle counts by about thirty-seven percent to forty-six percent compared to Cabral et al in page two of this paper.

Priya: So when you look at the numbers on the component level, it’s clear that using these specialized instructions is delivering very significant speedups for the underlying cryptographic primitives.

Nadia: Overall, what this paper presents is a set of optimized HQC implementations for AVXtwo and AVX-five hundred twelve plus GFNI, tested directly in TLS one point three on xeighty-six IoT gateways <ref:2610.11107#pg1,TLS 1.3 on x86 IoT gateways>.

Title and authors: Elias: And they also made some interesting observations about the relative cost of the FAFFT implementation over the Toom-Cook or Karatsuba multiplication design depending on whether you are using HQC-one HQC-three or HQC-five parameters <ref:2610.11107#pg2>.

Priya: That variation in cost—from seventy-seven percent down to forty-five percent at HQC-five for the FAFFT over Toom-Cook—is a key piece of data because it tells us that the best way to implement the math changes depending on which version of the algorithm you are using.

Nadia: So, to wrap up this discussion on "Accelerating HQC for Post-Quantum TLS one point three on xeighty-six IoT Gateways," we see strong performance gains in key generation and decapsulation when moving to AVX-five hundred twelve <ref:2610.11107#pg1,Accelerating HQC for Post-Quantum TLS 1.3 on x86 IoT Gateways>.

Elias: And while the raw computational speed is definitely up, the paper makes it very clear that on a constrained link, the size of the HQC ciphertext is what ends up dominating the overall handshake time.

Priya: That’s a big point because it shifts our focus from just making individual algorithms faster to also designing key exchanges with smaller outputs when deploying them in real-world, bandwidth-limited environments.

Nadia: So, while these AVXtwo and AVX-five hundred twelve implementations are good for reducing CPU cost on the gateway, the next big challenge is managing that massive ciphertext size over slow networks <ref:2610.11107#pg1>.

Elias: For future work, they suggest looking at applying SFAFFT to AVX-five hundred twelve and integrating caching techniques to see if we can get even more of those cycle count reductions <ref:2610.11107#pg1>.

Priya: That makes sense, since the paper itself flagged that as an area for further improvement, so pushing those specific architectural tweaks is the logical next step.

Nadia: So, looking at the whole picture of "Accelerating HQC for Post-Quantum TLS one point three on xeighty-six IoT Gateways," we have seen concrete performance improvements in key generation and decapsulation using AVX-five hundred twelve and GFNI <ref:2610.11107#pg1,Accelerating HQC for Post-Quantum TLS 1.3 on x86 IoT Gateways>.

Elias: And the study confirms that while implementation speed matters for the CPU, when we put it into a TLS one point three handshake on an IoT device, the network communication cost remains a major factor <ref:2610.11107#pg1>.

Priya: That means our next efforts shouldn't just focus on how fast we can compute things in theory, but also designing key exchanges that minimize the data sent over the wire for those constrained links.

Nadia: We’ve looked at how these AVXtwo and AVX-five hundred twelve implementations speed up HQC, and we’ve seen the trade-off between CPU efficiency and network communication cost on IoT gateways <ref:2610.11107#pg1>.

Elias: The paper "Accelerating HQC for Post-Quantum TLS one point three on xeighty-six IoT Gateways" shows that specific instruction set extensions can give us substantial gains in key generation and decoding cycle counts <ref:2610.11107#pg1,Accelerating HQC for Post-Quantum TLS 1.3 on x86 IoT Gateways>.

Priya: It’s a good reminder that even with hardware optimizations, the physical constraints of the network environment still dictate the final handshake experience for a user or an application.

The paper's summary: Nadia: So we're looking at "Accelerating HQC for Post-Quantum TLS one point three on xeighty-six IoT Gateways," and essentially, this paper shows how they took the standard HQC key exchange and made it much faster on those tiny devices using AVXtwo and AVX-five hundred twelve instructions.

Elias: Right, so we're talking about making the actual math behind that key exchange—the parts that happen during setup and closing a secure connection—run way quicker on hardware like xeighty-six processors.

Priya: What this means for us is that when you’re deploying something on an IoT gateway, you don't want the initial connection to take too long just because the math is heavy.

Nadia: Exactly, and they detail how they tackled the big computational bottlenecks in HQC, like polynomial multiplication and decoding functions, by using these advanced vector instructions from AVX-five hundred twelve.

Elias: They specifically mention extending multiplication methods like Toom-Cook or Karatsuba to use AVX-five hundred twelve for key generation, which gives them a measurable speedup there.

Priya: The data they present shows concrete improvements, like reducing key generation time by nearly thirty percent with the AVX-five hundred twelve version on some parameter sets.

Nadia: It’s not just theoretical; they tested this stuff in a real TLS one point three handshake environment and showed that the latency drops from over three milliseconds down to under three milliseconds when using AVX-five hundred twelve for local loopback tests.

Elias: That's a solid metric because it shows how much those low-level instruction optimizations actually translate into better protocol performance on the wire, even in ideal conditions.

Priya: But here’s the catch, and this is where things get interesting—they looked at constrained links, like a slow one with low bandwidth and extra delay.

Nadia: So if you look at those real-world scenarios, even with all these CPU speedups on the device side, the communication size of the HQC ciphertext ends up dominating how long it takes to complete the handshake.

Elias: They found that on a slow link with one Mbit/s bandwidth and a fifty millisecond delay, the HQC-five handshake took two hundred forty-three milliseconds compared to eight seventy milliseconds for something else.

Priya: That's a huge difference, so it tells us that making the individual math faster doesn't solve the problem of having to send massive amounts of data over a slow connection.

Nadia: So what this changes for us is that we have to look at both sides—the computation and the network traffic—when choosing post-quantum protocols for resource-limited devices.

Elias: They also pointed out how much component performance improved, like cutting the cycle count of the Reed–Solomon decoding by nearly eighty percent with their AVXtwo version.

Priya: So, in short, they proved that while we can make HQC run faster on the processor using these vector instructions, we still need to worry about how big the resulting keys and ciphertexts are when deploying this technology in the real world.

The paper's improvements: Tom: So we're looking at what the authors suggest they should do next for their work on accelerating HQC on xeighty-six gateways, and they're pointing toward some specific architectural tweaks.

Nadia: They’re suggesting applying something called SFAFFT to AVX-five hundred twelve, which is a different way to handle the Fast Fourier Transform that they used before.

Elias: That makes sense because they already showed in the paper that their FAFFT implementation gave them a lot of cycle count reduction over other methods, but this new technique is supposed to push it even further.

Priya: It sounds like they're trying to squeeze out every last bit of speed from the underlying mathematical operations, which is exactly what you want when you’re running heavy crypto on tiny chips.

Nadia: Exactly, and they also mentioned integrating caching techniques into the whole setup to reduce repeated calculations during a handshake.

Elias: Caching is smart; if you calculate a complex part of the key exchange once, storing it so the AI doesn't have to re-run that calculation every single time saves processing power.

Priya: That connects back to what we talked about earlier—if we can reduce the per-call cost of those functions by doing this caching, it means less heat and less battery drain on that IoT gateway.

Nadia: They also noted that they should look at how these optimizations scale with the size of the HQC parameter set, because their FAFFT over Toom-Cook advantage changes depending on whether you’re using HQC-one or HQC-five.

Elias: That’s important because it means there isn't just one perfect way to optimize; the best mathematical approach shifts depending on what kind of security level you need for your key exchange.

Priya: So, the implication is that future research shouldn't just focus on finding faster instructions; it needs to be smarter about *which* instruction set and *which* mathematical structure fits the specific problem at hand.

Nadia: Right, so they’re looking ahead to make these implementations even more flexible and efficient for different types of post-quantum security needs.

Elias: They are also pushing toward making sure these optimizations don't just work on one type of hardware but are robust enough to handle other platforms too.

Priya: That makes sense, because we want solutions that aren't just tailored to one specific chip architecture, but can be ported around.

Nadia: So, the next step is refining the math and adding smart memory tricks to keep those optimized operations running smoothly on these gateways.

Conclusion: Tom: So we’re wrapping up on "Accelerating HQC for Post-Quantum TLS one point three on xeighty-six IoT Gateways," by summarizing how they used AVXtwo and AVX-five hundred twelve to speed up the key exchange process.

Nadia: It really boils down to this: they showed that you can significantly cut down the computational cost of HQC on these specific processors using these instruction extensions.

Elias: And we saw those gains in key generation and decapsulation, especially with the AVX-five hundred twelve version, which was better than prior work by Cabral et al.

Priya: What this means for us is that if you’re deploying AI or secure systems on a gateway, having those math operations run fast saves power and keeps things snappy.

Nadia: True, and the paper made it very clear that while the CPU is getting faster, the physical size of the HQC-five ciphertext still dictates how much total time is spent communicating over constrained networks.

Elias: They pointed out that even with all these low-level instruction optimizations on local loopback tests, those massive ciphertexts still cause a big delay when you move to real network conditions.

Priya: So the data shows that optimizing the math isn't a magic fix for slow internet; you also have to consider how much data is actually flowing across that link.

Nadia: Exactly, and they highlighted that this specific work on "Accelerating HQC for Post-Quantum TLS one point three on xeighty-six IoT Gateways" provides a solid foundation by showing where the biggest computational wins are hiding in the algorithm itself.

Elias: It’s a good look at how hardware acceleration plays into post-quantum security, showing that AVX instructions can make these schemes viable for edge devices.

Priya: And it also flags the next big thing, which is pushing those optimizations further with things like SFAFFT and caching to really get every bit of efficiency out of the system.

Nadia: That’s what they are doing by suggesting future work, looking at how we can combine these techniques for even greater gains on these small hardware targets.

Elias: It's a solid paper for anyone building post-quantum crypto on embedded systems because it gives you concrete numbers on performance improvements versus the trade-offs.

Priya: The main thing I’m taking away is that speed is one part of the equation, but communication size is another huge factor we have to manage for these constrained links.

More episodes

← Home