E-ReCON: An Energy- and Resource-Efficient Precision-Configurable Sparse nvCIM Macro for Conventional and Spiking Neural Edge Inference

arXiv:2605.20717 · cs.NE, cs.AR, cs.CV, eess.IV · Submitted 2026-05-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "E-ReCON: An Energy- and Resource-Efficient Precision-Configurable Sparse nvCIM Macro for Conventional and Spiking Neural Edge Inference".

Tom: The gist E-ReCON presents E-ReCON, a 16 Kb energyand resource-efficient digital compute-in-memory (DCIM) macro based on a compact 3T1R ReRAM bitcell for edge-AI inference.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Let’s talk about who wrote this thing. The paper is titled E-ReCON: An Energy- and Resource-Efficient Precision-Configurable Sparse nvCIM Macro for Conventional and Spiking Neural Edge Inference, by Ankit Kumar Tenwar, Mukul Lokhande, and Santosh Kumar Vishvakarma from the Indian Institute of Technology Indore.

Jane: Those authors are from a strong background in electrical engineering, which makes sense considering they’re designing this kind of hardware. The title itself tells you the main goals: energy efficiency, resource efficiency, precision control for sparse networks, and support for both CNNs and SNNs.

Lu: What stands out is that they focus on making this a sixteen Kb macro using a compact 3T1R ReRAM bitcell that takes up only zero point eight five square micrometer of space <ref:2605.20717#pg1,a compact 3T1R ReRAM bitcell>. That tiny physical footprint is a big deal when you’re talking about deploying complex AI on very limited hardware.

Meng: The resource efficiency aspect is what I care about most from an engineering standpoint; fitting a functional macro into such a small area means we can integrate it much more densely into larger systems without creating huge power overheads just to manage the memory access.

Lalam: And the implication for AI culture is that if we can make inference this way, it opens up possibilities for running complex models where power and space are absolutely at a premium, which could change how we deploy things from cloud servers to handheld devices.

The paper's summary: Tom: So, what’s the core idea here? E-ReCON presents this digital compute in memory macro built on a 3T1R ReRAM bitcell designed for edge AI inference <ref:2605.20717#pg1>. They solve the ADC or DAC overhead problem by using this compact bitcell and an interleaved adder tree to keep things very low area.

Jane: In simple terms, they’ve created a way to do the math—the multiplication—right where it is stored in memory, which avoids those energy-hungry conversions we usually need between the memory and the processing unit.

Lu: They use a local AND-based multiplication within the bitcell for both conventional CNNs and spiking neural networks, meaning they support both types of workloads efficiently through that same mechanism.

Meng: That’s smart because it means you don't have to design separate hardware paths for different types of neural networks; one core structure handles everything.

Lalam: For SNNs, the paper highlights that because spike-based computation relies on binary activations, the proposed bitcell can directly support spike-weight multiplication using the same AND operation, which leads to low-power accumulation with less switching activity.

The paper's improvements: Tom: Now let’s look at what they actually improved. They introduced a few key things. First, this fully digital ReRAMCIM architecture eliminates the need for iterative tuning and ADC/DAC requirements, leading to deterministic low-latency operation.

Jane: That deterministic latency is crucial for real-time systems because you know exactly how long the calculation will take, which is a big step up from systems that rely on analog feedback loops.

Lu: They also have this dual-purpose 3T1R AND-type ReRAM bitcell, which they claim improves computational density and minimizes peripheral overhead compared to structures like the 9T4R.

Meng: Minimizing peripheral overhead is important because those extra transistors add complexity and power consumption that we don't want when designing for edge applications where every transistor counts.

Lalam: And then there’s the area-efficient interleaved adder tree, which they designed to reduce transistor count by nearly thirty-seven percent and lower power consumption by twenty-eight percent compared to a conventional 28T RCA-based structure <ref:2605.20717#pg1,compared to a conventional 28T RCA-based>. That efficiency in accumulation is a major win for the overall system energy profile.

Conclusion: Tom: So, to wrap this up, E-ReCON gives us a very compact macro that’s low latency at zero point four eight nanoseconds and can achieve throughputs between two point three one and three point one TOPS while maintaining an energy efficiency of up to four hundred nineteen TOPS/W in their simulation on a sixty-five nm CMOS process at one point two V <ref:2605.20717#pg1,energy efficiency of up to 419 TOPS/W>.

Jane: They also showed good accuracy results, like achieving over ninety-seven percent accuracy on LeNet-five and CNN-eight models under low precision settings, even pruning the network parameters by forty percent <ref:2605.20717#pg1>.

Lu: The paper also addresses robustness; they introduced plus or minus twenty percent variation into the ReRAM Verilog-A model for both high resistance state and low resistance state to show that the bitcell maintains reliable digital functionality because of its large sensing margin <ref:2605.20717#pg1>.

Meng: But what’s a real limitation they point out is that while it handles conventional CNNs and SNNs, they are still focused on the specific performance metrics shown, and you can't expect it to handle every single complex operation without further tuning or architectural changes.

Lalam: For me, the implication for AI culture is that this pushes us toward a future where we can run sophisticated inference models directly on battery-powered edge devices without needing massive power supplies.

Tom: So, E-ReCON is a strong example of how careful design at the bitcell and the adder tree level can lead to a highly efficient system for next-generation edge AI.

Jane: It’s definitely something worth following as we see more digital compute in memory approaches become standard tools for deploying neural networks everywhere.

Indian Institute of Technology Indore

cs.NE, cs.AR, cs.CV, eess.IV

Submitted: 2026-05-20

Updated: 2026-05-20

Journal ref: 2026 30th International Symposium on VLSI Design and Test (VDAT)

DOI: 10.1109/VDAT72243.2026.11704590

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: The gist E-ReCON presents E-ReCON, a 16 Kb energyand resource-efficient digital compute-in-memory (DCIM) macro based on a compact 3T1R ReRAM bitcell for edge-AI inference.

Key concepts

3T1R ReRAM Bitcell
This is a compact memory cell using three transistors and one resistor (ReRAM) to store data. It performs an in-memory AND operation between the stored weight and the input activation, allowing for efficient, parallel multiplication directly where the data is stored.
Interleaved Adder Tree
This is a specialized structure used to sum up the results from multiple bitcells efficiently. It alternates between a compact 10T adder and a standard 28T adder to reduce transistor count by nearly 37% and lower power consumption compared to conventional structures.
DCIM Macro
A DCIM macro is an integrated circuit design that performs computation directly within the memory array. E-ReCON uses this approach to eliminate slow data movement between memory and processing units, leading to very low latency and high energy efficiency for neural network inference.

Terminology

Summary

The gist

E-ReCON presents E-ReCON, a 16 Kb energyand resource-efficient digital compute-in-memory (DCIM) macro based on a compact 3T1R ReRAM bitcell for edge-AI inference.

Proposed Architecture and Bitcell

The proposed architecture eliminates ADC/DAC overhead through a compact 3T1R digital ReRAM bitcell and an interleaved 10T/28T adder tree for low-area accumulation The major contributions of this work are summarised as follows

[Proposed DCIM architecture]: A fully digital ReRAMCIM architecture that eliminates iterative tuning and ADC/DAC requirements for deterministic low-latency operation

[Support for conventional and spiking inference]: The proposed 3T1R bitcell supports both CNN and SNN workloads through efficient AND-based in-memory multiplication

[Dual-purpose 3T1R AND-type ReRAM bitcell]: The proposed bitcell improves computational density, minimises peripheral overhead, and provides a smaller footprint than advanced multiReRAM-cell structures such as 9T4R

[Area-efficient interleaved adder tree]: A resourceefficient interleaved 10T/28T adder tree reduces transistor count by nearly 37% and lowers power consumption compared to a conventional 28T RCA-based structure

ReRAM Bitcell Operation

The proposed ReRAM device operates between high-resistance state (HRS) and lowresistance state (LRS), controlled through SET and RESET voltages The SET/RESET voltage is chosen as 1.2 V with LRS and HRS values of 10 komega and 500 komega, respectively The bitcell performs local AND-based multiplication between the stored weight and the input activation when WL and CIM EN are enabled Input logic ‘1’ is represented by VBL=low and VSL=high, whereas logic ‘0’ is represented by the opposite condition When WL and CIM EN are enabled, the bitcell performs local AND-based multiplication between the stored weight and the input activation

Adder Tree Efficiency

The interleaved adder tree is designed to address the area overhead of conventional ripple-carry adders, which require 28 transistors per full adder This work uses a compact 10T full adder from [4] combined with conventional 28T full adders in an interleaved manner The interleaved adder tree reduces transistor count by nearly 37% and lowers power consumption by 28% compared to a conventional 28T RCA-based structure Since the 10T full adder introduces voltage degradation due to passtransistor logic, however alternates 10T and 28T full adders across successive RCA stages to restore signal integrity

Performance and Evaluation

The proposed macro was implemented in 65 nm CMOS at 1.2 V, achieving a minimum latency of 0.48 ns, throughput of 2.31-3.1 TOPS, and energy efficiency of up to 419 TOPS/W The proposed DCIM macro processes one layer at a time, with weights programmed into the array dynamically The architecture achieves 97.81% accuracy on LeNet-5, 93.23% on AlexNet for CIFAR-10, and 96.51% on CNN-8 for SVHN under lowprecision configurations Furthermore, pruning 40% of the network parameters preserves nearly 99.8% of the original accuracy while reducing MAC operations by 27%, memory bank usage proportionally, and computation cycles by 33% For SNN workloads, the proposed bitcell supports spikedomain multiplication and event-driven accumulation without additional conversion circuitry The 2A2W precision setting provides accuracy very close to FP32 for all evaluated networks, including VGG-8 on CIFAR-10, VGG-16 on CIFAR-10, and ResNet-18 on CIFAR-10

Robustness and Scalability

A major concern in ReRAM-CIM design is variability arising from filament formation and rupture, which affects the HRS/LRS resistance values To capture this effect, ±20% variation was introduced into the ReRAM Verilog-A model for both HRS and LRS states The proposed bitcell maintains reliable digital functionality because of its large sensing margin and digital readout behaviour The proposed 3T1R-AND bitcell occupies only 0.85 µm2, while the 1×4 layout demonstrates scalability toward larger memory arrays Overall, E-ReCON provides a scalable, low-latency, and energy-efficient nvCIM platform for next-generation edge-AI, IoT, biomedical sensing, and neuromorphic applications

Conclusion

E-ReCON demonstrates strong robustness against process and ReRAM variability while supporting 1-8 bit precision These characteristics make it well suited for batterypowered edge-AI, IoT, and neuromorphic applications E-ReCON provides a scalable, low-latency, and energy-efficient nvCIM platform for next-generation edge-AI, IoT, biomedical sensing, and neuromorphic applications

--- Page 1 ---

E-ReCON: An Energy- and Resource-Efficient Precision-Configurable Sparse nvCIM Macro for Conventional and Spiking Neural Edge Inference

Ankit Kumar Tenwar, Mukul Lokhande, Santosh Kumar Vishvakarma Dept. of Electrical Engineering Indian Institute of Technology Indore Simrol 453552, Indore, India

--- Page 2 ---

TABLE I: Truth table of the proposed 3T1R ReRAM bitcell showing AND-based in-memory multiplication for different input activations and stored weight states. Input Weight Output Value BL SL Value ReRAM State Value MULT (AND) -1 (0) VHIGH VLOW 1 LRS 1 VDD

--- Page 3 ---

Fig. 1: (a) Schematic of the proposed CIM-enabled 3T1R ReRAM bitcell implementing resistive AND-based multiplication, (b) Proposed digital ReRAM unit (1 × 4).

--- Page 4 ---

Fig. 2: Transient simulation waveforms of the proposed 3T1R bitcell during MAC operation for (a) HRS-stored weight (‘0’) and (b) LRS-stored weight (‘1’).

--- Page 5 ---

Fig. 3: Process, voltage, temperature, and ReRAM variability for the proposed 3T1R bitcell under TT, SS, FF, SF, and FS corners from-40◦C to 125◦C with ±20% HRS/LRS variation.

--- Page 6 ---

Fig. 4: (a) Proposed energy- and resource-efficient ReRAM-based DCIM macro, (b) accumulation flow of the interleaved adder tree, and (c) structure of the proposed 10T/28T interleaved adder tree.

--- Page 7 ---

Fig. 5: Distribution of peripheral power consumption across accumulation, activation, pooling, batch normalisation, and control circuitry in the proposed DCIM macro.

--- Page 8 ---

TABLE II: Comparison of different adder tree architectures in terms of average error, RMSE, power consumption, and delay in 65 nm CMOS at 1.2 V.

--- Page 9 ---

TABLE III: Comparison of the proposed 16 Kb ReRAM-based DCIM macro with prior CIM architectures in terms of technology, bitcell type, latency, throughput, and energy efficiency.

--- Page 10 ---

TABLE IV: Layer-wise mapping of the pruned LeNet-5 network onto the proposed DCIM macro, including kernel size, output dimensions, bank allocation, and operation cycles.

Improvements for AI systems

  1. Bold header: Digital compute-in-memory (DCIM) architecture eliminates ADC/DAC overhead. This allows for fully digital ReRAMCIM architecture that eliminates iterative tuning and ADC/DAC requirements for deterministic low-latency operation.

  2. Bold header: Support for both conventional CNN and SNN inference. The bitcell is suitable for SNN workloads because it can directly support spike-weight multiplication using the same AND-based operation, enabling low-power spike-domain accumulation with reduced switching activity and memory access overhead.

  3. Bold header: Enhanced energy efficiency through novel accumulation structure. The paper introduces an interleaved 10T/28T adder tree which reduces power consumption by 28% compared to a conventional 28T RCA-based structure.

  4. Bold header: High throughput and low latency for edge AI. The macro achieves a minimum latency of 0.48 ns, throughput of 2.31-3.1 TOPS, making it suitable for battery-powered edge-AI platforms.

  5. Bold header: Robustness against operational variations in hardware. The bitcell maintains reliable digital functionality under all process, voltage, temperature, and ReRAM variability conditions without requiring additional error-correction or reference-tuning circuitry.

Abstract

This work presents E-ReCON, a 16 Kb energy and resource-efficient digital compute-in-memory (DCIM) macro based on a compact 3T1R ReRAM bitcell for edge-AI inference. The proposed bitcell occupies only 0.85 um squared and supports reliable AND-based in-memory multiplication for both conventional convolutional neural network (CNN) and spiking neural network (SNN) workloads. To reduce accumulation overhead, a novel interleaved 10T/28T adder tree is introduced, reducing transistor count and power consumption by 37% and 28%, respectively, compared to a conventional 28T RCA-based design. Implemented in 65 nm CMOS at 1.2 V, the proposed macro achieves a minimum latency of 0.48 ns, throughput of 2.31-3.1 TOPS, and energy efficiency of up to 419 TOPS/W. When evaluated on LeNet-5, AlexNet, and CNN-8 models, the macro achieves 97.81%, 93.23%, and 96.51% accuracy on MNIST/A-Z, CIFAR10, and SVHN datasets, respectively. In addition, 40% pruning preserves nearly 99.8% of the original accuracy while reducing MAC operations and computation cycles. For SNN-oriented workloads, the proposed AND-type bitcell efficiently supports spike-weight multiplication with low switching activity, where the 2A2W configuration achieves accuracy close to the FP32 baseline across VGG-8, VGG-16, and ResNet-18 networks on CIFAR-10, CIFAR-100, and ImageNet-1K datasets. Compared to prior ADC-based ReRAM-CIM designs, the proposed architecture improves latency and energy efficiency by nearly 30-40% while maintaining robust operation under full PVT and ReRAM variability. Overall, E-ReCON provides a scalable, low-latency, and energy-efficient nvCIM platform for next-generation edge-AI, IoT, biomedical sensing, and neuromorphic applications.

Related papers