Carry-Through Checksum: A Lightweight Fault-Detection for CNN Inference at the Edge

arXiv:2609.16742 · cs.AR, cs.LG · Submitted 2026-09-15 · Read on arXiv

cs.AR, cs.LG

Submitted: 2026-09-15

Updated: 2026-09-15

Comments: Accepted at ATS'26. 6 pages, 3 figs and 3 tables

Code: https://github.com/chenyaofo/pytorch-cifar-models

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: Convolutional Neural Networks (CNNs) are increasingly deployed in safety-critical edge applications, where soft errors can silently corrupt inference outputs and lead to unsafe decisions.

Terminology

Abstract

Convolutional Neural Networks (CNNs) are increasingly deployed in safety-critical edge applications, where soft errors can silently corrupt inference outputs and lead to unsafe decisions. Such applications typically rely on resource-constrained embedded GPUs, requiring fault detection and mitigation techniques that add minimal compute, memory, and latency overhead while integrating seamlessly with the standard GPU inference pipeline. Existing algorithm-based fault tolerance techniques rely on matrix augmentation and per-operation checksum verification, imposing substantial overhead that is prohibitive for CNN inference on embedded GPUs. In this work, we propose carry-through checksum, a fundamentally new scheme for soft-error detection in CNN inference on embedded GPUs. The method embeds dedicated carry-through filters into the convolutional layers, which compute a checksum from the CNN's own operations and propagate it through inference, enabling end-to-end error detection with a single output verification. Experimental results on multiple CNN architectures show that the proposed method detects 95.86% and 86.56% of critical faults for FP32 and FP16, respectively, at almost no additional per-image overhead. Detected faults are mitigated through re-execution, incurring only 2.27% run-time overhead across the entire test set on an NVIDIA Jetson Orin NX GPU.

Sources

Related papers