LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service
Zhengzhe Xiang, Yinlin Chen, Fuli Ying, Binbin Zhou, Hailiang Zhao, Schahram Dustdar
Hangzhou City University · Zhejiang University · ICREA · TU Wien
cs.DC, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 75/100
The gist: LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service Abstract Summary: As edge-side vision services continue to expand toward low-latency, high-throughput
Terminology
Summary
LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service
Abstract Summary:
As edge-side vision services continue to expand toward low-latency, high-throughput scenarios, reducing the inference cost of vision models without sacrificing reliability has become a central concern. Existing semantic caching methods largely rely on empirical similarity thresholds; while such thresholds improve hit rates, they tend to introduce silent misclassifications near decision boundaries. To address this, we propose LipCache, a certified semantic caching framework for image classification. Without modifying the existing deployed main model, MainNet, the framework introduces a lightweight network, GuardNet, that maps inputs into a low-dimensional feature space subject to a Lipschitz constraint. It then computes a per-sample certified reuse radius from the local classification margin and the spectral norm of the classification head. At runtime, a cached result is reused only when the query feature falls inside the certified reuse ball; otherwise, the query falls back to MainNet. Thus, cache hits are transformed from empirical threshold tests into geometric certification decisions with explicit theoretical boundaries. Across standard image classification tasks like CIFAR, Tiny-ImageNet, and SVHN, LipCache achieves a measured speedup of up to 1.65× with limited end-to-end accuracy degradation, while all accepted cache hits satisfy the GuardNet-side certified-consistency condition. Furthermore, an enhanced GuardNet training recipe substantially improves cache hit rates in the Tiny-ImageNet multi-class extension while maintaining a certified-consistency rate of 100%. These results demonstrate that per-sample certified reuse can reduce main-model fallback while preserving theoretical consistency, providing a feasible approach to reliable cache-assisted inference at the edge.
Introduction Summary:
Deep-neural-network-based vision services have been widely deployed across diverse domains such as autonomous-driving perception, intelligent manufacturing inspection, real-time surveillance, and mobile augmented reality. The dual demand of these services for inference accuracy and response speed has made high-accuracy neural models a de facto standard component. However, the parameter count and computational complexity of such models keep growing, so that a single inference already consumes considerable compute and memory resources. This tension is further amplified in industrial edge scenarios. Unlike occasional photo recognition on a phone, industrial edge devices—such as production-line cameras and roadside perception units—typically operate continuously at high sampling rates, injecting a large volume of concurrent queries into the inference service within very short intervals. When these queries converge at the cloud, they readily trigger spikes in queries-per-second, network congestion, and server overload, severely threatening the overall service availability and tail latency. In other words, the root cause of the problem lies not only in the high cost of a single inference, but more critically in the fact that a large number of semantically redundant queries are repeatedly executed without discrimination.
The rise of the edge computing paradigm offers a structural opportunity to mitigate this risk. By pushing part of the computation or storage capability to edge nodes close to the data source, the system can potentially absorb a fraction of the query load locally, thereby reducing its dependence on the remote cloud link. New challenges arise, however: the latency and memory budgets of edge hardware are extremely tight and far from sufficient to host a full-sized model comparable to its cloud counterpart. Directly deploying a high-accuracy classifier on an edge device almost inevitably faces a sharp conflict between accuracy and resources.
To address this tension, existing work proceeds along roughly two lines. The first line aims to reduce the execution cost of the model at the edge, including compact architecture design, model compression, and early-exit strategies. These methods substantially improve edge deployability, yet their unit of optimization remains how to run the model more cheaply
; they do not touch upon a more fundamental question: for the many semantically repetitive queries, is it really necessary to run the model every time? The cache-assisted inference starts precisely from this gap. It exploits the pervasive redundancy of visual patterns in edge workloads and reuses previously computed results to skip part of the model invocations directly. However, current reuse strategies suffer from a structural weakness in reliability: exact matching is extremely sensitive to compression, cropping, and illumination changes, and can hardly enable effective reuse in real deployments; semantic caching, although it enlarges the reuse region, mostly relies on heuristic similarity thresholds and provides no formal guarantee on the correctness of the reused labels. Consequently, cache hits near decision boundaries may introduce silent misclassifications and uncontrollable accuracy loss. A caching strategy that combines an explainable safety boundary with practical reuse efficiency is still missing.
This paper proposes LipCache, a certified semantic caching framework that comprehensively exploits the storage and computation capability of edge devices. Its core design principle is modularity: rather than modifying the deployed high-accuracy main model (MainNet), it introduces a lightweight guard network, GuardNet, as a new branch to construct the semantic cache. Through operator-norm-controlled convolutions and spectral normalization, GuardNet maps inputs to a smooth low-dimensional feature space, and, starting from the local classification margin and the Lipschitz bound of the classification head, computes a per-sample certified reuse radius offline for each cached sample. Online, only when the query feature falls inside the certified ball of some cached sample, namely when their distance is less than the certified reuse radius, does the system reuse the stored label; otherwise it falls back to MainNet to preserve accuracy. Thereby, a cache hit is converted from an empirical threshold decision into a geometric certification decision with an explicit theoretical boundary.
We validate the core certification mechanism of LipCache on benchmarks with markedly different statistical structure: CIFAR (natural images), Tiny-ImageNet (higher-resolution fine-grained inter-class differences), and SVHN (domain-shifted digits and characters). This choice is designed to isolate the stability of the certified radius under different task geometries, rather than to exhaust all visual task scales; larger-scale settings (e.g., full ImageNet) and non-i.i.d. streaming scenarios are left as open directions discussed in Section 7. On these three tasks, the average hit rates under the tight certified radius reach 0.342, 0.472, and 0.502, respectively, with end-to-end accuracies of 0.921, 0.867, and 0.960, and corresponding measured speedups of 1.32×, 1.31×, and 1.65×; moreover, the certified-consistency rate of all results is 100%. Comparisons with empirical thresholds and the global nearest-neighbor baseline show that the certified radius is the only reuse boundary that can stably preserve certified consistency across the three tasks. Furthermore, in the Tiny-ImageNet multi-class extension, an enhanced GuardNet training recipe improves the hit rates of the 20/30/50-class tasks from 0.056/0.005/0.0004 to 0.423/0.254/0.124, while both the end-to-end accuracy and the certified consistency are preserved.
The main contributions are as follows:
-
We propose a dual-model cache-assisted inference framework that comprehensively exploits the storage and computation capability of edge devices and, without modifying MainNet, makes certified reuse decisions through a lightweight GuardNet.
-
Starting from the local classification margin and the Lipschitz bound of the GuardNet classification head, we derive a per-sample certified reuse radius in the feature space, providing a computable and interpretable rule for safe semantic reuse.
-
We systematically evaluate LipCache across multiple benchmark tasks, multi-class extensions, and edge-deployment boundaries, verifying the effectiveness of the certified radius as the only cross-task stable certified boundary and revealing the potential of the enhanced training recipe in enlarging the certified reuse region.
Related Work Summary:
Edge image classification faces a structural tension: high-accuracy neural models keep growing in their computational and memory demands, whereas edge devices must satisfy tight latency, energy, storage, and connectivity constraints. A practical edge system therefore requires more than just a smaller classifier. It must reduce computation, exploit redundancy among correlated inputs, and decide when a low-cost shortcut is safe rather than merely fast. The literature addresses these needs along three main lines and one important bridge. Efficient edge inference reduces the cost of running a model through compact architectures, adaptive execution, hardware–software co-design, or edge–cloud partitioning. Cache-assisted inference and semantic reuse avoid part of the execution by reusing results from temporally or semantically correlated inputs. Certified nearest-neighbor retrieval forms a bridge, studying when a nearest-neighbor decision in an embedding space remains invariant under perturbation. Lipschitz control and certification provide the tools that turn margins and continuity bounds into local stability guarantees. LipCache connects these lines: it uses a lightweight GuardNet as a cache-oriented front end to the high-accuracy MainNet, and accepts reuse only inside a per-sample certified reuse radius in the GuardNet feature space.
Problem Description Summary:
Consider an image classification system whose input space is X ⊆ Rn and whose label set is Y = 1,..., C. The system may invoke a high-accuracy classifier M: X → Y, but a single inference of M is costly—in edge scenarios this may amount to a combination of GPU time, cloud communication latency, and other overhead. For any input x ∈ X, the system ultimately needs to output a prediction ŷ ∈ Y.
A key observation is that edge vision workloads are not composed of mutually independent random samples. The input–label pairs (x, y) are drawn from a joint distribution D that often exhibits pronounced structural redundancy: influenced by factors such as a fixed camera viewpoint, periodic sampling, and slow scene changes, consecutively arriving queries are highly clustered in pixel space or semantic space, and adjacent inputs frequently share the same label. Under these conditions, executing M in full for every query means that a large amount of computation is wasted on samples whose answers could have been inferred from existing results. This observation points to a natural optimization path: cache a subset of historical predictions, and when a new query is sufficiently similar to some cached sample, directly reuse its label, thereby skipping the invocation of M.
However, the above path faces a fundamental tension. If the reuse condition is too loose, queries near a decision boundary may be incorrectly matched to an inappropriate cached label, leading to silent accuracy loss; if the reuse condition is too strict, cache hits will be so rare that the savings in M invocations become negligible. The core challenge is therefore not whether to introduce a cache,
but how to define a reuse rule that simultaneously achieves a sufficient hit rate to yield meaningful system gains and constrains the resulting prediction errors to a controllable range.
Cache-assisted inference can thus be formalized as the following constrained optimization problem: min InvokeRateM (π) subject to Accsys (π) ≥ AccM − ∆acc, where ∆acc ≥ 0 is a user-specified accuracy tolerance. A smaller ∆acc enforces a more conservative reuse strategy, tending to depress the hit rate to reduce accuracy risk; a larger ∆acc allows more aggressive cache reuse, trading potential accuracy for a lower M invocation frequency.
LipCache Framework Summary:
LipCache adopts a dual-model design, comprising a high-accuracy main model M and a lightweight guard model G. The main model provides the reference prediction: ŷM (x) = arg maxc∈Y Mc(x). The guard model maps each input to a low-dimensional feature space, G: X → Rd. Let z = G(x); an affine classifier operates in this space: s(z) = W z + b, where W ∈ RC×d, b ∈ RC, yielding the GuardNet prediction ŷG (x) = arg maxc sc(G(x)).
Each cache entry is a triple (zi, ỹi, ri), containing a GuardNet feature, a payload label, and a certified reuse radius. The training objective is designed as a layered framework: Ltotal = (1 − α)LCE + αLdistill + λmLmargin, where α ∈ [0, 1] balances direct supervision against an optional distillation signal from M, and λm controls the margin-regularization strength. When α = 0, the objective reduces to the pure cross-entropy training used in the 10-class main experiments of this paper.
Lipschitz control requires the feature map to satisfy Lip(G) ≤ 1. To this end, GuardNet imposes an operator-norm constraint on all parameterized layers. For a convolutional layer whose kernel is K ∈ Rcout×cin×k1×k2, the Lipschitz constant of the convolution operator is given by its frequency-domain representation: Lip(K) = maxω σmax(K̂(ω)), where K̂(ω) is the frequency-domain matrix obtained by zero-padding the kernel and applying a two-dimensional real FFT. Dividing the kernel by max(Lip(K), 1) before each forward pass enforces the 1-Lipschitz constraint. Linear layers use standard spectral normalization; ReLU, non-overlapping average pooling, and the final tanh projection are all non-expansive. Since the Lipschitz constant of each individual layer does not exceed 1, by submultiplicativity of operator norms under composition, the overall guard network satisfies Lip(G) ≤ 1.
Margin regularization directly serves cache efficiency: the certified reuse radius ri is proportional to the GuardNet classification margin, so explicitly encouraging large margins is a key lever for improving the hit rate. Let c1 and c2 denote the indices of the top-1 and top-2 logits: m(x) = [FG (x)]c1 − [FG (x)]c2. To avoid blindly enlarging the margin on already misclassified samples, the regularizer is applied only to the correctly classified subset of a batch B: Scorr = (x, y) ∈ B arg maxc∈Y[FG (x)]c = y, Lmargin = −(1/max(1, Scorr)) Σ(x,y)∈Scorr m(x).
For each candidate sample xi, we first compute zi = G(xi), the GuardNet label yiG = arg maxc [FG (xi)]c, and the payload ỹi selected by the construction strategy. This paper considers three payload semantics. The first is the proxy label, which directly stores ỹi = yiG; in this case the cached payload is naturally aligned with the certificate, and it is the default setting adopted in the main-text experiments. The second is the MainNet prediction scheme, which stores ỹi = ŷM (xi); the third is the ground-truth scheme, which stores ỹi = yi.
At runtime, each input xnew is encoded as znew = G(xnew), and the hit set is H(znew) = i ∥znew − zi∥2 < ri. If H ≠ ∅, it takes the payload from the nearest hit entry: i∗(xnew) = arg mini∈H(znew) ∥znew − zi∥2. The system predicts ŷsys(xnew) = ỹi∗(xnew) if H(znew) ≠ ∅, otherwise ŷM(xnew) if H(znew) = ∅. On a miss, the system falls back to M; if write-back is enabled, a missed sample may be added to the cache under the same construction strategy.
The computational overhead of the above pipeline directly determines the practical gain of the framework. Let CG and CM denote the per-query cost of GuardNet and MainNet, respectively. When maintaining N cache entries in a d-dimensional feature space, the expected cost per query is E[Cost] = O(CG + N d + InvokeRateM CM). When CG + N d is much smaller than CM, the system can reduce the overall overhead in expectation even if the hit rate is not high.
Theoretical Analysis Summary:
The structure of the certified reuse radius can be intuitively understood as classification margin divided by the Lipschitz bound
: the larger the margin, the larger the admissible perturbation; the smoother the mapping, the smaller the logit change induced by the same input variation. The feature encoder G has a finite Lipschitz constant LG ≜ Lip(G), and the per-layer spectral-normalization construction enforces LG ≤ 1. For the affine classification head f (z) = W z + b, noting that the bias term cancels automatically in a logit difference, we have ∥f (z1) − f (z2)∥2 = ∥W (z1 − z2)∥2 ≤ σmax(W)∥z1 − z2∥2, so f is σmax(W)-Lipschitz. By submultiplicativity, Lip(FG) = Lip(f ◦ G) ≤ LG σmax(W). Under the constraint LG ≤ 1, this simplifies to Lip(FG) ≤ σmax(W).
For a cached sample xi, let its feature be zi = G(xi) and its GuardNet prediction label be yiG = arg maxc fc(zi). Define the local classification margin of this sample as mi = [f (zi)]yiG − maxj≠yiG [f (zi)]j. Geometrically, mi measures the logit-space margin between zi and the nearest decision boundary—the larger the margin, the farther the sample lies from a decision boundary, and the larger the perturbation required to flip its label. Accordingly, the certified reuse radius is defined as ri = mi/(σmax(W)√2), where σmax(W) > 0.
Lemma 5.1 (Local label consistency): Let zi be a cached feature with GuardNet label yiG and positive margin mi > 0, and assume σmax(W) > 0. For any new sample xnew whose GuardNet feature is znew = G(xnew), if ∥znew − zi∥2 < ri = mi/(σmax(W)√2), then the GuardNet prediction at znew is still yiG.
The core idea of the proof is to use the Lipschitz bound to translate the displacement in feature space into an upper bound on the perturbation in logit space, and then to show that this perturbation is insufficient to cover the original margin mi. By the Lipschitz property of f, ∥δ∥2 ≤ σmax(W)∥znew − zi∥2. To preserve the label yiG, it suffices to show that for every competing class j ≠ yiG we have [f (znew)]yiG > [f (znew)]j. By the definition of the margin, [f (zi)]yiG − [f (zi)]j ≥ mi, so a sufficient condition for preserving the label is δj − δyiG < mi. Introducing the auxiliary vector ej,yiG that takes the value 1 at position j, −1 at position yiG, and 0 elsewhere, ∥ej,yiG∥2 = √2, and δj − δyiG = e⊤j,yiG δ ≤ ∥ej,yiG∥2∥δ∥2 = √2∥δ∥2, where the inequality follows from Cauchy–Schwarz. Combining with the Lipschitz bound, δj − δyiG ≤ √2σmax(W)∥znew − zi∥2. If ∥znew − zi∥2 < mi/(σmax(W)√2), then δj − δyiG < mi, i.e., the sufficient condition holds. Substituting gives [f (znew)]yiG − [f (znew)]j > 0 for all j ≠ yiG, and hence yiG remains the argmax label at znew.
Corollary 5.1 (Input-level sufficient condition): Given LG ≤ 1, any input satisfying ∥xnew − xi∥2 < ri is guaranteed to satisfy ∥znew − zi∥2 0, the sufficient condition is ∥xnew − xi∥2 < ri/LG.
Remark 5.1 (Tighter per-pair radius): The √2 factor in the radius ri of Lemma 5.1 comes from taking the worst case over all C − 1 competing classes simultaneously. By analyzing each class separately, a tighter radius can be obtained. For each competing class j ≠ yiG, decompose the logit difference as [f (z)]yiG − [f (z)]j = [f (zi)]yiG − [f (zi)]j + (WyiG − Wj)⊤(z − zi). By Cauchy–Schwarz, the cross term is bounded by −∥WyiG − Wj∥2∥z − zi∥2, so when ∥z − zi∥2 [f (z)]j, where ri,j = ([f (zi)]yiG − [f (zi)]j)/∥WyiG − Wj∥2. The tight per-sample radius takes the minimum: ri∗ = minj≠yiG ri,j. Since ∥WyiG − Wj∥2 ≤ √2 σmax(W) and [f (zi)]yiG − [f (zi)]j ≥ mi, we always have ri,j ≥ ri, and hence ri∗ ≥ ri.
Remark 5.2 (Certificate scope and capability boundary): The object certified by Lemma 5.1 is the local consistency of the GuardNet classifier in the feature space: it guarantees that the GuardNet prediction label yiG is invariant inside the certified ball. When the cached payload ỹi happens to equal yiG (the proxy-label mode), the label returned online is directly aligned with the certificate—this is the default setting adopted in the main experiments of this paper. However, the lemma does not provide formal guarantees for the following cases: (1) when the payload source is the MainNet prediction ŷM (xi) or the ground-truth label yi and it is inconsistent with yiG, the cached payload ỹi returned on a hit may differ from the certified yiG; (2) whether the MainNet prediction over the entire certified ball is consistent with the payload. In other words, the theoretical certificate of LipCache is on the GuardNet side, not on the MainNet side. The consistency between GuardNet and MainNet is indirectly promoted by the training objective and evaluated empirically in the experiments.
Experiments Summary:
The main text consistently reports three core metrics: the cache hit rate H (the frequency with which the system reuses the cache), the end-to-end accuracy (the agreement between the overall system output and the ground truth), and the certified-consistency rate (the fraction of all accepted cache hits that satisfy the certification condition; 100% means that no theoretical violation occurred). In addition, the measured system-level speedup is defined as Stime = TMain/(Tguard + Tsearch + Tfallback), where the numerator is the measured time of executing MainNet alone, and the denominator is the measured total time of GuardNet encoding, cache search, and, on a miss, the MainNet fallback.
The standalone accuracies of the three MainNet models are as follows: 94.78% on CIFAR-10, 87.0% on Tiny-ImageNet 10 classes (denoted Tiny-10), and 97.19% on SVHN. In the current final experiments, the certified-consistency rate of all certified-radius results is 100%, and no Lipschitz-constraint violation or theoretical violation at the proxy-decision level is observed.
Under the conservative certified radius, the average hit rates of CIFAR-10, Tiny-10, and SVHN are 0.341 ± 0.010, 0.303 ± 0.013, and 0.470 ± 0.007, respectively; switching to the tight certified radius, the hit rates rise to 0.342 ± 0.011, 0.472 ± 0.012, and 0.502 ± 0.013. The corresponding end-to-end accuracies remain at 0.921, 0.867, and 0.960, and all results pass the certified-consistency check.
A non-trivial certified reuse region is observed on all three tasks. Under the tight certified radius, the hit rates of CIFAR-10, Tiny-10, and SVHN are all significantly above zero, indicating that LipCache does not trigger theoretically safe reuse only in the vicinity of a tiny minority of samples, but rather can form stable usable hits over the full test set. Meanwhile, the end-to-end accuracy is only about 2.70, 0.27, and 1.19 percentage points lower than the respective standalone MainNet accuracy, showing that certified caching does not trade a large loss in prediction quality for reuse opportunities.
The differences across the three datasets reflect the systematic influence of task geometry on the efficiency of certified reuse. SVHN simultaneously exhibits the highest hit rate and the highest measured speedup—on tasks such as digit images, where intra-class variation is constrained, the local geometry learned by GuardNet is more stable, and the certified reuse region more easily covers test samples. The hit rate of Tiny-10 is notably higher than that of CIFAR-10, but its speed gain does not increase proportionally, because the higher input resolution raises the cost of GuardNet encoding and nearest-neighbor search, weakening the translation of hit-rate growth into real wall-clock gains. CIFAR-10 has the lowest hit rate, yet it still maintains a measured speedup above 1.3×—under a lower encoding cost, medium-scale certified hits are already sufficient to yield stable system gains.
Figure 5 compares three kinds of reuse boundaries: the tight certified radius, an empirical radius based on inter-class distance statistics, and a single global nearest-neighbor threshold. The three datasets exhibit a consistent pattern: the certified radius does not pursue the highest hit rate, but it is the only boundary definition that can stably keep the certified-consistency rate at 100% across all three tasks. Taking CIFAR-10 as an example, the hit rate, accuracy, and consistency corresponding to the tight certified radius are 0.344, 0.922, and 100%, respectively; the inter-class-distance empirical radius slightly raises the hit rate to 0.352, but the certified-consistency rate drops to 37.9%; the global nearest-neighbor threshold pushes the hit rate up to 0.494, yet it drives the end-to-end accuracy down to 0.777. On Tiny-10, the empirical thresholds are likewise more aggressive, but their corresponding certified-consistency rates are only 33.3% and 23.9%; on SVHN, the empirical thresholds do not even achieve a higher hit rate, yet they still depress the certified-consistency rate to 45.7% and 25.4%.
On the entry-selection strategy, radius-prioritized de-duplication gives the highest hit rate on CIFAR-10 and also outperforms random sampling and K-center on Tiny-10; on SVHN, the hit rate of K-means is slightly higher than that of radius-prioritized de-duplication, but the gap is small. Regarding return-value semantics, the proxy label is directly aligned with the theoretical guarantee and provides the clearest semantics. On CIFAR-10, the end-to-end accuracies of the proxy-label, MainNet-label, and ground-truth schemes are 0.9215, 0.9134, and 0.9134, respectively; on Tiny-10, they are 0.8680, 0.8680, and 0.8700, respectively; and on SVHN, they are 0.9617, 0.9241, and 0.9237, respectively. Cache capacity remains an effective lever. On all three datasets, increasing the cache from 100 to 400 samples per class raises the hit rate: CIFAR-10 from 0.288 to 0.392, Tiny-10 from 0.446 to 0.480, and SVHN from 0.434 to 0.531. The optimal feature dimensionality is dataset-dependent. On CIFAR-10, the hit rates for d = 32/64/128 are 0.348/0.344/0.336, respectively; on Tiny-10, the corresponding hit rates are 0.456/0.470/0.452; on SVHN, the hit rate continues to rise from 0.498/0.484 to 0.527.
On all three tasks, the normalized distance ratio ∥zq − zi∥2/ri of correct hits is consistently and markedly lower than that of incorrect hits. The mean distance ratios of correct hits on CIFAR-10, Tiny-10, and SVHN are 0.802, 0.752, and 0.767, respectively, whereas those of incorrect hits rise to 0.891, 0.910, and 0.908, systematically closer to the certification boundary. Meanwhile, the cache-center classification margin of correct hits is significantly larger: 2.43, 3.77, and 2.38 on the three datasets, versus only 1.63, 1.97, and 1.49 for incorrect hits. When the radius is progressively tightened on top of the conservative certified radius, all three datasets exhibit a decrease in hit rate and an increase in accuracy, while the certified-consistency rate remains at 100% throughout. For example, on CIFAR-10 the hit rate drops from 0.344 to 0.200 and the accuracy rises from 0.922 to 0.940; on SVHN the hit rate drops from 0.460 to 0.312 and the accuracy rises from 0.964 to 0.970.
Figure 9 gives the training-objective ablation results on CIFAR-10. Cross-entropy supervision alone provides the most robust default operating point among the current 10-class main experiments, with a hit rate and accuracy of 0.358 and 0.920, respectively. The margin-enlargement regularizer mildly raises the hit rate: under a stronger margin-regularization setting, the hit rate rises to 0.401 and the accuracy drops to 0.909. The center constraint that compresses the intra-class distribution is more aggressive: the hit rate can be further raised to 0.544, 0.596, and even 0.644, with the corresponding accuracy dropping to 0.894, 0.888, and 0.877. All training-objective variants pass the certified-consistency check.
On the cross-entropy-only multi-class baseline, the hit rates under the tight certified radius at 20/30/50 classes are only 0.056, 0.0047, and 0.0004, respectively; after adopting the enhanced training recipe, these three values rise to 0.423 ± 0.009, 0.254 ± 0.004, and 0.124 ± 0.001, respectively, while the end-to-end accuracy remains at 0.838, 0.847, and 0.836, and the certified-consistency rate stays at 100%. Under the 20/30/50-class settings, d = 32 achieves the highest hit rates of 0.431, 0.266, and 0.158, respectively, all outperforming d = 64 under the same settings.
Figure 11 gives the empirical stress-test results under query perturbation. All three datasets exhibit the same pattern: additive noise is the most stable, JPEG compression brings a mild decline, and cropping and scaling are the most damaging to the hit rate. For example, on CIFAR-10 the baseline hit rate under the tight certified radius is 0.343; it drops to 0.310 under strong JPEG perturbation, and drops to 0.093 and 0.152 under strong cropping and strong scaling, respectively. In all perturbation experiments, the certified-consistency rate remains at 100%.
The edge board executes GuardNet and cache lookup, while the remote MainNet latency is measured with ResNet50 on an RTX 4070 Ti. The device evaluation uses 2,000 queries and a 2,000-entry cache, obtaining a hit rate of 0.3295 and an end-to-end accuracy of 0.9170. The NPU takes 1.34 ms/query for GuardNet encoding and the cKDTree lookup takes 0.22 ms/query, giving an edge-local path of 1.56 ms/query. Cloud batch-1 MainNet inference takes 8.38 ms/query. Combining these measured components with LAN, WiFi, and cellular RTTs of 2/10/40 ms gives expected speedups of 1.22×, 1.32×, and 1.42×, respectively. The cloud GPU reduces its per-image time from 8.38 ms at batch 1 to 0.207 ms at batch 128, whereas the edge NPU encoding time follows a shallow U-shape (1.34/0.81/0.87/1.18 ms). A discrete-event simulation driven by these measured per-operation times shows a saturation boundary: at 1,000 queries/s, the edge reaches approximately 99% utilization and LipCache p99 latency rises to 177.4 ms, compared with 27.8 ms for all-cloud inference.
Conclusion Summary:
This paper proposed LipCache, a certified semantic caching framework for resource-constrained edge image classification. LipCache retains the high-accuracy MainNet as a fallback predictor and introduces a lightweight, Lipschitz-controlled GuardNet that makes cache-hit decisions in a compact feature space. By deriving a per-sample certified reuse radius from the local GuardNet margin and the Lipschitz bound of the classification head, the framework turns cache lookup into an interpretable local-consistency check, rather than an exact-match or global heuristic-threshold decision.
The experimental results show that, on three 10-class tasks with markedly different statistical structure, LipCache consistently exhibits the following pattern of results: non-zero certified hits, limited accuracy loss, measurable latency gains, and zero theoretical violations.
The comparison between the certified radius and empirical thresholds further shows that the per-sample safety criterion does not pursue the most aggressive hit rate, but provides the only reuse boundary that can stably preserve certified consistency across tasks. Entry-selection strategy, return-value semantics, feature dimensionality, and radius tightening mainly affect the operating point among hit rate, accuracy, and latency, rather than whether certification holds. The multi-class extension experiment further reveals that the main limitation of the current method lies not in the cache criterion itself, but in whether GuardNet can learn a sufficiently compact and separable certified feature geometry. The edge-deployment analysis further indicates that the system gain of LipCache is jointly determined by the hit rate, the guard-side encoding cost, and the network RTT, and that it is best suited to deployment scenarios where cloud invocations are expensive while the edge side can afford a lightweight proxy.
Future Work Summary:
The experiments in this paper validate the viability of the certification mechanism on three 10-class image classification tasks, designed to isolate the sensitivity of the certified radius to task geometry. Extension to larger-scale benchmarks such as full ImageNet (1000 classes) and CIFAR-100 is a natural next step. All experiments are conducted on i.i.d. test sets, targeting a controlled validation of the certification mechanism itself. Temporal redundancy, concept drift, and consecutive-frame reuse rates in streaming deployment involve sequential correlations and require dedicated benchmarks together with a temporal analysis of the hit rate. The goal of this paper is cache-reuse efficiency and inference acceleration, rather than end-to-end energy reduction.
An SNN-based GuardNet variant has been implemented to evaluate its structural compatibility with the LipCache pipeline. The current design retains the same teacher–student training philosophy as the ANN GuardNet, but replaces static activations with multi-step spiking dynamics. Preliminary observations on CIFAR-10 show that the SNN-based GuardNet is trainable and inherits part of the semantic structure learned by the ANN teacher, but they also clearly indicate that the current version is not yet sufficient to replace the ANN guard in the main LipCache pipeline. Specifically, the SNN-based GuardNet reaches a classification accuracy of 74.26% and an agreement rate of 74.78% with MainNet, whereas under the same feature dimensionality the ANN GuardNet reaches 76.84% accuracy and 77.13% agreement. In the current software implementation, the SNN-based GuardNet is also slower—the measured per-sample latency is 6.97 ms versus 0.50 ms for the ANN GuardNet, although the two checkpoints are comparable in size (338.8 KB vs. 355.9 KB).
Future directions include: energy-efficiency evaluation requiring a unified and physically meaningful accounting rule—ANN inference is measured in multiply–accumulate operations, and event-driven SNN inference is measured in accumulate operations; the SNN-based GuardNet itself needs significant improvement before it can be integrated as a reliable guard, including stronger distillation objectives, better temporal-encoding schemes, explicit margin-aware training tailored to spiking features, and meaningful architectural constraints capable of recovering the Lipschitz control used by the ANN GuardNet; the resulting energy and performance claims should be validated on neuromorphic hardware, including Loihi, TrueNorth, or SpiNNaker-class systems; and an improved SNN-based GuardNet should be integrated into the complete LipCache inference pipeline, so that its impact on hit rate, fallback rate, end-to-end latency, and overall throughput can be evaluated at the full-system level rather than only at the standalone level.
Improvements for AI systems
Based on this paper, here are specific improvements to AI systems and what the improved systems can do:
1. Certified Semantic Caching for Inference Acceleration
-
Improvement: Implement a dual-model architecture (MainNet + GuardNet) with Lipschitz-constrained feature extraction and per-sample certified reuse radii, replacing heuristic similarity thresholds with geometrically provable cache-hit decisions.
-
Improved System Capability: An edge inference system that accelerates image classification by 1.3–1.65× while guaranteeing that every cache hit is theoretically consistent with the guard model's prediction, eliminating silent misclassifications near decision boundaries.
2. Formal Safety Guarantees for Reuse Decisions
-
Improvement: Derive and enforce a certified reuse radius r i = m i / (sigma max(W) sqrt 2) from local classification margins and spectral norm bounds, with optional tighter per-class radii r i,j = (f y i G(z i) - f j(z i)) / W y i G - W j 2.
-
Improved System Capability: An AI system that can provably guarantee label invariance within a defined input neighborhood, enabling safe reuse of prior computations without risking incorrect outputs—critical for autonomous driving, medical imaging, or industrial inspection where silent errors are unacceptable.
3. Margin-Aware Training for Cache Efficiency
-
Improvement: Incorporate margin regularization L margin = -1 over S corr sum(x,y) in S corr m(x) into the guard network's training objective, applied only to correctly classified samples to avoid amplifying errors.
-
Improved System Capability: A classifier that learns feature spaces with larger decision-boundary margins, increasing cache hit rates by up to 7.5× (e.g., from 0.056 to 0.423 in 20-class tasks) without sacrificing certified consistency, making caching viable for multi-class and fine-grained recognition.
4. Modular Integration Without Retraining Main Model
-
Improvement: Deploy GuardNet as an external, lightweight branch that leaves the existing high-accuracy MainNet untouched, using distillation (L distill) to align GuardNet's predictions with MainNet's.
-
Improved System Capability: Organizations can retrofit existing production classifiers with certified caching by adding a small guard network (e.g., 355 KB), achieving immediate latency reductions without retraining or modifying their core models—enabling seamless adoption in legacy systems.
5. Adaptive Cache Management with Certified Entry Selection
-
Improvement: Use radius-prioritized de-duplication for cache population, prioritizing entries with larger certified radii, and support multiple payload semantics (proxy label, MainNet label, ground truth) with explicit trade-offs.
-
Improved System Capability: An AI system that automatically maintains a cache of high-value, certifiably safe samples, maximizing hit rates under memory constraints (e.g., increasing cache from 100 to 400 samples per class boosts hits from 0.288 to 0.392 on CIFAR-10) while allowing operators to choose between strict safety (proxy labels) or higher accuracy (MainNet labels).
6. Robustness to Input Perturbations
-
Improvement: Leverage the certified radius to provide inherent robustness guarantees against additive noise, JPEG compression, and mild geometric transforms, as demonstrated by 100% certified-consistency rates under stress tests.
-
Improved System Capability: An image classification service that maintains correctness guarantees even when inputs are degraded by real-world conditions (compression artifacts, sensor noise), with graceful performance degradation (e.g., hit rate drops from 0.343 to 0.310 under strong JPEG) rather than silent failures.
7. Edge-Cloud Load Balancing with Latency-Aware Decisions
-
Improvement: Implement a cost model E[Cost] = O(C G + Nd + InvokeRate M C M) to dynamically decide between edge-local cache lookup and cloud fallback based on measured RTT, encoding cost, and search overhead.
-
Improved System Capability: A distributed inference system that optimally balances edge and cloud resources, achieving speedups of 1.22–1.42× across LAN/WiFi/cellular connections and maintaining acceptable p99 latency (e.g., 177 ms at 1,000 queries/s) even under heavy load, by adaptively routing queries based on network conditions.
8. Scalable Multi-Class Certification
-
Improvement: Apply the enhanced training recipe (combining cross-entropy, distillation, and margin regularization) to scale certified caching from 10 to 50 classes, maintaining 100% certified consistency while improving hit rates from near-zero to 0.124–0.423.
-
Improved System Capability: A general-purpose classification system that can be extended to larger label spaces (e.g., product recognition, scene understanding) while preserving theoretical safety guarantees, making certified caching practical beyond toy benchmarks.
9. Energy-Aware Future Architecture (SNN-Based GuardNet)
-
Improvement: Explore spiking neural network (SNN) variants of GuardNet for event-driven, energy-efficient operation, with teacher-student distillation from ANN guards (currently 74.26% accuracy vs. 76.84% ANN, with potential for neuromorphic hardware deployment).
-
Improved System Capability: A future ultra-low-power edge system that performs certified caching using spiking neurons, reducing energy consumption by orders of magnitude compared to ANN-based guards, suitable for battery-powered IoT devices, while maintaining safety guarantees once SNN training and Lipschitz control are refined.
10. Interpretable Confidence for Cache Hits
-
Improvement: Provide per-sample certified radii as interpretable confidence scores, allowing operators to visualize the safety margin of each cached decision (e.g., correct hits have mean distance ratios of 0.75–0.80 vs. 0.89–0.91 for incorrect hits).
-
Improved System Capability: An AI system that outputs not just predictions but also provable safety bounds for each reused result, enabling human operators to audit decisions, set custom risk thresholds, and trust the system in high-stakes applications where explainability is mandatory.
Abstract
As edge-side vision services continue to expand toward low-latency, high-throughput scenarios, reducing the inference cost of vision models without sacrificing reliability has become a central concern. Existing semantic caching methods largely rely on empirical similarity thresholds; while such thresholds improve hit rates, they tend to introduce silent misclassifications near decision boundaries. To address this, we propose LipCache, a certified semantic caching framework for image classification. Without modifying the existing deployed main model, MainNet, the framework introduces a lightweight network, GuardNet, that maps inputs into a low-dimensional feature space subject to a Lipschitz constraint. It then computes a per-sample certified reuse radius from the local classification margin and the spectral norm of the classification head. At runtime, a cached result is reused only when the query feature falls inside the certified reuse ball; otherwise, the query falls back to MainNet. Thus, cache hits are transformed from empirical threshold tests into geometric certification decisions with explicit theoretical boundaries. Across standard image classification tasks like CIFAR, Tiny-ImageNet, and SVHN, LipCache achieves a measured speedup of up to 1.65 times with limited end-to-end accuracy degradation, while all accepted cache hits satisfy the GuardNet-side certified-consistency condition. Furthermore, an enhanced GuardNet training recipe substantially improves cache hit rates in the Tiny-ImageNet multi-class extension while maintaining a certified-consistency rate of 100%. These results demonstrate that per-sample certified reuse can reduce main-model fallback while preserving theoretical consistency, providing a feasible approach to reliable cache-assisted inference at the edge.
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing