Is INT8 Portable? A Cross-Platform Measurement Study of Quantized Inference on Embedded and Automotive Accelerators
cs.AR, cs.LG, cs.PF
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 22 pages, 3 figures, 8 tables. Artifact:https://github.com/yyshin-katech/embedded-ai-quantization-guide/tree/paper1-v1
Code: https://github.com/yyshin-katech/embedded-ai-quantization-guide
Project page: https://apple.github.io/coremltools
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell Ultra
- Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
- Quantizing deep convolutional networks for efficient inference: A whitepaper
- A White Paper on Neural Network Quantization
- A Survey of Quantization Methods for Efficient Neural Network Inference
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- Up or Down? Adaptive Rounding for Post-Training Quantization
- FBGEMM: Enabling High-Performance Low-Precision Deep Learning Inference
- Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications
- MLPerf Inference Benchmark
- MLPerf Tiny Benchmark
- MLPerf Mobile Inference Benchmark
- AI Benchmark: Running Deep Neural Networks on Android Smartphones
- AI Benchmark: All About Deep Learning on Smartphones in 2019
- A Comprehensive Evaluation of Deep Learning Object Detection Models on Heterogeneous Edge Devices
- Benchmarking Ultra-Low-Power $\mu$NPUs
- Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance
- MQBench: Towards Reproducible and Deployable Model Quantization Benchmark
- What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4