FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration
cs.AR, cs.AI, cs.DC
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/huawei-csl/wse-workload-generator
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM inference is increasingly constrained by accelerator memory capacity rather than compute throughput.
Terminology
Abstract
LLM inference is increasingly constrained by accelerator memory capacity rather than compute throughput. This constraint is especially acute in single-accelerator and small-node inference systems, where limited on-package memory capacity restricts the size of deployable models. HBF is an emerging 3D-stacked NAND flash technology that provides multi-terabyte near-accelerator capacity, making it a promising capacity tier for storing LLM weights. However, existing HBF-based proposals face three adoption challenges: they (1) rely on coarse-grained static prefetching for LLM weights aiming to hide the microsecond-level read latency of the NAND flash device while maximizing HBF's read throughput, (2) expose NAND flash management tasks (e.g., refresh operations) to the accelerator-visible critical inference path, and (3) miss optimization opportunities to specialize and optimize the flash-management mechanisms to the workload behavior. Our goal is to design an efficient HBF substrate that integrates HBF as a memory-capacity tier alongside HBM while addressing these three challenges. To this end, we propose FLINT, a workload-driven HBF substrate for capacity-scalable LLM inference. FLINT introduces three mechanisms: (1) a hardware burst-buffer controller that dynamically coalesces and pipelines HBF reads aiming to utilize existing NAND flash buffers while sustaining high HBF bandwidth, (2) a phantom-plane refresh mechanism, which removes refresh from the critical inference path by moving refresh-related NAND flash operations outside the read foreground back via low-cost resource duplication, and (3) a read-only FTL, which replaces SSD-class support for arbitrary writes with a compact table that translates logical weight bursts to physical HBF locations.
Sources
- GPT-4 Technical Report
- The Llama 3 Herd of Models
- Scaling Laws for Neural Language Models
- DeepSeek-V3 Technical Report
- Qwen3-VL Technical Report
- Kimi K2: Open Agentic Intelligence
- FP8 Formats for Deep Learning
- Microscaling Data Formats for Deep Learning
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- PowerInfer-2: Fast Large Language Model Inference on a Smartphone
- HAVEN: High-Bandwidth Flash Augmented Vector Engine for Large-Scale Approximate Nearest-Neighbor Search Acceleration
- MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs
- PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
- Measuring Massive Multitask Language Understanding
- Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
- Scaling Up On-Device LLMs via Active-Weight Swapping Between DRAM and Flash
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4