SISA: A Scale-In Systolic Array for GEMM Acceleration
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "SISA: A Scale-In Systolic Array for GEMM Acceleration".
Jane: The paper was written by N/A (Author list not present in the provided excerpt) from Swedish Foundation for Strategic Research.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: We’ve established that SISA is designed specifically for General Matrix-Matrix Multiplication, or GEMM, which is the core of almost everything an LLM does.
Jane: The paper highlights a major bottleneck: traditional systolic arrays are huge and efficient, but they can be quite wasteful when the input data isn't perfectly square.
Lu: I see this as a massive win for Lu’s creative possibilities because we’re moving from thinking about massive, fixed-size blocks to thinking about tailored, adaptable computational units.
Meng: From an engineering standpoint, the concept of "scale-in" needs clear definition; it implies we' aren't just making smaller chips but fundamentally changing how the resources are utilized.
Lalam: If we can make these complex calculations much more efficient, as this paper suggests, it means that the enormous power demands of current AI systems might finally become manageable for smaller organizations.
Summary: Tom: The paper provides a very clear summary of the problem: LLMs have input-dependent and highly skewed matrices, which makes traditional accelerators underutilized.
Jane: Think about a typical short prompt; the data dimensions are often small compared to the huge processor arrays, so much of that powerful hardware is just sitting idle.
Lu: That's where SISA comes in—it introduces this idea of parallelism through "slabs," allowing us to exploit those small dimensions in parallel.
Meng: It’s an elegant way to address the mismatch; Meng worries about complexity, though, how do these slabs interact when they are physically located next to each other?
Lalam: If we can use that extra capacity efficiently, Lalam believes it will allow AI to handle more diverse and complex tasks in a single request without needing massive parallel hardware.
Improvements: Tom: The results section of "SISA: A Scale-In Systolic Array for GEMM Acceleration" really stands out, showing up to an eight point five two times speedup for certain LLM workloads.
Jane: That huge speedup is achieved by exploiting that skewed data, which we know is common in prefill phases, and doing so with a massive reduction in energy-delay-product—ninety-three percent less!
Lu: I'm imagining the implications of this: Lu sees a future where training and inference times are dramatically shortened, allowing for iterative improvements to be made much faster than ever.
Meng: The ninety-three percent EDP reduction is critical; from an engineering view, that directly translates to smaller data centers needed to run the same AI capacity.
Lalam: If we can achieve this level of efficiency, Lalam believes it allows us to build more powerful and personalized models that truly reflect human culture and nuance.
Conclusion: Tom: So, we've covered a lot of ground with "SISA: A Scale-In Systolic Array for GEMM Acceleration," from the problem of wasted space to achieving impressive speedups.
Jane: It looks like this architecture successfully maintains peak performance for large tasks while delivering huge gains when dealing with those tricky, small matrices.
Lu: I think Lu is excited about the potential, especially seeing how the flexible design can handle anything from short prompts to massive batches without a huge penalty.
Meng: From a hardware perspective, Meng appreciates that SISA achieves this performance while keeping the area overhead relatively low compared to other reconfigurable designs.
Lalam: Lalam feels confident that "SISA: A Scale-In Systolic Array for GEMM Acceleration" sets us up for a much more efficient future where the power of AI can be realized by everyone, not just a few.
N/A (Author list not present in the provided excerpt)
Swedish Foundation for Strategic Research
cs.AR, cs.AI
Submitted: 2026-03-31
Updated: 2026-08-25
Journal ref: Euro-Par 2026: Parallel Processing, Lecture Notes in Computer Science, vol. 16781, pp. 359-372, Springer
DOI: 10.1007/978-3-032-35248-4_25
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 88/100
The gist: This paper proposes SISA (Scale-In Systolic Array), a novel hardware architecture designed to accelerate General Matrix-Matrix Multiplication (GEMM) operations, which are critical for Large Language
Key concepts
- GEMM (General Matrix-Matrix Multiplication)
- This is the core mathematical operation that powers almost everything an LLM does. The paper focuses on optimizing this calculation using specialized hardware accelerators, making it critical for improving AI system efficiency.
- Systolic Arrays
- These are powerful, traditional hardware structures designed for matrix calculations. However, they can be wasteful and underutilized when the input data dimensions are small or not perfectly square.
- SISA (Scale-In Systolic Array)
- SISA is an architecture that addresses the inefficiency of traditional arrays. By using a 'scale-in' approach and exploiting parallelism through 'slabs,' it efficiently handles skewed or small matrix dimensions.
Terminology
Summary
This paper proposes SISA (Scale-In Systolic Array), a novel hardware architecture designed to accelerate General Matrix-Matrix Multiplication (GEMM) operations, which are critical for Large Language Model (LLM) workloads. Traditional square systolic arrays suffer from significant under-utilization when processing the highly skewed
or small matrices common in LLM prefill and decode phases; SISA addresses this by introducing an adaptive architecture that can partition its processing elements into smaller, independently scheduled units to maximize efficiency.
The Problem of Under-utilization
Modern AI workloads, particularly LLMs, introduce specific computational challenges that traditional monolithic systolic arrays (SAs) are ill-equipped to handle. While large square arrays are efficient for dense GEMMs, they struggle with the input-dependent and highly skewed matrices
found in real-world applications. The paper identifies two primary phases of LLM inference:
Prefill phase:
Processes the entire input prompt using dense GEMMs which are compute-bound,
often involving skewed matrices due to short prompt lengths (median of 12 tokens).
Decode phase:
Generates tokens sequentially using GEMVs, making it predominantly memory-bandwidth bound.
While batching can help, factors like the KV cache and latency requirements prevent aggressively large batch sizes, leading to highly skewed matrix shapes.
How SISA Works
SISA extends the traditional square design with a scale-in mechanism
that partitions the array into horizontal rectangular sub-arrays called slabs.
These slabs are capable of parallel, independently scheduled execution,
allowing the hardware to adapt to various matrix dimensions. The architecture employs an Output-Stationary (OS) dataflow where each PE maintains a local accumulator, simplifying slab composition because partial sums do not traverse inter-PE interconnects.
The architecture utilizes several key mechanisms to maintain flexibility:
Slab Fusion:
Slabs can be fused along the height dimension using a simple buffer bypass implemented with multiplexers
to form larger logical arrays for large GEMMs.
Tiling and Scheduling:
The system decomposes GEMMs into tiles that map to different slab configurations based on the matrix dimensions (e.g., distributing computation along the N dimension for small-M matrices or fusing slabs for intermediate-M matrices).
Power-Gating:
To improve efficiency, unused slabs can be power-gated
to minimize idle energy consumption.
Performance and Evaluation
The researchers evaluated an instance of SISA with a 128×128 PE configuration using state-of-the-art LLM models like Qwen2.5 and Llama3.2. When compared to a monolithic TPUv4 baseline, SISA demonstrates significant advantages in the under-utilized regime, particularly for small sequence lengths (m ≤ 16), where it achieves up to 8.52× speedup and a reduction in energy-delay-product (EDP) up to 93%.
In comparison to other flexible architectures:
Versus ReDas:
SISA outperforms ReDas in most regimes, achieving up to 2.61× speedup while also exhibiting lower area overhead.
Area and Energy:
SISA incurs a modest total chip area increase of approximately 5.44% over a baseline TPU, as it preserves a simple PE design and concentrates overhead in the memory hierarchy.
Even in worst-case scenarios where SISA operates as a monolithic array, it remains competitive, incurring only an 8.47% higher EDP
than the TPU.
Improvements for AI systems
To improve AI systems based on the SISA architecture, I would implement the following hardware-software co-design improvements:
-
Implement a
Scale-In
Systolic Array (SISA) hardware backend instead of traditional monolithic systolic arrays. This involves partitioning the processing element (PE) grid into independently scheduled horizontal rectangularslabs
with dedicated local buffers and multiplexer-based slab fusion mechanisms. -
Integrate a runtime tiling and scheduling engine capable of dynamically switching between three execution modes: independent slab execution for small/skewed matrices, fused-slab execution for intermediate shapes, and full monolithic execution for large GEMMs.
-
Deploy granular power-gating at the slab level to disable unused PEs and local buffers during low-batch or short-sequence workloads.
By implementing these improvements, the AI system will be able to:
-
Execute LLM prefill phases with highly skewed matrices (e.g., short user prompts) up to 8.52× faster than current TPU-style architectures by maximizing PE utilization through slab-level parallelism rather than waiting for a full array to drain.
-
Perform high-efficiency LLM decoding at small batch sizes with up to a 93% reduction in Energy-Delay-Product (EDP), significantly reducing the operational cost of interactive chatbots.
-
Maintain peak throughput for massive GEMM operations through seamless slab fusion, ensuring that the transition from small to large workloads does not incur significant performance penalties.
-
Optimize Quality-of-Service (QoS) in multi-tenant environments by reducing Time To First Token (TTFT) through more efficient processing of variable-length input sequences.
Abstract
The currently dominant AI/ML workloads, such as Large Language Models (LLMs), rely on the efficient execution of General Matrix-Matrix Multiplication (GEMM) operations. Thus, most systems are equipped with dedicated matrix hardware accelerators based on square Systolic Arrays (SAs) of Processing Elements (PEs). While this organization was effective for traditional Deep Neural Networks (DNNs), LLMs introduce input-dependent and highly skewed matrices, leading to underutilized SA resources. To address this challenge, we propose SISA (Scale-In Systolic Array), a novel SA architecture that partitions the traditional square array into horizontal rectangular slabs. With minimal overhead, SISA exposes parallelism through independently scheduled slabs for efficient execution of small or skewed matrix shapes, while retaining full-array operation for large GEMMs. SISA achieves up to 8.52x speedup and 93% energy-delay-product (EDP) reduction for representative LLMs compared to a state-of-the-art monolithic SA with the same number of PEs.
Sources
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
- DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale
- FlexSA: Flexible Systolic Array Architecture for Efficient Pruned DNN Model Training
- LLM Serving Optimization with Variable Prefill and Decode Lengths
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4