Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration
cs.AR, cs.AI, cs.DC, cs.ET, cs.LG
Submitted: 2026-08-25
Updated: 2026-08-25
Terminology
Sources
- Reasoning Language Models: A Blueprint
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- The Llama 3 Herd of Models
- SpaDA: A Spatial Dataflow Architecture Programming Language
- Sparsity in Deep Learning: Pruning and growth for efficient inference and training in neural networks
- Ultra Ethernet's Design Principles and Architectural Innovations
- Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms
- Data Movement Is All You Need: A Case Study on Optimizing Transformers
- Microscaling Data Formats for Deep Learning
- Qwen2.5 Technical Report
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Cognitive Architectures for Language Agents
- A System Level Compiler for Massively-Parallel, Spatial, Dataflow Architectures
- FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
- SGLang: Efficient Execution of Structured Language Model Programs
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4