Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation
cs.CV, cs.PF
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/Azure/AzurePublicDataset
Terminology
Sources
- Cosmos World Foundation Model Platform for Physical AI
- Longformer: The Long-Document Transformer
- The Llama 3 Herd of Models
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
- ZipAR: Parallel Auto-regressive Image Generation through Spatial Locality
- ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
- GPT-4o System Card
- D'ej\`a Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
- Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
- Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
- Denoising Diffusion Implicit Models
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
- Gemma 3 Technical Report
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models