Lizard: An Efficient Linearization Framework for Large Language Models
cs.CL, cs.LG
Submitted: 2025-07-11
Updated: 2026-04-18
Code: https://github.com/EleutherAI/lm-evaluation-harness
Terminology
Sources
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Zamba: A Compact 7B SSM Hybrid Model
- RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
- The Llama 3 Herd of Models
- GPT-4 Technical Report
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Measuring Massive Multitask Language Understanding
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Mistral 7B
- LLaMA: Open and Efficient Foundation Language Models
- Language Models are Few-Shot Learners
- Linformer: Self-Attention with Linear Complexity
- RWKV: Reinventing RNNs for the Transformer Era
- PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering