Hydra: Phase-Aware Workload Characterization of LLM Inference across Edge SoC Generations, Backends, and Quantization Levels
cs.AR, cs.AI, cs.DC, cs.PF
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/amirtaherin/hydra
Terminology
Sources
- Evaluating Large Language Models Trained on Code
- No Language Left Behind: Scaling Human-Centered Machine Translation
- DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation
- Human Cognition in Machines: A Unified Perspective of World Models
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Splitwise: Efficient generative LLM inference using phase splitting
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
- ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
- Large Language Models for Robotics: A Survey
- VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
- A Survey on Efficient Inference for Large Language Models
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
- Squat: Quant Small Language Models on the Edge
- Generative AI on the Edge: Architecture and Performance Evaluation
- PalmBench: A Comprehensive Benchmark of Compressed Large Language Models on Mobile Platforms
- MELTing point: Mobile Evaluation of Language Transformers
- LLM in a flash: Efficient Large Language Model Inference with Limited Memory
- MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
- Mobile Foundation Model as Firmware
- MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from Microwatts to Megawatts for Sustainable AI
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4