RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems
cs.DC, cs.AI
Submitted: 2026-09-11
Updated: 2026-09-20
Code: https://github.com/yzygitzh/rooflanghttps:
Project page: https://yzygitzh.github.io/rooflang
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
- DeepSeek-V3 Technical Report
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
- Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization
- GLM-5: from Vibe Coding to Agentic Engineering
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
- Kimi K3: Open Frontier Intelligence
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems
- KernelSight-LM: A Kernel-Level LLM Inference Simulator
- InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
- LLM Inference Unveiled: Survey and Roofline Model Insights
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing