SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving
cs.AI, cs.SE
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/datacurve-ai/deep-swe
Project page: https://nvidia.github.io/TensorRT-LLM/latest/torch/arch_
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Gemma 4 Technical Report
- GLM-5: from Vibe Coding to Agentic Engineering
- DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks
- Kimi K3: Open Frontier Intelligence
- SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
- SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
- ISO-Bench: Can Coding Agents Optimize Real-World Inference Workloads?
- ProgramBench: Can Language Models Rebuild Programs From Scratch?
- InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection