SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops
cs.AR, cs.DC, cs.SY, eess.SY
Submitted: 2026-09-12
Updated: 2026-09-12
Code: https://github.com/WindChimeRan/SiliconBench
Project page: https://ranranhaoranzhang.com/siliconbench
Terminology
Sources
- Native LLM and MLLM Inference at Scale on Apple Silicon
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- LLM Output Drift: Cross-Provider Validation & Mitigation for Financial Workflows
- ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
- Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
- Orca: Progressive Learning from Complex Explanation Traces of GPT-4
- Production-Grade Local LLM Inference on Apple Silicon: A Comparative Study of MLX, MLC-LLM, Ollama, llama.cpp, and PyTorch MPS
- Bench360: Benchmarking Local LLM Inference from 360 Degrees
- Hermes 4 Technical Report
- Open-TQ-Metal: Fused Compressed-Domain Attention for Long-Context LLM Inference on Apple Silicon
- LiveBench: A Challenging, Contamination-Limited LLM Benchmark
- Qwen3 Technical Report
- Gated Delta Networks: Improving Mamba2 with Delta Rule
- Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference
- Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
- SWE-bench Goes Live!
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4