What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend
Shahed Masoudian, Passant Shafaei, Monorama Swain, Markus Schedl
cs.SE, cs.AI, cs.LG
Submitted: 2026-08-05
Code: https://github.com/langchain-ai/langchain
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Lessons from the Trenches on Reproducible Evaluation of Language Models
- Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
- The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties