Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model
cs.SE, cs.AI
Submitted: 2026-08-24
Updated: 2026-08-26
Terminology
Sources
- Apple Intelligence Foundation Language Models: Tech Report 2025
- Training Verifiers to Solve Math Word Problems
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- Gemma 3 Technical Report
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- Language Models (Mostly) Know What They Know
- What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models
- The Llama 3 Herd of Models
- SELFDOUBT: Uncertainty Quantification for Reasoning LLMs via the Hedge-to-Verify Ratio
- Language Models are Multilingual Chain-of-Thought Reasoners
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
- Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
- Reasoning Models Better Express Their Confidence
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties