Heaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language Models
cs.AI, cs.NE
Submitted: 2026-07-06
Updated: 2026-09-01
Comments: Major revisions are going to be implemented and the paper will be restored later
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) generate fluent outputs that can be wrong.
Terminology
Abstract
Large language models (LLMs) generate fluent outputs that can be wrong. Unlike humans, who often exhibit cues when providing false information, LLMs produce errors that are difficult to detect because autoregressive decoding provides no mechanism for verifying intermediate reasoning before state progression. We introduce Heaviside Continuity of Rolling Coefficients (HCRC), a verification-first execution framework that reformulates inference as predicate-gated state transitions governed by a Heaviside Gate. HCRC combines model confidence with independent verification signals from a parallel worker architecture, allowing execution to advance only when predefined correctness predicates are satisfied. This prevents invalid intermediate states from propagating, reducing epistemic entropy without modifying the underlying model. We evaluate HCRC on software-engineering and reasoning tasks across thirteen proposers from four providers. On capable proposers, the gate reduces the false-completion rate (FCR) from 4--7% to 0% while remaining latency-competitive and, in some settings, faster than the unwrapped model. On weaker proposers, it converts false completions into honest halts instead of corrupting downstream state. Beyond benchmarking, HCRC has operated for months as the production control plane of an agentic coding environment, authorizing file mutations, verification-driven progress reporting, and memory compaction. These results establish HCRC as a general framework for verification-driven LLM execution, showing that reliable reasoning can be achieved through principled execution control rather than model scale alone.
Sources
- Teaching Large Language Models to Self-Debug
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- Concrete Problems in AI Safety
- Program Synthesis with Large Language Models
- Constitutional AI: Harmlessness from AI Feedback
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- AI safety via debate
- Language Models (Mostly) Know What They Know
- MemGPT: Towards LLMs as Operating Systems
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- Hallucination is Inevitable: An Innate Limitation of Large Language Models
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection