NeuroProlog: Multi-Task Fine-Tuning for Neurosymbolic Mathematical Reasoning via the Cocktail Effect
cs.AI
Submitted: 2026-03-03
Updated: 2026-09-12
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) achieve strong performance on natural language tasks but remain unreliable in mathematical reasoning, frequently generating fluent yet logically inconsistent solutions.
Terminology
Abstract
Large Language Models (LLMs) achieve strong performance on natural language tasks but remain unreliable in mathematical reasoning, frequently generating fluent yet logically inconsistent solutions. We present NeuroProlog, a neurosymbolic framework that ensures verifiable reasoning by compiling math word problems into executable Prolog programs with formal verification guarantees. We propose a multi-task Cocktail training strategy that jointly optimizes three synergistic objectives in a unified symbolic representation space: (i) mathematical formula-to-rule translation (KB), (ii) natural language-to-program synthesis (SOLVE), and (iii) program-answer alignment. This joint supervision enables positive transfer, where symbolic grounding in formula translation directly improves compositional reasoning capabilities. At inference, we introduce an execution-guided decoding pipeline with fine-grained error taxonomy that enables iterative program repair and quantifies model self-debugging capacity. Evaluation on GSM8K across multiple model scales demonstrates that cocktail training improves accuracy over single-task baselines, with statistically significant gains for most evaluated models. Error analysis reveals scale-associated differences in repair behavior: larger models exhibit more readily correctable errors, whereas smaller models show reduced syntactic errors but persistent semantic failures. These findings suggest that model capacity influences the acquisition of reliable symbolic reasoning and self-correction capabilities.
Sources
- Mixing It Up: The Cocktail Effect of Multi-Task Fine-Tuning on LLM Performance -- A Case Study in Finance
- Evaluating Large Language Models Trained on Code
- Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
- Training Verifiers to Solve Math Word Problems
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning
- Faithful Chain-of-Thought Reasoning
- ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
- Solving Quantitative Reasoning Problems with Language Models
- An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
- Code Llama: Open Foundation Models for Code
- An Overview of Multi-Task Learning in Deep Neural Networks
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- Reflexion: Language Agents with Verbal Reinforcement Learning
- VERGE: Formal Refinement and Guidance Engine for Verifiable LLM Reasoning
- OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
- PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection