Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought
summary
The gist
The THOUGHT-LIKE-PRO framework introduces a novel learning framework designed to enhance the logical reasoning capabilities of Large Language Models by imitating verified reasoning processes from a
In short
THOUGHT-LIKE-PRO enhances Large Language Model reasoning by imitating verified Prolog logic engine processes. The framework allows LLMs to self-generate rules and facts, use Prolog to find logical paths, and then translate these paths into natural language Chain-of-Thought (CoT) examples. This method significantly improves the model's ability to reason logically and generalize well on new tasks.
Key concepts
- Prolog Logic Engine
- A symbolic logic system used for formal reasoning. In this framework, it is used to deduce all possible logical paths from given rules and facts to find a target result. This provides a verifiable, step-by-step logical structure that guides the LLM's thinking process.
- Self-Driven Architecture
- The framework operates autonomously without needing external advanced LLMs like GPT-4 for assistance. The base LLM formulates the initial rules and queries, then feeds these into the Prolog engine to derive results, creating a closed loop for reasoning enhancement.
- Reasoning Trajectories
- These are the complete sequences of logical steps or paths that lead from an initial problem state to a final target. The system collects all such paths derived from the Prolog engine and then selectively uses only the successful ones to train the LLM, ensuring it learns correct reasoning patterns.
- Model Averaging
- A technique used during training where results are averaged between a base model and a fine-tuned model (using $\alpha=0.5$). This helps balance specialization for specific tasks with generality, effectively mitigating 'catastrophic forgetting'—the loss of previously learned knowledge when learning new skills.
Terminology used across episodes
This episode discusses
- Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought · Paper Radio
- GPT-4 Technical Report
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
- Evaluating Large Language Models Trained on Code
- Challenges and Contributing Factors in the Utilization of Large Language Models (LLMs)
- Training Verifiers to Solve Math Word Problems
- How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition
- A Survey on In-context Learning
- Supervised Contrastive Learning for Pre-trained Language Model Fine-tuning
- Measuring Massive Multitask Language Understanding
- Measuring Mathematical Problem Solving With the MATH Dataset
- Mitigating the Alignment Tax of RLHF
- Large Language Model Guided Tree-of-Thought
- Faithful Chain-of-Thought Reasoning
- Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
- Continual Learning of Large Language Models: A Comprehensive Survey
- ProofWriter: Generating Implications, Proofs, and Abductive Statements over Natural Language
- A Comparative Study of Open-Source Large Language Models, GPT-4 and Claude 2: Multiple-Choice Test Taking in Nephrology
- Neuro-Symbolic Integration Brings Causal and Reliable Reasoning Proofs
The paper
Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought · Read on arXiv
INF Technology (shanghai) Co., Ltd. · Shanghai University of Engineering Science · Fudan University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought".
Tom: The THOUGHT-LIKE-PRO framework introduces a novel learning framework designed to enhance the logical reasoning capabilities of Large Language Models by imitating verified reasoning processes from a symbolic Prolog logic engine.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So the paper introduces THOUGHT-LIKE-PRO as a novel learning framework that uses imitation learning to mimic the Chain-of-Thought process generated from reasoning trajectories verified by a symbolic Prolog logic engine. The thesis is that this method helps LLMs improve their logical reasoning and generalizes better across outof-distribution reasoning tasks.
Jane: It specifically tackles the challenge of multihop inferential pathways and ensuring validity at every step in the LLM's thinking process, which is a big hurdle for current models. They propose using Prolog to ensure the correctness of these paths before training on them.
Lu: The methodology involves several steps: first, using an open-source chat LLM like Llama3-8B-Instruct to generate rules, facts, and queries based on instructions. Then they use the Prolog engine to find all logical paths leading to a target output.
Meng: And crucially, they only select those reasoning trajectories that actually reach the target output, filtering out any erroneous transformations or generation failures as described in Equation three <ref:2407.14562#pg1>. That selective filtering step is key for keeping the data clean and useful for imitation learning.
Lalam: After getting these verified paths, they create a new few-shot prompt called promptprocot to translate those complex reasoning trajectories into natural language Chain-of-Thought processes that the LLM can actually learn from.
Conclusion: Tom: So looking at the whole thing, THOUGHT-LIKE-PRO is this framework designed to boost the general reasoning abilities of LLMs in a self-driven way by using symbolic logic verification through Prolog. The authors claim it's simple and effective for both specialized and general benchmarks.
Jane: They focus on how this approach improves performance on tasks where the model has to reason logically, showing gains across GSM8K, MATH, MMLU, GPQA, ProofWriter, PrOntoQA, and HumanEval. They also mention that using model averaging helps manage catastrophic forgetting during fine-tuning.
Lu: What this means for the world is that we can start building AI assistants that demonstrate a more structured and verifiable form of thinking rather than just pattern matching based on massive datasets. It moves toward a system where the reasoning steps are explicitly checked against formal logic.
Meng: I see practical implications in reliability; if an AI agent is performing critical tasks, having its reasoning path verified by something like Prolog provides a level of assurance that's much higher than just relying on the model's raw output.
Lalam: For me, it means we can train models to exhibit more robust learning of multi-modal distributions because they are being trained on trajectories that are strictly logical and diverse. It builds better cultural intelligence into the AI.
Tom: So in short, this paper shows that by imitating these verified reasoning trajectories from Prolog, we get improved task performance and better generalization for LLMs. We'll be looking at how they plan to extend this framework further in future work.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck