Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought

arXiv:2407.14562 · cs.AI, cs.CL · Submitted 2024-07-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought".

Tom: The THOUGHT-LIKE-PRO framework introduces a novel learning framework designed to enhance the logical reasoning capabilities of Large Language Models by imitating verified reasoning processes from a symbolic Prolog logic engine.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So the paper introduces THOUGHT-LIKE-PRO as a novel learning framework that uses imitation learning to mimic the Chain-of-Thought process generated from reasoning trajectories verified by a symbolic Prolog logic engine. The thesis is that this method helps LLMs improve their logical reasoning and generalizes better across outof-distribution reasoning tasks.

Jane: It specifically tackles the challenge of multihop inferential pathways and ensuring validity at every step in the LLM's thinking process, which is a big hurdle for current models. They propose using Prolog to ensure the correctness of these paths before training on them.

Lu: The methodology involves several steps: first, using an open-source chat LLM like Llama3-8B-Instruct to generate rules, facts, and queries based on instructions. Then they use the Prolog engine to find all logical paths leading to a target output.

Meng: And crucially, they only select those reasoning trajectories that actually reach the target output, filtering out any erroneous transformations or generation failures as described in Equation three <ref:2407.14562#pg1>. That selective filtering step is key for keeping the data clean and useful for imitation learning.

Lalam: After getting these verified paths, they create a new few-shot prompt called promptprocot to translate those complex reasoning trajectories into natural language Chain-of-Thought processes that the LLM can actually learn from.

Conclusion: Tom: So looking at the whole thing, THOUGHT-LIKE-PRO is this framework designed to boost the general reasoning abilities of LLMs in a self-driven way by using symbolic logic verification through Prolog. The authors claim it's simple and effective for both specialized and general benchmarks.

Jane: They focus on how this approach improves performance on tasks where the model has to reason logically, showing gains across GSM8K, MATH, MMLU, GPQA, ProofWriter, PrOntoQA, and HumanEval. They also mention that using model averaging helps manage catastrophic forgetting during fine-tuning.

Lu: What this means for the world is that we can start building AI assistants that demonstrate a more structured and verifiable form of thinking rather than just pattern matching based on massive datasets. It moves toward a system where the reasoning steps are explicitly checked against formal logic.

Meng: I see practical implications in reliability; if an AI agent is performing critical tasks, having its reasoning path verified by something like Prolog provides a level of assurance that's much higher than just relying on the model's raw output.

Lalam: For me, it means we can train models to exhibit more robust learning of multi-modal distributions because they are being trained on trajectories that are strictly logical and diverse. It builds better cultural intelligence into the AI.

Tom: So in short, this paper shows that by imitating these verified reasoning trajectories from Prolog, we get improved task performance and better generalization for LLMs. We'll be looking at how they plan to extend this framework further in future work.

INF Technology (shanghai) Co., Ltd. · Shanghai University of Engineering Science · Fudan University

cs.AI, cs.CL

Submitted: 2024-07-18

Updated: 2026-10-08

Comments: 15 pages, including appendices. Accepted for publication in IEEE Transactions on Cognitive and Developmental Systems (TCDS)

Code: https://github.com/dave1010/tree-of-thought-prompting

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

The gist: The THOUGHT-LIKE-PRO framework introduces a novel learning framework designed to enhance the logical reasoning capabilities of Large Language Models by imitating verified reasoning processes from a

Key concepts

Prolog Logic Engine
A symbolic logic system used for formal reasoning. In this framework, it is used to deduce all possible logical paths from given rules and facts to find a target result. This provides a verifiable, step-by-step logical structure that guides the LLM's thinking process.
Self-Driven Architecture
The framework operates autonomously without needing external advanced LLMs like GPT-4 for assistance. The base LLM formulates the initial rules and queries, then feeds these into the Prolog engine to derive results, creating a closed loop for reasoning enhancement.
Reasoning Trajectories
These are the complete sequences of logical steps or paths that lead from an initial problem state to a final target. The system collects all such paths derived from the Prolog engine and then selectively uses only the successful ones to train the LLM, ensuring it learns correct reasoning patterns.
Model Averaging
A technique used during training where results are averaged between a base model and a fine-tuned model (using $\alpha=0.5$). This helps balance specialization for specific tasks with generality, effectively mitigating 'catastrophic forgetting'—the loss of previously learned knowledge when learning new skills.

Terminology

Summary

The THOUGHT-LIKE-PRO framework introduces a novel learning framework designed to enhance the logical reasoning capabilities of Large Language Models by imitating verified reasoning processes from a symbolic Prolog logic engine. The gist: this approach substantially enhances the reasoning abilities of LLMs and demonstrates robust generalization across outof-distribution reasoning tasks through imitation learning on strictly logical and diverse reasoning trajectories.

Introduction and Motivation

Large language models (LLMs) have shown exceptional performance as general-purpose assistants, excelling across a variety of reasoning tasks, which represents a significant step toward achieving artificial general intelligence (AGI) Despite these advancements, the effectiveness of LLMs often hinges on the specific prompting strategies employed, and there remains a lack of a robust framework to facilitate learning and generalization across diverse reasoning tasks. General logic reasoning tasks continue to present substantial challenges for LLMs, primarily emerging during the LLMs’ reasoning processes. At the semantic level, the generation of accurate multihop inferential pathways and the assurance of validity at each reasoning step remain significant obstacles.

THOUGHT-LIKE-PRO Framework Overview

The THOUGHT-LIKE-PRO framework is a novel approach designed to enhance the logical reasoning capabilities of LLMs by imitating the verified reasoning processes of the Prolog logic engine. The architecture of THOUGHT-LIKE-PRO is organized in a self-driven manner, eliminating the requirement for auxiliary services from other advanced LLMs, such as GPT4 (Achiam et al., 2023) and Claude 2 (Wu et al., 2023). The framework proceeds in a selfdriven manner, that enables LLMs to formulate rules and statements from given instructions and leverage the symbolic Prolog engine to derive results.

Methodology

The process involves several key steps:

  1. Utilizing open-source chat LLMs, such as Llama3-8B-Instruct (AI@Meta, 2024), to generate rules, facts, and queries following specific instructions and demonstrations. This is formulated as (Rx, Fx, Ty) ∼ pθ(· promptrf t, x, y).

  2. Employing the Prolog engine to perform logical reasoning by deducing the set of reasoning trajectories O(x,y), which encompasses all logical paths leading to the target Ty. This process is formalized as O(x,y) = Prolog(Rx, Fx, Ty).

  3. Selectively utilizing only those reasoning trajectories that conclusively reach the target Ty, excluding any data that result from erroneous transformations or generation failures as delineated in Equation 3. A meta-solver developed by (Yang et al., 2023) is implemented to retrieve all feasible reasoning paths.

  4. Designing a novel few-shot prompt, promptprocot (Listing 2), to translate existing reasoning trajectories into CoT-like reasoning processes expressed in natural language. This construction aids in creating the new dataset DRIF: cproi ∼ pθ(·promptprocot, x, y, oi), oi ∈ O(x,y), (x, cproi, y) ∈ DRIF, for i = 1, n.

  5. Optimizing the model using Supervised Fine-Tuning (SFT) as outlined in Equation 1 to align both the CoT-like reasoning trajectories and the target outputs conditioned on x autoregressively.

Experimental Setup and Results

The experiments utilized two A800-80GB (NVIDIA Corporation, 2023) GPUs for inference, fine-tuning, and evaluation tasks. The framework was tested across seven prominent LLM benchmark datasets including GSM8K (Cobbe et al., 2021), MATH (Hendrycks et al., 2021), MMLU (Hendrycks et al., 2020), GPQA (Rein et al., 2023), ProofWriter (Tafjord et al., 2019), PrOntoQA (Saparov and He, 2019), and HumanEval (Chen et al., 2018). The results show that performing imitation learning on strictly logical and diverse reasoning trajectories for a single instruction leads to improved and consistent reasoning performance. Furthermore, the application of the model averaging technique has proven effective in exploiting the trade-off between specialization and generality, addressing the issue of catastrophic forgetting (CF). The results show that OOD performance is comparable to Llama3-8B-Instruct itself and significantly exceeds that of methods not employing model averaging.

Conclusion

In this paper, THOUGHT-LIKE-PRO is introduced as a framework designed to enhance the general reasoning capabilities of LLMs in a self-driven way. The design is simple, straightforward, and effective, ensuring ease of use in industrial applications. Initial findings show improved task performance on both specialized reasoning and general benchmarks. This foundational framework will be further developed and extensively evaluated across various tasks to refine its methodology in future works.

Limitations

Our comparison of Llama3-8B-Instruct and THOUGHT-LIKE-PRO is constrained to moderatescale language modeling tasks due to limited computational resources. To comprehensively evaluate the performance potential of THOUGHT-LIKE-PRO, it may be beneficial to explore the generation of multiple reasoning paths for more challenging datasets. Additionally, we did not explore the performance of models under different α values when applying the model averaging technique, nor did we balance the performance of the models on specialized versus generalized tasks.

Prompt Demonstration

The framework utilizes a few-shot demonstration prompt (Listing 1) to guide the LLM to generate Prolog code containing rules R, facts F, and targets T based on the problem. This prompt structure guides the model pθ in generating relevant rules Rx, facts Fx, and deriving the reasoning target Ty. The resulting Prolog code is then used by the engine to deduce all logical paths O(x,y). Finally, a natural language description is generated by translating these trajectories into CoT-like reasoning processes expressed in natural language.

Experimental Details

The training process involves typical continual supervised finetuning based on the checkpoint of Llama3-8BInstruct. The Adam optimizer was employed with a consistent learning rate of 1e−6 for 3 epochs of training. For model averaging, α = 0.5 is implemented for a balanced averaging between the base model and the fine-tuned model.

Key Findings Summary

The results indicate that performing domain-specific fine-tuning and forgoing model averaging tend to enhance domain-specific expertise. Training on multiple correct trajectories toward the same target could promote robust learning of multi-modal distributions and significantly improve the robustness of LLM’s reasoning. Model averaging serves as a viable strategy to mitigate catastrophic forgetting, enhancing both the robustness and generalizability of our approach. The stability observed on the MATH dataset suggests that this approach successfully generalizes to MATH tasks and mitigates the negative impacts of catastrophic forgetting. This approach appears to successfully generalize to OOD reasoning tasks.

The THOUGHT-LIKE-PRO framework is a simple, straightforward, and effective method for enhancing LLM reasoning by leveraging symbolic logic verification through Prolog. The framework is straightforward to implement and replicate. Our initial findings show improved task performance on both specialized reasoning and general benchmarks. This foundational framework will be further developed and extensively evaluated across various tasks to refine its methodology in future works.

--- Page 1 ---

The gist

This approach substantially enhances the reasoning abilities of LLMs and demonstrates robust generalization across outof-distribution reasoning tasks through imitation learning on strictly logical and diverse reasoning trajectories.

How it works

The framework proceeds in a selfdriven manner, that enables LLMs to formulate rules and statements from given instructions and leverage the symbolic Prolog engine to derive results. Subsequently, LLMs convert Prolog-derived successive reasoning trajectories into natural language CoT for imitation learning. The architecture of THOUGHT-LIKE-PRO is organized in a self-driven manner, eliminating the requirement for auxiliary services from other advanced LLMs, such as GPT4 (Achiam et al., 2023) and Claude 2 (Wu et al., 2023).

Methodology Steps

The methodology involves constructing a few-shot demonstration prompt to guide the model pθ in generating rules Rx, facts Fx, and deriving the reasoning target Ty. Subsequently, the Prolog engine is employed to deduce the set of reasoning trajectories O(x,y), which encompasses all logical paths leading to the target Ty. During this phase, we selectively utilize only those reasoning trajectories that conclusively reach the target Ty. After acquiring the set O(x,y), we design a novel few-shot prompt, promptprocot (Listing 2), to translate existing reasoning trajectories into CoT-like reasoning processes expressed in natural language.

Improvements for AI systems

  1. Self-Driven Rule Formulation and Verification: The framework enables LLMs to formulate rules and statements from given instructions by leveraging a symbolic Prolog engine, allowing the LLM to formulate rules and statements from given instructions and leverage the symbolic Prolog engine to derive results. This allows for generating logically sound reasoning trajectories verified by an external logic system.

  2. Imitation Learning on Verified Trajectories: The system utilizes imitation learning to imitate the Chain-of-Thought (CoT) process which is verified and translated from reasoning trajectories generated by a symbolic Prolog logic engine. This ensures that the LLM is fine-tuned on strictly logical and diverse reasoning trajectories derived from the Prolog engine, leading to improved and consistent reasoning performance.

  3. Robust Out-of-Distribution (OOD) Generalization: The implementation of a model averaging technique, where α is a hyper-parameter that balances specialization and generalization, is employed to address the issue of catastrophic forgetting (CF), resulting in OOD performance that can be comparable to Llama3-8B-Instruct itself and significantly exceeds that of methods not employing model averaging.

  4. Domain Specialization via Training Configurations: The system allows for exploring different training strategies, such as comparing configurations like MA+multiple versus single trajectory without model averaging, which demonstrates how performing domain-specific fine-tuning and forgoing model averaging tend to enhance domain-specific expertise.

Sources

Related papers