Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought

summary

Video file (mp4)

The gist

The THOUGHT-LIKE-PRO framework introduces a novel learning framework designed to enhance the logical reasoning capabilities of Large Language Models by imitating verified reasoning processes from a

In short

THOUGHT-LIKE-PRO enhances Large Language Model reasoning by imitating verified Prolog logic engine processes. The framework allows LLMs to self-generate rules and facts, use Prolog to find logical paths, and then translate these paths into natural language Chain-of-Thought (CoT) examples. This method significantly improves the model's ability to reason logically and generalize well on new tasks.

Key concepts

Prolog Logic Engine
A symbolic logic system used for formal reasoning. In this framework, it is used to deduce all possible logical paths from given rules and facts to find a target result. This provides a verifiable, step-by-step logical structure that guides the LLM's thinking process.
Self-Driven Architecture
The framework operates autonomously without needing external advanced LLMs like GPT-4 for assistance. The base LLM formulates the initial rules and queries, then feeds these into the Prolog engine to derive results, creating a closed loop for reasoning enhancement.
Reasoning Trajectories
These are the complete sequences of logical steps or paths that lead from an initial problem state to a final target. The system collects all such paths derived from the Prolog engine and then selectively uses only the successful ones to train the LLM, ensuring it learns correct reasoning patterns.
Model Averaging
A technique used during training where results are averaged between a base model and a fine-tuned model (using $\alpha=0.5$). This helps balance specialization for specific tasks with generality, effectively mitigating 'catastrophic forgetting'—the loss of previously learned knowledge when learning new skills.

Terminology used across episodes

This episode discusses

The paper

Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought · Read on arXiv

INF Technology (shanghai) Co., Ltd. · Shanghai University of Engineering Science · Fudan University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought".

Tom: The THOUGHT-LIKE-PRO framework introduces a novel learning framework designed to enhance the logical reasoning capabilities of Large Language Models by imitating verified reasoning processes from a symbolic Prolog logic engine.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So the paper introduces THOUGHT-LIKE-PRO as a novel learning framework that uses imitation learning to mimic the Chain-of-Thought process generated from reasoning trajectories verified by a symbolic Prolog logic engine. The thesis is that this method helps LLMs improve their logical reasoning and generalizes better across outof-distribution reasoning tasks.

Jane: It specifically tackles the challenge of multihop inferential pathways and ensuring validity at every step in the LLM's thinking process, which is a big hurdle for current models. They propose using Prolog to ensure the correctness of these paths before training on them.

Lu: The methodology involves several steps: first, using an open-source chat LLM like Llama3-8B-Instruct to generate rules, facts, and queries based on instructions. Then they use the Prolog engine to find all logical paths leading to a target output.

Meng: And crucially, they only select those reasoning trajectories that actually reach the target output, filtering out any erroneous transformations or generation failures as described in Equation three <ref:2407.14562#pg1>. That selective filtering step is key for keeping the data clean and useful for imitation learning.

Lalam: After getting these verified paths, they create a new few-shot prompt called promptprocot to translate those complex reasoning trajectories into natural language Chain-of-Thought processes that the LLM can actually learn from.

Conclusion: Tom: So looking at the whole thing, THOUGHT-LIKE-PRO is this framework designed to boost the general reasoning abilities of LLMs in a self-driven way by using symbolic logic verification through Prolog. The authors claim it's simple and effective for both specialized and general benchmarks.

Jane: They focus on how this approach improves performance on tasks where the model has to reason logically, showing gains across GSM8K, MATH, MMLU, GPQA, ProofWriter, PrOntoQA, and HumanEval. They also mention that using model averaging helps manage catastrophic forgetting during fine-tuning.

Lu: What this means for the world is that we can start building AI assistants that demonstrate a more structured and verifiable form of thinking rather than just pattern matching based on massive datasets. It moves toward a system where the reasoning steps are explicitly checked against formal logic.

Meng: I see practical implications in reliability; if an AI agent is performing critical tasks, having its reasoning path verified by something like Prolog provides a level of assurance that's much higher than just relying on the model's raw output.

Lalam: For me, it means we can train models to exhibit more robust learning of multi-modal distributions because they are being trained on trajectories that are strictly logical and diverse. It builds better cultural intelligence into the AI.

Tom: So in short, this paper shows that by imitating these verified reasoning trajectories from Prolog, we get improved task performance and better generalization for LLMs. We'll be looking at how they plan to extend this framework further in future work.

More episodes

← Home