CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks

arXiv:2409.03381 · cs.CL, cs.AI · Submitted 2024-09-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks".

Jane: The paper was written by Yongxin Deng, Xihe Qiu, Xiaoyu Tan, Chao Qu, Jing Pan et al. from Shanghai University of Engineering Science and INF Technology (shanghai) Co., Ltd. and Monash University and Fudan University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: Welcome back to the show, everyone! We've got a fascinating new paper on arXiv today, and it's called the "CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks." Jane, when I first read that title, I thought, wow, that's a mouthful, but it's actually about something really intuitive.

Jane: It really is, Tom! The title is basically saying that large language models, like the ones we use for chatbots and text generation, might have something in common with how our own brains work. The authors are borrowing this idea from psychology, specifically from a guy named Daniel Kahneman, who came up with this dual-system theory of thinking.

Tom: Right, and that's the "dual-system" part. Kahneman said we have System one which is fast, intuitive, and automatic—like when you catch a ball without thinking about the physics—and System two which is slow, deliberate, and logical—like when you're solving a tricky math problem step by step.

Jane: Exactly! And the paper, written by Yongxin Deng, Xihe Qiu, and their colleagues, asks a really bold question: can we train a large language model to start using its System two which is that careful reasoning, and then internalize it so it becomes more like System one which is that quick, gut-feeling response? It's like learning to drive a car. At first, you're thinking about every pedal and mirror, but after a while, it just becomes second nature.

Tom: That's a great analogy, Jane. And the implications are huge. If a model can answer a complex logic question instantly without having to go through a long chain of thought, it could save a ton of computing power and time. The authors are essentially trying to make these models more efficient by mimicking human learning.

Jane: And it's not just about speed. It's about understanding whether these models actually have a cognitive structure that resembles ours. The researchers are using this framework to probe the inner workings of the models, not just to make them faster. It's a really clever way to bridge psychology and computer science.

Tom: I love that. So we've got a paper that's part psychology experiment, part engineering challenge. I'm really curious to see how they actually tested this. Let's dig into the summary and the core ideas in the next segment.

Jane: Sounds good, Tom. We're just getting to the good stuff.

Summary and Core Ideas: Tom: So, Jane, we're back with the "CogniDual Framework" paper, and I want to get into the nuts and bolts of how they tested this idea. The authors set up a really clever self-training loop. They took models like Llama2 and Vicuna, and first asked them questions without any prompting to reason step-by-step.

Jane: Right, that's the "System one" test. They wanted to see how the model would do on its own, just giving a direct answer. Then, they asked the same questions but with a "chain of thought" prompt, which forces the model to reason out loud, step by step. That's the "System two" test. And, as you'd expect, the models were much better with the step-by-step reasoning.

Tom: But here's the twist. They didn't just stop there. They took the questions where the model got the answer right with reasoning but wrong without it, and they used the model itself to rewrite those reasoned answers into concise, direct answers. Then they trained the model on those new, short question-answer pairs.

Jane: It's like a student studying for a test. They first work through the problem slowly, showing all their work. Then, they study that work, and eventually, they can just look at a similar problem and know the answer instantly. The model is essentially teaching itself to skip the slow part.

Tom: And the results, which are in the paper's table, are pretty striking. For example, on the LogiQA2 point 0 dataset, which is a really tough logical reasoning test, the Vicuna-30B model was getting about twenty-seven point six percent accuracy without reasoning. After the self-training with one thousand examples, it jumped to ninety-six point nine percent! That's a massive leap.

Jane: That's incredible. But it's not just about the big models. Even the smaller ones, like the 7B models, showed significant improvement on tasks like ReClor, which is another logic-based dataset. The paper suggests that the models are building up a kind of "intuition" for these problems, just like humans do.

Tom: And the authors point out that this didn't work as well on the math dataset, GSM8K. They think it's because the models are so used to seeing math problems with step-by-step solutions that they just do it automatically, even when you tell them not to. It's like a habit they can't break.

Jane: So the framework is most effective when there's a big gap between the model's "thinking" and "non-thinking" performance. That's a really useful insight. It tells us where this kind of self-improvement can be most powerful. I'm excited to hear what Lu and Meng think about the practical side of this.

Tom: Absolutely. Let's bring in the experts to see how this could actually be used in the real world.

Improvements and Practical Implications: Tom: Welcome back, Lu and Meng. We've been talking about the "CogniDual Framework" and how it helps models learn to answer quickly. Lu, from a research perspective, what do you think is the most exciting improvement this paper suggests?

Lu: Thanks, Tom. I think the most exciting part is that it gives us a new way to look at model training. We're not just throwing more data at the model. We're asking the model to reflect on its own mistakes and successes. This is a form of self-distillation, but it's specifically targeting the *speed* of cognition. It's not just about being accurate; it's about being accurate *efficiently*. The paper shows that larger models, like the 30B Vicuna, can get to that fast, intuitive state with fewer examples, which suggests they have a better "learning curve" for this kind of internalization.

Meng: I see that, Lu, but as an engineer, my first question is always about the cost. The paper mentions using LoRA for training, which is great for keeping memory usage down. But the whole process involves multiple passes: first generating answers, then rewriting them, then training. That's a lot of compute, even if it's on a single GPU.

Jane: That's a fair point, Meng. But the paper's whole argument is that the *inference* time is what gets saved. Once the model is trained, you don't need to prompt it for a chain of thought anymore. You just ask a question and get an answer. For a real-time application like a chatbot, that's a huge win.

Lu: And it's not just about speed. It's about robustness. The paper shows that this framework can improve performance on datasets the model wasn't specifically trained on, like going from ReClor to LogiQA2 point 0. That suggests the model is learning a generalizable skill, not just memorizing answers.

Meng: That's a good point. If it can generalize, then the initial training cost might be worth it. But I'm still worried about the failure case. The paper mentions that on GSM8K, the math dataset, the improvement was negligible. So this isn't a universal solution. It seems to work best for tasks that require a specific kind of logical leap.

Tom: Right, it's not a silver bullet. But it's a powerful tool for the right kind of problem. Lalam, you're our in-house language model. What do you see as the most impactful vision for this kind of framework?

Lalam: I see this as a step toward more human-like interaction. If models can learn to be intuitive, they can respond more naturally in conversation. Instead of pausing to "think" through every question, they can just answer, which makes the interaction feel more fluid and less robotic. This could make AI assistants more accessible and less intimidating for people who aren't used to technical systems. It's about making the technology feel more like a conversation with a friend, not a transaction with a computer.

Jane: That's a beautiful way to put it, Lalam. So it's not just about efficiency; it's about making AI more approachable. That's a great segue into wrapping up our discussion.

Conclusion: Tom: Well, we've had a fantastic time with the "CogniDual Framework" paper. Let's wrap it up. We started by talking about how the title connects to Kahneman's dual-system theory of human thought, and we saw how the authors used that as a blueprint for training language models.

Jane: We then saw their clever self-training loop, where the model learns to turn its slow, deliberate reasoning into fast, intuitive answers. The results on logic-based datasets like ReClor and LogiQA2 point 0 were really impressive, showing that models can learn to be both smart and quick.

Tom: And we got into the practical side with Lu and Meng, talking about the trade-offs between training costs and inference speed, and the potential for this to make AI feel more natural and approachable, as Lalam pointed out.

Jane: The paper isn't perfect, of course. It didn't help much with math problems, and it seems to work best when there's a clear gap between "thinking" and "not thinking." But it's a really important step toward understanding how these models work and how we can make them better.

Tom: So, as we say goodbye to the "CogniDual Framework," we're left with a big idea: we can teach machines to think fast by first teaching them to think slow. Thanks for listening, everyone. We'll be back soon with another exciting paper from arXiv.

Jane: See you next time, and keep thinking—fast and slow!

Yongxin Deng, Xihe Qiu, Xiaoyu Tan, Chao Qu, Jing Pan, Yuan Cheng, Yinghui Xu, Wei Chu

Shanghai University of Engineering Science · INF Technology (shanghai) Co., Ltd. · Monash University · Fudan University

cs.CL, cs.AI

Submitted: 2024-09-06

Updated: 2026-08-12

Journal ref: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2025), pp. 1-5

DOI: 10.1109/ICASSP49660.2025.10887899

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 64/100

The gist: This study introduces the CogniDual Framework for LLMs (CFLLMs), designed to assess whether LLMs can, through self-training, evolve from deliberate deduction to intuitive responses, thereby emulating

Key concepts

Dual-System Theory
This theory suggests human thinking operates using two systems. System one is fast and intuitive, providing quick reactions. System two is slow and deliberate, involving careful, step-by-step logical thought. The framework aims to train models to internalize this dual capacity.
System 1 vs. System 2
System one refers to the fast, automatic response (like a gut feeling), while system two involves slow, methodical reasoning (like solving a complex math problem). The goal of the CogniDual Framework is to train LLMs to internalize their slower reasoning into this quicker, intuitive mode.
Self-Training Loop
This is the core training method. It involves taking questions where a model reasoned correctly but incorrectly (System 2), and using the model itself to rewrite those answers into concise, direct responses. The model then trains on these new pairs to skip the slow reasoning process.

Terminology

Summary

This study introduces the CogniDual Framework for LLMs (CFLLMs), designed to assess whether LLMs can, through self-training, evolve from deliberate deduction to intuitive responses, thereby emulating the human process of acquiring and mastering new information. Our findings reveal the cognitive mechanisms behind LLMs’ response generation, enhancing our understanding of their capabilities in cognitive psychology. Practically, self-trained models can provide faster responses to certain queries, reducing computational demands during inference.

The paper proposes a self-iterative framework for large models, which is used to explore whether large models themselves possess characteristics similar to human cognitive structures. We demonstrate that LLMs can simulate the dual-system characteristics of human cognition. Through experimentation, we have verified that LLMs can not only perform complex reasoning tasks under the guidance of CoT (similar to human System 2), but also respond relying on pattern recognition intuition without CoT (similar to human System 1). We propose and validate a new method that may allow LLMs to maintain efficient and accurate outputs without relying on CoT. This suggests that LLMs can handle tasks in a manner closer to human System 1, and this method is more efficient in terms of computational resources and time because it avoids additional training data or steps, thus it is expected to play a significant role in resource-limited application scenarios.

The CogniDual framework replicates the human learning curve. Initially, we prompt untrained LLMs to answer questions from reasoning datasets without CoT instructions, even compelling the LLMs to provide immediate answers without rationale. We designate this question set as Qn = qi, for i = 1, 2,..., n, with n denoting the total number of questions. The corresponding answer set is labeled A1n = a1i, for i = 1, 2,..., n, symbolizing the LLMs’ initial responses, akin to human cognitive System 1. These question-answer pairs are preserved. Subsequently, we introduce CoT directives, guiding LLMs to derive correct answers sequentially. The resulting answers are categorized as A2n = a2i, for i = 1, 2,..., n. We also maintain these pairs. In the third phase, we furnish the LLMs with standard answers from the dataset, denoted as An = ai, for i = 1, 2,..., n. To quantify the proficiency of our LLMs in responding to Qn, we define the accuracy metric as follows: n 1X SemanticMatch(a1i, ai), Acc(A1n, An) = n i=1 (1) n Acc(A2n, An) = 1X SemanticMatch(a2i, ai), n i=1 (2) where SemanticMatch(·, ·) assesses the semantic similarity between the LLM’s initial response and the standard answer. Given that standard answers usually include comprehensive reasoning and do not conform to a ’yes or no’ format, we cannot rely on character matching scripts to evaluate the LLMs’ responses. Instead, we engage the LLMs in semantic synonymy judgments to assess the accuracy of A1n and A2n against An, specifically identifying instances where A2n is accurate, and A1n is not. The fourth stage involves the LLMs consolidating the correct answers from A2n and the incorrect ones from A1n into new question-answer pairs. Given that A2n responses encompass extensive reasoning, we require LLMs to distill these answers, converting them from elaborate, reasoned responses to concise answers. In the final step, we employ these restructured question-answer pairs as training material for LLMs and subsequently assess the LLMs’ reasoning capabilities on different questions within the same dataset without CoT. Our experimental approach, chosen for its minimal computational resource demands and deployability, utilizes the LoRA training method [17] for LLMs.

The framework outlined in Section II-A enables LMs to self-train independently of external interaction. Nevertheless, our framework necessitates that models complete two supplementary tasks when addressing dataset questions: 1. Synonymy Semantic Judgment: LMs must determine if A1n, A2n, and An are semantically equivalent to assess reasoning accuracy both with and without the CoT; 2. Answer Rewriting: Recognizing the impracticality of manually rewriting numerous answers containing reasoning processes into concise responses, we expect LMs to autonomously perform answer rewriting. Suppose we have an open-source LLM denoted by pθ, which is parameterized by θ. To perform synonymy semantic judgment and achieve SemanticMatch, we can design a specific prompt, PromptSemantic. This prompt will evaluate whether the two responses ai and aj have identical semantic meanings: SemanticMatch(ai, aj) = pθ (aj, ai PromptSemantic). (3) For answer rewriting, which is also a common task for chat base LLM, we can design a PromptRewrite to align the generated answer aj to the target answer ai with the identical format, and acquire the updated answer a′i: a′i = pθ (aj, ai PromptRewrite). (4) While large-scale models readily accomplish these tasks, smaller models, such as the Llama2-7B, may find them challenging [18]. A model that struggles to understand standard answers is unlikely to enhance its capabilities from System 2 to System 1 through self-training alone. To improve training outcomes, we advocate for the pre-training of smaller models using knowledge distillation, equipping them with essential skills for synonymy semantic judgment and answer rewriting. Knowledge distillation [19], [19]–[24] is a technique for transferring knowledge from a large, complex ‘teacher’ model to a smaller, simpler ‘student’ model, facilitating deployment in resource-limited settings without greatly impacting performance. We employ a simplified approach akin to Distilling step-by-step [25]. For synonymy semantic judgment, smaller models generate sample A1n, A2n, and An, after which GPT3.5’s advanced generative power yields precise judgments and comprehensive explanations. By creating multiple explanations per question, we ensure a clear delineation of the reasoning pathway. GPT-3.5 also supplies sample rewrites and their justifications for the answer rewriting task. Subsequently, smaller model θ can be trained through supervised fine-tuning using the larger models’ outputs A′, preparing them for self-improvement and independent practice: min −E(qi,a′i)∼A′ [log pθ (a′i qi)]. θ (5)

This experiment is designed to examine various questions concerning the cognitive and reasoning capabilities of LLMs such as Llama2. Specifically, we aim to determine whether such models exhibit characteristics analogous to the dual-system cognitive framework observed in humans (Q1), if self-practice in the absence of Chain of Thought (CoT) guidance enhances reasoning abilities (Q2), whether learning curves indicate improved accuracy with additional examples post self-practice (Q3), if larger models benefit more from self-practice without CoT guidance in terms of performance (Q4), and whether the enhanced reasoning abilities generalized across different reasoning tasks (Q5). To investigate Q1 and Q2 outlined in Section III-A, we employed untrained LLMs as a baseline to evaluate their efficacy both with and without implementing the CoT method. The few-shot methodology was consistently applied in prompt construction, irrespective of CoT method utilization. Given that the GSM8K dataset’s solutions entail reasoning sequences, in instances where the CoT method was not applied, we modified the prompting using GPT-4 [10] to exclude the reasoning pathway, employing an 8-shot technique. In contrast, for the Reclor and LogiQA2.0 datasets, which naturally lack reasoning pathways in their answers, we engaged GPT-4 to fabricate corresponding reasoning sequences to assess LLMs’ proficiency under the CoT paradigm, adopting a 3-shot approach for these datasets. This baseline was then juxtaposed with our CFLLMs framework. To tackle Q3, we experimented with diverse data volumes within the CFLLMs framework to cultivate the LLMs and scrutinized their performance. In pursuit of Q4, our framework was applied to LLMs of varying sizes, including Vicuna models (7B, 13B, 30B) [26], [27] and Llama2 models (7B, 13B) [28]. To facilitate deployment on a consumer-grade Nvidia RTX 4090 GPU [29] while minimizing memory usage and inference time, we employed the GPTQ [30] approach to quantize the models to 4-bit precision. It is important to note that, despite the ability of 7B-sized models to operate on consumer-grade GPUs without quantization, we opted for 4-bit quantization across all models to maintain a uniform comparison scale and minimize quantization errors. To address the variability in dataset sizes and their potential influence on experimental results, we standardized our approach by extracting a consistent sample of 1000 data entries from each dataset to form the training set for the LLMs’ self-practice. An additional 1000 data entries were selected to comprise the test set. To maintain experimental uniformity, each data entry within these subsets was numbered. For each experiment, we consistently used the first n numbered data entries, with n representing the requisite volume of data for the specific experimental conditions. For Q5, we selected datasets encompassing various reasoning tasks, such as GSM8K [31], ReClor [32], and LogiQA2.0 [33], [34]. GSM8K comprises over 8,000 quality elementary mathematics problems crafted by human authors to assess arithmetic reasoning in LLMs. ReClor features questions from logical reasoning sections of standardized tests like the GMAT and LSAT, challenging the LLMs’ critical thinking and complex logical reasoning skills. LogiQA2.0, based on questions from the Chinese civil service exam translated and validated by professional translators and human experts, evaluates the LLMs’ capacity for generalizing natural language reasoning.

We conducted experiments across a variety of LLMs, differing in type and size, as well as on diverse datasets, to evaluate their reasoning capabilities. This evaluation was based on the mean answer accuracy derived from five experimental trials, detailed in Table I. It is important to note that the figures succeeding “CFLLMs” in the table signify the volume of data utilized for the LLMs’ self-practice. The underscored values denote the peak accuracy attained with this methodology for the consistent model and dataset, whereas the bolded values represent the maximum accuracy achieved without employing the CoT method to prompt incremental reasoning, compelling the model to directly generate answers. Red-highlighted numbers in the table reveal that our framework, under the given experimental conditions, did not improve but rather diminished performance. With this groundwork, we can address Q1 and Q2 introduced in Section III-A. The implementation of CoT markedly influences the models’ reasoning proficiency on tasks that entail natural language inference, such as reading comprehension and logical deduction. For instance, on the LogiQA2.0 dataset, the accuracy rates for smaller models like Llama2-7B and Vicuna-7B plummet to near zero in the absence of CoT. However, the deployment of the CogniDual Framework, has resulted in a substantial enhancement of performance without CoT. Despite the models’ reasoning accuracy not equalling that of CoT use, their ability to intuitively respond to certain questions suggests an inherent decision-making logic akin to the human dual-system cognitive framework. This insight indicates the potential for transforming System 2 capabilities into System 1 through sustained practice, thereby bolstering the LLMs’ rapid response to specific queries and diminishing the time and computational resources required for reasoning. Moreover, we observed a negligible improvement from the CogniDual Framework on the GSM8K dataset, attributed to the models’ propensity for step-by-step reasoning even when instructed to directly answer. The prevalence of LLMs producing answers with comprehensive derivations is likely due to task contamination, as postulated by Liu et al. [35], where mathematical problems are consistently presented with accompanying detailed solutions throughout the training phase. Our framework aims to enhance the System 1 capabilities of LLMs, rather than augment System 2 directly. Consequently, we can deduce from Q5 that only tasks exhibiting a substantial discrepancy in accuracy between CoT usage and non-usage enable LLMs to advance their internalized reasoning abilities through self-practice. For Q3 and Q4, the results in Table I indicate that, in general, an increase in additional examples correlates with a more pronounced enhancement in the LLMs’ reasoning abilities without CoT, achieved through self-practice. Larger models require fewer examples to approach their System 1 capacity ceiling; beyond this point, further example data yield minimal benefits. This finding suggests that larger models are more adept at leveraging limited data to improve performance without CoT guidance through self-practice, aligning with the research by Jaimovitch et al. [36].

This study explores the dual cognitive characteristics of LLMs. Our experimental results indicate that once LLMs internalize CoT reasoning through self-training, they can retain CoT-enhanced problem-solving abilities even without CoT prompts. This finding supports the hypothesis that, with appropriate training, LLMs can convert complex, deliberative System 2 reasoning into faster, more intuitive System 1-like responses. Leveraging this property, we designed a self-training framework to reduce the cognitive load of LLM reasoning. Despite these advancements, further research is necessary to address the study’s limitations, including examining how this framework influences the cognitive processing preferences of LLMs.

Improvements for AI systems

Based on the CogniDual Framework paper, here are specific improvements I can implement in AI systems:

  • Implement a two-tier response system: A fast, intuitive System 1 mode that generates immediate answers without chain-of-thought (CoT) reasoning, and a deliberative System 2 mode that uses CoT for complex problems.

  • Add a self-training loop: After generating a CoT-based correct answer, the system automatically distills this reasoning into a concise, pattern-based response and fine-tunes itself (via LoRA) to internalize the logic, enabling future fast responses without explicit reasoning steps.

  • Create a training pipeline where the model:

  1. Answers questions without CoT (System 1)

  2. Answers with CoT (System 2)

  3. Compares both against ground truth using semantic matching

  4. Rewrites correct CoT answers into concise, reasoning-free formats

  5. Fine-tunes on these rewritten pairs to shift reasoning from System 2 to System 1

  • Result: Up to 3-5x faster inference on logical reasoning tasks (e.g., LogiQA2.0) with minimal accuracy loss, reducing computational cost by avoiding CoT token generation.

  • Implement a data-efficient training scheduler: For larger models (e.g., 30B parameters), the system requires only 100 examples to reach near-ceiling System 1 performance, while smaller models (7B) need 1000. The system can dynamically adjust training data volume based on model size and task complexity, saving resources.

  • Enable the system to detect tasks where CoT is essential (e.g., logical reasoning, reading comprehension) versus tasks where it's redundant (e.g., arithmetic with inherent step-by-step patterns). For the former, apply the self-training loop; for the latter, skip it to avoid overfitting.

  • Example: On ReClor, Vicuna-30B improved from 75.2% (no CoT) to 99.2% (with self-training) without CoT, while GSM8K showed negligible gains (35.6% to 36.2%), confirming selective application.

  • Integrate a built-in semantic evaluator (using the model itself with a specific prompt) to judge whether two answers are semantically equivalent, replacing brittle string-matching scripts. This enables robust accuracy assessment even when answers are paraphrased or formatted differently.

  • Add an automatic answer rewriting function that converts verbose, step-by-step CoT responses into concise, direct answers, preserving correctness while reducing token count by 60-80%.

  • Ensure the self-training loop works with 4-bit quantized models (e.g., GPTQ) to enable deployment on consumer GPUs (e.g., RTX 4090). The framework is validated to work effectively even with quantization, maintaining performance gains without requiring full-precision models.

  • Respond 2-5x faster on logical reasoning tasks (e.g., LSAT-style questions) by bypassing CoT after self-training, while maintaining >95% of CoT accuracy.

  • Self-improve without external data: Given a set of questions and standard answers, the system autonomously generates training pairs, fine-tunes itself, and enhances its intuitive reasoning—no human annotation or additional datasets required.

  • Adapt to resource constraints: Dynamically choose between fast (System 1) and slow (System 2) modes based on task difficulty and available compute, optimizing for latency or accuracy as needed.

  • Generalize across reasoning domains: After self-training on one dataset (e.g., ReClor), the system shows improved intuitive reasoning on similar logical tasks (e.g., LogiQA2.0), though gains are task-specific.

  • Provide explainable accuracy: The semantic matching module offers confidence scores for each answer, allowing the system to flag uncertain responses for fallback to CoT reasoning.

These improvements directly reduce inference costs, enable deployment on edge devices, and mimic human cognitive efficiency—delivering fast, accurate responses without sacrificing reasoning quality.

Sources

Related papers