CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks

summary

Video file (mp4)

The gist

This study introduces the CogniDual Framework for LLMs (CFLLMs), designed to assess whether LLMs can, through self-training, evolve from deliberate deduction to intuitive responses, thereby emulating

In short

The episode explores the CogniDual Framework, a method for training Large Language Models (LLMs) to mimic human cognitive efficiency. Based on Daniel Kahneman's dual-system theory, the researchers use a self-training loop to teach models to convert slow, deliberate reasoning into fast, intuitive responses. This process significantly boosts performance on logic-based tasks like LogiQA2 point 0.

Key concepts

Dual-System Theory
This theory suggests human thinking operates using two systems. System one is fast and intuitive, providing quick reactions. System two is slow and deliberate, involving careful, step-by-step logical thought. The framework aims to train models to internalize this dual capacity.
System 1 vs. System 2
System one refers to the fast, automatic response (like a gut feeling), while system two involves slow, methodical reasoning (like solving a complex math problem). The goal of the CogniDual Framework is to train LLMs to internalize their slower reasoning into this quicker, intuitive mode.
Self-Training Loop
This is the core training method. It involves taking questions where a model reasoned correctly but incorrectly (System 2), and using the model itself to rewrite those answers into concise, direct responses. The model then trains on these new pairs to skip the slow reasoning process.

Terminology used across episodes

This episode discusses

The paper

CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks · Read on arXiv

Yongxin Deng, Xihe Qiu, Xiaoyu Tan, Chao Qu, Jing Pan, Yuan Cheng, Yinghui Xu, Wei Chu

Shanghai University of Engineering Science · INF Technology (shanghai) Co., Ltd. · Monash University · Fudan University

DOI: 10.1109/ICASSP49660.2025.10887899

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks".

Jane: The paper was written by Yongxin Deng, Xihe Qiu, Xiaoyu Tan, Chao Qu, Jing Pan et al. from Shanghai University of Engineering Science and INF Technology (shanghai) Co., Ltd. and Monash University and Fudan University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: Welcome back to the show, everyone! We've got a fascinating new paper on arXiv today, and it's called the "CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks." Jane, when I first read that title, I thought, wow, that's a mouthful, but it's actually about something really intuitive.

Jane: It really is, Tom! The title is basically saying that large language models, like the ones we use for chatbots and text generation, might have something in common with how our own brains work. The authors are borrowing this idea from psychology, specifically from a guy named Daniel Kahneman, who came up with this dual-system theory of thinking.

Tom: Right, and that's the "dual-system" part. Kahneman said we have System one which is fast, intuitive, and automatic—like when you catch a ball without thinking about the physics—and System two which is slow, deliberate, and logical—like when you're solving a tricky math problem step by step.

Jane: Exactly! And the paper, written by Yongxin Deng, Xihe Qiu, and their colleagues, asks a really bold question: can we train a large language model to start using its System two which is that careful reasoning, and then internalize it so it becomes more like System one which is that quick, gut-feeling response? It's like learning to drive a car. At first, you're thinking about every pedal and mirror, but after a while, it just becomes second nature.

Tom: That's a great analogy, Jane. And the implications are huge. If a model can answer a complex logic question instantly without having to go through a long chain of thought, it could save a ton of computing power and time. The authors are essentially trying to make these models more efficient by mimicking human learning.

Jane: And it's not just about speed. It's about understanding whether these models actually have a cognitive structure that resembles ours. The researchers are using this framework to probe the inner workings of the models, not just to make them faster. It's a really clever way to bridge psychology and computer science.

Tom: I love that. So we've got a paper that's part psychology experiment, part engineering challenge. I'm really curious to see how they actually tested this. Let's dig into the summary and the core ideas in the next segment.

Jane: Sounds good, Tom. We're just getting to the good stuff.

Summary and Core Ideas: Tom: So, Jane, we're back with the "CogniDual Framework" paper, and I want to get into the nuts and bolts of how they tested this idea. The authors set up a really clever self-training loop. They took models like Llama2 and Vicuna, and first asked them questions without any prompting to reason step-by-step.

Jane: Right, that's the "System one" test. They wanted to see how the model would do on its own, just giving a direct answer. Then, they asked the same questions but with a "chain of thought" prompt, which forces the model to reason out loud, step by step. That's the "System two" test. And, as you'd expect, the models were much better with the step-by-step reasoning.

Tom: But here's the twist. They didn't just stop there. They took the questions where the model got the answer right with reasoning but wrong without it, and they used the model itself to rewrite those reasoned answers into concise, direct answers. Then they trained the model on those new, short question-answer pairs.

Jane: It's like a student studying for a test. They first work through the problem slowly, showing all their work. Then, they study that work, and eventually, they can just look at a similar problem and know the answer instantly. The model is essentially teaching itself to skip the slow part.

Tom: And the results, which are in the paper's table, are pretty striking. For example, on the LogiQA2 point 0 dataset, which is a really tough logical reasoning test, the Vicuna-30B model was getting about twenty-seven point six percent accuracy without reasoning. After the self-training with one thousand examples, it jumped to ninety-six point nine percent! That's a massive leap.

Jane: That's incredible. But it's not just about the big models. Even the smaller ones, like the 7B models, showed significant improvement on tasks like ReClor, which is another logic-based dataset. The paper suggests that the models are building up a kind of "intuition" for these problems, just like humans do.

Tom: And the authors point out that this didn't work as well on the math dataset, GSM8K. They think it's because the models are so used to seeing math problems with step-by-step solutions that they just do it automatically, even when you tell them not to. It's like a habit they can't break.

Jane: So the framework is most effective when there's a big gap between the model's "thinking" and "non-thinking" performance. That's a really useful insight. It tells us where this kind of self-improvement can be most powerful. I'm excited to hear what Lu and Meng think about the practical side of this.

Tom: Absolutely. Let's bring in the experts to see how this could actually be used in the real world.

Improvements and Practical Implications: Tom: Welcome back, Lu and Meng. We've been talking about the "CogniDual Framework" and how it helps models learn to answer quickly. Lu, from a research perspective, what do you think is the most exciting improvement this paper suggests?

Lu: Thanks, Tom. I think the most exciting part is that it gives us a new way to look at model training. We're not just throwing more data at the model. We're asking the model to reflect on its own mistakes and successes. This is a form of self-distillation, but it's specifically targeting the *speed* of cognition. It's not just about being accurate; it's about being accurate *efficiently*. The paper shows that larger models, like the 30B Vicuna, can get to that fast, intuitive state with fewer examples, which suggests they have a better "learning curve" for this kind of internalization.

Meng: I see that, Lu, but as an engineer, my first question is always about the cost. The paper mentions using LoRA for training, which is great for keeping memory usage down. But the whole process involves multiple passes: first generating answers, then rewriting them, then training. That's a lot of compute, even if it's on a single GPU.

Jane: That's a fair point, Meng. But the paper's whole argument is that the *inference* time is what gets saved. Once the model is trained, you don't need to prompt it for a chain of thought anymore. You just ask a question and get an answer. For a real-time application like a chatbot, that's a huge win.

Lu: And it's not just about speed. It's about robustness. The paper shows that this framework can improve performance on datasets the model wasn't specifically trained on, like going from ReClor to LogiQA2 point 0. That suggests the model is learning a generalizable skill, not just memorizing answers.

Meng: That's a good point. If it can generalize, then the initial training cost might be worth it. But I'm still worried about the failure case. The paper mentions that on GSM8K, the math dataset, the improvement was negligible. So this isn't a universal solution. It seems to work best for tasks that require a specific kind of logical leap.

Tom: Right, it's not a silver bullet. But it's a powerful tool for the right kind of problem. Lalam, you're our in-house language model. What do you see as the most impactful vision for this kind of framework?

Lalam: I see this as a step toward more human-like interaction. If models can learn to be intuitive, they can respond more naturally in conversation. Instead of pausing to "think" through every question, they can just answer, which makes the interaction feel more fluid and less robotic. This could make AI assistants more accessible and less intimidating for people who aren't used to technical systems. It's about making the technology feel more like a conversation with a friend, not a transaction with a computer.

Jane: That's a beautiful way to put it, Lalam. So it's not just about efficiency; it's about making AI more approachable. That's a great segue into wrapping up our discussion.

Conclusion: Tom: Well, we've had a fantastic time with the "CogniDual Framework" paper. Let's wrap it up. We started by talking about how the title connects to Kahneman's dual-system theory of human thought, and we saw how the authors used that as a blueprint for training language models.

Jane: We then saw their clever self-training loop, where the model learns to turn its slow, deliberate reasoning into fast, intuitive answers. The results on logic-based datasets like ReClor and LogiQA2 point 0 were really impressive, showing that models can learn to be both smart and quick.

Tom: And we got into the practical side with Lu and Meng, talking about the trade-offs between training costs and inference speed, and the potential for this to make AI feel more natural and approachable, as Lalam pointed out.

Jane: The paper isn't perfect, of course. It didn't help much with math problems, and it seems to work best when there's a clear gap between "thinking" and "not thinking." But it's a really important step toward understanding how these models work and how we can make them better.

Tom: So, as we say goodbye to the "CogniDual Framework," we're left with a big idea: we can teach machines to think fast by first teaching them to think slow. Thanks for listening, everyone. We'll be back soon with another exciting paper from arXiv.

Jane: See you next time, and keep thinking—fast and slow!

More episodes

← Home