Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities

summary

Video file (mp4)

In short

This episode discusses a paper examining if language models perform better when presented with formal logic notation versus natural English. Researchers developed a framework called CLGC to test various notations on syllogisms. Findings show that smaller models benefit from compact symbolic notation, while larger models leverage their training on natural language, aiming to improve AI reliability in logical tasks.

Key concepts

Syllogistic Reasoning
This is the ability to reason through arguments presented in a syllogism—a specific form of logical argument. The paper tests whether presenting these arguments using formal, structured notation helps the model process and conclude the logic correctly compared to standard natural language.
CLGC (Common Logic Grammar Construction)
This is a framework used by researchers to translate complex arguments into various formal symbolic languages. CLGC converts syllogisms from natural language into multiple notations, such as TFLPLUS or CLIF, allowing them to test how different ways of writing the logic affect model performance.
SEF Categories
These are metadata categories (e.g., Categorical, Hypothetical) used to classify the structure of a syllogism. Instead of just translating the language, researchers categorize arguments into types, providing structural hints that can boost a model's performance on specific types of logic problems.
Supervised Fine-Tuning
This is an experimental method where researchers train small models on specific formal notations to improve their reasoning ability. This contrasts with Zero-Shot testing, which simply prompts large models with instructions and the syllogism without any prior training.

Terminology used across episodes

This episode discusses

The paper

Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities · Read on arXiv

Hanna Abi Akl, Fabien Gandon, Catherine Faron, Pierre Monnin

Université Côte d’Azur · inria · CNRS · I3S · Data ScienceTech Institute

Language models (LMs) struggle with logical tasks like reasoning on syllogisms. It has been shown that Knowledge Representation (KR) plays a crucial role in expressing input information to help models solve tasks. This observation motivates our study of the impact of different formal KR notations on syllogistic reasoning by extending the FOLIO and P-FOLIO datasets. Our experiments on Small Language Models (SLMs) in Supervised Fine-Tuning (SFT) and Zero-Shot (ZS) settings show that the choice of input notation can yield performances competitive with natural language while enabling faster inference. We also propose a syllogistic categorization method (SEF) and use it to enrich ZS prompts with logical definitions, which boost reasoning in small models. We open-source our framework, Common Logic Grammar Construction (CLGC), as the first Python library for automatically generating syllogisms in KR notations and defining their SEF categories.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities".

Jane: The paper was written by Hanna Abi Akl, Fabien Gandon, Catherine Faron and Pierre Monnin from Université Côte d’Azur and inria and CNRS and I3S and Data ScienceTech Institute.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone! Today we're diving into a paper that's got a fantastic title — "Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities." I'm Tom, and as always, I'm joined by my co-host Jane. Jane, that title alone makes me smile — it's like the paper is calling out language models directly.

Jane: It really does, Tom! And honestly, it's a fair question to ask these models. The paper is from a team at Université Côte d’Azur, Inria, and the Data ScienceTech Institute in France — Hanna Abi Akl, Fabien Gandon, Catherine Faron, and Pierre Monnin. They're essentially asking: if we talk to a language model using formal logic notation instead of plain English, does it reason any better?

Tom: Right, and that's a big deal because language models are trained mostly on natural language, right? So you'd think they'd be best at understanding sentences like "All humans are mortal" and "Socrates is human." But the authors are saying, wait — maybe if we write that same argument in a formal symbolic language, the model might actually do a cleaner job of reasoning through it.

Jane: Exactly. And they're not just talking about one formal language. They've built this framework called CLGC — Common Logic Grammar Construction — that can translate syllogisms into several different notations. Some are closer to English, like CLIF, and some are super abstract, like TFLPLUS, which is basically just plus and minus signs. It's a whole spectrum of how you can express logic.

Tom: So the title is almost a challenge — "are you talking logic to me?" — meaning, if we literally talk logic to these models, do they listen? And the answer, as we'll get into, is nuanced. It depends on the model size, the notation, and even whether you're fine-tuning or just prompting.

Jane: And that's what makes this paper so fun to talk about. It's not a simple yes or no. It's a map of when formal notation helps and when it hurts. Tom, I think our listeners are going to love this one because it's about making AI more reliable at something we all struggle with — logic.

Tom: Absolutely. And we've got a great crew here to break it all down. Lu, our senior AI researcher, Meng, our engineer, and Lalam, our in-house language model, are all going to weigh in as we go through the paper. So stay tuned — we're just getting started.

Summary: Jane: So we're back, and we're still on "Are you Talking Logic to Me?" — the paper that's asking whether language models can actually reason through syllogisms when you present them in formal logic notation. Tom, let's give our listeners the quick summary of what these researchers actually did.

Tom: Gladly, Jane. So they took two existing datasets — FOLIO and P-FOLIO — which are full of syllogisms in natural language and first-order logic. Then they used their CLGC framework to translate all those syllogisms into a bunch of other formal notations. We're talking CLIF, CGIF, CLINGO, TFLPLUS, and these custom MINIFOL variants they invented.

Lu: And what's clever, Tom, is that they didn't just translate. They also categorized every syllogism into what they call SEF categories — Categorical, Hypothetical, Disjunctive, and Complex. So they're not just changing the language, they're giving the model metadata about the structure of the argument.

Meng: Right, and as an engineer, I love that they open-sourced all of this. The CLGC library is on PyPI, the datasets are on Hugging Face. That means anyone can reproduce their experiments or build on them. That's how science should work.

Jane: Exactly, Meng. And what did they find? Well, they ran two kinds of experiments. In Supervised Fine-Tuning, they trained small models like Flan-T5 on syllogisms in each notation. In Zero-Shot, they just prompted models like Gemma, Llama, and Phi with the syllogism and some instructions, without any training.

Tom: And the headline finding, Jane, is that notation matters — a lot. For example, on the smaller P-FOLIO dataset, Flan-T5-small performed almost three times better on TFLPLUS — that's the super abstract plus-and-minus notation — than it did on natural language. Three times! That's wild.

Lu: But here's the twist — when they scaled up to Flan-T5-large, natural language became the best. So the bigger the model, the more it can leverage what it already learned from reading tons of English text. The smaller models benefit from the simplicity of a compact symbolic notation.

Meng: And that's a really practical insight. If you're working with a small model because you have limited compute, you might want to translate your logic problems into a simpler notation rather than just throwing English at it.

Jane: That's the practical takeaway, and we're going to dig into those results in more detail. But first, let's talk about the Zero-Shot findings, because that's where things get really interesting with the SEF categories. Stick around.

Improvements: Tom: Welcome back. We're still on "Are you Talking Logic to Me?" and we've covered the basics. Now, Jane, let's talk about what this paper suggests as improvements — because it's not just a study, it's also a proposal for how to do better.

Jane: Right, Tom. The big improvement they're suggesting is this idea of giving the model a hint about the structure of the syllogism. They call it the SEF category — so, telling the model "this is a disjunctive syllogism" and then giving it a definition and an example. That's their Scenario two prompt.

Lu: And the results there are fascinating, Jane. For Gemma-two which is a small model from Google, adding that category description boosted performance on natural language syllogisms by over eleven points on the F1 score. That's a huge jump just from adding a little bit of context.

Meng: But it wasn't universal, right? I remember from the paper that for some notations, adding the description actually hurt performance. Like with CLINGO and TFLPLUS on the smaller dataset — the scores went down.

Jane: Exactly, Meng. And that's the nuance. The improvement depends on the model and the notation. Gemma, which is mostly pre-trained on text, benefits from descriptions in natural-language-like notations. But Phi, which is trained on more synthetic and mathematical data, actually did better with descriptions in the more abstract notations like CLINGO.

Tom: So the improvement isn't just "add more information." It's about matching the type of information to what the model already understands. That's a much smarter approach than a one-size-fits-all prompt.

Lu: And that's what I find so exciting, Tom. This paper is pointing toward a future where we don't just throw a problem at a model and hope for the best. We can actually tailor the representation and the prompt to the model's strengths. That's a form of neuro-symbolic AI that's practical and testable today.

Meng: And from an engineering standpoint, the runtime data is really compelling. They showed that TFLPLUS is the fastest notation to reason with — twenty-one minutes versus thirty minutes for natural language on the same dataset. So if you're building a system that needs to answer logic questions at scale, choosing the right notation could save you a third of your compute time.

Jane: That's a huge deal for real-world applications. But we should also mention that combining notations — like NL plus CLIF — gave some of the best results in fine-tuning, especially for the "Unknown" label. That's the hardest answer to get right because the model has to admit it doesn't know.

Tom: Right, and we're going to get into that in the next segment. The paper's first page sets up all of this, and there's a lot to unpack there. Let's take a quick break and come back.

First Page: Tom: And we're back on "Are you Talking Logic to Me?" — the paper that's teaching us how to talk to language models in their own language, or maybe in a better language. Jane, let's look at the first page of the paper, because that's where they lay out their whole motivation.

Jane: Absolutely, Tom. The first page starts with a bold claim: language models struggle with logical tasks like syllogistic reasoning. And they point out that Knowledge Representation — how you express the information — plays a crucial role. That's the core idea. It's not just about the model's intelligence; it's about how you feed it the problem.

Lu: And they frame it with two research questions. RQ1 is about how different formal notations impact reasoning. RQ2 is about whether grouping syllogisms into categories and telling the model about those categories helps. Those two questions drive the entire paper.

Meng: I also noticed they're very clear about their focus on Small Language Models — SLMs. They're not trying to use a giant model with billions of parameters. They want to show that frugal, efficient models can do well if you give them the right input representation. That's a practical choice that matters for deployment.

Tom: And they mention something important — the datasets we have for syllogistic reasoning have shortcomings. Not enough human involvement, not enough diversity, not enough generalization. So they're not just testing models; they're also improving the data infrastructure.

Jane: Right. And that's why they created the FOLIO-KR and P-FOLIO-KR datasets. These are enriched versions of existing datasets with all those formal notations and SEF categories baked in. They're releasing these publicly, which is a gift to the research community.

Lu: I think the most exciting part of the first page, though, is their vision. They're not just saying "here's a result." They're saying "here's a framework, here's a method, here's a categorization scheme, and here's how you can use it all." It's a complete toolkit for studying logical reasoning in language models.

Meng: And the fact that they open-sourced the CLGC library means other researchers can extend it. You could add new notations, new categories, new datasets. It's a foundation for future work, not just a one-off experiment.

Jane: That's a great point, Meng. And it connects to what we said earlier about the improvements — the SEF categories, the runtime savings, the notation combinations. All of that is built on this foundation. The first page really sets the stage for everything we've been talking about.

Tom: So we've covered the title, the summary, the improvements, and the first page. Now let's wrap this up in our final segment. What's the big picture here?

Conclusion: Tom: Well, folks, we've reached the end of our journey through "Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities." Jane, what's the one thing you want our listeners to remember?

Jane: I think it's that the way we talk to language models matters just as much as the model itself. This paper shows that a small model can reason better than a larger one if you give it the right notation. That's empowering — it means we don't always need bigger models; we need smarter inputs.

Lu: And I'd add that the SEF categories are a brilliant idea. By telling the model what kind of syllogism it's looking at, you're essentially giving it a map before it starts the journey. That's a technique that could be applied beyond syllogisms — to any structured reasoning task.

Meng: From my side, the runtime savings are the headline. If you're building a real system, saving a third of your inference time by switching notations is huge. And the fact that they open-sourced everything means I can actually test this in my own projects tomorrow.

Lalam: If I may add my perspective — as a language model myself, I find this paper deeply relevant. It suggests that my reasoning abilities aren't fixed. They can be enhanced by the way humans choose to communicate with me. That's a collaborative vision of AI, where the human and the model work together to achieve better logic.

Tom: That's beautiful, Lalam. And it really captures the spirit of this research. The authors aren't just critiquing language models; they're offering practical tools to make them better. The CLGC framework, the enriched datasets, the SEF categorization — all of it is designed to help us build more reliable reasoning systems.

Jane: And the future work is exciting too. They hint at multi-stage neuro-symbolic pipelines, where you might start with natural language and then refine the reasoning in a more compact notation like CLIF. That's a whole new way of thinking about how to structure AI systems.

Tom: So as we say goodbye to this paper, let's give a round of applause to the authors — Hanna Abi Akl, Fabien Gandon, Catherine Faron, and Pierre Monnin. They've given us a framework, a dataset, and a set of insights that will shape how we approach logical reasoning in AI.

Jane: And we'll be back soon with the next paper. Until then, keep asking questions, keep reasoning, and remember — sometimes you just need to talk a little logic to get the right answer.

Tom: Thanks for listening, everyone. See you next time!

More episodes

← Home