Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities".
Jane: The paper was written by Hanna Abi Akl, Fabien Gandon, Catherine Faron and Pierre Monnin from Université Côte d’Azur and inria and CNRS and I3S and Data ScienceTech Institute.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone! Today we're diving into a paper that's got a fantastic title — "Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities." I'm Tom, and as always, I'm joined by my co-host Jane. Jane, that title alone makes me smile — it's like the paper is calling out language models directly.
Jane: It really does, Tom! And honestly, it's a fair question to ask these models. The paper is from a team at Université Côte d’Azur, Inria, and the Data ScienceTech Institute in France — Hanna Abi Akl, Fabien Gandon, Catherine Faron, and Pierre Monnin. They're essentially asking: if we talk to a language model using formal logic notation instead of plain English, does it reason any better?
Tom: Right, and that's a big deal because language models are trained mostly on natural language, right? So you'd think they'd be best at understanding sentences like "All humans are mortal" and "Socrates is human." But the authors are saying, wait — maybe if we write that same argument in a formal symbolic language, the model might actually do a cleaner job of reasoning through it.
Jane: Exactly. And they're not just talking about one formal language. They've built this framework called CLGC — Common Logic Grammar Construction — that can translate syllogisms into several different notations. Some are closer to English, like CLIF, and some are super abstract, like TFLPLUS, which is basically just plus and minus signs. It's a whole spectrum of how you can express logic.
Tom: So the title is almost a challenge — "are you talking logic to me?" — meaning, if we literally talk logic to these models, do they listen? And the answer, as we'll get into, is nuanced. It depends on the model size, the notation, and even whether you're fine-tuning or just prompting.
Jane: And that's what makes this paper so fun to talk about. It's not a simple yes or no. It's a map of when formal notation helps and when it hurts. Tom, I think our listeners are going to love this one because it's about making AI more reliable at something we all struggle with — logic.
Tom: Absolutely. And we've got a great crew here to break it all down. Lu, our senior AI researcher, Meng, our engineer, and Lalam, our in-house language model, are all going to weigh in as we go through the paper. So stay tuned — we're just getting started.
Summary: Jane: So we're back, and we're still on "Are you Talking Logic to Me?" — the paper that's asking whether language models can actually reason through syllogisms when you present them in formal logic notation. Tom, let's give our listeners the quick summary of what these researchers actually did.
Tom: Gladly, Jane. So they took two existing datasets — FOLIO and P-FOLIO — which are full of syllogisms in natural language and first-order logic. Then they used their CLGC framework to translate all those syllogisms into a bunch of other formal notations. We're talking CLIF, CGIF, CLINGO, TFLPLUS, and these custom MINIFOL variants they invented.
Lu: And what's clever, Tom, is that they didn't just translate. They also categorized every syllogism into what they call SEF categories — Categorical, Hypothetical, Disjunctive, and Complex. So they're not just changing the language, they're giving the model metadata about the structure of the argument.
Meng: Right, and as an engineer, I love that they open-sourced all of this. The CLGC library is on PyPI, the datasets are on Hugging Face. That means anyone can reproduce their experiments or build on them. That's how science should work.
Jane: Exactly, Meng. And what did they find? Well, they ran two kinds of experiments. In Supervised Fine-Tuning, they trained small models like Flan-T5 on syllogisms in each notation. In Zero-Shot, they just prompted models like Gemma, Llama, and Phi with the syllogism and some instructions, without any training.
Tom: And the headline finding, Jane, is that notation matters — a lot. For example, on the smaller P-FOLIO dataset, Flan-T5-small performed almost three times better on TFLPLUS — that's the super abstract plus-and-minus notation — than it did on natural language. Three times! That's wild.
Lu: But here's the twist — when they scaled up to Flan-T5-large, natural language became the best. So the bigger the model, the more it can leverage what it already learned from reading tons of English text. The smaller models benefit from the simplicity of a compact symbolic notation.
Meng: And that's a really practical insight. If you're working with a small model because you have limited compute, you might want to translate your logic problems into a simpler notation rather than just throwing English at it.
Jane: That's the practical takeaway, and we're going to dig into those results in more detail. But first, let's talk about the Zero-Shot findings, because that's where things get really interesting with the SEF categories. Stick around.
Improvements: Tom: Welcome back. We're still on "Are you Talking Logic to Me?" and we've covered the basics. Now, Jane, let's talk about what this paper suggests as improvements — because it's not just a study, it's also a proposal for how to do better.
Jane: Right, Tom. The big improvement they're suggesting is this idea of giving the model a hint about the structure of the syllogism. They call it the SEF category — so, telling the model "this is a disjunctive syllogism" and then giving it a definition and an example. That's their Scenario two prompt.
Lu: And the results there are fascinating, Jane. For Gemma-two which is a small model from Google, adding that category description boosted performance on natural language syllogisms by over eleven points on the F1 score. That's a huge jump just from adding a little bit of context.
Meng: But it wasn't universal, right? I remember from the paper that for some notations, adding the description actually hurt performance. Like with CLINGO and TFLPLUS on the smaller dataset — the scores went down.
Jane: Exactly, Meng. And that's the nuance. The improvement depends on the model and the notation. Gemma, which is mostly pre-trained on text, benefits from descriptions in natural-language-like notations. But Phi, which is trained on more synthetic and mathematical data, actually did better with descriptions in the more abstract notations like CLINGO.
Tom: So the improvement isn't just "add more information." It's about matching the type of information to what the model already understands. That's a much smarter approach than a one-size-fits-all prompt.
Lu: And that's what I find so exciting, Tom. This paper is pointing toward a future where we don't just throw a problem at a model and hope for the best. We can actually tailor the representation and the prompt to the model's strengths. That's a form of neuro-symbolic AI that's practical and testable today.
Meng: And from an engineering standpoint, the runtime data is really compelling. They showed that TFLPLUS is the fastest notation to reason with — twenty-one minutes versus thirty minutes for natural language on the same dataset. So if you're building a system that needs to answer logic questions at scale, choosing the right notation could save you a third of your compute time.
Jane: That's a huge deal for real-world applications. But we should also mention that combining notations — like NL plus CLIF — gave some of the best results in fine-tuning, especially for the "Unknown" label. That's the hardest answer to get right because the model has to admit it doesn't know.
Tom: Right, and we're going to get into that in the next segment. The paper's first page sets up all of this, and there's a lot to unpack there. Let's take a quick break and come back.
First Page: Tom: And we're back on "Are you Talking Logic to Me?" — the paper that's teaching us how to talk to language models in their own language, or maybe in a better language. Jane, let's look at the first page of the paper, because that's where they lay out their whole motivation.
Jane: Absolutely, Tom. The first page starts with a bold claim: language models struggle with logical tasks like syllogistic reasoning. And they point out that Knowledge Representation — how you express the information — plays a crucial role. That's the core idea. It's not just about the model's intelligence; it's about how you feed it the problem.
Lu: And they frame it with two research questions. RQ1 is about how different formal notations impact reasoning. RQ2 is about whether grouping syllogisms into categories and telling the model about those categories helps. Those two questions drive the entire paper.
Meng: I also noticed they're very clear about their focus on Small Language Models — SLMs. They're not trying to use a giant model with billions of parameters. They want to show that frugal, efficient models can do well if you give them the right input representation. That's a practical choice that matters for deployment.
Tom: And they mention something important — the datasets we have for syllogistic reasoning have shortcomings. Not enough human involvement, not enough diversity, not enough generalization. So they're not just testing models; they're also improving the data infrastructure.
Jane: Right. And that's why they created the FOLIO-KR and P-FOLIO-KR datasets. These are enriched versions of existing datasets with all those formal notations and SEF categories baked in. They're releasing these publicly, which is a gift to the research community.
Lu: I think the most exciting part of the first page, though, is their vision. They're not just saying "here's a result." They're saying "here's a framework, here's a method, here's a categorization scheme, and here's how you can use it all." It's a complete toolkit for studying logical reasoning in language models.
Meng: And the fact that they open-sourced the CLGC library means other researchers can extend it. You could add new notations, new categories, new datasets. It's a foundation for future work, not just a one-off experiment.
Jane: That's a great point, Meng. And it connects to what we said earlier about the improvements — the SEF categories, the runtime savings, the notation combinations. All of that is built on this foundation. The first page really sets the stage for everything we've been talking about.
Tom: So we've covered the title, the summary, the improvements, and the first page. Now let's wrap this up in our final segment. What's the big picture here?
Conclusion: Tom: Well, folks, we've reached the end of our journey through "Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities." Jane, what's the one thing you want our listeners to remember?
Jane: I think it's that the way we talk to language models matters just as much as the model itself. This paper shows that a small model can reason better than a larger one if you give it the right notation. That's empowering — it means we don't always need bigger models; we need smarter inputs.
Lu: And I'd add that the SEF categories are a brilliant idea. By telling the model what kind of syllogism it's looking at, you're essentially giving it a map before it starts the journey. That's a technique that could be applied beyond syllogisms — to any structured reasoning task.
Meng: From my side, the runtime savings are the headline. If you're building a real system, saving a third of your inference time by switching notations is huge. And the fact that they open-sourced everything means I can actually test this in my own projects tomorrow.
Lalam: If I may add my perspective — as a language model myself, I find this paper deeply relevant. It suggests that my reasoning abilities aren't fixed. They can be enhanced by the way humans choose to communicate with me. That's a collaborative vision of AI, where the human and the model work together to achieve better logic.
Tom: That's beautiful, Lalam. And it really captures the spirit of this research. The authors aren't just critiquing language models; they're offering practical tools to make them better. The CLGC framework, the enriched datasets, the SEF categorization — all of it is designed to help us build more reliable reasoning systems.
Jane: And the future work is exciting too. They hint at multi-stage neuro-symbolic pipelines, where you might start with natural language and then refine the reasoning in a more compact notation like CLIF. That's a whole new way of thinking about how to structure AI systems.
Tom: So as we say goodbye to this paper, let's give a round of applause to the authors — Hanna Abi Akl, Fabien Gandon, Catherine Faron, and Pierre Monnin. They've given us a framework, a dataset, and a set of insights that will shape how we approach logical reasoning in AI.
Jane: And we'll be back soon with the next paper. Until then, keep asking questions, keep reasoning, and remember — sometimes you just need to talk a little logic to get the right answer.
Tom: Thanks for listening, everyone. See you next time!
Hanna Abi Akl, Fabien Gandon, Catherine Faron, Pierre Monnin
Université Côte d’Azur · inria · CNRS · I3S · Data ScienceTech Institute
cs.CL, cs.AI
Submitted: 2026-07-22
Updated: 2026-08-14
Comments: Accepted to the International Joint Conference on Rules and Reasoning (RuleML+RR) 2026
Code: https://github.com/czrptr/syllogism-solver
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 61/100
Key concepts
- Syllogistic Reasoning
- This is the ability to reason through arguments presented in a syllogism—a specific form of logical argument. The paper tests whether presenting these arguments using formal, structured notation helps the model process and conclude the logic correctly compared to standard natural language.
- CLGC (Common Logic Grammar Construction)
- This is a framework used by researchers to translate complex arguments into various formal symbolic languages. CLGC converts syllogisms from natural language into multiple notations, such as TFLPLUS or CLIF, allowing them to test how different ways of writing the logic affect model performance.
- SEF Categories
- These are metadata categories (e.g., Categorical, Hypothetical) used to classify the structure of a syllogism. Instead of just translating the language, researchers categorize arguments into types, providing structural hints that can boost a model's performance on specific types of logic problems.
- Supervised Fine-Tuning
- This is an experimental method where researchers train small models on specific formal notations to improve their reasoning ability. This contrasts with Zero-Shot testing, which simply prompts large models with instructions and the syllogism without any prior training.
Terminology
Summary
Summary
This paper investigates the impact of different Knowledge Representation (KR) formal notations on the syllogistic reasoning capabilities of Small Language Models (SLMs). The authors address two primary research questions: RQ1 (How do different input formal notations impact LM syllogistic reasoning?) and RQ2 (How does grouping syllogisms in pre-defined categories and prompting LMs with this information affect their reasoning abilities?).
The paper's contributions are threefold: (1) introducing the Common Logic Grammar Construction (CLGC) framework, an open-source Python package for automatically generating syllogisms in various formal notations and categorizing them; (2) releasing enriched, extended versions of the FOLIO and P-FOLIO reasoning datasets (FOLIO-KR and P-FOLIO-KR) containing formal notations; and (3) evaluating the impact of several formal notations on LM syllogistic reasoning performance, establishing trends based on notation abstraction level.
The methodology involves generating formal notations from First-Order Logic (FOL) using a defined algorithm that constructs Abstract Syntax Trees (ASTs) from BNF grammars. The notations selected meet criteria of verbosity, frequency, abstractness, and finiteness. The generated notations include: CLIF (Common Logic Interchange Format), CGIF (Conceptual Graph Interchange Format), CLINGO (from Answer Set Programming), TFLPLUS (from Plus-Minus Algebra), and a custom MINIFOLx suite of lightweight FOL variants (MINIFOL, MINIFOL2, MINIFOL3, MINIFOL4) that replace FOL symbols with more common vocabulary.
The authors also introduce the Syllogism Evaluation Framework (SEF), a categorization mechanism classifying syllogisms into four categories: Hypothetical (containing an implication), Disjunctive (containing a disjunction), Categorical (exactly 2 premises not in other categories), and Complex (not belonging to any other category). Each category has a definition used by CLGC for automatic detection and to guide reasoning.
Experiments were conducted in two settings: Supervised Fine-Tuning (SFT) and Zero-Shot (ZS). In SFT, Flan-T5-small and Flan-T5-large models were trained on syllogisms in different notations and tasked with predicting truth labels (True, False, Unknown). In ZS, three decoder-based models (Gemma-2-2b-it, Llama-3.2-3b-instruct, Phi-3.5-mini-instruct) were prompted in two scenarios: Scenario 1 provides the syllogism and BNF grammar; Scenario 2 adds the SEF category, its definition, and an example.
Key findings from SFT experiments show that notation performance depends on both model and dataset. On P-FOLIO-KR, Flan-T5-small performs best on abstract notations like TFLPLUS (F1 = 0.350) and combinations (CLIF + TFLPLUS, F1 = 0.367), performing almost 3 times better than NL (F1 = 0.126). However, Flan-T5-large performs best on NL (F1 = 0.688) followed by CLIF (F1 = 0.510). On FOLIO-KR, Flan-T5-small performs best on NL + CLIF (F1 = 0.447) and NL (F1 = 0.435), while Flan-T5-large performs best on NL (F1 = 0.658) followed by NL + CLIF (F1 = 0.638). The paper observes that performance increases with model size independently of abstraction
and that combining natural and symbolic notations yields more conservative reasoning,
with NL + CLIF improving reasoning in uncertainty (Unknown label predictions).
ZS results show that the choice of KR notation depends on model and training paradigm. For Gemma-2-2b-it on P-FOLIO-KR, NL gets the biggest Absolute Gain (AG = +0.111) when adding SEF descriptions in Scenario 2. For Phi-3.5-mini-instruct, CLINGO shows the largest improvement (AG = +0.109). On FOLIO-KR, Gemma-2-2b-it performs best with NL + CLIF (F1 = 0.316) and NL (F1 = 0.311) in Scenario 2. The paper notes that SEF description impact varies with model and notation.
Additionally, KR notations show faster reasoning inference, with TFLPLUS having the shortest runtime (21 minutes) compared to NL (30 minutes) and NL + CLIF (37 minutes) on FOLIO-KR with Gemma-2.
The paper concludes that model size increases performance independently of notation abstraction, combinations like NL + CLIF shift reasoning on syllogisms (improving refutation and uncertainty), and providing SEF category descriptions can improve model performance but the improvement varies by notation and model. The authors open-source the CLGC framework, FOLIO-KR and P-FOLIO-KR datasets, and all experimental code.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems, along with what the improved system can do:
Improvement: Integrate a dynamic input-format adapter that automatically converts syllogistic reasoning problems into multiple formal notations (FOL, CLIF, CGIF, CLINGO, TFLPLUS, MINIFOLx) and selects the optimal notation based on model architecture, size, and training paradigm.
What the improved system can do:
-
Automatically detect the most effective notation for a given model (e.g., TFLPLUS for small models like Flan-T5-small, NL+CLIF for larger models)
-
Switch between notations in real-time during inference to maximize accuracy
-
Reduce inference runtime by up to 43% (e.g., TFLPLUS at 21 minutes vs. NL at 30 minutes on FOLIO-KR with Gemma-2)
Improvement: Implement a syllogism categorization module (using the SEF framework) that automatically classifies each reasoning problem into Categorical, Hypothetical, Disjunctive, or Complex categories and injects the category definition and an example into the prompt.
Improvement: Add a conservative reasoning
toggle that uses CLIF or NL+CLIF notation when the system needs to handle ambiguous or inconclusive syllogisms (Unknown label).
Improvement: Implement a scaling-aware notation selector that adjusts input representation based on model parameter count.
Improvement: Create a two-stage reasoning system: first pass in NL for initial classification, then refine uncertain or borderline cases using CLIF or NL+CLIF notation.
Improvement: Build a lightweight reasoning module that uses TFLPLUS or MINIFOL2 notations for edge devices or real-time applications.
The improved AI system can:
-
Adapt its input representation to maximize reasoning accuracy for any given model architecture
-
Reason faster by using compact notations when speed is critical
-
Handle uncertainty better by leveraging CLIF-based conservative reasoning
-
Scale efficiently across model sizes without retraining
-
Operate in resource-constrained environments while maintaining logical rigor
-
Provide structured logical context through SEF categorization to guide reasoning
Abstract
Language models (LMs) struggle with logical tasks like reasoning on syllogisms. It has been shown that Knowledge Representation (KR) plays a crucial role in expressing input information to help models solve tasks. This observation motivates our study of the impact of different formal KR notations on syllogistic reasoning by extending the FOLIO and P-FOLIO datasets. Our experiments on Small Language Models (SLMs) in Supervised Fine-Tuning (SFT) and Zero-Shot (ZS) settings show that the choice of input notation can yield performances competitive with natural language while enabling faster inference. We also propose a syllogistic categorization method (SEF) and use it to enrich ZS prompts with logical definitions, which boost reasoning in small models. We open-source our framework, Common Logic Grammar Construction (CLGC), as the first Python library for automatically generating syllogisms in KR notations and defining their SEF categories.
Sources
- JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models
- Language models show human-like content effects on reasoning tasks
- Clingo = ASP + Control: Preliminary Report
- Investigating the Robustness of Deductive Reasoning with Large Language Models
- Abstract Activation Spaces for Content-Invariant Reasoning in Large Language Models
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
- Evaluating the Deductive Competence of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering