NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution

summary

Video file (mp4)

The gist

The NOTAI.AI framework addresses a critical limitation in machine-generated text detection: while current models achieve high classification accuracy, they often operate as "black boxes," failing to

In short

The episode discusses 'NOTAI.AI,' a system for detecting machine-generated text using a multi-signal ensemble model. Hosts explain that its revolutionary feature is moving beyond simple probability scores by providing detailed, evidence-based explanations of *why* the text was flagged as AI, increasing transparency and trust.

Key concepts

SHAP (Shapley Additive Explanations)
A method used in the system to mathematically quantify how much each component, such as Conditional Probability Curvature or Type-Token Ratio, contributes to the final detection decision. This allows users to see the precise evidence driving the score.
LLM-based explainer
This feature takes complex mathematical attributions (like raw SHAP numbers) and translates them into clear, plain language rationales. It constructs a coherent narrative justification memo that non-experts can immediately understand.

Terminology used across episodes

This episode discusses

The paper

NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution · Read on arXiv

University of Luxembourg Department of Computer Science, University of Luxembourg · University of Luxembourg

We present NotAI.AI, an explainable AI-generated text detection system. Instead of returning only a binary label or confidence score, the system shows which signals influenced the prediction and lets users inspect an attribution-based sensitivity estimate obtained by subtracting selected local contributions. NotAI.AI combines sentence-level conditional probability curvature, a neural detector score, and interpretable stylometric and readability features in an XGBoost meta-classifier. It explains predictions with TreeSHAP feature contributions and can turn the resulting evidence into a concise natural-language explanation. We evaluate the system on a category-balanced subset of RAID containing human-written, clean AI-generated, and attacked AI-generated texts. The full model outperforms variants based on individual feature families, reaching 0.9685 F1 on the held-out within-subset test split. In an automatic evaluation, two model judges rate 94.5-98.6% of generated explanations as faithful to the supplied detector evidence. The web interface (https://notai-ai.vercel.app), source code (https://github.com/Oleksandr-MB/EMNLP2026DEMO NotAI.AI), and demonstration video(https://youtu.be/l8Nk8kdBTHQ) are publicly available.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution".

Jane: The paper was written by Oleksandr Marchenko Breneur, Adelaide Danilov and Aria Nourbakhsh Salima Lamsiyah from University of Luxembourg Department of Computer Science, University of Luxembourg and University of Luxembourg.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Paper discussion segment 1: Tom: We've just finished discussing the general summary of "NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution," establishing that the model uses a multi-signal ensemble to detect structural patterns. Now, we want to dig into the specific design choices that make this system so revolutionary for listeners.

Jane: The biggest leap forward, as Jane noted earlier, is moving beyond simply presenting a probability score; it gives us a detailed breakdown of *why* the model decided something was machine-generated.

Lu: That’s achieved through their use of SHAP—Shapley Additive Explanations—which allows them to mathematically quantify exactly how much each component, like Conditional Probability Curvature or Type-Token Ratio, contributes to the final decision.

Meng: It's a massive improvement in practical utility because it lets us see which features are driving the score high or low. This makes the system highly transparent and incredibly valuable for real-time auditing of content, allowing us to check the evidence trail.

Lalam: Lalam finds this level of transparency crucial; we aren't just looking at an outcome, we're seeing a fingerprint, which helps us understand the subtle characteristics of digital creation rather than just making a blind guess about its origin.

Tom: You mentioned fingerprints, Lalam—and Lu was talking about quantifying contribution; so you are literally able to see the evidence for machine generation in this system. What does that practical visualization look like?

Jane: Exactly, and it doesn't stop at just listing the metrics. They use an LLM-based explainer to take those complex mathematical attributions and translate them into clear, plain language rationales that a non-expert can understand immediately.

Lu: I find that translation step fascinating; it bridges the gap between high-level statistical rigor and accessible human communication in a way that is genuinely innovative for this field. It’s translating math into narrative evidence.

Meng: From an implementation standpoint, having this automated natural language output means we can deploy this tool to non-experts without needing a second layer of analysis, which simplifies the user experience tremendously for adoption.

Lalam: It’s about giving users the authority to look at a piece of text and have a clear, evidence-based conversation with the machine about its origin, rather than just accepting an opaque verdict.

Tom: That's right, Jane; it moves us away from accepting black-box results and toward providing actionable information. Now that we know how this system provides such detailed evidence through attribution, let's see if its performance actually lives up to the hype in the next segment.

Improvements and Specific Features: Tom: We've seen how NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution is built, establishing its complex, multi-signal ensemble model. Now we need to focus on the most revolutionary aspect: its improvements in explainability beyond just listing the metrics.

Jane: It’s not enough to get a score; you must understand the reasoning behind that score. The standout feature, which I think is truly groundbreaking, is their integration of methods that translate complexity into comprehension.

Lu: To build on Jane's point about comprehension, Lu wants to reiterate how SHAP values are utilized here—they don't just show contribution; they provide a mathematically sound foundation for *why* the model weights certain features more heavily than others in the final decision.

Meng: From an engineering standpoint, this systematic explanation is revolutionary because it allows developers to pinpoint exactly which component—whether it's the Curvature metric or perhaps the Type-Token Ratio—is failing or succeeding in its detection task, enabling targeted improvements.

Lalam: What I find so valuable here is that this level of explanation builds trust. When a system can show its work, it shifts the conversation from "Is it fake?" to "Here is the evidence showing why you think it's machine-generated."

Tom: That concept of building trust through transparency is critical, Lalam. Jane, how does this LLM-based explainer actually perform that translation? Is it just summarizing, or is it doing something deeper with the mathematics?

Jane: It’s much deeper than summarization. The explainer takes the raw mathematical attributions—the numbers from SHAP—and constructs a coherent narrative. It’s essentially writing a justification memo based on statistical evidence.

Lu: And that narrative framing is what elevates it for the user. Instead of presenting a table of coefficients, the user reads, "The high curvature contributed significantly because..." which makes the science immediately actionable for non-statisticians.

Meng: This automation means that deployment is much simpler; we aren't requiring specialized analysts to interpret the output before a general editor can use it. The explanation *is* the user interface for the evidence.

Lalam: It gives users a powerful tool for discourse—the ability to challenge or confirm content based on quantifiable, articulated reasons rather than mere suspicion.

Tom: So we have moved from knowing *

Paper discussion segment 3: Tom: We’ve seen how NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution is built, but now we want to dig into what makes this specific design so revolutionary for the listeners.

Jane: The biggest leap forward is that it doesn't just provide a probability score; it gives us a detailed breakdown of *why* the model decided something was machine-generated. It really shows you the internal logic of AI.

Lu: That’s achieved through their use of SHAP—Shapley Additive Explanations—which allows them to mathematically quantify exactly how much each component, like Conditional Probability Curvature or Type-Token Ratio, contributes to the final decision.

Meng: It's a massive improvement in practical utility because it lets us see which features are driving the score high or low. This makes the system highly transparent and incredibly valuable for real-time auditing of content.

Lalam: Lalam finds this level of transparency crucial; we aren't just looking at an outcome, we're seeing a fingerprint, which helps us understand the subtle characteristics of digital creation rather than just making a blind guess about its origin.

Tom: You mentioned that fingerprints, Lalam—and Lu was talking about quantifying contribution; so you are literally able to see the evidence for machine generation in this system. What does that practical visualization look like?

Jane: Exactly, and it doesn's stopping there. They use an LLM-based explainer to take those complex mathematical attributions and translate them into clear, plain language rationales that a non-expert can understand immediately.

Lu: I find that translation step fascinating; it’ bridge the gap between high-level statistical rigor and accessible human communication in a way that is genuinely innovative for this field. It' turns math into narrative evidence.

Meng: From an implementation standpoint, having this automated natural language output means we can deploy this tool to non-experts without needing a second layer of analysis, which simplifies the user experience tremendously.

Lalam: It’s about giving users the authority to look at a piece of text and have a clear, evidence-based conversation with the machine about its origin, rather than just accepting an opaque verdict.

Tom: That's right, Jane; it moves us away from accepting black-box results and toward providing actionable information.

Lu: I’m particularly excited about how this enables a deeper academic dive into the specific patterns—the subtle stylistic choices—that differentiate human creativity from algorithmic output.

Meng: It suggests that we can build real-world systems where every single piece of content is accompanied by a justification for its authenticity, which is a huge change in workflow.

Lalam: This shift allows us to foster a culture of critical reading, where the public isn't just told something *is* AI, but knows exactly *how* it was generated.

Tom: It sounds like they have built a very powerful, multifaceted machine for analysis that is genuinely understandable by everyone.

Conclusion: Tom: We've spent time exploring exactly how NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution works, showing that we can't just rely on a single score to determine if a piece of writing is artificial.

Jane: It’s clear that the core value here is moving beyond just providing a judgment; it offers demonstrable evidence, which fundamentally changes how we approach content integrity in the digital age.

Lu: From an academic perspective, this research opens up fascinating pathways for us to analyze the specific patterns and subtle stylistic choices that differentiate human creativity from algorithmic output.

Meng: I’m impressed by how well-integrated the system is—it suggests that this architecture is incredibly robust and ready to handle real-world scaling challenges when we move into production systems.

Lalam: Lalam believes the most significant impact will be on how organizations build trust, allowing us to foster a culture of critical reading based on quantifiable, verifiable evidence.

Tom: That shift toward actionable transparency is something truly remarkable, Jane.

Jane: It makes the whole concept feel so much more grounded in reality rather than some abstract theoretical exercise.

Lu: I'm particularly excited about how this model could be adapted to other forms of generative media beyond just text, extending its reach into different creative domains.

Meng: If they can maintain this level of explainability at scale, the commercial utility for a large number of industries is genuinely enormous.

Lalam: We just need these tools to become universally adopted so that the public discourse can benefit from this new level of assurance about what we are reading.

Tom: This is truly a powerful combination of advanced mathematics and practical utility, isn't it?

Jane: It’s absolutely a turning point where the "how" something was created matters just as much as the "what."

Tom: Thank you all for joining us on this deep dive into NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution; we're looking forward to our next segment when we explore ethics in image synthesis.

More episodes

← Home