NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution".
Jane: The paper was written by Oleksandr Marchenko Breneur, Adelaide Danilov and Aria Nourbakhsh Salima Lamsiyah from University of Luxembourg Department of Computer Science, University of Luxembourg and University of Luxembourg.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Paper discussion segment 1: Tom: We've just finished discussing the general summary of "NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution," establishing that the model uses a multi-signal ensemble to detect structural patterns. Now, we want to dig into the specific design choices that make this system so revolutionary for listeners.
Jane: The biggest leap forward, as Jane noted earlier, is moving beyond simply presenting a probability score; it gives us a detailed breakdown of *why* the model decided something was machine-generated.
Lu: That’s achieved through their use of SHAP—Shapley Additive Explanations—which allows them to mathematically quantify exactly how much each component, like Conditional Probability Curvature or Type-Token Ratio, contributes to the final decision.
Meng: It's a massive improvement in practical utility because it lets us see which features are driving the score high or low. This makes the system highly transparent and incredibly valuable for real-time auditing of content, allowing us to check the evidence trail.
Lalam: Lalam finds this level of transparency crucial; we aren't just looking at an outcome, we're seeing a fingerprint, which helps us understand the subtle characteristics of digital creation rather than just making a blind guess about its origin.
Tom: You mentioned fingerprints, Lalam—and Lu was talking about quantifying contribution; so you are literally able to see the evidence for machine generation in this system. What does that practical visualization look like?
Jane: Exactly, and it doesn't stop at just listing the metrics. They use an LLM-based explainer to take those complex mathematical attributions and translate them into clear, plain language rationales that a non-expert can understand immediately.
Lu: I find that translation step fascinating; it bridges the gap between high-level statistical rigor and accessible human communication in a way that is genuinely innovative for this field. It’s translating math into narrative evidence.
Meng: From an implementation standpoint, having this automated natural language output means we can deploy this tool to non-experts without needing a second layer of analysis, which simplifies the user experience tremendously for adoption.
Lalam: It’s about giving users the authority to look at a piece of text and have a clear, evidence-based conversation with the machine about its origin, rather than just accepting an opaque verdict.
Tom: That's right, Jane; it moves us away from accepting black-box results and toward providing actionable information. Now that we know how this system provides such detailed evidence through attribution, let's see if its performance actually lives up to the hype in the next segment.
Improvements and Specific Features: Tom: We've seen how NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution is built, establishing its complex, multi-signal ensemble model. Now we need to focus on the most revolutionary aspect: its improvements in explainability beyond just listing the metrics.
Jane: It’s not enough to get a score; you must understand the reasoning behind that score. The standout feature, which I think is truly groundbreaking, is their integration of methods that translate complexity into comprehension.
Lu: To build on Jane's point about comprehension, Lu wants to reiterate how SHAP values are utilized here—they don't just show contribution; they provide a mathematically sound foundation for *why* the model weights certain features more heavily than others in the final decision.
Meng: From an engineering standpoint, this systematic explanation is revolutionary because it allows developers to pinpoint exactly which component—whether it's the Curvature metric or perhaps the Type-Token Ratio—is failing or succeeding in its detection task, enabling targeted improvements.
Lalam: What I find so valuable here is that this level of explanation builds trust. When a system can show its work, it shifts the conversation from "Is it fake?" to "Here is the evidence showing why you think it's machine-generated."
Tom: That concept of building trust through transparency is critical, Lalam. Jane, how does this LLM-based explainer actually perform that translation? Is it just summarizing, or is it doing something deeper with the mathematics?
Jane: It’s much deeper than summarization. The explainer takes the raw mathematical attributions—the numbers from SHAP—and constructs a coherent narrative. It’s essentially writing a justification memo based on statistical evidence.
Lu: And that narrative framing is what elevates it for the user. Instead of presenting a table of coefficients, the user reads, "The high curvature contributed significantly because..." which makes the science immediately actionable for non-statisticians.
Meng: This automation means that deployment is much simpler; we aren't requiring specialized analysts to interpret the output before a general editor can use it. The explanation *is* the user interface for the evidence.
Lalam: It gives users a powerful tool for discourse—the ability to challenge or confirm content based on quantifiable, articulated reasons rather than mere suspicion.
Tom: So we have moved from knowing *
Paper discussion segment 3: Tom: We’ve seen how NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution is built, but now we want to dig into what makes this specific design so revolutionary for the listeners.
Jane: The biggest leap forward is that it doesn't just provide a probability score; it gives us a detailed breakdown of *why* the model decided something was machine-generated. It really shows you the internal logic of AI.
Lu: That’s achieved through their use of SHAP—Shapley Additive Explanations—which allows them to mathematically quantify exactly how much each component, like Conditional Probability Curvature or Type-Token Ratio, contributes to the final decision.
Meng: It's a massive improvement in practical utility because it lets us see which features are driving the score high or low. This makes the system highly transparent and incredibly valuable for real-time auditing of content.
Lalam: Lalam finds this level of transparency crucial; we aren't just looking at an outcome, we're seeing a fingerprint, which helps us understand the subtle characteristics of digital creation rather than just making a blind guess about its origin.
Tom: You mentioned that fingerprints, Lalam—and Lu was talking about quantifying contribution; so you are literally able to see the evidence for machine generation in this system. What does that practical visualization look like?
Jane: Exactly, and it doesn's stopping there. They use an LLM-based explainer to take those complex mathematical attributions and translate them into clear, plain language rationales that a non-expert can understand immediately.
Lu: I find that translation step fascinating; it’ bridge the gap between high-level statistical rigor and accessible human communication in a way that is genuinely innovative for this field. It' turns math into narrative evidence.
Meng: From an implementation standpoint, having this automated natural language output means we can deploy this tool to non-experts without needing a second layer of analysis, which simplifies the user experience tremendously.
Lalam: It’s about giving users the authority to look at a piece of text and have a clear, evidence-based conversation with the machine about its origin, rather than just accepting an opaque verdict.
Tom: That's right, Jane; it moves us away from accepting black-box results and toward providing actionable information.
Lu: I’m particularly excited about how this enables a deeper academic dive into the specific patterns—the subtle stylistic choices—that differentiate human creativity from algorithmic output.
Meng: It suggests that we can build real-world systems where every single piece of content is accompanied by a justification for its authenticity, which is a huge change in workflow.
Lalam: This shift allows us to foster a culture of critical reading, where the public isn't just told something *is* AI, but knows exactly *how* it was generated.
Tom: It sounds like they have built a very powerful, multifaceted machine for analysis that is genuinely understandable by everyone.
Conclusion: Tom: We've spent time exploring exactly how NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution works, showing that we can't just rely on a single score to determine if a piece of writing is artificial.
Jane: It’s clear that the core value here is moving beyond just providing a judgment; it offers demonstrable evidence, which fundamentally changes how we approach content integrity in the digital age.
Lu: From an academic perspective, this research opens up fascinating pathways for us to analyze the specific patterns and subtle stylistic choices that differentiate human creativity from algorithmic output.
Meng: I’m impressed by how well-integrated the system is—it suggests that this architecture is incredibly robust and ready to handle real-world scaling challenges when we move into production systems.
Lalam: Lalam believes the most significant impact will be on how organizations build trust, allowing us to foster a culture of critical reading based on quantifiable, verifiable evidence.
Tom: That shift toward actionable transparency is something truly remarkable, Jane.
Jane: It makes the whole concept feel so much more grounded in reality rather than some abstract theoretical exercise.
Lu: I'm particularly excited about how this model could be adapted to other forms of generative media beyond just text, extending its reach into different creative domains.
Meng: If they can maintain this level of explainability at scale, the commercial utility for a large number of industries is genuinely enormous.
Lalam: We just need these tools to become universally adopted so that the public discourse can benefit from this new level of assurance about what we are reading.
Tom: This is truly a powerful combination of advanced mathematics and practical utility, isn't it?
Jane: It’s absolutely a turning point where the "how" something was created matters just as much as the "what."
Tom: Thank you all for joining us on this deep dive into NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution; we're looking forward to our next segment when we explore ethics in image synthesis.
University of Luxembourg Department of Computer Science, University of Luxembourg · University of Luxembourg
cs.CL
Submitted: 2026-03-05
Updated: 2026-09-03
Comments: 11 pages, 5 figures
Code: https://github.com/Oleksandr-MB/EMNLP2026DEMO_NotAI.AI
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 86/100
The gist: The NOTAI.AI framework addresses a critical limitation in machine-generated text detection: while current models achieve high classification accuracy, they often operate as "black boxes," failing to
Key concepts
- SHAP (Shapley Additive Explanations)
- A method used in the system to mathematically quantify how much each component, such as Conditional Probability Curvature or Type-Token Ratio, contributes to the final detection decision. This allows users to see the precise evidence driving the score.
- LLM-based explainer
- This feature takes complex mathematical attributions (like raw SHAP numbers) and translates them into clear, plain language rationales. It constructs a coherent narrative justification memo that non-experts can immediately understand.
Terminology
Summary
The NOTAI.AI framework addresses a critical limitation in machine-generated text detection: while current models achieve high classification accuracy, they often operate as black boxes,
failing to provide transparent evidence for their claims. This work introduces an explainable approach that moves beyond simple binary detection by integrating advanced mathematical concepts—specifically probability curvature and feature attribution—to not only classify text as machine-generated or human-written but also to generate a clear, inspectable rationale for that decision. This capability is vital for real-world deployment, ensuring that the system's outputs are better aligned with real-world use by non-experts.
Theoretical Foundation: Curvature Detection
The core detection mechanism relies on analyzing the underlying probability distribution of generated text. The authors leverage concepts related to information geometry, specifically calculating the Fisher Information Metric (FIM) or related measures of curvature. The paper posits that machine-generated text, due to its constrained statistical nature derived from large language models (LLMs), exhibits a measurable deviation in its local probability manifold compared to naturally occurring human writing. This difference is quantified by analyzing the curvature
of the text's embedding space. The detection process involves:
-
Calculating the curvature tensor across key segments of the input text.
-
Comparing this measured curvature against established baselines derived from known human and machine corpora (e.g., using benchmarks like RAID).
-
A significant deviation in expected curvature serves as a primary indicator of potential generation, providing a mathematically grounded score rather than merely relying on statistical patterns alone.
Explainable Detection via Feature Attribution
To transform the detection score into an actionable, understandable insight, NOTAI.AI incorporates sophisticated Explainable AI (XAI) techniques, primarily utilizing SHAP-based feature attributions. This module addresses the lack of interpretability inherent in complex neural networks. Instead of simply outputting a probability score, the system determines which specific features within the text contributed most significantly to the final classification decision. The methodology translates these technical attributions into natural language rationales, allowing users to understand why a piece of text was flagged.
The attribution process focuses on identifying:
-
Key Linguistic Markers: Specific phrases or word sequences that exhibit anomalous statistical properties (e.g., overly predictable transitions or unusual n-gram frequencies).
-
Structural Weaknesses: Patterns in syntax or coherence that deviate from human writing norms, such as highly uniform sentence length distribution or repetitive stylistic choices common to LLMs.
The NOTAI.AI Interactive System and Deployment
The final component is the user interface, designed to make complex detection science accessible to non-experts. The system synthesizes the results from the curvature module and the feature attribution module into a cohesive, interactive narrative. This design ensures that users do not merely receive a score but are guided through a transparent decision-making process.
The system provides several critical outputs for end-users:
-
Detection Score: A quantitative measure of machine likelihood.
-
Rationale Summary: A plain-language explanation detailing the primary reasons for the score (e.g.,
The text was flagged because of its consistently low lexical diversity in paragraph three
). -
Attribution Highlights: Visually highlighting specific segments of the text that are responsible for the detection outcome, thereby providing a clear audit trail for the classification decision.
By integrating these components, NOTAI.AI achieves a paradigm shift from opaque detection methods to transparent, inspectable decision-making,
making it a robust tool for academic review and content provenance verification.
Improvements for AI systems
(Internal Monologue: The proposed system, N OTAI.AI, has a strong foundation by combining ensemble detection with explainability. However, relying solely on supervised detector signals is inherently brittle. My focus must be on hardening the system against known failure modes—domain shift and adversarial manipulation—while elevating the explainability from mere feature attribution to true linguistic interpretability.)
1. Implementation of Domain-Agnostic Meta-Learning Module (DAMLM) for Detector Ensemble:
The current meta-classifier structure must be augmented with a meta-learning loop. Instead of treating domain shifts as novel failure modes, the system will be trained to quickly adapt its weighting schema across diverse, unseen generator distributions (D unseen). This involves using Model-Agnostic Meta-Learning (MAML) or similar few-shot adaptation techniques on the ensemble weights themselves.
- What the improved AI system can do: The resulting system will achieve zero-shot domain generalization. When presented with text generated by a novel model (e.g., a future, proprietary LLM not included in the training set) or operating under a significant stylistic shift (e.g., shifting from journalistic prose to highly technical scientific writing), the DAMLM can dynamically fine-tune the optimal combination of statistical and neural detector signals with minimal labeled examples, maintaining high detection accuracy (F 1 0.95) across cross-domain benchmarks like M4 or future extensions of RAID.
2. Integration of Counterfactual Feature Attribution (CFA) for Enhanced Explainability:
The current SHAP-based attributions provide what features are important, but not why they signal generation. We must evolve the explainability layer by integrating counterfactual reasoning into the attribution pipeline. Instead of merely reporting high feature weights, the system will generate a minimal set of synthetic textual edits (e.g., paraphrasing specific n-grams or replacing structural markers) that, when applied to the input text, are predicted to nullify the detection score.
- What the improved AI system can do: The system will provide Actionable Detection Rationales. For a non-expert user, instead of seeing
High weight on function words,
they will see: "The text exhibits a statistically low entropy in its use of transitional phrases (e.g., 'Furthermore,' 'Moreover'), which is characteristic of LLM output. To make this text appear human-written, consider replacing the phrase 'In conclusion' with a more varied transition like 'Ultimately.'This moves the tool from being merely an
inspectorto a sophisticated
linguistic diagnostic advisor."
3. Adversarial Robustness Training via Gradient Masking and Perturbation Injection:
To counter adversarial paraphrasing (where human-like edits are made specifically to fool the detector), the training regimen must incorporate both gradient masking and targeted perturbation injection during the meta-classifier's loss calculation. We will utilize a secondary, small auxiliary model trained explicitly to identify semantic perturbations that preserve the original meaning but maximize the distance in feature space from known human text distributions.
- What the improved AI system can do: The system will achieve Certified Detection Robustness. It will not only detect machine-generated text but will also provide a quantifiable measure of its robustness score against adversarial attacks. If the input text is flagged as suspicious, the system can report: "Detection Score: 0.92 (High). Predicted Robustness Margin against paraphrasing attacks: epsilon > 0.15 (Indicates high confidence that minor human edits will not successfully evade detection)." This level of quantifiable uncertainty reporting is critical for high-stakes applications.
Abstract
We present NotAI.AI, an explainable AI-generated text detection system. Instead of returning only a binary label or confidence score, the system shows which signals influenced the prediction and lets users inspect an attribution-based sensitivity estimate obtained by subtracting selected local contributions. NotAI.AI combines sentence-level conditional probability curvature, a neural detector score, and interpretable stylometric and readability features in an XGBoost meta-classifier. It explains predictions with TreeSHAP feature contributions and can turn the resulting evidence into a concise natural-language explanation. We evaluate the system on a category-balanced subset of RAID containing human-written, clean AI-generated, and attacked AI-generated texts. The full model outperforms variants based on individual feature families, reaching 0.9685 F1 on the held-out within-subset test split. In an automatic evaluation, two model judges rate 94.5-98.6% of generated explanations as faithful to the supplied detector evidence. The web interface (https://notai-ai.vercel.app), source code (https://github.com/Oleksandr-MB/EMNLP2026DEMO NotAI.AI), and demonstration video(https://youtu.be/l8Nk8kdBTHQ) are publicly available.
Sources
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Machine-generated text detection prevents language model collapse
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering