ABSENTIA: Detecting Broken Access Control Vulnerabilities in Web Applications

arXiv:2610.00977 · cs.CR, cs.SE · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "ABSENTIA: Detecting Broken Access Control Vulnerabilities in Web Applications".

Elias: Broken access control, which involves authorization failures where a principal acts on an unauthorized resource,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So, we’re starting with this paper titled "ABSENTIA: Detecting Broken Access Control Vulnerabilities in Web Applications." It looks like this work tackles the issue of broken access control, which is pretty prevalent in web security risks. The core idea seems to be that authorization isn't a simple data flow problem; it's more like a relation between who can do what on what resource. This makes it hard for traditional tools to find the flaws because no single rule applies universally across an application’s backend. Elias, you think this framing makes sense for understanding why this is so difficult to detect?

Elias: It does, Nadia. Because authorization isn't about how data moves in a sequence; it's about an application-specific relation—who may act on what. That means no rule written beforehand carries over to the next part of the system. Conventional taint analysis, which tracks input flow, just can’t capture that specific authorization context. This paper suggests that an LLM agent could actually infer this required relation directly from the code and data model, but without a systematic way to cover every route and prioritize what to inspect, those flaws remain hidden.

Priya: From a privacy and measurement standpoint, I’m interested in what this means for the actual data we can gather about these vulnerabilities. If the LLM agent is inferring these properties from source code, are we talking about a way to systematically audit security without needing dynamic testing or manual reviews for every single endpoint? What does this automation actually look like in terms of the information it extracts ?

Nadia: Exactly, Priya. The paper proposes ABSENTIA as a security scaffolding designed to turn general LLM agents into systematic vulnerability detectors for backend web applications. It claims this approach directs the agents to analyze source code route by route using invariant falsification. The central claim is that this methodology can recover the property each route is meant to guarantee directly from the source and code held to it, allowing a reviewer to verify findings instead of doing the heavy analysis themselves.

Elias: I see how they build that direction by first mapping out a graph of the application’s request surface, recording where each route is declared and linking it to the code and authorization checks. Then, they apply invariant falsification to audit each route one at a time, inferring what that specific route must guarantee and checking the code against that expectation. That sounds like a very targeted way to find those authorization gaps.

Paper summary: Priya: So, if we look at the data from this paper, what kind of empirical evidence do they present? Are we talking about just theoretical success or have they shown how effective this system is in practice against real vulnerabilities? What kind of results are we looking at regarding detection rates ?

Nadia: The paper validates ABSENTIA using a benchmark called BAC-BENC H, which contains thirty disclosed broken access control advisories across twenty-five repositories in Python, TypeScript, and JavaScript. The results show that this approach achieved a detection recall of nineteen out of thirty vulnerabilities. Furthermore, they noted that ABSENTIA outperformed dedicated static analysis tools like CodeQL and Semgrep, which recalled none.

Elias: That comparison against established tools is significant because it shows that systematic procedure can indeed recover application-specific authorization properties. The authors also mention that the performance gain comes from decomposition making vulnerabilities reachable, increasing recall from three instances to eighteen and then adding invariant falsification cuts reports by a third while raising verified precision.

Priya: That kind of performance metric—the increase in verified precision from forty point five percent to fifty point eight percent—tells us something about the reliability of these findings when we actually review them. What does that increased precision mean for the actual security engineering workflow? Does it reduce false alarms or increase trust in the reported issues ?

Nadia: It means that when ABSENTIA flags something, it’s much more likely to be a real issue because of that precision gain. This suggests that the systematic direction provided by ABSENTIA is highly effective at filtering out noise and focusing the review effort on actionable items. It really shows that automating the reading task of security properties can yield reliable results.

Elias: I think it points to a necessary evolution in how we approach authorization auditing in code, moving away from generalized flow analysis toward property-based verification guided by structural mapping. The system manages continuous analysis by reusing verdicts from previous runs, which helps keep the process manageable over time. This persistence is key for practical application within a development cycle.

Priya: And what about the cost of running this kind of systematic analysis? If we consider deploying this across a large enterprise codebase, how significant is the operational expenditure associated with running ABSENTIA compared to the potential reduction in costly post-deployment breaches ?

Nadia: The paper notes that the cost of running ABSENTIA is substantial, with a median cost of approximately forty-four per repository. However, they manage this by reusing verdicts from previous commits when files are byte-identical across those two versions. This continuity is what makes the analysis feasible for ongoing auditing rather than a one-off check.

Paper summary: Elias: The complexity of the scaffolding itself seems to be the main barrier, as it requires two agents for mapping and then sequential analysis by route. The authors admit that a stronger model underneath, like Claude Sonnet-five can recover more instances when reporting them, but they found that this doesn't improve the initial route extraction stage. So, the direction setting is crucial even if the inference engine is powerful.

Priya: It sounds like a trade-off between high cost and high specificity, where you invest more upfront in directing the analysis systematically to get more accurate results later on. This kind of systematic direction seems essential when dealing with authorization failures, where the context is so application-specific.

Nadia: So, to wrap up this discussion on ABSENTIA, we’ve seen how this scaffolding attempts to automate the arduous task of finding broken access control issues by systematically directing LLM agents through route-by-route analysis using invariant falsification. It shows that we can recover these application-specific authorization properties directly from the source code, which is a significant step forward in automated security auditing.

Elias: Indeed, the implications suggest a future where security engineers don't have to perform exhaustive manual checks on every single route; instead, they use this systematic approach as an audit tool. The core of ABSENTIA is turning general LLM agents into directed vulnerability detectors for backend web applications, which is a novel way to handle the nature of authorization.

Priya: When we think about the broader impact, this work suggests that we can start moving toward automated verification of intended security policies within an application’s source code rather than just searching for known patterns or flow signatures. This shifts the focus to what the route is *supposed* to guarantee, which is a much more robust way to approach authorization flaws.

Nadia: That shift in focus from flow tracing to property verification seems like the most important concept here for the future of security tooling. We've seen how ABSENTIA, despite its cost and complexity, managed to detect nineteen out of thirty disclosed vulnerabilities better than CodeQL or Semgrep.

Elias: The paper concludes that the systematic procedure itself is what makes those vulnerabilities reachable by an AI agent, demonstrating that structure and direction are as important as the underlying model's raw capability in this context. This points toward a future where security analysis relies heavily on architectural understanding rather than just pattern matching.

Paper summary: Priya: It really makes you wonder how far this direction-setting capability can be extended to other complex authorization scenarios that aren't just simple route checks, like multi-tenant access policies or fine-grained resource permissions. That seems like the next frontier for this kind of systematic property recovery.

Nadia: Exactly, Priya, because those complex relations are precisely where conventional static analysis tools struggle the most. ABSENTIA provides a framework to start mapping those application-specific relations systematically.

Elias: So, the paper on "ABSENTIA: Detecting Broken Access Control Vulnerabilities in Web Applications" presents a way to use structured scaffolding with LLMs to systematically audit web application routes using invariant falsification. It claims this method can recover the intended security properties of a route directly from the code, which has shown success in detecting nineteen out of thirty disclosed vulnerabilities compared to other static analysis tools.

Priya: The authors are suggesting that this systematic approach, building the request surface graph and then auditing route by route, offers a more reliable way for developers to verify authorization checks than just looking for conventional framework-specific guards. This methodology seems rooted in using AI as an engineer reading code to infer what the code is meant to enforce.

Nadia: The implications for the world of application security are that we might see a shift toward more automated, directed auditing processes that focus on verifying intended authorization relations rather than just chasing data flow patterns. This work highlights how essential systematic direction is when dealing with security properties that lack universal flow signatures.

Elias: That's a big idea, Nadia, because it tackles the fundamental difference between injection flaws and access control flaws—one being about data movement and the other being about a specific relation. If we can automate inference of that relation from the code, it could significantly improve our ability to secure complex web architectures.

Priya: I just think the most tangible impact is in how developers use these tools; instead of getting thousands of vague alerts, they get a targeted set of properties each route must satisfy, which helps them fix the root authorization mistake directly.

Nadia: It does sound like the title "ABSENTIA: Detecting Broken Access Control Vulnerabilities in Web Applications" is fitting because it describes a security scaffolding that systematically directs LLM agents to analyze application routes using invariant falsification. The authors are showing us how to use AI to perform this specific, directed reading task.

Conclusion: Nadia: So we've just looked at how ABSENTIA systematically scans code to find broken access control issues, and now we need to wrap up by talking about the title and authors of this paper, right?

Elias: Yeah, I think focusing on the title "ABSENTIA" is a good starting point because it hints at a scaffolding or a framework being built for this analysis.

Priya: From my side, I think we should really look at who wrote it and what kind of background those researchers have to understand their approach.

Nadia: Exactly, Priya; knowing the authors helps us gauge the credibility of this method when we're talking about real-world security fixes.

Elias: I noticed they were working with a problem where authorization isn't a simple data flow issue, which is a key distinction from what we usually see in these types of analyses.

Priya: That distinction is crucial because it means the authors are targeting a fundamentally different kind of vulnerability that standard tools often miss.

Nadia: And the authors are showing us how they use invariant falsification to check the properties each route should guarantee, which is a very specific technique.

Elias: That technique suggests they're trying to recover those application-specific rules directly from the source code, rather than relying on generic rules.

Priya: So, what does that recovery process actually look like in practice for an AI agent analyzing the codebase? What kind of data is being inferred?

Nadia: It’s about turning a general LLM into a systematic checker that builds a map of the application and then tests every path against its intended security requirements.

Elias: That mapping phase, building the request surface graph, seems like it's the most critical part for making sure the analysis is targeted and not just random.

Priya: I’m interested in how they handle continuous analysis; if this system can reuse verdicts from previous runs, does that make it practical for real development cycles?

Nadia: It definitely does; having that continuity means developers can see how their fixes affect the overall security posture over time, which is very useful.

Elias: The authors also pointed out the cost involved in running this kind of deep analysis, so we have to keep an eye on whether that cost scales effectively for large projects.

Priya: I think the real implication here is that we might start shifting our focus from finding simple injection patterns to verifying the intended security policy at a much deeper level.

Nadia: That sounds like a big shift; instead of just tracing data movement, we’re looking at what the application was supposed to enforce all along.

Elias: It moves us closer to understanding how complex authorization relations are actually implemented in code, which is where things get tricky for traditional methods.

Priya: And that's where the next big question is whether this systematic approach can handle those more complicated, multi-tenant access scenarios we see today.

André Vicente Duarte, Aditya Oke, Rui Melo, Shubham Gandhi, Nachiket Kotalwar, Charmi Khandor, Danqing Wang, Arlindo L. Oliveira

Carnegie Mellon University

cs.CR, cs.SE

Submitted: 2026-10-01

Updated: 2026-10-01

Code: https://github.com/indico/indico

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 84/100

The gist: Broken access control, which involves authorization failures where a principal acts on an unauthorized resource, remains difficult to detect in source code because its defect is defined by an

Key concepts

Broken Access Control
This defect occurs when an application fails to properly enforce 'who may act on what.' Unlike input flaws, it is a relationship between a user (principal) and a resource. Conventional tools struggle because the required security rule is semantic—it depends on the application's specific intended policy, not just syntactic code structure.
Invariant Falsification
This technique involves systematically testing every route to determine what properties it must guarantee. The agent constructs a request specifically designed to violate each property. This method moves beyond simply checking for existing guards; it actively tries to break the intended security logic of each application route.
Graph of the Application's Request Surface
This is a map built by ABSENTIA that records every route in an application. It links these routes to the specific code handling them, any authorization checks present, and sensitive operations. This mapping allows the system to understand how different parts of the application interact and where security boundaries lie.
Composition Across Routes
This stage addresses how actions on one route can be dangerous when combined with another. ABSENTIA tracks 'capabilities'—behaviors where an action on one route becomes critical due to a subsequent action on another route. This ensures that the security analysis considers the cumulative risk across multiple application paths.

Terminology

Summary

Broken access control, which involves authorization failures where a principal acts on an unauthorized resource, remains difficult to detect in source code because its defect is defined by an application-specific relation rather than a universal dataflow property. This paper introduces ABSENTIA, a security scaffolding that systematically directs general LLM agents to analyze web application source code route by route using invariant falsification. The core finding is that this approach detects 19 out of 30 disclosed broken access control vulnerabilities in a benchmark, outperforming dedicated static analysis tools like CodeQL and Semgrep, demonstrating that systematic procedure can recover application-specific authorization properties.

The Gist

ABSENTIA is a security scaffolding that turns general LLM agents into systematic vulnerability detectors for the backend of web applications by building a graph of routes and then working route by route, applying invariant falsification to infer and check the properties each route is meant to satisfy against the source code.

Motivation: The Challenge of Authorization Detection

Broken access control is fundamentally different from injection flaws because authorization is a relation—who may act on what—which lacks a universal flow signature that conventional taint analysis can trace. Unlike SQL injection, where an untrusted input reaches a dangerous operation in an alterable form, the defect in access control is a relation between a principal and a resource. This means no rule written in advance carries to the next application. Static Application Security Testing (SAST) tools struggle because they test for a conventional framework-specific guard, which is only syntactic, not semantic regarding the intended policy of the application. Detecting missing or wrong checks requires knowledge that comes from two sources: dynamic testing or manual review. ABSENTIA aims to automate this reading task by recovering the property each route is meant to guarantee directly from the source and code held to it.

How it Works: The ABSENTIA Scaffolding

ABSENTIA operates through a multi-stage process designed for systematic, directed analysis, addressing three core obstacles:

  1. Mapping the Application: The first stage builds a graph of the application’s request surface, recording where each route is declared and linking it to the code that serves it, authorization checks, and sensitive operations. This involves two agents: Profiling (to establish framework/language/data stores) and Extraction (to produce expressions that locate specific elements like routes or query values).

  2. Per-Route Invariant Falsification: ABSENTIA then analyzes each route independently by enumerating the properties the route must guarantee. The agent enumerates the properties the route is meant to guarantee and constructs a request that would violate each property, reporting a finding as the request that would violate that property rather than as a judgment.

  3. Composition Across Routes: Findings are linked across routes by recording capabilities, which are behaviors where one route's action matters if another route does something in particular. This stage ensures that what is minor on one route can be serious once a second route is in play.

  4. Re-Analyzing After a Commit: The system manages continuous analysis by reusing verdicts from previous runs, deciding between two commits which reuses the verdicts that a later commit cannot affect if the files read are byte-identical across the two commits.

Validation and Benchmarking

The effectiveness of ABSENTIA is validated using BAC-BENC H, a benchmark containing 30 disclosed broken access control advisories across 25 repositories in Python, TypeScript, and JavaScript. The evaluation measures recall on the vulnerable commit and a paired recall that credits an instance only when it is detected on the vulnerable commit and cleared on the patched commit. ABSENTIA achieved a detection recall of 19/30 at 51% verifier-adjudicated precision, significantly outperforming CodeQL and Semgrep, which recalled none. Furthermore, ABSENTIA leads dedicated analyzers in Python and trails only CodeQL and IRIS in Java on the OWASP Benchmark injection categories.

Component Contributions to Performance

The performance gain is attributed to the sequential nature of the components: Decomposition is what makes the vulnerabilities reachable, increasing recall from 3 instances to 18. Adding invariant falsification then cuts reports by a third while raising verified precision from 40.5% to 50.8%. The model underneath, MiniMax-2.5, is tested for its capacity, and experiments show that a stronger model (Claude Sonnet-5) can recover more instances when reporting the bug but does not improve the initial route extraction stage.

Cost and Continuous Analysis

The cost of running ABSENTIA is substantial, with a median cost of approximately "44" per repository. However, continuous analysis is managed by reusing verdicts from previous commits.

Improvements for AI systems

Here are specific improvements that can be made to existing AI systems, based on the methodology and findings presented in ABS ENTI A: Detecting Broken Access Control Vulnerabilities in Web Applications:


  1. The core improvement is transforming general-purpose LLM agents into a systematic, goal-oriented security auditing tool by implementing the ABS ENTI A scaffolding.

  2. The improved system can perform automated, per-route invariant falsification on web application source code to detect Broken Access Control (BAC) vulnerabilities (e.g., IDOR/CWE-639, missing authorization/CWE-862).

  3. The system will first build a comprehensive Security-Aware Source Code Graph by profiling the application's framework, language, data stores, and authentication mechanisms before analyzing specific routes. This ensures analysis is contextually relevant to the application's structure.

  4. The system will execute a two-stage auditing process:

Ease (Extraction) and Reasoning (Invariant Falsification). The Extraction stage identifies all relevant routes; the Reasoning stage infers the specific authorization properties each route is meant to guarantee, and then constructs a concrete request that violates that property to generate a falsifiable finding.

  1. The system can systematically compose findings across different routes, identifying chained vulnerabilities where no single route reveals the full flaw (e.g., combining an insecure file upload capability with an insecure retrieval endpoint).

  2. The system will implement a mechanism for Run-to-Run Stability and incremental analysis. It can determine which previous run's verdicts remain valid after a code commit, significantly reducing re-analysis costs by only re-evaluating routes affected by the changes in the diff.

  3. By using a model (like Claude Sonnet) as an adjudicator, the system can produce high-precision findings. The improved system can be tuned to maximize verified precision against human expert verification, ensuring that reported findings are substantiated and not just superficial patterns matched by a language model.

  4. The improved AI system will demonstrate superior generalization capabilities compared to traditional static analyzers (CodeQL/Semgrep) when applied to authorization flaws, as it derives the required security properties directly from the application's source code rather than relying on pre-written, framework-specific rules.

  5. The system can be deployed across multiple languages (Python, Java, TypeScript) and frameworks by leveraging its ability to dynamically profile and extract route definitions based on the specific codebase it encounters.

  6. The improved AI system will be optimized for cost efficiency in continuous auditing by reusing verdicts from previous analysis runs, providing a mechanism to scale security checks across a large codebase without incurring linear computational costs for every commit.

Sources

Related papers