From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage

arXiv:2610.10608 · cs.CR, cs.AI, cs.MA · Submitted 2026-10-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "From Investigation Failures to Reliable SOC Agents".

Nadia: The gist The study investigates five representative approaches to LLM-based alert triage using an interactive benchmark to determine how reasoning strategies affect reliability and performance in security operations centers;

Elias: First, who's behind it and why it matters.

Title and authors: Nadia: So we’re looking at this paper now, "From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage." It's about how these large language model agents actually handle the messy job of sorting through tons of security alerts in a security operations center.

Elias: Yeah, it tackles the core problem that most alerts are benign, but missing an actual attack is really bad for security teams. This paper sets up a way to test different ways these agents reason when they’re supposed to be gathering evidence.

Priya: It sounds like they’re trying to figure out if just having more reasoning power is enough, or if the structure of how the AI looks for clues actually matters more for making a reliable decision.

Nadia: Exactly. The researchers set up this interactive benchmark called ALERT-BENCH, which simulates a real security environment where the agent has to pull evidence from a live system before it can close an alert. They used one thousand two hundred forty-seven alerts from a four-day attack scenario to test this out <ref:2610.10608#pg1,1,247 alerts from a four-day>.

Elias: And they systematically compared five different reasoning strategies: single-pass tool use, iterative retrieval, sampled investigations, self-review, and explicit verification. It's like they put the AI through different training diets to see which one performs best at gathering information for triage.

Priya: When you look at the results from that study, it’s interesting because they found that when an agent’s search for surrounding activity comes up empty, attack alerts are much more likely to get dismissed, and same-context self-review actually makes things worse.

Nadia: That’s a big finding because it suggests that just letting the AI think through everything on its own isn't the answer; how they structure that evidence gathering is what really determines reliability in this kind of work.

Elias: And they built their solution, which they call AIDA, to address these failures by using a multi-agent architecture with a Dialectical Commitment Protocol. This means there are four agents—Strategist, Executor, Adversary, and Judge—that have to follow a specific process before any alert gets officially triaged.

Priya: I’m curious about what AIDA actually does in practice. It seems like it’s trying to enforce a more disciplined way of investigation instead of letting the AI just wander around looking for things.

Title and authors: Nadia: Well, AIDA introduces a lot of structure, like requiring an explicit proposed decision before challenge and assigning final authority to a separate Judge. They also added an append-only Investigation Ledger so they can keep track of every single step taken during the investigation process.

Elias: That ledger is key because it means if the final decision changes, the system can look back and see exactly what evidence was gathered along the way, which is important for debugging why a bad decision happened.

Priya: The performance numbers they report are quite strong; AIDA achieved an F1 score of zero point nine five eight, which they say is statistically significant compared to all the other approaches they tested <ref:2610.10608#pg2>. That’s a solid number for this kind of task.

Nadia: And it also showed a big jump in reducing false negatives, cutting that rate from forty point four percent down to just three point one percent. That’s a massive difference in catching actual threats that might have been missed before, which is what matters most for security.

Elias: It’s worth mentioning that the performance of this AIDA architecture isn't perfect across all different underlying models; for instance, one model reached an F1 score of zero point nine five two, while others were much lower <ref:2610.10608#pg2>. This suggests the structure itself is quite robust, even when you swap out the brain powering it.

Priya: It also sounds like they’re pointing toward a configuration that hits a certain level of capability—they mentioned reaching the operating point of a GPT-four point one reference configuration with an open-weight model in at least one case. That gives us some context on how much lift this architecture provides across different AI sizes.

Nadia: So, what does this mean for security teams who are actually using these tools day-to-day? It suggests that simply layering more complex reasoning isn't the solution; it’s about making the evidence gathering and the review process deliberate and structured.

Elias: The paper argues that reliability comes from how evidence is gathered and how conclusions are evaluated before an alert gets closed, not just having a smarter agent internally.

Title and authors: Priya: For someone who only listens to this show, what does this paper mean for their daily work? It means that if you’re using an AI to triage alerts, you need to build in these checkpoints where the system has to prove its findings before it makes a final call.

Nadia: That’s the point. If you want a reliable agent, you have to design the investigation process itself so it forces that structure, like AIDA does with its separate challenge and adjudication agents.

Elias: And they did push that further by proposing improvements like making the Strategist commit to a thesis before any challenge even starts, which directly counters the problem where self-review could make things worse.

Priya: I think that explicit commitment requirement is really smart because it stops the AI from just second-guessing itself in a way that introduces errors, which was a major issue in earlier methods.

Nadia: It’s about moving from an open-ended investigation to a very guided one where the steps and evidence collection are strictly defined by the protocol.

Elias: And they also suggested asymmetric evidentiary thresholds, meaning they require different levels of proof for closing an alert versus escalating it, which makes sense because you want to be careful about when you move from low-level dismissal to high-level escalation.

Priya: The idea that dismissal doesn't get a stronger investigation than escalation is a concrete insight into how the system prioritizes risk during triage.

Nadia: It shows that the architecture needs to handle those different levels of certainty very carefully, ensuring you don't accidentally close something serious because it didn't find one tiny piece of supporting evidence.

Elias: So, we’ve seen how they move from observing failures in various reasoning approaches to building a system where structure and separation—Strategist, Adversary, Judge—are the core ingredients for making triage reliable.

Priya: It sounds like the main implication is that we need to stop focusing only on the model's intelligence and start focusing on designing the investigation framework itself to enforce discipline.

Nadia: That seems to be exactly what this paper is proposing with its work on "From Investigation Failures to Reliable SOC Agents." It’s about fixing the process, not just upgrading the reasoning engine.

Elias: We’re going to wrap up this look at how AIDA structures that evidence gathering and adjudication process in our next segment.

The paper's summary: Nadia: So we’re looking at this paper now, "From Investigation Failures to Reliable SOC Agents." It’s about how we can stop these alert triaging AIs from just guessing and start making them actually reliable in a security center environment.

Elias: Yeah, the researchers set up an interactive test where they replay real system telemetry through a live SIEM, forcing the AI to pull its own evidence before it can make a final call on an alert.

Priya: It sounds like they’re trying to find out what makes an AI triage decision solid versus just being lucky or making a quick guess based on limited information.

Nadia: Exactly. They tested five different ways the AI tries to reason—things like single-pass tool use or even letting the AI self-review its own work—and they found that those methods actually hurt performance when things go wrong.

Elias: Specifically, when an agent’s search for surrounding activity comes up empty, attack alerts are way more likely to get dismissed, and sometimes the AI just making a decision based on what it already knows makes the error even bigger.

Priya: That tells us that having more complex reasoning power isn't enough; you have to fix the structure of how the AI gathers its clues.

Nadia: Right. So they built this new system called AIDA, which uses four distinct agents—Strategist, Executor, Adversary, and Judge—to force a much more disciplined way of doing things.

Elias: The core idea is to separate the planning from the execution and to have someone completely independent challenge the initial plan before a final decision is made.

Priya: And they added this append-only ledger that tracks every single step, which means if the final answer turns out to be wrong, you can look back and see exactly what evidence was collected at each stage.

Nadia: It sounds like AIDA’s big win here is that it hit an F1 score of zero point nine five eight, which they say is significantly better than all the other methods they tested in this benchmark.

Elias: And that improvement isn't just about the final score; they also saw a massive drop in false negatives, from over forty percent down to just about three percent.

Priya: That’s what really matters for security operations, because it means the AI is much better at finding actual threats that might have been completely ignored before.

Nadia: It shows that reliability doesn't come from making the internal reasoning engine smarter; it comes from designing a process where evidence gathering and conclusion evaluation are strictly structured.

Elias: So they’re pushing for this separation of concerns, with clear roles like Strategist for planning, Adversary for checking quality, and Judge for the final call.

Priya: And they even suggested setting different evidentiary thresholds; meaning you need a different amount of proof to dismiss an alert compared to when you decide it's serious enough to escalate.

Nadia: That makes sense because you don't want a system closing something important just because it couldn't find one tiny piece of supporting data.

Elias: This paper points toward fixing the process itself, not just upgrading the underlying model’s intelligence. It’s about building in checkpoints that force discipline during the investigation.

Priya: For someone listening who only cares about what this means for their job, it suggests that if you use an AI to triage alerts, you need to design a framework where the system has to prove its findings before it closes anything out.

Nadia: That’s the main point: stop focusing only on how smart the AI is and start designing the investigation framework so it forces that structure.

Elias: And they're already tweaking that structure by requiring the Strategist to commit to a plan before anyone challenges it, which directly addresses those self-review errors we saw earlier.

Priya: So we’ve seen how they move from observing failures in various reasoning approaches to building a system where defined roles and explicit decision points are the core ingredients for making triage reliable.

The paper's improvements: Nadia: So we’re looking at what they’re suggesting next for making these triage agents actually reliable. It sounds like they aren't just happy with their AIDA setup, and they want to push it further by fixing some of those initial structural problems.

Elias: Yeah, the paper points out that even with a good architecture, you still have issues if you don't define the decision-making process really clearly at the start.

Priya: It sounds like they are focusing on making sure there is a very firm commitment to an idea before any real investigation begins, which should stop agents from just wandering around and getting confused.

Nadia: Right. They’re proposing making the Strategist commit to a specific classification, supported by evidence, before the Adversary even gets involved in challenging it. That stops the same-context self-review problem they saw earlier when things got worse.

Elias: It’s about forcing that initial hypothesis to be written down and backed up immediately, instead of letting the AI build a theory and then trying to disprove itself later.

Priya: And I like that because it ties directly into the idea of asymmetric evidence thresholds; they suggest different levels of proof are needed for closing an alert versus escalating it.

Nadia: Exactly. So, if you want to dismiss an alert, you need a certain amount of benign evidence or a very complete investigation, but if you want to escalate, you need concrete proof that something malicious is happening.

Elias: That’s practical because it makes the system much more cautious about when it decides to move from low-level dismissal to high-level action.

Priya: And they’re also introducing this idea of an Investigation Ledger that keeps a record of every single step taken, which means the Judge can actually go back and ask for more evidence if they feel something is still missing.

Nadia: That ledger is crucial because it preserves the history, so if the final decision changes later on, you don't lose track of what was found along the way.

Elias: It sounds like they’re moving toward a system where structure and separation—Strategist, Adversary, Judge—are not just helpful suggestions but mandatory parts of how an alert gets handled.

Priya: So for someone listening who wants to know what this means for their daily work, it suggests that the next step isn't just adding more reasoning steps; it’s about rigorously defining the evidence gathering and review process itself.

Nadia: That’s right. It means building in these checkpoints where the system has to follow a very disciplined path, rather than letting it just run freely with its tools.

Elias: They are pushing for that disciplined flow by making sure every agent has a strictly defined role in the debate, which they found helps reduce errors in evidence-based discussions.

Priya: So the implication is that reliability comes from enforcing that structure—the separation of roles and the commitment to a thesis—rather than just relying on an AI being inherently smarter at reasoning.

Conclusion: Nadia: So we’re wrapping up this look at "From Investigation Failures to Reliable SOC Agents." Basically, this paper shows that making AI triage reliable isn't about just adding more brainpower; it's about building a disciplined process around how evidence is gathered and how decisions are checked.

Elias: Exactly. They took five different reasoning methods and found that the structure—the way the agents interact—is what separates a decent result from one that actually works in a real security operation center.

Priya: It really changes how we think about these systems, because it’s not enough to have an AI that can talk; you need to design the investigation framework so it forces that discipline.

Nadia: Right. They introduced AIDA, which uses those four distinct roles—Strategist, Executor, Adversary, and Judge—to ensure there’s always a check and balance on the decision-making process.

Elias: And the numbers back up that structure; AIDA got an F1 score of zero point nine five eight with a false negative rate slashed down to just three percent compared to forty percent in some of the older methods.

Priya: That drop in false negatives is huge for security because it means we're actually catching more threats that might otherwise get buried by a tired or lazy triage system.

Nadia: It shows that when you structure evidence retrieval and decision review this way, you substantially improve agentic triage performance.

Elias: The authors are pushing toward making these roles even tighter, suggesting things like requiring an explicit thesis commitment before the challenge phase starts to prevent those self-review errors they found earlier.

Priya: And that idea about different thresholds for dismissal versus escalation is really important because it shows the system needs to know when it’s allowed to close something based on less evidence.

Nadia: So, in short, this paper is a strong argument that reliability comes from process design rather than just model size or raw reasoning capability.

Elias: It sets a new benchmark for how we should architect these systems—focusing on separation of concerns and explicit decision points.

Priya: For someone listening who only cares about what this means for their work, it suggests that the next step isn't just upgrading the model; it’s about designing a framework where the investigation itself has built-in checkpoints.

Saimon Amanuel Tsegai, Alex Kantchelian, Danfeng (Daphne) Yao, Peng Gao

Virginia Tech · Google

cs.CR, cs.AI, cs.MA

Submitted: 2026-10-07

Updated: 2026-10-07

Comments: Preprint

Code: https://github.com/SigmaHQ/sigma

License: http://creativecommons.org/licenses/by/4.0/

The gist: The gist The study investigates five representative approaches to LLM-based alert triage using an interactive benchmark to determine how reasoning strategies affect reliability and performance in

Key concepts

ALERT-BENCH
This is an interactive benchmark that simulates real security operations by replaying enterprise telemetry through a live SIEM. It forces the AI system to actively retrieve evidence from the data before making a final triage decision, testing its investigative capabilities in a realistic setting.
Reasoning Approaches
The paper systematically compared five distinct methods LLMs use to investigate alerts, such as single-pass tool use or self-review. It showed that simple approaches often fail when evidence is missing, and some self-review techniques actually worsen the final decision quality.
AIDA (Adversarial Investigation and Dialectical Analysis)
A multi-agent architecture designed to improve triage reliability. AIDA uses four agents—Strategist, Executor, Adversary, and Judge—to structure evidence gathering, challenge assumptions independently, and ensure a final decision is made by a dedicated Judge after explicit proposal.
Dialectical Commitment Protocol (DCP)
A protocol within AIDA that dictates the workflow for the four agents. It mandates that an explicit decision must be proposed before the Adversary agent challenges it. This structured, adversarial process ensures thorough investigation and robust adjudication of security alerts.

Terminology

Summary

The gist The study investigates five representative approaches to LLM-based alert triage using an interactive benchmark to determine how reasoning strategies affect reliability and performance in security operations centers; this research is important because it shows that structuring evidence retrieval and decision review can substantially improve agentic SOC triage, achieving an F1 score of 0.958 compared to lower scores for the studied approaches.

How it works

The paper introduces ALERT-BENCH, an interactive alert-triage benchmark that replays enterprise telemetry through a live SIEM and requires each system to retrieve evidence. This environment turns recorded telemetry into an evaluation setting where the system investigates each alert through the SIEM before reaching a triage decision. The corpus consists of 1,247 alerts from a four-day multi-stage attack scenario, containing 225 attack-related and 1,022 benign alerts.

Systematic Study of Reasoning Approaches

The study systematically compares five representative reasoning approaches spanning single-pass tool use, iterative retrieval, sampled investigations, self-review, and explicit verification. These approaches include Direct-Prompt which uses single-pass tool use. Specifically, attack-related alerts are much more likely to be dismissed when the agents’ chosen searches for surrounding activity return no records. Furthermore, same-context self-review often makes the decision worse; Self-Refine changes 79 correct ones into errors.

Design of AIDA

Based on these findings, the paper designs AIDA (Adversarial Investigation and Dialectical Analysis), a multi-agent architecture that structures evidence gathering, independent challenge, and separate adjudication. AIDA introduces the Dialectical Commitment Protocol (DCP), which coordinates four agents: Strategist, Executor, Adversary, and Judge. The protocol requires an explicit proposed decision before challenge and assigns final decision authority to a separate Judge.

Performance of AIDA

When evaluated under the same conditions as the systematic study, AIDA achieves an F1 score of 0.958 with 0.948 precision and reduces the false-negative rate from 40.4% for the highest-F1 studied approach to 3.1%. This improvement is statistically significant against every studied approach. AIDA issues 4.79 tool calls before dismissal and 2.27 before escalation, a ratio of 2.11 with a p-value of 1.5×10−66.

Model Dependence

The performance of AIDA varies across underlying models; Qwen3.5 9B reaches an F1 score of 0.952, while the other open-weight models range from 0.098 to 0.777. AIDA therefore reaches the operating point of the GPT-4.1 reference configuration with an open-weight model in at least one case. The three-round limit bounds repeated continuation in these cases.

Conclusion

Overall, the results show that additional reasoning alone is insufficient for reliable triage. Reliability depends on how evidence is gathered and how conclusions are evaluated before an alert is closed.

ETHICS CONSIDERATIONS

The study does not involve human subjects or private user data, and the evaluation does not interact with external or operational systems. AIDA may process sensitive security telemetry and may make incorrect decisions when evidence is incomplete or adversarially influenced. Automated investigation techniques could also be misused. Accordingly, we position AIDA as a tool to support human analysts rather than enable autonomous decision-making, and we recommend access controls and human review for high-impact decisions.

REFERENCES

[1] B. A. Alahmadi, L. Axon, and I. Martinovic, “99% false positives: A qualitative study of soc analysts’ perspectives on security alarms,” in 31st USENIX Security Symposium (USENIX Security ’22), 2022, pp. 2783–2800 <ref:61B. A. Alahmadi, L. Axon, and I.

Improvements for AI systems

  1. Bold header: Explicit Thesis Commitment

This improvement requires the Strategist to commit a thesis, a proposed classification supported by cited evidence, before challenge begins and mandates that The Strategist may commit once it considers the collected evidence sufficient. This prevents the negative effect of same-context review observed in prior work by requiring an explicit decision before independent challenge.

  1. Bold header: Independent Challenge and Adjudication

The system should separate production and evaluation by having the Adversary receives the committed thesis and shared evidence without the Strategist’s private reasoning, and records a challenge to the thesis. This addresses findings that same-context self-review often makes the decision worse by providing an independent review context.

  1. Bold header: Asymmetric Evidentiary Thresholds

The AI should implement different evidentiary requirements to two outcomes, specifically requiring escalation requires concrete evidence supporting malicious activity while dismissal requires affirmative benign evidence or an investigation complete enough to justify closure. This directly addresses the finding that dismissal does not receive a stronger investigation than escalation.

  1. Bold header: Investigation Ledger for History Preservation

The system must use an append-only Investigation Ledger to record every step, preserving the history so that the Judge can request additional evidence when important information is still missing. This structure ensures that investigation history remains preserved even when the final decision reverses a previous one.

  1. Bold header: Structured Agent Roles

The multi-agent architecture should strictly define roles such as the Strategist (planning/thesis formation), Adversary (quality assurance/challenge), and Judge (final adjudication). This separation of concerns is designed to mitigate errors, as the separation also evaluated in evidence-based debate [60] was found to be beneficial.

Abstract

Security operations centers (SOCs) must triage large volumes of alerts, most of which are benign, while missed attacks can remain uninvestigated. Tool-using large language model (LLM) agents can retrieve evidence during triage, but it remains unclear how reasoning strategies determine what to gather and when an investigation is sufficient to close an alert. We study five representative approaches spanning single-pass tool use, iterative retrieval, sampled investigations, self-review, and explicit verification. To support this study, we build ALERT-BENCH, an interactive benchmark that replays enterprise telemetry through a live SIEM and requires each system to retrieve evidence. Across 1,247 alerts from a multi-stage attack scenario, every approach missed at least 40.4% of attack-related alerts. Trace analysis shows that attack alerts are more likely to be dismissed when searches return no records, same-context review has negative net correction, and dismissal receives no consistently stronger investigation than escalation. Based on these findings, we further design AIDA (Adversarial Investigation and Dialectical Analysis), a multi-agent framework that requires an explicit proposed decision before independent challenge and stronger evidentiary requirements before dismissal. AIDA preserves investigation history in an append-only Investigation Ledger and keeps the challenge in a separate reasoning context. A separate Judge adjudicates the proposed decision and challenge against evidence, resolving the alert or requesting another round when evidence is missing. On the same alerts, AIDA achieves an F1 score of 0.958, compared with 0.371-0.744 for the studied approaches, and reduces the false-negative rate from 40.4% to 3.1% while escalating 18.4% of alerts to analysts. These results show that structuring evidence retrieval and decision review can substantially improve agentic SOC triage.

Sources

Related papers