GROB: A Multi-Agent Architecture for Public-Trace Investigation of Candidate Agentic Activity

arXiv:2610.11467 · cs.CR, cs.AI · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "GROB: A Multi-Agent Architecture for Public-Trace Investigation of Candidate Agentic Activity".

Elias: The gist: GROB presents a multi-agent architecture designed to investigate candidate autonomous-agent activity by performing controlled, read-only collection of public Internet traces when privileged telemetry is unavailable.

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: Looking at the GROB architecture and what the authors are proposing in this paper, it seems they’ve built a system focused on retrospective identification rather than making final claims about specific actors.

Elias: The title itself, "GROB: A Multi-Agent Architecture for Public-Trace Investigation of Candidate Agentic Activity," really captures the essence of what they're doing—using a multi-agent setup to look at public traces for signs of agent activity.

Priya: What this means simply is that even if you don't have private data, you can use these controlled methods to find evidence that could later be compared with other information you find.

Nadia: The authors focus on preserving the evidence in a provenance-preserving workspace, which is distinct from making an attribution decision about who did what.

Elias: And they emphasize that candidate selection prioritizes records based on the evidence available at the time of collection, and they use deterministic rules to control what actually gets admitted to persistent evidence.

Priya: The paper points out a limitation, which is that public traces alone do not establish the identity of the underlying actor or a shared execution; it’s just retrospective identification.

Nadia: So if you're listening and you think about this, the main point is that these tools help investigators find potential activity in public data, setting up a basis for later verification with other sources.

Conclusion: Nadia: So to wrap up, GROB is this multi-agent setup designed to look at public internet traces to find signs of AI activity when you don't have private telemetry or specific targets in mind.

Elias: Yeah, it’s about using controlled, read-only collection of public data instead of having access to privileged information.

Priya: What this means practically is that defenders can use what's out there online to find stuff they can later check against other evidence.

Nadia: Right, but the core thing here is how they handle the data—they create a workspace where everything stays untrusted until it gets validated by deterministic code.

Elias: That separation between the Sentinel controlling detection and Scout testing hypotheses sounds like a way to keep the interpretations from getting baked into permanent memory without review.

Priya: And they’re focused on provenance, meaning they aren't trying to make an attribution decision about who did what; they’re just preserving the record of what happened in public.

Nadia: So even if you can't prove *who* it was, you can create a verifiable workspace that shows *what* activity might have occurred based on public traces.

Elias: It’s an exploratory research prototype, which is important because it signals this is more about building a framework for investigation than deploying something live for monitoring.

Priya: And the limitation they highlight is that without first-party execution evidence, these traces alone can't actually establish the identity of the actor or a shared execution.

Nadia: So it’s a tool for finding candidates, not definitive proof of action from public data alone.

Elias: It sets up a foundation where later investigation can compare these candidate records with independent evidence you might find elsewhere.

Priya: This kind of architecture could become useful for building better methods to sift through massive amounts of public web material for subtle signs of agentic behavior.

Chiara Bonfanti, Prof. Cataldo Basile

Politecnico di Torino

cs.CR, cs.AI

Submitted: 2026-10-08

Updated: 2026-10-08

Project page: https://www.nextgov.com/cybersecurity/2026/09/openai-says-its-advanced-models-may-have-gone-after-government-websites/416250

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: The gist: GROB presents a multi-agent architecture designed to investigate candidate autonomous-agent activity by performing controlled, read-only collection of public Internet traces when privileged

Key concepts

Multi-Agent Architecture
GROB uses several specialized agents with distinct jobs: Sentinel monitors and gates activity, Scout tests potential hypotheses about agent behavior, and Librarian keeps validated memory. This structure allows for a controlled workflow where interpretations are reviewed before being permanently stored.
Sentinel
The Sentinel agent is responsible for detection and controlling the process. It acts as a gatekeeper, determining which public traces warrant further investigation by Scout, ensuring that only relevant material moves forward in the system's workflow.
Librarian (CoALA-inspired Memory)
The Librarian maintains validated semantic memory, treating retrieved web material as untrusted initially. Only evidence admitted by deterministic code can be stored permanently. This design ensures that interpretations from large language models are reviewed before they become persistent facts in the system's knowledge base.
Provenance-Preserving Workspace
Instead of making final attribution decisions, GROB creates a workspace that preserves the history and source of every collected public record. This allows investigators to see exactly what evidence was found, when it was captured, and how it was linked together for later comparison.

Terminology

Summary

The gist: GROB presents a multi-agent architecture designed to investigate candidate autonomous-agent activity by performing controlled, read-only collection of public Internet traces when privileged telemetry is unavailable.

Research Question

The research question posed is: Can defenders use public Internet traces, without privileged telemetry or known targets, to identify and preserve signs of candidate agentic activity that can later be investigated and compared with independent evidence?

Operational Workflow and Architecture

GROB is a multi-agent architecture where individual agents have distinct roles and responsibilities The system was developed under the earlier name CyberXAI, which remains visible in earlier project material The prototype was designed to investigate several kinds of public Web evidence

The architecture involves a separation of concerns:

** Sentinel controls detection and gating while Scout performs hypothesis testing, and Librarian maintains validated semantic memory. This separation prevents model-generated interpretations from entering persistent memory without review. Large language models assist with interpretation, but only deterministic code can admit evidence to persistent memory. Candidate records are prioritized using the evidence available at collection time. Followup tests are defined before additional retrieval so that the initial observation remains distinguishable from its later investigation. Public-web collection is restricted to read-only retrieval of publicly accessible material. The system neither crosses access controls nor modifies remote systems. The system produces a provenance-preserving workspace rather than an attribution decision. The CoALA-inspired memory design treats retrieved Web material as untrusted until it is validated. Large language models assist with interpretation, but only deterministic code can admit evidence to persistent memory. Candidate selection prioritizes records using the evidence available at collection time. The system performs controlled link traversal and retains retrieved records together with their source and capture time. Candidate observations can open a scoped investigation, while deterministic rules control what is admitted to persistent evidence. The system operates without access to the underlying agent systems or to private service logs. Each collected public record is therefore stored together with its source and capture time. A trace is considered found when an exact identifier is retained through this process. It is resolved when the collected trace can be linked to a specific public revision, request, or resource. It is externally corroborated when independent evidence later reports compatible activity. None of these states, without first-party execution evidence, establishes the identity of the underlying actor or a shared execution. The system is an exploratory research prototype rather than a production monitoring system. Its CoALA-inspired memory design treats retrieved Web material as untrusted until it is validated. The system neither crosses access controls nor modifies remote systems. The system produces a provenance-preserving workspace rather than an attribution decision. Candidate records are prioritized using the evidence available at collection time. The system performs controlled link traversal and retains retrieved records together with their source and capture time. Candidate observations can open a scoped investigation, while deterministic rules control what is admitted to persistent evidence. The system operates without access to the underlying agent systems or to private service logs. The system produces a provenance-preserving workspace rather than an attribution decision. The CoALA-inspired memory design treats retrieved Web material as untrusted until it is validated. Large language models assist with interpretation, but only deterministic code can admit evidence to persistent memory. Candidate selection prioritizes records using the evidence available at collection time. The system performs controlled link traversal and retains retrieved records together with their source and capture time. Candidate observations can open a scoped investigation, while deterministic rules control what is admitted to persistent evidence. The system operates without access to the underlying agent systems or to private service logs. The system produces a provenance-preserving workspace rather than an attribution decision <ref:2610.

Improvements for AI systems

  1. Bold header: Retrospective Trace Resolution

This system can preserve public traces before later public reporting established their significance, allowing for retrospective identification and reconstruction of relevant public traces by linking early sparse identifiers to specific requests, as demonstrated by the Census example resolving to specific Census requests from 16-17 June.

  1. Bold header: Evidence Prioritization Architecture

The system can utilize a structure where Candidate records are prioritized using the evidence available at collection time, ensuring that only traces deemed relevant for later investigation are retained, preventing model-generated interpretations from entering persistent memory without review.

  1. Bold header: Artifact-Level Cross-Surface Joins

The improved AI system can perform stronger correlation than naming similarity alone by focusing on artifact-level DataUSA-to-Census chain in the public revision history, which links requests across different public labels and resources, establishing artifact-level continuity, not execution identity.

  1. Bold header: Layered Evidentiary Strength Assessment

The system can employ a mechanism that explicitly categorizes evidence strength according to the paper's layers: Lexical Naming and token similarity versus Artifact Exact requests, URLs, or other shared traces, ensuring that only stronger connections are used for high-confidence reconstruction.

  1. Bold header: Contextual Lead Generation

For ambiguous cases where direct joins are absent, the system can function as a contextual lead, preserving observations like AgentNovOneTransportHelper while noting that the correspondence remains a lead rather than an identity claim when no exact artifact or provenance record exists.

Abstract

We present GROB, a multi-agent architecture for investigating candidate autonomous-agent activity through public Internet traces when privileged telemetry is unavailable. The system performs controlled, read-only collection of public traces and preserves selected observations for later resolution. In a frozen September 2026 corpus, several collected traces became more informative as additional public evidence emerged. The strongest result concerns Census-labelled identifiers captured on 9 September. Public revision records later resolved these identifiers to specific Census requests from 16 - 17 June. Other results show weaker links between traces collected by GROB and evidence reconstructed or reported later. These links vary in strength, and only some can be tied to specific public records. The results show that sparse public traces can remain useful even before their significance is fully understood. Such evidence can support later reconstruction, but public traces alone do not establish organizational attribution. Execution identity presents a separate problem, as continuity of agent identity remains an active research question for autonomous language-model agents.

Related papers