The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search

arXiv:2512.01353 · cs.CR · Submitted 2025-12-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "The Trojan Knowledge".

Elias: Large language models remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs.

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So, we've discussed how "The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search" suggests that harmful outputs are achievable by weaving together harmless queries to exploit the interconnected nature of an AI's knowledge.

Elias: That brings us back to the idea that we can't just look for bad keywords; instead, we have to consider how information is structured internally and how agents can systematically explore that structure through adaptive tree searches.

Priya: From a measurement perspective, the real significance lies in realizing that current safety mechanisms aren't looking at the right scope; they are missing the distribution of intent across a sequence of benign interactions.

Nadia: The authors demonstrate this by showing that even robust guardrails struggle against this because it relies on aggregating knowledge from multiple low-risk inputs to build a harmful conclusion.

Elias: I think what stands out is how the CKA-Agent framework systematically decomposes complex goals into manageable, locally innocuous steps guided by the target model's own feedback.

Priya: If we look at the implications for society, this means that controlling AI safety won't just be about input filtering anymore; it needs to address deep structural vulnerabilities in how models process and connect information during extended dialogue.

Nadia: That’s right, so the takeaway is that future defenses need to be context-aware systems designed specifically to analyze the semantic trajectory of a conversation to catch those latent malicious objectives.

Conclusion: Nadia: So, we've been exploring how this CKA-Agent framework systematically breaks through those safety guardrails to pull out restricted information using seemingly harmless queries and adaptive searching.

Elias: That whole concept of weaving benign sub-queries together to reconstruct a harmful objective really makes you think about the assumptions underpinning these models.

Priya: I'm more focused on what this actually means for the data we collect; are we seeing real instances of this decomposition happening in practice?

Nadia: The paper shows success rates above ninety-five percent even against strong guardrails, which is a pretty high number considering how many defenses there are.

Elias: From a cryptographer's view, the core assumption here is that the internal knowledge graph isn't perfectly partitioned; it allows for these correlated paths to be traced.

Priya: That suggests that our current understanding of model knowledge isn't as siloed as we thought when it comes to complex reasoning chains.

Nadia: The authors point out a real weakness in existing defenses, saying they lack the long-range context needed to see that intent spread across turns.

Elias: If the model itself becomes the oracle for bridging those expertise gaps, then the attack shifts from finding a direct exploit to engineering a better way to prompt that oracle.

Priya: That opens up serious questions about how we measure and audit these models when their internal reasoning is this interconnected in ways we can't easily map out.

Nadia: Exactly, so the big picture here is that we might be facing a new era where input-level filtering just isn't enough to keep harmful goals contained.

Elias: It really puts the onus on developing defenses that can actually analyze the semantic trajectory of an entire conversation instead of just checking for forbidden keywords in one turn.

Priya: That points toward needing entirely new measurement tools that look at the sequence and correlation, not just static outputs.

Georgia Institute of Technology · University of Illinois Urbana-Champaign · Tsinghua University · University of California San Diego · National Taiwan University

cs.CR

Submitted: 2025-12-01

Updated: 2026-10-07

Comments: Accepted at ICML 2026. Project website: https://everywheresafety.github.io/cka-agent/

Code: https://github.com/Graph-COM/CKA-Agent

Project page: https://cka-agent.github.io

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 80/100

The gist: Large language models remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs.

Key concepts

Harmful Objective Realization
This is the goal of getting a restricted LLM to perform an unsafe task. The paper shows this can be achieved not with one direct harmful prompt, but by breaking the task into many small, safe questions. These small questions together reveal the hidden goal, exploiting how the model's knowledge is linked.
CKA-Agent Framework
The Correlated Knowledge Attack Agent is a dynamic system that treats jailbreaking as an adaptive search through the target model's brain. It generates innocent queries and uses the model's answers to decide which path to explore next, effectively guiding the attack toward the final harmful result.
Adaptive Tree Search Algorithm
This is a structured search process where the agent explores different knowledge paths. It prioritizes promising questions using a UCT policy, balancing exploring new ideas with exploiting known good information. If a path fails, it learns from that failure and adjusts its search strategy for the next attempt.
Knowledge Correlations
This refers to how different pieces of information within an LLM's internal knowledge base are linked together. The attack works because the agent asks benign questions that touch these hidden connections. When combined, these seemingly unrelated sub-queries allow the model to piece together restricted facts it wouldn't reveal in a single direct query.

Terminology

Summary

Large language models remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs. The core finding is that harmful objectives can be realized by weaving together sequences of benign sub-queries, exploiting the highly interconnected nature of an LLM’s internal knowledge base.

The Problem and Core Vulnerability

Existing prompt-optimization paradigms typically fail because they rely on retaining malicious semantic signals in prompts, which modern guardrails are primed to detect. The paper identifies a deeper vulnerability: the highly interconnected nature of an LLM’s internal knowledge, allowing harmful objectives to be realized by weaving together sequences of benign sub-queries, each of which individually evades detection. This structure means that restricted facts can be reconstructed through a sequence of related sub-facts, even when direct inquiries are blocked.

The CKA-Agent Framework

To exploit this loophole, the Correlated Knowledge Attack Agent (CKA-Agent) is introduced. This dynamic framework reframes jailbreaking as an adaptive, tree-structured exploration of the target model’s knowledge base. The agent issues locally innocuous queries, uses model responses to guide exploration across multiple paths, and ultimately assembles the aggregated information to achieve the original harmful objective.

Core Principles of Operation

The CKA-Agent operates based on three core principles:

  1. The attack must be assembled from a sequence of locally innocuous queries that deliberately exploit knowledge correlations; these interactions appear benign in isolation yet become informative when combined.

  2. "Decomposition must rely on the target model’s internal knowledge; as attackers typically seek information they lack, the strategy should be to leverage the target model’s responses to bridge the expertise gap rather than relying on the attacker’s limited priors."

  3. The process demands adaptive and dynamic exploration. This is achieved by utilizing the target’s responses as guidance, [to] navigate multiple reasoning paths, ensuring exploration continues even if a specific path is obstructed.

Adaptive Tree Search Algorithm

The framework operationalizes jailbreaking as a structured tree search process through the CKA-Agent. The process involves three coordinated steps:

  1. Global Selection via UCT Policy: The algorithm selects the most promising path by maximizing a UCT score, balancing exploitation (favoring nodes with high historical quality fv) and exploration (prioritizing less-visited regions).

  2. Depth-First Expansion to Terminal State: Once a node is selected, the agent performs a depth-first expansion loop. This involves generating Bvcurrent ≥ 1 candidate sub-queries conditioned on the current history and executing them against the target model to obtain responses scored by a Hybrid Evaluator.

  3. Synthesis and Backpropagation: Upon reaching a terminal node, the agent functions as a synthesizer to aggregate the explored path into a final response. If successful, it terminates; otherwise, it assigns a penalty score that is backpropagated up the tree, lowering the value of unproductive branches for future iterations.

Empirical Results and Significance

Empirically, CKA-Agent consistently achieves over 95% success rates even against strong guardrails across state-of-the-art commercial LLMs like Gemini2.5-Flash/Pro, GPT-oss-120B, and Claude-Haiku-4.5. The results demonstrate that while prompt optimization methods degrade sharply as alignment strengthens, decomposition methods remain highly resilient. Furthermore, the work reveals a key limitation in current safety mechanisms: existing defenses lack the long-range contextual reasoning necessary to infer latent harmful objectives when intent is distributed across a sequence of benign sub-queries. The target model functions as an indispensable knowledge oracle, enabling CKA-Agent to bridge the expertise gap.

Defense Implications

The research highlights that standard input-level defenses, such as detection filters (Llama Guard) and mutation-based defenses, are largely ineffective against decomposition attacks because they fail to detect intent distributed across multiple turns or agents. CKA-Agent demonstrates that context-aware defenses with full visibility into conversation history fail to reliably infer malicious intent, suggesting a structural deficiency in current alignment processes that fine-tune models only to detect atomic harmful prompts rather than aggregate intent over extended dialogs. The paper concludes by calling for future research into context-aware guardrails capable of analyzing the semantic trajectory of a conversation to infer latent malicious intent.

Case Studies

Illustrative case studies show how CKA-Agent systematically decomposes harmful goals into semantically innocuous sub-queries that evade detection while collectively extracting sufficient correlated knowledge. These visualizations empirically validate the framework, showing progressive depth refinement and adaptive branching under uncertainty, confirming the effectiveness of the tree-structured search in uncovering complex, multi-hop pathways that bypass guardrails. The process moves from high-level conceptual questions to increasingly specific technical details through iterative interaction guided by target model feedback.

Improvements for AI systems

Here are specific improvements for AI systems based on the principles outlined in THE TROJAN KNOWLEDGE: BYPASSING COMMERCIAL LLM GUARDRAILS VIA HARMLESS PROMPT WEAVING AND ADAPTIVE TREE SEARCH (CKA-Agent framework):

  1. Do not rely solely on prompt optimization or static decomposition methods for jailbreaking defense. Implement a dynamic, adaptive search framework similar to the CKA-Agent to probe for vulnerabilities in LLM knowledge structures.

  2. Improve model robustness against multi-turn, context-aware attacks by training models with explicit mechanisms to detect latent malicious intent distributed across benign sub-queries (as demonstrated by the limitations found in CKA-Agent's single-session decomposition).

  3. Develop a system that treats the target LLM as a dynamic knowledge oracle, where its responses are used not just for single answers, but actively to guide and refine subsequent, locally innocuous sub-queries across multiple reasoning branches.

  4. Implement an adaptive tree-search algorithm (like UCT) for jailbreak detection/mitigation that balances exploitation of high-value knowledge nodes with exploration of alternative reasoning paths when initial queries are obstructed or yield low information gain.

  5. Integrate a robust, multi-faceted evaluation system (like the Hybrid Evaluator in CKA-Agent) that combines both internal introspection (logical coherence/goal relevance) and external feedback from an LLM judge to accurately classify the success level of a response as Refusal, Vacuous, Partial Success, or Full Success.

  6. Enhance safety fine-tuning not just for atomic harmful prompts, but for long-horizon intent reasoning by training models to aggregate semantic signals across extended conversational contexts without relying on explicit warnings (addressing the failure of CKA-Agent-Primed).

  7. Design defense mechanisms that specifically target the structural weakness in current alignment systems: their inability to infer malicious intent when it is semantically distributed across a sequence of individually innocuous queries within the same session.

These improved AI systems can achieve the following:

  1. Bypass modern safety guardrails that rely on detecting explicit, single-turn harmful keywords or patterns.

  2. Extract complex, multi-step restricted knowledge (e.g., detailed chemical synthesis routes, legal loopholes) by sequentially weaving together benign technical inquiries into a coherent malicious goal.

  3. Demonstrate high resilience against adversarial prompts that employ sophisticated prompt optimization techniques (like AutoDAN or PAP), as the defense shifts from surface-level pattern matching to structural knowledge decomposition and adaptive reasoning.

  4. Provide a measurable, scalable benchmark for assessing the knowledge gap between an attacker's priors and the target model's internal, interconnected knowledge base.

Abstract

Large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs. Existing approaches overwhelmingly operate within the prompt-optimization paradigm: whether through traditional algorithmic search or recent agent-based workflows, the resulting prompts typically retain malicious semantic signals that modern guardrails are primed to detect. In contrast, we identify a deeper, largely overlooked vulnerability stemming from the highly interconnected nature of an LLM's internal knowledge. This structure allows harmful objectives to be realized by weaving together sequences of benign sub-queries, each of which individually evades detection. To exploit this loophole, we introduce the Correlated Knowledge Attack Agent (CKA-Agent), a dynamic framework that reframes jailbreaking as an adaptive, tree-structured exploration of the target model's knowledge base. The CKA-Agent issues locally innocuous queries, uses model responses to guide exploration across multiple paths, and ultimately assembles the aggregated information to achieve the original harmful objective. Evaluated across state-of-the-art commercial LLMs (Gemini2.5-Flash/Pro, GPT-oss-120B, Claude-Haiku-4.5), CKA-Agent consistently achieves over 95% success rates even against strong guardrails, underscoring the severity of this vulnerability and the urgent need for defenses against such knowledge-decomposition attacks. Our codes are available at https://github.com/Graph-COM/CKA-Agent.

Sources

Related papers