The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search
summary
The gist
Large language models remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs.
In short
Researchers developed a method called CKA-Agent to bypass LLM safety guardrails by weaving together many seemingly harmless questions. The agent uses an adaptive tree search strategy, asking benign sub-queries that exploit the model's interconnected knowledge. This approach succeeds by letting the model reconstruct harmful information through a sequence of safe interactions, showing current defenses fail against distributed intent.
Key concepts
- Harmful Objective Realization
- This is the goal of getting a restricted LLM to perform an unsafe task. The paper shows this can be achieved not with one direct harmful prompt, but by breaking the task into many small, safe questions. These small questions together reveal the hidden goal, exploiting how the model's knowledge is linked.
- CKA-Agent Framework
- The Correlated Knowledge Attack Agent is a dynamic system that treats jailbreaking as an adaptive search through the target model's brain. It generates innocent queries and uses the model's answers to decide which path to explore next, effectively guiding the attack toward the final harmful result.
- Adaptive Tree Search Algorithm
- This is a structured search process where the agent explores different knowledge paths. It prioritizes promising questions using a UCT policy, balancing exploring new ideas with exploiting known good information. If a path fails, it learns from that failure and adjusts its search strategy for the next attempt.
- Knowledge Correlations
- This refers to how different pieces of information within an LLM's internal knowledge base are linked together. The attack works because the agent asks benign questions that touch these hidden connections. When combined, these seemingly unrelated sub-queries allow the model to piece together restricted facts it wouldn't reveal in a single direct query.
Terminology used across episodes
This episode discusses
- The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search · Paper Radio
- A Survey of Large Language Models for Financial Applications: Progress, Prospects and Challenges
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- On the Opportunities and Risks of Foundation Models
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
- Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Red Teaming Language Models with Language Models
- Jailbreak Attacks and Defenses Against Large Language Models: A Survey
- Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
- Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Qwen3Guard Technical Report
- Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
- Evaluating Deep Unlearning in Large Language Models
- Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search
The paper
The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search · Read on arXiv
Georgia Institute of Technology · University of Illinois Urbana-Champaign · Tsinghua University · University of California San Diego · National Taiwan University
Large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs. Existing approaches overwhelmingly operate within the prompt-optimization paradigm: whether through traditional algorithmic search or recent agent-based workflows, the resulting prompts typically retain malicious semantic signals that modern guardrails are primed to detect. In contrast, we identify a deeper, largely overlooked vulnerability stemming from the highly interconnected nature of an LLM's internal knowledge. This structure allows harmful objectives to be realized by weaving together sequences of benign sub-queries, each of which individually evades detection. To exploit this loophole, we introduce the Correlated Knowledge Attack Agent (CKA-Agent), a dynamic framework that reframes jailbreaking as an adaptive, tree-structured exploration of the target model's knowledge base. The CKA-Agent issues locally innocuous queries, uses model responses to guide exploration across multiple paths, and ultimately assembles the aggregated information to achieve the original harmful objective. Evaluated across state-of-the-art commercial LLMs (Gemini2.5-Flash/Pro, GPT-oss-120B, Claude-Haiku-4.5), CKA-Agent consistently achieves over 95% success rates even against strong guardrails, underscoring the severity of this vulnerability and the urgent need for defenses against such knowledge-decomposition attacks. Our codes are available at https://github.com/Graph-COM/CKA-Agent.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "The Trojan Knowledge".
Elias: Large language models remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs.
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: So, we've discussed how "The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search" suggests that harmful outputs are achievable by weaving together harmless queries to exploit the interconnected nature of an AI's knowledge.
Elias: That brings us back to the idea that we can't just look for bad keywords; instead, we have to consider how information is structured internally and how agents can systematically explore that structure through adaptive tree searches.
Priya: From a measurement perspective, the real significance lies in realizing that current safety mechanisms aren't looking at the right scope; they are missing the distribution of intent across a sequence of benign interactions.
Nadia: The authors demonstrate this by showing that even robust guardrails struggle against this because it relies on aggregating knowledge from multiple low-risk inputs to build a harmful conclusion.
Elias: I think what stands out is how the CKA-Agent framework systematically decomposes complex goals into manageable, locally innocuous steps guided by the target model's own feedback.
Priya: If we look at the implications for society, this means that controlling AI safety won't just be about input filtering anymore; it needs to address deep structural vulnerabilities in how models process and connect information during extended dialogue.
Nadia: That’s right, so the takeaway is that future defenses need to be context-aware systems designed specifically to analyze the semantic trajectory of a conversation to catch those latent malicious objectives.
Conclusion: Nadia: So, we've been exploring how this CKA-Agent framework systematically breaks through those safety guardrails to pull out restricted information using seemingly harmless queries and adaptive searching.
Elias: That whole concept of weaving benign sub-queries together to reconstruct a harmful objective really makes you think about the assumptions underpinning these models.
Priya: I'm more focused on what this actually means for the data we collect; are we seeing real instances of this decomposition happening in practice?
Nadia: The paper shows success rates above ninety-five percent even against strong guardrails, which is a pretty high number considering how many defenses there are.
Elias: From a cryptographer's view, the core assumption here is that the internal knowledge graph isn't perfectly partitioned; it allows for these correlated paths to be traced.
Priya: That suggests that our current understanding of model knowledge isn't as siloed as we thought when it comes to complex reasoning chains.
Nadia: The authors point out a real weakness in existing defenses, saying they lack the long-range context needed to see that intent spread across turns.
Elias: If the model itself becomes the oracle for bridging those expertise gaps, then the attack shifts from finding a direct exploit to engineering a better way to prompt that oracle.
Priya: That opens up serious questions about how we measure and audit these models when their internal reasoning is this interconnected in ways we can't easily map out.
Nadia: Exactly, so the big picture here is that we might be facing a new era where input-level filtering just isn't enough to keep harmful goals contained.
Elias: It really puts the onus on developing defenses that can actually analyze the semantic trajectory of an entire conversation instead of just checking for forbidden keywords in one turn.
Priya: That points toward needing entirely new measurement tools that look at the sequence and correlation, not just static outputs.
More episodes
- 2610.10597-Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
- 2610.10608-From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage
- 2610.10612-PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits