Understanding and Exploiting Initialization Anchoring Weakness in Feedback-Based Agent Planning
cs.CR, cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/langchain-ai/langchain
Terminology
Sources
- ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent
- Detecting Language Model Attacks with Perplexity
- Thought Anchors: Which LLM Reasoning Steps Matter?
- Defeating Prompt Injections by Design
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification
- Understanding the planning of LLM agents: A survey
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- AgentBench: Evaluating LLMs as Agents
- Ignore Previous Prompt: Attack Techniques For Language Models
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
- From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents
- Cognitive Overload Attack:Prompt Injection for Long Context
- ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
- ReAct: Synergizing Reasoning and Acting in Language Models
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
- Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
- MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs