The Backdrop Exposes What the World Around an Agent Costs It
cs.CL, cs.AI
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/NusRAT-LiA/BackDrop3https:
Project page: https://nusrat-lia.github.io/BackDrop
Terminology
Sources
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
- Evaluating Large Language Models Trained on Code
- Defeating Prompt Injections by Design
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
- Kimi K2.5: Visual Agentic Intelligence
- RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
- DAGGER: Distractor-Aware Graph Generation for Executable Reasoning in Math Problems
- Qwen3 Technical Report
- Progent: Securing AI Agents with Privilege Control
- Authenticated Delegation and Authorized AI Agents
- OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Agent-SafetyBench: Evaluating the Safety of LLM Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering