Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents
cs.AI, cs.CL
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 30 pages, Accepted to FinNLP 2026 Workshop @ EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM) pricing agents may respond to how market data is presented, even when its numerical values remain unchanged.
Terminology
Abstract
Large language model (LLM) pricing agents may respond to how market data is presented, even when its numerical values remain unchanged. We introduce market signal injection (MSI), an attack that manipulates numerical formatting, competitor ordering, or qualitative market commentary without issuing explicit instructions. We evaluate nine open-weight models in simulated Bertrand duopoly and triopoly markets and three proprietary models in duopoly markets. Sentiment-based attacks produce the largest behavioral shifts, which propagate to other firms and alter profits and consumer surplus. Susceptibility varies across model families, and larger models are not consistently more robust. Matched neutral-text controls and a rule-based agent support a framing-based account of these shifts under the fixed demand parameters of our simulation. Episode-held-out probes distinguish baseline from attacked activations in all eleven re-evaluated model--condition pairs: linear AUC is 1.00 and MLP AUC ranges from 0.93 to 0.99. This separability does not by itself identify harmful pricing decisions. Input canonicalization removes the tested sentiment attacks, while decision boundary anchoring, which combines prompt constraints with output projection, provides partial mitigation under the tested adaptive attacks. These results identify data presentation as an attack surface for LLM pricing agents and motivate defenses that account for interactions among agents.
Sources
- Understanding intermediate layers using linear classifier probes
- Reasoning Models Don't Always Say What They Think
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
- From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents
- Algorithmic Collusion by Large Language Models
- The Llama 3 Herd of Models
- Mistral 7B
- Strategic Collusion of LLM Agents: Market Division in Multi-Commodity Competitions
- Qwen2.5 Technical Report
- Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs
- Gemma 2: Improving Open Language Models at a Practical Size
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection