SafeGEO: Understanding Generative Engine Optimization Risks in Recommendation Agents
cs.IR, cs.AI
Submitted: 2026-06-08
Updated: 2026-09-04
Comments: 41 pages,22 figures
Code: https://github.com/QianfengWen/SafeGEO
License: http://creativecommons.org/licenses/by/4.0/
The gist: Generative Engine Optimization (GEO) lets content owners rewrite web content to increase their visibility in generative systems.
Terminology
Abstract
Generative Engine Optimization (GEO) lets content owners rewrite web content to increase their visibility in generative systems. In recommendation agents, this creates a risk that sources controlled by sellers make flawed products appear better supported than they are. We study this risk at the generation stage by asking whether recommendation agents continue to make decisions that align with user utility when these sources are rewritten for GEO. To make this question measurable, we construct SafeGEO, an evaluation suite with 22 GEO attack variants across 600 recommendation cases. We empirically show that GEO attacks can promote flawed target products: they increase the rate at which such flawed products enter the recommendation set by up to 83.2 percentage points (pp). We further study whether agent-side design choices can mitigate this risk and show that simple defenses reduce harmful target promotion by up to 39.2 pp. These gains are substantial but do not restore the no-GEO performance, showing that GEO remains a serious risk despite mitigation.
Sources
- E-GEO: A Testbed for Generative Engine Optimization in E-Commerce
- ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement
- Bayesian Active Learning with Gaussian Processes Guided by LLM Relevance Scoring for Dense Passage Retrieval
- OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment
- SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization
- Manipulating Large Language Models to Increase Product Visibility
- GPT-4 Technical Report
- MemGPT: Towards LLMs as Operating Systems
- Evaluating Scene-based In-Situ Item Labeling for Immersive Conversational Recommendation
- Goal-Oriented Reasoning for RAG-based Memory in Conversational Agentic LLM Systems
- Multimodal Item Scoring for Natural Language Recommendation via Gaussian Process Regression with LLM Relevance Judgments
- Measuring Attribution in Natural Language Generation Models
- Temporal Order Matters for Agentic Memory: Segment Trees for Long-Horizon Agents
- Semantic XPath: Structured Agentic Memory Access for Conversational AI
- RecRanker: Instruction Tuning Large Language Model as Ranker for Top-k Recommendation
- Grounded Chess Reasoning in Language Models via Master Distillation
- Is Your LLM-as-a-Recommender Agent Trustable? LLMs' Recommendation is Easily Hacked by Biases (Preferences)
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- A Simple but Effective Elaborative Query Reformulation Approach for Natural Language Recommendation
- Elaborative Subtopic Query Reformulation for Broad and Indirect Queries in Travel Destination Recommendation
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG