Do System Prompts Leave Behavioral Fingerprints? A Large-Scale Empirical Study of Clone Detection via Output Similarity
cs.CR
Submitted: 2026-08-25
Updated: 2026-08-25
Terminology
Sources
- ProxyPrompt: Securing System Prompts against Prompt Extraction Attacks
- System Prompt Extraction Attacks and Defenses in Large Language Models
- Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability
- PathMark: Protecting Intellectual Property of Mixture-of-Expert LLMs via Path Watermarks
- Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- FragileFlow: Spectral Control of Correct-but-Fragile Predictions for Foundation Model Robustness
- SCOPE: Sequential Conformal Probing for Reliable OOD Rejection in LLM Services
- Agent Hacks Agents: Autoresearch Discovers Vulnerabilities in Production Agents
- A Fingerprint for Large Language Models
- MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory Systems
- Prompt Stealing Attacks Against Large Language Models
- RAFP: Identifying LLM Lineages via Rare-Region Fingerprints
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs