CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings
cs.CL, cs.IR
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/mrehacek/cg-probes
Terminology
Sources
- Understanding intermediate layers using linear classifier probes
- Refusal in Language Models Is Mediated by a Single Direction
- LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
- Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
- Language Models are Few-Shot Learners
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- BERTopic: Neural topic modeling with a class-based TF-IDF procedure
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- gpt-oss-120b & gpt-oss-20b Model Card
- Refusal Direction is Universal Across Safety-Aligned Languages
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering