RAPTOR: Ridge-Adaptive Logistic Probes
cs.LG, cs.AI
Submitted: 2026-01-29
Updated: 2026-09-25
Terminology
Sources
- Understanding intermediate layers using linear classifier probes
- Toward universal steering and monitoring of AI models
- What do Neural Machine Translation Models Learn about Morphology?
- Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
- Steering Llama 2 via Contrastive Activation Addition
- The Impact of Regularization on High-dimensional Logistic Regression
- Improving Instruction-Following in Language Models through Activation Steering
- Analyzing the Generalization and Reliability of Steering Vectors
- BERT Rediscovers the Classical NLP Pipeline
- Steering Language Models With Activation Engineering
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks