Same Probe, Different Numbers: Are Activation Probes Robust to Inference-Time Numerical Non-Determinism?
cs.LG, cs.AI
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/huggingface/transformers
Terminology
Sources
- Understanding intermediate layers using linear classifier probes
- Non-Determinism of "Deterministic" LLM Settings
- Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- Gemma 3 Technical Report
- Detecting Strategic Deception Using Linear Probes
- The Llama 3 Herd of Models
- Language Models Represent Space and Time
- Designing and Interpreting Probes with Control Tasks
- A Study of BFLOAT16 for Deep Learning Training
- Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families
- Similarity of Neural Network Representations Revisited
- BERT Busters: Outlier Dimensions that Disrupt Transformers
- Building Production-Ready Probes For Gemini
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Detecting High-Stakes Interactions with Activation Probes
- Mixed Precision Training
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks