A Function-Level Vulnerability Score Measures Flag Rate More Than the Model: Protocol Effects on Paired Benchmarks
cs.LG
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- Understanding intermediate layers using linear classifier probes
- The Internal State of an LLM Knows When It's Lying
- Probing Classifiers: Promises, Shortcomings, and Advances
- Discovering Latent Knowledge in Language Models Without Supervision
- Deep Learning based Vulnerability Detection: Are We There Yet?
- Data Quality for Software Vulnerability Datasets
- Vulnerability Detection with Code Language Models: How Far Are We?
- Vul-RAG: Enhancing LLM-based Vulnerability Detection via Knowledge-level RAG
- Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals
- Designing and Interpreting Probes with Control Tasks
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
- "Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization
- From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- ARVO: Atlas of Reproducible Vulnerabilities for Open-Source Software
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- Probing the Probing Paradigm: Does Probing Accuracy Entail Task Relevance?
- Top Score on the Wrong Exam: On Benchmarking in Machine Learning for Vulnerability Detection
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks