When Good Verifiers Go Bad: Silent Negative Transfer in Verifier-Guided VLM Training
cs.CR, cs.AI
Submitted: 2026-06-12
Updated: 2026-09-25
Comments: 15 pages
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Reinforced Self-Training (ReST) for Language Modeling
- Towards Continuous Intelligence Growth: Self-Training, Continual Learning, and Dual-Scale Memory in SuperIntelliAgent
- Iterative Reasoning Preference Optimization
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs