What Makes Synthetic Hard Negatives Work in Vision-Language Pretraining?
cs.CV, cs.AI
Submitted: 2026-10-07
Updated: 2026-10-07
Code: https://github.com/rom1504/img2dataset
Terminology
Sources
- SynCo: Synthetic Hard Negatives for Contrastive Visual Representation Learning
- ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives
- Explaining and Harnessing Adversarial Examples
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Do Better ImageNet Models Transfer Better?
- Scaling Language-Image Pre-training via Masking
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning
- Microsoft COCO: Common Objects in Context
- CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
- MLLMs-Augmented Visual-Language Representation Learning
- Fine-Grained Visual Classification of Aircraft
- SILC: Improving Vision Language Pretraining with Self-Distillation
- Representation Learning with Contrastive Predictive Coding
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- Perceptual Grouping in Contrastive Vision-Language Models
- Contrastive Learning with Hard Negative Samples
- EVA-CLIP: Improved Training Techniques for CLIP at Scale
- EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Manifold Mixup: Better Representations by Interpolating Hidden States
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models