A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions
cs.CV, cs.AI
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/BootsofLagrangian/audit-recap-t2i
Terminology
Sources
- HunyuanImage 3.0 Technical Report
- ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
- A Standardized Machine-readable Dataset Documentation Format for Responsible AI
- What If We Recaption Billions of Web Images with LLaMA-3?
- Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
- A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation
- From Pixels to Prose: A Large Dataset of Dense Image Captions
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- EmbeddingGemma: Powerful and Lightweight Text Representations
- Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
- Qwen-Image Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models