Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning
cs.CV, cs.AI
Submitted: 2026-02-01
Updated: 2026-08-27
Comments: Accepted to SIGGRAPH ASIA 2026; Project page: https://yuci-gpt.github.io/Beyond-Pixels/
Project page: https://yuci-gpt.github.io/Beyond-Pixels
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- The Mind's Eye: A Multi-Faceted Reward Framework for Guiding Visual Metaphor Generation
- HunyuanImage 3.0 Technical Report
- I Spy a Metaphor: Large Language Models and Diffusion Models Co-Create Visual Metaphors
- Emerging Properties in Unified Multimodal Pretraining
- Seedream 4.0: Toward Next-generation Multimodal Image Generation
- HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads
- TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models