Omni-IO Skills: Harnessing Your Agent Omni-Native
cs.CL
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/any2any-mllm/Omni-IO-Skill
Terminology
Sources
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- CUA-Skill: Develop Skills for Computer Using Agent
- MMSkills: Towards Multimodal Skills for General Visual Agents
- Towards Comprehensive Stage-wise Benchmarking of Large Language Models in Fact-Checking
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
- Qwen2.5-Omni Technical Report
- Moshi: a speech-text foundation model for real-time dialogue
- Qwen3-Omni Technical Report
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection
- MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
- OmniGAIA: Towards Native Omni-Modal AI Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering