COMIC: Agentic Sketch Comedy Generation
summary
The gist
The gist: COMIC is a fully automated agentic system that produces short comedic videos similar to sketch shows by employing multiple agents and iterative competition to generate funny scripts and
In short
COMIC is an agentic system that creates short comedic videos by mimicking human production studios using multiple specialized agents. It uses iterative competition between scriptwriters, critics, and editors to refine scripts and visuals. This multi-agent approach overcomes the limitations of fixed objectives by allowing diverse perspectives to generate high-quality, consistent sketch comedy.
Key concepts
- Agentic Flow
- The system models a human production studio using specialized agents like writers and directors. A forward pipeline moves from concept generation to script writing, followed by critical evaluation and iterative revision. This mimics the collaborative process of real creative teams.
- Multi-Agent Iterative Competition
- Instead of fixed rules, COMIC uses competition between different critic committees to improve scripts. Scripts are organized into 'islands,' each with a specialized committee that sets unique standards. Losing scripts are then refined by winners, allowing the system to explore a wider range of comedic styles.
- Script Writing Loop
- Scripts evolve through an iterative process where populations of potential scripts are partitioned into islands. Each island has its own critic committee defining the evaluation standards. Scripts are refined through round-robin pairwise evaluations, integrating feedback with semantic mutation for continuous improvement.
Terminology used across episodes
This episode discusses
- COMIC: Agentic Sketch Comedy Generation · Paper Radio
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- TAG: Tangential Amplifying Guidance for Hallucination-Resistant Sampling
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- MusicInfuser: Making Video Diffusion Listen and Dance
- DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation
- FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- LLM-grounded Video Diffusion Models
- VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
- Evolution of Heuristics: Towards Efficient Automatic Algorithm Design Using Large Language Model
- VISTA: A Test-Time Self-Improving Video Generation Agent
- Illuminating search spaces by mapping elites
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- AniMaker: Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation
- Wan: Open and Advanced Large-Scale Video Generative Models
- Accelerating Scientific Research with Gemini: Case Studies and Common Techniques
- Automated Movie Generation via Multi-Agent CoT Planning
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
- VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
The paper
COMIC: Agentic Sketch Comedy Generation · Read on arXiv
Susung Hong, Brian Curless, Ira Kemelmacher-Shlizerman, Steve Seitz
University of Washington
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "COMIC: Agentic Sketch Comedy Generation".
Tom: The gist:
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, looking at the whole thing with "COMIC: Agentic Sketch Comedy Generation", it seems like they've established a new way to build automated video production by using this multi-island topology where different critic committees drive iterative refinement.
Jane: It really moves away from just trying to hit one perfect target for humor and instead lets the system explore a wide range of comedic styles, which is what they claim is important for keeping things interesting.
Lu: The connection to the Red Queen hypothesis in evolutionary biology suggests that these systems need continuous evolution to maintain their fitness against co-evolving competitors.
Meng: For practical application, it shows how you can structure a complex creative task into manageable agentic loops that actually produce high-quality, diverse outputs on a reasonable computational budget.
Lalam: And I think the focus on using diverse critic committees really helps improve culture by showing how to build systems that can generate content with more variety and nuance than single-pass methods allow.
Tom: So, it's not just about making one video; it’s about creating a framework that can sustain a diverse range of comedic styles, which is what this work shows is possible.
Jane: It establishes a new state of the art for fully automated video production by leveraging this system where human-aligned critic committees drive iterative refinement.
Tom: That’s pretty much it on COMIC—a lot of structured iteration leading to videos that look and feel more like actual sketch comedy than anything we’ve seen before.
Conclusion: Tom: So, COMIC is this system that builds short comedy videos using multiple AI agents working together in a competition to make funny scripts and visuals that stay consistent.
Jane: It’s basically modeling a whole production studio with AI versions of writers and directors, trying to generate sketches just like real shows do.
Lu: The core idea is this forward pipeline where ideas go from concepts into full dialogues, then get critiqued by other agents, edited based on feedback, and finally turned into video sequences.
Meng: It’s interesting how they broke the big problem down into script generation and visual realization so they could tackle each part specifically.
Lalam: The system handles humor through this multi-agent iterative competition instead of just trying to find one perfect funny script upfront.
Tom: Yeah, and what I find really striking is how it tackles the issue of subjective humor by using these coupled subproblems for both writing and visuals.
Jane: It moves away from those old methods where you just set a fixed goal for the AI, because humor is so context-dependent and personal.
Lu: They found that instead of one perfect objective, they use an evolving population of scripts across different "islands," each with its own critic committee to shape the style.
Meng: That sounds like a way to explore a lot more comedic territory than just following one set of rules for what's funny.
Lalam: It really shows how this structure can sustain diverse comedic styles, which is something single-pass methods usually miss entirely.
Tom: So, the authors are basically proposing this framework as a new way to approach automated video production that aims for more variety and better alignment with what people actually find funny.
Jane: They’re calling it COMIC, and the whole point is that it uses this iterative competition to generate content that actually feels like it came from a real creative process.
Lu: It connects this idea to the Red Queen hypothesis in biology, suggesting that these systems need constant evolution to stay competitive in generating engaging content.
Meng: From an engineering side, it’s impressive how they managed the complexity of making those multiple agent interactions work together reliably on a reasonable computational setup.
Lalam: And for me, this means we can build systems that understand cultural nuances better because they're trained against a dynamic set of evaluators, not just static rules.
Tom: So, it’s less about finding the single funniest thing and more about building an engine that can keep producing a wide range of entertaining content consistently.
Jane: That’s the simple version—it’s a system designed to produce diverse, high-quality comedic videos through constant, structured competition between different parts of the AI.
Lu: And this structure opens up some really cool avenues for how we design these multi-agent systems in general because it shows that iteration is key.
Meng: It’s definitely a more realistic path than just giving one giant model a prompt and hoping for the best outcome every time.
Lalam: This whole approach is powerful because it addresses the difficulty of capturing human preference profiles, which is what makes humor so hard to automate.
Tom: So, COMIC sets up a new benchmark by showing that this kind of structured iteration leads to results that actually look and feel like real sketch comedy.
Jane: Next time we talk about this system, we’re going to dig into the specifics of how they're selecting those critics from the crowd.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck