COMIC: Agentic Sketch Comedy Generation

arXiv:2603.11048 · cs.CV, cs.AI, cs.CL, cs.MA, cs.NE · Submitted 2026-03-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "COMIC: Agentic Sketch Comedy Generation".

Tom: The gist:

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, looking at the whole thing with "COMIC: Agentic Sketch Comedy Generation", it seems like they've established a new way to build automated video production by using this multi-island topology where different critic committees drive iterative refinement.

Jane: It really moves away from just trying to hit one perfect target for humor and instead lets the system explore a wide range of comedic styles, which is what they claim is important for keeping things interesting.

Lu: The connection to the Red Queen hypothesis in evolutionary biology suggests that these systems need continuous evolution to maintain their fitness against co-evolving competitors.

Meng: For practical application, it shows how you can structure a complex creative task into manageable agentic loops that actually produce high-quality, diverse outputs on a reasonable computational budget.

Lalam: And I think the focus on using diverse critic committees really helps improve culture by showing how to build systems that can generate content with more variety and nuance than single-pass methods allow.

Tom: So, it's not just about making one video; it’s about creating a framework that can sustain a diverse range of comedic styles, which is what this work shows is possible.

Jane: It establishes a new state of the art for fully automated video production by leveraging this system where human-aligned critic committees drive iterative refinement.

Tom: That’s pretty much it on COMIC—a lot of structured iteration leading to videos that look and feel more like actual sketch comedy than anything we’ve seen before.

Conclusion: Tom: So, COMIC is this system that builds short comedy videos using multiple AI agents working together in a competition to make funny scripts and visuals that stay consistent.

Jane: It’s basically modeling a whole production studio with AI versions of writers and directors, trying to generate sketches just like real shows do.

Lu: The core idea is this forward pipeline where ideas go from concepts into full dialogues, then get critiqued by other agents, edited based on feedback, and finally turned into video sequences.

Meng: It’s interesting how they broke the big problem down into script generation and visual realization so they could tackle each part specifically.

Lalam: The system handles humor through this multi-agent iterative competition instead of just trying to find one perfect funny script upfront.

Tom: Yeah, and what I find really striking is how it tackles the issue of subjective humor by using these coupled subproblems for both writing and visuals.

Jane: It moves away from those old methods where you just set a fixed goal for the AI, because humor is so context-dependent and personal.

Lu: They found that instead of one perfect objective, they use an evolving population of scripts across different "islands," each with its own critic committee to shape the style.

Meng: That sounds like a way to explore a lot more comedic territory than just following one set of rules for what's funny.

Lalam: It really shows how this structure can sustain diverse comedic styles, which is something single-pass methods usually miss entirely.

Tom: So, the authors are basically proposing this framework as a new way to approach automated video production that aims for more variety and better alignment with what people actually find funny.

Jane: They’re calling it COMIC, and the whole point is that it uses this iterative competition to generate content that actually feels like it came from a real creative process.

Lu: It connects this idea to the Red Queen hypothesis in biology, suggesting that these systems need constant evolution to stay competitive in generating engaging content.

Meng: From an engineering side, it’s impressive how they managed the complexity of making those multiple agent interactions work together reliably on a reasonable computational setup.

Lalam: And for me, this means we can build systems that understand cultural nuances better because they're trained against a dynamic set of evaluators, not just static rules.

Tom: So, it’s less about finding the single funniest thing and more about building an engine that can keep producing a wide range of entertaining content consistently.

Jane: That’s the simple version—it’s a system designed to produce diverse, high-quality comedic videos through constant, structured competition between different parts of the AI.

Lu: And this structure opens up some really cool avenues for how we design these multi-agent systems in general because it shows that iteration is key.

Meng: It’s definitely a more realistic path than just giving one giant model a prompt and hoping for the best outcome every time.

Lalam: This whole approach is powerful because it addresses the difficulty of capturing human preference profiles, which is what makes humor so hard to automate.

Tom: So, COMIC sets up a new benchmark by showing that this kind of structured iteration leads to results that actually look and feel like real sketch comedy.

Jane: Next time we talk about this system, we’re going to dig into the specifics of how they're selecting those critics from the crowd.

Susung Hong, Brian Curless, Ira Kemelmacher-Shlizerman, Steve Seitz

University of Washington

cs.CV, cs.AI, cs.CL, cs.MA, cs.NE

Submitted: 2026-03-11

Updated: 2026-10-04

Comments: Project page: https://susunghong.github.io/COMIC/

Code: https://github.com/resemble-ai/chatterbox

Project page: https://susunghong.github.io/COMIC

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 88/100

The gist: The gist: COMIC is a fully automated agentic system that produces short comedic videos similar to sketch shows by employing multiple agents and iterative competition to generate funny scripts and

Key concepts

Agentic Flow
The system models a human production studio using specialized agents like writers and directors. A forward pipeline moves from concept generation to script writing, followed by critical evaluation and iterative revision. This mimics the collaborative process of real creative teams.
Multi-Agent Iterative Competition
Instead of fixed rules, COMIC uses competition between different critic committees to improve scripts. Scripts are organized into 'islands,' each with a specialized committee that sets unique standards. Losing scripts are then refined by winners, allowing the system to explore a wider range of comedic styles.
Script Writing Loop
Scripts evolve through an iterative process where populations of potential scripts are partitioned into islands. Each island has its own critic committee defining the evaluation standards. Scripts are refined through round-robin pairwise evaluations, integrating feedback with semantic mutation for continuous improvement.

Terminology

Summary

The gist: COMIC is a fully automated agentic system that produces short comedic videos similar to sketch shows by employing multiple agents and iterative competition to generate funny scripts and high-quality, consistent video.

How it works

  1. The system is loosely modeled on human production studios, using agentic versions of roles such as scriptwriters, editors, and directors (Fig. 2) <ref:2603.11048#pg3>.

  2. The overall agentic flow involves a forward pipeline where a writer generates concepts and expands them into full dialogues, a critic evaluates and compares scripts, and an editor revises scripts based on critic feedback <ref:2603.11048#pg6>.

  3. Script generation synthesizes a script s∗ ∈ S that establishes a compelling comedic premise, develops it through character interaction, and delivers a satisfying payoff <ref:2603.11048#pg6>.

  4. Visual realization translates s∗ into a shot sequence V = [v1,..., vN] that faithfully embodies the narrative while maintaining character identity and scene continuity <ref:2603.11048#pg6>.

Content Optimization via Multi-Agent Iterative Competition (COMIC)

The system addresses the problem of automated sketch comedy video generation by decomposing it into two coupled subproblems: script generation and visual realization <ref:2603.11048#pg6>. The approach is based on the observation that LLMs do occasionally produce humorous content when provided with the right structure, mirroring how real sketch comedy shows operate through brainstorming and iteration <ref:2603.11048#pg6>.

Why Fixed Objectives Fall Short

Traditional goal-based optimization fails because humor is inherently context-dependent and subjective, meaning a fixed scalar objective invites Goodhart’s Law [8] <ref:2603.11048#pg6>. Humor takes many different forms, and it is impossible to aggregate preference profiles into a single, consistent ranking without sacrificing desirable properties <ref:2603.11048#pg6>. Existing agentic strategies are limited because each role is defined by a fixed instruction, meaning the agent always applies the same evaluative lens with no mechanism to explore alternative perspectives <ref:2603.11048#pg6>.

Alignment to Real Viewers

Evaluation quality depends critically on critic selection, which is achieved through a generate-and-select strategy rather than fine-tuning a dedicated critic model <ref:2603.11048#pg8>. This involves synthesizing a large, diverse pool of candidate critics and retaining those whose preferences best align with empirical audience engagement signals <ref:2603.11048#pg8>. Engagement scoring is derived by fitting a per-channel logistic growth model to normalized view counts to project the carrying capacity Lproj, which defines each video’s engagement score <ref:2603.11048#pg8>. Task-specific selection involves defining two comparison tasks, Top vs. Bottom and Top vs. Middle, and selecting the highest-accuracy critic on those pairwise comparisons for each channel and sensitivity level <ref:2603.11048#pg8>.

Script Writing Loop

COMIC utilizes an approach to iteratively evolve a population of scripts by partitioning the global script population into K isolated islands, each governed by a specialized critic committee Ck drawn from Ctask <ref:2603.11048#pg9>. The fitness landscape on island k is shaped by two coupled elements: (1) the island-specific critic committee Ck, which defines evaluative standards, and (2) the evolving script population Sk, which determines the comparative baseline <ref:2603.11048#pg9>. Evolution proceeds through round-robin pairwise evaluation where losing scripts are refined using feedback from winners via an update operator U that integrates comparative feedback with semantic mutation <ref:2603.11048#pg10>.

Video Rendering Loop

The rendering stage translates critic-selected scripts into videos using a competition-based framework for video generation <ref:2603.11048#pg10>. This involves several steps:

  1. Script-conditioned critic generation, where a video metacritic agent prender generates critics with diverse personas conditioned on the script <ref:2603.11048#pg10>.

  2. Storyboarding, where a scene director agent generates D scene directions that specify rendering of Nj shots in text form <ref:2603.11048#pg10>.

  3. Iterative shot refinement with history tournament, where each shot is refined over Crender iterations based on feedback from script-conditioned critics, and a single-elimination tournament selects the best revision across history <ref:2603.11048#pg11>.

  4. Test-time scaling via scene-level tournament, which performs a final tournament among the full videos to select the best realization across diverse scene directions <ref:2603.11048#pg11>.

Comparison Studio

COMIC demonstrates superior performance compared to existing baselines when evaluated by human judges across multiple criteria, including Funniness, Watch More, Script Quality, Narrative Quality, Visual Realism, and Consistency <ref:2603.11048#pg24>. COMIC consistently outperforms agentic baselines by large margins across all dimensions <ref:2603.11048#pg24>. Furthermore, the automated ranking (COMIC > Sora > Veo > MA ≈ VGoT) aligns with human results, validating the benchmark as a proxy for human judgment <ref:2603.11048#pg24>. The framework also achieves the highest overall inter- and intra-diversity, demonstrating that its mechanism sustains a diverse range of comedic styles that single-pass methods do not <ref:2603.11048#pg24>.

Scale

The system exposes several scaling dimensions: number of islands K, scripts per island Sk, critics per island Ck, scene directions D, and rendering critics Crender <ref:2603.11048#pg22>. The computational complexity for the writing stage is O(G · K · S k squared · C k) evaluations per island <ref:2603.11048#pg22>. The dominant term for the rendering stage is O(D · N · Crender 2) render calls per script <ref:2603.11048#pg22>. The base configuration runs in approximately one day on a single H200 GPU with an API budget of around 5 dollars, which is orders of magnitude below the production cost of professional sketch comedy <ref:2603.11048#pg22>.

Conclusion

COMIC establishes a new state of the art for fully automated video production by leveraging a multi-island topology where diverse, human-aligned critic committees drive iterative refinement <ref:2603.11048#pg6>. The work connects to the Red Queen hypothesis in evolutionary biology, wherein species must continuously evolve to maintain their fitness against co-evolving competitors <ref:2603.11048#pg6>. Ultimately, this work establishes a new state of the art for automated, engaging, long-form video production.

References

  1. Arrow, K.J.: A difficulty in the concept of social welfare <ref:2603.11048#pg16>.

  2. Cantú-Paz, E., et al.: A survey of parallel genetic algorithms <ref:2603.11048#pg16>.

  3. Chan, C.M., Chen, W., Su, Y., Yu, J., Xue, W., Zhang, S., Fu, J., Liu, Z.: Chateval: Towards better llm-based evaluators through multi-agent debate <ref:2603.11048#pg16>.

  4. Cho, H., Ahn, D., Hong, S., Kim, J.E., Kim, S., Jin, K.H.: Tag: Tangential amplifying guidance for hallucination-resistant diffusion sampling <ref:2603.11048#pg16>.

  5. Dalal, K., Koceja, D., Xu, J., Zhao, Y., Han, S., Cheung, K.C., Kautz, J., Choi, Y., Sun, Y.: One-minute video generation with test-time training <ref:2603.11048#pg16>.

  6. Du, Y., Li, S., Torralba, A., Tenenbaum, J.B., Mordatch, I.: Improving factuality and reasoning in language models through multiagent debate <ref:2603.11048#pg16>.

  7. Fernando, C., Banarse, D., Michalewski, H., Osindero, S., Rocktäschel, T.: Promptbreeder: Self-referential self-improvement via prompt evolution <ref:2603.11048!

Improvements for AI systems

  1. Improve script generation by implementing Island k evolution to explore diverse comedic styles, as this yield[s] a Pareto frontier of diverse comedic styles. This allows for generating sketches across different tonal ranges (slapstick, dry wit, surrealism) rather than converging on a single style.

  2. Enhance humor evaluation accuracy by using Task-wise selection of critics based on pairwise comparison tasks like Top vs. Bottom, as this approach was shown to be superior to pooled or single-best critics across all channels and engagement tiers.

  3. Increase video rendering quality and consistency by implementing an Iterative shot refinement with history tournament, which uses a SingleElimination tournament across the full history of shots to ensure quality increases after each refinement step, guarding against over-refinement.

  4. Scale inference time efficiency by leveraging the multi-island topology for script writing, as this structure makes iterative refinement tractable compared to a globally pooled tournament, and allows for test-time scaling via scene-level tournament.

  5. Improve visual coherence and character fidelity by using specialized models like FLUX.2 [dev] alongside TAG [4] for generating canonical character appearances, ensuring that generated video maintains high Visual Consistency.

Abstract

We propose a fully automated AI system that produces short comedic videos similar to sketch shows such as Saturday Night Live. Starting from character references, the system employs a population of agents loosely modeled on roles in real production studios and structured to optimize the quality and diversity of ideas and outputs through iterative competition, evaluation, and refinement. A key contribution is the introduction of LLM-based critics aligned with real viewer preferences through the analysis of a corpus of comedy videos on YouTube, enabling automatic evaluation of humor. We further propose multi-island relativistic evolution for script improvement, along with script-conditioned rendering critics that perform shot- and video-level tournaments. In both human and automated evaluations, COMIC receives higher scores than agentic and video generation baselines.

Sources

Related papers