Creativity in LLM-based Multi-Agent Systems: A Survey

summary

Video file (mp4)

The gist

This survey is the first dedicated to creativity in LLM-based multi-agent systems (MAS), addressing a gap in existing surveys that focus on infrastructure but overlook creative output evaluation, the

In short

The episode surveys a paper on creativity in LLM-based multi-agent systems. Hosts discuss how collaboration across multiple specialized AI agents—using techniques like divergent exploration, iterative refinement, and collaborative synthesis—can unlock novel ideas. They conclude that creativity is about structured orchestration rather than a single model.

Key concepts

Divergent Exploration
This technique involves allowing each agent to generate many varied ideas from different perspectives without immediate judgment. It expands the idea space by letting agents run wild, similar to a brainstorming session where no idea is too silly.
Iterative Refinement
This is the feedback loop where one agent drafts, another critiques, and a third revises. This process continues until the output is polished. An example mentioned is using an Actor agent to check dialogue in a screenwriting project.
Persona Granularity
This refers to the level of detail given to an AI's personality beyond just a job title. Coarser personas yield broad ideas, while fine-grained personas with detailed backgrounds and traits lead to more predictable and focused creative output.

Terminology used across episodes

This episode discusses

The paper

Creativity in LLM-based Multi-Agent Systems: A Survey · Read on arXiv

Yi-Cheng Lin, Kang-Chieh Chen, Zhe-Yan Li, Tzu-Heng Wu, Tzu-Hsuan Wu, Kuan-Yu Chen, Hung-yi Lee, Yun-Nung Chen

National Taiwan University

Large language model (LLM)-driven multi-agent systems (MAS) are transforming how humans and AIs collaboratively generate ideas and artifacts. While existing surveys provide comprehensive overviews of MAS infrastructures, they largely overlook the dimension of creativity, including how novel outputs are generated and evaluated, how creativity informs agent personas, and how creative workflows are coordinated. This is the first survey dedicated to creativity in MAS. We focus on text and image generation tasks, and present: (1) a taxonomy of agent proactivity and persona design; (2) an overview of generation techniques, including divergent exploration, iterative refinement, and collaborative synthesis, as well as relevant datasets and evaluation metrics; and (3) a discussion of key challenges, such as inconsistent evaluation standards, insufficient bias mitigation, coordination conflicts, and the lack of unified benchmarks. This survey offers a structured framework and roadmap for advancing the development, evaluation, and standardization of creative MAS.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Creativity in LLM-based Multi-Agent Systems: A Survey".

Jane: The paper was written by Yi-Cheng Lin, Kang-Chieh Chen, Zhe-Yan Li, Tzu-Heng Wu, Tzu-Hsuan Wu et al. from National Taiwan University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the channel, everyone. We are diving into a brand new paper today, and it's called "Creativity in LLM-based Multi-Agent Systems: A Survey." I'm Tom, and as always, I'm here with the brilliant Jane. Jane, this title is a mouthful, but it's basically about getting multiple AI agents to work together to come up with new ideas, right?

Jane: Exactly, Tom. And I'm so excited to dig into this one. For years, we've been asking a single AI to write a story or design a logo. This paper asks a different question: what happens when you have a whole team of AIs, each with their own personality and job, brainstorming together?

Tom: So it's like moving from a solo artist to a full band. The paper is from a team at National Taiwan University, and they're making a big claim. They say this is the first survey dedicated to creativity in these multi-agent systems. No one has mapped this whole landscape out before.

Jane: That's a huge gap to fill. The authors point out that previous surveys looked at the plumbing of these systems—the architecture, the communication protocols, the infrastructure. But nobody was asking the fun question: are these systems actually producing anything novel and valuable?

Tom: Right, they were checking if the engine runs, but not if the car can win the race. This survey focuses on text and image generation, which is where a lot of the creative action is happening right now.

Jane: And the implications are wild. If we can understand how to structure these agent teams, we could unlock creativity that goes beyond what any single AI, or even a single human, can do on their own. The paper suggests that collaboration itself is a creative act.

Tom: I love that framing. They’re not just looking at the final product; they’re looking at the process. How do agents plan, execute, and decide? And how much control should they have versus the human user? That’s the core tension we’re going to explore.

Jane: It really is. And the authors break that down into a spectrum of "proactivity." On one end, you have agents that just wait for instructions. On the other, you have agents that are setting their own goals and critiquing each other’s work without any human input.

Tom: That spectrum is going to be key. So, stick around. We're going to unpack how these systems actually generate creative work, the techniques they use, and the big challenges that are still out there.

Jane: And we're going to make it all make sense. Let's get into the summary of the paper next.

Summary: Tom: So, Jane, we've set the stage. Now let's talk about what this survey actually found. The paper, "Creativity in LLM-based Multi-Agent Systems: A Survey," doesn't just say "more agents equals more creativity." It identifies three core techniques that make these systems work.

Jane: Right, and they’re like the three pillars of a creative team. The first is "Divergent Exploration." That’s when you let each agent run wild with a different perspective, generating as many varied ideas as possible before anyone judges them.

Tom: Think of it as a brainstorming session where no idea is too silly. The paper cites a study called Co-GPT Ideation where people working with an LLM generated more diverse and detailed ideas than people working alone. The AI expands the idea space.

Jane: But you can't just brainstorm forever. That's where the second technique comes in: "Iterative Refinement." This is the feedback loop. You have one agent draft something, another agent critiques it, and a third agent revises it. Round and round until it's polished.

Tom: The paper uses a great example called HoLLMwood for screenwriting. You have a Writer, an Editor, and an Actor. The Actor role-plays the character to check if the dialogue sounds right. That back-and-forth leads to richer stories than a single AI could write in one go.

Jane: And the third pillar is "Collaborative Synthesis." This is where you take all those diverse ideas and the refined pieces, and you merge them into one coherent whole. It’s not just about combining; it’s about integrating different perspectives into something unified.

Tom: Like a band where the drummer and the guitarist are playing different parts, but together they make a song. The paper mentions CollabStory, where multiple LLMs take turns writing paragraphs of a story, and they manage to keep it coherent.

Jane: So the summary is that creativity isn't a single spark. It's a structured process. The survey is saying that by dividing the cognitive workload—ideation, evaluation, coordination—across specialized agents, you get better results.

Tom: And they back this up with a whole taxonomy of techniques, datasets, and evaluation metrics. It’s a toolkit for anyone who wants to build one of these systems.

Jane: But here's the thing that really got me. The paper also talks about "persona." You can't just give an agent a job title. You have to give it a personality, a background, maybe even a writing style. The granularity of that persona—from a simple label like "marketing strategist" to a full biography—changes how creative the agent is.

Tom: That's fascinating. A coarse persona gives you broad, spontaneous ideas. A fine-grained persona, with a detailed career path and personality traits, gives you more predictable and focused output. It’s a trade-off between control and surprise.

Jane: And that trade-off is central to the whole paper. It’s not about building the most autonomous system. It’s about finding the right balance for the task at hand.

Tom: Okay, so we've got the techniques and the personas. But how do we actually know if any of this is working? That's the million-dollar question, and it's what we're going to tackle in the next segment.

Improvements: Tom: Welcome back. We've talked about how these multi-agent systems work, but now we have to ask the hard question: how do we know if they're actually creative? The paper, "Creativity in LLM-based Multi-Agent Systems: A Survey," spends a lot of time on this, and it’s a mess.

Jane: It really is a mess, Tom. There's no single "creativity meter." The survey lays out two main approaches. First, you have objective, metric-based measures. Things like counting unique words or measuring the distance between ideas in a mathematical space.

Tom: So, for text, you might use something like "Distinct-n" to see how many unique phrases the AI is using. Or for images, you'd use something like FID, which compares the generated images to real ones to see if they're diverse and high quality.

Jane: Those are great because they're scalable and reproducible. You can run them on a thousand outputs automatically. But they miss the soul of creativity. They don't capture whether a story is emotionally resonant or whether a design is truly surprising.

Tom: Exactly. That's why the paper also discusses subjective assessments. These are the human evaluations. You bring in experts or crowds of people to rate the output on things like originality, fluency, and elaboration. That's the Torrance Tests of Creative Thinking, or TTCT, applied to AI.

Jane: And now, there's a new player in the game: LLM-as-a-judge. You have one AI system rate the creativity of another AI system's output. It's fast and cheap, but the survey is careful to point out that these judges have their own biases.

Lu: If I can jump in here, Jane. This is where the survey gets really interesting for me. The inconsistency in evaluation is the biggest barrier to progress. If every paper uses a different rubric, you can't compare results. You can't tell if your new system is actually better than the last one.

Meng: And from a practical standpoint, that's a nightmare. I'm trying to build a product that helps designers generate concepts. If I can't measure whether my agent team is more creative than a single prompt, I can't justify the cost of running ten agents in parallel. The computational overhead is real.

Jane: That's a great point, Meng. The paper actually calls out "Resource-Efficient Orchestration" as a major challenge. You can't just spawn a hundred agents and hope for the best. You need to be smart about which agents are active and when.

Tom: So the improvements the paper suggests are about standardization. They want a unified benchmark, like a common set of tasks and scoring rubrics, so that everyone is playing the same game. They mention MultiAgentBench as a first step in that direction.

Lu: And beyond the benchmark, they want to move from static evaluations to real-time, interactive ones. Imagine a system that adapts its creativity based on your live feedback. That's the future they're pointing toward.

Meng: But that requires solving the trust problem too. The survey shows that when agents are too proactive, users feel like they've lost control. The system is generating ideas, but the user doesn't feel like it's their work anymore.

Jane: That's the balancing act. The paper suggests we need systems that can calibrate their proactivity in real-time, based on the task and the user's comfort level. That's a huge engineering challenge, but it's the key to making these tools actually useful.

Tom: Alright, so we've got the evaluation problem, the resource problem, and the trust problem. That's a lot to chew on. Let's wrap this up and see what the big takeaway is for the world.

Conclusion: Tom: We've covered a lot of ground on "Creativity in LLM-based Multi-Agent Systems: A Survey." Let's bring it all together. Jane, what's the one thing you want our listeners to remember?

Jane: I think it's that creativity in AI isn't about a single, magical model. It's about orchestration. This survey shows that by structuring collaboration—by giving agents distinct roles, personas, and feedback loops—we can unlock results that are more novel and more valuable than anything a single AI could produce.

Tom: And it's not just about the AI. The paper is really about human-AI collaboration. It's about designing systems that know when to step up and when to step back, to make the human partner more creative, not less.

Lu: I'd add that this survey is a roadmap. It identifies the three core techniques—divergent exploration, iterative refinement, and collaborative synthesis—and it gives us the vocabulary to talk about them. That's essential for the field to move forward.

Meng: And for the engineers out there, the message is clear: we need to build for evaluation and efficiency. We can't just throw compute at the problem. We need adaptive systems that know when to stop, when to refine, and when to let the human take the wheel.

Jane: The paper also warns us about the risks. Bias can creep in through the personas we design. Conflicts between agents can be destructive if not managed. And there's a big open question about authorship. If an AI team writes a novel, who owns it?

Tom: That's the deep stuff. The legal and ethical frameworks haven't caught up with the technology. But that's what makes this field so exciting. We're not just building tools; we're reshaping what it means to create.

Lalam: If I may add, the most impactful vision here is cultural. These systems can democratize creativity. They can help a child in a remote village write a story with the same resources as a professional in a big city. But only if we build them to be inclusive, to respect diverse voices, and to augment human imagination rather than replace it.

Tom: Beautifully said, Lalam. So, we're saying goodbye to this paper, but we're taking its framework with us. It's given us a lens to see the future of collaborative AI.

Jane: And we're excited to see what comes next. There are so many open challenges, from unified benchmarks to longitudinal studies on how users grow with these systems. The work is just beginning.

Tom: Thanks for joining us on the channel. We've been discussing "Creativity in LLM-based Multi-Agent Systems: A Survey." Until next time, keep questioning, keep exploring, and keep creating. See you soon.

More episodes

← Home