Scaling Participation in Modular AI Systems

arXiv:2606.07812 · cs.AI, cs.CL · Submitted 2026-06-05 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Scaling Participation in Modular AI Systems".

Jane: Humanity is a mosaic of multifaceted talents and needs, and any truly intelligent AI must reflect that richness.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: We've seen how they set up these systems by gathering models from different researchers and then connecting them, and now we look at what they claim are the actual improvements in performance and capability of this approach with "Scaling Participation in Modular AI Systems."

Jane: They show that scaling participation leads to a consistent boost across the board, with an average improvement of nearly twenty-nine percent when you go from just two models to thirty-two <ref:two thousand six hundred six point zero seven eight one two#pg4. This isn't just random luck; it’s tied directly to the diversity of the contributions.

Lu: What’s interesting is that they found this improvement isn't just about adding more raw compute power, but because those models bring different strengths to the table <ref:two thousand six hundred six point zero seven eight one two#pg4. They tested controlled experiments and showed that performance jumps significantly when you increase the diversity of the input models <ref:two thousand six hundred six point zero seven eight one two#pg4.

Meng: So, if we think about this practically for a company, what does that mean for our product development pipeline? Does it make our QA process more reliable?

Tom: It suggests that when you scale up the collaboration, you get systems that are much better at solving problems where any single model would fail <ref:two thousand six hundred six point zero seven eight one two#pg4. This collaborative emergence means the system can tackle much harder tasks than one giant model could manage alone.

Jane: They also looked specifically at things related to human diversity, and those improvements went up by over twenty percent when scaling participation from two to thirty-two models <ref:two thousand six hundred six point zero seven eight one two#pg4. This points toward a future where AI reflects the actual variety of human experiences better than current systems do.

Lalam: For culture, this means the AI won't just parrot one dominant viewpoint; it will be able to represent a much broader spectrum of ideas and values because it’s built from those varied inputs <ref:two thousand six hundred six point zero seven eight one two#pg5. That pluralistic nature is really something I think will make the output richer.

Tom: They also made a specific point about how individual strengths and compositional strengths aren't always correlated, with p-values consistently above zero point zero five <ref:two thousand six hundred six point zero seven eight one two#pg5. That’s a crucial caveat for anyone trying to design these systems.

Lu: It means you can't just pick the best individual models and hope they fit; you have to focus on designing the collaboration layer itself, because that’s where the real intelligence comes from <ref:two thousand six hundred six point zero seven eight one two#pg5.

Jane: So, it’s not just about gathering good models; it’s about engineering a system that knows how to effectively combine those different strengths to get a better result <ref:two thousand six hundred six point zero seven eight one two#pg4.

Meng: I wonder what the next step is for building those collaboration algorithms, since they tested so many different methods—API routing, text exchange, weight manipulation <ref:two thousand six hundred six point zero seven eight one two#pg3. Do we need a standardized way to do that?

Tom: That's the practical question right there. The paper lays out all these collaboration tools, and the next phase is figuring out which ones are actually the most effective for different types of problems <ref:two thousand six hundred six point zero seven eight one two#pg3.

Lalam: I hope that this direction helps cultivate an AI culture where diversity isn't just a buzzword, but actually a core part of the intelligence itself <ref:two thousand six hundred six point zero seven eight one two#pg5. It’s about building something that genuinely understands complexity.

The paper's summary: Tom: We’re wrapping up this look at "Scaling Participation in Modular AI Systems," which really shows us that building AI modularly, from the bottom up through diverse contributions, gives you a solid technical foundation <ref:two thousand six hundred six point zero seven eight one two#pg0.

Jane: It’s clear that this approach outperforms monolithic models by about fifteen percent on tasks like reasoning and factuality across many different areas <ref:two thousand six hundred six point zero seven eight one two#pg0. That kind of consistent jump across so many domains is pretty significant.

Lu: The main implication for me is that we’re looking at a way to build systems that are inherently more transparent and easier to update because they aren't tied up in one massive architecture <ref:two thousand six hundred six point zero seven eight one two#pg5. That structural difference is what makes the long-term vision appealing.

Meng: From an engineering standpoint, it means we can build systems that are cheaper to reuse and easier to maintain because you’re not trying to overhaul one gigantic black box <ref:two thousand six hundred six point zero seven eight one two#pg5. That reduces the deployment risk for real-world applications.

Lalam: I feel like this opens up a path where AI can be more pluralistic, reflecting different human priorities and values because it’s assembled by many hands <ref:two thousand six hundred six point zero seven eight one two#pg5. It suggests a culture where the AI isn't just one voice but a chorus.

Tom: It definitely moves us away from that monolithic status quo and toward something open and collaborative <ref:two thousand six hundred six point zero seven eight one two#pg1. This is about making sure we are building things that can be collectively owned by the community who contributed to them.

Jane: We’re talking about a future where AI is better at showing you where its answers come from and who is responsible for what it does <ref:two thousand six hundred six point zero seven eight one two#pg5. That level of accountability is something we need to push for in how these systems are designed.

Lu: The real challenge now, I think, is figuring out the next big hurdle for those collaboration algorithms—how to design the best way for these disparate models to actually communicate and work together efficiently <ref:two thousand six hundred six point zero seven eight one two#pg3.

Meng: Yeah, that’s where the engineering work gets really interesting. The paper gives us a lot of tools, but implementing those weight-level collaborations or the text-level refinement methods needs serious optimization before we see this at scale <ref:two thousand six hundred six point zero seven eight one two#pg3.

Lalam: I hope that this direction helps cultivate an AI culture where diversity isn't just a buzzword, but actually a core part of the intelligence itself <ref:two thousand six hundred six point zero seven eight one two#pg5. It’s about building something that genuinely understands complexity.

Tom: So, to recap on "Scaling Participation in Modular AI Systems," we have this new paradigm where intelligence comes from composed parts rather than one huge model, giving us up to fifteen percent better performance and a way to build things that are more transparent and pluralistic <ref:two thousand six hundred six point zero seven eight one two#pg0.

Jane: It’s a strong argument for moving toward open, bottom-up development where the system is built by different stakeholders <ref:two thousand six hundred six point zero seven eight one two#pg1.

Lu: The next big thing we need to watch is how those specific model collaboration algorithms evolve and what the practical limits of that diversity scaling are <ref:two thousand six hundred six point zero seven eight one two#pg3.

Meng: We’ll keep an eye on how these modular systems get deployed in real-world scenarios because that’s where we’ll see if this architecture actually delivers on the promise of being cheaper and easier to update <ref:two thousand six hundred six point zero seven eight one two#pg5.

Lalam: I hope this paper inspires a culture where building AI is seen as a collective effort, not just a solo race to build the biggest thing.

The paper's improvements: Tom: So we’ve talked about "Scaling Participation in Modular AI Systems," which shows how building AI from diverse contributions, from the bottom up, gives you a solid technical foundation <ref:two thousand six hundred six point zero seven eight one two#pg0.

Jane: It’s clear that this approach outperforms monolithic models by about fifteen percent on tasks like reasoning and factuality across many different areas <ref:two thousand six hundred six point zero seven eight one two#pg0. That kind of consistent jump across so many domains is pretty significant.

Lu: The main implication for me is that we’re looking at a way to build systems that are inherently more transparent and easier to update because they aren't tied up in one massive architecture <ref:two thousand six hundred six point zero seven eight one two#pg5. That structural difference is what makes the long-term vision appealing.

Meng: From an engineering standpoint, it means we can build systems that are cheaper to reuse and easier to maintain because you’re not trying to overhaul one gigantic black box <ref:two thousand six hundred six point zero seven eight one two#pg5. That reduces the deployment risk for real-world applications.

Lalam: I feel like this opens up a path where AI can be more pluralistic, reflecting different human priorities and values because it’s assembled by many hands <ref:two thousand six hundred six point zero seven eight one two#pg5. It suggests a culture where the AI isn't just one voice but a chorus.

Tom: It definitely moves us away from that monolithic status quo and toward something open and collaborative <ref:two thousand six hundred six point zero seven eight one two#pg1. This is about making sure we are building things that can be collectively owned by the community who contributed to them.

Jane: We’re talking about a future where AI is better at showing you where its answers come from and who is responsible for what it does <ref:two thousand six hundred six point zero seven eight one two#pg5. That level of accountability is something we need to push for in how these systems are designed.

Lu: The real challenge now, I think, is figuring out the next big hurdle for those collaboration algorithms—how to design the best way for those disparate models to actually communicate and work together efficiently <ref:two thousand six hundred six point zero seven eight one two#pg3.

Meng: Yeah, that’s where the engineering work gets really interesting. The paper gives us a lot of tools, but implementing those weight-level collaborations or the text-level refinement methods needs serious optimization before we see this at scale <ref:two thousand six hundred six point zero seven eight one two#pg3.

Lalam: I hope that this direction helps cultivate an AI culture where diversity isn't just a buzzword, but actually a core part of the intelligence itself <ref:two thousand six hundred six point zero seven eight one two#pg5. It’s about building something that genuinely understands complexity.

Tom: So, to recap on "Scaling Participation in Modular AI Systems," we have this new paradigm where intelligence comes from composed parts rather than one huge model, giving us up to fifteen percent better performance and a way to build things that are more transparent and pluralistic <ref:two thousand six hundred six point zero seven eight one two#pg0.

Jane: It’s a strong argument for moving toward open, bottom-up development where the system is built by different stakeholders <ref:two thousand six hundred six point zero seven eight one two#pg1.

Lu: The next big thing we need to watch is how those specific model collaboration algorithms evolve and what the practical limits of that diversity scaling are <ref:two thousand six hundred six point zero seven eight one two#pg3.

Meng: We’ll keep an eye on how these modular systems get deployed in real-world scenarios because that’s where we’ll see if this architecture actually delivers on the promise of being cheaper and easier to update <ref:two thousand six hundred six point zero seven eight one two#pg5.

Lalam: I hope this paper inspires a culture where building AI is seen as a collective effort, not just a solo race to build the biggest thing.

Conclusion: Tom: So we’ve covered "Scaling Participation in Modular AI Systems," which shows how building AI from diverse contributions, from the bottom up, gives you a solid technical foundation <ref:two thousand six hundred six point zero seven eight one two#pg0.

Jane: It’s clear that this approach outperforms monolithic models by about fifteen percent on tasks like reasoning and factuality across many different areas <ref:two thousand six hundred six point zero seven eight one two#pg0. That kind of consistent jump across so many domains is pretty significant.

Lu: The main implication for me is that we’re looking at a way to build systems that are inherently more transparent and easier to update because they aren't tied up in one massive architecture <ref:two thousand six hundred six point zero seven eight one two#pg5. That structural difference is what makes the long-term vision appealing.

Meng: From an engineering standpoint, it means we can build systems that are cheaper to reuse and easier to maintain because you’re not trying to overhaul one gigantic black box <ref:two thousand six hundred six point zero seven eight one two#pg5. That reduces the deployment risk for real-world applications.

Lalam: I feel like this opens up a path where AI can be more pluralistic, reflecting different human priorities and values because it’s assembled by many hands <ref:two thousand six hundred six point zero seven eight one two#pg5. It suggests a culture where the AI isn't just one voice but a chorus.

Tom: It definitely moves us away from that monolithic status quo and toward something open and collaborative <ref:two thousand six hundred six point zero seven eight one two#pg1. This is about making sure we are building things that can be collectively owned by the community who contributed to them.

Jane: We’re talking about a future where AI is better at showing you where its answers come from and who is responsible for what it does <ref:two thousand six hundred six point zero seven eight one two#pg5. That level of accountability is something we need to push for in how these systems are designed.

Lu: The real challenge now, I think, is figuring out the next big hurdle for those collaboration algorithms—how to design the best way for those disparate models to actually communicate and work together efficiently <ref:two thousand six hundred six point zero seven eight one two#pg3.

Meng: Yeah, that’s where the engineering work gets really interesting. The paper gives us a lot of tools, but implementing those weight-level collaborations or the text-level refinement methods needs serious optimization before we see this at scale <ref:two thousand six hundred six point zero seven eight one two#pg3.

Lalam: I hope that this direction helps cultivate an AI culture where diversity isn't just a buzzword, but actually a core part of the intelligence itself <ref:two thousand six hundred six point zero seven eight one two#pg5. It’s about building something that genuinely understands complexity.

Tom: So, to recap on "Scaling Participation in Modular AI Systems," we have this new paradigm where intelligence comes from composed parts rather than one huge model, giving us up to fifteen percent better performance and a way to build things that are more transparent and pluralistic <ref:two thousand six hundred six point zero seven eight one two#pg0.

Jane: It’s a strong argument for moving toward open, bottom-up development where the system is built by different stakeholders <ref:two thousand six hundred six point zero seven eight one two#pg1.

Lu: The next big thing we need to watch is how those specific model collaboration algorithms evolve and what the practical limits of that diversity scaling are <ref:two thousand six hundred six point zero seven eight one two#pg3.

Meng: We’ll keep an eye on how these modular systems get deployed in real-world scenarios because that’s where we’ll see if this architecture actually delivers on the promise of being cheaper and easier to update <ref:two thousand six hundred six point zero seven eight one two#pg5.

Lalam: I hope this paper inspires a culture where building AI is seen as a collective effort, not just a solo race to build the biggest thing.

Tom: We’ve got that for today on "Scaling Participation in Modular AI Systems." Next up, we’re looking at how different techniques are tackling video fusion with something called MambaVF.

Shangbin Feng, Yike Wang, Weijia Shi, Luke Zettlemoyer

University of Washington

cs.AI, cs.CL

Submitted: 2026-06-05

Updated: 2026-10-05

Code: https://github.com/BunsenFeng/model

Project page: https://yikee.github.io/open-model-collaboration

Importance score: 81/100

The gist: Humanity is a mosaic of multifaceted talents and needs, and any truly intelligent AI must reflect that richness.

Key concepts

Scaling Participation
This is a new paradigm where modular AI systems are constructed from the bottom up by involving many diverse stakeholders. Instead of one massive model, intelligence emerges from how these smaller, specialized models interact and collaborate.
Orchestrating Participation
This involves building compositional AI systems using the contributed models. This includes various collaboration algorithms across three levels: API-level (routing), text-level (like multi-agent refinement), and weight-level (modifying model parameters).
Collaborative Emergence
This occurs when a system solves complex problems that no single individual model could solve alone. This generalization happens by scaling the collaboration and representation of diverse models, proving that collective intelligence is greater than the sum of its parts.

Terminology

Summary

Humanity is a mosaic of multifaceted talents and needs, and any truly intelligent AI must reflect that richness. This paper introduces scaling participation, a new paradigm in which modular AI systems are built from the bottom up through the contributions of diverse stakeholders, which outperforms monolithic LLMs by up to 15.4% across 15 tasks such as reasoning and factuality

How it works

The core idea is to build modular, participatory AI systems where intelligence is composed from interacting components rather than compressed into one model and both the component models and the mechanisms governing their collaboration can be contributed by a broad range of stakeholders. This approach moves away from monolithic AI models developed by a few private firms toward an open, bottom-up, and collaborative future.

The process is operationalized in two main steps:

  1. Soliciting Participation: This involves reaching out to academic labs working on language models and asking for their participation by submitting the language models they trained in their research. The authors curated a pool of 61 models from 236 researchers around the world.

  2. Orchestrating Participation: This involves building bottom-up and compositional AI systems with these models and diverse model collaboration algorithms. These 32 contributed LMs collaborate, compose, and complement each other in 14 collaborative systems such as routing [11, 12], multi-agent debate [13, 14], and model fusion [15, 16].

Orchestrating Participation: Model Collaboration Algorithms

The paper employs a wide spectrum of model collaboration algorithms across three levels of information exchange to prototype participatory AI. These include API-level collaboration, text-level collaboration, and weight-level collaboration.

API-level collaboration features routing and selection among candidate models through methods such as Prompt Routing, Trained Router (using a causal language model to predict the best model), Graph Routing (employing a graph neural network for prediction), and Switch Generation (training a switcher model to govern how multiple candidate models take turns generating text patches).

Text-level collaboration involves exchanging generated texts across models, including Multi-Agent Refine (where each model refines its response based on others), Multi-Agent Finetuning (a training-based enhancement where both generation and critique datasets are iteratively generated), LLM Blender (training a ranker LLM to rank responses and a fuser LLM to aggregate the top-k responses), AggLM (an aggregator LLM trained with reinforcement learning with verifiable rewards for final response selection), Sparta Alignment, Heterogeneous Swarms, Dare-Ties, Greedy Soup, Weight Extrapolation (ExPO), and Model Swarms.

Weight-level collaboration involves arithmetic and search with model parameters through methods like Dare-Ties (a combination of pruning and sign consensus), Greedy Soup (iteratively adding models to a soup, averaging weights), Weight Extrapolation (extrapolating parameters into a new model using the top-k and bottom-k models), Model Swarms (collaborative search in parameter space via particle swarm optimization), Dare-Ties, and Greedy Soup.

Experiment Settings and Evaluation

The evaluation is conducted across 5 domains and 15 tasks in total: general QA (AGIEval [26], ARC-challenge [27], MMLU-redux [28]), reasoning (BigBench-Hard [29], GSM8k [30], MATH [31], Sciriff [32]), knowledge (WikiDYK [33], PopQA [34], BLEND [35]), safety (TruthfulQA [36], CocoNot [37]), and instruction following (AlpacaEval [38], Wildchat [39], Human Interest).

The performance is reported by scaling the number of contributed models from 2 to 32, and by evaluating different collaboration algorithms. The best setting in bold shows the superior performance of the participatory AI systems.

Analysis of Diversity and Emergence

Scaling participation via scaling model collaboration led to consistent improvements with larger and more diverse AI systems (Figure 3). The average improvement from 2 to 32 models was 28.78%.

Controlled experiments showed that the improvements are due to the increasing diversity and complementary strengths of contributed models, not merely compute, as an average improvement of 13.8% from the least diverse to most diverse settings indicates this. Furthermore, scaling participation continues to yield promising trends on tasks about human diversity, improving from 2 to 32 models by up to 21.58% compared to general-utility tasks.

Collaborative emergence is demonstrated where systems solve problems where individual models struggle, and this emergence scales with participation, indicating that generalization doesn’t need to come from excessive scaling of a single model, but from scaling the collaboration and representation of diverse contributed models.

The work suggests that evaluating and training AI models for compositional strength is an independent and critical research question, as individual and compositional strengths are mostly not correlated (with p-values consistently larger than 0.05).

Conclusion

Scaling participation in modular AI systems provides a technical foundation for transitioning from the monolithic status quo toward an open, bottom-up, and collaborative AI future. Participatory AI outperforms non-modular and/or non-participatory LLMs by up to 15.4% on average across evaluation domains. This paradigm creates systems that are more transparent, more pluralistic, easier to update, cheaper to reuse, and better able to show where outputs come from and who is responsible for them.

REFERENCES

[1] Ingrams, A., Kaufmann, W., Jacobs, D.: In ai we trust? citizen perceptions of ai in government decision making. Policy & Internet 14(2), 390–409 (2022)

[2] Educational Technology, U.S.D.o.E.: Artificial Intelligence and the Future of Teaching and Learning: Insights and Recommendations. Accessed: 2026-04-08 (2023). https://www.ed.gov/sites/ed/files/documents/ai-report/ai-report.pdf

[3] Chinta, S.V., et al.: Ai-driven healthcare: A review on ensuring fairness and mitigating bias. PLOS Digital Health 4(5), 0000864 (2025)

[4] Bughin, J., Seong, J., Manyika, J., Chui, M., Joshi, R.: Notes from the ai frontier: Modeling the impact of ai on the world economy. McKinsey Global Institute 4(1), 2–61 (2018)

[5] Asai, A., et al.: Synthesizing scientific literature with retrieval-augmented language models. Nature, 1–7 (2026)

[6] Lu, C., et al.: Towards end-to-end automation of ai research. Nature 651(8107), 914–919 (2026)

[7] Anthropic: Statement from Dario Amodei on our discussions with the Department of War. Accessed: 2026-04-01 (2026). https://www.anthropic.com/news/statement-department-of-war

[8] Sorensen, T., et al.

Improvements for AI systems

  1. Bold Header: Scale Participation via Model Collaboration

This improvement involves building modular, participatory AI systems from the bottom up through the contributions of diverse stakeholders, allowing participants to contribute small models trained on their own interests and priorities. This leads to systems that outperform monolithic LLMs by up to 15.4% across 15 tasks, such as reasoning and factuality.

  1. Bold Header: Enhance Human Diversity Representation

The system will be explicitly designed to reflect human diversity because as more models contributed by diverse stakeholders in AI join the system, the values, preferences, and priorities of them will be better reflected than monolithic models. This is achieved by evaluating systems on tasks like CultureBench [42] and Value Kaleidoscope [43], where performance improves over large monolithic models by up to 21.58%.

  1. Bold Header: Enable Collaborative Emergence

The system can solve problems that individual components cannot, as collaborative AI systems of multiple models could solve problems where individual models struggle, i.e., collaborative emergence. This generalization occurs through scaling the collaboration and representation of diverse contributed models, which is shown to scale with participation in solving problems where all participating models could not individually solve.

  1. Bold Header: Foster Participant Benefit

The system will ensure that contributions are valuable to stakeholders, as the compositional systems they took part in developing will perform better on their original priorities that motivated them to develop their models, thereby advancing their objective at the same time. This is demonstrated by evaluations showing that participatory and compositional AI systems advance on objectives that participants were originally pursuing.

  1. Bold Header: Implement Transparent Governance

The system promotes collective decision-making by ensuring No one will have unilateral control over participatory AI systems: since diverse stakeholders all contributed to the system, they also collectively own this artifact. This allows the community of stakeholders to jointly decide on important matters about this AI system, such as how to deploy it, which use cases are allowed and not, de-risking high-stakes applications.

Sources

Related papers