Token Distribution versus Data Volume: Domain Balancing in Multi-Domain Meeting Summarisation

arXiv:2608.15935 · cs.CL · Submitted 2026-08-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Token Distribution versus Data Volume".

Jane: Jointly fine-tuning an LLM on meeting summarisation corpora of widely varying size raises a question that prior work leaves confounded: when a domain-balanced training mixture helps,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So wrapping up this discussion on "Token Distribution versus Data Volume: Domain Balancing in Multi-Domain Meeting Summarisation," the main point is that when you're training an AI on meetings from different sources, simply throwing more data at everything doesn't always give you a uniform improvement across all of those meeting types Tom. Jane That we need to be thoughtful about *how* we mix that data based on whether the minority domains are important to us or not.

Lu: They conclude that the finding is that balancing redistributes quality, specifically raising minority domains while lowering majority domains, rather than just adding it uniformly across all of them Lu.

Meng: That means our allocation strategy shouldn't be one-size-fits-all; we should tailor the token distribution scheme to what we actually need the AI to perform well on Meng.

Lalam: I think this has big implications for how we deploy LLMs in varied environments; it suggests that strategic data mixing is a much more powerful lever than just raw volume alone Lalam.

Tom: Precisely. They also found that balancing by token count, rather than by the number of examples, gives domains different weights because equal meeting counts result in very unequal token amounts across the corpora Tom. Jane That distinction really matters for practical application because it tells us whether we're trying to match quality on a few key areas or just getting a general lift everywhere.

Lu: The paper highlights that practitioners should decide on allocation based on whether the minority domains matter, noting that their share under proportional allocation is fixed at one to two percent regardless of the total budget Lu.

Meng: So, for us building systems, if we know which domains are critical for our service, we can apply this token-level balancing to ensure those specific areas get the necessary attention during training Meng.

Lalam: It’s a really constructive conclusion because it gives us concrete guidance on how to approach imbalanced multi-domain mixtures in the future Lalam.

Conclusion: Tom: So we've seen how they meticulously crafted these experiments comparing different ways to balance data across multiple meeting types, and now we need to talk about what all this means in plain English.

Jane: Right, Tom, so the paper is really digging into this question of whether just having a massive pile of data helps or if the way that data is spread out actually makes a difference when you're summarizing different kinds of meetings.

Lu: I think the core insight they’re pushing is that for this specific task, it’s not just about quantity; it’s about the distribution itself—how tokens are allocated across those distinct domains.

Meng: From my side, what I find most interesting is how they separated balancing by token count versus balancing by example count; that distinction matters when you're trying to figure out the best way to build a system in reality.

Lalam: I’m seeing a potential shift here where the focus moves from just building bigger models toward building smarter, more thoughtful data mixtures for specific application needs.

Tom: Exactly, and looking at the title again—"Token Distribution versus Data Volume"—it really boils down to deciding whether you're maximizing volume or optimizing structure.

Jane: And that structure is what they show is crucial; it explains why some meeting types benefit from a specific kind of data mix while others might get drowned out if you just dump everything in equally.

Lu: It opens up so many creative possibilities for how we design these training pipelines, imagining bespoke mixtures tailored precisely to the nuances of different organizational communication styles.

Meng: I wonder how this translates practically into deploying an AI summarizer across a whole enterprise where you have all these varied meeting formats constantly coming in.

Lalam: If we can use this framework to intelligently guide data selection, it could fundamentally improve how we train models to understand and synthesize complex, diverse human interactions.

Tom: It sounds like the paper isn't just reporting findings; it’s giving us a blueprint for making smarter choices when preparing training sets for any multi-domain AI project.

Jane: That’s right, and understanding this level of detail helps everyone move past the general idea that more data is always better when dealing with varied sources.

Lu: The authors' conclusion suggests that the unit of balancing—whether it's tokens or examples—is just as important as the act of balancing itself.

Meng: So we need to think about which unit makes sense for our specific engineering constraints and domain priorities before we start allocating data.

Lalam: And that strategic decision-making process, driven by these results, is where the real long-term cultural impact lies for how AI tools are developed and trusted across different sectors.

Ashima Sood, Bryan Gardiner, Joan Condell

School of Computing, Engineering and Intelligent Systems, Ulster University

cs.CL

Submitted: 2026-08-16

Updated: 2026-09-28

Importance score: 84/100

The gist: Jointly fine-tuning an LLM on meeting summarisation corpora of widely varying size raises a question that prior work leaves confounded: when a domain-balanced training mixture helps, is the gain due

Key concepts

Token Mixture
This refers to the specific combination of text samples used to train a model. The researchers created mixtures where the total number of tokens was kept constant, but how those tokens were distributed across five different meeting corpora varied depending on whether they were balanced by token count or example count.
Balanced Allocation
This method distributes the fixed token budget equally among all five meeting corpora. It ensures every domain gets an equal share of training data in terms of the number of tokens, regardless of that domain's original size. This is one way to attempt to achieve domain balance.
Natural (Proportional) Allocation
This method allocates tokens based on each corpus's native token mass, meaning larger corpora receive a proportionally larger share of the budget. It reflects the actual volume of data available in each domain, serving as a volume-controlled counterpart to balanced allocation.
Balancing by Tokens vs. Examples
This distinction addresses whether balancing is done based on the total number of tokens or the number of individual examples (meeting transcripts). The study found that balancing by tokens weights domains differently than balancing by examples because equal example counts result in very unequal token amounts due to transcript length differences.

Terminology

Summary

Jointly fine-tuning an LLM on meeting summarisation corpora of widely varying size raises a question that prior work leaves confounded: when a domain-balanced training mixture helps, is the gain due to the distribution of tokens across domains, or merely to the volume of data seen? This paper disentangles these factors by constructing balanced and natural (nativeproportional) token mixtures at matched token budgets over five English meeting corpora, fine-tuning Mistral-7B with QLoRA, and evaluating per domain.

How it works

The study is structured around five research questions to isolate the effects of token distribution from data volume. The core methodology involves constructing training mixtures by allocating a fixed token budget B ∈ 2, 4, 8, 16, or 32M tokens across five English meeting corpora (AMI, ICSI, MeetingBank (MB), ELITR Minuting Corpus (ELITR), and EuroParlMin). The three allocation schemes tested are:

  1. Balanced (equal-token) allocation: Each corpus receives an equal share of the budget.

  2. Natural (proportional) allocation: Each corpus receives a share proportional to its own native token mass, which is the volume-controlled counterpart to balanced.

  3. Example-level (equal-count) allocation: This scheme balances example count rather than tokens, as equal example counts produce very unequal token amounts due to transcript length variations.

Key Experimental Setup and Model Selection

To control for confounding variables, the researchers held data volume fixed by matching balanced and natural distributions at matched token budgets. The model fine-tuned was Mistral-7BInstruct-v0.3 using QLoRA with a frozen 4-bit base into low-rank adapters (rank r=32, scaling α=16). A smaller model, Llama-3.2-3B Instruct, was retained as a scaling control for RQ5. The comparison was conducted across five corpora spanning project, academic, parliamentary, and municipal proceedings.

Pruning and Token Efficiency

The researchers investigated whether removing low-value transcript lines preserves quality at a reduced token count (RQ3). A pruning task was applied using gemma-3-27B-it3 under greedy decoding to remove pure conversational filler. The results showed that pruning low-value transcript lines removes ∼15% of tokens from the conversational corpora at no measurable cost. Furthermore, they found that pruning preserves quality while cutting ∼15% of tokens from the conversational corpora, and this saving is most useful exactly where balancing is most aggressive.

Unit of Balancing and Model Scale

The study distinguishes between balancing by token count versus balancing by example count (RQ4). They found that balancing by token rather than examples weights domains differently as equal meeting counts produce very unequal token amounts. The findings were validated on a smaller model, Llama-3.2-3B, showing that the redistribution effect holds across different backbones.

Evaluation and Findings

The evaluation used both automatic metrics (ROUGE1/2/L/Lsum and BERTScore-F1) and a fact-level LLM judge. The analysis revealed that balancing redistributes quality, raising minority domains while lowering majority domains, rather than adding it uniformly (RQ1). Specifically, at matched 32M budgets, Balanced leads every minority domain on both metrics, while the macro average favors balanced and the micro average favors natural. Finally, they concluded that practitioners should decide on allocation based on whether the minority domains matter: their share under proportional allocation is fixed at 1-2% regardless of budget, so matching balanced quality on those domains requires far more total data.

Conclusion

The results provide a basis for deciding when to balance an imbalanced multi-domain mixture, and on what unit. The paper suggests retaining the native distribution when the deployment is majority-weighted, balancing when all domains must be served, balancing by tokens rather than examples, and pruning prior to training in either case.

The gist

At matched token budgets, balancing does not lift every domain; it trades between them, improving all three minority domains (AMI, ICSI, ELITR) on both metrics while shifting quality from the majority domains (EPM/MB). The effect is visible only when reading results per domain rather than relying on aggregate averages. Balancing by tokens allocates by design, whereas balancing by examples allocates by accident because equal example counts produce very unequal token amounts across corpora.

Table 20: Full per-domain results (pruned, seed 42) across the complete metric family (ROUGE-1/2/L/Lsum and BERTScore-F1), with bootstrap 95% CIs in subscripts.

Dataset Budget Balanced Natural ∆ Bal. Nat.

:---::---::---::---::---:

Improvements for AI systems

Here are specific improvements for AI systems based on the findings of this scientific paper, categorized by the area of application:


)1. Dynamic Domain Allocation Strategy (For Joint Fine-Tuning LLMs):

The improved system should move beyond simple uniform or proportional allocation and implement a dynamic allocation strategy that adjusts token distribution based on domain criticality and budget constraints.

  • Improved System Capability: The AI can now intelligently decide whether to prioritize minority domains (e.g., academic discussions, parliamentary sessions) over data-rich domains (e.g., municipal proceedings) when the goal is maximizing quality in underserved areas, even at the expense of majority domain performance.

  • Specific Action: Implement a mechanism that identifies data-scarce domains and automatically shifts tokens towards them up to a predefined threshold (e.g., ensuring minority domains receive at least 10% of the total budget, rather than relying on the fixed 1-2% natural allocation). This directly leverages the finding that balancing redistributes quality.

)2. Token-Aware Summarization and Input Pruning Pipeline:

The system should integrate a pre-training pruning step specifically tuned to preserve content density across different meeting types, rather than just removing generic conversational filler.

  • Improved System Capability: The AI can generate summaries using significantly fewer tokens while maintaining high factual fidelity (BERTScore-F1) and conciseness, particularly in complex or long transcripts (like those from academic meetings).

  • Specific Action: Integrate the Pruning Prompt described in Appendix N.1 into the pre-training pipeline. Crucially, the system should be trained to recognize that pruning is most effective when combined with token balancing (i.e., removing filler lines from data-scarce domains where repetition is high), allowing it to achieve a target summary length (e.g., 2048 tokens) with higher semantic density than a standard, unpruned model.

)3. Unit of Balancing Calibration: Token vs. Example-Level Allocation:

The system must be able to distinguish between balancing by the number of examples (meetings) and balancing by the total token count, as this dictates the resulting quality shift.

  • Improved System Capability: When fine-tuning on a fixed compute budget, the AI can select whether to balance based on token mass or meeting count. This allows for targeted optimization when dealing with highly variable transcript lengths across domains.

  • Specific Action: Implement a control where users can specify the balancing unit (e.g., Balance by Token Count vs. Balance by Example Count). The system will then dynamically select the allocation scheme (Balanced, Natural, or Example-Level) that is most appropriate for the specific domain being summarized, as shown in RQ4 analysis.

)4. Robustness Verification via Model Scale Control:

The system's performance metrics must be validated across different model families and sizes to ensure its improvements are due to the data allocation strategy, not model architecture bias.

  • Improved System Capability: The AI can provide confidence scores on its summarization quality that are robust across different LLM architectures (Mistral vs. Llama), ensuring the findings generalize beyond a single backbone.

  • Specific Action: When deploying the summarization model, run a quick validation check against a smaller, different family model (like Llama-3.2-3B) using the same allocation strategy to confirm that the redistribution effect observed on large models holds true for smaller ones, as validated in RQ5.

)5. Fact-Level Ground Truth Generation and Self-Correction:

The final output should not just be a text summary, but a verifiable set of atomic facts with confidence scores derived from cross-referencing the transcript and reference minute.

  • Improved System Capability: The AI can generate summaries where every claim is backed by explicit evidence from the source material (high Faithfulness) and is highly relevant to what was actually recorded in the reference (high Conciseness).

  • Specific Action: Implement the three-stage LLM judge protocol (Atomic Fact Decomposition, Completeness, Faithfulness, Conciseness) as a mandatory post-generation step. The system will only output a summary if it can successfully pass the Completeness check against the reference and the Faithfulness check against the transcript. This moves summarization from a creative task to a verifiable information extraction task.

Sources

Related papers