A helps B while B hurts A: directed transfer in instruction-tuning mixture

arXiv:2609.39702 · cs.AI, cs.CL, cs.LG · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "A helps B while B hurts A".

Jane: A directed transfer map reveals that task selection for language model instruction tuning is not symmetric, demonstrating that certain tasks can actively help a target while others actively harm it.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we're diving into this paper titled "A helps B while B hurts A: directed transfer in instruction-tuning mixture." Basically, the core idea here is that picking which instruction-tuning tasks to train on isn't a simple matter of just adding more or picking similar sources; sometimes one task actively helps your target model while another one actively harms it.

Jane: It sounds like the paper is arguing against these common, straightforward ideas about how we should select training data for language models, suggesting that the relationship between a source and a target can be both beneficial and detrimental depending on the context.

Lu: I'm really intrigued by this concept of signed transfer; it moves beyond just counting sources or assuming everything is positive or symmetric in its effect on the model performance. It suggests there's a nuanced directionality we need to consider when building these specialized instruction-tuning mixtures.

Meng: From an engineering standpoint, if one task type can actively hurt the overall mixture performance, that means our heuristic for selecting tasks needs to be much more sophisticated than just looking at source similarity metrics. How does this translate into a practical selection strategy?

Lalam: I see this as crucial for developing models that actually improve culture in a meaningful way; if we tune them using sources that are actively counterproductive, we might end up with models that are brittle or perform poorly on the very things we want them to excel at. This paper suggests a much smarter way to curate those instruction-tuning sets.

Tom: Exactly, Lalam; it’s about moving from simple heuristics to something data-driven that predicts how a mixture will perform before you actually spend time fine-tuning. The authors introduce this transfer map, which is essentially a signed estimate of how much each source helps or hurts the held-out target model.

Jane: That transfer map is presented as the central contribution, defined by Equation two and it’s fitted across hundreds of fine-tuning runs on models ranging from 0 point 6B up to 32B parameters on both Qwen3 and Mistral models.

Lu: The idea that this map predicts a held-out target’s accuracy on unseen mixtures before any actual training occurs is significant because it allows for a form of pre-emptive selection based on this signed estimate. This seems like a big step toward optimizing the initial setup phase.

Meng: I'm interested in the prediction itself; if the map can predict performance, that suggests we could use it to filter out poor mixture candidates right at the start, saving significant compute resources during the fine-tuning process.

Paper summary: Lalam: It’s also interesting that this map is specific to its target and corpus, yet it transfers across different model scales; for instance, a mixture selected beforehand can beat training on all source tasks at every other size tested. That implies a robust selection mechanism that isn't tied to a specific model architecture.

Tom: That scale independence is really compelling; the paper shows this transfer map is stable across different parameter sizes, agreeing between Qwen3-32B and Mistral-24B at a correlation of zero point eight eight zero for non-MATCH sources. This stability gives us a lot of confidence in using this map as a general guide.

Jane: It’s also important to note that the map's selection strategy involves choosing sources with positive uncentered coefficients and dropping those with negative coefficients, which the paper claims outperforms similarity selectors because a similarity score carries no sign.

Lu: That contrast between signed estimation and simple similarity metrics really highlights why this approach is powerful; it acknowledges that just being related isn't enough to say one source will help another. The analysis shows that task A can help task B while B hurts A, which is a key insight into directed transfer.

Meng: So, the paper moves us away from just asking "how many sources?" and toward understanding the specific *effect* of each source type on the target model's capability. That shifts our focus from quantity to quality in instruction-tuning mixtures.

Lalam: And from a cultural perspective, if we can select mixtures that are demonstrably beneficial rather than just large or similar, it helps ensure the resulting AI reflects the desired behaviors and knowledge structures we are aiming for. It’s about intentionality in training data curation.

Tom: Moving on to the conclusions of "A helps B while B hurts A: directed transfer in instruction-tuning mixture," we see that this research refutes two common assumptions about how source tasks interact with target models: that transfer is always positive, and that transfer is always symmetric.

Jane: The authors conclude by emphasizing that helpfulness is a signed property of ordered source–target pairs, meaning the directionality matters significantly in this process. They show that the two directions differ in fourteen and thirteen of twenty-eight pairs across both model families, with some pairs having opposite signs where each direction's interval excludes zero.

Lu: This directional finding is substantial because it confirms that we absolutely need to consider the order of tasks when designing these instruction-tuning mixtures, as training sequence has a demonstrable effect that the map doesn't generally predict.

Meng: That means we can’t just treat all our source tasks equally when we build a mixture; sequencing them could be a critical design parameter if we want to maximize the positive contributions and minimize the negative ones.

Paper summary: Lalam: If we look at the broader implications, this suggests that building advanced instruction-tuning systems requires a much deeper understanding of these directed relationships than was previously assumed by simple heuristics. It points toward more sophisticated data selection pipelines for developing capable AI systems.

Tom: The paper really hammers home that which sources matter is not necessarily how many there are; the authors state clearly that no fit can separate mixture size from the average source effect, meaning adding more tasks doesn't automatically guarantee better results if some of those tasks are detrimental.

Jane: And they show that one interfering task type can actually cancel out the benefit of all the other sources, as seen when training on all other tasks in equal parts barely beats an untuned model on multi-hop questions or critiques. That shows how delicate these mixtures can be.

Lu: The paper also addresses limitations, pointing out that because every score comes from an automated rule and not human judgment, the method doesn't isolate supervision format because passages behind different task types overlap only partially. Furthermore, they found that the second corpus had much smaller training pools and they couldn't resolve the probe’s gain there.

Meng: That limitation about not being able to isolate supervision format is a real practical hurdle for implementation; it means we can't perfectly dissect why a certain task type works better in isolation versus in a mixture setting.

Lalam: So, while this transfer map is incredibly useful for predicting unseen mixtures before training, the paper also signals that we still have to investigate the underlying factors—like which specific task types act as interferers—to fully understand the system's behavior.

Tom: It sounds like "A helps B while B hurts A: directed transfer in instruction-tuning mixture" provides us with a much more rigorous framework for selecting instruction-tuning tasks by introducing this signed estimate of source influence.

Jane: It’s really about understanding that the relationship between components isn't just a simple addition; it has direction and nuance, which is something we need to keep in mind as we build more complex AI systems.

Lu: This work opens up avenues for designing instruction-tuning pipelines that are far more targeted and less reliant on brute-force task inclusion, focusing instead on the beneficial interactions between components.

Meng: For practical deployment, having a predictive tool like this map would allow us to prototype and select the optimal mixture much faster in our development cycles without running through hundreds of exhaustive fine-tuning runs initially.

Lalam: Ultimately, this research pushes us toward building AI not just by aggregating knowledge from many sources, but by intelligently managing the directed flow of influence between them to achieve better outcomes.

Conclusion: Tom: So we’ve just been digging into how this new research maps out the actual direction of transfer in instruction tuning mixtures, and now it’s time to wrap up with some big thoughts on what all this means.

Jane: It really boils down to understanding that these tasks aren't just randomly added together; they have a specific, sometimes opposing influence on the target model.

Lu: The core concept is that we can predict which mixture will work best before we even start training because we’re measuring how much each source helps or hurts the final result in a signed way.

Meng: From an engineering standpoint, this predictive power suggests we could build smarter selection pipelines that avoid wasting compute on mixtures destined to fail.

Lalam: And for us, it means we can curate instruction-tuning sets with a much higher degree of intentionality, ensuring the AI develops the specific cultural nuances we are aiming for.

Tom: Exactly! The title itself, "A helps B while B hurts A," perfectly captures that signed and directed nature of this transfer map.

Jane: It’s a simple way to say that the relationship between a source task and our target isn't always straightforward; there’s a back-and-forth dynamic at play.

Lu: The authors have done some really deep work fitting this map across hundreds of fine-tuning runs on various model sizes, which gives us a lot of confidence in its generalizability.

Meng: I’m impressed that the map remains stable even when comparing very different model architectures like Qwen3 and Mistral; that sort of consistency is hard to find.

Lalam: It implies we aren't just looking for the biggest set of sources, but rather the right *direction* of influence between them to get the best outcome.

Tom: Absolutely, so this isn't just about quantity anymore; it’s about understanding how these components interact and whether they are actively aiding or hindering each other.

Jane: This is a big step because it moves us away from simple rules and toward a data-driven measurement that anticipates performance on unseen data.

Lu: The prediction power here, where the map can estimate accuracy before training, opens up so many creative possibilities for designing these mixtures.

Meng: I’m thinking this could drastically speed up our iterative development cycle by allowing us to test mixture candidates much faster in the early stages of model building.

Lalam: It gives us a powerful tool to ensure that every piece of data we feed into the AI is contributing positively to its ultimate capabilities and the kind of intelligence we want it to exhibit.

Tom: So, this directed transfer map suggests a much more nuanced way for us to select and curate our instruction-tuning data.

Jane: And as we wrap up this section, what’s next is understanding how these predictive insights actually translate into practical choices for model development.

Nima H. Siboni, Vahid Rostami

Juna.ai · Computational Systems Neuroscience, Institute of Zoology, University of Cologne

cs.AI, cs.CL, cs.LG

Submitted: 2026-09-30

Updated: 2026-09-30

Code: https://github.com/Vahidrostami/directed-transfer

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: A directed transfer map reveals that task selection for language model instruction tuning is not symmetric, demonstrating that certain tasks can actively help a target while others actively harm it.

Key concepts

Transfer Map
A signed estimate of how much each source task helps or hurts a specific target task. It is calculated using a formula that incorporates the average effect and the directional influence of various source tasks on the target.
Directed Transfer
The finding that transfer between tasks is not symmetric. This means Task A can positively influence Task B, but Task B can also negatively impact Task A, depending on which model family or corpus is being considered.
Mixture-Agnostic Baseline
A comparison method where performance is measured without considering the specific mixture of tasks used for training. The transfer map's predictions significantly outperform this baseline, showing its predictive power in selecting optimal task combinations.

Terminology

Summary

A directed transfer map reveals that task selection for language model instruction tuning is not symmetric, demonstrating that certain tasks can actively help a target while others actively harm it. This finding shifts the paradigm from simple heuristics to a data-driven measurement that predicts unseen mixture performance before any fine-tuning occurs.

The Gist

Transfer is signed and directed: task A can help task B while B hurts A, on both model families, between task types from one corpus and inside fixed-budget mixtures.

How it works

The core contribution is the introduction of the transfer map, a signed estimate of how much each source helps or hurts each held-out target, defined by Equation 2:

(yT(M) = μT + ∑Sβ̂S→T zS)

This map is fitted in hundreds of fine-tuning runs on Qwen3 and Mistral models across various parameter scales. It predicts a held-out target’s accuracy on unseen mixtures, achieving less than half the error of a mixture-agnostic baseline. The map is specific to its target and corpus but transfers across model scale; a mixture selected in advance beats training on all source tasks at every other size tested.

Key Findings from the Transfer Map

The analysis addresses several questions regarding task selection heuristics:

  1. Transfer is signed and directed, with the two directions differ in 14 and 13 of 28 pairs across model families, and some pairs having opposite signs where each direction's interval excludes zero.

  2. Which sources matter, not how many; no fit can separate mixture size from the average source effect. The count heuristic is insufficient because a mixture’s size counts the sources it contains.

  3. One interfering task type can cancel the benefit of all others; for example, on Qwen3-32B, training on all other tasks in equal parts barely beats the untuned model on multi-hop questions (20% to 22%) and critiques (7% to 8%).

  4. The map predicts and selects mixtures before training; its recorded predictions have an MAE of 2.31 pp, which is less than half the error of a mixture-agnostic baseline.

Predictive Power and Selection

The map functions as a selector by choosing sources whose uncentered coefficients are positive, while dropping those with negative coefficients. This selection strategy outperforms other methods:

(The map predicts the effect of moving from the uniform mixture to the selected mixture before training.)

The recorded predictions outperform both baselines, leading with an MAE of 2.31 pp against a mixture-agnostic baseline MAE of 6.20 pp. The map's selection beats similarity selectors because a similarity score carries no sign.

Stability and Robustness

The transfer map is stable across model scale; it agrees between Qwen3-32B and Mistral-24B at a correlation of r = 0.880 for non-MATCH sources. Furthermore, the map is specific to its target, as it agrees across model families but did not carry over to our second corpus. The map's selection is specific to its target, and a selection serves one target at some cost to the other tasks.

Order Effects and Limitations

The experiment on training order tests whether the map predicts the effect of training sequence. The verdict is that the map therefore does not, in general, predict the effect of training order, refuting a magnitude law where effects are half the differential. Additionally, limitations include:

(Every score comes from an automated rule; no human judges answer quality.)

The design cannot isolate supervision format because passages behind different task types overlap only partially. The map is stable across model scale and corpus shift, but the factors that move the map—such as which task type acts as an interferer—remain open questions. The second corpus had much smaller training pools, and we could not resolve the probe’s gain there.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on this research, along with what those improved systems will be capable of:


  1. A new model training pipeline that uses a pre-computed Transfer Map instead of heuristic task selection (like the count or similarity heuristics).

  2. An AI system capable of making real-time, data-driven decisions about which instruction-tuning tasks to prioritize for a specific target instruction, based on the predicted signed transfer effect from the map.

  3. A system that can dynamically select an optimal mixture of source tasks (e.g., selecting a subset of 7 sources) to maximize accuracy for a given target task, instead of training on all available tasks or only similar ones.

  4. An AI capable of identifying and neutralizing interfering task types—those with negative transfer effects—during the fine-tuning process, leading to higher performance on sensitive reasoning targets (like causal explanation or multi-hop questions).

  5. A system that can predict the accuracy of a held-out target on unseen mixtures by consulting the Transfer Map before any fine-tuning is performed, outperforming mixture-agnostic baselines by more than 50%.

  6. An AI capable of choosing which source to include in a mixture for a given target based on its predicted positive transfer effect (positive coefficient in the map), effectively implementing a select the helpful, drop the harmful strategy.

  7. A system that can adapt its training strategy across different model sizes (0.6B to 32B parameters) by using scale-agnostic map entries, ensuring optimal task selection regardless of the model's parameter count.

  8. A system that can leverage recorded predictions from previous, distinct fine-tuning runs to predict the outcome of a new mixture before training begins, achieving predictive accuracy significantly better than baselines (MAE < 2.31 pp).

  9. An AI capable of understanding the directionality and asymmetry of task transfer—knowing whether Task A helps Task B or hurts it—allowing for nuanced optimization rather than assuming symmetry.

These improved AI systems can perform the following specific actions:

  1. They will achieve a significant, quantifiable boost in performance on complex, high-level reasoning tasks (causal explanation, multi-hop questions) by precisely selecting the most beneficial subset of training data sources.

  2. They will minimize performance degradation caused by negative transfer from certain instruction types within a fixed training budget.

  3. They will be able to select the exact combination of source tasks that yields the highest accuracy for a specific target, surpassing current methods that rely on simple heuristics (like adding more tasks or picking similar ones).

  4. They will function as highly accurate pre-training selectors, predicting how a model will perform on an unseen instruction mixture before any expensive fine-tuning runs are even initiated.

  5. They will be robust across different model architectures and scales, maintaining high predictive accuracy because the underlying transfer map is stable across model sizes and families.

  6. They will optimize resource allocation by intelligently dropping source tasks that interfere with the target's learning process, leading to more efficient training runs while maintaining or increasing target accuracy.

Abstract

Adapting a language model to a specialized corpus means choosing which instruction-tuning tasks to train on under a fixed budget, and testing one choice costs a fine-tuning run. Common heuristics add more source tasks or pick sources similar to the target. The first assumes transfer is never negative; the second, that it is symmetric. We show that both assumptions fail: task A can help task B while B hurts A, so helpfulness is a signed property of ordered source--target pairs. We introduce the transfer map, a signed estimate of how much each source helps or hurts each held-out target. We fit the map in hundreds of fine-tuning runs on Qwen3 and Mistral models from 0.6B to 32B parameters, with all sources drawn from one corpus and no training examples from the target. The map predicts a held-out target's accuracy on unseen mixtures: recorded before those runs, its predictions have less than half the error of a mixture-agnostic baseline. The map is specific to its target and corpus but transfers across model scale: a mixture selected in advance at one size beats training on all source tasks at every other size we tested. Transfer is thus a property of the data. The map selects the tasks that help and drops the one that interferes: accuracy on the reasoning targets (causal explanation, multi-hop questions and methodological critique) rises by up to 14 percentage points over training on all source tasks.

Sources

Related papers