Black-Box Knowledge Transfer across Distinct Feature Sets
Oh-Ran Kwon, Daeyoung Ham
The Ohio State University · University of Texas at San Antonio
stat.ML, cs.LG, stat.ME
Submitted: 2026-08-10
Updated: 2026-08-14
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: The paper introduces a method for transferring predictive knowledge from a pre-trained black-box function to a new, heterogeneous input space when the available target features differ from those the
Terminology
Summary
The paper introduces a method for transferring predictive knowledge from a pre-trained black-box function to a new, heterogeneous input space when the available target features differ from those the black box expects. The core idea is an elementary decomposition of the target regression function:
E[Y X] = E[fb(Z) X] + E[Y − fb(Z) X]. z z z =: g(X) =: h(X) =: δ(X)
Here, The transferable component h is the part of the regression g that can be explained by the black box. The non-transferable component δ is the new predictive information in X that the black box does not capture.
The method is a two-step neural network procedure: Since h is a regression of fb(Z) on X, it can be estimated from the abundant DA alone. Only δ requires the limited DT. Both steps use deep ReLU (rectified linear unit) networks.
Specifically, Step 1 estimates h from unlabeled paired data DA = (Zj, Xj) using least squares over a network class, and Step 2 estimates δ from limited labeled data DT = (Xk, Yk) by fitting residuals, with a validation split to decide whether to include the δ estimate (λ ∈ 0, 1).
The theory provides prediction risk bounds. Lemma 1 (oracle version) states: ER(b gλora, g) ≤ Ch ϕhnA log3 nA + min 2∥δ∥22, Cδ ϕδnT log3 nT,
where the first term is the cost of estimating h from nA auxiliary observations, and the second is the smaller of the bias from discarding δe and the cost of estimating δ from nT target observations. Theorem 1 adds a validation cost: ER(b gλb, g) ≤ Ch ϕhnA log3 nA + min 2∥δ∥22, Cδ ϕδnT log3 nT + Cval n−1/2 T.
Transfer is beneficial in two regimes: "(i) when the black box is accurate, so that ∥δ∥2 is small and little is left to estimate; and (ii) when the black box explains the complex part of g, so that δ is smooth and can be learned from few labels. Our estimator adaptively attains the favorable regime."
The paper compares against non-transfer learning (estimating g from DT alone). Theorem 2 gives the non-transfer rate: ER(b gnt, g) ≤ Cg ϕnT log3 nT, ϕnT = max ϕhnT, ϕδnT.
Theorem 3 establishes a minimax lower bound: inf sup EP R(b g, gP) ≥ Clb ϕnT,
showing the non-transfer rate is minimax-optimal. Under additional conditions (nA ≫ nT ρh/ρ, ρδ > ρh, ρ ≤ 1/2), the worst-case risk of our estimator is of strictly smaller polynomial order than that of any estimator based only on DT.
The paper also discusses imputation-based learning as a baseline, noting it requires the black box to be stable under perturbations of its inputs. Modern black boxes are highly nonlinear, so even small imputation errors can produce large deviations in the output.
The proposed method is expected to be better when fb is sufficiently nonlinear and Z is not almost surely determined by X.
The framework extends to multiple black boxes. Theorem 4 shows the ensemble estimator converges at the rate of the best single black box, without requiring prior knowledge of which it is. Moreover, the ensemble can strictly improve on it when the black boxes are comparably accurate but distinct.
Specifically, part (i) gives conditions for strict improvement, and part (ii) guarantees safety: ER(b gens, g) ≤ minS∈ Z,W ERS + C eval n−1/2 val.
Simulations corroborate the theory: "The proposed method performs at least as well as the non-transfer NN in every setting, with the benefit of transfer increasing as s grows (a), as M shrinks (b), and as nA grows (c); these trends are consistent with Theorem 1. For two black boxes,
The ensemble performs best by exploiting both hZ and hW, and transferring either black box alone still improves on non-transfer NN and Imp+NN."
Real data analysis on chlorophyll-a concentration estimation shows: Transferring either black box alone achieves lower error than non-transfer NN and than its Imp+NN counterpart, and the ensemble attains the lowest error across all training fractions.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:
Improvement 1: Adaptive Black-Box Knowledge Transfer Across Heterogeneous Input Spaces
I can implement a two-stage neural network that decomposes any target regression function into a transferable component (explained by a pre-trained black-box model) and a non-transferable residual. The system first learns the mapping from new features (X) to the black-box's expected output (Z) using abundant unlabeled paired data, then fits the residual using limited labeled target data. This allows the AI to leverage any existing black-box model (e.g., a physics simulator, a large pre-trained model) even when the new task has different input features, without retraining the black box.
What the improved AI can do:
-
Predict outcomes in a new domain (e.g., medical imaging with different sensor modalities) by transferring knowledge from a black-box trained on a related domain (e.g., satellite imagery), using only a few labeled examples in the new domain.
-
Automatically decide whether to include the residual term based on validation performance, avoiding negative transfer when the black box is irrelevant.
Improvement 2: Risk-Aware Model Selection with Theoretical Guarantees
I can use the validation-based estimator (λ ∈ 0,1) to automatically choose between (a) relying solely on the black-box transfer, (b) adding the residual estimator, or (c) falling back to non-transfer learning. The system will have provable risk bounds (from Theorem 1) that guarantee performance is never worse than the best of these options up to a small validation cost (C val n T-1/2). This makes the AI robust to unknown black-box quality.
Improvement 3: Ensemble Transfer from Multiple Black Boxes
I can implement the multi-source extension (Theorem 4) that trains separate transfer components for each black box and combines them via an ensemble. The system will automatically converge to the best single black box’s performance without knowing which one is best, and can strictly improve when multiple black boxes are comparably accurate but capture different aspects of the target function.
Improvement 4: Nonlinearity-Aware Transfer (Superior to Imputation)
Unlike imputation-based baselines that require stable black boxes under input perturbations, my improved system directly models the conditional expectation E[f b(Z)X] using deep ReLU networks. This makes it robust to highly nonlinear black boxes (e.g., deep neural networks, chaotic simulators) where small imputation errors would otherwise cause large output deviations.
Improvement 5: Sample-Efficient Learning with Theoretical Rate Advantage
I can exploit the regime where the black box explains the complex part of the target function (i.e., δ is smooth). In such cases, the system achieves a strictly better polynomial-order convergence rate than any non-transfer estimator, provided the auxiliary data (n A) is sufficiently large relative to target data (n T). This means the AI can reach high accuracy with far fewer labeled target examples.
Sources
- PPI++: Efficient Prediction-Powered Inference
- Transfer Learning for High-dimensional Quantile Regression with Distribution Shift
- ChauffeurNet: Learning to Drive by Imitating the Best and Synthesizing the Worst
- A Recent Survey of Heterogeneous Transfer Learning
- Transfer Learning for Nonparametric Regression: Non-asymptotic Minimax Analysis and Adaptive Procedure
- Heterogeneous transfer learning for high-dimensional regression with feature mismatch
- SADA: Safe and Adaptive Aggregation of Multiple Black-Box Predictions in Semi-Supervised Learning
- Prediction-Powered Conditional Inference
- SMART: A Spectral Transfer Approach to Multi-Task Learning
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey