Forward and Inverse Virtual Metrology for Phototransistor Gain: A Hierarchical, Uncertainty-Aware Approach for Small Production Datasets
Mahshid Amirabgir, Lorenza Ferrario, Paolo Conci, Mahdieh Amirabgir, Giancarlo Orengo
Fondazione Bruno Kessler · University of Rome Tor Vergata · Microfab Solutions · Politecnico di Milano
cs.LG, cs.CE, cs.SY, eess.SY
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 43 pages, 15 figures
Code: https://github.com/mahshid-amirabgir/phototransistor-virtual-metrology
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: The paper addresses the challenge of predicting and optimizing the current gain of silicon bipolar phototransistors from fabrication process parameters in a small-sample, hierarchically structured
Terminology
Summary
The paper addresses the challenge of predicting and optimizing the current gain of silicon bipolar phototransistors from fabrication process parameters in a small-sample, hierarchically structured setting. The authors note that "The customization, optimization and stabilization of the process flow of a silicon bipolar phototransistor commits months of cleanroom time before a finished device can be measured, so a model that predicts device gain from process parameters before a run has value out of proportion to its accuracy. The study is conducted on
a real fabrication history, thirteen to fourteen process runs of a single device: a small-sample, hierarchically structured setting unlike the large-corpus regime of conventional virtual metrology."
The central finding is that Decomposing the variance of device gain, we find that roughly half of it lies between process runs rather than within them, so recipe-only prediction is bounded by construction.
The paper provides three main contributions: "a forward gain predictor with a relative, uncertainty-aware signal, an inverse search that returns recipes for a target gain, and, as the foundation for all of it, a multi-level data-quality assessment tailored to the nested physical entities of fabrication (batch, wafer, die) with an explicit cross-level linkage score." The normalized dataset and analysis code are released for full reproducibility.
The paper is situated in the context of semiconductor research institutions contributing to the lab to fab
path, where highly innovative laboratory devices are optimized to become part of application systems based on stabilized, repeatable fabrication process flows.
The specific device studied is a silicon bipolar phototransistor, which can be addressed to industrial applications such as decoders, which require a specified collector–emitter breakdown voltage (BVCE), or to applications requiring controlled, high response speed.
The key practical problem is that Simulation, design and fabrication of a silicon phototransistor customized to a specific application commits months of cleanroom time and irreversible material before a single finished device can be measured.
This creates an asymmetry "between the speed of a decision and the cost of learning whether it was right, which is what makes predictive modelling of fabrication outcomes valuable: a model that forecasts gain from process parameters before a run, even imperfectly, can convert blind iteration into informed choices and remove clearly infeasible recipes before they consume fab time."
The paper contrasts its setting with mainstream virtual metrology: "That literature, however, has developed largely in high-volume processes where tens of thousands of wafers are available, and where natural groupings in the data (chambers, tools, products) are commonly handled by training a separate model for each group, an approach that limits generalization across groups. The present setting is
the opposite on both counts. The data come from a real but limited fabrication history, on the order of fourteen process runs of a single device."
The paper lists three main contributions:
-
Prediction and recipe-design capability: "a forward model that predicts gain from process parameters with leave-one-batch-out honest accuracy and a relative uncertainty signal, an inverse search that returns the recipes capable of reaching a target gain, and a Gaussian-process uncertainty layer that tells an engineer how far to trust each prediction."
-
Scientific finding about the process: "that between-run variance dominates recipe effects (the intraclass correlation, in the sense of Nakagawa, Johnson, and Schielzeth (2017), is near one-half), and that the within-run position effect is real but sign-inconsistent across runs, a structure invisible to any single pooled model and recovered only by modelling the run explicitly."
-
Multi-level data-quality methodology: "that extends per-level quality scoring to the nested physical entities of fabrication data and introduces an explicit cross-level linkage score, with each level's dimensions chosen, and each exclusion justified, rather than applied uniformly."
The device is "an npn silicon bipolar phototransistor fabricated by a planar silicon process, in which the emitter and base are defined by ion implantation while the collector is provided by an n-type epitaxial layer on the substrate rather than by implantation. It descends from
the optical-encoder phototransistor developed at the institute (then IRST, now Fondazione Bruno Kessler) and optimized through coupled process and device simulation by Dalla Betta et al. (1998), who showed that increasing the emitter implant dose raises the current gain."
The fabrication process involves a long sequence of material depositions, selective removals, thermal and doping steps, and metrology, interspersed with lithography steps that transfer the device layout onto the wafer.
The emitter is defined by phosphorus ion implantation, with a screen oxide (target thickness 100 nm) controlling the quantity and depth profile of the implanted phosphorus and limits crystal damage from the incident ions.
The models use four primary process parameters: the screen-oxide thickness, the emitter implant dose and the emitter sheet resistivity after activation. A fourth feature, the within-batch resistivity spread, captures the uniformity of resistivity control.
All features are released in normalized form, each is expressed as a ratio to a per-arm reference (a robust within-arm central value).
The response variable is Device current gain hFE,
measured as the median gain over the dies that pass quality screening, those flagged GOOD and surviving the cleaning steps.
Gain values are subject to a physical-validity filter of [100, 5000]: readings outside this window are treated as non-physical and removed before aggregation.
The device-acceptance range for the encoder application is [600, 1500],
which is a design target rather than a data filter.
The modelling corpus spans [644.0, 1444.5].
The data are hierarchical with three strata: fabrication batch, wafer and die.
The corpus contains Thirteen gain-bearing phototransistor batches,
each comprising 16 to 25 wafers except batch 3 (45),
with more than 20,000 phototransistors and 16 test strips per wafer, on the order of 9.3 million die-level gain measurements in total.
Batch 3 is split into two runs: "Batch 3 was fabricated as 45 wafers, but a data-driven analysis identifies two sub-populations within it: re-indexing wafer position at wafer 26 reverses the position–gain correlation, a signature independent of where the boundary is placed, and a corroborating discontinuity in median gain appears at the same location. The runs are labelled
b3a (wafers 1–25) and b3b (wafers 26–45)."
Two analysis arms are defined: "The strict arm (Arm A) excludes batch 2 entirely, yielding 13 runs over 260 wafers. The inclusive arm (Arm B) instead imputes the missing resistivity, on the reasoning that batch 2 retains valid thickness, gain and other data, and need not be discarded over a single absent feature, yielding 14 runs over 285 wafers."
The paper argues that Fabrication data does not fit
the shape of single-level data-quality frameworks: "It is organized in strict physical strata: a batch is processed under one recipe, a wafer is one substrate processed within that batch, and a die is one measured device, and the same quantity behaves differently at each stratum. The failures that matter most are
cross-level integrity failures (a wafer tracing to no recipe, an implausible die count, a batch identifier concealing two process runs) that are properties of the links between levels."
The MLDQ framework scores nine data-quality dimensions
: Traceability, Completeness, Accuracy, Consistency, Redundancy, Balance, Uniqueness, Usefulness, and Quantity. Each level scores a tailored subset, with every exclusion carries a reason rather than a silent omission.
The cross-level linkage dimension is "the new scored quantity, measuring how cleanly the hierarchy joins through three sub-scores: L1 (wafer-to-batch), the fraction of wafers joining exactly one recipe; L2 (wafer-to-die), the fraction with a non-empty die set; and L3 (count plausibility), the fraction whose die count falls within a robust band around the corpus median." The combined score is:
Linkage = 0.40 L1 + 0.40 L2 + 0.20 L3
The L3 plausibility band is band = m ± max(3 MAD, 0.05 m),
where MAD is the median absolute deviation of the per-wafer die counts and the five-percent-of-median floor prevents the band from collapsing when the MAD is zero.
The headline result is that MLDQ is robust to the imputation choice by construction. The baseline and Arm B scopes produce identical scores at every level mean and all nine wafer-level dimensions, to three decimal places.
The batch level mean is 72.9, low by design, dragged by the Quantity score that encodes small N as a measurable property
; the wafer mean is about 95,
the canonical 'is the dataset usable' figure
; and the linkage score of about 98 is high because L1 and L2 are both perfect, with only L3 below 100.
The L3 score "surfaces not one anomaly but three structurally distinct patterns, all in batches 5, 8 and 11: a deliberate double-measurement protocol at exactly twice the modal count; extra measurement-letter scans 14–23% above modal; and a small set of extreme overcounts several times modal."
The paper decomposes the variance of wafer-level gain into between-run and within-run components
using a one-way decomposition with the process run as the grouping factor
giving an intraclass correlation coefficient (ICC).
The ICC is 0.537 in Arm A and 0.501 in Arm B; that is, roughly half of the total variance in gain lies between runs rather than within them at the point estimate.
The variance components are comparable in magnitude (Arm A: between-run 7987, residual 6893; Arm B: 7336 and 7319 in the same standardized units).
A cluster bootstrap places the 95% interval at [0.03, 0.74] (Arm A) and [0.03, 0.73] (Arm B), with 65% and 63% of resamples respectively retaining an ICC above 0.40.
The modelling consequence is that "A model whose only inputs are recipe parameters, which are constant within a run, can by construction explain only the between-run share of variance, and only to the extent the recipe actually drives it; it has no access to the within-run half at all."
The model is specified as yij = x⊤ij β + ui + vi pij + εij,
where "xij are the standardized recipe features (emitter dose, screen-oxide thickness, emitter sheet resistivity and its within-batch spread) with fixed effects β, pij is the within-run position, ui ∼ N(0, σu2) is the per-run random intercept, vi ∼ N(0, σv2) the per-run random slope on position, and εij ∼ N(0, σε2) the residual."
The random slope is strongly warranted
: A likelihood-ratio test of the random-slope model against the random-intercept-only model gives χ2 = 60.2 (Arm A) and 52.1 (Arm B).
Among fixed effects, "emitter dose and screen-oxide thickness carry the largest positive standardized coefficients (dose +46.3, thickness +37.8 in Arm A; +43.0 and +30.6 in Arm B, all with p < 0.005 in both arms), and the resistivity spread carries a significant negative coefficient (−32.7, p = 0.037 in Arm A; −37.6, p = 0.009 in Arm B)."
The marginal R2 is 0.468 (Arm A) and 0.466 (Arm B), while the conditional R2, fixed plus random effects, is 0.678 and 0.640.
The gap carries the finding in two numbers: a substantial share of what the model captures, between a quarter and a third of the explained variance, is run-specific structure the fixed effects cannot see.
The headline finding is the position effect: the pooled fixed-effect position coefficient, the single slope a recipe-plus-position model would report, is significant and negative (−3.68 raw gain units per position step, p ≈ 10−5, Arm A; −3.96, p ≈ 10−6, Arm B).
However, "Fitting position against gain within each run separately, four runs reach Bonferroni-corrected significance in both arms, and they do not agree in sign: run 1 is strongly positive (Spearman ρ = +0.89) while runs 6, 9 and 11 are strongly negative (ρ = −0.83, −0.72, −0.81)."
The model's per-run random slopes span both signs and recover exactly this: run 1's slope is about +80 gain units while run 6's is about −80.
The per-run slopes and the independent per-run Spearman correlations agree closely (Pearson r ≈ 0.88–0.90 between the two).
The ridge regression minimizes β̂ = arg min ∥y − Xβ∥22 + λ∥β∥22, with penalty λ selected by leave-one-batch-out cross-validation.
Key results: "The recipe-only model's pooled LOBO-CV R2 is only marginally positive (0.062 in Arm A, 0.108 in Arm B); adding within-run position raises the pooled value slightly but drives the mean-of-folds R2 strongly negative (−1.22 in Arm A, −1.06 in Arm B). Adding position
raises the pooled leave-one-batch-out R2 from 0.062 to 0.108 (Arm A) and from 0.108 to 0.155 (Arm B), a difference of +0.046 and +0.047. The random forest
generalizes worse than ridge on every cross-validated metric: its pooled LOBO-CV R2 is negative (−0.28 in Arm A, −0.22 in Arm B)."
The deployable forward model is the mixed-effects model of Section 5.2, with the wafer-level ridge regression of Section 5.3 as a transparent cross-check.
The value is "not high point accuracy, which the variance hierarchy of Section 5.1 caps, but an honest forecast that, combined with the uncertainty quantification of Section 5.6, tells an engineer both a predicted gain and how far to trust it."
Feature importance analyses agree: "permutation importance ranks emitter dose and within-run position as the two strongest features in both arms (mean ∆R2 ≈ 0.47 and 0.46 in Arm A; in Arm B the order reverses, position 0.45 and dose 0.33), followed by the two resistivity features and thickness; and SHAP attributions return the same leading set: emitter dose, within-run position and mean resistivity."
The inverse search "evaluat[es] the mixed-effects forward model across a grid of candidate recipes (3,750 points spanning the three normalized emitter-dose set-points, fifty resistivity values and the twenty-five within-run positions, with thickness and resistivity spread held at their central values) and retain[s] the recipes whose marginal prediction band intersects the target band."
The results show The number of admissible recipes is unimodal in the target: zero at 600 and 700, rising sharply to a peak near a target of 1000 (about 1,970–1,985 admissible recipes), and falling back to zero by 1400–1500.
At a target of 1100, "the low dose set-point (normalized 0.929) admits no recipe in either arm; the mid set-point (1.000) admits an intermediate fraction (662 of 1,250 grid recipes in Arm A, 682 in Arm B); and the high set-point (1.071) admits all 1,250. This reveals
a three-tier identifiability structure in the emitter dose": The low dose is eliminative... The mid dose is identifying... The high dose is saturated.
The GP uses a Matérn kernel with ν = 2.5 plus a white-noise term
over emitter dose, mean resistivity and thickness
(excluding resistivity spread because its fitted length scale ran to the boundary
).
Under five-fold cross-validation its R2 is 0.501 (Arm A) and 0.447 (Arm B), close to the mixed-effects model's marginal R2.
However, Under leave-one-batch-out validation, however, the GP's R2 falls to strongly negative (mean-of-folds −1.79 and −1.69).
The predictive standard deviation is "translated into three tiers against the typical in-sample spread: a prediction is HIGH trust when the GP standard deviation is within 1.2× the training-region spread (a data-dense region), MEDIUM up to 1.5× (the edge of the training data), and LOW beyond that (an extrapolation zone)."
The paper states: "The random forest generalizes worse than ridge on every cross-validated metric... We therefore prefer ridge as the wafer-level cross-check: a flexible, high-capacity learner has no advantage when the binding constraint is between-run variance rather than functional complexity."
The honest limits are explicit: "The forward model's cross-run R2 is low by construction, not by under-fitting. The uncertainty tiers promise a relative signal, denser versus sparser regions of recipe space, not a calibrated coverage guarantee... The inverse search returns recipes the model predicts will hit a target, inheriting all of the forward model's uncertainty; it is a screening aid, not a guarantee of a fabricated outcome."
The assistant couples a local, open-weights large language model (Llama 3.1, served through Ollama) to the validated models through a small set of typed tools.
Two tools are exposed: "a forward tool that takes a recipe and returns the mixed-effects prediction... together with the Gaussian-process standard error and trust tier, and an inverse tool that takes a target gain and returns the admissible recipes."
The central design property is that "the language model itself performs no arithmetic and fits nothing... Because every gain value, every recipe and every trust tier originates from a model call rather than from the language model's own parameters, the assistant cannot fabricate a numerical result. Each response
carries a provenance record showing which tool was called, the arguments it received and the values it returned."
The paper concludes that "the dominant structure in the data is between-run rather than recipe-driven: roughly half of gain variance lies between process runs, and the within-run position effect changes sign across runs so that a pooled model averages it away."
The transferable contribution is methodological: decompose variance before choosing a model, score data quality across physical levels with explicit cross-level linkage, and surface scarcity and exclusions rather than hiding them.
Future work includes: "New datasets on bipolar phototransistors are expected in the next two years, supporting further validation of the proposed methodology. Furthermore, given the multi-level structure of the device, another exploitation of this work will be in considering different devices and the characterization parameters on which they should be optimized."
Improvements for AI systems
Based on this paper, here are the specific improvements I can make to AI systems, and what the improved systems can do:
Improvement: Before training any predictive model, the AI system automatically decomposes target variance into between-group and within-group components (ICC calculation) and uses this to set an upper bound on achievable prediction accuracy. It then selects model complexity accordingly.
What it can do: In any hierarchical manufacturing or experimental setting, the system will tell you upfront: Your recipe-only model is capped at 50% explainable variance by construction
— preventing wasted effort on over-parameterized models and setting honest expectations for stakeholders.
Sources
- Machine Learning based CVD Virtual Metrology in Mass Produced Semiconductor Process
- Gaussian Process Regression for Uncertainty Quantification: An Introductory Tutorial
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks