Demand Transfer Estimation at Scale via Restricted Logit Modeling

arXiv:2608.12680 · cs.LG, cs.AI · Submitted 2026-08-13 · Read on arXiv

Lakshya Garg, Deep Narayan Mishra, Swapnil Yadav, Haoan Wang, Sujal Alugubelli, Karthik Kumaran, Anupriya Sharma

Walmart Global Tech

cs.LG, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 8 pages. Accepted in the Main Conference of IEEE ICMLA 2026

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: Item demand forecasting is an integral component of store assortment optimization (SAO).

Terminology

Summary

Item demand forecasting is an integral component of store assortment optimization (SAO). Existing literature focuses on learning a suitable customer choice model and using this model to determine the value of an objective function (i.e. expected demand) with respect to an assortment proposal. However, for large item universe with many categories, this approach can prove inefficient, needing a separate demand forecast for every possible item assortment. An alternate approach to SAO exists whereby we combine the efficiency of forecasting item demand independently, while at the same time applying adjustments to the independent forecasts that account for the relations between item demand and the availability of other similar items on the shelf.

Central to this approach is the estimation of Demand Transfer (DT) coefficients. These DT coefficients represent the percent of a particular target item’s (Item that the customer walked in the store to buy) demand that is redirected to each other item in the universe should the target item be removed from the shelf. We introduce an approach that allows us to compute these DT coefficients on large item universes (assortments having 1 million+ items). Experiments on data as well as historical transaction data for multiple locations within categories demonstrate that when certain reasonable assumptions about substitution behavior are satisfied, our procedure is able to accurately estimate underlying DT coefficients and lead to improvements in demand forecasting.

Most approaches to solve demand of similar items have used a Choice Model as a component in a broader function that describes Revenue or some suitable objective function, which is then optimized as seen in Abdallah and Vulcano [3], Fisher and Vaidyanathan [7], this is computationally unsuitable for large item universe and for our downstream SAO use case of identifying how many units of each SKU should be put on the shelf. For our use case explicit transfer coefficients between items are needed, which some works such as Blanchet et al. [5], Şimşek and Topaloglu [14] have dealt with using a Markov Chain model. However, the methods used by Blanchet et al. [5], Şimşek and Topaloglu [14] to estimate transition probabilities suffer from scaling and stability issues for very large item universes. The work that we propose in this paper is a modification of the Markov Chain model to estimate explicit transfer coefficients between items leveraging a substitution framework that makes the original algorithm scalable for a huge item universe (1 million+ items).

One way to account for DT effects within the demand modeling process is to control for them by separately modeling an item’s demand Di,a under each potential assortment a ∈ X X ⊆ U, where U denotes the (universe) set of items under analysis. This formulation has been adopted in much of the relevant literature, in which a general revenue function r: 2U → R is maximized by identifying an optimal subset of items from the universe. DT is incorporated through the inclusion of a customer choice model that provides the probability that a customer will purchase item j given that a subset S ⊆ U is available on the shelf. These approaches are inadequate for our purposes for two main reasons. 1) A large retail operation may contain thousands of categories, each with very large item counts, which makes such an approach computationally impractical. For a single category U, performing an optimization over all possible item-availability combinations would require as many as U · 2U separate time-series forecasts. 2) Most possible assortments will not have been historically implemented in any stores due to the large item count, leading to a lack of data to make reliable demand forecasts corresponding to these assortments. 3) Finally, the above approaches only consider the question of which items to make available on the shelf and not how many of each of those selected items should be put on shelf. As our optimization engine must answer both of these questions we require explicit DT coefficients.

A key result shown in Blanchet et al. [5] is that a Markov Chain choice model becomes mathematically equivalent to a Multinomial Logit (MNL) model when the Markov Chain’s transition matrix has rank one—that is, when every item has the same pattern of transition probabilities to all other items in the assortment. This equivalence is important because the MNL model is far more scalable and computationally efficient than a general Markov Chain model. Therefore, we leverage this result—together with reasonable assumptions about customer demand—to estimate the parameters of an MNL model. Once the MNL parameters are estimated, we apply a simple transformation to interpret them as the transition probabilities of an equivalent Markov Chain choice model.

Item substitution is a symmetric, anti-reflexive, and non-transitive relation corresponding to the similarity of two items from a customer’s perspective. High substitutability between two items means that in general, customers are willing to purchase one item in the others place, or consider them functionally equivalent. This suggests that all substitutable items with respect to a given target item can serve as an implicit characterization of the target item within the same need state. We develop a comprehensive framework that addresses the complexity of real-world substitution patterns. Our approach to detect substitutability between items is to introduce a so-called Substitution Score for every pair of items. We define sij, ∀(i, j) ∈ U as follows: sij = I(θij ≥ τ); θij, τ ∈ R where θij represents the result of a decision variable for items i, j and τ represents a threshold value with U being the item universe. It is this θij which we refer to as our substitute score, and is a continuous variable expressing the degree to which i, j are related. If θij is above our chosen threshold, we can conclude that i, j are substitutes of one another.

The most reliable kind of data with respect to estimating substitute scores for an DT context is the historical store customer purchase data itself. Methodologies that use this kind of data can only be relied upon when we have sufficiently many purchases of both items (target and substitutable item) under question to decide whether their purchase patterns are (statistically) significantly associated. This condition is satisfied by many item pairs in our universe, and the methodology we use to measure their substitutability is called the Store Yule’s Q (YQ) score. The first step in calculating a pair’s Store Yule’s Q score is to compute its Odds Ratio (OR). The OR is a measure of association between two events. It is a metric ranging from 0 to ∞ and is defined as the relative probability between event 1 given event 2 and the probability of event 1 given event 2 does not happen. One of the key inputs into the substitute Yule’s Q metric is the Customer Odds Ratio. For items A and B, it is calculated as follows: ORcust = (nCustA∩B × nCust(A∪B)C) / (nCustB∩AC × nCustA∩B C) where nCustx indicates the number of customers that have purchased x. Although the Customer odds ratio is a useful metric for detecting items with a high association from a customer purchase perspective, its utility is hampered by the fact that items that are highly dissimilar but are often purchased together (i.e., complements such as bread and butter) will also tend to have high association along with true substitutes. In order to filter these complement pairs out, we utilize another metric called a Basket Odds Ratio, which measures the association of two items from a cart/basket perspective. The formula for this basket odds ratio is: ORbasket = (nBasketA∩B × nBasket(A∪B)C) / (nBasketB∩AC × nBasketA∩B C) where nBasketx indicates the number of baskets that contain x in the data. Before a substitute odds ratio is created, an adjustment is made to the customer odds ratio: ORCust-Adj = min(ORCust, ORBasket + 10). This adjustment is used to reduce the effect of extremely large customer odds ratios on final substitute scores, particularly when the basket odds ratio is also large. A substitute odds ratio is then calculated from the adjusted customer and basket odds ratios using the following formula: ORsubs = ORCust-Adj / (ORBasket + 1). By substituting Equation 10 in Equation 6 a Substitute Yule’s Q score is created. Therefore we get a final continuous substitute score (which corresponds to our decision variable θij for all i, j ∈ U mentioned in the methodology) in the range [−1, 1]. Throughout this work, items with substitution scores greater than 0.6 are considered substitutes. The threshold was selected based on empirical evaluation conducted during model development. Lower thresholds increased substitute-set coverage but introduced a larger number of weak substitution relationships that frequently corresponded to complementary rather than substitutable products. Conversely, higher thresholds improved precision but reduced substitute-set coverage, limiting measurable demand-transfer effects. New/low selling items often lack the historical sales data required to generate reliable substitution scores through Store Yule’s Q. To address this cold-start/low velocity item problem, we employ SBERT-based similarity models that leverage product attributes, including item descriptions, brand, and price, to rank the most similar items within the assortment and infer potential substitute relationships.

As our DT proportions are used as multiplicative factors of expected demand forecasts, they can be interpreted as probabilities. The estimation of ρ:= ρij i,j∈U amounts to estimating a set of conditional probabilities ρij:= P(Bj Ii), where Bj denotes the event that an arbitrary customer ultimately purchases item j, and Ii denotes the event that the customer initially intended to purchase item i, and that every item in U was available except for i. In our characterization of DT effects via the stochastic coefficient matrix ρ ∈ [0, 1]U ×U, the value of each element ρij is determined by two phenomena: 1) Item Loyalty: For certain items, when they become unavailable for purchase, some customers are unwilling to switch to any alternative. Also referred as the no-purchase option. 2) Item switching: If a customer is willing to switch to an alternative item, their preferences over the remaining available items are distributed according to the need state fulfilled by the item they originally intended to purchase. With respect to these phenomena, we make the following modeling assumptions: Assumption 1: Item loyalty and item switching preferences are independent factors. That is, Dij = (1 − ρiϕ) · ρij, where Dij represents the final percentage of demand transferred from item i to item j after accounting for the no-purchase option. Assumption 2: Item switching preferences are entirely determined by the underlying need state associated with an initial interest in the deleted item. That is, ρij = P(Bj Ii) = P(Bj Ni) where Ni denotes the event that the customer has the need state associated with item i, and that all items in the universe except i are available. A consequence of assumption 2 is that an item’s Demand Transfer profile depends not on the item itself, but on that item’s need state. Assumption 3: The need state corresponding to a customer’s initial interest in a particular item is fully characterized by the set of items that are sufficiently substitutable for it. That is, P(Bj Ii) = P(Bj Ni) = P Bj only m ∈ U σmi = 1) where σmi is an indicator function equal to 1 whenever m and i are substitutable items. Moreover, these substitutable items are the only items to which a customer initially interested in purchasing item i would consider switching.

In the MNL model, the mean utility parameter of an item i ∈ U is ηi. If we offer a subset O ⊆ U of items, then the probability that a customer purchases item i is e ηi / P j∈O e ηj. The likelihood function corresponding to the parameter vector η = (ηi)i∈U is given in Abdallah and Vulcano [3]: L(η) = P T t=1 P n j=1 Kj ηj − mt log P i∈St exp(ηi), where mt = P n i=1 Zit represents the total number of purchases in period t (or a single purchase if the data is generated at a per-purchase level, in which case mt = 1), and Kj = P T t=1 Zjt represents the total number of purchases of item j ∈ U over the observed historical period. One can efficiently maximize this function using the minorization–maximization method described in Abdallah and Vulcano [3].

Logit Models (LM) such as the MNL, despite their widespread use have two important drawbacks which hinders their usefulness for our specific optimization approach, especially at the scale at which we operate. Firstly with respect to computing DT coefficients, the output of a random utility model such as the MNL consists in a vector of item preference utilities. These utilities can be used to determine, given an offer set O ⊆ U, the probability distribution of customer purchase over all of the items in O. However, the output of a choice model yields only a marginal distribution of probabilities with respect to an item’s need state, since the conditional output of a choice model is conditional only on the availability of items on offer. In particular, a customer’s initial item preference determines the particular need state they are trying to fulfill, which will affect the set of items they are willing to consider for purchase in the items place, if the initial item happens to be off-shelf. Therefore, the need state of an item needs to factor into the choice modeling analysis. Prior work has typically addressed this issue by performing analysis itself at an item hierarchy level that includes only items fulfilling the same need state. For our use case, this solution was impractical, because existing business hierarchy levels (1) are too numerous to render computation practical, (2) fail to either directly separate items by need-state, or (3) fail to satisfy the needed mutual disjointness property to keep the analysis well-defined. Thus, in our solution, we perform the analysis at a category level, which allows us to fit more precise models to item sub-universes instead of a general model fit to all categories (which would be inaccurate), but also does not yield an impractical number of choice models to be estimated. However, since item categories contain multiple need states which can overlap (are not disjoint), we leverage our item substitution scores. We then use the category-level utility estimates and leverage the IIA condition to allow us to restrict attention only to highly-substitutable (with substitution score > 0.6) items during inference, which is enabled by assumptions 1,2, and 3 above. The IIA condition allows us to use MNL parameters to construct a valid probability distribution over the substitutes of the deleted item i using our category-level utilities, and interpret that as the probability that an arbitrary customer purchases any of the substitutes given that the substitutes are the only ones on offer. Finally, because we assume all DT coefficients across different need states are zero assumption 3, a Markov Chain model’s transition probabilities for the items in a particular need state could be entirely characterized by its submatrix for that need-state. Furthermore, because of assumption 2, the rows of this matrix would be identical because the transition probabilities from any item state i is fully determined by i’s need-state. So the matrix would have rank one, and thus as established in Theorem 3.1 of Blanchet et al. [5], we can interpret these MNL probabilities as transition probabilities, i.e. Demand Transfer coefficients.

The discussion in subsection ”Multinomial Logit (MNL) Model” presents a natural way of incorporating our need-state information, as determined by our substitution indicators, into the Demand Transfer coefficient calculation of our Restricted Logit Model. This means, that given utility estimates η̂ = (η̂n)n∈U, the final DT coefficient estimates for any given i ∈ U should be given by: ρ̂ij = exp η̂j / P k∈Si exp η̂k, ∀j ∈ S; ρ̂ij = 0, ∀j ∈ U S; ρ̂ii = 0, ∀i ∈ U where S ⊆ U is the set of all substitutes for i. We implemented the MNL fitting procedure described in (Abdallah and Vulcano [3]) for our Restricted Logit model across all categories using NumPy [8] and parallelized this algorithm across categories using Spark [15], parallelized across categories using internal distribute system.

To evaluate our model we perform offline backtesting on historical transaction data for items across a representative sample of locations for multiple categories. Logic: 1) Post training the model on one year of transaction data, we keep the last two months of data as a test set. 2) For this test set, we source the forecast and the actual sales observed for each item. 3) We compute two WMAPE metrics: a) Forecast MAPE: MAPEforecast = X Forecast − Actual / Actual; b) Adjusted MAPE: MAPEadjusted = X Adjusted Demand − Actual / Actual; c) Adjusted WMAPE weighted against the actual units sold for the item: WMAPE = P (Actual · MAPE) / P (Actual); d) The adjusted demand for item i is defined as D̃i = D̂i + ρji D̂j + ρki D̂k, where D̂i, D̂j, and D̂k denote the forecasted demands of items i, j, and k, respectively. Given items i, j, k form a subset S ⊆ U of the item universe and are assumed to be mutually substitutable. Table II reports offline backtesting results for a representative sample of locations across 50 categories (selected to span consumables and general merchandise with varying demand characteristics). The general trend observed was that appreciable reductions in forecast error occurred more frequently at locations with a higher baseline wMAPE, supporting our hypothesis that a significant portion of the raw forecast error is attributable to unaccounted demand transfer. Across our test categories we observed consistent reductions in wMAPE, supporting the efficacy of DT correction, which helps to establish the efficacy of our approach to DT calculation at scale. This is particularly important for our use case, where we need to estimate DT coefficients at scale for a large number of items in a large number of categories. Because actual customer substitution decisions are not directly observable in transaction data, forecast improvement serves as a proxy validation metric for demand-transfer estimation.

In this paper, we introduced and described our Restricted Logit Model pipeline which is a modification of the original MNL model proposed by Abdallah and Vulcano [3]. Our Restricted Logit Model pipeline is created to solve the problem of estimating Demand Transfer at scale when used in conjunction with the proposed substitution framework. Our method is particularly useful in cases: 1) Where it is known that the underlying need state of the items in the universe can differ. 2) Where item need-states are difficult to characterize. 3) Need states do not correspond perfectly with a set of disjoint partitions of the item universe. 4) Where the item universe is huge having more than millions of items. To the best of our knowledge, few existing choice-model-based approaches simultaneously provide explicit demand-transfer coefficients, support overlapping need states, and remain computationally tractable for item universes of the scale considered in this work. While the proposed framework has promising results, several opportunities remain for future research. In particular, relaxing the Independence of Irrelevant Alternatives (IIA) assumption through more flexible choice models, such as Nested Logit or Mixed Logit formulations, may better capture substitution behavior while preserving scalability. Further benchmarking against alternative demand transfer methods, including unrestricted MNL models, Markov Chain choice models, and modern machine learning approaches, as well as component-level ablation studies, would help provide a more comprehensive evaluation of the framework’s effectiveness.

Improvements for AI systems

Improvements to AI Systems:

  1. Scalable Demand Transfer Estimation for Large Item Universes
  • Implement a Restricted Logit Model (a modified MNL) that estimates explicit demand-transfer coefficients (ρ̂ij) between items without requiring separate forecasts for every possible assortment.

  • The system can handle item universes with 1 million+ items by leveraging category-level utility estimation and substitution scores, avoiding the computational explosion of 2 U assortment scenarios.

  • It outputs explicit transfer probabilities (e.g., ρ̂ij = exp(η̂j) / Σ k∈Si exp(η̂k)) that can be directly used in downstream optimization (e.g., shelf-space allocation) without re-running forecasts.

  1. Substitution-Aware Demand Forecasting with Cold-Start Handling
  • Integrate a hybrid substitution-score framework:

  • Use Store Yule’s Q (YQ) scores from historical transaction data to identify substitutable item pairs (threshold > 0.6), filtering out complements via a Basket Odds Ratio adjustment.

  • For new/low-selling items lacking transaction history, use SBERT-based semantic similarity on product attributes (descriptions, brand, price) to infer substitute relationships.

  • The improved system can adjust independent demand forecasts by adding transferred demand from unavailable items (e.g., D̃i = D̂i + ρji·D̂j + ρki·D̂k), reducing forecast error in real-world retail settings.

  1. Overlapping Need-State Modeling without Disjoint Partitions
  • Replace rigid item-hierarchy-based need-state segmentation with a flexible approach that uses substitution scores to define overlapping need states.

  • The system can estimate MNL utilities at the category level, then apply the Independence of Irrelevant Alternatives (IIA) property to restrict inference to only highly substitutable items, yielding valid probability distributions for demand transfer even when need states overlap.

  • This enables accurate DT coefficients in scenarios where items serve multiple customer intents (e.g., a snack that substitutes for both chips and cookies).

  1. Parallelized, Production-Ready Implementation
  • Deploy the MNL fitting algorithm (minorization-maximization) parallelized across categories using distributed computing frameworks (e.g., Spark, internal distributed systems).

  • The system can process historical transaction data (e.g., one year) for millions of items across thousands of categories, producing DT coefficients in a time-efficient manner suitable for large-scale retail operations.

  1. Proxy Validation via Forecast Error Reduction
  • Use offline backtesting with WMAPE (weighted mean absolute percentage error) to validate DT estimates: compare raw forecast error vs. adjusted forecast error (after applying DT corrections).

  • The system can automatically identify locations/categories where demand transfer is most impactful (higher baseline error) and prioritize DT correction there, improving overall forecast accuracy.

  1. Future Extensibility to More Flexible Choice Models
  • The framework is designed to be extended beyond MNL: relax IIA by integrating Nested Logit or Mixed Logit models while preserving scalability.

  • This would allow the AI system to capture more complex substitution patterns (e.g., correlation between similar brands) without sacrificing the ability to handle large item universes.

What the Improved AI System Can Do:

  • Accurately forecast item demand under any assortment scenario (including item removals) by explicitly modeling demand transfer between substitutes, without needing to simulate all possible assortments.

  • Optimize both which items to stock and how many units of each (shelf-space allocation) using explicit DT coefficients, improving revenue and reducing stockouts/overstocks.

  • Handle cold-start items (new products) by leveraging semantic similarity to infer substitution relationships, enabling DT estimation from day one.

  • Scale to retail operations with millions of SKUs and thousands of categories, processing historical transaction data in a distributed, parallelized manner.

  • Provide interpretable, probabilistic DT coefficients that can be integrated into existing demand forecasting and inventory optimization pipelines.

Sources

Related papers