What Makes a Peer? Valuation-Anchored Similarity in Private Markets

arXiv:2608.12594 · q-fin.ST, cs.AI, cs.LG · Submitted 2026-08-12 · Read on arXiv

Sebastian Frank, Jingrao Lyu, Max Jarmey, Preetha Saha, Mingshu Li, Sweet Kaur, Sola Akinola, Dhagash Mehta

BlackRock, Inc.

q-fin.ST, cs.AI, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: The paper introduces a supervised similarity learning framework for identifying economically meaningful peer companies in private markets, where limited transparency, sparse disclosures, and

Terminology

Summary

The paper introduces a supervised similarity learning framework for identifying economically meaningful peer companies in private markets, where limited transparency, sparse disclosures, and infrequent transactions make traditional peer identification challenging. The authors propose defining company similarity through the lens of market valuation rather than static feature matching or semantic descriptions. Specifically, they train a CatBoost gradient-boosted decision tree model on observed private company valuations and derive a valuation-aware similarity metric from importance-weighted leaf-node co-occurrences across the ensemble.

The framework is trained using post-money valuation as the supervisory target, but its primary objective is not standalone valuation prediction—rather, it is the extraction of a valuation-conditioned similarity structure among private companies. The authors formalize company similarity as a supervised learning problem, where similarity is learned from observed market valuations rather than predefined distance functions, industry classifications, or textual descriptions. By anchoring similarity to valuation outcomes, the approach generates peer relationships directly aligned with private market pricing dynamics.

The dataset comprises approximately 270,000 private companies globally, spanning more than 50 industry groups, hundreds of product and service categories, and multiple geographic dimensions. Deal history extends from 1977 to 2025, with the median deal having an age of five years as of calibration (2025), and 75% of deals having taken place since 2015. Approximately 53,000 firms have an observed or derivable post-money valuation and at least one PE or VC transaction, forming the calibration universe. The final feature set contains 27 variables, of which 23 are categorical, reflecting the inherently taxonomic nature of private market data. Financial variables such as revenue, EBITDA, and net income exhibit the highest levels of missing values (75–80%), whereas categorical attributes including country, deal stage, and customer type are largely complete.

The similarity metric is constructed as follows: for companies X1 and X2, let Z1,t and Z2,t denote the leaf nodes to which they fall while traversing an assigned tree t in the ensemble. The tree-weighted leaf-node co-occurrence similarity function S(X1, X2) is formulated as the sum over all trees of the tree weight wt multiplied by an indicator function equal to 1 when the two companies occupy the same terminal node and 0 otherwise. Tree weights are assigned based on the incremental reduction in training loss, with each tree's step-size predictive contribution St calculated via the absolute difference in sequential training loss, then normalized to obtain wt. The final bounded supervised dissimilarity measure D(X1, X2) is derived directly from the complement of the similarity matrix score: D(X1, X2) = 1 − S(X1, X2), mapping the pairwise relationship into a symmetric matrix where D ∈ [0, 1].

The model is trained to minimize root mean squared error in log space, with the target variable log-transformed and winsorized at the 1st and 99th percentiles. The dataset is divided into training and test sets using an 80:20 split stratified by internal subindustry classification. Recency-based weighting is applied to training samples to emphasize more recent transactions. Hyperparameters are optimized using Optuna with tree-structured Parzen Estimator sampling over 50 trials using five-fold cross-validation. The optimal configuration includes a tree depth of 13, 1,859 estimators, 108 max leaves per tree, 242 min data in leaf, and a learning rate of 0.023.

The CatBoost model outperforms a multivariate linear regression baseline, achieving a mean absolute error of 1.08 versus 1.19, a root mean squared error of 1.44 versus 1.56, an R2 of 0.46 versus 0.36, and a mean absolute percentage error of 6% versus 7%. Prediction errors are concentrated in the lowest and highest valuation deciles, with out-of-sample RMSE rising from 0.94 in Q8 to 2.35 in Q1 and 2.05 in Q10. Higher prediction errors are observed in cybersecurity and selected software segments, and for later-stage private equity transactions and deals classified as unspecified, whereas early-stage VC financings tend to exhibit lower errors.

To correct for retransformation bias when exponentiating log-space predictions, multiplicative calibration factors are applied by deal type: 1.74 for Series transactions, 5.06 for Unspecified transactions, and 2.19 for all other transactions. Conformal prediction is used to construct distribution-free prediction intervals from calibration residuals, normalized and grouped by deal type to account for heteroskedasticity. At a 95% confidence level, coverage is 93% with an interval mean ratio of 3.35; at 75% confidence, coverage is 74% with a ratio of 1.80; at 60% confidence, coverage is 60% with a ratio of 1.60.

SHAP analysis reveals that deal type is the dominant valuation driver, with a mean absolute SHAP contribution of 1.31, followed by country (1.17), region and city (0.28), revenue (0.22), and average time elapsed since most recent financials (0.20). The remaining 22 features contribute a combined 1.62. These top five features play the largest role in shaping both valuation predictions and the induced similarity structure.

Neighborhood analysis evaluates whether the learned metric produces economically coherent peer groups. For the top three numeric features as measured by SHAP (revenue, average time elapsed since last financials, and net income), the learned metric consistently shows lower deviation than cosine distance (between embeddings derived from internal company descriptions) and mixed performance against Gower and Euclidean distances on Financial Services companies. For the top three categorical features (deal type, country, and region and city), the learned metric shows a lower mismatch rate than other metrics.

In k-NN benchmarking, the learned similarity metric is compared against Euclidean, Gower, and embedding-based similarity measures. For each metric, a weighted k-nearest-neighbor estimator is constructed to assess valuation prediction accuracy across multiple neighborhood sizes. Importantly, this evaluation does not reuse CatBoost valuation predictions: the CatBoost model is used only to construct the similarity structure, whereas valuation estimates are generated independently through a non-parametric k-NN procedure operating on the resulting neighborhood graph. In Financial Services, the MAE and RMSE of the post-money valuation resulting from the learned metric are consistently lower than the other metrics across all neighborhood sizes.

The authors conclude that the framework offers a data-driven alternative to comparable-company analysis based on fixed industry, geographic, or deal-stage filters. Rather than matching firms by surface characteristics, it identifies peers that share common valuation drivers. Future work may extend the approach to scalable cross-industry similarity, softer tree-based proximity measures, and richer textual or graph-based representations of private companies.

Improvements for AI systems

Improvements to AI Systems:

  1. Valuation-Conditioned Similarity Learning: Replace static feature-based or text-embedding similarity metrics with a supervised, outcome-driven similarity function derived from gradient-boosted tree leaf co-occurrences. The AI system learns peer groupings directly from market pricing signals, not predefined categories, enabling dynamic, economically meaningful clusters.

  2. Handling Sparse, Heterogeneous Data: Implement robust preprocessing for high-dimensional categorical data (23 of 27 features) and heavy missingness (75–80% for financials). The system uses categorical boosting (CatBoost) to natively handle categorical variables and missing values, reducing imputation bias and improving generalization in low-transparency domains.

  3. Recency-Weighted Training: Incorporate time-decay weighting on training samples to prioritize recent transactions, allowing the AI to adapt to shifting market conditions and avoid overfitting to outdated deal dynamics.

  4. Retransformation Bias Correction: When predicting in log-space, apply multiplicative calibration factors by deal type (e.g., 1.74 for Series, 5.06 for Unspecified) to correct systematic underestimation after exponentiation, improving point-estimate accuracy.

  5. Distribution-Free Uncertainty Quantification: Use conformal prediction with heteroskedasticity-aware normalization (grouped by deal type) to generate valid prediction intervals without assuming error distributions. The system provides calibrated confidence levels (e.g., 93% coverage at 95% confidence) for risk-sensitive applications.

  6. Interpretable Peer Justification: Leverage SHAP-based feature attribution to explain why two companies are peers (e.g., deal type, country, revenue), enabling users to audit and trust the similarity structure. The system can highlight the top valuation drivers per cluster.

  7. Non-Parametric Valuation via Learned Graph: Decouple similarity construction from prediction by using the learned metric only to build a k-NN graph, then perform independent non-parametric valuation (weighted k-NN). This avoids circularity and yields lower MAE/RMSE than Euclidean, Gower, or embedding-based metrics in private market settings.

  8. Scalable Cross-Industry Matching: Extend the framework to handle over 50 industry groups and hundreds of product categories without manual taxonomy mapping, enabling cross-industry peer discovery that traditional filters miss.

What the Improved AI System Can Do:

  • Identify economically coherent peer companies in private markets using valuation-driven similarity, not just industry codes or text descriptions.

  • Generate accurate post-money valuation estimates with calibrated confidence intervals, even when financial data is sparse or missing.

  • Provide explainable peer recommendations, showing which features (e.g., deal stage, geography, revenue) drive the similarity.

  • Adapt to temporal shifts in market pricing by weighting recent deals more heavily.

  • Operate across diverse geographies, industries, and deal types (VC/PE) with minimal manual feature engineering.

  • Produce robust predictions and peer groups for early-stage startups, late-stage private equity, and unspecified transaction types, with error metrics tailored to each segment.

Sources

Related papers