Improving Forecasts of Suicide Attempts for Patients with Little Data
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Improving Forecasts of Suicide Attempts for Patients with Little Data".
Jane: Ecological Momentary Assessment (EMA) studies provide real-time data on suicidal thoughts and behaviors, but predicting suicide attempts remains challenging due to their rarity and patient heterogeneity.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So folks, we're talking about this paper today titled "Improving Forecasts of Suicide Attempts for Patients with Little Data," which tackles the tough problem of predicting suicide attempts when you have scarce data from Ecological Momentary Assessment studies. Jane, can you give us the quick rundown on what this research is all about?
Jane: Absolutely, Tom. This paper addresses a huge hurdle in forecasting suicide attempts, which is that these events happen so rarely and patients are so different that standard single models just don't work well across everyone. The authors show that while individual models tailored to each patient perform better than one big model for all patients, those individual models end up overfitting when they only have limited data to work with.
Lu: That heterogeneity is the core issue, isn't it? It suggests that we can't rely on a one-size-fits-all approach for these complex behaviors, which opens up some really interesting avenues for how we model patient risk.
Meng: From an engineering standpoint, I wonder if this means we need incredibly robust ways to handle sparse data inputs before any prediction even starts.
Lalam: If this paper works, it means our AI can start learning subtle patterns that are specific to individuals without needing massive datasets for every single person, which is a big step for personal care applications.
Tom: Exactly, Lalam. The paper introduces Latent Similarity Gaussian Processes, or LSGPs, as the solution to capture that patient heterogeneity and let those with less data benefit from the trends of similar patients. It claims this approach allows patients with limited data to intelligently draw on similar patients’ trends when making forecasts.
Jane: So instead of trying to train one model for everyone, they create a latent space where patients are mapped, and distance in that space corresponds directly to how similar their forecasting trends are. This lets the model use those similarities to make smarter predictions even when data is sparse.
Lu: That concept of a latent similarity space is fascinating because it moves us from modeling every patient as an isolated point to understanding them within a structure where relatedness matters for prediction accuracy. It really opens up possibilities for how we categorize these risk profiles in a more meaningful way than simple demographics alone.
Paper summary: Meng: But if we're mapping patients into this latent space, how do you actually define that similarity metric mathematically so the model doesn't just pick arbitrary connections? I mean, that's where the engineering gets tricky.
Lalam: Well, they formalize it by modeling each patient’s latent variable as a multivariate normal distribution and defining how observations relate to those latent variables within a specific likelihood function. This gives them a solid mathematical foundation for connecting those patients.
Tom: It sounds like they are setting up this sophisticated structure to manage the complexity of patient differences, and they’re using some advanced inference techniques to make it computationally feasible given the large amount of data involved in their experiments, which is quite impressive.
Jane: They also tackle the difficulty of analytical inference because the likelihood function isn't perfectly Gaussian, so they use Sparse Variational LSGPs and Stochastic Variational Inference to approximate what a full analysis would look like. It’s a smart move to make this complex idea actually runnable on real data.
Lu: The SV-LSGPs formulation using inducing points and SVI seems like a practical way to manage the computational load when dealing with thousands of observations across many patients, which is crucial for scaling up this idea. It shows how theoretical elegance can be translated into an efficient computational framework.
Meng: So, they’ve shown that their SV-LSGP approach outperforms several baselines, even without extensive kernel design or hyperparameter searching, which suggests a certain level of robustness in their setup. That's something practical for deployment.
Lalam: I think what’s really exciting here is the discovery about patient similarity structure they found when exploring different demographic groups; it suggested that random groupings actually boosted performance more than grouping by demographics, implying these factors aren't the primary drivers of risk clustering in this context.
Tom: That finding about modularity being close to zero across different graphs is telling, Jane. It suggests that the underlying similarity structure isn't neatly explained by standard demographic labels, which is a big piece of information for future research.
Jane: Right, and it ties back to the main point: even with this understanding of similarity, they still have that limitation where idiographic models can't be used to predict outcomes for a completely new patient because you need that individual patient-specific data to define their location in the latent space.
Paper summary: Lu: That inability of idiographic models to predict for new patients is a key limitation they explicitly address in Section three which is important context when we talk about deploying this system into actual clinical settings where you're constantly encountering novel cases <ref:2511.18199#pg2>.
Meng: So, if we want to take this forward practically, the next big hurdle for us would be developing a way to robustly define those patient similarity kernels so that the mapping into the latent space is as accurate and useful as possible.
Lalam: I think the long-term implication here is that we can move toward personalized risk assessment systems where even if a patient's data stream is small, they still get an informed forecast based on what similar individuals have experienced. This could genuinely help tailor preventative strategies.
Tom: That’s a powerful vision, Lalam. So, to wrap up this paper on "Improving Forecasts of Suicide Attempts for Patients with Little Data," the main idea is that by using Latent Similarity Gaussian Processes, we can model patient heterogeneity in a way that lets patients with little data tap into the trends of similar patients.
Jane: And the conclusion they draw is that this approach shows promise because it outperforms all but one baseline and offers a new understanding of how patients are similar, even without heavy kernel design. It’s about using structure to solve the scarcity problem in these critical predictions.
Lu: The paper demonstrates that while single models fail when applied broadly, individualized models run into overfitting issues unless you have enough data per patient, which LSGPs help mitigate by leveraging latent similarity.
Meng: From an implementation perspective, the use of Sparse Variational LSGPs provides a computationally tractable way to make these complex models work with the scale of real-world EMA data.
Lalam: And ultimately, this work suggests that we can build more resilient AI systems for sensitive areas by focusing on patient relationships rather than just isolated data points.
Tom: We've got a lot to unpack here about how mathematical modeling can actually help us handle the messy reality of human behavior in these situations. That’s what we’re talking about today as we wrap up our discussion on "Improving Forecasts of Suicide Attempts for Patients with Little Data."
Conclusion: Tom: So we've been digging into how this new paper tackles predicting suicide attempts using Latent Similarity Gaussian Processes, and now it's time to look at what this whole thing really means for the listeners.
Jane: It’s true that the paper focuses on overcoming the challenge of patient data scarcity by using similarity metrics to predict risk when individual data points are thin.
Lu: The authors, who have deep roots in statistical modeling, put together a really clever way to map patients into a latent space where their forecasting tendencies cluster together naturally.
Meng: It’s interesting because this moves the prediction away from needing massive datasets for every single person, which is a big practical consideration for any real-world application.
Lalam: For me, the most exciting implication is that we can start building systems that offer personalized support based on patterns from people who are structurally similar to them, even with limited initial input.
Tom: Exactly, Lalam. It’s not just about better math; it’s about making the prediction process accessible to patients who aren't well-represented in large datasets.
Jane: The title itself, "Improving Forecasts of Suicide Attempts for Patients with Little Data," really highlights the paper's core contribution: tackling a very difficult problem with limited information.
Lu: It suggests that patient heterogeneity, which we thought was an insurmountable barrier, can actually be leveraged as a source of predictive power instead of just being seen as noise.
Meng: I see it as a way to build more robust AI for sensitive health prediction where the data streams are inherently messy and intermittent, which is how most real-life monitoring happens.
Lalam: That capability to draw on the patterns of similar individuals offers a pathway toward preventative care that feels much more tailored than current general models.
Tom: It opens up a whole new way to think about risk assessment, moving past the limitations of those single models we discussed earlier.
Jane: So, this paper isn't just an academic exercise; it suggests a tangible path toward creating more nuanced and helpful predictive tools for vulnerable populations.
Lu: And the methodology they used, incorporating Sparse Variational LSGPs with Stochastic Variational Inference to handle the complexity, gives us a really strong blueprint for how to build these kinds of scalable models.
Meng: That blueprint is what matters; it shows us how to keep the computation manageable while still capturing that necessary patient-specific nuance.
Lalam: Moving forward, I see this as a chance for AI to become a tool that supports more intimate, data-informed interventions in mental health settings.
Tom: Absolutely, Lalam. We’re going to keep exploring how this structure of latent similarity can be applied in the next segment as we look at the experimental findings on patient similarity structure.
Genesis Hang, Annie Chen, Hope Neveux, Matthew K. Nock, Yaniv Yacoby
Wellesley College · Harvard University
stat.ML, cs.LG
Submitted: 2025-11-22
Updated: 2026-10-05
Comments: Accepted at the TS4H Workshop at NeurIPS 2025
Code: https://github.com/AmazaspShumik/sklearn-bayes
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 79/100
The gist: Ecological Momentary Assessment (EMA) studies provide real-time data on suicidal thoughts and behaviors, but predicting suicide attempts remains challenging due to their rarity and patient
Key concepts
- Patient Heterogeneity
- Suicidal risk varies significantly between individuals; there are many subtypes of at-risk patients. A single model fails because each patient's path to risk is unique and conflicts with others, requiring individualized approaches.
- Latent Similarity Gaussian Processes (LSGPs)
- This model places patients in a hidden 'latent space.' In this space, the distance between patients reflects how similar their forecasting trends are. This allows researchers to infer a patient's location based on similar individuals, helping those with little data leverage collective trends.
- Sparse Variational LSGPs (SV-LSGPs)
- Since exact mathematical calculations are too complex for the large dataset, this technique uses an approximation. It introduces 'inducing points' to summarize the data efficiently, making the complex model computationally feasible and allowing for faster learning.
Terminology
Summary
Ecological Momentary Assessment (EMA) studies provide real-time data on suicidal thoughts and behaviors, but predicting suicide attempts remains challenging due to their rarity and patient heterogeneity. The central finding is that single models perform poorly when applied across all patients, while individualized models overfit for patients with limited data; this paper introduces Latent Similarity Gaussian Processes (LSGPs) to capture patient heterogeneity, allowing those with little data to leverage trends from similar patients.
The Challenge of Patient Heterogeneity and Data Scarcity
The primary difficulty in forecasting suicide attempts lies in the low base-rate of these events and the inherent heterogeneity among patients. Prior work on suicidal ideation suggests that patients’ paths to suicide ideation are heterogeneous, suggesting that, at the very least, there are many subtypes of at-risk patients,
which necessitates moving away from single models. The authors demonstrate that a single model trained on data to predict suicide attempts from all patients performs worse than individualized, per-patient models.
This is because each patient exhibits a different forecasting trend, that, when combined with one another, conflict with one another.
Consequently, while idiographic models outperform single models in many metrics (e.g., ROC-AUC), they are prone to severe overfitting for patients with little data,
as they require us to collect enough data per-patients to make informed forecasts.
The Latent Similarity Gaussian Processes (LSGP) Model
To address the need for capturing patient heterogeneity, the authors formalize their observations into a single model using Latent Similarity Gaussian Processes (LSGPs). The LSGP posits that patients lie in a latent space in which distance corresponds to similarity in forecasting trends.
By inferring patients’ locations in this latent space, the model enables those with little data to intelligently draw on trends from similar patients.
The mathematical formulation involves:
-
A distribution over latent variables: A patient's latent variable is modeled as a multivariate normal distribution:
A zn ∼ p(z) = N (0,IDz).
-
The likelihood function for the outcome: The probability of an event is modeled via a sigmoid function applied to the linear combination of inputs and latent variables:
FXb; θ ∼ N (0, Kθ(X, b Xb)), D yifi ∼ Bernoulli(sigmoid(fi)).
-
The input to the likelihood: The observation vector includes the patient's latent variable:
xˆi represents the concatenation of the inputs xi with the latent variable zni corresponding to patient ni.
Model Implementation and Inference Techniques
The LSGP model is implemented using several advanced techniques to handle its complexity, especially given that analytical inference is impossible due to the non-Gaussianity of the likelihood and the large number of observations (14763 from N = 77 patients).
The authors employ Sparse Variational LSGPs (SV-LSGPs) by replacing the exact likelihood with a variational approximation. This involves:
-
Defining an inducing point formulation: They introduce a matrix of inducing points, W, to
summarize the training data,
which enables more efficient inference. -
Using Stochastic Variational Inference (SVI): They learn the parameters of this approximation by minimizing the divergence between an approximate and true posterior using the Evidence Lower Bound (ELBO). This process is computationally tractable because
the first term of L can be estimated via mini-matching with just O(M3) per gradient step.
-
Computing Latent Similarity: To visualize patient similarity, they treat the LSGP's covariance as a graph and compute the covariance between patients by applying the latent-space kernel, Kzθ, to the means of q(zn; θ). This covariance matrix is then used as an
adjacency matrix
whereedge weights equal the covariance between the patients.
Experimental Findings on Patient Similarity
The experiments reveal significant insights into patient similarity and model performance. Key findings include:
-
Model Performance: The SV-LSGP shows promise, as it
outperforms all baselines except for VB-LR
across various metrics when compared to idiographic models. Furthermore, the approach shows thatwithout significant kernel design or hyperparameter search, our approach already matches the best-performing baseline nearly all metrics.
-
Patient Similarity Structure: When exploring similarity graphs across different demographic groups, the authors find that
modularity for all graphs is close to 0, indicating balanced similarity within/between groups.
This suggests thatthese demographic factors do not explain patient similarity,
aligning with the observation thatrandom groupings outperform demographic groupings.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on this research, and what those improved systems could do:
) System Improvement: Transition from Single-Model Forecasting to Latent Similarity Gaussian Processes (LSGP) for Personalized Prediction.
The core improvement is moving away from single, monolithic models trained on all patients toward an individualized approach that leverages patient similarity in a learned latent space. This directly addresses the failure of single models to capture patient heterogeneity and the overfitting issue in per-patient models with limited data.
-
The LSGP model will be used to map patients into a continuous latent space where proximity reflects similarity in suicidal risk trajectories (forecasting trends).
-
For a new patient with little data, the system will not rely solely on their own sparse history but will intelligently
borrow
predictive trends fromneighboring
patients in the latent space (similar patients), leading to more robust forecasts than single or idiographic models.
) System Improvement: Implementation of Graph-Based Visualization for Risk Profiling and Mechanism Discovery.
The system will incorporate a graph structure where nodes represent patients and edge weights represent the covariance (similarity) between them in the learned latent space. This moves beyond simple prediction to interpretability.
-
The system will generate dynamic similarity graphs, allowing researchers to visualize how patient risk profiles cluster based on underlying data characteristics (e.g., demographics, clinical features).
-
By analyzing graph modularity metrics (as shown in Section 4), the system can identify whether patient similarity is driven by shared demographic factors or inherent behavioral mechanisms, helping to uncover subtypes of at-risk patients that might be missed by purely predictive models.
) System Improvement: Development of a Kernel Design Strategy for Enhanced Generalization.
The current preliminary results suggest the LSGP performs well without significant kernel design, but the paper explicitly states that future work should investigate inductive biases of different kernels to potentially outperform all baselines and generalize better.
-
The system architecture will be designed to allow for systematic exploration and optimization of kernel functions (e.g., state-dependent linear kernels defined in Equation 8) based on patient characteristics or latent space positions, aiming to tailor the model's sensitivity to different risk profiles.
-
This allows the system to move beyond a
one-size-fits-all
latent space assumption toward one that is dynamically optimized for specific subgroups of patients.
) System Improvement: Robust Inference via Sparse Variational Techniques (SV-LSGP).
To make the complex LSGP model computationally feasible for real-time or near real-time forecasting, the system must utilize the sparse variational formulation with inducing points (Equation 3 and 4) and Stochastic Variational Inference (SVI).
-
The system will be optimized using SVI to efficiently learn the latent structure and kernel parameters without requiring intractable analytical inference.
-
This enables deployment of a scalable model capable of handling large-scale, heterogeneous datasets typical in real-world EMA applications while maintaining high predictive accuracy across patient subgroups (as demonstrated by the superior performance of SV-LSGP in Table 1).
) Improved AI System Capabilities: What the Enhanced System Can Do.
The improved system will evolve from a simple risk score predictor to a sophisticated, personalized clinical decision support tool capable of:
-
Predicting the probability of an imminent Suicide-Related Event (SRE) for any individual patient based on their real-time EMA data, even when that patient has very little historical data.
-
Providing an explanation for the forecast by highlighting similar patients in the latent space and indicating which observed trends from those neighbors are influencing the prediction (leveraging the learned similarity graph).
-
Identifying novel, previously unobserved subtypes of suicide risk by analyzing clusters in the patient similarity graph, potentially linking specific behavioral patterns to underlying demographic or clinical features.
-
Guiding personalized intervention strategies by understanding not just
what
is happening now, buthow
this individual patient's trajectory relates to the broader population of at-risk individuals.
Sources
- Gaussian Process Regression with Heteroscedastic or Non-Gaussian Residuals
- Meta Reinforcement Learning with Latent Variable Gaussian Processes
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey