What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization

arXiv:2601.17609 · cs.CL · Submitted 2026-08-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization".

Jane: The paper was written by Mohammad Ghassemi, Sara Rezaeimanesh and Michigan State University from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: So, we’re looking at the title again, "What LLMs Know but Don’t Say," which hints at this internal knowledge base. The authors are proposing a way to use this knowledge without relying on the typical outputs of AI models—that's what "Non-generative" means.

Jane: It's a huge step away from prompting an LLM to just predict an answer; instead, they are extracting structural information from the logits. This is crucial because generating text is inherently variable, so unreliable for precise scientific modeling.

Lu: The implication here is that we can use AI not just as a source of answers, but as a source of reliable constraints or guiding rules for Bayesian statistical models. It’s about using the structure of knowledge rather than the content.

Meng: If we are only looking at the logits, it means this method is deterministic and repeatable, which is exactly what any real-world engineer needs to deploy a model reliably in a high-stakes environment.

Lalam: A deterministic way to know what’s true about a feature's impact—that’s huge for confidence. It moves us away from the "hallucination" problem and toward using AI as a structured knowledge base.

Summary of Paper's Claims: Tom: The paper summarizes how LoID, or Logit-Informed Distributions, works by probing the LLM with specific semantic pairings for each feature. It’s not just asking "what is the impact," but using carefully constructed sentences to see if the model favors a positive or negative effect.

Jane: It measures that preference by looking at how often the LLM predicts specific tokens, like "positive" or "negative." This helps quantify exactly how much of a belief an LLM has about a feature’s influence.

Lu: And the paper claims that this quantification—the strength and reliability of the belief—is what forms Normal priors for the coefficients in regression models. It’s turning raw LLM probability into statistical parameters.

Meng: The core claim is that these generated priors are much more informative than just using standard, uninformative Bayesian priors, especially when we are dealing with out-of-distribution data. That's the central hypothesis we need to test practically.

Lalam: It suggests that the LLM’s internal knowledge about how things usually work—like how age influences income—can be systematically extracted and applied to guide our statistical models toward better outcomes.

Improvements Suggested by the Paper: Tom: So, we've seen how it works; now let's talk about the results. The paper suggests that LoID significantly improves performance when trained on out-of-distribution OOD data. This is a critical improvement in generalizing beyond the training set.

Jane: They found that LoID often outperforms existing methods, like LLMProcesses and AutoElicit, across fifteen different datasets, which are benchmarks in healthcare and finance. The paper shows consistent performance gains compared to the original OOD model.

Lu: The best part of the improvement is that it recovers up to fifty percent of the performance gap relative to a perfect oracle model. That means we’re getting much closer to peak performance than we thought possible with current methods.

Meng: I found that especially impressive; achieving that much of the oracle's potential without actually having access to the full, true distribution is a massive win for practical deployment under uncertainty.

Lalam: It implies that by incorporating this external domain knowledge, we are making our AI models far more robust and dependable when they will inevitably face real-world data shift.

Conclusion: Tom: We’ve covered a lot of ground today, from the technical details of LoID to its impressive results on out-of-distribution datasets. It really shows that LLMs are much more than just chat bots; they contain structured, useful knowledge.

Jane: I think the major takeaway is that we can use this internal AI knowledge base—the "what LLMs know but don't say"—to build much more reliable statistical models for real-world applications.

Lu: It’s a paradigm shift, turning latent knowledge into actionable data parameters, and it opens up huge possibilities for the future of predictive modeling.

Meng: My final thought is that this method is efficient, which is vital for practical implementation in any complex system design.

Lalam: To summarize our discussion on "What LLMs Know but Don’t Say: Non-generative Prior Extraction for Generalization," we see a future where AI acts as an intelligent guide to improve the reliability of our most critical predictive models.

Mohammad Ghassemi, Sara Rezaeimanesh, Michigan State University

cs.CL

Submitted: 2026-08-22

Updated: 2026-08-25

Importance score: 81/100

The gist: The following is a detailed summary of the scientific paper: Problem and Motivation In domains such as medicine and finance, "large-scale labeled data is costly and often unavailable," leading to

Key concepts

Logits
Logits are the raw output scores of a Language Model before they are converted into probabilities or text. Instead of prompting the LLM for a direct answer, researchers analyze these underlying scores to determine if the model favors a positive or negative effect regarding a specific feature.
LoID (Logit-Informed Distributions)
This is the core method where structural knowledge is extracted from LLM logits. It uses carefully constructed semantic pairings to measure an LLM's internal belief about a feature's influence, allowing researchers to quantify the strength and reliability of that knowledge.
Non-Generative Extraction
A technique where AI is used not to predict or generate text, but its internal structural knowledge is extracted instead. This method is deterministic and repeatable, providing reliable constraints for statistical models rather than relying on variable output.
Out-of-Distribution (OD) Data
This refers to data that differs significantly from the data a model was originally trained on. The paper demonstrates that using the LoID method helps AI models perform much better when encountering this real-world shift or uncertainty.

Terminology

Summary

The following is a detailed summary of the scientific paper:

Problem and Motivation

In domains such as medicine and finance, large-scale labeled data is costly and often unavailable, leading to models trained on small, non-representative samples that struggle to generalize to real-world populations. While Bayesian methods offer a principled way to encode prior knowledge, manually eliciting high-quality priors at scale remains a longstanding challenge. Large language models (LLMs) offer a promising source of domain knowledge. However, existing methods often rely on generative sampling, which introduces variability. The authors identify a gap in current research: few works have systematically explored how to extract priors from the internal structure of LLMs (e.g., hidden states, logits, etc.) for tabular tasks under distribution shift.

Proposed Solution: LoID (Logit-Informed Distributions)

The authors propose LoID, a deterministic method for extracting informative prior distributions by directly accessing token-level predictions from an LLM. Instead of relying on generated text or examples, LoID probes the model’s confidence in opposing semantic directions (positive vs. negative impact) through carefully constructed sentences.

Methodology

LoID is designed to derive "informative priors for Bayesian logistic and linear regression that generalize robustly in low-data settings:

  1. Querying LLM Belief: For each feature f j, the authors construct paired prompts expressing opposing semantic claims about its effect on the target variable. They use the probabilities from the LLM’s softmax over its vocabulary to extract P i+ (probability of positive token) and P i- (probability of negative token).

  2. Logit Preference Score: The logit preference score quantifies this directional belief: logit j(p) = (p i+ over 1-p i+). This score is used to represent the LLM’s normalized preference for a positive versus negative association.

  3. Constructing Priors: The authors model each coefficient beta j with a Gaussian prior: beta j about N(mu j, sigma j 2).

  • Mean (mu j): This is calculated by averaging the logit preference scores across a set of paraphrased prompts T: mu j = 1 over T sum i in T logit j(p).

  • Variance (sigma j squared): This captures the uncertainty in the LLM’s belief: sigma j squared = std(logit 1(p),..., logit T(p)). Higher variance is assigned to features where the LLM shows inconsistent preferences across paraphrases.

  1. Bayesian Inference: Posterior inference is performed using a Laplace approximation, centered at the maximum a posteriori (MAP) estimate, and then locally approximated via a second-order Taylor expansion of the log-posterior.

Evaluation and Results

The method was evaluated on 15 real-world tabular datasets under synthetic out-of-distribution (OOD) settings characterized by covariate shift. The performance was measured using Area Under the Curve (AUC) and Mean Squared Error (MSE).

  • Performance: Across datasets, LoID significantly improves performance over logistic and linear regression models trained on OOD data, recovering up to 50% of the performance gap relative to the oracle model.

  • Success Modes: LoID excels when feature-target relationships are stable and generalizable, such as in Adult Income (83.8% gap closed) or Concrete Strength (93.9% gap).

  • Failure Modes: LoID struggles when relationships are contextual or behavioral, such as in Bank Marketing (-130.6% gap) or Bike Sharing (-1.8% gap, though the latter is near-zero).

Comparison to Baselines

The performance of LoID was compared against four baselines: (1) standard uninformative priors (N(0, 1)), (2) AutoElicit (generative sampling), (3) LLMProcesses (in-context learning/direct prediction), and (4) the oracle benchmark (P IID).

  • LoID vs. LLMProcesses: LoID outperforms LLMProcesses on 12/14 datasets, demonstrating that extract[ing] priors is more effective than direct prediction.

  • LoID vs. AutoElicit: LoID outperforms AutoElicit on 12/14 datasets, showing that its logit-based approach is more reliable than free-form elicitation and provides more calibrated uncertainty.

  • Uninformative Baseline: On 11/15 tasks, LoID substantially outperforms the uninformed N(0, 1) baseline, confirming the value of LLM knowledge.

Conclusion

The authors conclude that LoID is highly effective: "LoID demonstrates that extracting distributional priors is more robust than direct prediction (compared to LLMProcesses) and more reliable than free-form elicitation (compared to AutoElicit), outperforming both baselines on 12 out of 15 datasets. The method's effectiveness relies on the ability to leverage embedded world knowledge" captured in LLMs, closing a significant performance gap when feature-target relationships are semantically stable.

Improvements for AI systems

As a diligent AI researcher, I have analyzed the proposed methodology in this paper. The core contribution—LoID (Logit-Informed Distributions)—is robust and highly valuable for leveraging embedded domain knowledge while mitigating the variability inherent in generative models.

To transition this research into a high-performance, production-ready AI system that surpasses the current benchmarks, I propose the following specific engineering and algorithmic improvements:

We must formalize LoID as a dedicated preprocessing module within a robust Bayesian inference framework.

Implementation:

  • Feature Probing Module (LoID Engine): Implement the core logic for generating N sent paraphrased prompts for each feature f j. The system will query the LLM (e.g., Gemma-2-27B) and extract raw logits, calculating the logit preference score logit j(p) for every token.

  • Prior Synthesis Module: Calculate the prior mean mu j (average of logit scores) and the prior variance sigma j squared (variance across all iterations). This creates a structured set of Gaussian priors N(mu j, sigma j 2) for each coefficient beta j.

  • Integration: Feed these derived priors into the Bayesian inference engine (e.g., PyMC/MCMC), replacing the standard uninformative or correlation-based prior assumptions.

What the System Can Do: The system will automatically inject domain knowledge into its mathematical structure, allowing it to perform robust posterior inference even when training data is limited or biased, significantly improving OOD generalization compared to baseline models.

The current use of standard deviation (sigma j squared = std(logit i(p j))) is a solid measure of epistemic uncertainty, but we can refine this to improve robustness against LLM overconfidence.

The paper identifies failure modes (e.g., Bank Marketing, Combined Cycle Power) where LLM knowledge is either absent or contextually misaligned.

The paper notes the complexity T LoID = O(d times N sent times T LLM). This is manageable, but we can optimize further.

Sources

Related papers