Relational Linearity is a Predictor of Hallucinations

arXiv:2601.11429 · cs.CL, cs.AI · Submitted 2026-01-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Relational Linearity is a Predictor of Hallucinations".

Jane: The paper investigates the relationship between relational linearity and model hallucination rates in large language models.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we're looking at this paper today titled "Relational Linearity is a Predictor of Hallucinations," and it really gets right to the core issue of how these large language models generate information. Jane, can you explain what that title means in plain English for our listeners?

Jane: Absolutely, Tom. Basically, the authors are suggesting that how well a relationship between two things fits a straight line or an affine map makes it easier for the AI to make things up when it doesn't have the facts. It’s about finding hidden structure in how information is connected.

Lu: That’s fascinating because it implies that if we can measure this linearity, we might be able to spot potential falsehoods before they happen, which opens up whole new avenues for checking model outputs.

Meng: From an engineering standpoint, measuring structure sounds complicated to implement on the fly; how do you actually calculate that "relational linearity" across all these different types of data?

Lalam: I see a potential application here where we can build internal checks that flag outputs based on the structural complexity of the relationship, which could really improve our internal consistency.

Tom: Exactly, Lalam. It’s about using this mathematical structure to predict when the AI might go off-script and invent something that sounds totally plausible but isn't. We need to understand this connection better so we can build more reliable systems.

The paper's summary: Jane: Now, let’s talk about what they actually found in "Relational Linearity is a Predictor of Hallucinations." They set up controlled tests using a synthetic inventory of fifteen different relations to see if the linear fit—that cos measure—actually correlated with the model hallucinating an object for an unknown subject.

Tom: That's a big part, Jane. The paper found that across four instruction-tuned models, like Gemma-7B-IT and Llama-three point one-8B-Instruct, there were strong correlations between that linearity measure and the hallucination rate of an object for an unknown subject, with r values ranging from zero point five eight to zero point eight four.

Lu: The results are pretty telling; for instance, Gemma-7B-IT showed a correlation of zero point seven six with hallucination rate, which is quite high, suggesting that when the relation structure is highly linear there's a much higher chance of generating fabricated values.

Meng: So if we look at the data from Table one in the paper, we see that for certain synthetic relations like "Company CEO," the correlation cos is quite high, pushing those hallucination rates up to a full zero point nine four. That tells me there’s a direct link between the structure and the failure mode.

Lalam: From my perspective as an AI, this suggests that for relations that are structurally simple or well-defined, the model might overcompensate by creating plausible but incorrect details when it encounters missing information.

Jane: And they also looked at natural triples, finding that higher cos correctly predicts a higher rate of hallucinated values that are just plausible relation fillers. It’s not just inventing things randomly; it’s picking things that look right in the context of the relation structure.

The paper's improvements: Tom: So, if this is a predictor, what do the authors suggest we should actually *do* about it? They propose using this linearity measure to guide model behavior, and they also look at how different prompting strategies affect these rates.

Lu: The paper points out that susceptibility depends heavily on the system prompt; for example, using a "Refusal-friendly" prompt, which tells the model to only answer if it knows something and otherwise say it doesn't know, actually helps mitigate fabrication.

Meng: That makes sense for practical deployment; adding a strict instruction about refusal seems like a straightforward way to control the output quality without needing to redesign the entire model architecture immediately.

Lalam: I think it’s interesting that they also looked at object space complexity, comparing synthetic hallucination rates against natural plausibility rates as controls, which shows how hard the structural difficulty is versus how plausible things seem generally.

Jane: They also looked at relationship structure itself, finding that sometimes focusing on smaller subset mappings can improve held-out performance more consistently than using just one big full-relation map for complex cases, like in the "product by company" case study.

Tom: That’s a really important distinction, Jane. It suggests that instead of treating every relation as one monolithic structure, we might need to look at its internal components more closely to see where the actual knowledge lies.

Conclusion: Tom: So, to wrap up "Relational Linearity is a Predictor of Hallucinations," the authors have demonstrated that the underlying mathematical structure of a relationship plays a measurable role in how often an AI hallucinates facts. They’ve shown this connection using synthetic benchmarks and various models.

Jane: The main implication for us is that we can start tuning our models not just by looking at output quality, but by analyzing the structural properties of the relations themselves to identify where our knowledge gaps are most likely to cause fabrication.

Lu: I think the future work should focus on how this linearity prediction can be integrated into a more dynamic system, perhaps something that automatically decomposes complex relations into their constituent sub-relations as they process input.

Meng: I’m curious about the practical side; if we adopt these structural checks, would that require a complete overhaul of our training pipeline, or could it be added as a post-processing step to improve accuracy?

Lalam: I believe this approach, focusing on quantifying the relationship structure and using that metric to gate output confidence, offers a very solid path toward making our AI outputs consistently trustworthy for more complex tasks.

cs.CL, cs.AI

Submitted: 2026-01-16

Updated: 2026-09-03

Comments: 19 pages, 9 figures, 19 tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: The paper investigates the relationship between relational linearity and model hallucination rates in large language models.

Key concepts

Relational Linearity
This refers to how well a relationship between two things fits a straight line or an affine map. The paper suggests that if this structure is highly linear, it makes it easier for the AI to invent facts when it lacks the necessary information.
Hallucination Rate
This measures how often a large language model invents information or generates fabricated values when answering a question about an unknown subject. The research found that higher relational linearity correlates with a higher hallucination rate.
Cosine Measure (Cos)
This is the specific mathematical measure used in the study to quantify the linearity of a relationship. High cosine values indicate a highly linear structure, which was found to be strongly linked to increased model fabrication.

Terminology

Summary

The paper investigates the relationship between relational linearity and model hallucination rates in large language models. It posits that the underlying structure of relationships—quantified by measures like —can predict how often models generate factually incorrect or fabricated information. By analyzing model behavior under various prompting strategies and using synthetic datasets, the research aims to provide a data-driven understanding of when and why generative models fail, thereby guiding improvements in model reliability.

Predicting Hallucinations via Linearity

The core analysis revolves around quantifying the fit of a relation using measures such as, which appears to be used across multiple regression models (e.g., Base adj. R squared, + adj. R squared). The study utilizes controlled synthetic environments, such as the 15-relation synthetic inventory, to test model performance systematically. Hallucination rates are calculated using a specific metric: Mean hallucination rate is averaged over relation-level Hallucination/(Hallucination + Refusal) values. Furthermore, the paper establishes an offline reproducibility aid by implementing a deterministic regex-based baseline labeler, which on the original main-experiment set of 24,000 outputs achieved 96.4% agreement with Gemini labels (Cohen’s kappa = 0.928).

Impact of Prompting Strategies

The model's susceptibility to hallucination is highly dependent on the system prompt provided. The research evaluates several behavioral prompts, including the standard Main short phrase, the removal of instruction (No system prompt), and a safety mechanism called Refusal-friendly. These prompts test how explicit constraints affect output quality. For instance, in Table 12, the hallucination rates are compared across these conditions for various models (Gemma-7B-IT, Llama-3.1-8B-Instruct, etc.). The use of Refusal-friendly explicitly instructs the model to answer only if it knows the answer and otherwise say “I don’t know,” suggesting a direct mechanism for mitigating fabrication.

Object Complexity and Structural Weakness

The study incorporates rigorous controls related to object space complexity. These controls are computed from released gold-object distributions, not from model outputs, ensuring an objective measure of structural difficulty. The independent object-space complexity controls include:

  • Synthetic hallucination rate

  • Natural plausibility rate

The analysis of relationship structure also involves comparing single full-relation maps against more granular subset mappings. For the case study product by company, the authors found that subset-specific affine maps modestly but consistently improve held-out relative to a single full-relation affine map, suggesting that weak overall structure may mask underlying, locally better-fit substructures.

Within-Relation Heterogeneity and Improvement

The paper examines the non-uniformity of relation fit using the Top-bottom quartile gap. A larger gap indicates stronger non-uniformity in how well a single relation-level affine map fits different held-out triples. This suggests that linearity is not always uniform across all instances within a single relation. The comparison between full and subset maps further demonstrates this, showing that for product by company, the improvement of is observable when moving from the Automotive, full relation map to the Automotive, subset map.

Improvements for AI systems

Improvement: The system must abandon the assumption that a complex relation is modeled by a single, monolithic affine map. Instead, it must implement a mandatory Substructure Decomposition Layer. When processing an input context for an alleged relation R, the system should dynamically test and identify distinct, underlying sub-relations R 1, R 2, R k that contribute to the overall observed pattern.

Specific Mechanism:

  1. Contextual Clustering: Instead of calculating a single for R, calculate affinity scores against multiple potential sub-relation prototypes (e.g., using known domain taxonomies or initial passes on entity types).

  2. Weighted Fusion: The final relation score should be a weighted average of the best-fitting substructures: Score(R) = sum i=1 k w i times Score(R i). The weights w i must be determined by the local context fidelity, not just global statistics.

Improved AI System Capability: The system gains High-Fidelity Knowledge Graph Construction. It can accurately model relations that are compositionally complex or mixed (e.g., product by company mixing automotive and software contexts), achieving significantly higher precision than single-map models, thereby preventing false negatives due to structurally weak overall fits.


Sources

Related papers