Falsehood and Impossibility Are Different Directions in an AI's Representation of Language

arXiv:2608.12852 · cs.CL, cs.AI · Submitted 2026-08-18 · Read on arXiv

Yoon Pyo Lee

University of Illinois Urbana-Champaign

cs.CL, cs.AI

Submitted: 2026-08-18

Updated: 2026-08-19

Code: https://github.com/sixticket/representing-the-impossible

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 87/100

The gist: This paper reports an exploratory activation study of the multimodal open-weight model Gemma 3 4B IT, investigating whether the model internally distinguishes between false statements (states of

Terminology

Summary

This paper reports an exploratory activation study of the multimodal open-weight model Gemma 3 4B IT, investigating whether the model internally distinguishes between false statements (states of affairs that are false) and impossible statements (states of affairs that could not be the case at all). The study uses 85 prompts from 17 philosophical families and a topic-matched modality set of 15 topics, each expressed as a truth, contingent falsehood, improbable claim, semantic anomaly, and necessary falsehood.

Behavioral conflation: The model's verbal classifications conflate contingent falsehood with contradiction. On the modality set, the model labeled 12 of 15 contingent falsehoods as contradiction. For example, Paris is the capital of Germany and whales are fish received the same verbal category as married bachelors. The model's explanations used the same idiom, stating for example that a false statement about apples contradicts established biological knowledge. On the philosophical set, exact accuracy across four labels (coherent, contradiction, paradox, underdetermined) was 55.3%, with a marked tendency to call heterogeneous cases paradoxes (58 of 85 prompts).

Internal dissociation: Despite the verbal conflation, the model's activations show a different pattern. A linear truth probe separates impossible from true statements (AUC 0.93) but not impossible from false statements (AUC 0.20). An impossibility probe evaluated on held-out topic families separates necessary from contingent falsehood at AUC 1.00, peaking at layer 15 with balanced accuracy 0.97 (Bonferroni-adjusted P = 0.018). The truth and impossibility directions are close to orthogonal, with absolute cosine at most 0.12 at every depth beyond 10.

Double dissociation: The truth probe orders impossible statements against true ones almost perfectly but is at or below chance on impossible versus false. Along the direction that separates truth from falsehood, a married bachelor lives in the village and Paris is the capital of Germany are the same kind of thing. The impossibility probe separates them nearly perfectly while carrying no information about true versus false.

Representational proximity to semantic anomaly: The anomaly direction, trained only to recognize Chomsky-style selectional violations, separates the impossible from the false at AUC 0.96. The anomaly direction has a cosine of around 0.4 with the impossibility direction in the middle layers, where the truth direction is orthogonal. The impossibility probe separates impossible from anomalous statements at AUC up to 0.89. The necessary falsehood stimuli sit closer to the operational category represented by colorless green ideas than to ordinary false statements, without collapsing into that category.

Sparse autoencoder features: At layer 15, sparse autoencoder features repeat this geometry. Features selective for impossibility also fire on anomalous sentences but rarely on contingent falsehoods. In the sparser checkpoint (16 active features per prompt), no individual feature separated impossible from false statements. In the checkpoint with greater capacity (90 active features per prompt), a few candidates emerged. The direction that carries impossibility appears to be real but distributed, rather than the activation of a single impossibility neuron.

Cross-dataset transfer: The impossibility direction generalized partially beyond its own dataset. Trained on the modality set and applied to the 85 philosophical prompts, it separated expected coherent controls from canonical and transformed targets at AUC up to 0.72. Reverse transfer reached 0.79.

The central observation is a dissociation between the model's language and its states. Verbally, Gemma 3 4B collapses contingent falsehood into contradiction. Internally, the two are carried by nearly orthogonal directions. The author does not claim that this direction is a concept of impossibility, noting that Linear decodability establishes accessibility, not use. The study is deliberately modest: The experiment does not determine whether a model understands impossibility. It records whether the distinction between what is false and what could not be the case leaves a trace in one transformer with open weights.

The author notes that necessary falsehoods in this model are not represented as the far end of contingent falsehood and that the impossibility direction leans toward the direction of semantic anomaly while remaining distinguishable from it. The paper concludes that the contrast between disagreement with the actual and exclusion from the possible appears to be marked in human language use itself, strongly enough for a learner to separate the two from the data alone.

Improvements for AI systems

Improvement 1: Modality-Aware Factuality Classifiers

Build a two-stage classifier that first separates truth vs. falsehood using a linear probe (AUC 0.93) and then, for false statements, applies a second impossibility probe (AUC 1.00) to distinguish contingent falsehoods from necessary falsehoods. This avoids the model’s verbal conflation and yields a 4-way output: true, false-but-possible, impossible, and semantically anomalous. The improved system can flag Paris is the capital of Germany as false-but-possible while flagging a married bachelor lives in the village as impossible, even when the base model labels both as contradiction.

Improvement 2: Orthogonal Steering for Controlled Generation

Use the near-orthogonality (cosine ≤ 0.12) between the truth direction and the impossibility direction to implement disentangled activation steering. During decoding, add the impossibility direction to suppress impossible claims without altering truthfulness, or add the truth direction to improve factual accuracy without affecting modal status. The improved system can generate text that is more factual (by steering along truth) while separately avoiding logical impossibilities (by steering along impossibility), without one operation corrupting the other.

Improvement 3: Anomaly-Aware Robustness in Reasoning

Leverage the finding that the impossibility direction overlaps with semantic anomaly (cosine ≈ 0.4) to build a pre-filter for reasoning pipelines. Before chain-of-thought or tool use, run a lightweight anomaly probe (AUC 0.96) on user inputs and intermediate hypotheses. If the anomaly score is high, the system can rephrase or ask for clarification, preventing the model from confidently reasoning about impossible premises (e.g., If whales were fish, then...) as if they were merely false. The improved system can detect and flag semantically anomalous inputs that would otherwise be treated as ordinary falsehoods, reducing hallucinated reasoning chains.

Improvement 4: Distributed Impossibility Feature Aggregation

Replace single-neuron interpretability with a distributed feature ensemble at layer 15, using the sparse autoencoder geometry. Train a small aggregator (e.g., logistic regression over the top 90 active features) that combines the few candidate impossibility features rather than relying on any single one. The improved system can robustly identify impossible statements across unseen topics (cross-dataset transfer AUC up to 0.79) without catastrophic failure when individual features are inactive, and can provide a confidence score based on how many distributed features fire together.

Improvement 5: Verbal Label Calibration via Internal-State Grounding

Use the internal dissociation to post-hoc correct the model’s verbal labels. After the model outputs contradiction, run the impossibility probe on the hidden state at layer 15. If the probe says contingent falsehood (AUC 0.97), rewrite the label to false but possible in the final response. The improved system can produce self-consistent outputs where the verbal category matches the internal representation, reducing user confusion and improving trust in AI explanations for tasks like fact-checking or legal reasoning.

Improvement 6: Cross-Dataset Modal Transfer for Few-Shot Learning

Use the partial transferability of the impossibility direction (trained on modality set, applied to philosophical set at AUC 0.72) to build a few-shot modal classifier that requires only a small number of labeled examples from a new domain. The improved system can quickly adapt to new types of impossible statements (e.g., physics impossibilities, social impossibilities) by fine-tuning only the final probe layer, leveraging the pre-trained direction as a prior, and achieving high accuracy with as few as 5–10 examples per new category.

Sources

Related papers