Auditing a KB Elicitation of Frontier LLM Knowledge: A Multi-dimensional Analysis of GPTKB v1.5
cs.CL, cs.AI
Submitted: 2025-10-08
Updated: 2026-09-18
Comments: Accepted at AKBC@EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLMs are remarkable artifacts that have revolutionized a range of knowledge-intensive tasks.
Terminology
Abstract
LLMs are remarkable artifacts that have revolutionized a range of knowledge-intensive tasks. A significant contributor is their factual knowledge, which, to date, remains poorly understood, and is usually analyzed from biased samples. In this paper, we provide a framework and the results of a multi-dimensional analysis of GPTKB v1.5 (Hu et al., 2025a), a recursively elicited Knowledge Base (KB) of 100 million facts (or beliefs) of a frontier LLM, namely, GPT-4.1. Given the scale of the elicited facts, we provide a multi-dimensional approach to qualitatively and quantitatively analyze these facts as opposed to the mainstream fact completion benchmarks, which are prone to availability bias. We find that the models' factual knowledge differs quite significantly from established knowledge bases, and that its accuracy is significantly lower than indicated by previous benchmarks. We also find that inconsistency, ambiguity and hallucinations are major issues, shedding light on future research opportunities in neuro-symbolic AI concerning extraction, consolidation and verification of factual LLM knowledge.
Sources
- GPTKB v1.5: A Massive Knowledge Base for Exploring Factual LLM Knowledge
- Exploring Wikipedia Gender Diversity Over Time $\unicode{x2013}$ The Wikipedia Gender Dashboard (WGD)
- How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering