GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base

arXiv:2608.06992 · cs.CL, cs.AI, cs.DB · Submitted 2026-08-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base".

Jane: The paper was written by Yujia Hu, Tuan-Phong Nguyen and Simon Razniewski from ScaDS.AI Dresden/Leipzig and Technische Universität Dresden and Institute for AI, VNU University of Engineering and Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We just met the paper on the desk, and the title already signals a big promise: a knowledge base pulled from a language model, but with entity identity cleaned up. That's not a small ask.

Jane: The author list matches the ambition. Yujia Hu, Tuan-Phong Nguyen, and Simon Razniewski, split between Dresden and Hanoi. This is a group that has clearly been building toward this for a while.

Tom: The word "disambiguated" is the star of that title. Most LLM knowledge bases treat a string like "Munich" as if it were one thing. The paper wants to separate homonyms and merge synonyms.

Lu: So two different Munichs become two entities, while The Big Apple lands under New York City. Does the demo actually show that happening?

Tom: It does. Every disambiguation step has an evidence panel you can open. You see the candidates considered, and you see the context that decided between them.

Meng: That's what the word "auditing" promises. You can check why a decision was made, rather than taking it on faith.

Lalam: And that's where the big-picture impact lives. If eye knowledge comes with an audit trail, it stops being an oracle and starts being infrastructure.

Jane: The authors also come from a long line of materialization work. This isn't a cold-start experiment; it's the next step after earlier versions.

Tom: Good point. So before we get lost in the infrastructure talk, let's look at what the summary abstract says is actually inside.

Summary: Jane: So from the title we moved to the abstract. The numbers hit first: 38.4 million triples, 1.6 million entities, 207 thousand relations, and 66 thousand classes.

Tom: Those are big numbers, but the more interesting part is the process. It starts from a seed entity and expands recursively, adding facts while the KB is under construction.

Jane: Each new mention gets disambiguated right away. Homonyms get separated, synonyms get merged, and the context of the source triple guides every call.

Lu: That's much smarter than dumping everything and cleaning up later. The context is still warm when a fact appears.

Tom: The pipeline runs through elicitation, named entity recognition, and disambiguation. A description of Budapest tells the model which Budapest it means.

Meng: And because the whole thing is materialized, you can browse entities and click through links. It feels like a real knowledge graph, not a chat window.

Jane: The demo is also live at gptkb.org, and the full KB is downloadable. That makes it a reusable resource, not just a poster.

Tom: The interface supports SPARQL, natural-language questions, and entity linking from user text. They even point to the components: GRASP for questions, LELA for linking.

Lalam: That completeness is what turns an LLM's private memory into public infrastructure. The scale matters less than the ability to query and audit it.

Jane: Okay. So that's the 30,000-foot view. Next we need to ask what actually improved over earlier versions and other LLM-derived KBs.

Improvements: Jane: So the summary gave us the scale. Now let's push on the real improvements over previous systems.

Tom: The paper's biggest move is disambiguation during construction. Older LLM-derived KBs mostly used surface strings as identifiers.

Jane: That sounds abstract, but it means these systems trip on names. Munich the film and Munich the city end up as one row.

Lu: This KB instead uses context from the source triple. The film Munich shows up in a movie fact, so it stays separate.

Meng: And synonyms like The Big Apple get folded into New York City as an alias. That's the kind of consolidation users can actually inspect.

Tom: One number really caught my eye: 36.8 percent of entities here are novel to Wikidata. It's not just copying an existing graph.

Jane: They also measured how clean the disambiguation is. Human judges found 94.5 percent of sampled triples true, and 96 percent of sampled entities verifiable.

Lu: On the disambiguation side, 98 percent of same-label merges were correct, and every same-label split in their 100-case sample was correct.

Meng: The main error mode was false synonym merges. That's useful honesty for the next version.

Jane: And they claim it's the first demo to combine SPARQL, natural-language queries, entity linking, and provenance on an LLM-derived KB.

Lalam: The bigger implication is for eye accountability. If a model's knowledge is materialized and indexed, you can check it like a library instead of interrogating an oracle.

Tom: The paper also emphasizes that resolving entities during construction avoids fragmentation. A post-hoc cleanup would first have to detect the mess, then repair it.

Jane: That ordering is what makes their alias pages work. One entity can carry many surface forms without losing its identity.

Tom: And the paper's opening section sets up that whole argument. Let's walk through the first page and see how they frame the problem.

First Page: Jane: We've been talking about the fix. The first page explains why the fix is necessary.

Tom: It starts with a classic problem: surface strings are bad identifiers. The paper gives two failure directions, homonyms with the same name and synonyms with different names.

Jane: The Munich example makes it concrete.

Europe, hasMajorCity, Munich: points to the city, while

Mathieu Amalric, notableWork, Munich: points to the film.

Lu: And on the synonym side, New York City and The Big Apple describe the same place. A naive KB would keep them apart forever.

Tom: The first page also gives the core observation. The source triple plus the descriptions of existing entities usually provides enough context.

Jane: That's a refreshingly simple engine. No external Wikipedia mapping, no gazetteer. The LLM itself decides whether this mention is new or known.

Meng: That's why the demo feels necessary. When you materialize a KB this way, you need to show the disambiguation trail or nobody will trust the graph.

Tom: The first page emphasizes that disambiguation happens on the fly. That ordering is what makes later queries reliable.

Jane: And the demo's transparency is meant to make that process inspectable. You can see every trajectory from the seed entity, plus the surface forms and candidate matches.

Lu: The authors also note that descriptions matter for interpretability. In their ablation, removing the subject description dropped the true triple rate from 92.8 percent down to 80 percent.

Lalam: This framing is important. The authors are saying LLM knowledge doesn't have to be a black box. It can be turned into something with an audit trail.

Tom: So that's the opening argument. Now it's time to wrap up what this means for the rest of the field.

Conclusion: Jane: So we started with a title promising a disambiguated LLM knowledge base. We end with a demo that really tries to deliver that.

Tom: The KB holds 38.4 million triples and 1.6 million entities, with 207 thousand relations and 66 thousand classes. More importantly, each fact has a trace.

Jane: Their evaluation reports strong precision on triples and entities, and solid split and merge behavior for homonyms and synonyms.

Lu: The residual errors are mostly false synonym merges. That's an honest finding, and it gives future work a clear target.

Meng: The interface makes it hard to ignore: SPARQL, natural language, entity linking, and provenance all in one place.

Lalam: This is the direction eye needs to go. Not bigger prompts, but structured, auditable knowledge that humans can verify.

Tom: Good point. So we'll say goodbye to this paper and get ready for the next one.

Jane: Thanks for listening. Next up, something new on the arXiv desk.

ScaDS.AI Dresden/Leipzig · Technische Universität Dresden · Institute for AI, VNU University of Engineering and Technology

cs.CL, cs.AI, cs.DB

Submitted: 2026-08-07

Updated: 2026-09-02

Comments: 7 pages, 11 figures

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 69/100

The gist: The paper presents a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM).

Key concepts

Disambiguation
Separating homonyms (e.g., Munich city vs. Munich film) and merging synonyms (e.g., The Big Apple into New York City) during knowledge base construction, using context from the source triple to decide identity.
Materialized knowledge base
A knowledge base where facts are explicitly stored as triples (subject, predicate, object), allowing browsing, querying, and auditing, unlike a chat interface that generates answers on demand.
Audit trail
Evidence panels showing the disambiguation decisions made for each entity, including candidate matches and context, so users can verify why a fact was included or how a name was resolved.

Terminology

Summary

The paper presents a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM). The authors argue that while LLMs have emerged as implicit repositories of factual knowledge and prior work has constructed KBs directly from LLMs, "these approaches largely rely on surface strings as entity identifiers. Yet surface strings are inadequate identifiers, exhibiting two complementary failure modes: without further triple context, they conflate homonymous entities that share identical entity label (e.g., Munich in [Europe, hasMajorCity, Munich] vs. Munich in [Mathieu Amalric, notableWork, Munich]) and fragment synonymous strings that denote the same entity (e.g., New York City and The Big Apple)."

GPTKB 2.0 addresses this by performing context-guided disambiguation during recursive KB construction, separating homonyms and merging synonymous mentions as facts are elicited. The demo's distinguishing feature is transparency: users can browse entities, follow links across the KB, and audit the provenance of individual facts, including surface forms, candidate matches, source triples, and disambiguation decisions. The interface additionally supports structured SPARQL queries, natural-language questions translated to SPARQL, and entity linking from user-provided text to canonical GPTKB 2.0 entries. The KB is available at https://gptkb.org/.

According to Table 1, GPTKB 2.0 contains 1,592,185 entities, 38,450,135 triples, 207,633 relations, and 66,523 classes. It has 15.8M triple objects that are entities and 22.6M triple objects that are literals, with an average of 38.4 triples per entity and 15.8 outlinks per entity. Notably, 36.8% of entities are novel to Wikidata (validated on 1,000 samples).

GPTKB 2.0 represents entities, relations, and classes as first-class elements with unique identifiers, so that disambiguation can operate across homonymous and synonymous mentions. Following curated KBs such as Wikidata, besides unique ID, each element also carries a textual description, which serves as the context signal for disambiguation. The pipeline is recursive, context-guided and expands the KB along the subject–object frontier, disambiguating each new mention on the fly, proceeding through context-guided steps of elicitation, named entity recognition (NER), and disambiguation.

In the elicitation step, a bare label is ambiguous, so we condition the LLM on context — the LLM is prompted for triples with the entity's description included, so elicited facts pertain to the intended entity. In NER, Each elicited object is then classified as a literal or a named entity... using the source triple as context (e.g., 01099 is a literal in [Neustadt Dresden, postalCode, 01099] but an entity in [Frisch, partOfBand, 01099]). In disambiguation, "Named-entity objects are resolved against the KB to merge synonyms and separate homonyms. The LLM is prompted with the target entity label, its source triple, and the labels and descriptions of the nearest candidates by embedding similarity; it either matches an existing entity, marks it new, or abstains. A match inherits the incoming label as an alias. For instance, Munich (city) and Munich (film) are kept distinct through their differing source-triple context; The Big Apple is mapped to New York based on the source triple [Emerald City, contrastsWith, The Big Apple]."

The interface is built with the Django framework and served via Nginx, with the KB stored in an OpenLink Virtuoso triple store. It is described as the first demo interface to make the consolidation process fully traceable while combining SPARQL querying, natural-language (text-to-SPARQL) querying, and entity linking over an LLM-derived KB.

Access. Entities can be reached through several routes: the start page links to featured entities such as Vannevar Bush and San Francisco; a search field supports string-based lookup; and pages can be accessed directly at https://gptkb.org/entity/ (e.g., E14 for the USA).

Traceable decision-making. Each entity's page exposes the disambiguation decisions behind it, making both homonymy and synonymy auditable. To trace homonymy handling, users can inspect the candidates presented to the LLM when this entity was disambiguated, revealing how it was kept distinct from same-label candidates and judged new to the KB at the point it first appeared — for example, clicking Show under How this entity was disambiguated reveals how Hyde Park in London is separated from its homonyms existing in GPTKB 2.0. To trace synonymy handling, the surface forms consolidated into this entity are listed as its aliases; clicking one opens the prompt for the corresponding disambiguation step, showing how that form was merged. In the example of New York City, clicking the alias The Big Apple reveals the mentions mapped under that label, and the info button opens the prompt used at that step, listing the target entity, its source triple, and the retrieved candidates.

Entity linking. The interface integrates LELA... an entity-linking system, over GPTKB v2. Users enter text and click Link entities; LELA detects named-entity mentions and links each to a GPTKB v2 entity or marks it NIL. Linking is also context-dependent: in the example, the two occurrences of Mercury resolve to distinct entities, the planet (Mercury) and the Roman deity (Mercury), despite sharing a surface form. Each decision is auditable through a disambiguation-evidence panel that lists every candidate the linker weighed, where candidates are retrieved via surface-form matches ranked by observation frequency and a BM25 text-relevance fallback. "Notably, context can override raw frequency: although the planet is the most frequent entity linked to Mercury (observed 53 times), the mythological mention is correctly resolved to the Roman deity (observed only 16 times), which the surrounding context favors."

SPARQL query. The KB is stored in an OpenLink Virtuoso triple store and exposed through a SPARQL endpoint at https://gptkb.org/query/, supporting structured queries, filtering, and aggregation over all 38.4M triples. This turns questions about the model's knowledge that would otherwise require bespoke, large-scale prompting... into single declarative queries evaluated with mature database technology. Example analyses include most frequent classes (label: human 166,328; municipality 66,232; tourist attraction 51,891; fictional beaver 29,751; song 25,408; neighborhood 24,317; building 23,436; film 23,376) and entities with most aliases, where "128 distinct surface labels map to the entity United States of America. The authors note this illustrates the value of resolving entities during construction rather than afterward: because each merge is available to subsequent elicitation steps, later facts are gathered against the consolidated entity, avoiding the fragmentation that a post-hoc pass would first have to detect and then repair."

Natural-language query. The demo integrates GRASP, "an agent that answers plain-English questions over GPTKB 2.0. Given a question, the agent iteratively explores GPTKB 2.0 through tool calls, resolving the entities, properties, and classes it needs, then composes and executes a SPARQL query to produce the answer. For example, What is Budapest known for? is resolved by locating the Budapest entity (E13406) and the traditionallyKnownFor property (P67817), yielding a concise list of answers. Consistent with the rest of the interface, every step is transparent: the demo exposes the agent's full reasoning trace, the generated SPARQL query (which users can open in the query editor), and the KB entries the agent resolved, each linked to its entity page."

Disambiguation. The authors manually assess both error directions on samples of 100 cases each. Merge precision measures whether entities combined into one are truly the same: "98% of same-label objects merged into an existing entity are correct (homonymy), as are 91% of differently-labeled aliases resolved to an existing entity (synonymy); false merges (9%) in the latter are the main error mode. Split precision measures whether entities kept separate are truly distinct: all 100 same-label pairs assigned distinct IDs are genuinely different (homonymy), and for 95% of entities no synonymous duplicate appears among their top-20 label-embedding neighbors (synonymy; an approximation, since duplicates with dissimilar surface forms may be missed). The conclusion is that the pipeline thus handles both homonymy and synonymy reliably in both directions, with false synonym merges the largest residual error."

Overall quality. The authors assess randomly sampled subsets: 200 items validated by humans and 1,000 items judged by an agentic pipeline, verified against web-retrieved evidence. Precision exceeds 90% for both triples and entities. For triples (labeled true, plausible, implausible, or false), 94.5% are true and 2.0% false on the human sample, and 92.8% true and 0.5% false on the automatic sample. In an ablation where the judge sees each triple without the subject entity's description, the true rate drops to 80%, underscoring how much descriptions contribute to interpretability. For entities (labeled verifiable, plausible, or unverifiable), 96% are verifiable on the human sample and 92.3% on the automatic sample. Additional figures in Table 2 show human NED precision of 94.5% (n=400) and merge-correct/split-correct rates of 97.5%/100% (n=100 each), while LLM-based evaluation (n=1,000) gives triple precision of 92.8% true / 5.5% plausible / 1.2% implausible / 0.5% false, and entity factuality of 92.3% verifiable / 6.7% plausible / 1.0% unverifiable.

The paper positions GPTKB 2.0 against three lines of work. Disambiguated knowledge bases: "Curated KBs (DBpedia, YAGO, Wikidata) inherit canonicalization from human-maintained resources, and mention-grounding methods reuse such inventories, but both bound coverage to what those resources contain; text-based systems broaden coverage yet often set consolidation aside; LLM-derived KBs are tied to neither a fixed corpus nor a closed schema, but inherit no external identifiers and must disambiguate from scratch." Traceable KB interfaces: "Public KBs are typically exposed through browsers, entity pages, and SPARQL endpoints that present the KB as a finished artifact showing what it contains, not how each entry arose. For LLM-derived KBs, whose entries are generated and consolidated automatically, the provenance behind each fact and disambiguation decision is itself evidence of trustworthiness. GPTKB 2.0's interface makes this process traceable, surfacing for each fact its derivation trajectory, the candidates weighed during disambiguation, and the context behind each decision."

The paper concludes: "We presented GPTKB 2.0, a disambiguated knowledge base materialized entirely from LLM, together with a web interface that makes both its knowledge and the disambiguation decisions behind it explorable. The KB comprises 38.4M triples over 1.6M entities, 207.6K relations, and 66K classes, all consolidated on the fly during a recursive construction pipeline. The distinguishing feature of this demo is transparency: every fact carries provenance recording the surface forms, candidate matches, and context behind its disambiguation. Through the interface, users can browse and trace entities, query the KB in SPARQL or natural language, and link named entities in their own text against it. By turning an LLM's latent parametric knowledge into a structured, queryable, and inspectable resource, GPTKB v2 offers both a usable general-domain KB and a step toward making the knowledge inside LLMs open to scrutiny."

Improvements for AI systems

1. Canonicalize on the fly instead of post-hoc. Interleave fact elicitation with contextual disambiguation: condition each triple-elicitation prompt on the target entity's description and source-triple context, resolve each object mention immediately, inherit incoming labels as aliases, and keep gathering later facts against the consolidated entity. Result: the system avoids fragmentation that post-hoc repair would have to detect — it achieves 98% same-label merge precision, 91% synonym-merge precision, and 100% split precision, with later facts correctly attached to one canonical entity rather than duplicated across surface variants.

2. Classify literals vs. entities with triple context. Decide whether an extracted object string is a literal or a named entity based on the source triple, not the string alone (postal code 01099 vs. band 01099). Result: fewer type-confusion errors; literals are stored as values while entities are routed through disambiguation.

3. Ground every step in entity descriptions. Give each entity a textual description and include it in elicitation, disambiguation, candidate comparison, and verification prompts. Result: fact truth-rate holds at 94.5%, versus 80% when the subject description is withheld — the system produces more accurate and interpretable knowledge when descriptions anchor reasoning.

4. Make disambiguation evidence-based with abstention. When resolving a mention, retrieve nearest candidates by embedding similarity, present their labels and descriptions alongside the target mention and its source triple, and let the LLM either match, mark-new, or abstain. Persist the candidates, prompt, and decision as provenance. Result: the system keeps Munich (city) and Munich (film) apart, merges The Big Apple into New York from triple context, and abstains rather than forcing a wrong merge — with every decision auditable.

5. Materialize parametric knowledge into a queryable store. Export the LLM's elicited facts into a triple store (e.g., Virtuoso/SPARQL) so aggregate analyses become single declarative queries instead of bespoke prompting. Result: the system can compute class frequencies, alias counts, and property statistics over 38.4M triples reproducibly — e.g., finding that 128 surface labels map to United States of America.

6. Answer natural-language questions via a transparent tool-use agent. Resolve the entities, relations, and classes in a question by iteratively exploring the KB, then compose and execute SPARQL, exposing the full reasoning trace, the generated query, and the resolved KB entries. Result: the system answers multi-step questions like What is Budapest known for? with verifiable, rerunnable steps.

7. Let text context override entity-linking frequency. Rank candidate entities by surface-form observation frequency with a BM25 fallback, but allow surrounding text to override raw frequency. Result: Mercury in a mythological sentence links to the Roman deity (observed 16 times) instead of the planet (observed 53 times), with an evidence panel showing every candidate weighed.

8. Self-verify against web evidence. Use an agentic judge that retrieves web evidence to label triples true/plausible/implausible/false and entities verifiable/plausible/unverifiable, and feed that signal back into construction or surface it as confidence. Result: the system self-monitors at >92% true/verifiable rates, and can flag low-confidence facts for human review.

9. Expand recursively along the subject–object frontier. From seed entities, elicit facts about each discovered entity at first mention, disambiguating it on the spot. Result: broad coverage (36.8% of entities novel to Wikidata) with consolidation happening at first appearance rather than requiring a later repair pass, and 94.5% human-verified triple precision.

Combined capability: these changes produce an AI system that builds a large, disambiguated knowledge base end-to-end with auditable provenance — users can browse any entity, trace why it was merged or split, query the whole KB in SPARQL or plain English, and link mentions in their own text — while maintaining high precision through context-grounded decisions and web-verified self-judgment.

Abstract

We present a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM). GPTKB 2.0 contains 38.4M triples over 1.6M canonical entities, together with 207.6K consolidated relations and 66K consolidated classes. Unlike prior LLM-derived knowledge bases that largely identify entities by surface strings, GPTKB 2.0 performs context-guided disambiguation during recursive KB construction, separating homonyms and merging synonymous mentions as facts are elicited. The demo makes this process inspectable: users can browse entities, follow links across the KB, and audit the provenance of individual facts, including surface forms, candidate matches, source triples, and disambiguation decisions. The interface further supports structured SPARQL queries, natural-language questions translated to SPARQL, and entity linking from user-provided text to canonical GPTKB 2.0 entries. GPTKB 2.0 is available at https://gptkb.org/, with the full KB downloadable for offline use.

Sources

Related papers