REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs

arXiv:2608.10963 · cs.CL · Submitted 2026-08-11 · Read on arXiv

VNU University of Engineering and Technology

cs.CL

Submitted: 2026-08-11

Updated: 2026-09-02

Code: https://github.com/yammdd/AKBC-Shared-Task-2026

Project page: https://lm-kbc.github.io/challenge2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper presents R EAP (Relation-aware Elicitation And Parsing), a system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting.

Terminology

Summary

The paper presents R EAP (Relation-aware Elicitation And Parsing), a system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting. The system operates under strict constraints: no retrieval-augmented generation or external knowledge, at most 32B parameters, and no model fine-tuning. The task requires returning complete sets of objects (which may be empty, single-valued, or multi-valued) as valid JSON arrays for six relations: countryLandBordersCountry, personHasCityOfDeath, hasCapacity, awardWonBy, companyTradesAtStockExchange, and hasArea.

The R EAP system is a two-stage pipeline that decouples factual elicitation from answer serialization. Stage 1 prompts the LLM using strategies tailored to each relation, while Stage 2 uses deterministic JSON parsing when possible and LLM-based extraction only as a fallback. The system combines structured chain-of-thought reasoning, relation-specific query strategies, and a reasoning-based empty-set gate to elicit parametric knowledge, followed by direct extraction into valid JSON arrays.

The relation-specific strategies include: for awardWonBy, a chronological multi-pass approach with one general prompt and five auxiliary prompts decomposed by time periods (from the award's inception through the 1970s, the 1980s–1990s, the 2000s–2010s, 2020 onward, and a final pass for less prominent or non-Western recipients), using temperature T = 0.7 to diversify recall; for countryLandBordersCountry, a four-direction geographical scan (eastern, western, southern, and northern boundaries) that excludes maritime borders and returns BORDERS: NONE for island countries; for companyTradesAtStockExchange, a four-step CoT procedure including a Strict Public Check (empty-set gate) to determine if the entity is private, nonprofit, or a non-listed subsidiary; for hasArea, a four-step entity disambiguation and unit normalization procedure; for hasCapacity, a four-step reasoning procedure with capacity-range verification (1,000–35,000 seats for small venues and 35,000–100,000 for large national stadiums); and for personHasCityOfDeath, a five-step reasoning procedure with an empty-set gate that returns FINAL ANSWER: [] if the person is alive or information is unavailable.

The post-processing stage includes direct parsing of FINAL ANSWER: lines, balanced-bracket scanning for truncated JSON arrays, numeric extraction and normalization for quantitative relations (removing thousands separators and measurement units), title and noise filtering for awardWonBy (removing year prefixes, HTML tags, work-title suffixes, and honorifics), parenthetical filtering, and case-insensitive deduplication.

The authors evaluate three instruction-tuned models—Gemma-2-9B-it, Llama-3.1-8B-Instruct, and Mistral-Small-24B-Instruct-2501—against the organizer's Qwen3.5-9B baseline. Experiments are conducted on Kaggle using a TPU v5e-8 with eight TPU v5e cores, serving models with vLLM TPU Server and bfloat16 precision. One complete run over all six relations in the validation set (475 records) takes approximately 2–5 minutes with batch size=32.

In zero-shot settings, all models achieve similar low macro-F1 scores (Llama-3.1-8B: 0.38, Gemma-2-9B: 0.42, Mistral-24B: 0.43), suggesting they struggle when required to recall facts and satisfy output format simultaneously. The two-stage pipeline improves macro-F1 for all models: Llama-3.1-8B improves from 0.38 to 0.48, Gemma-2-9B from 0.42 to 0.51, and Mistral-24B from 0.43 to 0.65. Mistral-24B yields the largest absolute gain (+0.22) and is selected as the primary model.

An ablation study on the validation set using Mistral-24B shows: removing task decomposition drops macro-F1 from 0.65 to 0.59 (awardWonBy collapses from 0.25 to 0.00 and countryLandBordersCountry falls from 0.98 to 0.63); removing the empty-set gate lowers macro-F1 to 0.57 (hurting precision on personHasCityOfDeath from 0.50 to 0.36 and companyTradesAtStockExchange from 0.76 to 0.57); and removing CoT reasoning entirely causes the largest drop to 0.43 (degrading hasArea from 0.82 to 0.55 and countryLandBordersCountry from 0.98 to 0.27).

On the official test set, the primary system (built on Mistral-Small-24B-Instruct-2501) achieves a macro-F1 score of 0.62, with relation-level scores: countryLandBordersCountry (0.95), hasArea (0.77), companyTradesAtStockExchange (0.73), personHasCityOfDeath (0.53), awardWonBy (0.37), and hasCapacity (0.23). This significantly outperforms the Llama (0.48) and Gemma (0.48) pipelines and more than doubles the organizer's Qwen3.5-9B baseline (0.30).

The discussion highlights that structured reasoning is important for reliably eliciting knowledge encoded in model parameters, and that a larger model (Mistral-24B) benefits more from the pipeline than smaller models (Gemma-2-9B), particularly on long-tail entities such as small islands, club-level stadiums, and less prominent awards. The limitations include incomplete parametric knowledge for very rare long-tail entities, predictions for quantitative relations sometimes falling outside the 5% relative-tolerance threshold, and slight run-to-run variation (about ±0.005 macro-F1) due to stochastic sampling for awardWonBy and non-deterministic distributed computation on TPU v5e-8. hasCapacity is identified as the most challenging relation (test macro-F1 0.23) because venues are frequently confused with larger namesakes in the same city and even small numeric errors exceed the 5% tolerance.

Improvements for AI systems

Improvements to AI Systems:

  1. Relation-Aware Prompt Decomposition: Implement a two-stage pipeline that separates factual recall from output formatting. Stage 1 uses relation-specific query strategies (e.g., chronological multi-pass for awards, four-direction geographical scans for borders, multi-step CoT for quantitative relations). Stage 2 uses deterministic parsing (e.g., extracting FINAL ANSWER: lines, balanced-bracket scanning) before falling back to LLM-based extraction. This decoupling reduces format errors and improves recall.

  2. Empty-Set Gate for Precision: Add a reasoning-based gate that explicitly checks whether the answer should be empty (e.g., person alive, private company, island nation). This gate prevents hallucinated false positives, improving precision on relations like personHasCityOfDeath (precision +0.14) and companyTradesAtStockExchange (precision +0.19).

  3. Structured Chain-of-Thought with Verification Steps: Replace free-form reasoning with fixed-step CoT procedures that include numeric range checks (e.g., stadium capacity 1,000–100,000 seats), entity disambiguation (e.g., distinguishing stadiums from namesakes), and unit normalization. This boosts performance on quantitative relations like hasArea (F1 +0.27) and countryLandBordersCountry (F1 +0.71).

  4. Multi-Pass Temporal Decomposition for Long-Tail Facts: For relations with many possible objects (e.g., awardWonBy), use multiple prompts segmented by time periods (inception–1970s, 1980s–1990s, 2000s–2010s, 2020+, and a final pass for obscure recipients). Use higher temperature (T=0.7) to diversify recall. This prevents omission of less prominent or non-Western entries.

  5. Deterministic Post-Processing Filters: Apply rule-based cleaning: remove year prefixes, HTML tags, honorifics, and work-title suffixes; filter parenthetical noise; perform case-insensitive deduplication; strip thousands separators and units from numeric answers. This reduces format errors without additional LLM calls.

  6. Model-Size-Aware Pipeline Tuning: The pipeline yields larger gains for larger models (Mistral-24B: +0.22 F1) than smaller ones (Gemma-9B: +0.09). Therefore, allocate more computational resources to larger models when using this pipeline, and expect diminishing returns for models under 10B parameters.

What the Improved AI System Can Do:

  • Reliably extract complete, valid JSON arrays for multi-valued relations (e.g., all countries bordering another, all awards won by a person) without external retrieval or fine-tuning.

  • Achieve high precision on empty-set cases (e.g., correctly returning [] for living persons or private companies), reducing hallucination rates.

  • Handle quantitative facts with tolerance-aware accuracy (e.g., area within 5% relative error) by using verification steps and unit normalization.

  • Recall long-tail entities (small islands, club stadiums, obscure award recipients) via temporal decomposition and diversified sampling.

  • Operate under strict constraints (≤32B parameters, no RAG, no fine-tuning) while outperforming baseline systems by 2× on macro-F1 (0.62 vs. 0.30).

  • Provide fast inference (2–5 minutes for 475 records on TPU v5e-8) suitable for batch knowledge-base construction tasks.

Abstract

We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a budget of at most 32B parameters and no model fine-tuning. Our system combines structured chain-of-thought reasoning, relation-specific query strategies, and a reasoning-based empty-set gate to elicit parametric knowledge, followed by direct extraction into valid JSON arrays. On the test set, the system, built on the Mistral-Small-24B-Instruct-2501 model, achieves a macro-F1 score of 0.62, with particularly strong results on countryLandBordersCountry (F1 = 0.95), companyTradesAtStockExchange (F1 = 0.73), and hasArea (F1 = 0.77). Our code is publicly available at https://github.com/yammdd/AKBC-Shared-Task-2026.

Sources

Related papers