TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models
cs.DB, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: Preprent version of the paper accepted to SIGMOD 2027 Round 2
Code: https://github.com/shirley-wu/text_to_table
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Table extraction from texts is an important task for information systems, and recent approaches that prompt large language models (LLMs) with instructions have drawn great attention for their strong
Terminology
Abstract
Table extraction from texts is an important task for information systems, and recent approaches that prompt large language models (LLMs) with instructions have drawn great attention for their strong performance. Existing works have assumed the input texts to be table descriptions or specialized documents. However, these efforts have largely overlooked another prevalent category of texts, commonly found in news reports and social media: naturally occurring texts. Extracting tabular information from such texts poses two distinct challenges. First, high variability and the absence of explicit structural cues make fixed heuristic LLM prompts limited in precisely delineating extraction boundaries. Second, manually predefined schemas cannot capture open-ended, unseen attributes in naturally occurring text. In this paper, we propose a framework, TEAR, to address these challenges. It comprises two synergistic workflows: a Table Extraction Workflow that dynamically adapts instructions to overcome the limitation of heuristic instructions, and an Attribute Recommendation Workflow that discovers new attributes from texts to complement the heuristic schema. To our knowledge, TEAR is the first framework that supports automated text-driven attribute recommendation, enabling exploratory schema design for table extraction. To evaluate TEAR, we establish the benchmark for table extraction and attribute recommendation on naturally occurring texts, including two real-world datasets, manual annotations, appropriate metrics, and baseline comparisons. Experiments show that TEAR achieves state-of-the-art performance on both tasks, and the recommended attributes effectively enhance extraction performance in exploratory scenarios.
Sources
- Map&Make: Schema Guided Text to Table Generation
- The E2E Dataset: New Challenges For End-to-End Generation
- Instruction Tuning with GPT-4
- When does In-context Learning Fall Short and Why? A Study on Specification-Heavy Tasks
- Challenges in Data-to-Document Generation
- A Survey on Neural Open Information Extraction: Current Status and Future Directions
Related papers
- Vibe Coding on Trial: Operating Characteristics of Unanimous LLM Juries
- Human-Level Text-to-SQL via Reinforcement Learning on Verified Data, Without Pipeline Engineering
- Bridging Business Intent and Data: A Benchmark for Automatic Relational Data Product Generation
- DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation
- MaDI-Bench: An End-to-End Data Integration Benchmark
- Eigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning