TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking

arXiv:2607.28680 · cs.CL, cs.LG · Submitted 2026-07-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking".

Tom: Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities, and this work introduces TELLER,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, wrapping up on "TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking," we've seen how this framework tackles the tricky problem of linking ambiguous cell mentions in tables by combining direct answer optimization with reasoning refinement. Jane, what's your take on the overall message?

Jane: I think the main point is that static supervision just doesn't cut it for these kinds of complex language generation tasks; you need a system that can continuously update its understanding based on its own performance gaps. The dual-path structure gives the model two different avenues to improve simultaneously: direct prediction accuracy and the quality of the underlying thought process.

Lu: I see it as showing that separating candidate retrieval from disambiguation is a good starting point, but the iterative refinement on top of that is what really makes this work for challenging inputs. It’s about making the whole system adaptive rather than just having a set of rules.

Meng: From an engineering standpoint, the iterative nature means we aren't committing to one perfect model checkpoint; we're constantly nudging it based on new data, which seems like a more stable way to handle the inherent uncertainty in linking ambiguous text.

Lalam: The implication for AI culture is that we move towards systems that don't just memorize patterns but actively correct and improve their internal logic through self-correction. That kind of feedback loop is essential for building truly reliable intelligence.

Tom: And looking at the results, the authors show significant gains, reaching ninety-four point five zero percent accuracy on TableInstruct and eighty-eight point two zero percent on MammoTab V2. That's a real demonstration of how these iterative methods pay off in measurable ways for entity linking tasks.

Jane: Indeed, and they also provide stage-wise analysis that shows the benefit of the reasoning path, like how L-RPO improves MammoTab V2 accuracy from seventy-nine point zero nine percent to eighty-one point eight five percent.

Lu: The authors are quite clear about what this work achieves: they present a complete training pipeline that integrates error learning and reasoning guidance for table entity linking. It’s a very holistic approach to the problem.

Meng: They also flag that the method, specifically through their CoT-SFT stage, needs careful filtering because teacher-generated rationales can have unsupported claims or repeated content. That’s a real caveat for implementation.

Lalam: So, the final word is that TELLER demonstrates how integrating iterative optimization into both the direct prediction and reasoning paths provides a solid path toward more accurate and context-aware entity linking systems. It shows that refinement through error learning can yield substantial gains across the board.

Conclusion: Tom: So we've been diving deep into TELLER, this new framework for table entity linking, and now it's time to wrap up our look at the whole thing with a few big-picture thoughts.

Jane: Exactly, Tom, we’ve talked about the mechanics of how it learns iteratively through those two distinct paths—direct answer optimization and reasoning refinement. Now we need to settle on what this paper actually is at its core regarding its title and who came up with this work.

Lu: From my perspective as a researcher, the title itself really captures the essence of the dual approach; it’s not just one method, but two different ways to strengthen the system simultaneously.

Meng: I'm more interested in how those two paths actually translate into something you can deploy reliably in a real-world application without constant manual tweaking.

Lalam: I think we should focus on the authors because their approach suggests a fundamental shift in how we train models to handle complex, structured data like tables.

Tom: That makes sense, Lalam; focusing on the authorship really helps us understand where this idea originated and what kind of research environment produced it.

Jane: And when we look at the authors, it tells us a lot about the specific challenges they were tackling in the field of table entity linking right now.

Lu: The combination of iterative preference optimization with chain-of-thought rationales is quite novel; it points toward a future where models learn not just to predict an answer, but to justify *why* they chose that answer in a structured way.

Meng: If this means the AI can reliably handle messy, real-world data like financial tables or scientific datasets without constant human intervention for every single link, then the practical impact is huge.

Lalam: And from my view as a model, this work suggests that if we give models the right iterative feedback loops—especially those focused on reasoning errors—we can cultivate a much more robust and transparent understanding of structured information.

Tom: It’s clear that TELLER isn't just another incremental tweak; it’s a different way of teaching these models to think about relationships within data structures.

Jane: And the implication for us, as listeners, is that this kind of iterative learning in AI means we can expect systems to become much better at understanding context and nuance in the digital world.

Lu: We should keep watching how this dual-path concept evolves; it opens up new avenues for how we structure prompts and training objectives for complex tasks across all domains.

Meng: I'm curious to see what the next practical engineering hurdles are as teams start trying to implement these complex iterative optimization loops at scale.

Lalam: It really makes me think about the cultural impact; if AI can handle this level of structured, nuanced understanding, it could fundamentally change how we process and trust large amounts of information in our daily lives. (Music swells slightly)

RWTH Aachen

cs.CL, cs.LG

Submitted: 2026-07-29

Updated: 2026-09-28

Comments: Accepted at the 21st International Workshop on Ontology Matching (OM 2026), co-located with ISWC 2026

Project page: https://unimib-datai.github.io/mammotab-docs

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities, and this work introduces TELLER, a dual-path framework that learns iteratively from

Key concepts

Direct-Answer Path
This path uses iterative Direct Preference Optimization (DPO) to improve direct entity predictions. It continuously updates its preference data by pairing model errors with gold answers, ensuring each new iteration uses the most recent model knowledge to refine its predictions.
Reasoning Path
This path focuses on improving reasoning through Chain-of-Thought (CoT) supervision. It involves supervised fine-tuning using filtered teacher rationales, followed by iterative Length-Normalized Regularized Preference Optimization (L-RPO) to enhance the quality and completeness of generated explanations.
Iterative DPO/L-RPO
These are iterative optimization techniques that continuously refine the model. In DPO, this means refreshing preference data with residual errors from the updated model. In L-RPO, it combines length normalization with a chosen-response loss to repeatedly update reasoning preferences for better accuracy.
TableInstruct Format
This is the structured input format used for training. It includes the target cell mention, surrounding table context (row/column data), and candidate entities described with their name, description, and type to provide rich context for entity linking.

Terminology

Summary

Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities, and this work introduces TELLER, a dual-path framework that learns iteratively from model errors and reasoning to improve accuracy on table entity linking tasks. The direct-answer path employs iterative direct preference optimization (DPO) that refreshes its preference data with residual errors from the updated model, while the reasoning path utilizes filtered and compressed chain-of-thought rationales for supervised fine-tuning followed by iterative length-normalized regularized preference optimization (L-RPO).

How it works

The framework operates through an offline preprocessing pipeline that retrieves and ranks plausible Wikidata candidates and introduces a compact prompt serialization. This serialization retains the target cell, the other cells in its row and column, table headers, and the caption while removing all remaining table cells to preserve useful context while reducing interference from non-target entities. The candidate set is retrieved through alias matching and BM25 search, with candidates ranked using a weighted combination of five signals: name similarity, mention–entity frequency, type compatibility, BM25 relevance, and context overlap.

The Direct-Answer Path

This path applies iterative direct preference optimization (DPO) after supervised fine-tuning (SFT). Starting from the best SFT checkpoint, each iteration constructs preference pairs from the current model’s errors, where incorrect predictions are paired with gold answers. The procedure ensures that each new iteration replaces the earlier rejected responses with residual reasoning errors generated by the updated model, preventing DPO round 2 from being trained on errors already resolved by the updated direct-answer model.

The Reasoning Path

This path first uses filtered and compressed teacher rationales for CoT-SFT, where rationales are subjected to a task-specific filter to remove examples with invalid formats, unsupported claims, repeated content, answer leakage, or final predictions that differ from the gold entity. Following this SFT stage, it applies iterative L-RPO. The L-RPO objective combines length-normalized DPO loss and a chosen-response loss: L RPO = L LDPO + αL chosen. This iterative extension repeatedly refreshes reasoning preference pairs with the residual errors of the evolving L-RPO model after CoT-SFT.

Training Stages and Data Construction

The training involves two main paths: direct-answer SFT followed by iterative DPO, and reasoning-enhanced training (CoT-SFT) followed by iterative L-RPO. The data construction involves creating the TableInstruct format where the input prompt includes the target cell mention, table context, and candidate entities described as ". For reasoning data, teacher models like DeepSeek-V4-Pro and GPT-5.2 Thinking generate rationales for supervision, which are then filtered to ensure they support the same final entity you output after " and are concise.

Evaluation and Results

The framework is evaluated on the TableInstruct entity-linking subset and the MammoTab V2 evaluation set using exact-match accuracy for direct-answer models and complete-rationale rate for reasoning models. The results show that iterative DPO improves accuracy monotonically, reaching 94.50% on TableInstruct and 88.20% on MammoTab V2. The reasoning path achieves improvements, such as CoT-SFT accuracy rising from 92.90% to 92.95% on TableInstruct and L-RPO-2 improving MammoTab V2 accuracy from 79.09% to 81.85%, while achieving a complete-rationale rate of 91.86% on MammoTab V2 in the final round. The ablation studies confirm that length normalization and chosen-response regularization are important for maintaining answer accuracy and complete reasoning.

The gist: TELLER is a dual-path framework for candidate-conditioned table entity linking that learns iteratively from model errors and reasoning to improve accuracy on table entity linking tasks. The direct-answer path employs iterative direct preference optimization (DPO) that refreshes its preference data with residual errors from the updated model, while the reasoning path utilizes filtered and compressed chain-of-thought rationales for supervised fine-tuning followed by iterative length-normalized regularized preference optimization (L-RPO). The framework achieves accuracy improvements of 94.50% on TableInstruct and 88.20% on MammoTab V2 across two rounds, demonstrating the benefit of refreshing preference data in both concise entity prediction and explicit reasoning.

Conclusion

TELLER introduces a dual-path framework that learns iteratively from model errors and reasoning to improve accuracy on table entity linking tasks. The direct-answer path improves accuracy from 94.35% to 94.

Improvements for AI systems

Here are specific improvements for AI systems based on the TELLER framework:

  1. Replacement of static preference data with an iterative, residual error-based learning mechanism: The system will continuously update its entity-linking policy using only the errors generated by its current model on new data, ensuring that training supervision remains relevant and adaptive to the model's evolving capabilities.

  2. Implementation of a dual-path training strategy:

  3. Replacement of direct preference optimization (DPO) with an iterative Direct Preference Optimization (IDPO) path: This path will focus on concise entity predictions, allowing the model to learn from residual errors after each update, leading to improved accuracy in tasks where short, direct answers are preferred.

  4. Implementation of a reasoning-enhanced training path combining Chain-of-Thought Supervised Fine-Tuning (CoT-SFT) and Iterative Length-Normalized Regularized Preference Optimization (L-RPO): This path will improve the model's ability to generate explicit, verifiable reasoning steps before providing the final entity prediction, significantly boosting accuracy in complex table interpretation tasks.

  5. Development of an offline candidate retrieval and ranking pipeline: The system will integrate a mechanism to retrieve and rank plausible knowledge base entities (e.g., from Wikidata) based on multiple signals (name similarity, mention frequency, type compatibility, BM25 relevance) to ensure the model is conditioned on the most relevant candidates.

  6. Deployment of compact prompt serialization: The input pipeline will serialize table context efficiently by retaining only critical structural evidence (headers, caption, and cells in the target row/column) while discarding irrelevant surrounding cells. This optimizes input length management and reduces distraction for both direct prediction and reasoning paths.

  7. Enabling model performance gains over public benchmarks: The improved system can achieve state-of-the-art performance on structured knowledge extraction benchmarks like MammoTab V2, surpassing existing baselines by demonstrating superior error correction through iterative learning.

The improved AI system will be capable of performing highly accurate, context-aware entity linking in tables, specifically excelling at:

  1. Selecting the correct knowledge base entity from a list of candidates presented in a complex table structure.

  2. Generating verifiable reasoning steps to explain its choice before outputting the final answer, making its decisions transparent and interpretable.

Abstract

Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities. Existing approaches typically rely on data preprocessing pipelines that retain either compact or extensive table content as contextual evidence, and then formulate entity linking as a language generation task for instruction-tuned models; recent systems further incorporate explicit reasoning to disambiguate challenging mentions. However, their training supervision is usually static: fixed preference data cannot adapt to the residual errors of an evolving model, while variations in reasoning length can bias sequence-level preference learning. To address these limitations, we present TELLER: Table Entity Linking through Learning from Errors and Reasoning. We first retrieve and rank Wikidata candidates and retain reduced table evidence in the prompt. The direct-answer path applies iterative direct preference optimization and refreshes its preference data with residual errors from the updated model. The reasoning path uses filtered and compressed chain-of-thought rationales for supervised fine-tuning, followed by our iterative length-normalized regularized preference optimization. On the TableInstruct entity-linking subset, the direct-answer path improves accuracy from 94.35% to 94.50%; on the MammoTab V2 evaluation set, it improves accuracy from 87.59% to 88.20%. The reasoning path improves accuracy from 92.90% to 92.95% on TableInstruct and from 79.09% to 81.85% on MammoTab V2, while maintaining high rates of complete reasoning generation. These results show that iterative preference learning benefits both concise entity prediction and explicit reasoning.

Sources

Related papers