Labeling Training Data for Entity Matching Using Large Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Labeling Training Data for Entity Matching Using Large Language Models".
Jane: Recent large language models (LLMs) achieve strong performance on entity matching without requiring task-specific training data,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we're looking at this paper today called "Labeling Training Data for Entity Matching Using Large Language Models." It’s about how we can use those powerful large language models for entity matching without having to manually label all that training data ourselves.
Jane: That sounds like a big deal because getting that task-specific training data is usually the hardest part of using these LLMs for this kind of work. The paper basically asks if we can skip labeling by using a knowledge distillation workflow where an LLM acts as a teacher to label pairs for a smaller student model.
Lu: It's interesting because the core idea is that you use the big, slow LLM to generate high-quality labels for training data, and then you train a much faster, smaller student model on that machine-labeled set instead of labeling everything manually yourself.
Meng: From an engineering standpoint, that makes sense if we think about the cost and time involved in getting task-specific data versus using a distillation pipeline to get high-quality labeled examples quickly. But how do they actually structure this workflow?
Tom: Well, they lay out this whole process starting with building a candidate pool by embedding records and ranking them by similarity. They then have several ways to pick which pairs to send to the LLM teacher for labeling, which is where things get pretty detailed.
Jane: It looks like they compare three different pair selection strategies: one based on similarity search, another using active learning, and a feature-based active learning variant. This choice of strategy seems crucial for how much quality data you end up getting.
Lu: The paper tests several ways to select those pairs, including a similarity search that looks at batches of five candidates and stops when it finds certain positive and negative pairs, or an active learning approach where a committee of models scores unlabeled pairs and sends the most disagreed-upon ones to the LLM.
Meng: That sounds like a lot of complexity to manage in practice. I wonder how stable those selection methods are when you're trying to get that high quality input for the teacher model.
Paper summary: Tom: The LLM labeling itself is also constrained by a very specific instruction, forcing it to output only a JSON field saying whether two records match or not, and they even add a conservative prompt telling the teacher model to only predict a match if the evidence strongly supports it without contradiction.
Jane: That post-processing step is important because even with an LLM teacher, you still get labels you have to check. They look at two methods: relabeling, which uses another LLM to verify the original label, and closure-based dropping, which removes pairs that would split a connected component in a graph of source records.
Lu: The results show that this whole setup is quite effective; student models trained on these machine-labeled sets perform about as well as models trained on the benchmark training sets, with only differences staying below two F1 points in both directions.
Tom: And the paper points out that active learning with a Ditto strategy was actually the strongest pair selection method on four of the five benchmarks they tested, and it even shifted the class balance toward positives by up to twelve point seven four percentage points on DBLP-ACM.
Meng: That's a significant shift in class balance if you're trying to train a model where one outcome is rare, which is often the case in matching tasks. It suggests that smart selection can really help guide the training process.
Jane: Regarding cost, they quantified it and found that using GPT-five point two for labeling all five datasets costs between US twenty-eight point three one to US forty point eight eight, which they say is a small fraction of the estimated cost of doing manual labeling yourself in many cases <ref:2606.28823#pg0>.
Tom: And inference time is where things get really interesting because Ditto was shown to be forty-one point five to five hundred thirty-four times faster than just using an LLM directly for matching tasks, which really speaks to the efficiency gain here <ref:2606.28823#pg0>.
Lu: It seems like this knowledge distillation workflow isn't just about getting a better label; it's about making the overall process much cheaper and quicker than traditional methods. The paper, "Labeling Training Data for Entity Matching Using Large Language Models," shows how you can substantially reduce or even eliminate the manual effort needed to label use case-specific training data for entity matching when combined with a suitable pair-selection method.
Paper summary: Meng: So, what does this mean practically for someone who is actually building these systems? It suggests that instead of spending all that time and money creating custom training sets, you can leverage existing LLMs as a powerful labeling engine to bootstrap your student models.
Tom: Exactly, it means the focus shifts from generating data to optimizing the distillation pipeline itself, focusing on those four design choices: pair-selection strategy, teacher model, label post-processing method, and the student model. It’s a holistic approach.
Jane: Looking at the authors and their work with these large models like GPT-five point two and Kimi K2 point 6 as teachers versus smaller models like RoBERTa for comparison, it highlights how different sizes of AI can be leveraged for different parts of the pipeline here <ref:2606.28823#pg0>.
Lu: The potential here is that we can train specialized entity matching tools much faster, which could improve the culture around data integration by making it more accessible and less reliant on massive manual annotation efforts.
Meng: From my side, I see this as a way to move from slow, high-cost training cycles to a much more iterative process where you can quickly test different distillation settings without needing huge human annotation teams ready for every iteration.
Tom: It's definitely moving the needle on how we approach entity matching data preparation. So, what are the bigger implications of this knowledge distillation workflow for the future of NLP applications?
Jane: The implication is that we don't have to rely solely on massive, expensive human labeling efforts to fine-tune models for specific entity matching tasks anymore; we can use LLMs as a scalable way to generate synthetic, high-quality training data.
Lu: It means the barrier to entry for applying advanced matching techniques with LLMs is lowered because you don't need a huge team of annotators just to get started on your next project.
Meng: It makes the whole process more accessible and less dependent on having perfect, task-specific datasets ready from day one.
Tom: That’s the essence of this paper, "Labeling Training Data for Entity Matching Using Large Language Models." It shows a practical path to getting better entity matching models using distillation workflows rather than just relying on raw LLM performance alone.
Conclusion: Tom: So we’re wrapping up this look at "Labeling Training Data for Entity Matching Using Large Language Models." Basically, they figured out a way to use big language models as teachers to generate the training data needed for smaller, faster matching AI models instead of labeling everything by hand.
Jane: It boils down to using knowledge distillation. The authors take a powerful LLM and have it label pairs that are too much work for humans, and then they train a much smaller student model on those machine-labeled examples.
Lu: What's really interesting is how they select which pairs get labeled by the LLM teacher. They test several strategies—similarity search, active learning, even feature-based methods—to figure out the best way to pick the data.
Meng: From an engineering view, that makes sense because you can’t just feed a huge dataset into a small model and expect good results without high-quality labels upfront. The choice of selection method directly impacts how much better that student model ends up being.
Lalam: I see this as a way to scale up our matching capabilities dramatically. Instead of waiting for endless human annotation cycles, we can leverage existing LLM intelligence to bootstrap our training sets quickly and efficiently.
Tom: It really changes the workflow from being data-heavy to process-heavy, focusing on how you structure that distillation pipeline instead of just drowning in manual labeling tasks.
Jane: And the results are solid; student models trained this way perform nearly as well as those trained on perfectly labeled benchmark sets, with only tiny differences remaining.
Lu: The authors found that a specific active learning method was particularly strong at shifting the class balance toward matches, which is a significant detail for how entity matching systems actually operate in the real world.
Meng: And the cost analysis shows that this whole process is way cheaper than doing all of that manual work yourself, and it’s much faster when you consider how quickly these LLMs can generate those labels compared to direct use.
Lalam: So, we’re moving toward a future where entity matching AI isn't bottlenecked by the time it takes to get perfectly labeled data.
Tom: Exactly. This means the focus shifts from just training a model to building an efficient system that leverages LLMs for high-quality data generation on demand.
Jane: It opens up possibilities for applying advanced matching techniques to much more complex, niche datasets where human labeling would otherwise be prohibitively expensive or slow.
Aaron Steiner, Christian Bizer
University of Mannheim
cs.CL
Submitted: 2026-06-27
Updated: 2026-10-05
Comments: 13 pages, 2 figures, 11 tables
Code: https://github.com/wbsg-uni-mannheim/Automaticdata-labeling
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 86/100
The gist: Recent large language models (LLMs) achieve strong performance on entity matching without requiring task-specific training data, but applying these models to large sets of candidate pairs remains
Key concepts
- Knowledge Distillation Workflow
- This is the core process where a large 'teacher' LLM is used to generate high-quality labels for candidate pairs. These labeled pairs are then fed into a smaller 'student' model to learn how to perform entity matching efficiently, saving computational resources.
- Candidate Pool Generation
- This initial step creates a large set of potential matching pairs from two tables. It involves embedding records and ranking them using similarity scores derived from these embeddings to select the most relevant candidates for subsequent labeling.
- Active Learning (Ditto)
- This is an advanced pair selection strategy where a committee of matching models ranks unlabeled pairs. The system intelligently selects pairs where the committee members disagree, sending these most uncertain pairs to the LLM teacher for expert labeling.
Terminology
Summary
Recent large language models (LLMs) achieve strong performance on entity matching without requiring task-specific training data, but applying these models to large sets of candidate pairs remains slow and costly, necessitating methods like knowledge distillation to train smaller student models using machine-labeled data. This paper investigates whether the need to provide task-specific training data can be avoided by using knowledge distillation workflows where an LLM serves as a teacher model to label training pairs that are subsequently used to train a smaller student model.
How it works
The core methodology involves a knowledge distillation workflow
illustrated in Figure 1, which consists of several sequential steps: building a candidate pool, selecting pairs for labeling, invoking an LLM teacher model for labeling the pairs, optionally post-processing the generated labels, and finally exporting a machine-labeled training set which is subsequently used to train a student matcher. This process allows for the use of computationally cheaper student models trained on high-quality data generated by an LLM teacher.
Pair Selection Strategy
The workflows begin with Candidate Pool Generation,
where records are embedded using text-embedding-3-small, and each left-table record is ranked against every right-table record using cosine similarity of the embeddings. The candidate pool is filled with 20 pairs per left-table record: 18 pairs generated by combining the record with its nearest neighbors in the right table, plus two records sampled at random from the lower-ranked half of the similarity ranking
(Page 3).
The paper compares three pair selection strategies for labeling:
-
Similarity search: This scans candidate pool candidates in similarity-ranked batches of five, stopping after a batch produces at least one positive and four negatives, with specific rules for labeling
dense positive regions
(Page 3). -
Active learning: This strategy starts from a
similarity search seed that targets 100 pairs with a 30/70 positive/negative split,
and in each round, it trains a committee of matching models to score up to 20,000 unlabeled pairs, passing on those where the committee disagrees most (Page 3). -
Active learning (ML): This feature-based variant uses an ensemble of logistic regression, random forests, etc., and ranks candidates by
disagreement over match-probability scores
before sending the most disagreed upon pairs to the LLM teacher model for labeling (Page 3). -
Active learning (Ditto): This strategy trains a bagged ensemble of five Ditto matchers during the active-learning phase, where ranking is
lexicographic: variance of the predicted match probabilities first, then entropy of the thresholded match/non-match votes, distance of the mean probability from 0.5, and number of split votes
(Page 3).
LLM-based Labeling
For every selected pair, the teacher receives two serialized entity descriptions and returns a JSON output. The system prompt constrains the answer to a JSON match field: You are an expert entity matcher. Decide if two records refer to the same real-world entity. Return only valid JSON with exactly one field: 'match': truefalse
(Page 3).
The teacher models used include GPT-5.2, Qwen 3.6 Plus, and the openweight Kimi K2.6, with GPT-5.2 being used for most experiments (Page 3). The relabeling prompt adds a conservative constraint: Be conservative: predict match=true only when the evidence strongly supports that both records are the same real-world entity and there is no meaningful contradiction
(Page 3).
Post-Processing the Labels
The goal of post-processing is to correct wrong labels or drop pairs whose labels are likely wrong. Two methods are studied:
-
Relabeling: This passes each pair through a second LLM prompt that verifies the assigned label using a conservative, evidence-based instruction (Figure 3). It can be in a
relabel
variant where the reviewed decision replaces the original, or arelabel-drop
variant where an unstable label is dropped (Page 4). -
Closure-based dropping: This filter identifies
bridge edges,
which are match pairs whose removal would split a connected component in an undirected graph of source records (Page 4).
Experimental Evaluation and Results
The experiments compare the performance of student models trained on machine-labeled sets against those trained on the benchmark training sets, judging fitness for use by checking if they achieve the same F1 score as the same student model trained on the benchmark training set
(Page 1).
Key findings include:
student models trained using machine-labeled sets perform approximately on par with models trained on the benchmark training sets, with the remaining differences in both directions staying below two F1 points
Active learning (Ditto) was shown to be the strongest pair selection strategy on four of the five benchmarks,
and it shifts the class balance toward positives
by up to 12.74 percentage points on DBLP-ACM (Page 10).
Post-processing showed that +Relabel is the best variant on Abt-Buy (+1.56 F1 over the unprocessed Active learning (Ditto) baseline)
while +Closure drop alone falls below the unprocessed baseline on every benchmark
(Page 6).
The cost analysis shows that Active learning (Ditto) is the cheapest at US28.31, ahead of Active learning (ML) at US30.16 and similarity search at US40.88
for labeling costs, while inference time shows Ditto is 41.5 to 534 times faster than directly using an LLM to perform the matching tasks
(Page 8).
Overall, the results indicate that current LLMs, when combined with a suitable pair-selection method, can substantially reduce or even eliminate the manual effort required to label use case-specific training data for entity matching
(Page 1).
REFERENCES
[1] Parikshit Bansal and Amit Sharma. 2023. Large Language Models as Annotators: Enhancing Generalization of NLP Models at Minimal Cost. arXiv:2306.15766 [cs.CL]
[2] Nils Barlaug and Jon Atle Gulla. 2021. Neural Networks for Entity Matching: A Survey. ACM Transactions on Knowledge Discovery from Data 15, 3 (2021), 52:1–52:37.
[3] Ursin Brunner and Kurt Stockinger. 2020. Entity Matching with Transformer Architectures - A Step Forward in Data Integration. In Proceedings of the 23rd International Conference on Extending Database Technology. OpenProceedings.org, Konstanz, Germany, 463–473.
[4] Peter Christen. 2012. Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection. Springer, Berlin, Heidelberg.
[5] Vassilis Christophides, Vasilis Efthymiou, Themis Palpanas, George Papadakis, and Kostas Stefanidis. 2020. An Overview of End-to-End Entity Resolution for Big Data. Comput. Surveys 53, 6 (2020), 127:1–127:42. doi:10.1145/3418896
[6] Vassilis Christophides, Vasilis Efthymiou, and Kostas Stefanidis. 2015. Entity Resolution in the Web of Data. Morgan & Claypool, San Rafael, CA, USA.
[7] Ahmed K. Elmagarmid, Panagiotis G. Ipeirotis, and Vassilios S. Verykios. 2007. Duplicate Record Detection: A Survey. IEEE Transactions on Knowledge and Data Engineering 19, 1 (2007), 1–16.
[8] Luyang Fang, Xiaowei Yu, Jiazhang Cai, Yongkai Chen, Shushan Wu, et al. 2026. Knowledge distillation and dataset distillation of large language models: Emerging trends, challenges, and future directions. Artificial Intelligence Review 59, 1 (2026), 17. doi:10.1007/s10462-025-11423-3
[9] Ivan P. Fellegi and Alan B. Sunter. 1969. A Theory for Record Linkage. J. Amer. Statist. Assoc. 64, 328 (1969), 1183–1210
[10] Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. ChatGPT Outperforms Crowd Workers for Text-Annotation Tasks. Proceedings of the National Academy of Sciences 120, 30 (2023), e2305016120
[11] Robert Isele and Christian Bizer. 2012. Learning Expressive Linkage Rules using Genetic Programming. Proceedings of the VLDB Endowment 5, 11 (2012), 1638–
[12] Pradap Konda, Sanjib Das, Paul Suganthan G. C., AnHai Doan, Adel Ardalan, et al. 2016. Magellan: Toward Building Entity Matching Management Systems.
Improvements for AI systems
-
The knowledge distillation workflow enables
student models trained on machine-labeled sets perform approximately on par with models trained on the benchmark training sets, with the remaining differences in both directions staying below two F1 points,
allowing for high-quality entity matching training data generation at a fraction of the manual effort. -
Active learning strategies, particularly
Active learning (Ditto),
canshift the class balance toward positives
and select more hard positives than benchmark sets, meaning the resulting training sets are composed of pairs that better capture difficult corner cases, leading to improved downstream performance on product benchmarks. -
Post-processing methods like
relabeling
ordrop
can improve downstream F1 scores by up to+1.56 F1
on Abt-Buy and WDC, indicating that applying a conservative, evidence-based review LLM step can refine the quality of machine-labeled data before training the final matcher. -
The system allows for inference time scaling where
Ditto is 41.5 to 534 times faster than directly using an LLM to perform the matching tasks,
enabling high-volume entity matching at asmall fraction of the cost of human labeling
by deploying compact student models like RoBERTa. -
The paper demonstrates that
the open-weight Kimi K2.6 stays within 0.85 F1 of the best hosted teacher model on every benchmark,
showing that knowledge distillation is effective even when using non-proprietary, open models as teachers, reducing reliance on closed-source LLMs for data labeling.
Abstract
Large language models (LLMs) achieve strong entity matching performance without task-specific training data, but applying them to large sets of candidate pairs is slow and costly. Matchers built on pretrained language models (PLMs), such as BERT, offer faster inference but require training data. We systematically study knowledge-distillation workflows in which an LLM teacher labels training pairs for a smaller student matcher. We vary pair selection, labeling budget, teacher model, correspondence post-processing, and student model across eight benchmarks, including unseen entities and non-English data. We compare students trained on machine-labeled data with matchers trained on the original benchmark training sets. In most cases, PLM-based matchers trained on LLM-labeled data perform similarly to those trained on benchmark sets. Pair selection matters most for small labeling budgets, where active learning is often most effective. An open-weight teacher trains competitive students, so distillation requires no closed-weight models. Compact PLM-based students compete with much larger LLM students on most tasks while requiring 34 to 459 times less inference time than direct LLM matching. On the two benchmarks with high shares of unseen products, PLM-based students substantially underperform their teachers, as do students trained on benchmark data. Under GPT-5.2 pricing, LLM labeling costs per training set average 5.86 to 8.11. These findings support knowledge distillation as a practical approach to reduce the effort of labeling task-specific training data while enabling efficient inference.
Sources
- Large Language Models as Annotators: Enhancing Generalization of NLP Models at Minimal Cost
- A Survey on Knowledge Distillation of Large Language Models
- DistillER: Knowledge Distillation in Entity Resolution with Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering