CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs

arXiv:2510.06039 · cs.CL, cs.AI · Submitted 2026-08-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs".

Jane: The paper was written by Chengwei Wu, Xingrui Zhuo, Mingyang Gao, Xinghe Cheng, Zhichao Yan et al. from Nanjing University of Science and Technology and Beijing Academy of Artificial Intelligence and University of Science and Technology Beijing and Hefei University of Technology and Beijing University of Chemical Technology and Renmin University of China and Beihang University and Beijing University of Posts and Telecommunications and Griffith University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. I'm Tom, and as always, I'm joined by my co-host, Jane. Today we've got a fascinating paper that's all about Chinese language models, and it's got a title that's a mouthful: "CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs."

Jane: And I'm Jane. Tom, I have to say, when I first saw this title, I thought, "Okay, another dataset paper." But the more I dug into it, the more I realized this is actually tackling a really specific and thorny problem. It's not just about making a bigger dataset; it's about making a *better* one for Chinese specifically.

Tom: Exactly. And the core idea is this thing they call "data-text pairs." So, imagine you have a fact, like "The Great Wall is in China." In their dataset, that fact is stored as a structured triple—something like (Great Wall, located in, China)—and it's also paired with a natural language sentence that says the same thing. The paper's whole bet is that this kind of alignment is what Chinese LLMs are missing.

Jane: Right, and that's a big deal. Because if you think about English, there are tons of resources that link structured knowledge to text. But for Chinese, the paper argues, most of the data is just unstructured text. It's a huge pile of words with no explicit structure. So the models have to figure out the relationships between entities all on their own, which is really hard.

Lu: And that's where the unique challenges of Chinese come in. I'm Lu, by the way, for those just tuning in. The paper gives a great example of this. Take the word "行" – it can mean "to go" or "industry" depending on context. Or even more famously, "武汉市长江大桥" can be parsed as "Wuhan City / Yangtze River Bridge" or "Mayor of Wuhan / Jiang Daqiao." Without structured data, a model can easily get these ambiguities wrong.

Meng: So, they're basically saying that the lack of structured data is a root cause of errors in Chinese LLMs? I'm Meng, by the way. From an engineering standpoint, that makes sense. If you train a model on a pile of ambiguous text, it's going to learn ambiguous patterns. But if you give it clear, structured examples, it has a much better chance of learning the correct mappings.

Jane: Precisely, Meng. And that's why this paper is so exciting. They've built a massive resource to fix that. We're talking about seven million text pairs, which correspond to fifteen million triples. And they've organized it into four big domains: History and Politics, Humanities and Society, Technology and Economics, and Nature and Environment.

Tom: So it's not just big, it's also broad. And the authors are from a whole bunch of institutions—Nanjing University of Science and Technology, Beijing Academy of Artificial Intelligence, and several others. It feels like a real community effort to address this gap.

Lu: It is, and that breadth is crucial. If you only build a dataset about, say, history, you're only testing one slice of a model's knowledge. By spanning these four domains, they can start to ask questions about generalization. Does a model that's good at tech also understand humanities? That's a much more comprehensive evaluation.

Meng: So, for someone like me who has to actually deploy these models, this dataset isn't just an academic exercise. It's a tool. It gives us a way to stress-test a model's understanding of Chinese in a structured way before we put it in front of users.

Jane: And that's the promise. But building the dataset is only half the story. The other half is how they actually use it to evaluate the models, and that's where things get really interesting. We'll get into the nitty-gritty of their benchmark design next.

Tom: Stay with us, folks. We're just getting started.

Summary: Tom: Welcome back. We're digging into "CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs." So, Jane, we've established that the dataset is huge and well-organized. But what did they actually *do* with it?

Jane: Right, Tom. They didn't just build it and say, "Here you go." They built a whole evaluation framework on top of it, which they call CB-ECLLM. And they use it to test eight different Chinese LLMs on three specific tasks. The goal was to see how well these models can actually use structured knowledge.

Lu: And the three tasks they chose are really smart. You've got Knowledge Graph Completion, which is like a fill-in-the-blank for facts. You give the model a partial triple, like (Aisin-Gioro Xibao, date of birth, ?), and it has to pick the right year from a list of options. Then there's Question Answering, which is a more natural language version of the same idea. And finally, Triple-to-Text generation, where you give the model the structured facts and ask it to write a coherent sentence.

Meng: So, they're testing both understanding and generation. That's a solid approach. But what did they find? I'm guessing not all models are created equal.

Tom: You guessed right, Meng. The results in the paper are fascinating. First, they found that performance varies wildly depending on the task and the dataset. For example, InternLM2-7B is a champ at Triple-to-Text generation, but it struggles with Question Answering. So a model can be great at one thing and mediocre at another.

Jane: And here's the counter-intuitive part that really stood out to me. They found that almost all the models did *better* on Knowledge Graph Completion than on Question Answering. That's the opposite of what you usually see in English, where models are better at open-ended questions because they've seen so much unstructured text during pre-training.

Lu: That's a brilliant observation, Jane. It really highlights the unique challenge of Chinese. The unstructured text is so ambiguous that the models get tripped up. But when you give them the structure, the explicit relationships, they can anchor themselves and perform much more reliably. The structure actually *helps* them in Chinese, whereas in English, it might be a crutch.

Meng: So the structured data is acting like a safety net for the models. That's a really important insight for anyone trying to build a Chinese LLM. It suggests that injecting structured knowledge during training or inference could be a huge win.

Tom: And they also looked at model size. Generally, the bigger models, like Yi-9B, did better than smaller ones. But that wasn't a hard and fast rule, especially for Question Answering, where a smaller model like DeepSeek-7B could hold its own. It seems like for open-ended questions, raw parameter count isn't everything.

Jane: So, the base models have clear strengths and weaknesses. But the paper doesn't stop there. They wanted to see if they could improve these models, which brings us to their fine-tuning experiments.

Lu: And that's where the dataset's real value as a training resource, not just an evaluation tool, comes into play.

Tom: Exactly. Let's talk about what happens when you actually train these models on the CDTP data.

Improvements: Tom: We're back with "CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs." So, we saw the base models had some issues, especially with unstructured tasks. What's the fix?

Jane: The fix, Tom, is Supervised Fine-Tuning, or SFT. They took those same eight models and fine-tuned them on the CDTP dataset. And the results are pretty dramatic. After fine-tuning, almost all the models saw a significant boost in performance across all three tasks.

Meng: So, the dataset isn't just a good test; it's also good training data. That's a double win. But how much of an improvement are we talking about?

Tom: We're talking about big jumps. For example, look at the GLM-four-9B model on the History and Politics dataset. Its F1 score for Knowledge Graph Completion went from around zero point five one to over zero point six two. And on Question Answering, the accuracy jumped from about eighteen percent to over fifty-six percent. That's a massive leap.

Lu: And it's not just about the average. The paper points out that fine-tuning narrows the performance gap between the models. The smaller models, like Phi-two which were struggling, improved a lot. After SFT, models of similar sizes start to perform much more consistently. It's like the fine-tuning data is a great equalizer.

Jane: That's a really good point, Lu. It suggests that a lot of the performance difference we see in base models isn't necessarily about inherent capability, but about how well they've aligned their knowledge. The CDTP data helps them get that alignment right.

Meng: So, if I'm building a product, I could take a smaller, cheaper model, fine-tune it on this dataset, and get performance that rivals a much larger model? That has huge practical implications for cost and latency.

Tom: Exactly, Meng. But the most exciting part for me is the robustness test. They took a fine-tuned model, Yi-9B, and tested it on data it had never seen before—out-of-distribution data like YAGO3-ten for KGC and WebNLG for text generation.

Jane: And it held up. The fine-tuned model performed better on these unseen datasets than the base model. This is huge because it shows the model isn't just memorizing the training data. It's learning generalizable skills about how to handle structured knowledge and turn it into language.

Lu: That's the key. It's one thing to ace a test you've studied for. It's another to apply that knowledge to a new situation. The fact that SFT on CDTP improves performance on OOD data suggests the model is learning a deeper understanding of the relationship between facts and text, not just surface patterns.

Meng: So, the improvements are real and they generalize. That's a strong signal that this dataset could be a foundational resource for the Chinese NLP community. It's not just a benchmark; it's a training ground.

Tom: And that's what makes this paper so impactful. It's not just pointing out problems; it's providing a solution. But we should also think about what this means for the future of AI, and not just in a technical sense.

Jane: Absolutely. Let's bring in Lalam to give us a broader perspective on the cultural impact.

Conclusion: Tom: Well, we've reached the end of our discussion on "CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs." It's been a fantastic conversation. Jane, can you wrap it up for us?

Jane: I'd love to, Tom. So, to recap, this paper gives us a massive, high-quality dataset of seven million aligned text-triple pairs. It's a tool for both evaluating and training Chinese LLMs. The evaluation showed that Chinese models struggle with ambiguity in unstructured text, but they perform much better when given structured knowledge. And the fine-tuning experiments proved that using this dataset to train models makes them not only better but also more robust to new, unseen data.

Lu: And that robustness is the real prize. It means we're moving towards models that can handle the messy, ambiguous reality of human language, not just clean, curated examples. This is a big step for Chinese NLP.

Meng: From my side, the practical impact is clear. This dataset gives us a reliable way to build and test Chinese language models that are actually ready for real-world applications, whether that's customer service, information retrieval, or content generation.

Lalam: I agree with all of you. But I'd like to add that the cultural impact is just as profound. Chinese is a language rich with history, idioms, and context-dependent meanings. A dataset like CDTP, which forces models to grapple with that complexity, is essential for creating AI that doesn't just translate Chinese, but truly understands it. It's about building technology that respects and preserves the nuance of the culture it serves. This isn't just about better chatbots; it's about making sure that as AI becomes more integrated into our lives, it can do so in a way that is culturally competent and aware.

Tom: That's a beautiful way to put it, Lalam. And it's a great note to end on. The paper is a significant contribution, and we're excited to see where this line of research goes. We'll be keeping an eye on the future work they've outlined, like expanding the dataset to more domains and exploring cross-modal signals.

Jane: So, we'll say goodbye to this paper and get ready to dive into the next one. Thanks for listening, everyone.

Tom: See you next time.

Chengwei Wu, Xingrui Zhuo, Mingyang Gao, Xinghe Cheng, Zhichao Yan, Jiapu Wang

Nanjing University of Science and Technology · Beijing Academy of Artificial Intelligence · University of Science and Technology Beijing · Hefei University of Technology · Beijing University of Chemical Technology · Renmin University of China · Beihang University · Beijing University of Posts and Telecommunications · Griffith University

cs.CL, cs.AI

Submitted: 2026-08-16

Updated: 2026-08-18

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 40/100

The gist: "(i) enriching Chinese corpora with high-quality structured information; (ii) enabling fine-grained evaluation tailored to knowledge-driven tasks; and (iii) supporting multi-task fine-tuning to

Key concepts

Data-Text Pairs
This is a core concept where a piece of structured knowledge, like a fact stored as a triple (Subject, Predicate, Object), is paired with the corresponding natural language sentence that expresses the same fact. This alignment helps models learn relationships that are otherwise missing in unstructured text.
Knowledge Graph Completion (KGC)
This is an evaluation task where a model must fill in missing information within a structured knowledge graph. For example, given partial facts, the model selects the correct missing piece of information from a list of options. It tests the model's ability to retrieve and complete explicit facts.
Supervised Fine-Tuning (SFT)
SFT is a training method where existing models are further trained on specific datasets like CDTP. This process significantly improves performance by aligning the model's knowledge with the structured data, leading to substantial boosts in accuracy and consistency across different tasks.
Robustness Test
This involves testing a fine-tuned model on data it has never seen before, called out-of-distribution data. The goal is to see if the model learns generalizable skills rather than just memorizing the training examples, which indicates a deeper understanding of language relationships.

Terminology

Summary

Summary

The paper introduces the Chinese Data-Text Pair (CDTP) dataset and the Comprehensive Benchmark for Evaluating Chinese Large Language Models (CB-ECLLM). The authors state that "Large Language Models (LLMs) have achieved remarkable success across a wide range of natural language processing tasks. However, Chinese LLMs face unique challenges, primarily due to the dominance of unstructured free text and the lack of structured representations in Chinese corpora. They argue that existing benchmarks for LLMs partially assess Chinese LLMs, but they are still predominantly English-centric and fail to address the unique linguistic characteristics of Chinese, lacking structured datasets essential for robust evaluation."

To address these challenges, the paper presents CB-ECLLM, which is based on the newly constructed Chinese Data-Text Pair (CDTP) dataset. The CDTP dataset comprises over 7 million aligned text pairs, each consisting of unstructured text coupled with one or more corresponding triples, alongside a total of 15 million triples spanning four critical domains. These four domains are History and Politics, Humanities and Society, Technology and Economics, and Nature and Environment. The core contributions of CDTP are described as threefold: "(i) enriching Chinese corpora with high-quality structured information; (ii) enabling fine-grained evaluation tailored to knowledge-driven tasks; and (iii) supporting multi-task fine-tuning to assess generalization and robustness across scenarios, including Knowledge Graph Completion, Triple-to-Text generation, and Question Answering."

The paper highlights the unique challenges of the Chinese language, noting that a single Chinese expression can yield multiple plausible interpretations depending on context. It provides examples such as "行 (meaning to do/go or profession/industry) and 武汉市长江大桥 (which can be interpreted as the mayor of Wuhan named Jiang Daqiao or the Yangtze River Bridge in Wuhan). The authors state that This absence of structured signals constrains performance on a wide range of knowledge-intensive tasks, including knowledge retrieval, relation inference, and text generation."

The dataset construction process involves two phases: data collection and data processing. For data collection, the authors compile a raw data corpus of approximately 27 million entries from various online encyclopedic sources, such as Baidu Baike and Sogou Baike. The process begins with aggregating a set of seed entries from three sources: OwnThink, KgClue, and WudaoCorpora, yielding a total of 56 million seed entries. A customized Baike crawler then systematically retrieves both textual descriptions and triples for each seed entry from a precompiled list, employing a depth-first search strategy to recursively traverses hyperlinked entities. This process yields approximately 30 million triples involving around 15 million unique entities.

The data processing phase includes data cleaning and quality enhancement. For cleaning, the authors target two prevalent types of noise: redundant triples—relation assertions unsupported by textual evidence and redundant descriptions—repetitive or irrelevant content that deviates from the corresponding structured knowledge. For redundant triples, they examine each data pair by checking whether the semantic content of each triple is supported by the accompanying text and remove triples that lack textual grounding. For redundant descriptions, they compute a position index by locating the character offset of the last matched triple and normalizing it by the total length of the text, retaining pairs with an index above 0.75. For quality improvement, they employ Relation Filtering (retaining only high-frequency relation types), Manual Verification (with seven trained annotators and double annotation ensuring over 90% agreement), and Search Validation (retaining a triple only if both the head and tail entities co-occur in the top 10 external web results).

The benchmark evaluates eight Chinese LLMs: GLM-4-9B, Yi-9B, Qwen1.5-7B, InternLM2-7B, Llama-3-8B, Baichuan2-7B, Phi-2, and DeepSeek-7B. The evaluation covers three tasks: Question Answer (QA), Knowledge Graph Completion (KGC), and Triple to Text Generation (T2T). For QA and KGC, each query is cast as a multiple-choice problem with one correct answer and nine distractors, with distractors drawn from KG neighbors and type-/frequency-matched entities, filtered for semantic feasibility, and enriched with polysemy-targeted candidates to capture Chinese-specific ambiguities. Metrics used include MRR, Hits@1, and F1 Score for QA and KGC, and BLEU, ROUGE, and METEOR for T2T.

The experiments address three research questions. For RQ1 (Effectiveness), the results show that Performance varies significantly across tasks and datasets, with some models excel in T2T tasks but underperform in KGC or QA tasks. The authors observe that all Chinese LLMs consistently outperform in KGC compared to QA tasks within the MRR metric, which stands in contrast to trends commonly seen in English. They also find that the larger-scale LLMs exhibit better performance, although this trend does not hold as strongly for QA tasks.

For RQ2 (Supervised Fine-Tuning), the authors fine-tune the models using DeepSpeed with ZeRO Stage 2; Batch Size: 8; Learning Rate: 9.65 × 10−6 with a cosine learning rate scheduler; Max Sequence Length: 2048; Mixed Precision: BF16; Number of Epochs: 3 on 8 NVIDIA H100 GPUs. Results show that SFT models consistently outperform their base counterparts on the T2T task across all datasets, and that Fine-tuning improves the performance by reducing the performance gaps between different models, leading to more balanced results across multiple tasks.

For RQ3 (Robustness), the authors test the robustness of Chinese LLMs Yi-9B on OOD data, before and after SFT. They use YAGO3-10 for KGC, HotpotQA for QA, and WebNLG for T2T. Results show that SFT significantly enhances model performance across all tasks (KGC, QA, T2T) when tested with OOD data, and that the extent and nature of the improvements vary by task type, with The QA task exhibit[ing] the most substantial gains.

The paper concludes that CDTP comprises over 7 million meticulously aligned text pairs... totaling 15 million triples spanning four critical domains and that This dataset offers a comprehensive benchmark for evaluating reasoning, generalization, and structured knowledge comprehension in Chinese LLMs, paving the way for AGI. The authors acknowledge one limitation: the proposed CDTP dataset only contains textual inputs and structured triples, without incorporating cross-modal signals such as images or audio, which constrains its applicability in multimodal settings. Future directions include Enhanced Domain and Entity Coverage, Cross-Domain Generalization, and Linguistic and Cultural Enrichment.

Improvements for AI systems

Based on the scientific paper, here are the specific improvements I can make to AI systems and the resulting capabilities:

1. Implement a Structured Knowledge Alignment Layer

  • Improvement: Integrate a data-processing pipeline that filters out triples not explicitly grounded in accompanying text (redundant triples) and removes text descriptions lacking triple support (redundant descriptions), using position-based indexing and cross-referencing with relation tables.

  • Resulting Capability: The AI system will produce outputs strictly aligned with structured knowledge, reducing hallucination and improving factual consistency in knowledge-intensive tasks.

2. Add Chinese-Specific Disambiguation Mechanisms

  • Improvement: Incorporate polysemy-targeted distractors and context-dependent disambiguation features into the evaluation framework, drawing from the paper’s examples (e.g., “行” meaning “to do” vs. “profession/industry”).

  • Resulting Capability: The AI system will better handle Chinese linguistic ambiguities, such as character-level polysemy and syntactic ambiguity (e.g., “武汉市长江大桥” vs. “武汉市长/江大桥”), improving accuracy in QA and relation inference tasks.

3. Develop a Multi-Task Fine-Tuning Framework

  • Improvement: Use the CDTP dataset (7 million text-triple pairs, 15 million triples across four domains) to fine-tune models jointly on Knowledge Graph Completion (KGC), Triple-to-Text Generation (T2T), and Question Answering (QA) tasks, following the paper’s SFT configuration (DeepSpeed, batch size 8, learning rate 9.65e-6, 3 epochs).

  • Resulting Capability: The AI system will generalize better across structured and unstructured tasks, with improved performance on out-of-distribution (OOD) data, as demonstrated by the paper’s robustness experiments (e.g., YAGO3-10, HotpotQA, WebNLG).

4. Implement a Domain-Aware Evaluation and Training Strategy

  • Improvement: Categorize training and evaluation data into four domains (History and Politics, Humanities and Society, Technology and Economics, Nature and Environment) and apply domain-specific fine-tuning and stratified sampling.

  • Resulting Capability: The AI system will show more consistent performance across diverse domains, reducing performance variance and improving reliability in specialized applications (e.g., legal, medical, or technical text).

5. Add a Robustness Verification Module

  • Improvement: Integrate a post-training evaluation step that tests the model on OOD datasets (e.g., YAGO3-10 for KGC, HotpotQA for QA, WebNLG for T2T) to verify stability under distributional shifts.

  • Resulting Capability: The AI system will maintain accuracy and reliability when deployed in real-world scenarios where input data deviates from training distributions, as evidenced by the paper’s finding that SFT improves OOD performance across all tasks.

  • Generate Factually Consistent Chinese Text: Produce natural language descriptions from structured triples (T2T) with high BLEU/ROUGE/METEOR scores (e.g., >0.78 ROUGE L after SFT), ensuring every generated sentence is grounded in verifiable facts.

  • Complete Knowledge Graphs with High Precision: Predict missing entities or relations in Chinese KGs with Hits@1 >0.70 and MRR >0.80 after fine-tuning, even for ambiguous or polysemous entities.

  • Answer Chinese Questions with Contextual Awareness: Handle open-ended QA tasks with improved accuracy (ACC >0.60) by leveraging structured knowledge and disambiguation features, reducing errors from linguistic ambiguity.

  • Generalize to Unseen Domains and Data: Maintain performance on OOD datasets (e.g., YAGO3-10, HotpotQA, WebNLG) after fine-tuning, demonstrating robustness in real-world deployment where data distribution shifts.

  • Support Multi-Domain Applications: Operate effectively across history, politics, humanities, technology, economics, nature, and environment domains, with balanced performance and reduced domain-specific biases.

  • Provide Interpretable and Reliable Outputs: Align structured knowledge with unstructured text, enabling traceable reasoning and reducing hallucination in knowledge-intensive tasks, which is critical for high-stakes applications like legal or medical AI.

Sources

Related papers