CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph
cs.CL
Submitted: 2026-06-29
Updated: 2026-09-18
Comments: EMNLP 2026 Main
Code: https://github.com/goto456/stopwords
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- DeepSeek-V3 Technical Report
- ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
- Textbooks Are All You Need
- FastWARC: Optimizing Large-Scale Web Archive Analytics
- Distilling the Knowledge in a Neural Network
- ChineseWebText: Large-scale High-quality Chinese Web Text Extracted with Effective Evaluation Model
- Training Compute-Optimal Large Language Models
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- OpenAI GPT-5 System Card
- Scaling Laws for Neural Language Models
- CFBenchmark: Chinese Financial Assistant Benchmark for Large Language Model
- Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources
- The Llama 3 Herd of Models
- CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- C-Pack: Packed Resources For General Chinese Embeddings
- Qwen2.5 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering