Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning
cs.CL, cs.AI
Submitted: 2025-05-14
Updated: 2026-09-24
Terminology
Sources
- Entanglement generation in capacitively coupled Transmon-cavity system
- Tamil-Llama: A New Tamil Language Model Based on Llama 2
- ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model
- An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference
- Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
- Airavata: Introducing Hindi Instruction-tuned LLM
- Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP
- SuperBPE: Space Travel for Language Models
- WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models
- FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models
- Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learning
- The ultraspherical rectangular collocation method and its convergence
- Attention Is All You Need
- Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering