Chunk Twice, Embed Once: A Systematic Study of Segmentation and Representation Trade-offs in Chemistry-Aware Retrieval-Augmented Generation
cs.IR, cs.AI, physics.chem-ph
Submitted: 2025-06-13
Updated: 2026-09-17
Code: https://github.com/langchain-ai/langchain
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Language agents achieve superhuman synthesis of scientific knowledge
- Benchmarking Retrieval-Augmented Generation for Chemistry
- RAG-Enhanced Collaborative LLM Agents for Drug Discovery
- Financial Report Chunking for Effective Retrieval Augmented Generation
- ChemTEB: Chemical Text Embedding Benchmark, an Overview of Embedding Models Performance & Efficiency on a Specific Domain
- MMTEB: Massive Multilingual Text Embedding Benchmark
- MTEB: Massive Text Embedding Benchmark
- ChemQuests: A Curated Chemistry Question-Answer Database Extracted from ChemRxiv papers
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- MPNet: Masked and Permuted Pre-training for Language Understanding
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- GLU Variants Improve Transformer
- Unsupervised Dense Information Retrieval with Contrastive Learning
- ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- One Embedder, Any Task: Instruction-Finetuned Text Embeddings
- SciBERT: A Pretrained Language Model for Scientific Text
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG