MolLangData: A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method
cs.CL, cs.AI, q-bio.BM
Submitted: 2026-02-02
Updated: 2026-09-26
Code: https://github.com/TheLuoFengLab/MolLangData
Terminology
Sources
- Galactica: A Large Language Model for Science
- ChemLLM: A Chemical Large Language Model
- A Molecular Multimodal Foundation Model Associating Molecule Graphs with Natural Language
- MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- RetroDFM-R: Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning
- MolTextNet: A Two-Million Molecule-Text Dataset for Multimodal Molecular Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering