ViLegalExpert: A Large-Scale Benchmark for Vietnamese Legal Retrieval and Question Answering from Real-World Consultations
cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Terminology
Sources
- LEGAL-BERT: The Muppets straight out of Law School
- MiniRAG: Towards Extremely Simple Retrieval-Augmented Generation
- LightRAG: Simple and Fast Retrieval-Augmented Generation
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Multilingual Denoising Pre-training for Neural Machine Translation
- A Replication Study of Dense Passage Retriever
- PhoBERT: Pre-trained language models for Vietnamese
- VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering
- Multi-style Generative Reading Comprehension
- S-Net: From Answer Extraction to Answer Generation for Machine Reading Comprehension
- BARTpho: Pre-trained Sequence-to-Sequence Models for Vietnamese
- mT5: A massively multilingual pre-trained text-to-text transformer
- QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension
- DCMN+: Dual Co-Matching Network for Multi-choice Reading Comprehension
- JEC-QA: A Legal-Domain Question Answering Dataset
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering