GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval
cs.LG, cs.AI, cs.CL
Submitted: 2026-08-23
Updated: 2026-08-23
Comments: 7 pages, 2 figures, IEEE dual-column format. Submitted to arXiv for preprint distribution
Code: https://github.com/isaacus/mleb
License: http://creativecommons.org/licenses/by/4.0/
The gist: We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval.
Terminology
Abstract
We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval. GreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1),demonstrating competitive performance among models under 1B parameters. Our approach combines a two-stage training pipeline that first distills knowledge from a larger teacher model into a compact student architecture, then applies domain-specific fine-tuning with hard negative mining; a carefully curated dataset of 3.4 million query-passage pairs, including 150,000 human-curated samples across diverse legal jurisdictions; and an efficient inference architecture supporting multiple quantization levels (BF16, INT8, binary) enabling deployment in resource-constrained environments. We provide detailed analysis of our training methodology, architectural choices, and comprehensive evaluation across legal retrieval tasks. Our results demonstrate that domain-specific training with high-quality data can improve performance for specialized domain applications
Sources
- LEGAL-BERT: The Muppets straight out of Law School
- MTEB: Massive Text Embedding Benchmark
- Distilling the Knowledge in a Neural Network
- Contrastive Learning with Hard Negative Samples
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- C-Pack: Packed Resources For General Chinese Embeddings
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- Efficient Estimation of Word Representations in Vector Space
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
- RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering
- SciBERT: A Pretrained Language Model for Scientific Text
- CodeBERT: A Pre-Trained Model for Programming and Natural Languages
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks