Efficient LLM Distillation for Bangladesh Legal Context: A Smartphone-Compatible Retrieval-Augmented Generation Model
cs.CL
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 10 pages, 6 figures, 8 tables
Code: https://github.com/ggml-org/llama.cpp
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- TinyBERT: Distilling BERT for Natural Language Understanding
- BanglaBERT: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in Bangla
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation
- A Hybrid Approach to Information Retrieval and Answer Generation for Regulatory Texts
- Distilling the Knowledge in a Neural Network
- Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid Search
- LawPal : A Retrieval Augmented Generation Based System for Enhanced Legal Accessibility in India
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
- Memorization Dynamics in Knowledge Distillation for Language Models
- QLoRA: Efficient Finetuning of Quantized LLMs
- Which Quantization Should I Use? A Unified Evaluation of llama.cpp Quantization on Llama-3.1-8B-Instruct
- Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
- LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
- LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
- Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering