Language Models for Portuguese: A Systematic Mapping Study
cs.CL, cs.AI
Submitted: 2026-08-03
Updated: 2026-09-11
Comments: 37 pages; 7 figures; 8 tables
Code: https://github.com/diego-feijo/bertpt
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: In recent years, the rapid development of language models has transformed the field of Natural Language Processing through a wide range of applications.
Terminology
Abstract
In recent years, the rapid development of language models has transformed the field of Natural Language Processing through a wide range of applications. However, the development of language models has not progressed uniformly across all languages. In the case of the Portuguese language, there has recently been a growing effort by academia and companies to develop language models and create data resources for Portuguese. These efforts have resulted in the rise of an increasingly diverse ecosystem of language models for Portuguese. However, information on these models remains dispersed in scientific publications, technical reports, model repositories, and project documentation. This survey presents a systematic mapping study of language models developed for Portuguese, providing a comprehensive overview of the current state of the field. We map a total of 46 models, characterizing them by various aspects, including base model, architecture, computational resources, training datasets, licensing, code availability, data, and model weights. Furthermore, we analyzed the evolution and relationships among these models through a phylogenetic perspective, identified current research gaps and opportunities, and discussed future directions for the development of language models for Portuguese.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Sabi'a-3 Technical Report
- A Survey of Large Language Models for European Languages
- Sabi'a-2: A New Generation of Portuguese Large Language Models
- Curi'o-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
- PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data
- Tucano 2 Cool: Better Open Source LLMs for Portuguese
- Carolina: a General Corpus of Contemporary Brazilian Portuguese with Provenance, Typology and Versioning Information
- Amadeus-Verbo Technical Report: The powerful Qwen2.5 family models trained in Portuguese
- PeLLE: Encoder-based language models for Brazilian Portuguese based on open data
- Mono vs Multilingual Transformer-based Models: a Comparison across Several Language Tasks
- Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Introducing Bode: A Fine-Tuned Large Language Model for Portuguese Prompt-Based Task
- DeBERTa: Decoding-enhanced BERT with Disentangled Attention
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
- Mistral 7B
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- Cabrita: closing the gap for foreign languages
- Bactrian-X: Multilingual Replicable Instruction-Following Models with Low-Rank Adaptation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering