Arabic Morphosyntactic Tagging and Dependency Parsing with Large Language Models

arXiv:2603.16718 · cs.CL · Submitted 2026-03-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Arabic Morphosyntactic Tagging and Dependency Parsing with Large Language Models".

Jane: The paper was written by Mohamed Adel, Bachar Alhafni and Nizar Habash from New York University Abu Dhabi and Mohamed bin Zayed University of Artificial Intelligence.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Methodology and ICL: Tom: The researchers designed several ways to improve the LLM's performance, particularly focusing on retrieval-based in-context learning or ICL, which is a huge methodological improvement for how we guide these systems.

Jane: They are testing different selection methods, like retrieving examples based on how similar the input sentences look in character structure versus how similar they are in meaning, and that's a very smart way to approach the problem.

Meng: From a practical standpoint, choosing a high-similarity example over a random one is much more efficient because it gives the model specific context that is directly relevant to solving the problem at hand instead of guessing randomly.

Lu: It suggests that we can teach LLMs not just by giving them general knowledge, but by strategically showing them exactly how to handle specific types of linguistic structures they are struggling with through a targeted example.

Lalam: This ability to improve targeted linguistic competence allows us to move beyond just a surface-level translation and start toward a truly comprehensive understanding of the language's logic.

Tom: The focus is clearly on making the model more capable through context, rather than just relying on its raw training data that can be vague or unorganized.

Jane: It’s fascinating how they are measuring both the visual similarity of words and their semantic content to see which approach helps the LLM learn better in this difficult language.

Lu: It gives us a way to probe whether an LLM is truly learning a rule, or if it's just memorizing patterns it saw during its massive training phase.

Meng: The engineering implication here is that we are building a guidance mechanism—a highly specific set of instructions—to ensure the AI performs reliably in real-world applications where context matters.

Lalam: And this allows us to build tools for Arabic speakers that go beyond simple automated text generation, allowing for genuine cultural and linguistic depth in our technology.

Tom: It’s clear they found that ICL is a powerful lever that helps overcome the inherent limitations of just zero-shot prompting, which is a major breakthrough.

Jane: Now, let's look at the actual results to see how much these improvements translate into real performance gains on the test sets.

Results and Trade-offs: Tom: When looking at the test results, we see that while LLMs are getting better with ICL, there’s a clear gap between achieving feature-level accuracy and hitting the "all tags" score.

Jane: That "all tags" metric is super strict—it requires every single one of those fourteen features to be perfectly correct for every word in the sentence, which is incredibly difficult in practice.

Meng: And that's why I think operationalizing these results is challenging; even if we get high accuracy on individual features, hitting that perfect the full-bundle score is a lot of work for a complex language system.

Lu: It points to a fundamental limitation in how LLMs process deep structural dependencies versus their ability to memorize patterns, suggesting they struggle with the long-term logic of the sentence structure.

Lalam: The paper also shows that the trade-offs between the performance of open-weight models and proprietary models are very different, which is something we need to consider when deciding how we build future language tools for Arabic speakers.

Tom: It's a clear distinction—the open-source models improve significantly with context, while the proprietary ones are already so strong that they reach a baseline level of performance on their own.

Jane: The difficulty in achieving "all tags" shows that even when we guide the model perfectly, it is still struggling with deep structural consistency in the way Arabic requires it.

Lu: It tells us that we need to build more sophisticated parsers than simple statistical models, because the current AI is still missing the full picture of dependency chains.

Meng: This data helps us decide whether to invest in massive open-source infrastructure or if a specialized proprietary service is needed for peak performance on tasks like this.

Lalam: And knowing these strengths and weaknesses allows us to tailor solutions that truly serve the specific needs of Arabic speakers, rather than using one generic tool.

Tom: It's clear this research is providing a very honest look at where current AI capabilities lie in the realm of advanced linguistic understanding today.

Jane: Let’s wrap up by summarizing what all these findings mean for the future of Arabic NLP and technology.

Conclusion and Future Outlook: Tom: So, looking across all this data from "Arabic Morphosyntactic Tagging and Dependency Parsing with Large Language Models," it’s clear that while Large Language Models are making steady progress in understanding Arabic's complex structure, the challenges remain significant.

Jane: I think the most important thing is that we have a detailed map of where LLMs succeed—especially with guided learning—and where they are still struggling to achieve complete structural accuracy.

Lu: This research shows us exactly where future work needs to focus, pushing us toward improving how models handle those deep syntactic constraints in morphologically rich languages.

Meng: It gives us actionable data on which specific retrieval strategies are most effective, which is crucial for anyone trying to build a scalable and efficient AI system for the Arabic language.

Lalam: We have to remember this paper's full title—"Arabic Morphosyntactic Tagging and Dependency Parsing with Large Language Models"—and recognize that we are moving toward a world where our technology can better serve everyone's right to understand their native language.

Tom: It sounds like we are seeing a future where the AI is not just predicting words, but truly understanding the grammatical relationships between them.

Jane: It’s a testament to how much progress is being made, even though the path forward requires careful, targeted refinement of the methods used.

Lu: The paper shows that we are no longer just memorizing patterns; we are building systems that can reason about linguistic structure itself.

Meng: I feel like this is the moment where AI moves from a theoretical concept to a practical tool, giving us clear pathways for real-world implementation in Arabic language processing.

Lalam: We hope that these findings pave the way for cultural technologies that fully reflect and support the beauty of Arabic language.

Tom: It's been a fascinating discussion on what we can expect from LLMs in such an intricate language like Arabic, and we are so excited about what's next.

Conclusion: Tom: So, after diving into all these results for "Arabic Morphosyntactic Tagging and Dependency Parsing with Large Language Models," we can say that while AI is incredibly powerful, it’s still learning the full nuances of Arabic structure. It's not just about having a big model; it's about making sure that the guidance we give it—the in-context examples—truly helps make a model smarter.

Jane: Exactly, and I think this research showed us that we are moving far beyond simple word-by-word translation. It’s really helping us build a foundation where the deep grammatical relationships in Arabic can truly be understood by AI tools for the people who speak them.

Lu: This work demonstrates that we're no longer just relying on raw patterns seen in training data; we are actively building systems that can reason about linguistic structure itself, which is a huge conceptual leap.

Meng: The practical impact of having clear guidance—seeing exactly how retrieval-based ICL works—means we can now build much more efficient and reliable AI systems for real-world applications.

Lalam: We truly hope that the advances discussed in this paper pave the way for cultural technologies that fully reflect and support the beauty and complexity of Arabic language, making them accessible to everyone.

Tom: It’s been a fascinating journey through these findings, seeing how much we can push these models in such an intricate language like Arabic.

Jane: And even though the path forward requires careful refinement—especially in handling those tricky tokenization errors—it' is definitely a breakthrough that this sets the bar for what we expect from advanced AI.

Lu: I think it’s exciting to see that we are moving toward building systems that can reason about linguistic structure, rather than just memorizing patterns.

Meng: It gives us actionable data, and I see a clear pathway for real-world implementation in how these tools will run efficiently.

Lalam: We want to make sure the machine understands the spirit of the language, not just its individual parts, and this research helps us get there.

Tom: It’s been a great discussion about how far we've come with LLMs in handling complex languages like Arabic.

New York University Abu Dhabi · Mohamed bin Zayed University of Artificial Intelligence

cs.CL

Submitted: 2026-03-17

Updated: 2026-09-03

Comments: Accepted to EMNLP 2026

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 84/100

The gist: The paper investigates advanced Natural Language Processing tasks—specifically morphosyntactic tagging and dependency parsing—for Arabic using Large Language Models (LLMs).

Key concepts

In-Context Learning (ICL)
This is a methodological improvement where LLMs are guided by strategically showing them targeted examples within the prompt. Instead of relying only on massive training data, ICL teaches the model how to handle specific linguistic structures, allowing it to learn complex patterns and improve its competence.
Morphosyntactic Tagging and Parsing
This refers to AI analyzing a language, such as Arabic. Tagging identifies the grammatical function of every word, while parsing maps the structural relationships between those words. This allows technology to move beyond simple translation toward understanding the language's deep logic.
'All Tags' Metric
This is an extremely rigorous performance measure. It requires that all fourteen specific grammatical features associated with a word must be perfectly accurate within the entire sentence. Achieving this score indicates a high level of structural consistency for AI tools.
Retrieval-Based Selection
This method selects training examples by comparing the new input against existing data. The system retrieves relevant examples based on either visual similarity (character structure) or semantic similarity (meaning), ensuring the model receives highly pertinent context.

Terminology

Summary

The paper investigates advanced Natural Language Processing tasks—specifically morphosyntactic tagging and dependency parsing—for Arabic using Large Language Models (LLMs). Given Arabic's rich morphological complexity, accurate linguistic analysis is critical for downstream applications. The research evaluates how different LLMs, such as Llama4, Qwen3, GPT5.2, and Gemini3, perform on these tasks under varying prompting conditions to determine the optimal approach for robust Arabic language understanding.

Morphosyntactic Tagging Accuracy Across Features

The study provides a detailed analysis of Per-feature morphosyntactic tagging accuracy across models and prompting settings using Table 13. This table measures performance across numerous linguistic features, including POS (Part-of-Speech), PER (Person), GEN (Gender), NUM (Number), ASP (Aspect), and MOD/VOX/STT/CAS/PRC. The results demonstrate a clear dependency on both the model used and the number of shots provided in the prompt.

Key observations regarding feature accuracy include:

  • Impact of Shots: Performance generally improves as the number of shots increases, although specific features show varying sensitivities. For instance, across all models tested, performance metrics are presented for 0 shots and 10 shots.

  • Model Comparison: Models like Gemini3 consistently achieve high scores across most features when provided with sufficient prompting context. For example, in the 10-shot setting, Gemini3 achieves an F1 score of 99.8 for the CAS feature and 99.8 for the PRC feature.

  • Feature Difficulty: Certain features appear to be more challenging or variable across models compared to others, requiring careful evaluation of model robustness across all linguistic dimensions.

Dependency Parsing Performance Across Genres and Settings

The research evaluates dependency parsing capability using the Labeled Attachment Score (LAS) as the primary metric, detailed in Table 14. This analysis spans multiple Arabic genres, including Odes, Quran, Hadith, and texts from various historical periods (6th–12th through 21st centuries). The goal is to identify generalized performance across diverse linguistic domains.

The LAS breakdown highlights several critical factors influencing parsing accuracy:

  • Genre Variation: Performance varies significantly depending on the source material. For example, when comparing aggregate results, the Macro Average LAS scores provide a comprehensive view of how performance changes across data conditions.

  • Model Superiority: The study compares multiple parsers—CamelParser, Gemini3, and an Oracle setting—to benchmark current capabilities. The Oracle setting represents the best achievable performance for simple genre-based system selection.

  • Structural Factors: Performance is analyzed based on structural characteristics, such as sentence length (Short, Mid, Long) and training data size (Train S to Train XL). For instance, the aggregate results show that while models perform well overall, specific combinations of these factors define the boundaries of current NLP capabilities.

Comparative Analysis and Limitations

The study employs a rigorous comparative framework to assess LLM suitability for Arabic NLP tasks. The inclusion of multiple prompting settings (0 shots vs 10 shots) and diverse datasets (e.g., CA vs MSA, different centuries) allows researchers to isolate the effects of domain, supervision level, and input complexity.

The analysis also introduces specific selection techniques for evaluation:

  • Oracle: This setting identifies the best possible simple genre-based system selection.

  • Formula: This technique assumes a simple selection rule: if Length=Mid OR Train=XL OR Period=21st, then select CamelParser’s output, else Gemini3’s output.

These comprehensive results allow researchers to understand not only the peak performance of state-of-the-art models but also the specific conditions under which linguistic analysis—both morphosyntactic tagging and dependency parsing—is most reliable for Arabic.

Improvements for AI systems

Based on a rigorous analysis of the provided data—which spans few-shot scaling, deep morphosyntactic tagging, and cross-genre robustness in Arabic NLP—the following specific improvements must be implemented to elevate the AI system from state-of-the-art benchmarking to production reliability.


The Improvement: Implement a Hierarchical, Feature-Weighted Prompting Mechanism rather than treating feature tagging (POS, PER, NUM, GEN) as independent prediction tasks fed by a single prompt structure. The system must learn the interdependence between features.

System Capability:

The improved system will perform Constraint-Guided Morphosyntactic Tagging. For example, if the model predicts a specific Person feature (PER) and a specific Number feature (NUM), it will dynamically constrain the allowed values for other related features (like GEN or POS) based on established Arabic grammatical rules before final output generation. This moves beyond simple accuracy reporting to enforcing structural validity, drastically reducing hallucinated or grammatically inconsistent tags, especially in highly nuanced texts like Quranic verses.

By integrating these three components, the resulting AI system moves beyond being a sequence tagger and becomes a Contextually Aware, Grammatically Constrained Arabic Linguistic Engine. It will not only achieve high average accuracy but will provide guaranteed structural integrity across vastly different linguistic domains and input complexities, which is the critical difference between an academic benchmark and mission-critical industrial deployment.

Sources

Related papers