Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model
cs.CL
Submitted: 2026-01-11
Updated: 2026-09-18
Comments: 15 pages, 3 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multimodal aspect-based sentiment analysis (MABSA) aims to identify aspect-level sentiments by jointly modeling textual and visual information, which is essential for fine-grained opinion
Terminology
Abstract
Multimodal aspect-based sentiment analysis (MABSA) aims to identify aspect-level sentiments by jointly modeling textual and visual information, which is essential for fine-grained opinion understanding in social media. Existing approaches mainly rely on discriminative classification with complex multimodal fusion, yet they lack explicit sentiment explainability. In this paper, we reformulate MABSA as a generative and explainable task, proposing a unified framework that simultaneously predicts aspect-level sentiment and generates natural language explanations. Based on multimodal large language models (MLLMs), our approach employs a prompt-based generative paradigm, jointly producing sentiment and explanation. To further enhance aspect-oriented reasoning capabilities, we propose a dependency-syntax-guided sentiment cue strategy. This strategy prunes and textualizes the aspect-centered dependency syntax tree, guiding the model to distinguish different sentiment aspects and enhancing its explainability. To enable explainability, we use MLLMs to construct explanation-augmented datasets for fine-tuning. Experiments show that our approach not only achieves overall gains in sentiment classification accuracy, but also produces coherent and aspect-grounded explanations.
Sources
- GateMABSA: Aspect-Image Gated Fusion for Multimodal Aspect-based Sentiment Analysis
- Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis
- hyperdoc2vec: Distributed Representations of Hypertext Documents
- Efficient Estimation of Word Representations in Vector Space
- ASTE Transformer Modelling Dependencies in Aspect-Sentiment Triplet Extraction
- LLaMA: Open and Efficient Foundation Language Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- AdaptiSent: Context-Aware Adaptive Attention for Multimodal Aspect-Based Sentiment Analysis
- ZEN 2.0: Continue Training and Adaption for N-gram Enhanced Text Encoders
- Learning Span-Level Interactions for Aspect Sentiment Triplet Extraction
- Large Language Models Enhanced by Plug and Play Syntactic Knowledge for Aspect-based Sentiment Analysis
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Qwen3 Technical Report
- BERTScore: Evaluating Text Generation with BERT
- Revisiting Structured Sentiment Analysis as Latent Dependency Graph Parsing
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering