Adapting Vision-Language Models for Human-Readable XAI in Industrial Object Detection
cs.CV, cs.AI
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/freddyfernandes/Finetuned_VLM_assistant
Terminology
Sources
- Comparison of Open-Source and Proprietary LLMs for Machine Reading Comprehension: A Practical Analysis for Industrial Applications
- Integrating LLMs for Explainable Fault Diagnosis in Complex Systems
- On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
- Efficacy of Synthetic Data as a Benchmark
- LLMs for Explainable AI: A Comprehensive Survey
- Training language models to follow instructions with human feedback
- Synthetic Data Generation Using Large Language Models: Advances in Text and Code
- Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026)
- Qwen2.5-VL Technical Report
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- What matters when building vision-language models?
- Black-box Explanation of Object Detectors via Saliency Maps
- A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
- LoRA: Low-Rank Adaptation of Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models