Spatial Action Review: A Visual Analytics Dashboard for Auditing Language-to-Action Hand-offs in Electron Microscopy
cs.CV
Submitted: 2026-09-21
Updated: 2026-09-21
Code: https://github.com/Mohinta2892/SpatialActionReview
Terminology
Sources
- Qwen3-VL Technical Report
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
- MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models
- {\mu}-Bench: A Vision-Language Benchmark for Microscopy Understanding
- VISTAQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Evidence-Backed Video Question Answering
- Ferret: Refer and Ground Anything Anywhere at Any Granularity
- Multimodal Large Language Models for Bioimage Analysis
- From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models