LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning
Yupan Ding, Jing Xiao, Zhenyuan Zhang, Chaofeng Chen, Liang Liao, Gui-Song Xia, Mi Wang
Wuhan University · Xi'an University of Electronic Science and Technology
cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 100/100
The gist: Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image
Terminology
Summary
Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarthBench, a benchmark containing approximately 120k question-answering samples derived from 117k unique images. Its sequences average 15.14 frames and extend to 30 frames, covering 12 tasks across evolution summarization, spatial reasoning, anomaly identification, and logical prediction. A 30k-sample subset further provides structured reasoning traces linking key frames and changed regions to final answers. We develop LongEarth through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought supervision. Building on LongEarth, LongEarth-R1 applies group relative policy optimization with format, temporal, and spatial rewards. LongEarth-R1 achieves the best results on all 12 long-sequence tasks while remaining competitive on standard remote sensing benchmarks.
Remote sensing analysis is increasingly moving beyond isolated observations toward understanding how geographic regions evolve over time. With the growing availability of high-revisit Earth observation imagery, it is now possible to observe multi-stage processes such as urban construction, disaster evolution, and ecosystem recovery. For example, given observations before, during, and after a flood, a model should identify its onset and affected regions, track its evolution, and infer the recovery stage. This requires more than pairwise change detection: the model must recover a temporal trajectory, localize supporting frames and regions, and separate geographic evolution from seasonal and acquisition-induced variations. We refer to this capability as long-horizon Earth observation reasoning: reasoning about multi-stage geographic evolution from extended Earth observation sequences.
Recent remote sensing vision-language models (RSVLMs) have progressively extended language-guided interpretation from individual observations to image pairs and temporal sequences. Single-temporal models connect an individual observation with language through image description, visual question answering, and visual grounding. Bi-temporal models further compare paired observations to describe and localize geographic changes. More recent temporal RSVLMs accept multiple observations and support sequence-level dialogue or change understanding. Nevertheless, these capabilities are still insufficient for long-horizon reasoning. A single image contains no evolution trajectory, an image pair exposes mainly the endpoints of change, and current sequence models generally emphasize recognition or descriptive question answering rather than reconstructing a complete multi-stage process. They consequently remain limited in tracing critical transitions, grounding conclusions across frames and regions, and inferring how an evolving process may continue.
To address these limitations, we formulate long-horizon Earth observation reasoning as reasoning over extended sequences. This capability requires models to identify how geographic entities evolve, localize and characterize their changes, assess whether the observed evolution is temporally consistent, and infer unobserved states. We therefore organize it into four progressive cognitive dimensions: 1) Evolution summarization organizes major stages into coherent temporal trajectories; 2) Spatial reasoning localizes changes and models their evolving spatial relations; 3) Anomaly identification detects chronological violations, repetitions, and contextual interference; and 4) Logical prediction infers subsequent or missing states from accumulated evidence. Figure 1(a) illustrates this hierarchy through 12 fine-grained tasks.
Based on the four cognitive dimensions, we construct LongEarth-Bench, a large-scale benchmark for long-horizon remote sensing spatiotemporal reasoning. It contains approximately 120k question-answering (QA) samples derived from 117k images, with an average sequence length of 15.14 frames. The benchmark covers diverse geographic regions, land-cover categories, temporal spans, change rates, and multi-stage processes. Its QA samples are generated from segmentation annotations, geometric relations, and temporal change trajectories. Moreover, 30k samples contain structured reasoning annotations connecting frame selection, temporal localization, spatial evidence, and final conclusions. Figure 1(b) compares LongEarth-Bench with existing benchmarks according to average sequence length and coverage of the four cognitive dimensions.
To enable the RSVLMs with long-horizon spatiotemporal reasoning, we first develop LongEarth through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought (CoT) traces that connect key frames, temporal changes, and changed regions to final answers. The former establishes stable frame-level temporal anchors, while the latter teaches the model to select relevant observations and integrate temporal and spatial evidence across the sequence. Building on LongEarth, LongEarth-R1 applies group relative policy optimization (GRPO) with complementary format, temporal, and spatial rewards to optimize output completeness, evolution-stage localization, temporal ordering, key-frame selection, and changed-region consistency beyond final-answer correctness. Together, LongEarth and LongEarth-R1 establish reproducible supervised and reinforcement-learning baselines on LongEarth-Bench. Figure 1(c) compares LongEarth-R1 with representative remote sensing methods across six model capabilities, with LongEarth-R1 covering all six.
The main contributions of this paper are as follows.
• We formulate long-horizon Earth observation reasoning as reasoning over multi-stage geographic evolution through four cognitive dimensions covering evolution summarization, spatial reasoning, anomaly identification, and logical prediction.
• We introduce LongEarth-Bench, a large-scale benchmark with approximately 120k samples. It covers diverse geographic processes and includes 30k samples with structured annotations for evidence-grounded reasoning.
• We develop LongEarth through supervised fine-tuning and LongEarth-R1 through GRPO-based temporal and spatial rewards, and systematically evaluate current RSVLMs across sequence lengths, evolution processes, and reasoning dimensions.
RSVLMs have progressed from static image-language alignment to temporal Earth observation understanding. Early models such as RemoteCLIP, GeoRSCLIP, GeoChat, and EarthGPT focus on aligning a single observation with language. VHM extends this capability to scene classification, visual question answering, and visual grounding, while EarthVQA advances relational visual question answering and spatial reasoning. Beyond remote sensing, SpaceVLLM studies frame-specific spatiotemporal video grounding. RSICCformer, GeoLLaVA, Change-Agent, and CCExpert extend language-based analysis to bi-temporal image pairs through change captioning, detection, or difference-aware integration. DisasterM3 provides a bi-temporal remote sensing vision-language dataset and benchmark for disaster assessment and response. However, existing bi-temporal settings cannot explicitly represent the intermediate stages of an evolving geographic process.
Recent multi-temporal RSVLMs process extended Earth observation sequences. EarthDial supports multispectral, multi-temporal, and multi-resolution conversational inputs, TEOChat enables dialogue and question answering over temporal observations, using sequences that average 2.07 frames and extend to at most 8 frames, and UniRS unifies single-image, bi-temporal, and video tasks. TimeSenCLIP learns spectral-temporal representations from Sentinel-2 time series, while DynamicVL benchmarks dynamic city understanding with multi-temporal scenes averaging 6.73 frames and extending to at most 10 frames, while VLRS-Bench evaluates cognition, decision, and prediction with sequences averaging 1.59 frames and covering up to eight temporal phases. These studies establish important foundations for multi-temporal remote sensing understanding. However, extended trajectories spanning many evolution stages remain less studied, particularly when evaluation requires anomaly localization, cross-stage spatial reasoning, and identification of the observations that support a conclusion.
Recent RSVLMs have incorporated structured reasoning and reinforcement learning to support more interpretable and evidence-grounded geospatial analysis. MultimodalCoT separates rationale generation from answer inference, while Geo-CoT constructs perceptually grounded reasoning traces and trains RSThinker through supervised fine-tuning followed by GRPO. Beyond remote sensing, IAD-R1 similarly combines CoT-based supervised fine-tuning with structured GRPO for consistent vision-language reasoning. SAMChat adopts CoT supervision and GRPO for small-scale remote sensing analysis, while RemoteReasoner applies reinforcement learning to a unified workflow covering object-, region-, and pixel-level geospatial reasoning. GeoReason aligns reasoning and answers through logical-consistency reinforcement learning, while GeoChain studies multimodal CoT geographic reasoning. In parallel, VLRS-Bench evaluates complex reasoning over single-temporal and multi-temporal remote sensing inputs. These studies show the value of intermediate supervision and task-specific rewards for reasoning beyond final-answer prediction. However, existing methods primarily reason over individual observations or restricted temporal settings, without explicitly aligning intermediate reasoning with evidence distributed across long Earth observation sequences.
LongEarth-Bench defines this hierarchy through 12 fine-grained tasks. We abbreviate the four cognitive dimensions as evolution summarization (EvoSum), spatial reasoning (Spatial), anomaly identification (AnomID), and logical prediction (LogPred). Figure 2(a) reports the sample distribution. The 120,367 samples are broadly balanced across EvoSum (26.6%), Spatial (22.8%), AnomID (26.7%), and LogPred (23.9%), while preserving diversity across the 12 tasks.
Figure 3 presents the construction pipeline of LongEarth-Bench: multi-source data integration, sequence filtering and cognitive task construction, followed by quality-controlled structured reasoning annotation. Specifically, LongEarth-Bench is derived from SpaceNet 7 (SN7), SDSU MidWest Flood (SDSU), DynamicEarthNet (DynEarth), FLAIR#2 (FLAIR), and PASTIS-R (PASTIS). These sources cover urban construction, flood dynamics, natural land-cover evolution, cloud–snow interference, and crop phenology. As shown in Figure 2(b), they span North America, South America, Europe, Africa, Asia, and Oceania, reducing dependence on a single geographic region or evolution process. Coordinate systems are normalized, while spatial resolutions, temporal metadata, and annotation formats are standardized to construct temporally ordered and spatially aligned remote sensing image sequences. Samples with severe observation corruption, incomplete temporal coverage, registration failures, or insufficient long-term variation are filtered.
Task annotations are generated from temporal trajectories together with segmentation, polygon, land-cover, and image-level flood evidence. Stage boundaries, trends, locations, directions, extents, observation quality, and phenological states are derived from cross-frame changes. Controlled order perturbations, repeated-frame insertions, and contextual disturbances are used to construct anomaly tasks, while prediction tasks are formed from historical trajectories and adjacent temporal states. Rule-based validation verifies sequence integrity, answer consistency, frame indices, spatial labels, and image paths, followed by human review for visual support and ambiguity.
To complement answer-level supervision, we construct a balanced subset of 30k structured reasoning samples across the 12 tasks. Initial traces are generated by Qwen3-VL-8B-Thinking in a unified format: the field contains visual scanning, feature identification, and integrated analysis, and the field retains the reference answer. We automatically verify tag completeness, stage ordering, non-empty reasoning fields, and answer consistency; invalid outputs are regenerated. Human experts then assess visual grounding, chain coherence, and potential answer leakage. The resulting subset provides reliable process-level supervision for learning temporal stages, spatial dynamics, anomalies, and logical prediction.
Figure 2(c) represents complementary temporal regimes across the five data sources. SDSU provides the shortest sequences, with a mean length of 6.24 frames, whereas DynEarth provides the longest, averaging 21.31 frames. SN7, FLAIR, and PASTIS cover intermediate and long contexts, with mean lengths of 14.15, 15.84, and 18.56 frames, respectively. These source-specific distributions preserve the temporal characteristics of disaster events, urban construction, natural-surface evolution, and crop phenology.
As shown in Figure 2(d), LongEarth-Bench contains 120,367 samples in total with an average sequence length of 15.14 frames, a median of 16 frames, and a maximum of 30 frames. Compared with TEOChatlas, DVL-Instruct, and VLRS-Bench, LongEarth-Bench is 7.3×, 2.2×, and 9.5× longer on average, and supports maximum sequences that are 3.75×, 3.0×, and 3.75× longer, respectively. Moreover, 21.1% of the samples contain at least 24 observations. LongEarth-Bench therefore covers a broad spectrum from compact event sequences to extended multi-stage trajectories, providing a controlled setting for evaluating sequence-grounded long-horizon reasoning.
Figure 2(e) summarizes the sequence-length distribution for each task of LongEarth-Bench and shows that all 12 tasks cover multiple sequence-length intervals, while their dominant temporal horizons differ according to their evidence requirements. Change-direction reasoning and spatial-location prediction rely on extended observations, with 53% and 57% of their samples falling within 22–25 frames, respectively. In contrast, contextual robustness and extent assessment contain more compact event-centered sequences. The substantial overlap across tasks also prevents sequence length from serving as a simple task shortcut.
This section presents a two-stage framework for long-horizon spatiotemporal reasoning. As shown in Figure 4, the framework is built on Qwen2.5-VL-7B. Stage 1 applies supervised fine-tuning (SFT) with two complementary forms of supervision. Sequence-aware answer supervision uses explicit sequence identifiers to establish stable frame-level temporal anchors, while structured CoT supervision teaches the model to select relevant observations and integrate temporal and spatial evidence across the sequence, yielding LongEarth. Stage 2 initializes from LongEarth and applies GRPO with complementary format, temporal, and spatial rewards to improve reasoning structure and spatiotemporal consistency, yielding LongEarth-R1.
The first stage contains two supervised components. The first component performs sequence-aware answer supervision, which adapts the base model to long-term remote sensing inputs and establishes frame-level temporal anchors. The second component injects structured reasoning traces, which further guides the model to organize cross-frame evidence before producing the final answer.
Sequence-Aware Answer Supervision. This component adapts the base model to long-term remote sensing data and establishes the basic mapping from image sequences and questions to answers. Given a remote sensing sequence X = I1,..., IT with T temporal observations and a question q, the model is required to generate an answer a based on the full sequence. Since long-term tasks require both single-frame recognition and cross-frame change understanding, we assign each image an explicit sequence identifier (Seq. ID), such as Image 1, Image 2,..., Image T. Let Vt = fv (It) denote the visual tokens of the t-th frame encoded by the visual encoder fv. The ordered multimodal input is represented as: H = [q; Image 1, V1;...; Image T, VT]. This representation provides stable frame-level anchors and reduces ambiguity in key-frame reference, temporal interval localization, and cross-frame relation modeling. During training, the visual encoder is frozen, while low-rank adaptation (LoRA) modules are applied to the language model layers for parameter-efficient alignment. For answer-only samples, this objective mainly constrains the final answer and does not explicitly supervise how the model selects cross-frame evidence or organizes intermediate reasoning.
Structured CoT Supervision. Within the supervised stage, we use the structured reasoning subset of LongEarth-Bench to supervise evidence-grounded responses. For each reasoning sample (X, q, r, a), the target response contains a reasoning trace r and final answer a, formatted as r a. The reasoning trace follows three steps: visual scanning, feature identification, and integrated analysis. Visual scanning describes the main land-cover states and spatial layouts across temporal observations. Feature identification extracts key frames, changed regions, directions, or anomalies related to the question. Integrated analysis combines cross-frame evidence and produces the final conclusion. This design shifts the model from direct answer learning to process-level supervision. Specifically, given the target sequence y = (y1,..., yN), the structured-reasoning objective is: LCoT = − sum n=1 N log pθ (yn y<n, H). This objective jointly supervises the reasoning trace and final answer, but reference-trace imitation alone cannot directly penalize temporal ordering errors or spatial inconsistencies. The second stage therefore applies GRPO to further optimize structural validity and spatiotemporal grounding.
Starting from LongEarth, the second stage applies GRPO to obtain LongEarth-R1 and align reasoning with long-term remote sensing objectives. Unlike SFT, which fits a single reference output, GRPO compares a group of candidate responses for the same input and updates the policy using relative rewards. It therefore improves answer quality together with reasoning structure and spatiotemporal consistency. For multimodal input H, the old policy πθold samples G responses yi = (ri, ai), each comprising a reasoning trace and final answer. We assign each response a weighted reward: Ri = λf Rifmt + λt Ritime + λs Rispace, where the three terms measure format validity, temporal grounding, and spatial consistency, respectively. The format reward checks the required reasoning–answer structure and task-specific answer parsability. The temporal reward combines matching of referenced frames with their chronological consistency, while the spatial reward evaluates agreement on changed locations, directions, and extents. For tasks without temporal or spatial labels, inapplicable terms are omitted and the remaining weights are renormalized; detailed reward definitions are provided in the supplementary material.
GRPO normalizes rewards within the response group as Ai = (Ri − µR)/(σR + ϵ) and uses the policy ratio ρi = πθ (yi H)/πθold (yi H). Let ρ̄i = clip(ρi, 1 − ε, 1 + ε). The clipped objective is JGRPO = (1/G) sum i=1 G [min(ρi Ai, ρ̄i Ai) − βDKL (πθ ∥πref)], where ε is the clipping coefficient, β is the Kullback–Leibler (KL) weight, and πref is the policy obtained after the supervised stage. The KL term preserves the visual-language ability acquired during SFT. Together, the rewards shift learning from answer imitation toward structured, temporally ordered, and spatially grounded reasoning.
We train LongEarth-R1 on 8× NVIDIA A800 GPUs, freezing the visual encoder and tuning language-side LoRA modules with r = 128 and α = 256 in both stages. SFT runs for two epochs with bfloat16, gradient checkpointing, cosine decay, a peak learning rate of 2×10−5, a 0.03 warmup ratio, and no weight decay. Inputs are limited to 8,192 tokens and interleave square-resized images with temporal prefixes and instructions. GRPO starts from the SFT checkpoint and optimizes the same LoRA modules with group-sampled responses and the proposed rewards.
We evaluate final answers using Accuracy, Temporal F1, and Spatial F1. Accuracy applies to closed-set answers, including categorical judgments, single-frame localization, fixed spatial labels, and discrete extent levels. Temporal and free-form spatial answers use set-based F1 over predicted and reference elements. Temporal F1 operates on parsed frame indices, whereas Spatial F1 uses S = object, change, region, direction, extent. Temporal F1 evaluates T1, T3, T8, and temporal-set variants of T7, T9, and T12; Spatial F1 evaluates free-form T4–T6 and T11. Remaining tasks use Accuracy.
We evaluate the 12 LongEarth-Bench tasks spanning EvoSum, Spatial, AnomID, and LogPred. Table 1 shows that LongEarth-R1 ranks first on all tasks, with particularly strong gains on AnomID and tasks requiring long-range temporal evidence. The consistent improvements show that sequence grounding, structured reasoning, and reward-based alignment jointly strengthen temporal understanding, spatial grounding, and prediction. Figure 5 further shows that LongEarth-R1 avoids temporally misaligned intervals by grounding its answer in multi-frame evidence.
We further test whether specialization for long-term reasoning preserves general remote sensing understanding. Following the standard evaluation protocol, we evaluate single-image scene recognition on AID and UCM, bi-temporal change understanding on ABCD-CD, CDVQA-QA, xBD, and S2Looking, and short-sequence multi-image reasoning on Qfabric and fMoW. For each benchmark, LongEarth-R1 is fine-tuned on the corresponding training split and evaluated using the task-native protocol; detailed sources and task descriptions are provided in the supplementary material. Table 2 shows that LongEarth-R1 achieves the best result on six of nine datasets, including all bi-temporal change-understanding benchmarks, indicating effective transfer from long-horizon supervision to conventional change analysis. Its performance remains close to the strongest specialist on the remaining benchmarks, showing that the proposed training preserves general remote sensing understanding across single-image, bi-temporal, and short-sequence settings.
We conduct cumulative component and reward ablations across the four cognitive dimensions of LongEarth-Bench. The component study progressively adds SFT, Seq. IDs, R1, and GRPO, whereas the reward study removes one term at a time from the full objective. Table 3 shows that the complete configuration, LongEarth-R1, achieves the strongest macro-average and the most balanced performance across all four dimensions. Relative to LongEarth, adding GRPO yields consistent gains, with the largest improvement on AnomID (+12.61), indicating that reward-based optimization strengthens anomaly-sensitive evidence selection and spatiotemporal reasoning. Removing any reward lowers the macro-average by 5.19–6.68%; although removing the format reward slightly improves LogPred, the full objective remains best overall, confirming the complementarity of the three rewards.
Figure 6 separates two sources of temporal difficulty: the number of input frames processed by the model and the temporal span of evidence required by the question. In the top panel, all methods degrade as input sequences become longer, reflecting the increasing need to suppress irrelevant observations and localize the informative frames. LongEarth-R1 nevertheless leads every input-length interval by 10.7–31.9%, retaining 51.4 at 26–30 frames after reaching 87.8 on 2–5-frame inputs. This trend indicates that explicit sequence grounding and reward-based alignment improve robustness to long-context distraction, although very long inputs remain challenging. The bottom panel shows a different pattern. LongEarth-R1 achieves the best score in six of seven ground-truth evidence-span intervals, and its performance generally improves as the answer can be supported by evidence distributed across a broader temporal span. Thus, a broad evidence span is not necessarily harmful: when multiple observations provide complementary evolution cues, cross-frame reasoning can benefit from them. The only exception is the 23–30-frame interval, where TEOChat* attains a higher score.
Long-term remote sensing understanding requires reasoning over evolving geographic evidence rather than isolated observations. We present LongEarth-Bench, which defines this setting with four cognitive dimensions and structured reasoning supervision. We further develop LongEarth through supervised sequence grounding and LongEarth-R1 through reward-driven spatiotemporal alignment. Results improve performance across the 12 long-sequence tasks while retaining transfer to conventional remote sensing tasks. These findings support explicit modeling of temporal order, intermediate evidence, and spatial consistency for long-horizon Earth observation reasoning.
Improvements for AI systems
Improvements to AI systems:
-
Long-horizon temporal reasoning with explicit frame anchoring: The system can process sequences of 15–30 Earth observation images, assign explicit sequence identifiers (Image 1, Image 2, …) to each frame, and maintain stable temporal anchors throughout reasoning. This enables tracking multi-stage geographic evolution (e.g., flood onset → spread → recovery) rather than only comparing two endpoints.
-
Structured chain-of-thought supervision for evidence grounding: The system is trained on 30k samples with structured reasoning traces that decompose answers into visual scanning, feature identification, and integrated analysis. It learns to select key frames, localize changed regions, and cite temporal intervals before producing a final answer, reducing hallucinated or temporally misaligned conclusions.
-
Reinforcement learning with temporal and spatial rewards: The system optimizes not just final-answer correctness but also temporal ordering (chronological consistency of referenced frames) and spatial consistency (agreement on changed locations, directions, extents). This improves anomaly detection (e.g., detecting repeated frames or chronological violations) by +12.61 points over supervised-only training.
-
Cross-task transfer to conventional remote sensing benchmarks: The system retains and improves performance on single-image scene recognition (AID, UCM), bi-temporal change understanding (ABCD-CD, CDVQA-QA, xBD, S2Looking), and short-sequence reasoning (Qfabric, fMoW), achieving state-of-the-art on 6 of 9 datasets. This shows long-horizon training generalizes to standard tasks without specialization loss.
-
Robustness to long-context distraction: The system maintains 51.4% accuracy on 26–30-frame inputs (vs. 87.8% on 2–5 frames), outperforming baselines by 10.7–31.9% across all input-length intervals. It learns to suppress irrelevant observations and focus on informative frames, which is critical for real-world satellite archives with noisy or redundant imagery.
-
Evidence-span exploitation: The system improves as the ground-truth evidence spans more frames (up to 22–30 frames), leveraging complementary cues across observations. This enables more confident predictions when multiple frames provide corroborating evolution signals, rather than degrading with longer inputs.
-
Four cognitive reasoning dimensions: The system can perform evolution summarization (organizing stages into trajectories), spatial reasoning (localizing changes and modeling spatial relations), anomaly identification (detecting temporal inconsistencies), and logical prediction (inferring future or missing states) across 12 fine-grained tasks, covering the full spectrum of long-horizon Earth observation reasoning.
What the improved AI system can do:
-
Given a 30-frame satellite sequence of a flood event, it can identify the onset frame, track the affected region’s expansion, detect anomalous frames (e.g., cloud interference or duplicated observations), and predict the recovery stage—while citing the specific frames and regions that support each conclusion.
-
For urban construction monitoring, it can summarize the multi-stage development (clearing → foundation → structure → completion), localize each change’s spatial extent, and infer the next likely stage based on historical trajectories.
-
It can transfer this temporal reasoning ability to standard tasks like change detection in bi-temporal image pairs, achieving better accuracy than specialized models, without requiring retraining from scratch.
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing
- FLAIR #2: textural and temporal information for semantic segmentation from multi-source optical imagery
- GeoReason: Aligning Thinking And Answering In Remote Sensing Vision-Language Models Via Logical Consistency Reinforcement Learning
- UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models
- VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- CCExpert: Advancing MLLM Capability in Remote Sensing Change Captioning with Difference-Aware Integration and a Foundational Dataset
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection