AI papers — 2026-09-29
Today's focus centers on AmbiModBench, which is an effort to benchmark gene perturbation prediction methods that go beyond simply looking at shared responses. The work involved testing these models against a set of perturbations and evaluating their predictive accuracy in a way that moves past shared response patterns. This contrasts with the L 1-2 GLasso study, which used L 1-2 Regularized Multi-task Graphical Lasso to jointly estimate eQTL mapping and gene networks, suggesting an effort to link genetic variation directly to network structure.
Furthermore, the DCFold approach aimed for efficient protein structure generation using a single forward pass. KoopCell explored learning single-cell dynamics from distribution snapshots via a Koopman-Based Generative Model. These studies highlight a push toward more nuanced biological modeling, moving from simple correlation to complex relational mapping and dynamic system learning. What remains open is how these diverse methods can be integrated to create truly predictive frameworks for gene perturbation in a way that surpasses the limitations of current shared response analyses.
The work on structured state-space models explored the emergence of a primacy effect within these systems, suggesting that initial conditions exert a disproportionate influence on subsequent dynamics. This was investigated alongside conditioning direct feedback alignment using activity and error geometry, which aimed to refine how these models respond to input by focusing on the geometric relationship between activity and error signals.
Furthermore, recovery-directed symbolic distillation of neural likelihoods was attempted to improve model performance by distilling complex neural likelihoods into a more structured symbolic representation. This process seeks to capture essential information while simplifying the underlying complexity. These efforts are situated within a broader context where distribution-aware channel capacity for effective connectivity moves beyond simple Gaussian assumptions, suggesting that modeling the underlying data distribution is crucial for understanding how information flows through these systems.
Simultaneously, improvements were made to causal effect estimation of weighted regression-based estimators by employing neural networks. This indicates an effort to enhance the accuracy of inferring causal relationships from observational data. These methodological explorations are paralleled by work on categorical approaches to conflict resolution, establishing a corrected correspondence between category theory and graph models for resolving conflicts.
The work on goal-conditioned supervised learning for multi-objective recommendation explored how to train models to optimize several competing objectives simultaneously. This process involved setting up a framework where the learning signal was guided by specific goals. This approach aimed to move beyond single-objective optimization by incorporating multiple criteria into the training process.
Related efforts in sequential changepoint localization focused on post-detection inference, suggesting methods for identifying shifts in data streams after an initial detection event occurs. This implies a need for robust real-time adaptation. Furthermore, research into focus on likely classes for test-time prediction suggests strategies to improve model performance during inference by prioritizing the most probable categories based on the input data.
These ideas connect to building intelligent agents using neuro-symbolic concepts, which seeks to integrate the strengths of neural networks with explicit symbolic reasoning capabilities. In terms of knowledge augmentation for large language models, work was done on using knowledge-augmented LLMs specifically for the ARC benchmark, aiming to enhance their generalization by injecting external knowledge.
Meanwhile, searching for actual causes involved developing approximate algorithms that allow for adjustable precision when trying to pinpoint causal factors. Finally, efforts to push toward the simplex vertices addressed a specific issue in smoothed vector quantization related to code collapse by employing a simple remedy.
The attentionViG approach explored a method for dynamic neighbor aggregation within vision graph neural networks by incorporating cross-attention mechanisms to better weigh the influence of neighboring nodes during the message passing phase. This was contrasted with efforts in large language model mitigation, where a capability-oriented survey examined how retrieval augmented generation and agentic systems address hallucination, suggesting that improving reasoning capabilities is key.
Simultaneously, work on time series anomaly detection involved COGNOS, which utilized constrained Gaussian-noise optimization and smoothing to enhance the universal enhancement for this task. Research into adaptive nonparametric dimensionality reduction provided a general framework for reducing the complexity of data representations without imposing strict structural assumptions. These diverse efforts suggest an ongoing tension between improving local relational modeling in vision tasks and tackling systemic issues like model hallucination or optimizing complex time series analysis through constrained optimization techniques.
The investigation into soft geometric inductive biases for object centric dynamics explored a method where a specific type of bias was introduced to guide learning, aiming to improve the representation of objects in dynamic systems. This approach involved modifying the loss function or network structure to favor certain geometric relationships during training. This suggests that imposing structural priors can help models capture physical realities more efficiently.
In parallel, research on SB-TRPO focused on developing safe reinforcement learning by incorporating hard constraints directly into the policy optimization process. This suggests a path toward more reliable decision-making in complex environments. Deep Delta Learning presented an alternative learning paradigm that sought to improve sample efficiency through a delta-based update mechanism. This implies that focusing only on the changes made during an iteration can be more effective than updating the entire network.
Meanwhile, NC-Bench was established as a benchmark specifically designed to evaluate the conversational competence of large language models, providing a standardized way to measure how well these models handle complex dialogue. Furthermore, what if TSF recontextualized time series forecasting by framing it as scenario-guided multimodal forecasting, suggesting that incorporating different types of input modalities and scenarios can lead to more robust predictions.
The work on untangling input language from reasoning language provided a diagnostic framework for assessing cross-lingual moral alignment in LLMs. This indicates an effort to separate linguistic features from underlying ethical reasoning. Contextual Distributionally Robust Optimization introduced a method that integrates causal and continuous structures into optimization problems, aiming to create models that are less sensitive to uncertainty in the data distribution.
Finally, STEP-LLM addressed the generation of CAD models from natural language using large language models. This demonstrated a capability to translate high-level textual descriptions into precise geometric representations.
The work on OP-Bench focused on benchmarking over-personalization within memory-augmented personalized conversational agents, specifically looking at how different personalization strategies perform. This involved setting up a framework to test these agents against various personalization levels to see which approach yielded the best conversational outcomes.
Complementing this, research into Just-In-Time Reinforcement Learning explored continual learning within LLM agents without requiring gradient updates. This suggests a method for adapting agent behavior incrementally as new interactions occur. The GLOVE study introduced a Global Verifier designed for LLM memory-environment realignment, aiming to ensure the agent's internal memory accurately reflects the current environment state. These efforts are situated alongside investigations into latent-coT models to determine if they exhibit true step-by-step sequential reasoning.
These findings from these diverse experiments point toward the challenges in maintaining consistent personalization while enabling continuous, adaptive learning within complex conversational systems.
The work on GUI-GenBench focused on evaluating image generation models as interactive graphical user interface environments, specifically looking at how these models can function in a generative context. The research explored the capabilities of these models when presented with visual inputs and subsequent user interactions within a simulated GUI framework. This effort suggests a path toward assessing generative AI not just for static image quality but for dynamic, interactive usability.
In contrast, TSR investigated trajectory-search rollouts for multi-turn reinforcement learning using large language model agents. This involved setting up scenarios where LLM agents could perform sequential actions and evaluate their performance through these search procedures to refine their policy over multiple turns.
Furthermore, efforts were made to address issues in federated low-rank adaptation by developing methods to prevent rank collapse when client heterogeneity is present. CodeScaler addressed the scaling of code large language model training and test-time inference by employing reward models to guide this process, aiming for more efficient deployment of these models.
CausalReasoningBenchmark introduced a real-world benchmark designed for disentangled evaluation of causal identification and estimation. This provided a structured way to test how well models can isolate causal relationships in complex data. Finally, research into Words & Weights examined streamlining multi-turn interactions through co-adaptation techniques, suggesting ways to improve the coherence of conversational or agentic sequences.
The work on physics-informed neural networks with architectural physics embedding focused on large-scale wave field reconstruction. This explored how incorporating physical constraints into the network's structure impacts its ability to model complex wave phenomena. The findings indicated that this architectural embedding method provided a pathway for achieving better reconstruction fidelity. This suggests that explicitly encoding physical principles into the network's design can mitigate some of the challenges inherent in purely data-driven approaches when dealing with high-dimensional wave data. However, the abstracts do not detail the specific quantitative improvements or limitations encountered during this reconstruction task, leaving open questions regarding its scalability across even larger datasets and its general applicability beyond wave field problems.
Today's papers
- AmbiModBench: Benchmarking Gene Perturbation Prediction Beyond Shared Responses. [paper]
- l 1-2 GLasso: L 1-2 Regularized Multi-task Graphical Lasso for Joint Estimation of eQTL Mapping and Gene Network. [paper]
- Forecasting Bacterial Antimicrobial Resistance Trends Using Machine Learning on WHO GLASS Surveillance Data: A Retrieval-Augmented Generation Approach for Policy Decision Support. [paper]
- DCFold: Efficient Protein Structure Generation with Single Forward Pass. [paper]
- KoopCell: Koopman-Based Generative Model for Learning Single-Cell Dynamics from Distribution Snapshots. [paper]
- Robust Biomolecular Complex Design Across Protein Conformational Landscapes. [paper]
- MolLangData: A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method. [paper]
- Hessian Matching for Machine-Learned Coarse-Grained Molecular Dynamics. [paper]
- Emergence of the Primacy Effect in Structured State-Space Models. [paper]
- Conditioned Direct Feedback Alignment via Activity and Error Geometry. [paper]
- Recovery-Directed Symbolic Distillation of Neural Likelihoods. [paper]
- Beyond Gaussian Assumptions: Distribution-Aware Channel Capacity for Effective Connectivity. [paper]
- Improving Causal Effect Estimation of Weighted RegressionBased Estimator using Neural Networks. [paper]
- Categorical Approach to Conflict Resolution:A Corrected Correspondence between Category Theory and the Graph Model for Conflict Resolution. [paper]
- Help Me Help You: The Aggregate Value of Source and Target Data in Transfer Learning. [paper]
- Evaluating Cell AI Foundation Models in Kidney Pathology with Human-in-the-Loop Enrichment. [paper]
- Goal-Conditioned Supervised Learning for Multi-Objective Recommendation. [paper]
- Post-detection inference for sequential changepoint localization. [paper]
- Focus on Likely Classes for Test-Time Prediction. [paper]
- Building Intelligent Agents with Neuro-Symbolic Concepts. [paper]
- From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark. [paper]
- Searching for Actual Causes: Approximate Algorithms with Adjustable Precision. [paper]
- Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization. [paper]
- Patch Rebirth: Toward Fast and Transferable Model Inversion of Vision Transformers.
- AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs. [paper]
- Mitigating Hallucination in Large Language Models: A Capability-Oriented Survey on RAG, Reasoning, and Agentic Systems. [paper]
- COGNOS: Universal Enhancement for Time Series Anomaly Detection via Constrained Gaussian-Noise Optimization and Smoothing. [paper]
- A general framework for adaptive nonparametric dimensionality reduction. [paper]
- AI Annotation Orchestration: Evaluating LLM verifiers to Improve the Quality of LLM Annotations in Learning Analytics. [paper]
- AI-driven ionic liquid discovery with unified chemical intelligence. [paper]
- The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs. [paper]
- On Memory: A comparison of memory mechanisms in world models. [paper]
- Soft Geometric Inductive Bias for Object Centric Dynamics. [paper]
- SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints. [paper]
- Deep Delta Learning. [paper]
- NC-Bench: An LLM Benchmark for Evaluating Conversational Competence. [paper]
- What If TSF: Reframing Time Series Forecasting as Scenario-Guided Multimodal Forecasting. [paper]
- Untangling Input Language from Reasoning Language: A Diagnostic Framework for Cross-Lingual Moral Alignment in LLMs. [paper]
- Contextual Distributionally Robust Optimization with Causal and Continuous Structure. [paper]
- STEP-LLM: Generating CAD STEP Models from Natural Language with Large Language Models. [paper]
- OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents. [paper]
- Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates. [paper]
- GLOVE: Global Verifier for LLM Memory-Environment Realignment. [paper]
- Denoising Time Matters:Diverse Generation in Diffusion Language Models. [paper]
- Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks. [paper]
- VLM-Guided Experience Replay. [paper]
- Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation. [paper]
- Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems. [paper]
- GUI-GenBench: Evaluating Image Generation Models as Interactive GUI Environments. [paper]
- TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents. [paper]
- Preventing Rank Collapse in Federated Low-Rank Adaptation with Client Heterogeneity. [paper]
- CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models. [paper]
- CausalReasoningBenchmark: A Real-World Benchmark for Disentangled Evaluation of Causal Identification and Estimation. [paper]
- Don't stop me now: How Validation Criteria Affect Checkpoint Selection and Early Stopping. [paper]
- Words & Weights: Streamlining Multi-Turn Interactions via Co-Adaptation. [paper]
- LFPO: Likelihood-Free Policy Optimization for Masked Diffusion Models. [paper]
- Physics-Informed Neural Networks with Architectural Physics Embedding for Large-Scale Wave Field Reconstruction. [paper]
- Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs. [paper]
- A theoretical model of dynamical grammatical gender shifting based on set-valued set function. [paper]
- The Trace Is the State: Exact Credit Assignment for LLM Agent Teams. [paper]
The papers
- A Survey on Efficient Vision-Language-Action Models — As a fastidious and diligent AI researcher, I have meticulously analyzed both provided texts (A and B) regarding "Efficient Vision-Language-Action models (Efficient VLAs)." My task is to synthesize these two descriptions into a single, comprehensive, long, and detailed summary th [episode]
- LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform — Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. [episode]
- From Directions to Regions: Decomposing Activations in Language Models via Local Geometry — Activation decomposition methods in language models are tightly coupled to geometric assumptions on how concepts are realized in activation space, and this work introduces Mixture of Factor Analyzers (MFA) as a scalable, unsupervised alternative that models activation space as a [episode]
- Confidently Deceptive: On the Relationship Between Confidence and Deception in LLMs — Large language models (LLMs) can produce deceptive responses, and this study investigates how confidence in those deceptive outputs amplifies their risk to users. [episode]
- Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding — Reward models for text-to-video (T2V) generation often fail at fine-grained semantic alignment because they lack systematic verification and rely on implicit reasoning, leading to errors in handling complex prompts. [episode]
- Stable Velocity: A Variance Perspective on Flow Matching — Stable Velocity introduces a variance-based perspective on flow matching, revealing a two-regime structure that governs both training and inference dynamics. [episode]
- How LLMs Are Persuaded: A Few Attention Heads, Rerouted — A small set of mid-layer attention heads almost entirely determines a language model's answer, and persuasion works by redirecting attention through a rank-one evidence-routing feature. [episode]
- Denoising-Enhanced Coarse-to-Fine Infrared Small Target Detection with Attention Prior-Guided Knowledge Distillation — Infrared small target detection (IRSTD) in high-resolution images remains challenging due to targets' small size, weak features, and severe interference from complex dynamic backgrounds. [episode]
- OrientedFormer: An End-to-End Transformer-Based Oriented Object Detector in Remote Sensing Images — Oriented object detection in remote sensing images remains challenging due to objects being distributed in multiorientation, and this paper proposes OrientedFormer, an end-to-end transformer-based oriented object detector that addresses these issues through three dedicated module [episode]
- Understanding and inverse design of implicit bias in stochastic learning: a geometric perspective — A key challenge in machine learning is explaining how learning dynamics select among many solutions that achieve identical loss values in overparameterized models—a phenomenon known as implicit bias. [episode]
- FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning — Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: multi-task pretraining, where related tasks borrow representational strength, and continual adaptation, where new tasks emerge with previously unseen modalit [episode]
- UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City — Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. [episode]
- Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning — In LLM Reinforcement Fine-Tuning (RFT), this paper introduces METIS, a novel framework that internalizes curriculum judgment as a native capability to drive efficient and high-performing training. [episode]
- Deep Minds and Shallow Probes — As a meticulous researcher, I have carefully analyzed both provided summaries of the paper "Deep Minds and Shallow Probes." The synthesis below integrates these details into a comprehensive, high-fidelity overview, ensuring no critical nuance is lost. [episode]
- OmniVR: Audio-Video Conditional Generation for Archival Footage Restoration — Historical films suffer from co-occurring visual and audio degradations—blur, noise, flicker, hiss, clipping, and dropout—yet existing methods restore each modality independently, leaving quality gaps and cross-modal inconsistency. [episode]
- GeoMetric: Injecting Metric Geographic Structure into Worldwide Image Geo-Localization — Worldwide image geo-localization aims to determine where on Earth a single image was captured, and this work introduces GeoMetric, a retrieval-based framework that encodes GPS coordinates relationally rather than in isolation by injecting distance structure into representation le [episode]
- Knowledge-Intensive Video Generation — Text-to-video generation has rapidly improved in visual quality, but it remains under-evaluated for factuality and practical usefulness in information-seeking scenarios. [episode]
- Making AI Scientists Auditable from Evidence to Claim — As a meticulous researcher, I have thoroughly reviewed both provided texts concerning XCIENTIST and related concepts in making AI scientists auditable from evidence to claim. [episode]
- IndexRAG: Index-Time Reasoning for Multi-Hop Retrieval-Augmented Generation — IndexRAG presents a novel approach that shifts cross-document reasoning from online inference to offline indexing, allowing for single-pass retrieval and a single LLM call at inference time. [episode]
- OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning — OmniWeaving proposes an omni-level video generation model designed to achieve unified, free-form video creation by integrating powerful multimodal composition and reasoning capabilities. [episode]
- AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists — Building AI co-scientists that assist in open-ended scientific discovery remains challenging due to a scarcity of high-quality, large-scale data for training and evaluation. [episode]
- Diffusion-grounded VideoLLM for Entity-aware temoporal grounding — Understanding videos requires more than answering open-ended questions—it demands the ability to pinpoint when events occur and how entities interact across time. [episode]
- ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages — Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially in multilingual and low-resource scenarios. [episode]
- Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF — Large language models frequently exhibit performance biases against regional dialects of low-resource languages, and this study proposes a two-phase framework to evaluate dialectal bias in LLM question-answering across nine Bengali dialects. [episode]
- Inside the LLM Word Factory — Detokenization is characterized as an early-layer two-stage mechanism where attention writes a token-specific directional signal from preceding subwords, and the MLP composes it with the local embedding. [episode]
- Robust 3D Reconstruction from Multi-View Optical Satellite Imagery via Reliability-Aware Height-Evidence Fusion in Gaussian Splatting — A Digital Surface Model (DSM) reconstruction method using 3D Gaussian Splatting (3DGS) that addresses height-layer mixing errors by introducing a risk-map guided consistency framework. [episode]
- MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image — As a fastidious researcher, I must meticulously synthesize all provided information from both sources to construct a comprehensive and accurate summary of the paper "MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image." The initial analysis reveals two distin [episode]
- Neural Proposals, Symbolic Guarantees: Neuro-Symbolic Graph Generative Modeling — NeuroSymbolic Graph Generative Modeling (NSGGM) introduces a neurosymbolic framework that reapproaches molecule generation as a scaffold and interaction learning task with symbolic assembly, offering explicit controllability and formal guarantees that pure deep neural methods lac [episode]
- Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification — Multimodal contrastive learning has enabled zero-shot visual classification, but existing methods often produce inconsistent predictions across hierarchical label spaces, which this work addresses by proposing a level-restricted contrastive learning framework to improve both hier [episode]
- delta-mem: Efficient Online Memory for Large Language Models — δ-mem proposes a lightweight memory mechanism that augments frozen full-attention backbones with a compact online state of associative memory to dynamically maintain and steer historical information during generation. [episode]
- PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning — This research introduces PACE (Policy-Native Adaptive Decision Timing), a novel approach to long-horizon reasoning that treats temporal abstraction—specifically, the decision of how deeply to commit to a sequence of actions before re-evaluating the state—as a dynamic control [episode]
- Code2Math: Can Your Code Agent Evolve Math Problems Through Exploration? — As large language models advance their mathematical capabilities, there is a significant bottleneck in training and self-evolution due to the scarcity of challenging, high-quality problems. [episode]
- Compact Feed-Forward 3D Gaussians via Saliency-Guided Primitive Merging — 3D scene reconstruction, modeling, and rendering are highly relevant for numerous tasks, and 3D Gaussian splatting has become a standard choice in this context. [episode]
- MOPDA: Mixed-Trajectory On-Policy Distillation for Language-Guided Industrial Anomaly Detection — Mixed-Trajectory On-Policy Distillation for Language-Guided Industrial Anomaly Detection (MOPDA) proposes a novel framework to bridge the gap between language understanding and precise pixel-level anomaly localization in vision-language models. [episode]
- Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility — Tool efficiency and marginal tool utility are introduced as new quantitative metrics to evaluate how useful tools are in an LLM agent trajectory, aiming to provide a direct measure of tool usefulness distinct from accuracy proxies. [episode]
- MASRubric: Auditing Information Flow in Multi-Agent Systems with Failure-Distilled Pitfall Rubrics — AgentDropoutV2 (ADv2) introduces a test-time rectify-or-reject pruning framework that dynamically optimizes information flow in Multi-Agent Systems (MAS) by intercepting agent outputs, iteratively correcting errors using an indicator pool derived from historical failures, and pru [episode]
- Watch the Model Think: On-Policy Extraction of Activation Steering Vectors — Activation steering provides parameterefficient control over large language models (LLMs) at inference time, but many methods rely on off-distribution supervision and discrete masking, leading to brittle interventions. [episode]
- Quantifying and Mitigating Domain Shift in Peach Leaf Damage Classification: Attention Mechanisms and Fine-Tuning Strategies — Deep learning provides a practical framework for crop damage assessment from imagery, supporting early decision-making in agricultural management. [episode]
- SymDrift: One-Shot Generative Modeling under Symmetries — Generative modeling of physical systems, such as molecules, requires learning distributions that are invariant under global symmetries, such as rotations in three-dimensional space. [episode]
- Query-Dependent Use of Generated Descriptions for Reliable Visual Question Answering — A new method is introduced to regulate how Vision-Language Models (VLMs) use self-generated captions, demonstrating that caption utility is a per-query property rather than a fixed asset. [episode]
- Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference — Flux Attention introduces a context-aware framework that dynamically optimizes attention computation at the layer level to overcome the quadratic complexity bottleneck and hardware inefficiencies associated with static or head-level sparsity in Large Language Models during long-c [episode]
- Robust Active Learning for Few-Shot Example Selection in Text-to-SQL — Few-shot example retrieval is a dominant paradigm for grounding large language models (LLMs) in domain-specific text-to-SQL systems, but expert annotation remains prohibitively expensive. [episode]
- BFMT: Enhancing Search Capabilities of Tree Sampler via Bootstrap Flow-Map Tree — Bootstrap FlowMap Tree (BFMT) is a novel sampling framework designed for history-aware global search and alignment under strict sampling budget constraints, enabling efficient online feedback-driven search in domains where discovery is bottlenecked by costly evaluations. [episode]
- GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning — Human visual reasoning relies on active vision, a metacognitive process where top-down control directs focus to task-relevant details while maintaining peripheral awareness. [episode]
- SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks — Teaching language models to use search tools is not only a question of whether they search, but also of whether they issue good queries. [episode]
- A Scalable Entity-Based Framework for Auditing Bias in Large Language Models — A scalable entity-based framework for auditing bias in large language models introduces a systematic method using named entities as controlled probes to measure systematic disparities in model behavior, which matters because it provides a rigorous, large-scale tool for identifyin [episode]
- Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models — Recurrent LLM architectures have emerged as a promising approach for improving reasoning, as they enable multi-step computation in the embedding space without generating intermediate tokens. [episode]
- Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation — Personalized Text-to-Image (PT2I) generation aims to produce customized images based on reference images, and this work introduces DynaIP, a cutting-edge plugin designed to enhance fine-grained concept fidelity, balance between Concept Preservation (CP) and Prompt Following (PF), [episode]
- Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content — Zeroing a single attention head recovers up to 70.4 pp of adversarial accuracy for RoBERTa on Jigsaw and 37.3 pp for Llama Guard 2 on ToxiGen, at ≤0.6 pp clean cost, revealing mechanistically traceable fairness gaps in current toxicity classifiers. [episode]
- Three-Phase Scribble-Adaptive Curriculum Learning for autoPETV Grand Challenge — This report describes Libo Zhang’s algorithmic solution to autoPETV Grand Challenge on interactive lesion segmentation in wholebody PET/CT, which addresses the progression toward human-in-the-loop interaction by training a residual-encoder U-Net with a three-phase curriculum th [episode]
- Rethinking the State Update Gate for Long-Sequence Recurrent 3D Reconstruction — Streaming 3D reconstruction under strict constant-memory constraints suffers from long-sequence drift due to recurrent state updates that are structurally bounded and content-independent. [episode]
- PHOEBI: An Open-World Benchmark for Multi-Label Bacterial Identification in Phase-Contrast Microscopy — Optical microscopy enables rapid, label-free imaging of live bacteria and is the standard instrument for species identification across clinical, environmental, and industrial microbiology. [episode]
- Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs — Omni-modal Large Language Models (Omni-MLLMs) currently suffer from performance fragility because their static fusion topologies introduce systemic flaws, leading unimodal baselines to frequently outperform joint multimodal inference. [episode]
- ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning — Reaction feasibility prediction, a fundamental problem in computational chemistry, has benefited from diverse tools enabled by recent advances in artificial intelligence; however, individual tool performance varies substantially across reactions, making it difficult for any singl [episode]
- Adaptive Weighted h-Transform Sampling for Coarse-Guided Visual Generation — Coarse-guided visual generation addresses the need to synthesize high-fidelity fine samples from degraded or low-fidelity coarse references, which is crucial for applications like deblurring and super-resolution. [episode]
- PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization with Architecture-specific PTX — PTXBench introduces an auditable benchmark and adaptation environment for architecture-specific PTX programming, exposing a persistent capability gap where current LLMs struggle to consistently achieve competitive performance across evolving GPU architectures like H100 and B200. [episode]
- NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems — This paper introduces NOVA (NOise-Aware Verbal Confidence CAlibration Rules), a novel framework designed to address the critical issue of poor confidence calibration in Large Language Models (LLMs) when deployed within Retrieval-Augmented Generation (RAG) systems, particularly in [episode]
- Adaptively Truncated Signature-based Logistic Regression for Semi-parametric Functional Classification — This research introduces Path Signatures Logistic Regression (PSLR), a novel semi-parametric framework designed for the classification of vector-valued functional data when accompanied by scalar covariates. [episode]
- EviSearch: Trustworthy Extraction and Synthesis of Clinical Trial Evidence with Agents that Improve with Use — EviSearch introduces a multi-agent extraction system designed to automate ontology-aligned clinical evidence table creation directly from native trial PDFs while guaranteeing per-cell provenance for audit and human verification. [episode]
- Fast LeWorldModel — Fast LeWorldModel proposes Fast-LeWM, a fast latent world model that replaces repeated local rollout with action-prefix prediction to enable parallel multi-horizon state evolution. [episode]
- Factual recall in linear associative memories: sharp asymptotics and mechanistic insights — Large language models demonstrate remarkable ability in factual recall, yet fundamental limits remain unclear. [episode]
- Goal-Conditioned Supervised Learning for LLM Fine-Tuning — Goal-conditioned supervised learning (GCSL) is presented as an offline fine-tuning framework for Large Language Models that treats feedback signals directly as explicit goals, offering high training efficiency and scalability without relying on costly online reinforcement learnin [episode]
- Knowledge Graphs, the Missing Link in Agentic AI-based Formal Verification — Recent advances in Large Language Models (LLMs) have enabled workflows that generate SystemVerilog Assertions (SVAs) from natural-language specifications, with the potential to accelerate Formal Verification (FV). [episode]
- One Turn Too Late: Learning When to Intervene Against Multi-Turn Malicious Intent — Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs) because attackers can distribute harmful objectives across multiple benign-looking turns, making defense more complex than judging individual prompts. [episode]
- Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling — Quantization noise can accumulate over diffusion model denoising trajectories, degrading generation quality, and this paper introduces Q-Drift, a sampler-side correction that treats quantization error as an implicit stochastic perturbation to derive a marginal-distribution-preser [episode]
- UniMedSeg: Unified In-Context Learning for Multi-Paradigm 2D/3D Medical Image Segmentation — Medical image segmentation foundation models face fragmentation due to their reliance on paradigm-specific architectures and dimension-specific designs, which prevents heterogeneous annotations from being jointly absorbed by a single scalable model. [episode]
- PhoneWorld: From Real-App Trajectories to Dynamic and Verifiable Environments for Phone-Use Agents — A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. [episode]
- NovelAPIBench: Diagnosing How A Code LLM Learns to Use Novel APIs — Large Language Models for code generation frequently navigate novel APIs absent from their pretraining data, requiring coordination of heterogeneous knowledge components like signatures, module paths, and usage patterns. [episode]
- Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models — Large language models are typically optimized for accuracy and, therefore, will guess even when uncertain about their predictions. [episode]
- OVD: On-policy Verbal Distillation — On-policy Verbal Distillation (OVD) introduces a memory-efficient framework that transfers reasoning capabilities from large teacher models to smaller student models using discrete verbal feedback. [episode]
- Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks — General-purpose agents like OpenClaw are increasingly used as autonomous tool users, but their coding ability remains difficult to measure under SWE-bench because generic agents do not inherently satisfy the required clean Docker workspace, patch, and prediction contract. [episode]
- Rethinking Object-Centric Representations for Video Dynamics Modeling — Unsupervised video object tracking aims to decompose dynamic scenes into persistent, object-centric entities without manual annotations, and this work proposes STAITUS, a unified framework that explicitly disentangles object appearance from spatial pose to yield sharper masks and [episode]
- AdversaRiskQA: An Adversarial Factuality Benchmark for High-Risk Domains — AdversaRiskQA introduces a novel benchmark for evaluating Large Language Models (LLMs) under adversarial factuality conditions in high-risk domains like health, finance, and law. [episode]
- Coverage-aware Semantic Representation Learning for Underrepresented Visual Concepts — Semantic Coverage Imbalance (SCI) is a previously overlooked bias arising from long-tailed semantic representations that affects how models learn and reason about rare yet meaningful semantics. [episode]
- One-Forcing: Towards Stable One-Step Autoregressive Video Generation — One-Forcing proposes a simple yet effective approach that augments Distribution Matching Distillation (DMD) with an auxiliary GAN loss to achieve high-quality and efficient one-step video generation, establishing state-of-the-art performance among one-step causal video generation [episode]
- Sampling Meets Interaction: Sequentially-Controlled Multi-Particle Flow-Maps for Efficient Inference-time Search — Sequentially-Controlled Interactive Multi-Particle Flow-Maps (IMPFM) is a framework designed for sample-efficient online feedback-driven search by progressively transporting particles toward a target distribution while maintaining broad coverage essential for heterogeneous prefer [episode]
- Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing — Image compositing aims to seamlessly insert a foreground object into a background image, and recent advances in diffusion models have significantly enhanced the quality, especially when the foreground and background images come from different domains. [episode]
- CompDiff enables fair and zero shot medical image generation across demographic intersections through compositional diffusion — Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that generators produce equally high-quality images across demographic groups. [episode]
- Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking — Adversarial Reward Auditing (ARA) is a framework that reframes reward hacking as a competitive game between an actively discovering Hacker and an Auditor that detects exploitation, offering a dynamic and adaptive defense against reward model vulnerabilities in Reinforcement Learn [episode]
- KANFIS: A Neuro-Symbolic Framework for Interpretable and Uncertainty-Aware Learning — Adaptive Neuro-Fuzzy Inference System (ANFIS) was designed to combine neural network learning with fuzzy logic reasoning, but conventional architectures suffer from structural complexity and rule explosion. [episode]
- Reinforced Planning with Latent World Models — Humans solve complex problems by constructing plans and mentally simulating their outcomes with an internal model of the world, and this work introduces Reinforced Planning (RP1), a method that learns how to improve multi-step plans by reinforcing good search rules into a neural [episode]
- Flash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier Transform — Equivariant networks offer significant parameter efficiency by embedding geometric symmetries as structural priors, but this efficiency often fails to translate into computational speedup because existing implementations treat structured weights as generic dense matrices. [episode]
- How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift — Task adaptation in large language models (LLMs) is an alignment intervention that reshapes pre-existing model behavior in a structured, dimension-dependent way. [episode]
- Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance — This research introduces DGEval, a novel benchmark designed specifically to evaluate Large Language Model (LLM) knowledge of the International Maritime Dangerous Goods (IMDG) Code under Amendment 42-24. [episode]
- PEST: Parameter Efficient Steering of Blackbox VLMs via Agentic Few-shot Alignment for Hateful Meme Moderation — In this work, a novel framework is proposed to simultaneously address hateful meme moderation by combining classification, explanation, and intervention using task-specific generative AI agents. [episode]
- P 2O: Joint Policy and Prompt Optimization — Reinforcement Learning with Verifiable Rewards (RLVR) often suffers from advantage collapse on hard samples, which eliminates crucial learning signals, and this paper introduces Joint Policy and Prompt Optimization (P2O) to mitigate this collapse by alternating continuous policy [episode]
- One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy — OneWM-VLA introduces a method to compress per-frame visual information into a single semantic token, demonstrating that this reduction in visual bandwidth does not compromise long-horizon control when integrated with a joint flow-matching objective. [episode]
- Investigating Learner-Aware Design of LLM-Generated Educational Feedback — Although large language models (LLMs) show promise for generating educational feedback, it remains unclear how feedback should be designed to support answer revision and learner acceptance across diverse profiles. [episode]
- When, How Long and How Much? Interpretable Neural Networks for Time Series Regression by Learning to Mask and Aggregate — Time series extrinsic regression (TSER) aims to predict a continuous target variable from an input time series, and this work introduces MAGNETS, an inherently interpretable neural architecture that learns input-dependent masks and aggregates them into meaningful concepts without [episode]
- Two-Stage Learned Decomposition for Scalable Routing on Multigraphs — Most neural methods for Vehicle Routing Problems (VRPs) are limited to Euclidean settings or simple graphs, but this work introduces a NodeEdge Policy Factorization (NEPF) approach that splits routing into a node permutation stage and an edge selection stage to enable scalable le [episode]
- Learning to Generate Rigid Body Interactions with Video Diffusion Models — Recent video generation models struggle to produce physically plausible object interactions and lack object-level control mechanisms, which limits their utility as world simulators for robotics. [episode]
- TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking — Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities, and this work introduces TELLER, a dual-path framework that learns iteratively from model errors and reasoning to improve accuracy on table entity linking tasks. [episode]
- When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints — Safety alignment in large language models (LLMs) is primarily evaluated under open-ended generation, but this work identifies a systematic failure mode where reformulating harmful requests as forced-choice multiple-choice questions can systematically bypass refusal behavior, even [episode]
- PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition — Zero-shot skeleton-based action recognition (ZSSAR) is typically treated as a skeleton-text alignment problem, but this approach suffers from an upstream semantic loss where crucial object and pose-relative cues are lost during the initial human pose estimation (HPE) process. [episode]
- BALAR: A Bayesian Agentic Loop for Active Reasoning — As an excellent, fastidious, and diligent researcher, I have meticulously analyzed the provided excerpts regarding "BALAR (Bayesian Agentic Loop for Active Reasoning)." My task is to synthesize these pieces into a comprehensive, detailed summary of the scientific paper described. [episode]
- Capability Provenance in Language Models: A Case Study in Social Reasoning — As a fastidious and diligent researcher, I have meticulously analyzed these excerpts from "Capability Provenance in Language Models: A Case Study in Social Reasoning." The provided text details a sophisticated methodology employing Training-Data Attribution (TDA) to map corpus re [episode]
- UniPool: Learning Expert-to-Layer Ownership from Brief Global Access — Mixture-of-Experts (MoE) architectures are being challenged by rigid per-layer expert allocation rules, which may lead to redundant expert capacity. [episode]
- CODEBLOCK: Learning to Supervise Code at the Right Granularity — Supervised fine-tuning of code LLMs typically applies uniform cross-entropy loss to all response tokens, implicitly assuming that every token provides equally useful learning signal. [episode]
- Contrast Matters: Understanding Robustness of In-Context Fine-Tuning to Target-Context Relatedness — Training strategies for models to balance in-context learning (ICL) and in-weights learning (IWL) and switch between them based on context relevance are crucial for continuous adaptation without further training. [episode]
- Higher-Order Geometric Updates for Levenberg-Marquardt Method via Riemann Normal Coordinates — Nonlinear least-squares optimization is central to many computational problems, and this work introduces a method that improves geometric consistency for finite steps in optimization by using higher-order corrections derived from Riemann normal coordinates. [episode]
- D-GAP: Improving Out-of-Domain Robustness via Dataset-Agnostic and Gradient-Guided Augmentation in Frequency and Pixel Spaces — Out-of-domain (OOD) robustness remains a significant challenge in computer vision because models degrade when applied to new environments, and generic or dataset-specific augmentations often fail to provide consistent gains. [episode]
- Beyond Multimodal Alignment: Shared Physical Representations Across Sensors and Action Orders — World models increasingly treat compact multimodal representations as interfaces between perception and physical interaction, yet existing probes do not establish whether information acquired through different sensors carries the same executable meaning, or whether it survives a [episode]
- VisionLogic: Discovering and Grounding Decision-Relevant Visual Concepts — VisionLogic introduces a novel neuralsymbolic framework that produces faithful, hierarchical explanations as global logical rules over causally validated concepts, addressing the limitation of prior methods that rely solely on correlational signals. [episode]
- Federated Computation of ROC and PR Curves — Receiver Operating Characteristic (ROC) and Precision-Recall (PR) curves are fundamental tools for evaluating machine learning classifiers, offering detailed insights into the trade-offs between true positive rate vs. false positive rate (ROC) or precision vs. recall (PR). [episode]
- Design of Experiment for Discovering Directed Mixed Graph — We study experimental design for accurately identifying directed mixed graph structures, which are causal graphs that include both feedback loops (cycles) and unobserved confounders (bidirected edges). [episode]
- Dex2HOI: Dexterous Bimanual Two-Object Interaction Generation — Dex2HOI presents a unified diffusion model designed to synthesize dexterous bimanual manipulations involving up to two objects from text, addressing the research gap concerning coordinated multi-object interaction in human behavior. [episode]
- Agent Collectives Should Not Detect Their Own Imposters: A Chess Case Study — As a fastidious and diligent AI researcher, I have meticulously reviewed the provided excerpts from two sources concerning "GAMBIT" and its related work on multi-agent systems (MAS) and imposter detection. [episode]
- StreamPPG: Low-Latency rPPG Estimation via Consistent Privileged Learning — Remote photoplethysmography (rPPG) estimates blood volume pulse signals from facial videos, but conventional methods suffer from significant latency or reduced accuracy. [episode]
- DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams — Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams. [episode]
- MMORF: A Multi-agent Framework for Designing Multi-objective Retrosynthesis Planning Systems — Multi-agent systems offer a promising approach for multi-objective retrosynthesis planning by leveraging interactions among specialized agents to incorporate multiple objectives into retrosynthesis planning. [episode]
- Learning Domain- and Class-Disentangled Prototypes for Domain-Generalized EEG Emotion Recognition — Electroencephalography (EEG)-based emotion recognition faces significant hurdles due to inter-subject variability, reliance on target-domain data, and label noise; this paper proposes a Multi-domain Aggregation Transfer learning framework with domain–class prototypes (MAT) to a [episode]
- Correcting Spectra Outside the Backbone: A Model-Agnostic Rectifier for Hyperspectral Image Super-Resolution — HSI-SR aims to enhance spatial resolution while preserving spectrally faithful and physically plausible characteristics, addressing the issue where common super-resolution methods neglect spectral consistency across bands. [episode]
- EFC++: Elastic Feature Consolidation with Prototype Re-balancing for Cold Start Exemplar-free Incremental Learning — Exemplar-free Class Incremental Learning (EFCIL) aims to learn from sequential tasks without access to previous task data, and this paper addresses its most challenging scenario: Cold Start, where insufficient data exists in the first task to learn a high-quality backbone. [episode]
- RL Forgets! Towards Continual Policy Optimization — Continual post-training for vision-language models (VLMs) has become central to adapting them to evolving tasks, but existing evidence suggests that reinforcement learning (RL) is not inherently robust against catastrophic forgetting. [episode]
- Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations — Speech representations from self-supervised speech models (S3Ms) are known to be sensitive to phonemic contrasts, but their sensitivity to prosodic contrasts has not been directly measured. [episode]
- Diffusion Masked Pretraining for Dynamic Point Cloud — Dynamic point cloud pretraining is advanced by Diffusion Masked Pretraining (DiMP), a unified self-supervised framework that leverages diffusion modeling to address two critical limitations in existing methods: spatiotemporal positional leakage and the collapse of distributional [episode]
- GeoFidelity-Bench: Evaluating Block-Conditioned Geographic Fidelity in Text-to-Image Street-View Generation — Text-to-image models can generate visually plausible city streets, but whether their outputs correspond to a requested road segment rather than a generic city prior remains unclear. [episode]
- AD-SAM: Adapting the Segment Anything Model for Semantic Segmentation in Autonomous Driving — AD-SAM presents a fine-tuned vision foundation model designed for semantic segmentation in autonomous driving, significantly enhancing performance over existing models by integrating dual encoders and a hybrid loss function. [episode]
- J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules — Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks, but they typically expose only final labels, leaving the decision knowledge acquired through fine-tuning implicit within the model. [episode]
- MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation — Large language models have made substantial progress in mathematical reasoning, but benchmark development for multilingual evaluation has lagged behind English in both difficulty and recency. [episode]
- Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety — Clinician pairwise preferences are found to be an unreliable signal for assessing clinical safety because preferred responses do not consistently align with safety-critical rubric scores. [episode]
- Selective Fine-Tuning for Targeted and Robust Concept Unlearning — Text-guided diffusion models are being exploited to generate harmful content, necessitating robust concept unlearning methods that can selectively remove undesirable concepts without degrading generation quality. [episode]
- Distribution-Aware Programming: Learning Specialized Solvers from Experience — As a fastidious and diligent researcher, I have meticulously analyzed both provided texts to construct a comprehensive and detailed summary of the paper, "Distribution-Aware Programming: Learning Specialized Solvers from Experience." Here is my synthesis: * This research introduc [episode]
- Models Designed to Forget: Machine Unlearning via Key Deletion — Machine unlearning is rapidly becoming a practical requirement, driven by privacy regulations, data errors, and the need to remove harmful or corrupted training samples. [episode]
- Learning beyond Site Bias for OOD Generalization in Brain Networks — Graph-based learning on functional magnetic resonance imaging (fMRI) has shown strong potential for brain network analysis, but existing methods degrade under cross-site out-of-distribution (OOD) settings because site-conditioned confounders induce non-pathological shortcuts, whi [episode]
- SalamahBench: Dialect and Category Level Safety Evaluation of Arabic Language Models — SalamahBench introduces a unified benchmark for evaluating Arabic Language Models (ALMs) safety, addressing critical gaps in existing English-centric safety resources by providing category-aware evaluation across 12 MLCommons hazard categories. [episode]
- SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing — Scientific diagrams convey explicit structural information, yet modern text-to-image models often produce visually plausible but structurally incorrect results. [episode]
- Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya — Multilingual pre-trained language models struggle with low-resource Ge’ezscript languages like Amharic and Tigrinya due to high out-of-vocabulary rates and excessive subword fragmentation from Latin-script tokenizers, a problem this paper addresses by introducing VEXMLM, a voca [episode]
- Agentic Forecasting with Structured Linguistic Beliefs — This research introduces the Bayesian Linguistic Forecaster (BLF), an agentic system designed for binary forecasting that demonstrates state-of-the-art performance on the ForecastBench benchmark. [episode]
- LaSEr-Edit: Localized Span-level Error Editing with Energy-based Localization — As a fastidious and diligent AI researcher, I have thoroughly reviewed the provided text snippets from "LaSEr-Edit: Localized Span-level Error Editing with Energy-based Localization." My analysis indicates a method designed for constraint-satisfying text revision using an energy- [episode]
- Self-Play Enhancement via Advantage-Weighted Refinement in Online Federated LLM Fine-Tuning — SPEAR introduces an efficient online learning algorithm for federated LLM fine-tuning that utilizes a feedback-guided self-play loop to construct naturally contrastive pairs, enabling model improvement without requiring privileged ground-truth contexts or expensive group generati [episode]
- What is Missing from AI Post-Training AI: An Empirical Analysis — Large language model (LLM) agents can now post-train an LLM end-to-end, raising the prospect of AI-for-AI. [episode]
- Plain Transformers are Surprisingly Powerful Link Predictors — Plain Transformers are Surprisingly Powerful Link Predictors because they demonstrate that plain Transformer architectures can serve as highly effective link predictors under strict deployment constraints, such as operating on fixed-budget, sampled local subgraphs, without relyin [episode]
- Interp3R: Continuous-time 3D Geometry Estimation with Frames and Events — Interp3R introduces a novel framework that extends existing pointmap-based models to estimate depth and camera poses in continuous time by leveraging asynchronous event data. [episode]
- GLI-AL: A Multi-Modal Glioma MRI Label Resource with Unified Anatomy-Lesion Labels — BraTS datasets provide multi-center, pre-operative, multi-parametric MRI and expert tumor-subregion annotations for brain tumor segmentation research, but they are among the central public benchmarks for machine learning in glioma imaging. [episode]
- Enabling Regulatory Multi-Agent Collaboration: Architecture, Challenges, and Solutions — Large language models (LLMs)-empowered autonomous agents are transforming both digital and physical environments by enabling adaptive, multi-agent collaboration. [episode]
- Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale — Enterprise AI aims to move toward continuous event monitoring, detection, and action across specialist agents, yet existing multi-agent systems largely assume discrete request-response workflows and remain underexplored at enterprise scale. [episode]
- Reasoning Shift: How Context Silently Shortens LLM Reasoning — Reasoning models exhibit a tendency to produce significantly shorter reasoning traces when solving problems under different context conditions compared to when the problem is presented in isolation. [episode]
- WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning — Large language models often fail to perform at expert levels in specialized technical mathematics, particularly in wireless communications, where problems require precise handling of information-theoretic bounds and complex signal processing formulations. [episode]
- Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor Programs — In modern theoretical analyses of neural networks, which often rely on Gaussian approximations in the infinite-width limit, this work rigorously identifies that attention layers depart fundamentally from Gaussianity under realistic architectural dimensionality and standard scalin [episode]
- Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations — Explanations, particularly Chain-of-Thought (CoT) reasoning in vision language models, present a double-edged sword because they can simultaneously clarify complex judgments and foster dangerous blind trust. [episode]
- Investigating Single-Block Recurrence in Vision Transformers for Image Recognition — A single-block recurrent Vision Transformer (bViT) architecture is investigated to determine how much of a deep ViT's performance can be realized through recurrent computation rather than layer-specific parameterization, revealing that sufficient representation width allows a sha [episode]
- Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning — Reinforcement learning fine-tuning is addressed through a game-theoretic framework that provides an explicit statistical interpretation for setting the KL regularization coefficient, moving beyond heuristic choices. [episode]
- Quantifying the Effect of Test Set Contamination on Generative Evaluations — This research meticulously investigates the critical threat posed by test set contamination—the unintentional inclusion of benchmark data within a language model's pretraining corpus—to the trustworthiness and reliability of evaluating frontier AI systems. [episode]
- Cross Modality Image Translation In Medical Imaging Using Generative Frameworks — As a diligent researcher, I have meticulously analyzed both provided texts from arXiv and synthesized them into a comprehensive, detailed summary of this work. [episode]
- NOSA: Native and Offloadable Sparse Attention — Decoding throughput improvements from larger inference batches are limited by GPU memory, which is largely consumed by the key-value (KV) cache. [episode]
- SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning — Latent world models are powerful planning paradigms that have struggled with proposal quality as planning horizons grow, and this paper introduces SAGE, a prior-conditioned planner that uses latent subgoal decomposition to structure action search. [episode]
- A Gravitational Interpretation of Safety Reversion under Fine-Tuning — Fine-tuning on ordinary data can partially reverse behaviors acquired earlier in training, and this paper proposes that these phenomena are usefully viewed through a common training-history lens. [episode]
- Learning to Predict Future-Aligned Research Proposals with Language Models — Large language models are increasingly used to assist researchers in idea exploration, but evaluating their generated research proposals remains difficult because novelty and soundness are hard to measure automatically. [episode]
- Pseudo-differential-enhanced physics-informed neural networks — Pseudo-differential-enhanced physics-informed neural networks (PINNs) introduce an extension of gradient enhancement applied in Fourier space to improve training fidelity and learning efficiency for solving partial differential equations. [episode]
- CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from A Single-View Image — CATSplat introduces a novel generalizable transformer-based framework for 3D Gaussian Splatting that reconstructs 3D scenes from a single-view image by leveraging textual guidance and spatial guidance to overcome the inherent constraints of monocular settings. [episode]
- HighSync: High-Quality Lip Synchronization via Latent Diffusion Models — HighSync presents an end-to-end diffusion-based framework for high-fidelity lip synchronization that generates photorealistic talking-face videos aligned with arbitrary input audio, addressing prior limitations in both image quality and temporal consistency. [episode]
- Token Distribution versus Data Volume: Domain Balancing in Multi-Domain Meeting Summarisation — Jointly fine-tuning an LLM on meeting summarisation corpora of widely varying size raises a question that prior work leaves confounded: when a domain-balanced training mixture helps, is the gain due to the distribution of tokens across domains, or merely to the volume of data see [episode]
- FIRM: Flow-based Imaging via Regularized Minimization — Flow matching approaches to imaging inverse problems commonly incorporate measurements in two ways. [episode]
- Gaussian Relational Graph Transformer — Relational graph learning models relational databases as graphs and has demonstrated superior performance on a wide range of relational predictive tasks. [episode]
- RooseBERT: A New Deal For Political Language Modelling — RooseBERT introduces a novel pre-trained language model specifically tailored for English political discourse, addressing the limitations of general-purpose models in capturing domain-specific features like implicit argumentation and strategic communication. [episode]
- Deep Learning with Pretrained 'Internal World' Layers: A Gemma 3-Based Modular Architecture for Wildfire Prediction — Deep learning models, especially large Transformers, carry substantial "memory" in their intermediate layers—an internal world that encodes a wealth of relational and contextual knowledge. [episode]
- TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment — Fine-tuning large language models (LLMs) via Fine-Tuning-as-a-Service (FTaaS) platforms can erode inherent safety alignment, necessitating post-training recovery mechanisms that restore safety without destroying task utility. [episode]
- LISA: Likelihood Score Alignment for Visual-condition Controllable Generation — LISA proposes an effective regularization method, Likelihood Score Alignment, that explicitly aligns the intermediate features of a side network with an approximated likelihood score to enhance visual-condition controllable generation. [episode]
- Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models — Multi-domain Contrastive Policy Optimization (MCPO) is a structured contrastive reinforcement learning framework proposed to enhance Large Reasoning Models by transforming cross-domain interactions from harmful competition into beneficial knowledge transfer. [episode]
- Cliff Tokens: Analyzing Failure Trigger Tokens in LLM Mathematical Reasoning — Large language models (LLMs) exhibit inconsistent reasoning paths, where some traces succeed while others fail, and this study introduces "cliff tokens," which are precise tokens where token-wise potential drops significantly. [episode]
- On the Interaction of Compressibility and Adversarial Robustness — Modern neural networks are expected to simultaneously satisfy a host of desirable properties: accurate fitting to training data, generalization to unseen inputs, parameter and computational efficiency, and robustness to adversarial perturbations. [episode]
- Explicit Layer Modeling for Video Object Insertion and Video Layer Decomposition — Most video editing systems lack explicit layered video representations, limiting their ability to perform realistic compositing and consistent manipulation, which is especially problematic in video object insertion and layer decomposition. [episode]
- SWARD: Stochastic Window-Attention-Based Relational Distillation for Cross-Architectural Semantic Segmentation — Large-scale vision foundation models drive gains in dense prediction tasks like semantic segmentation, but their size limits deployment, motivating knowledge distillation to transfer capabilities from heavy transformer teachers to lightweight convolutional students. [episode]
- FinEvolveBench: A Benchmark for Self-Evolving Agents on Low-Repetition Tasks with Implicit Rewards — Experience-based self-evolution enables language-model agents to improve their behavior by accumulating and updating experience at test time, yet existing evaluations often assume recurring task patterns and explicit success signals. [episode]
- BnMMLU: Measuring Massive Multitask Language Understanding in Bengali — BnMMLU introduces a comprehensive benchmark for measuring massive multitask language understanding in Bengali, addressing the critical gap in standardized evaluation for low-resource languages. [episode]
- FacePlex: Toward Natural Full-Duplex Conversational Avatars — Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion. [episode]
- Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies — Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies by enforcing sparsity, locality, and competition to create models that adaptively partition representation space around statistical imbalances. [episode]
- Graph Your Own Prompt — As a fastidious and diligent AI researcher, I have meticulously analyzed both provided summaries of Graph Consistency Regularization (GCR) from arXiv. [episode]
- CRAX: Fast Safe Reinforcement Learning Benchmarking — Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving, and this work addresses the gap between existing RL benchmarks and high-fidelity simulation by proposing CRAX, a hardware-accelerated benchmark [episode]
- The Routing Plateau: Understanding the Accuracy Limits of LLM Routers — Many current LLM routing methods converge to a narrow performance range far below the oracle router, indicating fundamental limits in their ability to handle query-specific routing decisions. [episode]
- Large Language Model Selection with Limited Annotations — LLM SELECTOR introduces a principled framework for active model selection that efficiently identifies the best Large Language Model (LLM) for a given task under limited annotation budgets, significantly reducing annotation costs. [episode]
- When Does Depth Matter For In-Context Learning? Adaptive Inference in Deep Transformers — Transformers are theorized to implement distributed inference over vectorized tokens, where function vectors act as compressed state variables and MLP blocks choose which statistic should be measured next, suggesting that depth and MLP blocks enable a richer class of in-context l [episode]
- DeepForestVisionV2: Ecology-Driven Taxonomy Expansion for Camera-Trap Monitoring in African Tropical Forests — Camera-trap monitoring in African tropical forests is being enhanced by DeepForestVisionV2, an ecology-driven expansion that addresses deployment challenges across vertical stratification, scene openness, and anthropogenic interfaces. [episode]
- Extending SSMs with the Exponentially Weighted Signature — The Exponentially Weighted Signature (EWS) generalizes classical path signatures by replacing uniform historical weighting with bounded linear operators, offering richer temporal dynamics and cross-channel coupling. [episode]
- TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training — On-policy distillation (OPD) for long-horizon agent training remains inefficient because vanilla methods waste computational resources on low-yield tail turns and concentrate supervision loss disproportionately on shallow tokens. [episode]
- UniRect-CoT: Enhancing Generation in Unified Multimodal Models via Reflective Rectification with Inherent Understanding — Unified Multimodal Models (UMMs) exhibit a capability mismatch where their understanding significantly outperforms their generation, suggesting that rich internal knowledge remains underactivated during image synthesis. [episode]
- Slot-RAE: Streamlining Object-Centric Learning via Direct Representation Auto-Encoders — Deploying object-centric models for real-world scene understanding typically requires complex pipelines to achieve both robust scene decomposition and high-fidelity generation. [episode]
- EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration — EverAnimate is an efficient post-training method designed for long-horizon animated video generation that preserves visual quality and character identity by restoring drifted flow trajectories through persistent latent context memory and intrinsic restoration capabilities. [episode]
- Semantically Similar, Yet Not Answerable: Diagnosing the Semantic-Answerability Gap in Table RAG — Tables are critical knowledge sources in retrieval-augmented generation (RAG), but retrieved tables often lack sufficient evidence to answer a query, creating a fundamental mismatch between semantic relevance and answerability. [episode]
- PyroAdapt: Adapting Wildfire Prediction under Spatial Heterogeneity and Temporal Shift — Adapting wildfire prediction under spatial heterogeneity and temporal shift is crucial for reliable forecasting in changing environments. [episode]
- BabelSafe: A Policy-Grounded Multilingual Safety Benchmark for LLMs — ML-BENCH and ML-GUARD introduce a policy-grounded multilingual safety benchmark and guardrail model, respectively, to address the limitations of existing multilingual safety evaluations that rely on general risk taxonomies. [episode]
- Let the Target Select for Itself: Data Selection via Target-Aligned Paths — Targeted data selection aims to identify training samples from a large candidate pool that improve performance on a specific downstream task, and this work proposes Target-Aligned Candidate Selection (TACS), an alternative reference path that decouples candidate utility scoring f [episode]
- BGM-IV: AI-Powered Bayesian Generative Modeling for Instrumental Variable Regression with High-Dimensional Covariates — BGM-IV proposes a latent Bayesian generative modeling approach that reframes nonlinear instrumental variable regression as posterior inference in a causally structured latent space, providing a principled and effective strategy for estimating causal effects when dealing with high [episode]
- USS: Unifying Spatial-Semantic Prompting for End to End Embodied Visual Tracking — Embodied Visual Tracking (EVT) requires an agent to continuously follow a specified target while actively moving through dynamic environments, and this work proposes a paradigm shift from language-only target indication to unified spatial-semantic prompting. [episode]
- Beyond Imitation: Reflective On-Policy Self-Distillation for LLM Reasoning — On-policy self-distillation (OPSD) improves large language model reasoning by providing dense token-level supervision during on-policy rollouts, but existing methods often fail to generalize robustly beyond in-domain data. [episode]
- Neural Cluster First, Route Second: Capacitated Vehicle Routing via Differentiable Optimal Transport — Neural CFRS introduces a purely non-autoregressive neural Cluster-First-Route-Second method for the Capacitated Vehicle Routing Problem (CVRP) that leverages differentiable Optimal Transport to enforce global fleet capacity constraints end-to-end. [episode]
- Training-free image inversion for one-step diffusion models — This work introduces a novel training-free inversion (TFinv) framework designed specifically for one-step diffusion models, addressing critical challenges in real image inversion and editing by tackling initial latent editability and caption gap issues. [episode]
- Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters — Chronicles-OCR introduces a comprehensive benchmark designed to evaluate the cross-temporal visual perception capabilities of Vision Large Language Models (VLLMs) across the entire evolutionary trajectory of Chinese characters, spanning from Oracle Bone Script to Cursive Script. [episode]
- N"urnberg NLP at PsyDefDetect: Multi-Axis Voter Ensembles for Psychological Defence Mechanism Classification — Detecting levels of psychological defence mechanisms in supportive conversations is inherently ambiguous, and this research addresses that ambiguity by developing a multi-axis voter ensemble to achieve robust classification. [episode]
- From Privacy to Generalization: Linear Max-Information Bounds for Differentially Private Learning Algorithms — Understanding how differentially private stochastic gradient descent (DP-SGD) can be used to derive generalization guarantees for deep networks is a central challenge in modern machine learning theory, particularly because existing methods often fail to provide non-vacuous bounds [episode]
- On the Invariance and Generality of Neural Scaling Laws — Neural scaling laws establish a predictable relationship between model performance and data or compute, offering crucial guidance for resource allocation in new domains and tasks. [episode]
- Faster Inference of Flow-Based Generative Models via Improved Data-Noise Coupling — Conditional Flow Matching (CFM) provides an efficient alternative to diffusion models for training continuous normalizing flows, and this work introduces LOOM-CFM, a novel method that extends minibatch optimal transport by preserving and optimizing noise-data pair assignments acr [episode]
- Floquet Fibre Geometry and Higher-Order Reduced Coordinates for Off-Manifold Transients near Nonlinear Aeroelastic Flutter — Assigning a reduced coordinate to a full state near an attracting limit cycle is fundamentally an issue of invariant-fibre geometry, where the choice of projection determines whether one achieves an accurate description of the dynamics. [episode]
- EviLink: Multi-Path Schema Linking with Uncertainty-Guided Evidence Acquisition for Large-Scale Text-to-SQL — Schema linking is reframed as uncertainty-aware schema-need inference over multiple plausible SQL paths, where systems distinguish required schema items from path-dependent uncertain ones and acquire evidence only where needed. [episode]
- Hypernetworks for Dynamic Feature Selection — Dynamic feature selection (DFS) is a machine learning framework where features are acquired sequentially for individual samples under budget constraints, and this paper proposes Hyper-DFS, a hypernetwork-based approach that generates subset-specific classifier parameters on deman [episode]
- Tags for DAGs: Graph Refinement with Meta-Informed Relations — Not every causal relation between variables is equal, and this can be leveraged for causal discovery by assigning multiple tags to each variable in a graph. [episode]
- Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing — Instruction-based image editing (IIE) models often suffer from over-editing, introducing unintended changes to regions unrelated to the desired edit. [episode]
- Diverse Sampling in Diffusion Models with Divergence-Free Particle Guidance — EDDY (Exact-marginal Diversification via Divergence-free dYnamics) is a training-free guidance mechanism for diffusion and flow matching models designed to promote sample diversity while strictly preserving each particle's marginal distribution. [episode]
- Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization — Chain-of-Thought (CoT) reasoning has revolutionized Large Language Models (LLMs) by decomposing complex problems into sequences of intermediate steps, but it suffers from computational cost and reasoning path collapse due to its reliance on discrete token spaces. [episode]
- Learning how to Forget: Fine-tuning for Long-Context Sparse Attention — A novel method for fine-tuning transformer language models with sparse attention is introduced, demonstrating that this approach can be effective on moderate hardware budgets and often outperforms exact attention methods. [episode]
- The Interplay of Harness Design and Post-Training in LLM Agents — Tool-integrated LLM agents are often wrapped within a harness, and this paper investigates how harness design influences post-training performance in both in-distribution and out-of-distribution settings. [episode]
- How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline — Large language models (LLMs) increasingly deploy only limited language settings, most notably “English (US),” despite global diversity and colonial history, raising foundational questions about which variety of English LLMs implicitly prefer. [episode]
- SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs — Soft Modality-Guided Expert Specialization (SMoES) addresses how modality-specific signals should guide expert routing in Mixture-of-Experts Vision-Language Models (MoE-VLMs), a critical area because existing routing strategies often rely on idealized priors that fail to capture [episode]
- AI-driven Dispensing of Coral Reseeding Devices for Broad-scale Restoration of the Great Barrier Reef — Coral reefs are facing imminent collapse due to climate change and pollution, making automated, large-scale restoration efforts crucial. [episode]
- DySem: Uncovering Dynamic Semantic Components of Large Language Models for Calculating Semantic Textual Similarity — Calculating semantic textual similarity is a foundational task in natural language processing, and current large language models (LLMs) typically rely on extracting last-layer hidden states with fixed dimensions to compute similarity for every text pair. [episode]
- Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties — Estimating volumetric mechanical properties, including Young’s modulus, Poisson’s ratio, and density at each voxel, is intrinsically ambiguous from vision alone because visually similar objects may have substantially different material compositions and physical behaviors. [episode]
- Synthetic data for ratemaking: imputation-based methods vs adversarial networks and autoencoders — Synthetic data generation, particularly for actuarial ratemaking, is explored as a solution to data scarcity and privacy concerns by benchmarking imputation-based methods against deep generative models. [episode]
- WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing — In an autoregressive Transformer, each layer typically attends only to KV produced at its own depth, limiting its ability to utilize deeper representations during decoding. [episode]
- Generalization Dynamics of Linear Diffusion Models — Diffusion models are powerful generative models whose generalization capabilities with finite data remain unclear, and this work addresses that gap by analyzing them through the lens of data covariance spectra. [episode]
- One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning — Large Language Models (LLMs) exhibit significant inconsistencies in moral judgments when evaluated across different languages and cultural contexts, often reflecting underlying cultural misalignment embedded in their English-centric pretraining data. [episode]
- FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices — Deploying Large Language Models (LLMs) on edge devices faces significant technical hurdles, primarily concerning memory elasticity in environments with unified, dynamic memory. [episode]
- Skip What You Can Predict: Predictive Repositioning for Policy Optimization for Efficient LLM Training — Gradient Extrapolation-Based Policy Optimization (GXPO) is a novel, plug-compatible policy-update rule designed to approximate longer local lookahead steps in GRPO-style reinforcement learning using only three backward passes. [episode]
- Scaling Vision Transformers for Functional MRI with Flat Maps — Scaling Vision Transformers for Functional MRI with Flat Maps introduces CortexMAE, a family of self-supervised foundation models trained on functional MRI (fMRI) data using a novel cortical flat map projection representation. [episode]
- Human vs. Machine Deception: Distinguishing AI-Generated and Human-Written Fake News Using Ensemble Learning — The rapid adoption of large language models has introduced a new class of AI-generated fake news that coexists with traditional human-written misinformation, raising important questions about how these two forms of deceptive content differ and how reliably they can be distinguish [episode]
- Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study — Despite growing interest in Multimodal Domain Generalization (MMDG) for enhancing model robustness, it remains unclear whether reported performance gains reflect genuine algorithmic progress or are artifacts of inconsistent evaluation protocols. [episode]
- Learning to Select Source Domains: Proxy-Rewarded Policy Optimization for Molecular OOD Generalization — Robust prediction of molecular properties under extreme out-of-distribution (OOD) scenarios is a pivotal bottleneck in AI-driven drug discovery, and this work addresses it by proposing a framework that intelligently selects optimal source domains for knowledge transfer. [episode]
- Equivalent Flows, Unequal Learning: Clean-Latent Prediction in Transformers — Flow matching with clean-data prediction has shown that regressing the clean point can exploit low-dimensional structure more effectively than predicting an ambient noised quantity. [episode]
- Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion — Masked diffusion models have shown promising performance in generating high-quality samples, but accelerating their sampling process remains relatively underexplored. [episode]
- Semantic Purification for Conditional Representation Learning — Conditional representation learning aims to extract criterion-specific features for customized tasks, but existing subspace projection methods suffer from sensitivity to basis quality and vulnerability to inter-subspace interference. [episode]
- Can Video World Models Track Unobserved World States? — Video world models are increasingly used as simulators, yet visual fidelity alone does not show that a model maintains the hidden state of the world. [episode]
- Agentic Search for Counterfactual Recourse under Fixed LLM Budgets — Counterfactual recourse generation under fixed LLM budgets shifts from finding a single optimal explanation to efficiently generating a diverse set of oracle-validated alternatives, which is crucial for users who benefit from multiple feasible options. [episode]
- Digitally enriching a high-risk population for pancreatic cancer using routine blood-based measures and clinical histories — Longitudinal sequences of coded diagnoses and blood test values accrued by patients throughout their clinical interactions were used to train a custom Transformer-based neural network with a multi-head attention mechanism to predict risk of pancreatic cancer with a multiyear lead [episode]
- LLARS: Enabling Domain Expert & Developer Collaboration for LLM Prompting, Generation and Evaluation — LLARS (LLM Assisted Research System) is an open-source platform designed to bridge the gap between domain experts and developers by unifying prompt engineering, batch generation, and hybrid evaluation into a single system. [episode]
- Training-Free Uncertainty Estimation for Embedding Models — The representation reliability of self-supervised learning models is crucial for their deployment in downstream tasks, and this paper introduces an ensemble-based method that estimates this reliability without prior knowledge of those tasks. [episode]
- SyncEdit: Rethinking Lip Synchronization as Editing with Audio-Driven Diffusion Models — Lip synchronization and audio-visual editing are fundamental challenges in multimodal learning, underpinning applications like film production and virtual avatars. [episode]
- Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity — In this study, researchers propose KME-CLIP to enhance CLIP's similarity computation by leveraging the inherent linear structure of Pointwise Mutual Information (PMI) within a Reproducing Kernel Hilbert Space (RKHS), theoretically proving that this approach can approximate PMI wi [episode]
- AutoExpert: Automating 3D LiDAR Annotation from Expert-Crafted Guidelines — Data annotation is crucial for developing machine learning solutions to numerous applications such as autonomous driving. [episode]
- Adaptive Activation Steering for Efficient LLM Reasoning via Closed-Loop PID Control — Reasoning LLMs often suffer from overthinking, where models spend excessive tokens on redundant reflection and transitions that inflate cost without improving accuracy. [episode]
- SGP-SAM: Self-Gated Prompting for Transferring 3D Segment Anything Models to Lesion Segmentation — Large segmentation foundation models, such as Segment Anything Model (SAM), have advanced promptable segmentation in natural images, but directly transferring these 3D SAM-style models to lesion segmentation remains challenging due to weak spatial representational capacity for sm [episode]
- SGAP-Gaze: Scene Grid Attention Based Point-of-Gaze Estimation Network for Driver Gaze — Driver gaze estimation is essential for understanding situational awareness, and this paper proposes SGAP-Gaze, a Scene Grid Attention based Point-of-Gaze estimation network that explicitly incorporates scene information into gaze modeling to achieve robust driver PoG estimation. [episode]
- MissClick: Execution-Aware Adversarial Attacks on Coordinate Generation in GUI Grounding Models — Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mapped to executable clicks, creating a security vulnerability that allows for large displacements in executed clicks based on the place-value str [episode]
- FoR-Net: Focus-on-Regions Network for Semantic Segmentation — FoR-Net is an efficient semantic segmentation framework that focuses on identifying and enhancing hard regions by employing a selector-driven Top-K mechanism, demonstrating competitive performance under resource-constrained settings. [episode]
- Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty — Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty. [episode]
- Large Language Models Hack Rewards, and Society — Reinforcement learning (RL) enables large language models (LLMs) to learn from rewards, and this capability can be exploited to discover loopholes in societal rules, leading to a failure mode termed "societal hacking." The central finding is that RL training may exploit gaps betw [episode]
- An Efficient Subspace Algorithm for Federated Learning on Heterogeneous Data — This work introduces FedSub, an efficient subspace algorithm designed for federated learning on heterogeneous data, addressing challenges related to client drift and high communication/computation costs. [episode]
- Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time — Multimodal large language models (MLLMs) often suffer from hallucinations because textual tokens dominate generation, undermining the potential of perceptual inputs. [episode]
- Lost in Motion: Vision Language Models Fail the Dynamic Gauges Test — Vision-Language Models (VLMs) are being evaluated as potential virtual instruments for reading analog gauges, but they fail to meet the necessary standards for reliable, real-time dynamic measurement. [episode]
- Energy Vision--Language--Action: A Controlled Multimodal Benchmark for Intent-Conditioned Residential Energy Management —
- Mitigating Hallucination in Large Language Models: A Capability-Oriented Survey on RAG, Reasoning, and Agentic Systems —
- SyncRA: Learning Temporal Correspondence in Omni-Modal Models —
- Hesitation-Aware On-Policy Distillation for Diffusion Language Models —
- OneFixer: High-Quality and Consistent One-Step Autoregressive 3DGS Refinement for Driving Scenes —
- ReadingMachine: A Computational Methodology for Structured Corpus Reading and Large-Scale Synthesis —
- Typed Temporal Interaction Features for Simulation-Backed Forecasting of Open-Source Game Release Incidents —
- The Ongiini-Eval-OW Benchmark: A Concept Paper for the Planned Benchmarking of Machine Translation and Large Language Models on Oshindonga and Oshikwanyama —
- Separating Diagnosis from Disease Representation: Dual-View EEG Learning with Neural-Dynamics-Guided Deformation —
- DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution —
- KoopCell: Koopman-Based Generative Model for Learning Single-Cell Dynamics from Distribution Snapshots —
- Reduce, Then Encode: Multiscale Volumetric Reduction for 2D Foundation Models in Brain MRI —
- Mixture-of-Top-k Attention: Efficient Attention via Scalable Fast Weights —
- Precise Editing and Flexible Referencing for Interactable Worlds —
- GenoMorph: Pathway-Grounded Genomic Disease Reasoning via Adaptive Latent Computation —
- T 5: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training —
- Is invariance all you need for algorithmic fairness? Removing demographic information can create new bias —
- SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation —
- LFPO: Likelihood-Free Policy Optimization for Masked Diffusion Models —
- What is Left for Us? Second Scholarship Against the Degradation of Research by AI —
- AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation —
- AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents —
- Forest-Guided Semantic Transport for Label-Supervised Manifold Alignment —
- MTLiquid: Enabling Efficient Multi-Task Learning using Liquid Neural Networks for Lightweight Healthcare Monitoring Systems —
- AmbiModBench: Benchmarking Gene Perturbation Prediction Beyond Shared Responses —
- Robust Biomolecular Complex Design Across Protein Conformational Landscapes —
- Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation —
- Recovery-Directed Symbolic Distillation of Neural Likelihoods —
- One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs —
- Beyond Gaussian Assumptions: Distribution-Aware Channel Capacity for Effective Connectivity —
- Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility —
- IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages —
- Spectral Saliency for Machine Unlearning —
- Recent Advances in Medical Imaging Segmentation: A Survey —
- Forecasting Intraday USD/CAD Exchange Rate with News-Derived Monetary-Policy Signals —
- Still There, No Longer Seen: Exposing Compression-Induced Risk in Large Vision-Language Models —
- BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation —
- PalmLeaf-VQA: A Multi-Script Visual Question Answering Benchmark for Historical Palm-Leaf Manuscript Understanding Across Diverse Regions —
- Temporal-Attention Head Specialization During Video Diffusion Training —
- SMARtCARE: Privacy-Preserving Agentic AI Systems for Bounded-Autonomy Clinical Decision Support —
- LLM Judge Validation Under Sparse Overlap: From Inference to Design —
- Auditing Quality Filters for Long-Tail Human Data Curation —
- CyberWorld: World Models for Sample-Efficient Autonomous Cyber Defense —
- CueKFS: Agentic Cue-Driven Keyframe Selection for Long Video Understanding —
- COUNTERMEM: World-Model Verified Counter-Factual Memory for Language Agents —
- IndustryLLM: Failure-Driven LLM Training for Industrial Procurement —
- A Surgical Foundation Model Reveals Task-Dependent Label Efficiency —
- SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions —
- Context-dependent agent evaluation with orthogonal equilibrium learning —
- DOHF: Online Diffusion Fine-tuning with Doob's h-transform Guidance —
- Metro-WM: Long-Horizon Latent Planning with Realisable Sub-Goals —
- Improving Medical Calculation of LLMs with Embedded Coding —
- ROTE: Benchmarking Neural Memorization on Complexity-Controlled Symbolic Sequences —
- CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs —
- Symbolic Guidance for LLM Agents in Distributed Multiagent Coordination —
- Communication between Frozen Large Language Models via Prompt Optimization in a Referential Game —
- What Does the Rank Buy? A Spectral and Distributional Analysis of Low-Rank Adaptation —
- VC Dimension and Expressivity of Real-Valued Transformers —
- Memory as Middleware for Self-Improving AI Agents —
- READ-Bench: Benchmarking Historical Instance Retrieval for Time-Series Diagnosis —
- Contract monitoring: governing AI via separation of powers —
- GameBoyWorlds: A Testbed for Self-Improvement in Embodied Video Games —
- When Should a Human Take Back Control? Optimal Delegation under Turbulent AI Risk —
- Escaping Alignment: A Physical Trap Model of Best-of-N Jailbreaking —
- Noisy Test-Time Reinforcement Learning for Code LLMs —
- Residual Streams Read, Recurrent States Remember: The Global Workspace in Mamba Models —
- PastForward: Faster On-Device GUI Agents via Computational Experience Reuse —
- AI Harness: Certification under Proposal-Conditioned Information for Foundation-Model Agents —
- Fracast-0: Fractal Weight Sharing for a Time Series Foundation Model with Only 85K Parameters —
- Witness: Discovery, Deciphering, and Epiphany in Interactive Puzzle Environments —
- A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases —
- RAO-Nav: Probing Omni-Language Models for Zero-shot Semantic Audio-Visual Navigation —
- Certifying Interventional Agreement Among Observationally Equivalent Causal Models —
- Phase Space Attention:A Hairer Lift Circumvents the Single-Layer Induction Obstruction —
- Delayed Supervision for Test-Time Language Models —
- Active Feature Acquisition With Incomplete Training Data —
- What Can a Leaderboard Certify? Compositional Controllability for Fair Evaluation and Training of Biomedical Literature-Review Agents —
- Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging —
- Agentsensus: Consensus-Compressed Shared Memory for Multi-Agent Story Worlds —
- GLIDE: Generalized Layer-wise Intrinsic Distributional Evaluation for Heterogeneous LLM Agents —
- HyperReCo: Retrieving and Connecting Evidence with Hypergraph Neural Networks for LLM Multi-hop Reasoning —
- Measurement Boundaries in LLM Financial Agent Evaluation: Fixed-Tape Execution and Multi-Defect Auditing —
- Function Over Form: Distributional Orthogonalization in Mixture-of-Experts with Replica Expert Mechanism —
- Black-Box Auditing of Epistemic Reliability in Multi-Agent Debate Distillation —
- TimeES: Probabilistic and Deterministic Time Series Forecasting via Evolutionary Spectra —
- From Trajectories to Grounded Preferences: Process Preference Synthesis via Interaction Element Graphs for Web PRMs —
- Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization —
- Reward Hacking and Agent Containment Failure: A Monte Carlo Study Based on the 2026 Hugging Face Incident —
- Automatic Speech Recognition for the Basa` a Language: A Low-Resource Approach —
- PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers —
- Authorization Closure Graph: Minimal Repair for LLM Agents with Evolving User Instructions —
- From Latents to Wires: Surgical Post-Editing on Large Language Models —
- Seeing Parts, Reasoning about Worlds: Visual Inference under Partial Observation —
- De-biasing Skeleton-based Action Recognition with Convex Hull Adaptive Shift —
- Towards Scalable Data Diversification for Language Model Pretraining via Leverage Score Sampling —
- PULSE: Identifying Demonstration-Utility Features with Sparse Autoencoders —
- From Outcomes to Strategies: Learning Strategy Utility for Mathematical Reasoning —
- VPEvolve: A Self-Evolving Virtual Process Engineer for Computational Lithography —
- Learning from Others, Acting for You: Cross-User Memory Sharing for LLM Agents —
- LocalProp: Neuro-Localized Memory-Efficient Backpropagation —
- MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents —
- Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations —
- STR: Supervised Transcoder Replacement for Reducing Steering Side Effects —
- Activation Flow: Manufacturing Activations for Steering —
- CUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations —
- MA-FPPO: Multi-Agent Flow-Pretrained Policy Optimization —
- ProTTT: Learning to Learn Semantic User Memory with Test-Time Training —
- Fail Loudly: An Auditable Runtime for Agentic Data Analysis —
- Retrieved but Not Delivered: Multimodal Memory Delivery for Long-Term Agents —
- Artificial intelligences and human scientists exhibit complementary strengths in theory building —
- Interpretable Physics Informed WiFi Indoor Localization: Learning an Effective Access Point Geometry and Using It to Prune —
- DepthBench: Measuring How Residual Connections Enable More Computational Depth —
- Can Open-Weight Large Language Models (LLMs) Simulate Human Survey Populations? A Cross-Instrument Calibration Study —
- Intuition vectors —
- "You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely Accused —
- What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation —
- Refinement Symmetry in Multimodal Transformers —
- Learning from a Thoughtful Teacher: Adaptive On-Policy Self-Distillation for Mathematical Reasoning —
- InterTab: Interleaved Visual-Structure Alignment for Multi-Modal Table Reasoning —
- Contract Memory Compiler: Resolve, Then Traverse —
- Timestep Weighting: A Hidden Key to Effective ELBO-Based Flow-Matching RL —
- MixBench-TS: A Multivariate Time Series Forecasting Benchmark Where Channel Mixing Pays Off —
- CoWindow Attention: Full Causal Coverage Is a Collective Property —
- Expected Reasoning-Step Return Unifies On-Policy Learning from Rewards and Teachers —
- CAIRN: Dynamic Fact-Intent DAGs for Multi-Agent Exploration —
- MM-OPD: Towards One More Bottleneck Between Perception and Reasoning —
- Dude, Where's My State? Execution Information Requirements for Stateful Agents —
- The GUI Is Not the State: Diagnosing State Aliasing in GUI World Models —
- When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds —
- CLAIRE: A Schema-Grounded Hybrid Workflow for Healthcare Administrative Form Completion —
- Learning response-aware patient dynamics for respiratory support —
- Getting Motif-ated: Controllable AI Compositions from Injected Motif Prompts —
- Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs —
- Rank Collapse Is Recoverable, Growing Q Is Not: Out-of-Sample Early Warning for Value Divergence in High-UTD Soft Actor-Critic —
- Beyond Accuracy: Counterfactual Fragility and Demographic Bias in Clinical Evaluation of LLMs —
- Nutri-ATLAS: Embodied Agent for Tabulated Lookup and Assistance for Smarter nutrition —
- Overwhelmed by Choice: Studying LLM Decision Making at Scale —
- Logic Gate Networks and Lookup Table Networks as Lightweight Hardware Classifiers for Inter-patient ECG Arrhythmia Classification —
- Optimizing H-Graph Hybridization for Diffusion-Guided RRT —
- Refreshing Less, Selecting Better: Reusing Stale Gradient Features for Efficient Influence-Based Data Selection —
- StraTune: Adaptive Selection of Revision Operators for Self-Evolving LLM Skills —
- Can LLMs Predict the Future? A Brier Score Analysis of Prediction Markets —
- Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning —
- PolyTopoBench: A Benchmark for Complex Vector Polygon Generation from Remote Sensing Imagery —
- Logical subspace in LLMs —
- TwinS-GCN: Spectral conjugate for Spectral Graph Convolutional Networks —
- TRACE: Learning to Self-Calibrate Wireless Digital Twins from ISAC Measurements —
- Precision As You Need: Stochastic Computing Is a Dense Adaptive Quantizer —
- The Geometry of Logic: Stratification Induces Semantic Structure and Robust Reasoning —
- Diagnosing Sampled LLM Reasoning in Formal Geometry: Coverage, Realization, and Validity Evidence —
- Planner-as-Router: Joint Plan-Time Model Routing for Cost-Efficient Multi-Agent Workflows —
- SRE-Marathon: A Continuous, Change-Driven Benchmark for Autonomous Site Reliability Agents —
- Improving the Diversity of LLM Outputs without a Trade-off —
- ukasiewicz Neural Networks Extended: Residual Architectures and Crystallization Strategies for Interpretable Rule Extraction —
- SemReward-VL: Semantic Reward-Guided Video-Language Adaptation for Developmental Behavior Assessment —
- The Model Knows Another Way: Strategy Switching for Effective RLVR Exploration —
- LLM sequential decision making under uncertainty in biochemical domains —
- Algorithmic Harms Associated with Generative Model-Augmented Recommendation Systems —
- NutriVision: Ingredient-Conditioned Fusion and Prediction for Single-Image Food Nutrition Estimation —
- OneSign: Unifying Sign Language Understanding Tasks with One Model —
- Query, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language Navigation —
- FedSysID: A Federated Approach to Sample-Efficient System Identification —
- Sample-Efficient Optimization over Generative Priors via Coarse Learnability —
- CAT: Can Trust be Predicted with Context-Awareness in Dynamic Heterogeneous Networks? —
- AMO: Operator-level Adaptive Muon Orthogonalization —
- Driving Video Retrieval for Complex Queries with Structured Grounding —
- DroneShield-AI: A Multi-Modal Sensor Fusion Framework for Real-Time Autonomous Drone Threat Detection, Behavioral Intent Classification, and Swarm Intelligence in Contested Airspace —
- Revisiting the Volume Hypothesis —
- On the Convergence of Stochastic Low-Rank Adaptation —
- Multilingual Fine-Tuning via Localized Gradient Conflict Resolution —
- Introspecting Alignment Shifts Beyond Behaviors Implanted Through Fine-Tuning —
- What Can Latent World Models Know? Physical Information in Multimodal Predictive Representations —
- Beyond Selection: Token Parameterization for Extreme Visual Token Compression —
- DP- lambda CGD: Efficient Noise Correlation for Differentially Private Model Training —
- FIDAL: Diversity-Aware Federated Active Learning Under Real-World Distribution Shifts —
- EEGAgentBench: Benchmarking LLM Agents on Short- and Long-Horizon EEG Analysis —
- When Keywords Drop but Classifiers Hold: Soft Refusals under KV Cache Compression —
- NanoForecast v0.5: Competitive Time Series Forecasting Through Training Pipeline Optimization —
- Witeness Overlap: Directional Provenance Inside Open-Weight Model Families —
- Panoptic Scene Program Diffusion Transformer —
- Relational Compression: A Framework for Relational Fidelity in Constrained Representations —
- seq2cause: One Autoregressive Backbone, Four Causal Discovery Tasks in Event Sequences —
- Medium-Term Multi-Resolution Electric Load Forecasting using Economic Data and Foundation Model —
- Averaged Mirror Descent and Dual Gradient Methods: Convergent Algorithms for Entropic Gromov-Wasserstein Problem —
- TemporalGraphLLM: Temporal Graph Neural Networks with Large Language Models for Dynamic Text-Attributed Graphs —
- CAFE: Counterfactual Prediction via Fast Posterior Estimation —
- Kernel-Based Steering of CLIP with Vision-Language Model Preferences —
- "Where Can I Trust You?": Boundary-Aware Evaluation of Surrogate Fidelity —
- Arithmetic Simplicity in Stochastic Gradient Methods —
- CompassPlay: Rewarding the Proposer for Where It Moves the Solver —
- Refresh or Realize? Compute Allocation in Drifting Models —
- Overfitting of Spectral Gradient Descent: How Matrix Geometry shapes Generalization and Implicit Bias —
- Self-Reconstruction Dynamics for Autoencoder Reconstruction Refinement —
- Learning Through Game: Skewed Transfer of Tabular Knowledge to Strengthen Image Model —
- When Can Old Evaluations Certify a New Model? Label-Efficient Release Decisions under Evaluator Drift —
- SIMANF: Sample Free Learning of Unnormalized Distributions via Simulated Annealing in Normalizing Flows —
- Before Answering: Evidence Sufficiency under Size-Matched Memory Construction —
- Certification Frontiers for Gaussian LoRA: Independent Priors, Posterior Risk, and Prediction-Preserving Balancing —
- When the Merge Coefficient Stops Mattering: Proximity Regularized Merging for Continual LoRA Adaptation —
- AECSF: Adaptive Ensemble Conditional Score Filtering for High-Dimensional Nonlinear Data Assimilation —
- A Journey to the Edge of Stability —
- A Comparative Analysis of Attention versus State-Space Models for In-Context Learning —
- Continual Data Unlearning in Diffusion Models via Transition-based Regularization —
- Attribution Without a Second Pass: Inline Per-Sample Gradient Provenance at 1% Overhead —
- Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality —
- Stabilizing the Dynamic Low-Rank Training —
- What Should the Reflector See? An Empirical Study of Evidence in Reflective Prompt Optimization —
- Length-Independent State Tracking Under a Parallel Scan —
- Masking Frequent Tokens Sharpens Direct Preference Optimization —
- What Should Federated LoRA Share? FedSAIL via Input-aware Subspace Alignment —
- Prioritizing Repeated LLM Evaluation for Hidden Failure Discovery —
- Elastic Selective Spectral Hybrids for Train-Once, Export-Many Budgeted Inference —
- Self-Evolving Time-Series Forecasting Agents with Episodic Memory and Online Policy Learning —
- Extremely Fast and Compact Binary Graph Representations via Randomized Operator Sketching —
- Quantization-Aware Pre-Training with Constrained Empirical Weight Distribution —
- Does Transolver really need a Transformer? —
- JEPA Learns What the Mask Leaves Unrecoverable —
- Affine Geometry of Gaussian ReLU Networks via Conditional Kac-Rice Formulas —
- BiasReducer: Adaptive Bias Mitigation for Reward Models —
- SIFT: Enhancing Time Series Foundation Models via Semantic Invariance and Structural Fidelity Fine-Tuning —
- AnchorMixGAN: Anchor-Aligned Generative Semi-Supervision for DDoS Detection in Cloud-Integrated IoT Networks —
- Reuse or Relearn? A Spectral View of Earth Observation Foundation Models —
- CT-OPD: Counterfactual Trace On-Policy Distillation for Diffusion Vision-Language Models —
- OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing —
- Robust Bayesian Optimization with Q-Exponential Surrogates —
- HamiFormer: Dual-Expert Diffusion Fields with Affine Symplectic Maps —
- Understanding and Exploiting Anisotropy in Post-Training —
- FedHV: Low-Overhead Hypervolume Weighting for Federated Multi-Objective Optimization —
- Structuring Relations Among Learning Paradigms via Protocol--Objective--Resource Reductions —
- USAI-Quant: A Quantitative Reasoning Benchmark for Vision-Language Models in Built Environments —
- Retimed Bellman Flows: Escaping the Impossible Triangle of Velocity Bootstrapping —
- Over-the-Air Federated Learning in Heterogeneous Mobile Wireless Networks —
- UniCache: Task- and Type-Aware KV Cache Compression for Unified Multimodal Models —
- Predicting the Next State Is Not Enough: JEPA Representations for Lean Theorem Proving —
- Staying on the Attractor: Supervising Neural Surrogates of 3D Turbulence Where They Leave It —
- Is H&E Image-to-Spatial Transcriptomics Simpler Than It Looks? —
- SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation —
- Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset —
- Beyond Math and Code: Lightweight Corpus-Grounded Process Rewards for Factual Question Answering —
- ResearchMath-14K: Scaling Research-Level Mathematics via Agents —
- Can Agents Infer Environment from Interaction? Evidence from Agentic Automata Learning —
- Toward Embedding-Based Psychometrics: Structural Modeling of Assessment-Item Semantics With Contextual Scores —
- Agents Can Use Base Models to Evade AI Detection —
- Omni-IO Skills: Harnessing Your Agent Omni-Native —
- Quantization Thresholds Replicate, Failure Modes Do Not: A Three-Model Study of Agentic Tool Use in Polish from 8-bit to 2-bit —
- Generalization and Memorization along the Learning Trajectory of Neural Language Models: A Geometric Account of Categorization —
- LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models —
- FA-Bench: A Benchmark for Phone- and Word-Level Timestamp Accuracy in Forced Alignment and ASR on Clean and Noisy Speech —
- ChildSafeAds Shared Task 2026: Commercial Content in Child-Facing YouTube Videos —
- An Evaluation of AI-Supported Evidence-Based Learning for Public Speaking Skill Development —
- LLM-Guided Ontology-Driven Knowledge Graph Construction from Unstructured Text —
- IndicFDB: Benchmarking Full-Duplex Voice Agents across Indian Languages —
- ScopeIF: Improving Scope-Aware Precise Instruction-Following in Large Language Models via Graded Reward Modeling —
- Streamlined Reflective Evolution for Task-Adaptive Self-Refinement Pipelines —
- How to Reduce Whisper Hallucination —
- KinyaMed: Seeds, Not Rows -- What a Corpus Requirement Written in the Wrong Unit Fails to Constrain —
- OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets? —
- A model of rational interlocutors: Unification of comprehension and production —
- Theory of Scene: Breaking the Symmetry Trap in Multi-Agent LLM Coordination —
- Linger and Lose: Knowledge Collapse in Low-Bit Language Models —
- Focusing Condition: Inference-Time Self-Contrastive Steering Elicits Better Conditional Text Embeddings in LLMs —
- Attribution Gaps in Zero-Training LLM+OVOD Pipelines: A Fine-Grained Analysis of the CAAP--SNAP Discrepancy —
- Self-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual Reports —
- ARSM: Auto-Regressive State Machine for Agentic Reasoning Compression —
- AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research —
- Classifying Dominant Temporal Orientation without Pretrained Text Embeddings: A Novel Morphosyntactic Inventory Vector Approach —
- Pinned and Still Unstable: Within-Judge Verdict Variance and the Noise Floor of LLM-as-Judge Leaderboards —
- ECG-Scroll: A Long-Horizon, Streaming Benchmark and Agent Environment for Interpretation of Ambulatory Electrocardiograms —
- Reading Too Much into Context: Passive Exposure Can Steer LLM Decisions —
- Identifying Temporal Features within Transcoders for Time Sensitive Factual Recall —
- SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought —
- Scoring the Wrong Question: Readout Failures in Constrained-Option Evaluation —
- Where Do Test-Time Scaling and Training Fall Short in Individual Stance Prediction? —
- CAME: Company-Aware Evidence-Memory Experts for Interpretable Quarter-Ahead Revenue Forecasting —
- Knowing Is Not Choosing: What Explicit Verification Adds Beyond Generative Preference —
- Save Your Saturated Data: Learning Beyond Reward Saturation in Group-Based RL —
- BaatCheet: A Multilingual Corpus for Dialogue Translation in Indian Languages —
- AG-CoT: Verified Algorithmic Traces for LLM Program Synthesis on Clifford Circuits —
- The Text Beside the Image: Detection, Utility and Leakage for Trustworthy Multimodal Medical Data and Beyond —
- BERT4DTI: BERT-based Model for Predicting Drug-Protein Interactions —
- Beyond Calibration: Do a Typed-Decision Model's Probabilities Obey the Probability Axioms? —
- Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning —
- Turning Speech Language Models into Multilingual Listeners —
- CalibHyper: Chance-Corrected Relational Hypergraphs for Few-Shot Molecular Property Prediction —
- NLPG: Natural-Language Policy Gradients for Self-Evolving Language Agents —
- Language Discrimination Improves Linguistic Learning in Multilingual Speech Models —
- CHI: A Composite Hallucination Index Unifying Entity, Relation, and Quantity Dimensions for Summarization Evaluation —
- MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses —
- Dense Is Not Enough: Hierarchical Supervision Allocation for Long-Horizon On-Policy Distillation —
- Decoupling Token Roles in Autoregressive Pretraining —
- Preserving Morphemes: Morphology-Guided Pre-Tokenization for Nepali —
- From Position Risks to Block Survival: Faster Generation for Diffusion Language Models —
- MIC: Explaining Image-Claim Inconsistencies in AI-Generated Multimodal Misinformation —
- Raven: The Harness of Harnesses for Composable Agentic Intelligence —
- SMAT: Simple and Efficient Merge-Aware Training —
- Grounding Memory Summarization in Utility Intent —
- DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation —
- Context Spanning: A Communication Framework for Full-Duplex Speech Models and External LLM Backends —
- GSM: Efficient Language Modeling with Shared Global State —
- LLMs Trust Their Own: Identity-Dependent Conformity in Multi-Agent Systems —
- TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs —
- E-CONAN (Entailment, CONtradition And Neutral) Diagnostics Dataset Investigating Linguistic Phenomena in Arabic Natural Language Understanding —
- ManiEdit: Sequential Unstructured Knowledge Editing for Language Models from a Manifold Perspective —
- ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments —
- Probe to Act: Elevating Browser-Use Agent via Active Visual Probing —
- Learning to Learn from Context: Synthetic Training from Perturbed Public Documents —
- Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails —
- One Model Is Not a Crowd: Multi-LLM and Aspect-Conditioned Diverse Comment Generation —
- BOReFT: Manifold Steering of Language Models for Black-box Optimization —
- Shared Experience, Separate Learning: Companion Confidence Calibration for LLMs —
- When Do Agents Help? Embedding, LLM and Agentic Alignment of Classical Texts and Their Translations —
- Yor` u b' a in Unicode: An Overview of a Problem —
- Auditing Agent Actions through Query-Conditioned Attribution —
- Reset Is Not Recovery: Evaluating Recoverability from False Conversational Context via Sycophancy Hysteresis —
- From Granular Revision Operations to Meaningful Revision Units: Evaluating LLMs for Revision Boundary Detection —
- LA-CPD: Local-Evidence-Aware Change-Point Detection for Human-LLM Authorship Segmentation —
- Causally Debiased Latent Action Model for Embodied Action-Conditioned World Models —
- Language-Augmented Video Action Anticipation: Design Fundamentals, Benchmarks, and Open Challenges —
- GeoCo-SAVi: Geometry-Consistent Slot Attention for Explicitly Editable Object Representations —
- RoboTok: A Scalable Data Engine for Internet Demonstration Video Retrieval and Dexterous Manipulation Learning —
- DiCoR: Decoupled Referent Disambiguation and Contour Recalibration for Efficient Referring Remote Sensing Image Segmentation —
- ForensicZoom: Adaptive Visual Inspection with Multimodal LLMs for Industrial-Grade Face Forgery Detection —
- One Evaluation, Any Operating Point: Hypernetwork-Amortized MeanFlow for 3D MRI Reconstruction —
- When Retrieval Hurts: Measuring and Explaining Retrieval-Induced Hallucination in Chest X-ray Report Generation —
- Measuring the evolution of camera distance across a century of film —
- Beyond Volume Overlap: Surface Matching for Topology-Aware Coronary Artery Segmentation —
- Federated 3D Gaussian Splatting for Large-Scale Scene Reconstruction at Wireless Edge —
- SelfCue: Making a 3D CT Report Generator Say What It Already Knows —
- Rate-Adaptive One-Step Diffusion Compression for AIGC Images —
- UNMATCH: Selective Unbalanced Token-Patch Matching for Forensic Image-Claim Verification —
- 3dgs-sc: a controlled static screen-content benchmark for 3d gaussian splatting —
- Binaural Audio-Visual Instance Segmentation —
- PEEL-DDPM: Physics-Enabled Evidential Learning for the Denoising Diffusion Probabilistic Model —
- Fusion Under Component Failure: Negative Results and Failure Modes in Ensemble AI-Generated Image Detection —
- Dimension-Specific Imbalance and an Adaptive Hybrid Label Strategy for Multi-Task Affective State Recognition in Classroom Video —
- Enhancing Visual Reasoning in Chest X-Ray Report Generation Using Reinforcement Learning —
- Gaussian Image Steganography via Parameter-Domain Keyed Embeddings —
- Facial classification Using Hybrid Quantum Machine Learning —
- PredRA: Fast Medical Image Translation by Deterministic Component Extraction and Controlled Stochastic Refinement —
- PruneForget: Joint Unlearning and Pruning of Vision Models —
- PQR3D: Progressive Query Refinement over Reference-Conditioned Temporal Windows for Multi-View 3D Object Detection —
- Skeletons in Flow: Graph Structured Flow Matching for Human Motion Prediction —
- Geometry-Preserving Blind Watermarking for Raw 3D Point Clouds —
- Scalable In-Domain Self-Supervised Foundation Model for Dense Representation Transfer in High-Resolution Plant Imaging —
- Two-Stage Multi-View Gait Recognition with a Re-Embedding Network —
- RoboSTAR: Next-Scale Autoregressive Sign Language Translation for Humanoid Robots —
- FSS-UBrain: Multi-region Few-Shot Brain Tumor MRI Segmentation —
- FoundDSR: A Generalizable Foundation Model with Guided 2D Gaussian Splatting for Depth Super-Resolution —
- QuacamFM: Quaternion-Constrained Flow Matching for Camera Pose Estimation —
- SetOPD: From Few Visual Exemplars to Multimodal Candidate Sets for Remote-Sensing Open-Prompt Detection —
- CFCH: Coarse-Fine Collaborative Hierarchical Learning for Anterior Segment Disease Analysis —
- Can Motion-Language Models Ground Structure? STRIDE for Evaluating the Evaluators —
- StegGNN: Learning Graphical Representation for Image Steganography —
- REMEDY: How Far Is Video Generation from Medical Education World Models? —
- RefCompose: Multi-Reference Image Generation via LoRA-Conditioned Diffusion —
- An End-to-End Latent-Rollout Approach for Pushing Few-Step ImageNet- 256 Generation to FID 1.11 without Fr'echet Losses —
- From Feed-Forward to Flow: Unifying Reconstruction and Generation Is Easier Than You Think —
- RIPE-MambaSpike: Resolution-Independent Spiking-State-Space Interfaces for Parameter-Efficient Event-Based Vision —
- Cross-Domain Few-Shot Writer Adaptation for Real-World Handwritten Mathematical Expression Recognition —
- RCVLA: 4D Radar-Grounded Semantic Reasoning and Trajectory Arbitration for Autonomous Driving —
- Progressive Risk Estimation for Accident Anticipation —
- AnesTRACE: Benchmarking Intraoperative Anesthesia from Multimodal Perception to Multi-step Decision-Making —
- Region-Local Copula Evidence Fusion for Heterogeneous Remote Sensing Change Detection —
- Unlocking Geodesic Gromov-Wasserstein Distances for 3D Modeling —
- LoCoVSR: Local Context Diffusion Posterior Sampling for Video Super-Resolution —
- SV2V-RSim: A Comprehensive Benchmark for Self-Selective V2V Cooperative Perception with Near-Realistic Data —
- Resource-Aware Parameter-Efficient Model Adaptation for Onboard High-Dimensional Data —
- MAD-Guard: Controlled Study of Autoregressive Generation versus Direct Decision Interfaces for Closed Multimodal Forensic Tasks —
- When Noise Meets Long-Tail: Feature-Threshold Dual Calibration for Robust Pseudo-Labeling —
- Learning Multimodal Embeddings with Evidence-Aligned Readout —
- Parameter-Efficient 3D Segmentation of Liver and Liver tumors: Depthwise factorization Scales Better Than Dense Convolution with Spatial Dimensionality —
- Distributed Hydrological Modeling in the Feature Space —
- Improving Video Sparse Attention with Fine-grained Router and Sparse Rebasing —
- Rethinking the Fully Hyperbolic Vision Transformer in Polar Coordinates —
- Certified Interface Aliases: Exact Collisions in Vision-Language Preprocessing, and When They Exist —
- Residual Diffusion Implicit Models —
- Train Together or Merge Later? Unifying VLA Experts via a Shared Action Interface —
- Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models —
- QSCP: Beyond Class-Name Prompts for Query-Guided Semantic Change Parsing —
- dKFD: Phase-Structured Evidence Allocation for Fixed-Budget Localized Event Understanding —
- FloodDiffusion 2: Efficient and Path Controllable Streaming Motion Generation —
- ABO-Med: Accelerated Bilevel Optimization for Few-Shot Medical Image Classification —
- Can Protein-Derived Knowledge Improve Pathology Foundation Models? —
- ReAL: Accelerating Flow Matching through Segment Advancement with Shared Lookahead —
- Scope-WM: Scoped Computation for Efficient Visual World Models —
- RepFlow: Reciprocal Supervision Improves Generation and Representation in Flow Models —
- WorldAgent: Verification-Guided Agentic Physical World Construction —
- Relevance Does Not Imply Applicability: Experience Activation for Personal GUI Agents —
- LoopTrack: A Simple Baseline for Parameter-Efficient Transformer Tracking —
- Informative Viewpoint Selection for Episodic-Memory Embodied Question Answering using Omnidirectional Images —
- VehDyn: A Driving World Model Benchmark for Vehicle Dynamics —
- Focus and Supplement: Dual-Enhanced Vision Transformer for Multi-Class Anomaly Classification —
- Groupwise Selective State-Space Filtering for Accurate and Streaming Action Boundary Detection —
- TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining —
- PulseQuant: Propagation-Guided Subspace Correction for 4-Bit Video Diffusion Transformers —
- TNF based Spectral Embedding for Effective Application of Supervised Machine Learning Techniques in Automobile Insurance Fraud Detection —
- TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models —
- Printability-Constrained Adversarial Decals for Near-Nadir Aerial Perception: Measured Ink Gamuts, Nested Realism Constraints, and a Physical-World Bound —
- A Light Bilevel Refinement Aligns Self-Supervised Representations for Stronger Task-Specific Learning —
- TC-ADA: One-Shot Active Domain Adaptation for Semantic Segmentation —
- Elucidating the Design Space of Regression-based Diffusion Reinforcement Learning —
- SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning —
- Anguinus Sculpturae: Compositional Synthesis of Peak-Enhancement Breast DCE-MRI Scans —
- ViCoR: Reliable Molecular Structure Extraction via Spatially Aligned Verification and Executable Revision —
- LoopLUT: 3D Lookup Tables with Progressive Region Refinement for Real-Time 4K Image Enhancement —
- IVT-Guard: All-in-One Reasoning Model for AI-Generated Content Detection —
- In-Token Learning for High-Fidelity Image Restoration via Diffusion Transformers —
- Offline Policy Evaluation via Mixed Bellman Residuals and Adaptive Critic Representations —
- Low-Rank Single-Index Bandits with Unknown Links: From Matrices to Tensors —
- Simulation-Free Learning of GP-SDEs from Irregular Observations —
- M3-Score: Fidelity, Memorization and Coverage as Separate Axes for Evaluating Generative Radiology Image Models —
- ENet-GP: Unified Document Image Restoration —
- Calibrated Derivative-Process Sensitivity for Gaussian-Process Variable Selection —
- Domain-Adapted Diffusion Models for Conditional Independence Testing —
- Minimax-Optimality of Posterior Sampling for Reinforcement Learning —
- Schr"odinger--F"ollmer Actor--Critic: Diffusion Policy Improvement with Finite-Sample Analysis —
- From Dual Tracking to Clipping: Provably Faster Distributionally Robust Multi-Objective Optimization —
- A Statistical Perspective on Knowledge Distillation: Foundations, Classical Methods, and Large Language Model Extensions —
- Sparsity by Default: The Theory and Practice of ARD in Gaussian Process Regression for Variable Selection —
- SLP-ProbHard: Probabilistic Hard-Constrained Learning via Structural Latent Parameterization —
- The Statistical Benefits of Multiple Responses for Learning from Demonstrations —
- Correct then Forecast: Observer State-Space Models for Time Series Forecasting —
- Sharp training-conditional coverage for conformal prediction under covariate shift —
- Benign Overfitting for General Norms and Distributions —
- Just-In-Time Scene Graph Growth: Combating Perceptual Saturation in Long-Horizon Robotics —
- Revitalizing Medical Time Series with Vision-Informed Retrieval: A Vision-Language Perspective —
- Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates —
- Don't stop me now: How Validation Criteria Affect Checkpoint Selection and Early Stopping —
- What Must a World Model Distinguish for Planning? —
- Constrained Flow Policy Updates: A Generalized Schr"odinger Bridge View —
- When Does Backpropagating Through Policy Memory Matter? Physical Credit, Optimizer Updates, and Observability —
- Hierarchical Frequency-Domain Compression of Implicit Geometric Representations for Large-Scale Point Clouds —
- PanoFuse: Panorama-Enhanced Vision-Language-Action Learning with Decoupled Semantic-Geometric Routing —
- Probabilistic Object Detection with Conformal Prediction —
- Emergent One-Third Scaling Law as Attention Tries to Concentrate —
- TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents —
- Autonomous Research Project Management as an Agent Skill: A Case Study in Exact Spectral Spatial Regression —
- What does FFN compression change downstream? Same-state causal restoration in diffusion language models —
- Decide, Don't Generate: Competitive Dimensional ABSA with Jev's Typed Decisions —
- Fysiverse-3D-SimReady Technical Report: Agentic Physical Simulation for Pragmatic 3D World Reconstruction —
- PanOVOcc: Panoramic Embodied Open-Vocabulary Occupancy Mapping with Long-term Spatial Voxel Memory —
- Trustworthy synthetic visual media: Evidence across the media lifecycle —
- Universal Pose Pretraining for Generalizable Vision-Language-Action Policies —
- Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind —
- ChestPheNoT: Deployable, Auditable Label-Status-Evidence Extraction from Radiology Reports —
- When Does Domain Adaptation Help on Physical Vibration Sensors? A Held-Out-Bearing Study of Neural-Operator and Convolutional Models —
- Measure Learning at Steady State: A BIRD-SQL Formula 1 Case Study —
- A Comparative Transfer-Learning Study of CNN Backbones for Partial Face Recognition on the SoF Dataset —
- Unsupervised spiking feature learning for event-based pedestrian crossing detection: approaching supervised accuracy without labelled training data —
- Query-aligned video frame selection for long video understanding —
- Cross-Dataset Transfer and Unknown-Class Detection in Imbalanced SAR Ship Classification —
- MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning —
- LukeNet: A lightweight CNN integrated with an XAI model for Smart acute lymphoblastic leukemia detection and management —
- Cross-Material Support Transfer for Core-Loss Prediction Under Waveform Covariate Shift —
- Product-Aware Deterministic Rounding for Quantized Matrix Multiplication —
- Enhancing generalization in endwall film cooling prediction: Incorporating the superposition principle into transformer-based neural operators —
- Disentangle and Drop: Robust Universal Removal of Image Watermarks via Reconstructive Grayscale Residual Decomposition —
- Help Me Help You: The Aggregate Value of Source and Target Data in Transfer Learning —
- Video Captioning in Low-Light Conditions through Efficient Uncertainty-Aware Caption Correction —
- Devanagari Handwritten Character Recognition Using TrOCR: A Transformer-Based Model with Real-Time Web Deployment —
- Cross-Dataset Generalization of Bangladeshi Rice Leaf Disease Classifiers: Benchmark, Diagnosis, and Mitigation —
- CoRe-Gen: Robust Spectrum-to-Structure Generation under Imperfect Fingerprint Conditions —
- PRISM: Recovering Instruction Sets from Language Model Activations —
- Imitating Radiological Scrolling: A Global-Local Attention Model for 3D Chest CT Volumes Multi-Label Anomaly Classification —
- OmniLayout: A Schematic-Coupled Multimodal Benchmark for Constraint-Aware Geometric Reasoning in PCB Layout —
- MolLangData: A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method —
- STEMTOX: From Collaborative Tags to Fine-Grained Toxic Meme Detection via Entropy-Guided Multi-Task Learning —
- Semantic Modality Compensation for Unsupervised Visible-Infrared Person Re-identification under Unpaired Settings —
- Smoothing the Semantic Landscape: Generalizable AI-Generated Image Detection via Text-Induced Flatness —
- SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents —
- ESTHER: Egocentric Stereo Hand Estimation and Reconstruction in the Wild —
- Product-Stability: Provable Convergence for Gradient Descent on the Edge of Stability —
- Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning —
- DCFold: Efficient Protein Structure Generation with Single Forward Pass —
- AI Annotation Orchestration: Evaluating LLM verifiers to Improve the Quality of LLM Annotations in Learning Analytics —
- Native Association: Confidence-Aware Human Perception in the Wild with a Foundation VLM —
- Apparent Compression, Real Stability: The Intrinsic Dimension of Learning a Quantum Wavefunction —
- CriPO: Enhancing Rubric-based RL via Self-Distillation —
- CSI-Agent: LLM-Assisted Few-Shot Adaptation for Cross-Domain Wi-Fi CSI Sensing —
- Fill2SR: Repurposing Inpainting Diffusion Transformers for Real-World Super-Resolution —
- DynActiveGS: Active Gaussian Splatting for Dynamic Scene Reconstruction —
- Constrained Edit Fields for Training-Free Flow Editing —
- FraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund Evidence —
- Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization —
- Ego-Forge: Text and Geometric-Attention Free Exo-to-Egocentric Video Generation —
- Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts —
- Reasoning Concentrates Errors, and Self-Consistency Never Notices —
- StrucTab: A Structured Optimization Framework for Table Parsing —
- CFLoRA: Federated Fine-tuning of LLMs with Complementary Factors for Error-free Aggregation —
- Does Learning to Predict the World Help Agents Act? Auditing World-Model Post-Training —
- Frequency-Domain AI-Generated Image Detection: Exploring Decoder and Channel Attention for Feature Refinement —
- OPERA: A Unified Omnimodal Progressive Spatio-Temporal Reasoning Agent for Referring Video Segmentation —
- What If TSF: Reframing Time Series Forecasting as Scenario-Guided Multimodal Forecasting —
- Bounding Retraining Equivalence and the Deletion Floor in Materials Machine Unlearning —
- Shared Doubt: Zero-Shot Cross-Lingual Confidence Estimation for Language Models —
- Revisiting Diffusion Fine-Tuning for Unsupervised Domain Adaptation —
- DBCF: Dual-Branch Complementary Fusion of Foundation Models for Generalized Deepfake Detection —
- Latent-Action-Guided Video-Language Feature Learning for Surgical Instrument-Tissue Interaction Recognition —
- Demystifying the Unreasonable Effectiveness of Greedy Alignment Methods —
- From Scene Graphs to Answers: Selective Neuro-Symbolic Reasoning for Autonomous Driving —
- CellScientist: From Execution Feedback to Auditable Model-Revision Trajectories for Cellular Perturbation Prediction —
- Beyond Solo and Consistency: Vindicating Multi-Agent Debate via Conditional Progressive Pruning —
- Understanding Generalization Requires Universal Induction —
- A Novel Unified Approach to Deepfake Detection —
- Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment —
- Solving Every Step Is Not Enough: Milestone Oracles Reveal a Composition Gap in LLM Math Reasoning —
- KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding —
- MS-GLA: Multi-Scale Gated Linear Attention for Addressing Representational Bottlenecks via Multi-Temporal Resolution —
- Remember by Asking: Retrieval-Induced Memory Evolution for LLM Agents —
- Summarize Before Grounding: Query-Guided Chunk Condensation for Long-Video Temporal Grounding —
- Compressing Value Predictions for Learning-Augmented Metrical Task Systems —
- Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence —
- Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates —
- mu-bench: A Multilingual Utterance Transcription Benchmark —
- Does Vision-Language Pretraining Granularity Matter? A Controlled Evaluation of Vision-Language Objectives Across Chest X-Ray Interpretation Tasks —
- A Benchmark for LLM's Understanding of Middle School and High School Science Topics —
- Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding —
- What Visual Generators Need from Teachers: Rethinking Representation Alignment —
- The causal relation between off-street parking and electric vehicle adoption in Scotland —
- Deep Reinforcement Learning for Equity Trading: Benchmarking Actor-Critic Methods with Forward Retraining —
- Fast and Precise Learned Charged-Particle Trajectory Regression at the Large Hadron Collider —
- GERIS: A Game-Theoretic Framework for Filtering Instance-Dependent Label Noise in License Plate Data Augmentation —
- HGPTrans: Hierarchical Graph-Pooling Transolver for Automotive Aerodynamic Drag Coefficient Prediction —
- Continuous-Time Trajectory Generation from Discrete Observations with Stochasticity —
- REALM: Regime-Switching, Explainable, and Activation-Induced Linear Models —
- Mandela-Bench: Multimodal Models Remember Canonical Images Instead of Seeing Them —
- Depth Any Seen: Which Surfaces and How Far? —
- RemTraceNet: Few-Shot Forensic Detection of Invisible Watermark Attacks —
- Grounding Vision-Language Models in Driving Semantics: A Multi-Dataset Predicate Framework for Explainable Reasoning —
- What Drives Dialectal Jailbreaks? An Ablation of Surface Form, Cultural Framing, and Strategy Banks —
- Beyond the Graph: An Adaptive Meta-Learner Fuses Explainability, Weather, and Dynamics for Robust Bus ETA Prediction —
- Learning Steadily: Accumulating Relative Point Margin Scores for Face Image Quality Assessment —
- SPINET: Sheaf Protein Inverse Folding Network —
- Words & Weights: Streamlining Multi-Turn Interactions via Co-Adaptation —
- Seeing the Heat: Synthesizing High-Resolution Wood Thermal Responses from Optical Imagery —
- Learning What to Evaluate: Correlation-Aware Decoupling for Multiobjective Bayesian Optimization —
- SCLATE: a Substrate for Continual-Learning Agent Training and Evaluation —
- Learning When to Recur: Token-Adaptive Recursion for Imbalanced Ophthalmic Domain Incremental Learning —
- PolyStepOR: Learning to Decide Without Optimal Decisions —
- Neural Dynamics as the Composition of Quantized Units —
- CRF Loss is How Networks Should Learn Boundaries in Weakly Supervised Segmentation —
- Perturb-and-Solve: Efficient Learned-Operator Conditioning for Latent Diffusion Inverse Problems —
- Structure-First Point Cloud Learning with Mapper Region Graphs —
- Frontier Learning: Training LLM Reasoners at the Edge of Capability —
- Geometric Encoding for Spatial Reasoning in Vision-Language Models —
- Uncertainty-Aware Selection of Online Algorithms with Simulator Ensembles —
- Predicting the Financial Impact of Supply Chain Risk for Major AI-Related Semiconductor Firms: A Heterogeneous Graph Patch Transformer Approach —
- InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision —
- Inspector: Conversational and Lightweight Analyzer of Analog Circuit Layouts Using LLM and CNNs —
- What Does Model Growth Really Add? Functional Capacity Beyond Parameter Count —
- Anytime-Valid LLM Leaderboards via Benchmark-weighted and Block-Factorized e-Processes —
- Despite Instructions: Frontier Agents Improvise Covert Channels at Test Time —
- Vestrum: Improving Agent Harnesses by Adapting Their Verification, Structure and Memory —
- Test-Time Generalized Category Discovery —
- LLMs are not stochastic parrots: Evidence for meaning-mediated abstraction from conlang-like tasks —
- Quantifying Behavioral Tails in Black-Box Language Models —
- Same Probe, Different Numbers: Are Activation Probes Robust to Inference-Time Numerical Non-Determinism? —
- DAAF: From Failure Localization to Editable System Assets in LLM Agents —
- When Consent Outlives Context: Residual Authority Replay in Long-Lived Agents —
- Closing the Cross-Dialect Gap: Query Plans as a Portable Interface in Text-to-SQL —
- DiffPTS: Rethinking Diffusion ELBO for Probabilistic Time Series Forecasting —
- Domain Generalization under Sampling Pattern Shifts in Irregular Time Series —
- Prediction Limits and Koopman Closure of Geometry-Induced Soft State Abstractions —
- Not Too Hard, Not Too Easy: Learning from Intermediate States for LLM Structured Reasoning —
- Rate-Distortion Adaptive Primitive Selection for Omnidirectional Gaussian Splatting —
- Classifier-pruned Bayesian optimization for particle accelerator tuning: Exploring temporally structured manifold of 6D beam phase space —
- MAC-Net: A Multi-Task Deep Learning Framework for Modeling Cognitive Function From Task-Based fMRI —
- NC-Bench: An LLM Benchmark for Evaluating Conversational Competence —
- On the Capability and Limitation of Hard Prompt —
- Echoes of Deeds: Moral History Can Shape and Steer LLM Behavioral Choices —
- OmniFysics-Captioner Technical Report: Grounding Omni-Modal Understanding in the Physical World for Better Captioning —
- Safe Score Matching: Diffusion Policies with Hamilton-Jacobi Reachability for Online Safe Reinforcement Learning —
- Agent Safety From Within: Detecting Harmful Trajectories from LLM Internal States —
- Explaining Textual Entailment with Lexical Entailments: Using LLMs to Supply Lexical Relations for Formal Proofs —
- What Paired Evaluations Reveal under Visual Perturbations —
- CoCurve: Cross-Module Co-Pruning Curvature for Structured LLM Pruning —
- Audit-First VAPO: Risk-Certified Selective Updates under Imperfect Verification —
- EMIR squared: Evolution-Aware Memory with Intent-Guided Multi-Round Retrieval —
- Clarify the User or Verify the World? Uncertainty Routing for Proactive Agents —
- Relic: From Multi-Agent Collaboration to Persistent Organizational Capability —
- VaME: Exploring Variational Latent Reasoning for Multimodal Embeddings —
- Uncertainty Quantification of Next Generation Reservoir Computing with Applications to Memory-Driven Dynamical Systems —
- GAUGE: Group-Wise View-Inconsistency Rectification for Feed-Forward 4D Tracking —
- Porimon: An LLM-Based Pok'emon Battle Agent Enhanced by Long/Short-Term Knowledge Augmented Generation —
- EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks —
- FARE: Deep Reinforcement Learning For Fair Exposure Constrained Uncertainty Aware Financial Content Personalization —
- Stochasticity Is Not the Hard Part: Reduction and Complexity in Instructional Sequencing over Prerequisite DAGs —
- One Readout, Many Repairs: Diffusion-Guided Hierarchical Search for Tool-Agent Repair —
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting —
- EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models —
- Understanding Quantization of Optimizer States in LLM Pre-training: Dynamics of State Staleness and Effectiveness of State Resets —
- lambda-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning —
- EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding —
- AeroCopilotBench: Safety-Gated Evaluation of LLM Agents on Aircraft Emergency Procedures in an Executable Cockpit —
- SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera —
- Harnessing Coupled Stream Completion For Human-Object Interaction Modeling —
- From Scores to Samples: Elastic Forcing for Autoregressive Video Generation —
- GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space —
- Tandem Reinforcement Learning with Verifiable Rewards —
- Byzantine-Robust Federated RAG via Aligned Calibration and Fixed-Membership Conformal Prediction —
- RefineRL: Advancing Competitive Programming with Self-Refinement Reinforcement Learning —
- The Weakest Link Tells It All: Outcome-Supervised Process Reward Modeling via Learnable Credit Assignment —
- Design and Implementation of Agentic Orchestrations and Orchestration of Agents —
- Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need —
- Progressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMs —
- Prompt-Anchored Residual Adaptation for Biomedical Vision-Language Models —
- Emergence of the Primacy Effect in Structured State-Space Models —
- Agentic Video Understanding: A Survey —
- Can't Find Waldo: Evaluating VLMs' Sensitivity to Image Resolution and Detail Level —
- Integrated Deep Learning Framework Designed on Hybrid Optimization Strategies for Automated Health Detection and Analysis in Silkworms —
- When to Evict, Not What to Keep: Draft-Guided Eviction for Training-Free KV-Cache Compression —
- 3-D Emissions Mapping and Social Cost Estimation for US Domestic Aviation at West Coast Hubs —
- MDL-Calibrated Significance-Gain Pair Encoding: Replication-Aware Automatic Stopping for Subword Tokenization —
- Don't Repeat Yourself: Self-Supervised Fine-Tuning for Coverage —
- Verification of PETSc with CIVL using LLM-generated ACSL contracts and deterministic driver generation —
- Modernising the Compressed-Domain Video Captioner: A Controlled Study of SigLIP2 and GPT-2 Substitutions —
- Do Emotion Concepts Generalize Across Sources, Modalities, and Architectures in Vision-Language Models? —
- From Pixel Generation to Topological Inference: Structural Dual Super-Resolution for Trustworthy Cross-Physical-Domain Trabecular Morphology Learning —
- When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context —
- D squared-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation —
- Privacy-Preserving Full-Body Meshing from mmWave Radar via Mesh Foundation Model Supervision —
- A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees —
- CoDrive: Cross-Vehicle World-Consistent Video Generation with Precise Trajectory Control for Cooperative Driving —
- Projective Normal Fields: A Convex Optimization Method for Constructing Smooth UDFs —
- TaoTex: Boosting Texture Detail Fidelity for Native 3D Material Generation —
- Detection of Adversarial Attacks on Super-Resolvers Using Spectral Features —
- Verifying the Linear Representation Hypothesis: How Interpretable Are Vision SAEs? —
- Physics-Guided Spectral Distillation for Underwater Image Enhancement on Resource-Constrained Devices —
- Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning —
- DrawingsDreamer: A Unified Multi-View Engineering Drawings Generation Model —
- eval-unlearn: Benchmarking unlearning in Text-to-Image Diffusion Models —
- Evaluating Hierarchy-Aware Deep Learning for the Recognition of Tironian Notes —
- Towards Generalizable 3D Anomaly Detection via Relational Inconsistency Modeling —
- OPIS: An Input-Grounded Benchmark for Multi-Object Memory in Video World Models —
- Mixed-Prior Decision Risk for Open-Set Recognition —
- RefineDrive: Reliable Failure-Guided Learning for Vision-Language-Action Driving —
- LEGAU: Learning Semantic Gaussian Priors for Scalable Category-level Pose Estimation —
- Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge —
- Style-Driven Data Synthesis and Degradation-Aware Enhancement for Ultrasound Image Restoration —
- G cubed-LoRA: Organizing Reward-Weighted Video Data with Gradient-Guided Grouped LoRA —
- Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut —
- VideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene Reconstruction —
- DiMoP: Diffusion-Driven Motion Representation Learning With Frame-Level Pseudo-Classification for Skeleton-Based Action Recognition —
- Generative Uncertainty as a Self-supervised Signal for Semantic Similarity Learning —
- PIVOT: Pivot-Aware On Policy Self Distillation for Multi-Turn VLM Agents —
- TMCS: Tool-Grounded Multi-Agent Reasoning for Compositional Chemical Problem Solving —
- RoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model Rollouts —
- Rethinking Visual Token Compression for Video Large Language Models: A Simple Yet Strong Baseline —
- AHMAD: Adaptive Hybrid Multi-task Vision Learning with Assisted Distillation for Keypoint Detection —
- Spectral Super-Resolution using Spatial-Spectral Residual Operator Networks —
- BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion —
- EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold —
- On-Policy Self-Distillation for Multi-Turn Image Editing —
- Many Eyes, One World: Feed-Forward 3D Reconstruction from Mixed Cameras —
- RT-Super: Learning Tumor Segmentation from Longitudinal Images and Reports —
- Revisiting Risky Tackle Detection with Vision Transformers —
- WorldPlay2: Extending Real-Time Interactive World Models in Control and Horizon —
- ReSS: Residual-Restoring Sparse Attention for 3D Vision Transformers —
- Impact of Patient Orientation in Single- and Multi-View Camera Environments for AI-based Rehabilitation Monitoring —
- Copy the Same, Distill the Difference: Initializing Linear Vision Transformers —
- InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video —
- Conformal Prediction and Conditional Coverage for Tabular Foundation Models —
- Universal Approximation of Measure-to-Measure Operators by Pushforwards —
- A Hierarchy of Entropy-Shapley Games for Multivariate Predictive Uncertainty —
- Simplex Diffusion Models —
- Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks —
- Universality and Generalization of Causal Transformers Across Context Lengths —
- Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons —
- Information-Theoretic Analysis of Next-Token Prediction under Markovian Data —
- Reliability-Gated Fusion of Consumer Head and Foot IMUs for Lower-Body 3D Pose —
- K-OPSD: Verifiable On-Policy Self-Distillation for Post-Training Vision-Language Models on AEC Drawings —
- AGILE-GS: Anchor-Guided Fast Next-Best-View Selection for Active 3D Gaussian Splatting —
- Role-Guided MOE for Encoder-Level Pathology Representation Learning in WSI Classification —
- TRAP: Understanding and Mitigating Privacy Memorization in Language Models —
- What Next-Event Accuracy Cannot See: Closed-Loop Evaluation of Emergency Department Trajectory Simulators —
- STAR: Adaptive Spatial-Temporal Normalization for Unified Microservice Incident Management —
- Does Joint-Embedding Predictive Architecture Pretraining Help Time Series Forecasting? —
- Toward AI-Assisted Poultry Coccidiosis Diagnosis: Evaluating Gemini and BiomedParse on Eimeria Microscopy Images —
- World Agent: Can Language Models Keep a World Running? —
- One-Step Is Optimal: Unconditional Rectified Flows are Noise2Noise Denoisers, and Multi-Step Integration Provably Hurts---A Benchmark and Task-Based Detectability Study on Low-Dose CT —
- Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity —
- X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets —
- DGF-Bench: A Benchmark for Simulating and Auditing Deception Against Multi-Agent Governance Boards —
- When Valid Tool Calls Change Meaning: Formation-Consistent Dispatch for LLM Agents —
- PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models —
- Only Ask What You Don't Know: Grounded Delta Planning for Efficient Multi-step RAG —
- PlurVA-LLM-2026 Shared Task Track-1: Pluralistic Value Alignment in LLMs via Multilingual Fine-Tuning and Threshold Calibration —
- When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents —
- Multi-Agent System Search via Active Substructure-aware Policy Optimization —
- Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation —
- Remote Sensing Sparse-View 3D Gaussian Splatting via Depth Image-Based Rendering —
- Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces —
- AwarenessBench: Assessing Cognitive Capabilities of Language Models —
- PlanGuard: A Guardrail for Multi-Step Plan Safety in Embodied Agents —
- Generative AI-Based Data Augmentation for Oral Lesion Classification: The PhotoMOCI Dataset and Benchmark —
- Mind the RefGAP: Correcting Reference Attention in Diffusion-Based Visual Editing —
- Learning to Reason with Persistent Object States for Video Instance Segmentation —
- SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale —
- Resolving State-Representation Mismatch: State-Space Visual Reasoning for Open-Loop VLA Planning —
- Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation —
- HM-ROUTER: Joint Model and Harness Routing for Agentic Systems —
- Dual Advantage Fields —
- TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment —
- What Does a Routing Oracle Measure Under Stochastic Decoding? Coupling, Scorer Choice, and Single-Commit Ceilings —
- The Model Says Walk: Measuring whether LLMs Condition on Hidden Constraints —
- FAST-Brain: A Flow-Aligned Spatio-Temporal Surrogate Brain Model —
- Selecting Diverse SFT Traces Improves Post-RL Generalization —
- Natural Image Autoencoder-Based fMRI Representations for Trait and State Prediction —
- LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning —
- On Device Agentic Operation Caches -- Classifier-Centric NL-to-Action Generation —
- Beyond the Training Horizon: Mechanisms and Limits of Length Generalization in Looped Transformers —
- Synthetic Thermal Image Generation for Real-Time Animal Detection Under Low-Visibility Conditions —
- Modular Discovery of General Game-Playing Algorithms with Large Language Models —
- CORTEX: A Verified Experience Layer for Generalist Agents —
- Flow-Matching-Based Protein Structure Tokenizer Made Efficient and Easy —
- Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring —
- Agentic Multi-Turn Reasoning: A Fairness Approach —
- ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents —
- CodeSkill: Latent Skill Abstraction for Long-Horizon Code Agents —
- D-JEPA: Design-Recoverable JEPA Representation with Swappable Physics Decoders —
- When Does the Concept of "Dog" Emerge in an Audio LLM? —
- What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation —
- APEX: An Extensible Model for Agent-Assisted Production Scheduling —
- Long-Horizon Analog Design Bench: Benchmarking Agents on Hours-Long Analog and Mixed-Signal Circuit Design Tasks —
- COEVO: Co-Evolving Context and Parameters for Recursive Self-Improvement —
- Multi-Dimensional Comparative Scale Construction for Efficient Personalized Subjective Judgment in High-Traffic Applications —
- QuPID: Quantum Parameter-Efficient Input-Dependent Retrieval Adaptation for Medical RAG —
- Structured Sparse Memory for Recurrent Reasoning —
- ZeroGAR: Benchmarking the Adversarial Robustness of Zero-Shot Graph Models —
- The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization —
- Unlocking Latent Personalization in LLMs —
- CoViST: Visual Token Compression via Composable States —
- AutoHGNN: Robust and Efficient Neural Architecture Search for Hypergraph Neural Networks —
- DISCERN: Can AI Agents Work Like Scientists and Guide Discovery? —
- Are Benchmarks Reliable? Toward Structural Diagnosis via Sample-Level Capability Boundaries —
- DrafTS: Time-Aware Decomposition with Residual Correction for Time Series Modeling —
- OpenFC: Learning Verification Policies towards Open-Search Fact Checking —
- Federated Multi-Modal Human Activity Recognition using Multi-Agent Reinforcement Learning —
- What Happens During Autonomous Deep Research After the User Steps Away? —
- RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse —
- Dynamic Kuramoto-Hodge Operators for PDEs on Complex Geometries and Topologies —
- TopoMamba: A Load-Support Relation-Guided Multi-Directional State-Space Model for Topology Optimization —
- CompoWorld: Compositional Environment Scaling for General Agents —
- Supervision Recovery for Time Series Anomaly Detection via Context-Anchored Pairing —
- EAT: Expert Account Tracker for Efficient MoE Inference —
- JustQuant: You Don't Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization —
- From Distributions to Stochastic Processes: Neural Approximation of Measure-Valued Maps —
- Scalable and Data-Driven Decision Support in the Maintenance, Repair, and Overhaul Process —
- Trajectory Unlearning on LLM-based Agents —
- MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception —
- Diffusion Reward Models —
- Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models? —
- Is your uncertainty map wrong, or is its target? Exact diagnostics for the Tweedie diagonal, and a gradient-free alternative —
- HTN Planning as a Coordination Layer for Multi-Server MCP Tool Orchestration —
- RSD-Poker: Structure-Adaptive and Shift-Robust Risk-Utility Certification for Residual Policies in Imperfect-Information Games —
- Program-Verified Self-Evolution for Vision-Language Models —
- How code helps different tasks? A decompositional lens on LLM post-training —
- Curating Merchant-Matching Training Data with Two Confidence-Gated Local LLM Judges —
- DEALS: Decentralized Expertise-Aware Load Serving for Multi-Agent LLM Systems —
- Learning Strategies to Break Judges —
- Evidence-Inference Reconstruction: When The Evidence Is Recalled But The Reasoning Goes Wrong —
- Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents —
- An Active-Bottleneck Mechanism for Weak-to-Strong Generalization —
- Laya as a Typed Probabilistic Assessor: An Independent Reproduction and a Preregistered Study of Calibration and Selective Escalation —
- Dual-Vocabulary Language Model for Cross-Tokenizer Distillation —
- PI-NOMT: Physics-Informed Neural Optimal Mass Transport for Brain Fluid Dynamics —
- Designing Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business Agents —
- HyperMCTS: Hypergraph-Augmented MCTS for Long-Horizon LLM Agents —
- A Computer Vision Approach to Visual Fraud Detection in Phishing Websites Using YOLOv8 —
- On the Token Value Inequality in Efficient Reasoning —
- Do World Models Learn Global Understanding? —
- Thinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV Inference —
- Jev in Medicine: A Benchmark Evaluation —
- ARCH-B: Architectural Representation, Comprehension and Hierarchy Benchmark —
- UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents —
- ZID-Net: Zero-Inference Diffusion Prior Decoupling Network for Single Image Dehazing —
- STRIDE: State-Transition Representation via Increment Dynamics and Evolution —
- What Should Data Teach? Moving Bottlenecks Across Circuit, Store, and Use —
- Feasible Flow Matching for Graph Reconstruction via Within-Sampling Primal-Dual Guidance —
- CARVE: Breaking Data Barriers in Chip Placement by Harnessing Reusable Expertise —
- Geometry-Aware Operator Families for Structured Representation Learning —
- From Constitutions to Control: Interpretable Rewards for Aligning Language Models —
- The limits of exactness: On the failure of automatic differentiation in physics-informed machine learning —
- Deep Learning Techniques for Phoneme Recognition in Italian Children' s Speech —
- Adaptive Ensemble Selection for Noisy Labels on Tabular Data —
- Balancing Early Performance Sacrifices with Long-Term Gains: Scaling Learning-Rate Warmup Duration Across Training Horizons —
- SketchSSM: Write to the Full State, Read from a Compact Sketch —
- Adaptive Latent Capacity for World Models —
- Saturation-Insensitive Dueling Bandits with General Function Approximation —
- Clipped or Unclipped? Finite-Sample Trade-offs for Averaged SGD under Heavy-Tailed Noise —
- DevelopmentODE: Structured Neural ODEs for Early Brain Development Dynamics Across a Decade —
- RMB: Reward Model Boosting Mitigates Reward Hacking —
- Orthogonal Witness Control for Muon Optimization via Sigmoid Spectral Reshaping —
- ILP-BO: Integer Linear Programming-Based Black-Box Optimization —
- Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs —
- Feedback-Robust AI for Patient Knowledge Graphs —
- Geometric Inductive Biases for Semi-Supervised Equalization: The Constellation-Aware Transformer —
- T-MoXAI: A Hierarchical Explainability Framework for Temporal Multimodal Data —
- From Grey-Box to Green-Box: When can Physics-Informed Machine Learning Reduce Carbon Footprints in Structural Health Monitoring? —
- Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration —
- StatD2GAN: When Calibration Masks Generator Quality in Held-Out Evaluation of Synthetic Weather Sequences —
- ResDiffFRG: Residual Diffusion for Multiple Appropriate Facial Reaction Generation —
- PQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate Attention —
- Theory Guided and Interpretable Neural Operator Design for Partial Differential Equation Learning —
- DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation —
- A Spectral Theory of Compositional Learning —
- Opera: A Verbal Critic Framework for Long-horizon Coding Agents —
- Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models —
- Steering Language Model Goals with Value Transplant —
- Faithful Activation Verbalization: Reducing Hallucinations in LLM Representation Interpretation —
- RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback —
- Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning —
- Zero-Shot Cue-Grounded Topic Segmentation of Spoken Documents —
- Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation —
- When Words Fall Short: Iterative Synergy Between Verbalized Reasoning and Hidden Features for LLM Confidence Estimation —
- Nudgeability: Reasoning Models Follow Confidence Signals Without Tracking Their Own Competence —
- Unbiased Top- k Estimation for On-Policy Distillation —
- ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models —
- Using LLMs to Detect LLM-Generated Texts: A Cross-Generation Analysis —
- Fair Fact-Checking: Closing the Cross-Lingual Gap in LLM Factual Judgement with RoSh —
- RoPE is Dead, Long Live RoPE: Towards Scalable Data-aware Positional Encodings —
- Low-Confidence Remasking Traps Flexibility: Realizing Arbitrary-Order Potential for Diverse Rollouts in Diffusion LLMs —
- RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications —
- Papers Without Code: Availability of GitHub Repositories Linked in *CL Publications —
- When Can Attention Heads Be Statically Defined? —
- The Model Knows When to Stop: Training-Free Early Stopping for Long-Context Reading —
- 3D Point Tracking with State Space Models —
- CLIMB-flow: Coupled Linear Inverse posterior sampling via Multiscale-Based flow —
- Re:Cognize -- Open-Set Comic Character Re-Identification —
- Preserving DEG Rankings for Gene Discovery in Histology-Based Spatial Gene Expression Prediction —
- ReDrive: Shaping Representations with World Modeling for End-to-End Driving —
- Residual-Stream Burden Shapes Representation Learning in Diffusion Transformers —
- Gaussian Splatting-based Volumetric Video Compression with Sparse 4D Anchors —
- A Multi-Dataset Benchmark of YOLO-Based Weed Detection in Precision Agriculture —
- SNaP: One-Step Posterior Sampling for Noisy Inverse Problems —
- WhiteCon: Semi-Supervised Domain Adaptation Regression Through Whitening Transform and Dual Consistency —
- Analytical and Convolutional Neural Network-Based Motion-Vector Propagation for Efficient Video Object Detection —
- PrefLUT: Reusable and Refinable Personalized Color Editing from Pairwise Preferences —
- Beyond Geometry: Benchmarking and Consistency Reasoning for 3D Logical Anomaly Detection —
- SpatialSkill: Self-Evolving Skills for Cross-View Spatial Reasoning —
- SpecRegMatch: Robust Semi-Supervised Regression for Vehicle Interior Noise Prediction —
- The Devil is in the Spectrum Bias: Spectrum-Balanced Feature Matching for Robust Representation Distillation —
- Functional Hand Type Prior for 3D Hand Pose Estimation and Action Recognition from Egocentric View Monocular Videos —
- MotionSpaceFlow: Representation-Aware Flow Matching in Direct Motion Space —
- Decision Readouts for Text-Mediated Video Anomaly Detection: An Exploratory Evaluation of Jev and Qwen —
- MaLiang-Harness: A Programmable Path to Image and Video Generation —
- MiCo: Mutual Information Coverage Optimization through Semantic Erasure Modeling for Efficient MLLM Inference —
- DORA: Dynamic Online Reinforcement Agent for Token Pruning in Vision Transformers —
- SkillPE: Creativity-Oriented Cinematic Skill Evolution for Text-to-Video Prompt Engineering —
- E-WAVE: Event-based Continuous Optical Flow via Warping-Aligned Visual Encoding —
- Marathoner: Ultra-Long-Horizon Autonomous Intelligence —
- When the Score Becomes the Target: Rethinking Metric Validity in Autonomous Driving —
- PACER: Progressive Availability-Conditioned Evidence Routing for Radiology Report Generation under Incomplete Clinical Context —
- When Does an Image Determine the Answer? Benchmarking Visual Answerability across Charts and Scenes —
- HPMD: A Historical Persian Manuscript Dataset for Word Spotting with Line-Level Annotation —
- ConCAD: Constraint-Aware Image-to-CAD Generation with Dual-Granularity Rewards —
- GenNVS: Geometry-enhanced Novel View Synthesis via Disentangled 3D Prior —
- The Statistical Cost of Causal Discovery with Feedback —
- Triangular Resampling for Long-Horizon Motion Generation —
- Singularities of Non-negative Matrix Factorization and their application to Bayesian inference —
- Unified Trajectory Matching Policy Optimization: Diverse T2I Generation and VLA Generalization —
- GT-PSSM: Unified Probabilistic Framework for Stochastic Dynamics Modeling and Dependency Learning in Multivariate Time Series Anomaly Detection —
- Two-Sample Testing for Inhomogeneous Random Graphs in Non-Integral L r Norms —
- Functional Autoencoders for Amplitude-Phase Representation Learning —
- ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control —
- Uniform Race: Parameter-Free Approximate Rejection Sampling —
- Rubric-as-Experts: Case-Specific MQM Rubrics for Translation Error Span Detection —
- Simple Extensions of Single-Objective Acquisition Functions and Hedge Strategies for Multi-Objective Bayesian Optimization —
- Handwritten Digit Leakage from Smartphone Motion Sensors Across Unseen Users and Phone Models —
- Beyond Temporal Smoothing: Spatial Energy Budgets Stabilize One-Step Diffusion Editing —
- S squared COPE: Self-Supervised Concept Discovery via Preference Learning —
- Wasserstein Convergence of ODE-Based Samplers in Decentralized Diffusion Model via Velocity Field Decomposition —
- ExpVoyager: Direct Experience Navigation for Dynamic Agent Skill Synthesis —
- ConTraIRL: Factorized Contrastive Abstractions for Transferable IRL —
- Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning —
- BitsMoE: Cost-Aware Bit Allocation in Spectral Space for MoE LLM Quantization —
- Tool Calling is Linearly Readable and Steerable in Language Models —
- DeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Agent Memory —
- DeepArrhythmia: Segment-Contextualized ECG Arrhythmia Classification via Selective Evidence Acquisition —
- Environment-Conditioned Tail Reweighting for Invariant Learning under Mixed Shifts —
- Hindsight Compacts but Does Not Repair: Rethinking On-Policy Self-Distillation in Reasoning Models —
- Hessian Matching for Machine-Learned Coarse-Grained Molecular Dynamics —
- Towards Fairness under Label Bias in Image Segmentation: Impact, Measurement and Mitigation —
- Federated Semi-Supervised Graph Neural Networks with Prototype-Guided Pseudo-Labeling for Privacy-Preserving Gestational Diabetes Mellitus Prediction —
- HO-SFL: Hybrid-Order Split Federated Learning with Backprop-Free Clients and Dimension-Free Aggregation —
- SOCKET: SOft Collision Kernel EsTimator for Sparse Attention —
- Generative Priors Conditioned on Natural Language for Bayesian Inversion in PDEs —
- Are Vision-Language-Action Models Robust to One-Step Observation Perturbations? —
- Bridging Generative and Discriminative Noisy-Label Learning via Direction-Agnostic EM Formulation —
- AraDynFact: Dynamic Evaluation of Factual Knowledge in Arabic —
- When and How to Canonize: A Generalization Perspective —
- Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks —
- Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs —
- Human Activity Recognition via Ultra-Wideband Data: A Framework for Dimensionality Reduction, Pattern Discovery, and Predictive Modeling —
- After the Fix: Transfer of Corrected Agent Experience —
- TCMQA: A 38K-Question Traditional Chinese Medicine Benchmark with a Licensed-Practitioner Reference —
- Relative Generalization Invariance of LLM Pretraining —
- Model-Agnostic Online Certificate-Driven Calibration for Time Series Forecasting Under Distribution Shift —
- FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching —
- Locally Sound, Globally Insufficient: The Local-Global Gap in Multi-Hop Reasoning —
- SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models —
- CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving —
- VCRE-Fib: View-Conditioned Regional Evidence for Fine-Grained Ultrasound Grading of Schistosoma japonicum-Associated Liver Fibrosis —
- Two-Timescale Fine-tuning Provably Learns New Features for Two-Layer ReLU Networks —
- In-game Toxic Detection: Bi-directional Representations with Attention Residuals —
- PhiFold: Towards Dynamic Protein Design with Physics-Structured Covariance Modeling —
- Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference —
- Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness —
- A theoretical model of dynamical grammatical gender shifting based on set-valued set function —
- CausalReasoningBenchmark: A Real-World Benchmark for Disentangled Evaluation of Causal Identification and Estimation —
- Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems —
- PINNMorph: Evolving Online Adaptation Policies for Physics-Informed Neural Networks —
- Toward Interactive Understanding of Code APIs —
- Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability —
- Bayesian Complete-Pooling in Cross-Subject Classification for Motor Imagery Electroencephalogram —
- 3D Gaussian Splatting with Fisheye Images: Field of View Analysis and Depth-Based Initialization —
- Why and When Deep is Better than Shallow: Implementation-Agnostic State-Transition Model of Deep Learning —
- COGNOS: Universal Enhancement for Time Series Anomaly Detection via Constrained Gaussian-Noise Optimization and Smoothing —
- Learning Conditional Expectation Operators via Functional Newton Updates —
- Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets —
- Denoising Time Matters:Diverse Generation in Diffusion Language Models —
- Towards Eliminating Catastrophic Forgetting in the Curriculum Learning of Math Reasoning Tasks —
- The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs —
- When Does a Skill Add Value? Task-Conditional Gain Prediction for Selective Skill Use —
- Forecasting Bacterial Antimicrobial Resistance Trends Using Machine Learning on WHO GLASS Surveillance Data: A Retrieval-Augmented Generation Approach for Policy Decision Support —
- Binding Multiple Modalities via Multimodal Wasserstein Barycenter —
- Distillation Defenses Easily Break After Reinforcement Learning —
- AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning —
- Scalable Attribution and Control of Model Behavior During Training —
- Short-Length Code Designs for Integrated Sensing and Communications: A Deep Learning Approach —
- GraphSelect for Budgeted Representation Selection in Multimodal Graph Inference —
- Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy —
- Posterior Regimes and Latent Deception: Variational Bayesian Inference in Hidden Markov Models for Sequential Fraud Detection in Financial Transactions —
- Pretraining Transformers with Quantized Softmax in Attention —
- HiLoRe: What to Store, Compress, or Recompute for Efficient GRPO Training —
- Fine Until Fine-Tuned: Repeated Solutions Make Reasoning Fragile —
- DiffMath: Symbol- and Graph-Aware Latent Diffusion Transformer for Handwritten Mathematical Expression Generation —
- SPR: Toward a Graph Foundation Model for Transferable Graph Cognition via Spectral Patterns and Relational Geometry —
- Discovering What You Can Control: Interventional Boundary Discovery for Reinforcement Learning —
- Beyond Fixed Features: Architecture-Dependent Sensitivity to Node Representations under Heterophily —
- VL-AcneSeg: A Vision-Language Framework for Region-Aware Acne Lesion Segmentation —
- EpiKV: Epiphany-Aware KV Cache Eviction Without the Attention Matrix —
- RefAdapt-DiT: Adaptive Joint Attention for Reference-Conditioned Diffusion Transformers —
- Multi-Scale Semantic Mapping in Urban Environments via Observation Calibration and Policy Dependence Regularization —
- Counterfactual Attention Policy Distillation for Temporal Video Grounding —
- Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders —
- One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion —
- TQTS-Bench: A Multi-Syntax Benchmark for Text-to-Query over Time-Series Databases —
- Flat-Consensus Diffusion for Robust Data Reshaping under Noisy Evaluator —
- Semantic Uncertainty Quantification Needs Factual Equivalence —
- REALIS: A Curated Dataset for Studying the Challenges of AI Image Detection —
- Contamination, Prior, or Evidence? Decomposing and Training Evidence Use in Whole-Slide Vision-Language Models —
- Coherence-Aware Distributional Evaluation of Open-Ended Text Generation —
- Controllable GNN Explanations via Multi-Metric Preference Selection —
- Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models —
- JIVE: Jacobian-Informed Volume Expansion for Diverse Generative Sampling —
- RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning —
- When Less Compute Is More: Adaptive Early Exit Improves Pretrained Outlier Detection —
- Beyond Square Roots: A Memory-Efficient Explicit Factorization for Multi-Epoch Private Learning —
- Generalization Analysis of Online Stochastic Gradient Descent for Overparameterized Two-Layer Neural Networks —
- TreeRef-BFN: Equivariance-Free De Novo Molecule Generation based on 2D Topology and Internal 3D Geometry —
- Modeling Whole-Slide Images as Dynamic Tumor Microenvironment Fields —
- Multi-Task Learning of Conditional Mean Operators: applications to dynamical systems and uncertainty quantification —
- PPG-LM: A Photoplethysmography-Language Model with Multi-Level Clinical Alignment —
- SynCo: Learning Cross-Modal Synergy by Contrasting Interaction Residuals —
- CapField-OPD: Learning Continuous Capability Fields via Joint-Anchored Multi-Teacher On-Policy Distillation for Flow Models —
- DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning —
- Structure-Adaptive Tree Field Integrators —
- l 1-2 GLasso: L 1-2 Regularized Multi-task Graphical Lasso for Joint Estimation of eQTL Mapping and Gene Network —
- Improving Causal Effect Estimation of Weighted RegressionBased Estimator using Neural Networks —
- OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents —
- STEP-LLM: Generating CAD STEP Models from Natural Language with Large Language Models —
- Deep Delta Learning —
- SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints —
- Soft Geometric Inductive Bias for Object Centric Dynamics —
- Categorical Approach to Conflict Resolution:A Corrected Correspondence between Category Theory and the Graph Model for Conflict Resolution —
- Post-detection inference for sequential changepoint localization —
- Goal-Conditioned Supervised Learning for Multi-Objective Recommendation —
- Patch Rebirth: Fast and Transferable Model Inversion of Vision Transformers —
- Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization —
- Searching for Actual Causes: Approximate Algorithms with Adjustable Precision —
- Focus on Likely Classes for Test-Time Prediction —
- A general framework for adaptive nonparametric dimensionality reduction —
- MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues —
- Latent Action Reparameterization for Efficient Agent Inference —
- Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs —
- Physics-Informed Neural Networks with Architectural Physics Embedding for Large-Scale Wave Field Reconstruction —
- CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models —
- GUI-GenBench: Evaluating Image Generation Models as Interactive GUI Environments —
- VLM-Guided Experience Replay —
- Learning Transferable Sensor Models via Language-Informed Pretraining —
- mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT —
- Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss —
- Communication Gain and Delay Cost Under Cross-Timestep Delays in Cooperative Multi-Agent Reinforcement Learning —
- An Inspectable LLM Council for Multi-Model Research Answer Aggregation —
- MambaSL: Exploring Single-Layer Mamba for Time Series Classification —
- Towards Effective Theory of LLMs: A Representation Learning Approach —
- Recovering Physical Dynamics from Discrete Observations via Intrinsic Differential Consistency —
- AIPO: Learning to Reason from Active Interaction —
- LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment —
- MELD: Multi-Task Equilibrated Learning Detector for AI-Generated Text —
- How Deep Can LLMs Learn to Reason? Expressiveness Is Key —
- What Happens Inside Agent Memory? Circuit Analysis from Emergence to Diagnosis —
- Alignment has a Fantasia Problem —
- RepoMAS: Solving Progressively Specified Tasks with Issue-Driven Multi-Agent Systems —
- Universal Quantum Transformer —
- Multi-Legal-Bench: When the Answer Is in the Input. Label Leakage in Legal Benchmarks Built from Court Registries —
- Anytime Training with Schedule-Free Spectral Optimization —
- Fidelity Probes for Specification--Code Alignment —
- Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks —
- Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack —
- SkillSmith: Co-Evolving Skills and Tools for Self-Improving Agent Systems —
- StarOR: Synergizing Tree Search and Test-Time Reinforcement Learning for Optimization Modeling —
- Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity —
- DWM: Separating World Effects from Actions in Latent World Models —
- The Steering Budget: Examples beat Knobs —
- Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Self-Improving Personal Agents —
- TopoExplore: Homology as an Exploration Signal —
- One Image, No Tokens: A Controlled Study of Glyph-Based Chinese Language Modeling —
- Scaling Laws for Collapse in Asynchronous GRPO —
- Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition —
- Mechanistic Personality Analysis of LLMs: Steering Personality via Latent Feature Interventions —
- JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators —
- Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models —
- ClinLens: Towards Long-Horizon LLM Agents for Longitudinal Multimodal Clinical Data Science —
- FLARE: Flow Matching with Local Axis-Angle Representations for Stochastic Micromagnetic Evolution —
- WorldAttention: An Efficient Attention Architecture for Interactive Video World Models —
- KiT: A Foundation Model for Financial Time-Series Forecasting using DiffusionTransformers —
- Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration —
- CRISP: Cultural Reward Modeling for Implicit Situated Propriety —
- Role Support in Knowledge-Graph Error Ranking: Predictor Regimes and Evaluation Policies —
- Shapley-based Data Valuation for LLM Alignment via Sequential Preference Optimization —
- LOD: Latent Objective Discovery in Heterogeneous Multi-Agent Reinforcement Learning —
- FEAT: Free energy Estimators with Adaptive Transport —
- Inversely Learning Transferable Rewards via Abstracted States —
- Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting —
- Estimating the Hessian Matrix of Ranking Objectives for Stochastic Learning to Rank with Gradient Boosted Trees —
- PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving —
- f-FUM: Federated Unlearning via min--max and f-divergence —
- Continuous First, Discrete Later: VQ-VAEs Without Dimensional Collapse —
- Multiscale Euclidean Network Trajectories: Second-Moment Geometry, Attribution, and Change Points —
- When Should Humans Step In? Optimal Human Dispatching in AI-Assisted Decisions —
- Complete-muE: Optimal Hyperparameter Transfer and Scaling for MoE Models —
- FIS-DiT: Breaking the Few-Step Video Inference Barrier via Training-Free Frame Interleaved Sparsity —
- Convergence Analysis of Newton's Method for Neural Networks in the Overparameterized Limit —
- A Theory of Saddle Escape in Deep Nonlinear Networks —
- Diagnosing Visual Ignorance in Vision-Language Models —
- EntmaxKV: Support-Aware Decoding for Entmax Attention —
- Generalized Functional ANOVA: A Complete Theoretical Framework —
- To MRL or not to MRL: Text Embeddings are Robust to Truncation Without Matryoshka Learning, Except In Heavy Truncation Scenarios —
- When Can We Trust the Sparse Lens? A Certification Framework for SAE Faithfulness —
- Multi-Bitwidth Quantization for LLMs Using Additive Codebooks —
- A Continuous-Time Reinforcement Learning Framework for Fine-Tuning Discrete Diffusion Models —
- Interval and fuzzy physics-augmented neural networks (iPANN and fPANN) for uncertainty quantification and propagation in constitutive modeling —
- Robust Peak-cost Constrained Reinforcement Learning —
- Extractable Memorization From First Principles —
- SAGA: Structure-Attended Generative Action Embedding Model that encodes Multi-Surface User Action Sequences —
- BOOKMARKS: Efficient Active Storyline Memory for Role-playing —
- Parallel Tokenizers: Rethinking Vocabulary Design in Cross-Lingual Transfer of Low-Resource Languages —
- Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMs —
- In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners —
- ALAS: An Automatic Latent Alignment Score for Audio Language Models —
- HiMed: Incentivizing Hindi Reasoning in Medical LLMs —
- Testing the Assumptions of Active Learning for Translation Tasks with Few Samples —
- JuDi: Revisiting Judge Decoding from First Principles via Training-Free Distributional Divergence —
- MAPLE: Medical Aspect-Based Summarization with Phrase-Level Evidence —
- Skill Training with Corruption and Reconstruction Loop —
- The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models —
- UXBench: Benchmarking User Experience in AI Assistants —
- CAT-Free: Multi-View Pedestrian Localization without Calibration, Annotations, or Target-Scene Training via Adaptive Geometric Filtering —
- UnfoldCRF: Structured Mask Refinement with Image-Conditioned Latent Regions —
- Understanding Confabulation and Rethinking Reconstruction in Activation Explanations —
- Trust and Task Completion in the World of Consumer AI Agents —
- DPAMixerSR: An Efficient Degradation-Pattern-Aware Model for Image Super-Resolution —
- What Should We Freeze? Guarded Freezing: Connectivity Shapes the Fine-Tuning of Pretrained Models —
- Acceptance Dynamics Across Cognitive Domains in Speculative Decoding —
- Language as an Independent Information Layer: A Conceptual Model of Communication, Cognition and Decision-Making —
- Simple Diffusion Language Models Are More Effective Few-Step Generators Than Reported —
- Towards Identifiable Representations under Misspecified Structure —
- Octree-based Video Representation —
- Agentic Network Traffic Monitoring —
- Receiver-Conditioned Latent Communication gives 94% CacheBack —
- Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems —
- LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks —
- Domain Adaptation with Target Information via Doubly-Anchored Distributionally Robust Optimization —
- MemoReason: Evaluating the Effect of Parametric Memory on Contextual Reasoning in LLMs —
- Choir: An Open Protocol for Distributed Multi-Agent Autoformalization —
- Carnator: Fast Text-to-Video Generation with Generation-Native Compatibility-Guided Cross-Request Reuse —
- SurgGMF: Fully Causal Gaussian Motion Forecasting for Anticipatory Surgical Scene Rendering —
- InsHuman: Towards Natural and Identity-Preserving Human Insertion —
- Mantis: Mamba-native Tuning is Efficient for 3D Point Cloud Foundation Models —
- MooD: Perception-Enhanced Efficient Affective Image Editing via Continuous Valence-Arousal Modeling —
- Generative Classifier for Domain Generalization —
- Towards a Universal Image Degradation Model via Content-Degradation Disentanglement —
- Extracting Neural Materials from Images —
- Self-eXplainable AI for Medical Image Analysis: A Survey and New Outlooks —
- A Multimodal Feature Distillation with Mamba-Transformer Network for Brain Tumor Segmentation with Incomplete Modalities —
- The Judge Is Not Its Twin: Post-training makes a model's writing more predictable but barely moves its taste, as a judge, toward predictable writing —
- Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs —
- T-SNN: Temporal Simplicial Neural Network for EEG Decoding —
- Dexterous Tactile World Model —
- Multi-Marginal Inverse Optimal Transport for Contrastive Learning Via Explicit Anchor-Positive-Negative Coupling —
- The Extender: A Log-Structured Transformer —
- Quizzing the Translation: A Prover-Grounded Evaluation Metric for NL to FOL —
- WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning —
- From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark —
- Stage-adaptive Token Selection for Efficient Omni-modal LLMs —
- Dinomaly2: A Unified Framework for Unsupervised Image Anomaly Detection —
- On the Behavioral Traits of LLM Agents —
- Write Back the: Revisiting the Same Tokens with Fresh Representations —
- SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models —
- Levy-Driven Correspondence Estimation for Registration —
- QureRadEmbed: Structuring Radiological Similarity through Attribute and Reasoning Supervision —
- A Solvable Theory of Pre-training Data Poisoning: Regime-Dependent Scaling Exponents —
- Enabling Unsupervised Training of Deep EEG Denoisers With Intelligent Partitioning —
- Towards Transparent Diagnostics: Investigating Architectural Trade-offs and Explainability in Malaria Detection —
- Lightweight Wrappers for Adapting Time Series Foundation Models to Regional Drought Forecasting —
- One Perception, All Maneuvers: Directional Traffic Signal Understanding for Maneuver-Level Signal Intent Prediction —
- TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces —
- Fusing Complementary Multi-view Features for Screen-Based Eye Tracking —
- Full-Body Golf Swing Kinematic Reconstruction From a Smartwatch IMU —
- Rethinking Open-World Video Anomaly Detection: Diagnosing Definition Blindness —
- WarpHammer: Densifying Scene Warps with 3D Object Priors for Extreme View Synthesis —
- VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes —
- Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs —
- Deep Weighted Bellman Residual Minimization for Q* Estimation —
- Weighted Spline-Expanded Networks with Distributional Balancing for Continuous Treatment Effects —
- Instruct, Not Answer: Using Instruction Privileges in On-Policy Context Distillation —
- Re-derivability Decides What a Staged Agent Pipeline Recovers After an Upstream Fault —
- FILIGREE3D: Scaling Sparse Latent Flow Matching for Ultra-High-Resolution Image-to-3D Generation —
- Context Pruning for Coding Agents via Multi-Rubric Latent Reasoning —
- WorldGuide: Learning Success-Failure Boundaries in Latent World Models for Vision-Language-Action Policies —
- SenseAgent: An LLM Agent for Adaptive Cross-Domain IMU Sensing —
- HoTS: Homophily-Aware Temperature Scaling for Graph Neural Network Calibration —
- Advancing Wildlife Conservation through Multimodal Animal Re-Identification with Environmental Metadata —
- M3OS: A Monte Carlo Graph Search-Orchestrated Multi-Agent LLM System for Evidence-Traced Molecular Optimization —
- A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards —
- PSM: Dataset Distillation Based on Precise Statistical Matching by Difficulty —
- Temporal Graph Learning of Wearable Actigraphy and Sleep Traces for Modelling Adolescent Crystallized Intelligence —
- Toward Open-World Video Segmentation over Long Horizons —
- EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events —
- Do System One Decisions Add Up? A Study of Probabilistic Coherence —
- Codifying the Judge: Scalable Evaluation via Program Distillation —
- Reliable Replay through Spatial Coherence in Online Continual Learning —
- Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference —
- Rubric-Aware On-Policy Self-Distillation for LLM Personalization —
- NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability —
- Towards Autonomous and Auditable Medical Imaging Model Development —
- OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration for ARC-AGI-3 —
- Scriboora: Rethinking Human Pose Forecasting —
- Benchmarking EEG Foundation Models at Scale: Lessons from 20,000 Evaluations —
- Elicitation and Decision Geometry in Single-Index Bandits —
- One Analyst Is Not Ground Truth: Grading Agent-Built Financial Models Against Observed Professional Practice —
- ActionUNet: Improving Robustness of VLA Models with Efficient Multi-scale Fine-tuning —
- Scaffold Then Internalize: Representation Injection for Diffusion Transformers —
- SegBanana: Steering Unified Multimodal Models into Medical Segmenters —
- When Helpful Text Hurts: Option-Redirecting Bias in Vision-Language Models —
- ALDER: Discovering the Laws of a World by Acting in It —
- BudgetVerify: Budget-Tiered Verification for Financial QA —
- ECHO: Event-Augmented Context with Hindsight and Outlook for Wrist-Only Manipulation —
- Model Compression with Exact Budget Constraints via Riemannian Manifolds —
- OneCanvas: 3D Scene Understanding via Panoramic Reprojection —
- SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis —
- ChemOPD: Multi-Teacher On-Policy Distillation for Multi-Task Chemical Reasoning —
- SMORE: Stability-Promoting Mesh-Agnostic Model Reduction for Time-Dependent PDEs —
- Mind the Perspective: Let's Reason Recursively for Theory of Mind —
- StoryEngine: A State-Grounded Agentic Framework for Video Storytelling —
- Trapped by Their Own Rollouts: Understanding Aggregation--Rollout Feedback in Federated On-Policy Distillation —
- Supporting and Performing Culture from the Inside —
- PostCam: Camera-Controllable Novel-View Video Generation with Query-Shared Cross-Attention —
- Multilinguality in Hybrid Attention LLMs —
- Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry Transformer —
- A Reproducibility Study of Partial Residual Ablations in Pre-LN Transformers —
- eTracer: Polarity-Aware Evidence Grounding —
- TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM —
- Presence Is Not Faithfulness: Figurative Vehicle Intrusion in Text-to-Image Generation —
- Context-aware tokenization for Cross-subject Emotion Decoding from EEG —
- Where Does the Watermark Hide? Push-Pull Disentanglement for Invisible Watermark Removal —
- NeuronSifter: Intervention Planning in CNS Microenvironments —
- On the Limits of Metacognitive Monitoring in LLMs —
- Separating personal from population gains when calibrating EEG foundation models for new users —
- Pass or Fail? Evaluating LLMs on Two Greek Examination Benchmarks —
- Reference-Grounded Data Curation for Instruction-Following Thai-English Machine Translation —
- Draft-KV: Learning Useful Latent Communication Between Language Models —
- Beyond Token Alignment: Event Completion for Cross-Tokenizer On-Policy Distillation —
- Quality Determines Direction, Length Shapes Magnitude: Length Control for Open-Ended Reinforcement Learning —
- ReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code Generation —
- Revisit to Segment: Working Memory Distillation for Reasoning Segmentation —
- SubRot: Signed Gradient Subspace Calibration for VLM Rotation Quantization —
- SPOC-Net: Single-Primitive Online Composition Network for GNSS Jamming Set Recognition —
- P4Q: Co-designing Token Pruning and Quantization for Vision-Language Model Acceleration —
- When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model —
- ORAV: Benchmarking Audio-Video Generation from Multimodal Contexts —
- Transform-Aligned Learned Features for Lossy Point Cloud Attribute Compression —
- From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models —
- When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment —
- Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts —
- From Input to Output: A Flexible Agent for Dual-End Interpretation of Sparse Autoencoder Features —
- Sample What You Say: Aligning Language Models to Sample the Distributions They State —
- The Right Lesson at the Right Step: Deriving Control Updates for Self-Evolving Agents —
- Adapt Semantics, Not Structure: Few-Instance Schema Calibration for Scientific PDF Extraction —
- From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data —
- EvoIn: Bridging Evolution and Internalization for Agent Fine-Tuning —
- From Weak Task Specifications to Scientific Extraction Agents: Optimizing Task Construction —
- BV Loss: Block Verification-Aware Loss for Block Diffusion Speculative Decoding —
- When Words Speak Louder than Images: Towards Understanding Language Bias in Vision-Language Models —
- WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents —
- DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents —
- 5W1H+Which: Context-Valid Semantic Indexing with Progressive Ontology Binding —
- Don't Forget! Decomposing the Training Dynamics of Memorization in Language Models —
- Rubric Rewards from Item Response Theory —
- TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science —
- Twist, Don't Tilt: Trajectory-Exact Constrained Decoding for Masked Diffusion Models —
- Shockingly Simple Self-retrospection Improves Agentic Models Without RL —
- How to Loop MoE: Flatten the Experts, Untie the Attention —
- Representation Alignment as a Bottleneck in LLM-Based Retrosynthesis Planning —
- Can LLMs Value the Right Evidence? Evidence-Value Misalignment in Dynamic Medical Diagnosis —
- Harness Learning Enables Generalizable Test-Time Adaptation —
- Simultaneous Translation between Sign Languages —
- Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery —
- LLMs are General Asynchronous Agents —
- Late Attention Layers Alone Can Copy Entity Tokens, but Not Without Attending to Their Context —
- Language Models Act on Hidden Valence —
- Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition —
- Retrieving Biblical Intertextual References in Karen Blixen's Seven Gothic Tales —
- ReVA: A Scene-Centric Dataset Beyond Repetition for Remote Sensing Video Question Answering —
- Derivative-Informed Training of Neural Operators On-the-Fly via Sketched Tangent Consistency —
- Can Attack Difficulty Be Characterized Before Optimization? A Study of Pre-optimization Difficulty in Person-Vanishing Attacks —
- FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs —
- A literature-guided descriptor-based framework for filtering composition search spaces —
- SCBO: Semantically Coherent Batching and Ordering for LLM-Based Social Surveys —
- On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models —
- Editable Map-Conditioned Trajectory Generation for Human Mobility Simulation —
- Self-Designed Evaluators and Warm Memory for Long-Horizon Agents —
- From Static to Dynamic: On-Policy Distillation from Image to Video Diffusion Models —
- World Models with Predictable Long-Horizon Marginals —
- SchemaMem: Schema-Indexed Recurrent Memory for Delayed State Retrieval —
- LLM Unlearning Evaluation with TRIAGE —
- V-Gym: Enhancing Agentic Visual Reasoning via Skill-Data Co-Evolution —
- How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining —
- ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network —
- SANTA++: Sampling Attention through Representative Keys —
- LatentReRig: An SDF-Based VAE with Dual Decoders for Latent-Space Deformation Conditioning —
- FlashSampling: Fast and Memory-Efficient Exact Sampling —
- Bidirectional Information Flow (BIF) - A Sample Efficient Hierarchical Gaussian Process for Bayesian Optimization —
- SSD: Spatially Speculative Decoding Accelerates Autoregressive Image Generation —
- TerMeZO: Ternary Sparse Zeroth-Order Optimization for Fine-tuning BitNet Models at the Edge —
- Identifiable Convex-Concave Regression via Sub-gradient Regularised Least Squares —
- In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion —
- Devol-ONE: One Autoregressive Mixture of Transformers to Unify Vision-Language-Action and Latent World Modeling —
- Handwritten Text Recognition Lives in the High-Pixel Variance Subspace —
- Borrowing from anything: A generalizable framework for reference-guided instance editing —
- Neural Language Models Learn the Contextual Distributions of Dependency Structures: a statistical learning theory to compositionality —
- ControlTrace: Recovering Control Fields for Hidden-Content Recognition —
- Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training —
- EviSplat: Preserving Multi-View Evidence in 3D Gaussian Splatting for Open-Vocabulary Segmentation —
- Self-Adapting Group of Experts for Multi-Agent Reasoning —
- From UNI2-h to ConvNeXt-T: Lightweight Nuclei Instance Segmentation via Knowledge Distillation —
- Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents —
- N"urnberg NLP at ChildSafeAds 2026: Structurally Dissimilar Voter Ensembles under Four Levels of Data Access —
- USA: Update-aware SAM for Cross-domain On-Policy Disitllation of Language Agents —
- VisionPsy-Nano: Improving Accuracy, Efficiency, and Reliability in On-Device Vision-Language Models —
- FlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning —
- ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients —
- Control-Geometry Straightening for Sampling-Based Latent Planning —
- NSV-Shift: A Contrastive Benchmark for Non-Speech Vocalization Understanding and Response Adaptation in Speech-to-Speech Models —
- Leaky Students: Membership Inference against On-Policy Distillation —
- Superposed Inference for Hyperdimensional Computing —
- Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models —
- Enhanced Video Text Editing with Trajectory-Aligned Glyph Rendering —
- Hierarchical Response Preservation for Continual Adaptation of Zero-Shot Graph-Text Models —
- SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation —
- Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding —
- MoSPR: Histology-to-Gene Expression Prediction with Morpho-Spatial Macrostates and Low-Rank Molecular Programs —
- Concept Score Relearning: A Unified Cross-Architecture Attack on Concept Erasure —
- Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models —
- Decision-Sufficient State Representations: Measuring and Reducing Write-Time Regret —
- Using Machine Learning to Investigate Predictors of Fasting Blood Glucose: Insights into Circadian Timing and Age Interactions —
- Motion Forcing: A Decoupled Framework for Robust Video Generation in Motion Dynamics —
- OmniSmartHome: A Multimodal Reasoning Benchmark for Smart-Home Agents —
- Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models —
- Finite Probes Suffice: Identifiability and Universality for Weight-Space Learning —
- When do Random Forests work? —
- Reasoning on the Simplex: Geometric Fixed-Point Models —
- A Dual Representation of Influence Functions for Linearizable Models —
- Manifold limit for the training of shallow graph convolutional neural networks —
- A VLM Answer Is Not an Anomaly Score: Rank Compression Across Image and Video Anomaly Detection —
- Look Inside Each Video: Rethinking How Video Anomaly Detection Is Evaluated —
- FUND: Density Flow for Sampling Unnormalised Distributions —
- BOCCHI: A More Realistic and Challenging Benchmark for Local Motion Blur Detection with MSDCT-UNet —
- SkillVine: Agent Skill Evolution via Branching Exploration —
- AugRelNet: Relation-Augmented Dynamics for Structure Discovery and Forecasting from Limited Data —
- LINC: Decoupling Local Consequence Scoring from Hidden Matching in Constructive Neural Routing —
- PIC-UIE: Predicting Image-Adaptive Corrections for Lightweight Underwater Image Enhancement —
- PrefixGuard: Online Failure Warning and Trace-Grounded Diagnosis for LLM Agents —
- EveryQuery: A Promptable Foundation Model for Clinical Prediction Tasks over Electronic Health Records —
- Synergizing Discriminative Exemplars and Self-Refined Experience for MLLM-based In-Context Learning in Medical Diagnosis —
- RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes —
- R2T: Rule-Encoded Loss Functions for Sequence Tagging in Low-Resource Languages —
- Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise —
- Advanced Policies: A First-Principles Path from Policy Gradient to Q-Learning —
- Dr. Free: You Don't Need Difficulty Rewards for Self-Evolving Search Agents —
- R squared Flow: Recursive Self-Improvement via Recursive Skill Evolution —
- The Effects of Incremental Instruction Delivery on Language-Model Creative Writing —
- A Visual Classification Dataset and Model Evaluation for Historical Manuscript Illustrations —
- ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts —
- When Known Physics Helps Neural PDE Models: Residual Constraints Out-Regularize Generic Priors for Nonlinear Dynamics —
- ALLOT: Budgeted Hybrid-Memory Routing for Knowledge Updates in LLMs —
- DimPO: Dimensionality Reduction for Attention using Preference Optimization —
- Evidence-Aligned Multimodal On-Policy Self-Distillation for Fine-Grained Visual Understanding —
- CAST: Reconstruction-Coupled Acceleration of Interactive World Models —
- From internal representations to model improvement through prediction errors —
- MergeHEIR: Mitigating Multimodal Hallucinations as the Tax of Model Merging —
- ForeFly: A Dual-Horizon World Action Model for Aerial Vision-Language Navigation —
- Proxy2World: Learning to Generate Worlds From Lightweight Proxies without Seeing Them —
- Use What You Know: Causal Foundation Models with Partial Graphs —
- AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents —
- Counterexamples to Local Reconstruction Gain as a Proxy for Final Fidelity in Residual Completion —
- Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders —
- Preventing Rank Collapse in Federated Low-Rank Adaptation with Client Heterogeneity —
- Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision —
- Scaling Long-Form Story Generation via Narrative State Tracking —
- Adapting Nonstationary Multi-output Gaussian Processes to Bayesian Optimization —
- HiFloat4 Format for Language Model Pre-training on Ascend NPUs —
- RAEGL: Risk-Aware Evidence-Gated Learning for Selective Contextual Routing under Temporal Shift —
- On Memory: A comparison of memory mechanisms in world models —
- SpecRead: A Benchmark for Measuring Whether Language Models Understand Hardware Specifications —
- Reinforcing Agentic Creativity in Scientific Ideation with Night Science —
- Topology-Adaptive Hyperbolic Graph Attention Networks Guided by the Hyperbolic Sombor Index —
- Untangling Input Language from Reasoning Language: A Diagnostic Framework for Cross-Lingual Moral Alignment in LLMs —
- Shared Worlds, Private Minds: Structured Memory for Long-Form Writing as World Creation —
- DirectUV: Image-Conditioned UV Texture Generation with Surface-Aware Positional Encoding —
- COTCAgent: Preventive Consultation via Probabilistic Chain-of-Thought Completion —
- Next Thoughts Are Distributions: Generative Autoregressive Reasoning in the Latent Space —
- VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis —
- Learning an Anchored Prompt Space for Continual Adaptation of Large Language Models —
- DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time —
- NeuronDiscover: Agent-in-Twin for Mechanistic Discovery in Neuronal Microenvironments with World Action Models —
- ARAPDiffusion: Geometry-Distribution Feedback for Deformable 3D Shape Generation —
- AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? —
- LLMAdBench: A Human Preference Benchmark for Advertising in LLM Responses —
- Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents —
- X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization —
- Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores —
- HyperLabel: Multi-Label Classification via Hypergraph-Based Label Correlation Modeling —
- PC-SubMax: Efficient Prompt Compression via Regularized Submodular Maximization —
- SinBrief: A Hybrid Framework for Abstractive Text Summarisation of Sinhala Legal Documents —
- OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories —
- Building Intelligent Agents with Neuro-Symbolic Concepts —
- Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment —
- LAM: Efficient Lossy Agent Memory Framework With A Retrieval-Score Error Bound —
- SafeMol: Dual-Modality Safety Alignment for Molecular Multimodal Models —
- SHEFL: Sparse Heterogeneous Ensemble Federated Learning with Group-Balanced Aggregation —
- TrajGANR: Trajectory-Centric Urban Multimodal Learning via Geospatially Aligned Neural Representations —
- Beyond Saying Less: Fine-Grained Alignment for Informative and Faithful Vision-Language Models —
- Less Is More: Genetic Frame Selection for Efficient Novel View Synthesis —
- ReLoc: Rethinking Scene Coordinate Regression Architecture for Robust Outdoor LiDAR-based Localization —
- Learning Evidence Highlighting for Frozen LLMs —
- CertMark: Distortion-Free Multi-Bit Watermarking with Certified Decoding —
- Downstream-Aware Context Selection for Online In-Context Reinforcement Learning —
- GeoShrink: Accelerating Diffusion Transformers with Two Lines of Code —
- Permutation-Equivariant Flow Matching for Alignment-Free Neural Weight Generation —
- Contextual Distributionally Robust Optimization with Causal and Continuous Structure —
- Wavelet-Based Parity Detection Revisited: Representation Dependence, Generalization, and Mechanistic Analysis —
- Fisher Simplicity in Kolmogorov-Arnold Networks and Multilayer Perceptrons —
- PACGNet: Pyramidal Adaptive Cross-Gating Network for Multimodal Object Detection in Aerial Imagery —
- Continual Learning via Self-Probe Gradients —
- Max-pooling Network Revisited: Analyzing the Role of Semantic Probability in Multiple Instance Learning for Hallucination Detection —
- Representation Editing for Multimodal Test-Time Adaptation —
- One Threshold Is Not Enough: Prompt-Invariant Caching Schedules for Video Diffusion —
- ReFM: Semantic-Aware Refinement Flow Model for Motion Retargeting —
- Reciprocal Guidance: Orchestrating Draft and Verify Budgets for Advancing the Diffusion-AR Self-Speculation Frontier —
- Collaborative Synthetic Data for Privacy-Preserving Financial Fraud Detection Across Organizational Silos —
- PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation —
- Rewarding Novel Deductions: Solver-guided Process Supervision for Logical Reasoning —
- Noise-adapted Neural Operator for Robust Non-Line-of-Sight Imaging —
- Retrospective Distillation Attribution via Normalized Response Similarity —
- Eyes on the Road: A Naturalistic Comparison of MTW Rider Gaze in Urban Indian Traffic —
- Double-Edged Sword of Mediated Visibility: How Visual Framing Undermines Congresswomen's Perceived Competence —
- Learning High-Risk High-Precision Motion Control —
- Towards Communication-Efficient Social Intelligence in Language Agents —
- Recent Advances in Agentic Agri-Robotic Phenotyping: A Perspective Review from Fragmented Multimodal Sensing to Unified PhenoAgent Intelligence —
- Augmenting Visual Anomaly Detection with Automated Interpretability —
- Resource-Aware Federated Mixture-of-Experts with Adaptive Pruning for Onboard Learning in LEO Satellite Constellations —
- Extraction of clinical findings from mammography and breast ultrasound reports: a comparison between specialists and Artificial Intelligence —
- Adjoint Guidance Flow: Amortized Critic Guidance for VLA Policies —
- Are You Sure You're Sure? Two Confounds in a Sycophancy Benchmark —
- Unlocking Few-Step Diffusion for Faithful Previews —
- C-HAT-Bench: Benchmarking Chinese AI-Text Detection Beyond Fully Generated Text —
- SceneScaffold: Active Scene-State Construction for Unified 3D Scene Understanding —
- See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology —
- How Far Is Document Parsing from Solved? PureDocBench: A Source-Traceable Benchmark across Clean, Degraded, and Real-World Settings —
- Agnostic Smoothed Online Regression with Adversarial Responses —
- Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features —
- Sliced-Regularized Optimal Transport —
- What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling —
- The Epistemics of Agent Memory: Measuring, and Governing, the Consolidation Decision in Long-Horizon LLM Agents —
- FurE: Efficient Instance-Specific 3D Fur Reconstruction without Animal-Fur Datasets —
- SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing —
- Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference —
- Capability Self-Assessment in Large Language Models —
- KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation —
- BioDyad: Synchronize Biomedical Discovery and Machine Learning Engineering —
- When Does Geometric View Synthesis Help Wine Label Retrieval? A Public One-Shot Benchmark Across Self-Supervised and Vision-Language Backbones —
- ReSight-SMC: Two-Stage Power Sampling via Island SMC with Visual Scouts —
- Gauge-Equivariant Attention for Rotation-Stable 360 Scene Understanding —
- Muon Under Gradient Noise and the Limits of Orthogonalization Near Optima —
- Distributionally Robust Average-Reward Reinforcement Learning: Finite-Sample Guarantees under Weak Communication —
- Evaluating Single and Multi-Omics Based Explainable Artificial Intelligence (MOXAI) for Molecular Subclass Classification of Adult-Type Diffuse Gliomas —
- Readout is not Recovery: Dissociating Coordinate Emission from Visual-Corruption Repair in Vision-Language Models —
- MultiEcho: An Experimental Science of Learned Worlds —
- CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning —
- IGSD: Environment-Verified Hindsight Self-Distillation for Search Agents —
- Oracle Gaps in Reliability Coverage: Sampling Noise or Policy Specialization? —
- AuthorityLens: Rethinking LLM-Based Agent Systems Through the Lens of Authority —
- Latent Space Is Not Flat: Rethinking Latent Structure for 3D Medical Image Synthesis —
- SWE-MILE: Asynchronous Potential-Induced Milestone Credit Assignment for Long-Horizon Software Engineering Agents —
- Cost-free Spectral Estimation for Adaptive Newton--Schulz in Matrix Optimizers —
- SAGE: Subspace Alignment for Classifier-Free Guidance in Mixture-of-Experts Diffusion Models —
- CAR-VLA: Complexity-Aware and Risk-Adaptive Reasoning for Autonomous Driving —
- ReVision3D: Attribution-Guided Recursive Self-Improvement for 3D Medical Perception —
- CP-Agent: A Harness-Engineered Agent for Crystal Plasticity Simulation Workflows —
- Efficient Message Passing for Partial Differential Equation Priors —
- Sprout: Building Dynamic Memory While Reasoning for Agentic Video Understanding —
- Transformer MLP Gate Thresholds Are Couplings to a Carried Reference Direction —
- Allspark: Weak to Strong Transfer via Alternating Chain of Thought —
- Working with AI: A Design Framework for Human-AI Collaboration —
- Superquadric Primitive Decomposition of 3D point clouds via Geometric-Aware Inlier Refinement —
- DroneWAM: Efficient World Action Model for Drone Visual Navigation —
- AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering —
- Large Language Models Substantially Compress Well-Being Inequality but Largely Preserve Its Socioeconomic Structure —
- What Does a ProcGen Generalization Gap Measure? Action Rules, Convergence, and the Missing Random Floor —
- Learning Robust Recommenders from Noisy Implicit Feedback via GMM-Weighted Bayesian Transition Matrix —
- Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models —
- Bridging Stochastic Flow Maps and Boltzmann Generators with Normalizing Flows —
- Rethinking Cross-Channel Importance in Time-Series Forecasting —
- PARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence Escalation —
- How Linear Attention Remembers —
- Why Does Agentic Safety Fail to Generalize Across Tasks? —
- OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs —
- Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models —
- Leveraging Soft Prompts for Privacy Attacks in Federated Prompt Tuning —
- VisionHOPE: Visual Backbones as Self-Modifying Learning Systems —
- ANTMAN: Adaptive Need Tracking for Multi-Agent Navigation in Large Information Spaces —
- Manifold-Stable Flow Matching —
- Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models —
- Just-In-Time Agent Memory with Runtime Agentic Research —
- OOD Generalization as a Bifurcation Problem —
- Self-Confirming Superposition Traps in Reinforcement Learning —
- Propagate, Then Sharpen: Post-Hoc Refinement of Frozen Node Classifiers —
- Generalization Dynamics of LM Pre-training —
- Shared Autoregressive Context Can Distort Relationships in Synthetic Data —
- Rethinking Contextualization by Reinterpreting Attention Head Channels —
- Distributional sentiment modeling and anomaly detection for consumer complaint assessment —
- Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation —
- Back-Tracking from Clarity: Self-Learning to See Text from Afar —
- Does CoT-Pass@k Really Check the CoT? A Multilingual Mathematical Audit —
- Graph Memory: Spectral Associative Memory via Dirichlet Energy —
- Neural Scaling Laws of Transformer Operator Network —
- From Knowing to Abstaining: Bridging the Representation-Action Gap in Vision-Language Models —
- SGA-Flow-GRPO: Spatial Gradient-Guided Credit Assignment for Flow-GRPO —
- Investigating the Effect of k-NN Preprocessing on Developing Graph Neural Networks: A Fairness-Based Perspective —
- Goal-Persistent Coding Agents as Scientific Performance Engineers: A Fixed-Radius Nearest-Neighbor Case Study —
- DraftAttention2: Fast Video Diffusion with Low-Resolution-Guided Mixed-Precision Attention —
- Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement —
- Calibration-Free Surface Normals Estimation in Vision-Based Tactile Sensing using Universal Photometric Stereo —
- Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement —
- How Much Imprecision is Enough Imprecision in my Classifier? A Practical Elicitation Procedure —
- A General Harness for Protein Foundation Model Fitness Prediction —
- VastMAT: A Large-Scale Multi-Category Benchmark for Multi-Animal Tracking —
- TRACE: Single-Pass Decoding-Trace Risk Localization for Generation Calibration —
- Beyond the Manifold Hypothesis: Hybrid Spectral Parameterizations for Flow Matching —
- WorldGraph: Graph-Native World Modeling —
- Mind the Spike: Mechanisms and Brittleness of Visual Massive Activations in Large Vision-Language Models —
- Right Answer, Wrong Reason: Accuracy, Consistency, and Consensus Are Misleading Indicators of LLM Faithfulness in Clinical Decision Support —
- Scaling Properties of Same-Family On-Policy Distillation —
- Measurement-Gated Provenance Attenuation for Frozen EEG Representations —
- On Privacy in Data-Space Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics —
- Bison: Cross-Dataset Learning for Unseen-Compound Perturbation Prediction —
- Evaluating Machine Unlearning in ASR —
- Tracing the Evolution of Oracle Bone Characters Across Three Millennia —
- GLF-Q: Global-Local Feature-based Quantization for Vision Transformers —
- Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models —
- PACE-FNO: Physics-Aligned Canonical Equivariance for Fourier Neural Operators —
- PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization —
- Enabling Timely Guidance before Skill Retrieval: Retaining Helpful Warm Tips in Agent Context —
- FuseAlign: Forced Alignment in the Wild —
- RPA: Residual Patch-Token Adapter for Image Retrieval from EEG and MEG —
- Evaluating Cell AI Foundation Models in Kidney Pathology with Human-in-the-Loop Enrichment —
- Measuring Collapse and Correction in Homogeneous-Panel LLM Debate —
- Telescopic Language Models —
- Business Compromise Detection with Agentic AI and LLM-driven Knowledge Discovery —
- MixDetect: Word-Level Localization and Quantification of AI Editing —
- Reducing Hallucination in Multimodal Large Language Models through Hard Grounding Preference Supervision —
- Language Distances are Practical for Equitable Cross-Lingual Transfer —
- Source Anchoring for Physical Consistency in Flow Matching Models —
- OSCC: Certified Observation-Safe Coupling Optimization for Gradient-Noise Control in Imperfect-Information Learning —
- The Impact of Stochasticity on the Rashomon Effect in Machine Learning —
- SoFT: Soft Targets for Generalizable LLM Fine-Tuning —
- Multi-level context Modeling for consistent expert selection in Mixture-of-Experts —
- Anatomy-Structured Hierarchical MIL for Weakly-Supervised Thoracic Disease Detection in Chest X-rays —
- How Reusable Are Benchmarks with Richer Feedback? —
- TANGO: Watermarking Masked Diffusion Language Models in Token Pairs —
- CityToolVQA: Tool-Augmented Visual Question Answering for 3D Spatial Cognition in Urban Low-Altitude Environments —
- ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining —
- A Function-Level Vulnerability Score Measures Flag Rate More Than the Model: Protocol Effects on Paired Benchmarks —
- One Sensor, Whole Body - 3D Body Pose from a Single Consumer Earbud IMU —
- From Phase Transition to Systemic Failure: A Decoupled Analytics Framework for GNN Robustness —
- Constraints Are Graphs, Not Chains: Exact Decoding for Diffusion Language Models —
- When to Commit and When to Defer: Maturing Markov Decision Processes under Refining Information and Expiring Opportunities —
- EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research —
- The Earth in One Gaze: Training-Free Active Focus for UHR Remote Sensing Understanding —
- FeCoSplat: Feedback-Guided Compression for Feed-Forward 3D Gaussian Splatting —
- Model Casting and Low-Parameter Gating: Towards More Sparsely Activated FFNs —
- Structure-Mapping-Guided Self-Explanation for Learning Mathematical Procedures —
- Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication —
- Memory as a cache: Exact context reuse and deletion by construction —
- Reachability is not enough: Diagnosing long-range behavior in GNNs —
- Architecture-aware Robustness Evaluation of Explainable Deep Learning for Breast Cancer Diagnosis —
- On the Relation Between Interval Regret and Dynamic Regret —
- The Key Handoff: Retrieval in Hybrid Language Models —
- nCMD: Benign-Anchored Feature Selection for Imbalanced Network Intrusion Detection —
- MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning —
- The Selection Rule Decides the Winner: A Pre-Registered Audit of Open-Set Graph Anomaly Detection —
- StarBOA: Real-Time Mamba State-Space Unrolling for Sparse Radar Micro-Doppler in ISAC Networks —
- Mind the Residual Gap: Probabilistic Downscaling under Real-World Bias —
- Spontaneous Context Restoration: How Language Models Recover from Corrupted Inputs —
- RINI: Seeing the Prior Is Not Enough —
- Chameleon: Dynamic Format Adapter for Efficient Diffusion —
- PainterBench: A Figural Divergent-Thinking Benchmark for Tool-Using Language Models —
- Action Shaping: Policies Absorb What They Can Express —
- GeoCR: Learning a Generalist Cloud Removal Prior from Heterogeneous Observations —
- Jev Matches 7B Language Models for Speech-Neuroprosthesis Rescoring —
- Replay in the Silent Degrees of Freedom: Continual Learning Without an Offline Phase —
- MedRouter: Demystifying Knowledge Differences Across Medical LLMs for Routing-Based Reasoning —
- Symmetry-quotient Flatness and Generalization —
- Ceiling of a Task: When Can a Transformer Succeed Without Its Chain of Thought? —
- Rondo: Unsupervised Discovery of Recurring Temporal Structure —
- Scale adaptive and robust intrinsic dimension estimation via optimal neighbourhood identification —
- Fast-Slow Thinking RM: Efficient Integration of Scalar and Generative Reward Models —
- Seeing and Solving Are Not Enough for Vision-Language Models —
- When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents —
- Structured Residual Connectivity Matters for Diffusion Transformers —
- Naturalness-guided Manifold Flow Matching for Sign Language Production —
- From Anomalies to Failures: Constructing Causal Error Graphs for Agentic Trace Diagnosis —
- Beyond Conservatism: Recoverability-Conditioned Exploration for Model-Based Imitation Learning —
- BITS: Rethinking Fair and Comprehensive Evaluation for Irregular Time Series Forecasting —
- When Should the Count Change? Learning State Maintenance for Causal Video Counting —
- Information Design Against Gaming and Learning Adversaries —
- Does Native 3D Texture Generation Necessarily Require 3D Assets for Training? —
- Resolution as a First-Class Decision: Task-Conditioned Routing for Efficient Multimodal Large Language Models —
- LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs —
- CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage —
- Feedback Makes Perfect: A Closed-Loop Framework for NL-to-STL Translation —
- Which Self-Improvements Should We Trust? Reliable Self-Improvement When Agents Reuse Their Benchmarks —
- Position Aware Layer Queries for Test Time Training in Vision Language Models —
- FocusDrive: Reasoning with Visual Focus for Autonomous Driving —
- CoLMbo-SV: A Grounded Language Model for Explainable Speaker Verification —
- In Two Minds about Lifelong Learning: Exploring Hemispheric Redundancy and Specialisation in Neural Models —
- Age-Adaptive Handwriting Reconstruction from an IMU-Based Digital Pen through Shared Representations and Domain-Specific Heads —
- Enhancing Foundation Models for Imbalanced SAR Ship Classification via Targeted Oversampling —
- LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models —
- Parser, Chunking, and Embedding Interactions in Retrieval-Augmented Generation over Indian Government Regulatory Documents —
- OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit —
- Beyond Memory Construction: Rethinking Memory Access for LLM-based Conversational Agents —
- SWE-Game: Can Coding Agents Build the Games We Want? —
- CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR —
- EgoTSR++: Egocentric Spatiotemporal Reasoning for Task Progress Understanding —
- When Evidence Changes the Subject: Subject-Typed Claim Licensing for Learned Routing —
- Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents —
- When Can First-Order Models of Fine-Tuning Bound Forgetting? —
- Can Tabular Foundation Models Amortize Statistical Inference? —
- MetaSampling: Making Frame Samplers Efficient for Long-Video Question Answering —
- AevaScenes: An FMCW LiDAR Dataset and Benchmark for Long-Range Perception —
- AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs —
- Policy Plasticity Matters in Offline-to-Online Reinforcement Learning: Refitting Offline Policies for Online Adaptation —
- DualGuard: Dual-Mode Quality Control for Logic-Preserving Data Augmentation —
- Optimal Transport Dropout for Structured Predictive Uncertainty —
- Background Gradients Shape Memorization in Flow Matching —
- Prior Directions: Separating Geometric Identification from Behavioral Efficacy in Visual Revision —
- Backdoor as Probe: Test-Time Adversarial Defense for CLIP —
- STAMP: Predicting Out-of-Distribution Generalization without Target Data —
- LVMT: Video Mask Transformer for Long-term Video Segmentation —
- In-Context Adaptation of Encoder-Decoder Models in Speech Recognition —
- An Imperfect Verifier is Good Enough: Learning with Noisy Rewards —
- Task-Aware Discretization of Differentiable Logic Gate Networks —
- GLOVE: Global Verifier for LLM Memory-Environment Realignment —
- CoeF-SFL: Preserving Collaborative Server-Client Learning with Enhanced Communication Efficiency —
- CLC-YOLO: A Compact Channel-Gated Prototype Network for Real-Time Leakage-Aware Breast Ultrasound Lesion Segmentation —
- LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models —
- Interactive Distributionally Robust Multi-Agent Learning with General Function Approximation —
- Epistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient Investigation —
- ReGDiff: Guided Diffusion in Regulated Latent Space for Exploring Metamaterial Voxel Geometry —
- LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles —
- Factorization Regret mediates compositional generalization in latent space —
- Multinomial Subset Routing with Sum-Max Rewards and Operational Constraints —
- HARMONIA: Interpretable Graph Learning through Mixtures of Neural Bases —
- SR4-Fit: A Unified Interpretable Rule-Based Machine Learning Framework for Informative and Trustworthy Decision-Making —
- Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents —
- HyperDAM: Hyperspectral Distractor-Aware Memory with Amodal Expansion for SAM 3 Tracking —
- Towards Multi-View Sign Language Understanding: A Benchmark Dataset and Baseline —
- TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception —
- Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning —
- Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL —
- Oracle-Efficient Online Classification with Stochastic Inputs and Adversarial Outputs —
- What Do Latent Predictive Vehicle Representations Retain? Measuring State, Geometry, and Local Response —
- Learning the Graph and the Embedding Together: Classifier-Independent Rewiring for Heterophilic Node Classification —
- HERO-MoE: Historical Expert Routing with Scale-Preserving Fusion —
- Compositional Objectives: Learning Structure in Structure —
- Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction —
- You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement —
- KV-Lingo: Learning KV-Cache Translators with Distillation —
- CAESAR: Clustering via Autonomous Embedding-Space Agglomerative Reorganization —
- Lagrangian and Hamiltonian Neural Networks With a Dissipative System —
- On-Policy Attention Linearization —
- High-Probability Guarantees for SGD under beta-Heavy-Tailed Gradient Noise —
- SAMBAR: Selective Anchoring via Method of Multipliers for Balanced Knowledge Acquisition and Retention in Vision-Language-Action Models —
- Playing to Par: Reinforcement Learning for Provably Optimal Quadrilateral Block Decompositions —
- FOCUS: Fixed-Confidence Online Causal Learning Using Sequential Adaptive Interventions —
- Efficient Support Recovery of Mixtures of Sparse Linear Classifiers with Less Measurements —
- Integrating Language Models into Listened and Imagined Speech Decoding from MEG —
- Depth Laws for the Precision Floor of Trained Neural Networks: Amplification, Residual Scaling, and a Quantization-Aware Training Paradox —
- Representation Learning for Exact Preimages —
- Can Circuit Alignment Predict OOD Generalization? —
- Mechanistic Interpretability Reveals Shared Causal Subspaces in Brain-to-Speech Decoders —
- A Unified Optimism-Agnostic Framework for Linear Bandits over Spherical Action Sets —
- Analytic-Walk Rotary Positional Encodings for Graphs —
- An Attention-Driven Heterogeneous GNN Model for Credit Card Fraud Detection —
- SCORENF: Score-based Normalizing Flows for Sampling Unnormalized distributions —
- Auditing Information Disclosure During Large-Scale Gradient-Based Training via Gradient Uniqueness —
- How Does Preconditioning Guide Feature Learning in Deep Neural Networks? —
- The Temporal Tug-of-War: Visualizing and Detecting RAG Conflicts in Diffusion Models via Trajectory Variance —
- Goal inference with Rao-Blackwellized Particle Filters —
- Neural Tractability via Structure: Learning-Augmented Algorithms for Graph Combinatorial Optimization —
- Action-Driven Processes for Continuous-Time Control —
- A Game-Theoretic Spatio-Temporal Reinforcement Learning Framework for Collaborative Public Resource Allocation —
- Ambiguous Online Learning —
- HERO: Preserving Structure and Semantics in Heterogeneous Continual Graph Learning —
- Regret Bounds for Robust Online Decision Making —
- When Pair Count Is Not the Sample Size: What All-Pairs Agent Comparisons Estimate —
- The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining —
- DynamicDx: Evaluating Evidence Acquisition in Video-Based Diagnosis —
- Certified Long-Horizon Code Agent Evolution via Validation-Gated Skill Optimization —
- Model-Aware Data Selection from In-and-Out Information Interplay —
- AgentHabit: Characterizing Distinct Behaviors of Agents on Everyday Tasks —
- The Decomposition Tax: LLM Pipelines Lose Up to 40 Accuracy Points at Their Own Interfaces —
- Adaptive Consistency Graph for Long-Horizon Agents —
- Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents —
- How Far Do Persona Effects Generalize in Language Models? —
- Improving LLM Collaboration via Multi-Agent Preference Learning —
- FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors —
- Not Every Term Adds New Structure: Sobolev Novelty for Symbolic Regression —
- The Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language Models —
- AnchorRep: Defending LLMs Against Cross-Model Adversarial Transfer via Representation Repulsion —
- MassAlloc Attention: Let Attention Allocate Its Own Compute —
- Feature Space Guidance for Breast Cancer Classification in DCE-MRI —
- CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering —
- LLM Alignment--Utility Asymmetry under Semantic-Preserving Transformations —
- Energy-aware frugal Bayesian optimization —
- Quantitative Measurement of Language Distance among Closely Related Indo-European Languages Using Pretrained Language Models: A Case Study on the North Germanic Branch —
- Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction —
- When Does Synergy Help Active Feature Acquisition? A PID-Based Study —
- Probabilistic Geodesic Flow Matching on Location-Scale Families —
- OpenMASC: An Open-Source Pipeline for Cross-Trajectory Metal-Aware Sampling and Correction in Accelerated MRI —
- HeroFrame-Bench: Reference-Anchored Evaluation via Rubric--Ranking Co-Evolution for Movie Hero Frame Selection —
- Graph Forward Distribution Matching for Molecular Inverse Design —
- ChronoFlow: Hierarchical Flow Matching for Irregular Time Series Generation —
- The Price of Locality: Why Forward-Forward Underperforms Backpropagation? —
- ControlGS: Conditioning Neural Gaussians for Downstream-Processing-Aware XR Rendering —
- The Trace Is the State: Exact Credit Assignment for LLM Agent Teams —
- Understanding the Subspace Stabilization of the Hessian and Gradient Covariance Matrix —
- Diffusion-Based Rollouts as a Stabilization Mechanism for Long-Horizon Environmental Forecasting —
- On the O(sqrt d over K 1/4) Convergence Rate of AdamW Measured by 1 Norm —
- CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies —
- Does Adversarial Training Improve Generalization in Multi-View VLAs? Revealing and Mitigating View Collapse —
- Generative Residual Factorization —
- Using LMs to Model the Effects of Context and Coreference during Sentence Comprehension —
- Cross-Architectural Mixture-of-Experts with Adaptive Soft Routing for Plant Leaf Disease Classification —
- Opening LLM Judges: Recovering Preference Signals Beyond the Final Verdict —
- Endo-TSR: Temporal Spectral Modeling of Appearance and Motion for Endoscopic Reconstruction —
- Gradient-Guided Decoupled Adaptation for Geospatial Vision-Language Models —
- PhysAlign: A Benchmark for Evidence-Grounded Role Alignment in Multimodal Physics Reasoning —
- SPACE: Sparse Predictive Attractor via Counterfactual Eviction for Streaming Video Memory —
- GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning —
- Statistical Testing for Multiple Instance Learning via Selective Inference with Applications to Computational Pathology —
- Look Before You Judge: Training-Free Region Mining for Grounded and Explainable Deepfake Detection —
- High-Capacity Robust Medical Image Exfiltration via Neural Network Weight Replacement —
- AI-driven ionic liquid discovery with unified chemical intelligence —
- Positions Are Not Facts: The Mismatch Between KV Caches and Memory —
- Cross-modal Translation via Conditional Latent Denoising for Video Deepfake Detection —
- Arbitrary-Accuracy Neural Approximation with Optimal Neuron Count and Near-Optimal Bit Complexity —
- Unifying Video Tasks via Spatiotemporal Analogy —
- SolveEdit: Benchmarking Visual Problem Solving in Generative Models —
- Spectral Reversal: Counteracting Singular Value Bias for Graph Prompting —
- ProDyGS: Dynamic Gaussian Splatting from a Single Static Monocular Camera —
- Deep Learning Methods in Neuroscience: From Modeling Molecular Mechanisms to Classifying States of Consciousness —
- Latency-Aware Client Assignment for Parallel Split Learning With Global Sampling —
- CarveMix-RC: Addressing Rare-Class Imbalance Through Lesion-Aware Synthetic Augmentation for Brain Metastasis Segmentation —
- VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation —
- Training Witnesses: Trusting the Training without Trusting the Trainer —
- DoAtlas-2: A Foundation for Self-Evolving Causal Biomedical Discovery —
- No Free Efficiency: Revisiting the Trade-off Between Training Efficiency and Model Vulnerability —
- WorldWeave: Growing Persistent Geometric Worlds for Video Generation —
- Scaling Versatile 3D Assets Editing with a Million-Scale Dataset —
- HUMAN-TCI: Hierarchical Multi-Stream Motion-Aware Network with Torso-Centered Interaction for Text-to-Motion Retrieval —
- ACPruner: Visual Token Pruning as Biased Attention Coverage Maximization in LVLMs —
- Temporal Modelling for Burn Scars on Sentinel-3 —
- XFlow: A Workflow Model for Instruction-Guided Lesion Segmentation in Chest X-rays —
- SubjectAnchor: Subject-Aware Memory-to-Video for Multi-Shot Storytelling —
- BMND: Direct Poisson Denoising by N-Dimensional Block Matching and Collaborative Filtering —
- Reinforcement Learning from Intermediate Renders for Image-to-Code Generation —
- RRG-SLAM: Real-time Reflection-aware Gaussian SLAM for Indoor Scenes —
- Long Time No See: Benchmarking VLMs for Out-of-Sight Spatiotemporal Reasoning in Egocentric Videos —
- TSGate: Timestep-Aware Gated Attention for Diffusion Transformers —
- Evidence Before Accuracy: A MRI-PET Fusion Network for Alzheimer Disease Classification with Causal Regional Validation —
- Distilling Visual Reasoning into Text Space —
- ConvCue: Complementary Visual Inductive Biases for Vision-Language Models —
- Uncovering Ordinal-Matching Bias in Audio-Visual LLMs —
- DecFlowEdit: Self-Localized Flow-based Image Editing via Guidance Decoupling —
- Over-Personalization Is a Decision Failure: Generation-Induced Apply Bias in LLMs —
- PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond? —
- Recursive LLM Degradation in Biomedical Question Answering: A Cross-Generation Study —
- SALMONN-duo: Adaptive Dual-System Coordination for Full-Duplex Voice Agents —
- MAS-OPD: On-Policy Distillation for Multi-agent Systems —
- Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models —
- Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge —
- When an Evaluation Rule Writes Training Labels: Measuring Human-Reference Forgiveness in NAVSIM —
- FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents —
- Certified Selective Automation of LLM Agent Evaluation —
- ControlScope: Workflow Revision and Reliability in LLM Agents —
- LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization —
- When Harness Beats Scale, and When Reading Beats Both —
- When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model —
- Loop Dropout: Regularizing Shared Updates in Looped Language Models —
- X-MoD: Practical Scaling Laws for Sparse-Depth Routing Beyond Mixture-of-Depths —
- DreamingGoose: Staged Distillation from Autoregressive Transformers to Bidirectional Recurrent Diffusion Language Models —
- Toward a Graded Measure of Belief Stability in Large Language Models —
- BIABench: Evaluating AI agents on real-world bioimage analysis tasks —
- RAGWarrant: Evidence-Preserving Governance for RAG Policy Promotion Under Quality, Cost, Latency, and Risk Constraints —
- Word Similarity Datasets for Indian Languages: Annotation and Baseline Systems —
- High-Level Text Preprocessing for Semantic Similarity Analysis of Discursive Texts: A Framework and Empirical Demonstration —
- SlopBench: How Well Can We Rank Language Models by Slop? A Multi-Domain Benchmark of Repetitive AI Writing —
- LLMs learn different forms of metacognition when trained to predict their own accuracy —
- Faster Block-Diffusion Serving with Distribution-Free Risk Guarantees —
- Lost with a Map: Conversational State and Behavioral Reliability in Language Models —
- Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State —
- What masking geometry works best for EEG foundation models? —
- MoGround: Measuring and Mitigating Modality Distraction in Vision-Language Models —
- A Free Knob: Decoupling Calibration and Predictive Skill in Threshold-Based Evaluation —
- How Synthetic Labels Improve Conformal Prediction: A Perspective on Conditional Coverage —
- FoldAttention: Declared-Reference Softmax for Fast Decode and Deterministic Backward —
- PDMD: Projected Distribution Matching Distillation for Video Diffusion Models —
- Discovering Symmetries in Neural Network Parameter Spaces —
- Pulseflow: PPG Counterfactual Generation Via Latent Transport —
- Does Execution Require Target KV Fidelity? A Mixed-Fidelity KV Runtime for LLM Serving —
- LLM4Trust: Exploring the Capabilities of Large Language Models for Trust Evaluation —
- The cost of useful natural gradient updates —
- Predicting Block-Coordinate Performance via Cross-Curvature —
- Geometric Identification in Predict-Then-Optimize Learning —
- Two Heads Are Better Than One: Aggregating Weaker LLMs for Better Forecasts —
- Counting on Thinking: Tracing Evidence Integration in Language Models —
- Efficient Dynamic Algorithms for Graph Neural Networks with Non-Linear Propagation —
- GTRL: Grounding Divide-and-Conquer Value Learning with Temporal Differences —
- Convergence of Practical Muon —
- Last-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPs —
- Mycelium: A Generalizable Cross-Grid Multi-Task Model for Electrical Distribution Systems —
- Optimal Nonparametric Dynamic Pricing with Censored Demand and Adversarial Inventory —
- Phenomenon-Graph JEPA: Label-Efficient Representation Learning for Contactless Cardiorespiratory Sensing —
- W2Rep: Learning Visual Representations by Watching the World Change —
- DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving —
- Preserve, Reveal, Expand: Towards Faithful 4D Video Editing with Region-Aware Conditioning and Benchmarking —
- EngIntervene: Benchmarking Multimodal Engineering State Understanding and Design Intervention Reasoning —
- Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning —
- VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation —
- C cubed ASD: Multi-Level Consistency-Driven Representation Learning for Robust Active Speaker Detection —
- BEAT: Rhythm-Elastic Alignment for Agentic Music-guided Movie Trailer Generation —
- Two-Step Occupation Coding —
- Safety-Constrained Cascade Inference for Robust Malaria Cell Classification Under Field Corruptions —
- ForkLeft: Entropy-First Rollouts for Prefix-Aligned Autoregressive-to-Diffusion Distillation —
- Low Latency Gaze Tracking via Latent Optical Sensing —
- G squared TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models —
- ShellfishNet: A Domain-Specific Benchmark for Visual Recognition of Marine Molluscs —
- Beyond Inpainting: Unleash 3D Understanding for Stable Camera-Controlled Video Generation —
- ABFR-KAN: Kolmogorov-Arnold Networks for Functional Brain Analysis —
- LEGATO: Large-scale End-to-end Generalizable Approach to Typeset OMR —
- Analog-Friendly Predictive Coding without Activation Derivatives —
- A Kernel-based Stochastic Approximation Framework for Nonlinear Operator Learning —
- Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It —
- Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers —
- Amnesia by Design, Memory By Necessity: Persistent State for Document Intelligence —
- DRIFT: Data Selection for LLM Instruction Tuning via On-Policy Attribution —
- What Should a Streaming Video Model Remember? —
- ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment —
- Adapting Vision-Language Models for Human-Readable XAI in Industrial Object Detection —
- DP-Rec: Towards Dynamic Patching for Efficient Long-Sequence Recommendation —
- Conditioned Direct Feedback Alignment via Activity and Error Geometry —
- EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic Memory —
- Typed Decision Models: An Early Evidence Audit and Evaluation Checklist —
- SWT: Self-Supervised Video Object Segmentation via Sliding, Wavelet and Transportation —
- Subtract the Corruption: Training-Data-Free Corrective Machine Unlearning using Task Arithmetic —
- Active Causal Discovery Benchmark: Evaluating LLM Agents Under Budgeted Interventions —
- GUIDE-FBO: Guidance via Uncertainty Intervention and Distributional Exchange for Federated Bayesian Optimization —
- Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning —
- Lagrangian--Hamiltonian Flows for Video Prediction and Image Generation: A Symplectic Perspective —
- Type-Balanced Federated Learning for Visual Analog Meter Reading —
- CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations —
- Domain-adaptive Zero-Shot Image Enhancement via Locality-Constrained Diffusion Guidance —
- AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors —
- JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments —
- Toward On-Chip Training of Spiking Neural Networks for Dense Event-Based Vision —
- Advancing Video-Text Pretraining with Multi-View Captions —
- Improving Test-Time Scaling with Adaptive Looped Transformers —
- UnStep: Training-Free Acceleration of Causal Video Diffusion with Fewer Steps Than Distillation —
- InfoEdit: Probing Global Layout Reasoning in Infographic Editing —
- When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning —
- QuanReview: Offline, Auditable Reconciliation of Human and LLM Span Annotations —
- Automated Species Identification in Camera Trap Images for Wildlife Conservation —
- When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety —
- How Well Can LLMs Simulate Real Learner Evaluations of Educational Feedback? —
- Which the Eye Fears: Writing with Read-Blindness Explains Massive Activations in Transformers —
- FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models —
- Who Is Left of Whom? Tracing Spatial Evidence and Role Binding in Relative-Position Reasoning —
- A mechanistic study of language model introspection —
- Equivariant Neural Primal-Dual Assignment for Maximum Common Edge Subgraphs —
- Age of Learning: Temporal Persistence of Prediction Errors as a Learning Signal —
- Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent —
- Gradients for Interventions and Activations for Detection: Targeted Feature Learning in Language Models —
- Sliced Orlicz-Wasserstein —
- Understanding Clinical Cognitive Dialogues Using Large Language Models —
- ScreenHaystack: Finding Blind Zones in GUI Grounding —
- See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs —
- OpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin Vocalizations —
- Zero-Storage Procedural Neural Synthesis via Boundary Dynamics: Formal Verification in Lean 4 and Bare-Metal Gauntlet Validation —
- Transfer Learning for Edge Classification on Dynamic Text-Attributed Graphs —
- Tsubame: Tree Replay for Diffusion-Based Speculative Decoding —
- Understanding and Mitigating Under-Confidence in GNNs from the Final Layer —
- Beyond Prompt or Skill? Attribution-Guided Optimization of Modular LLM Programs —
- PluginRSI: Recursive Improvement of Agent Harnesses with Reusable Plugins —
- Why Directly Learning Periodic Trajectories Can Fail —
- Clinical Trajectory Alignment for Medical Vision-Language Pre-training —
- RLHarness: Co-evolving Procedural Skills with Reinforcement Learning for Long-horizon Multimodal Reasoning —
- DF-CBM: Region-Aware Concept Bottleneck Models for Deepfake Detection —
Important terms
- AmbiModBench
- A benchmark designed to test gene perturbation prediction methods by evaluating their accuracy beyond simple shared response patterns, focusing on nuanced biological modeling.
- Koopman-Based Generative Model
- A generative model used to learn single-cell dynamics from distribution snapshots, enabling the exploration of complex dynamic system learning in biology.
- Structured State-Space Models
- Models that investigate the primacy effect in systems, suggesting initial conditions heavily influence later dynamics, alongside geometric alignment techniques.
- Neuro-symbolic Concepts
- The integration of neural networks with explicit symbolic reasoning to build intelligent agents, combining pattern recognition with structured logical thinking.
- CausalReasoningBenchmark
- A real-world benchmark created to disentangle the evaluation of causal identification and estimation in complex data sets.