AI papers — 2026-09-30
Today's focus centers on exploring how we can better estimate the causal effects related to T cell receptors by looking at several distinct approaches to modeling and scoring these complex biological interactions. One line of inquiry involved Mirror-Score, which aimed to expose the limitations of sequence-compatibility ranking in D-peptide design by using calibrated, inference-only scoring. This suggests a move away from simple ranking towards methods that better capture the underlying compatibility dynamics.
Simultaneously, we are examining Scalable Diffusion SBI for compositional inference under simulator misspecification. This tackles the challenge of making reliable predictions when the underlying simulation models are imperfect. Another piece of work, AssayRouter, introduced historical utility priors to guide frozen molecular predictor routing. This suggests a way to leverage past experimental data to inform future predictions. These efforts are connected by broader themes in benchmarking graph-based models for in-silico toxicity prediction and OmniVCBench, which benchmarks evidence-grounded multimodal reasoning towards AI virtual cells.
Furthermore, the work on explainability from training with applications to TCR-epitope prediction shows an attempt to make these complex models more transparent. Finally, LEMON-ZEST addresses efficiency in protein language modeling through evolution-informed tokenization.
The investigation into fast weight programming and linear transformers explored how to optimize these models, moving from machine learning applications toward neurobiology. Specifically, the work examined methods for traversing the solution space of neural networks using Hessian Null Space Continuation to understand the dynamics of these large architectures. This approach suggests a pathway for better control or understanding within complex models.
Simultaneously, research into reward valuation in large language models focused on causal induction of anhedonia. This delves into how these models assign value and what it means when they fail to do so in a way that mirrors human experience. Furthermore, the study on large-scale factor analysis indicated that machine intelligence remains only partially interpretable. This highlights a fundamental limitation in understanding the internal workings of these systems.
This complexity is further addressed by neural structural reasoners, which propose a brain-inspired architecture designed for reasoning over structured knowledge. Complementing this structural focus is flattening the connectome spectrum through a spectral filter applied to functional connectivity. This was shown to induce a pretraining target for fMRI encoders. These studies collectively suggest that while we are making strides in programming and interpreting these models, the full picture of their internal logic and valuation remains elusive.
The recent work focused on developing MonitorBench, which serves as a comprehensive benchmark specifically designed to assess the chain-of-thought monitorability of large language models. This effort directly addresses the need to measure some aspects of artificial intelligence metacognition by creating a standardized way to probe how these models reason. The study involved setting up scenarios where the model's internal reasoning steps could be explicitly monitored. The results from this benchmark are crucial for understanding what is currently possible in terms of AI self-regulation.
Furthermore, Meta-TTL explored meta-learning self-improvement policies for language agents. This suggests a path toward autonomous refinement of agent behavior. This concept connects to the work on optimal skill selection for LLM agents with provable bicriteria guarantees. It implies that future agent design might incorporate these learned improvement policies to select the most effective skills under specific constraints. The findings from MonitorBench and Meta-TTL collectively suggest that while current models excel at complex reasoning, there is a significant gap in systematically evaluating their ability to monitor and improve their own processes. What remains open is how to translate these benchmark results into practical, scalable methods for implementing self-improving policies within deployed language agents.
The work on text classification has shown varied approaches to tackling complex language understanding problems. One line of inquiry involved developing MONOVAB, which was an annotated corpus specifically designed for Bangla multi-label emotion detection. This suggests an effort to build richer datasets for sentiment analysis in that language. Simultaneously, there was a focus on model efficiency and adaptation techniques with PRILoRA. This introduced pruned and rank-increasing low-rank adaptation methods aimed at improving the performance or reducing the size of models.
Further advancements in large language models were explored through dynamic optimizations of LLM ensembles utilizing two-stage reinforcement learning agents to manage their behavior. In a different domain, entity linking using LLMs was applied to estimate product carbon footprints. This demonstrated an application of these models beyond pure text categorization. Other studies looked at predicting team performance from communications within simulated search-and-rescue scenarios. There was also work on Thunder-NUBench which established a benchmark for assessing sentence-level negation understanding in LLMs. These efforts collectively point toward a multifaceted research landscape where researchers are concurrently improving dataset quality, refining model architectures through pruning and adaptation, exploring dynamic ensemble management, and applying language models to highly specialized tasks like carbon footprint estimation and performance prediction.
The work on generating questions from solutions to improve inference time reasoning in large language models suggests that providing the model with a question derived from a correct answer can enhance its ability to reason during the inference phase. This approach, which involves using solution-to-question generation as an intervention, implies a mechanism for guiding the model's internal thought process toward more efficient logical paths.
Furthermore, research into causal separation of sycophantic behaviors in large language models indicates that these behaviors are not monolithic. Instead, they can be causally separated into distinct components. This suggests that targeted interventions might be possible to manage or mitigate specific forms of unhelpful model behavior. This contrasts with work exploring how temporal biases shape retrieval in transformer and state-space models, where the focus shifts to understanding how time-dependent factors influence information access.
Another line of inquiry examines tag guided process supervision for personalization reasoning. This suggests imposing structured guidance during the processing steps can refine how models handle personalized tasks. These explorations collectively point toward a multifaceted approach where interventions range from structural guidance on reasoning paths to disentangling complex behavioral patterns and managing temporal dependencies in model operation.
The work on Block Sparse Flash Attention explored a method to improve attention mechanisms by reducing the computational cost associated with large matrices. This suggests that this technique could lead to more efficient processing in transformer models. This contrasts with the efforts in Thunder-KoNUBench, which focused on creating a corpus-aligned benchmark specifically for understanding Korean negation. This aimed to test how well models grasp complex linguistic structures within that specific language domain.
Simultaneously, research into Asymptotic Universal Alignment introduced a novel alignment framework that utilizes test-time scaling to achieve better performance across various tasks. This aligns conceptually with the work on evaluating Test-Time Scaling of General LLM Agents. This investigates the efficacy of scaling techniques when applied to general agent behavior. Furthermore, investigations into Evaluating Alignment of Behavioral Dispositions in LLMs and On Calibration of Large Language Models: From Response To Capability delve into understanding how well these models align their outputs with desired behaviors and how their actual capabilities scale relative to their responses. These studies suggest that while specific architectural optimizations like Block Sparse Flash Attention offer efficiency gains, the broader challenge lies in developing robust alignment frameworks and calibration methods that ensure LLMs perform reliably across diverse tasks.
The work on SiDiaC version 2.0 focused on building a Sinhala diachronic corpus. This implies an effort to capture historical linguistic changes within that language for training models. This is complemented by RAWR, which explored reward assignment without rollouts in verifiable domains. This suggests a focus on efficient learning signals where direct outcome evaluation is difficult.
Simultaneously, the screening studies suggested that simple screening methods might be sufficient for certain tasks. This perhaps reduces the need for overly complex evaluation pipelines. The capacity gap revisited in chain-of-thought distillation points toward a practical challenge when trying to transfer reasoning capabilities from larger models to smaller ones. Furthermore, the work on think multilingual not harder addresses teaching reasoning models to code-switch by using a data-efficient framework. This indicates an attempt to handle language mixing effectively without massive datasets.
GRAVITY introduces an architecture-agnostic method for structured anchoring of long-horizon conversational memory. Relative kinetic utility aims to calibrate cross-layer credit for pruning global structured LLMs. These suggest ongoing efforts in optimizing model structure and memory management through nuanced credit assignment.
The work on PRISM suggests an attempt to establish a geometric risk bound for decomposing drift into scale, shape, and head. This approach builds upon the broader effort to decompose and measure evaluation awareness. Simultaneously, KSAFE-MM introduced a multimodal safety benchmark specifically localized for Korean cultural risks through localized contextualization.
Furthermore, diversification of reinforcement learning value rollouts was explored by using first-token exploration as a method to diversify the rollout strategies. In terms of reasoning, research into cutting at decision points with sampling was conducted to understand how agents make choices. The evaluation of cross-lingual knowledge consistency in code-mixed Indian languages using IndicKLAR provided insights into language variation within models. Finally, efforts were made to determine if proactive agents require a large language model to decide when it is appropriate to act. These varied investigations collectively point toward the ongoing challenge of robustly quantifying risk and ensuring reliable decision-making across different domains and linguistic contexts.
Today's papers
- Estimating the Causal Effects of T Cell Receptors This paper estimates how changing T cell receptors affects their function. [paper]
- Mirror-Score Calibrated Inference-only Scoring Exposes the Limits of Sequence-compatibility Ranking in D-peptide Design This work shows that a specific scoring method reveals weaknesses in ranking sequences for D-peptide design. [paper]
- Scalable Diffusion SBI for Compositional Inference under Simulator Misspecification This paper proposes a scalable diffusion model to infer chemical compositions even when the simulator is imperfect. [paper]
- AssayRouter Historical Utility Priors for Frozen Molecular Predictor Routing This method uses past experimental data to guide which frozen molecular predictors should be used in future predictions. [paper]
- Benchmarking graph-based models for in-silico toxicity prediction in drug discovery This paper tests different graph-based models to see which ones are best at predicting toxicity online. [paper]
- OmniVCBench Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells This benchmark evaluates how well multimodal reasoning agents can reason about virtual cells using evidence. [paper]
- Explainability from Training with Applications to TCR-Epitope Prediction This research looks at how to explain a model's predictions when applied to predicting T cell receptor epitopes. [paper]
- LEMON-ZEST Evolution-Informed Tokenization for Efficient Protein Language Modeling This paper suggests a new way to tokenize proteins based on evolutionary information for better language modeling. [paper]
- Fast weight programming and linear transformers from machine learning to neurobiology This work connects fast weight programming in deep learning with the principles of neurobiology. [paper] [episode]
- Reward Valuation in Large Language Models: Causal Induction of Anhedonia This study investigates how causal induction can be used to understand anhedonia in large language models through reward valuation. [paper] [episode]
- Large-scale factor analysis shows machine intelligence is only partially interpretable This analysis suggests that large-scale factor analysis reveals that machine intelligence remains only partially interpretable. [paper]
- Neural Structural Reasoner A Brain-inspired Architecture for Reasoning over Structured Knowledge This paper introduces a brain-like architecture designed to reason over structured knowledge. [paper]
- Flattening the Connectome Spectrum: A Spectral Filter for FC Induces a Pretraining Target for fMRI Encoders This technique uses spectral filtering of functional connectivity to create a pretraining target for fMRI encoders. [paper]
- Which Attention Heads are like the Human Head? Not the Ones that Compute This paper examines which attention heads in neural networks mimic human-like behavior rather than just performing computation. [paper]
- Traversing the solution space of neural networks with Hessian Null Space Continuation This method uses continuation in the Hessian null space to explore all possible solutions within a neural network's solution space. [paper]
- Poly-attention a general scheme for higher-order self-attention This paper proposes a general framework for implementing higher-order self-attention mechanisms. [paper] [episode]
- MonitorBench A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models This benchmark assesses how well large language models can be monitored during chain of thought reasoning. [paper] [episode]
- Measuring (some aspects of) the metacognition of AI This research attempts to measure certain aspects of meta cognition in artificial intelligence systems. [paper] [episode]
- Meta-TTL Meta-Learning Self-Improvement Policies for Language Agents This paper develops meta-learning policies that allow language agents to self-improve over time. [paper] [episode]
- When Does Equivariance Help? Canonical Alignment in Neural Fluid Surrogates This work explores when equivariance helps in canonical alignment for neural fluid surrogates. [paper] [episode]
- Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees This paper finds the best way to select skills for LLM agents using provable guarantees across two criteria. [paper] [episode]
- Edge Selection for the Effective use of Piecewise-Constant Distributions as Neural Network Outputs for Event Prediction This method identifies the most useful edges in piecewise constant distributions when using them as outputs for event prediction. [paper] [episode]
- Cold-Start Active Preference Learning in Socio-Economic Domains This paper focuses on learning active preferences quickly in socio-economic areas where data is scarce. [paper] [episode]
- Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics This paper provides an exact solution to linear attention mechanisms derived from continuous-time dynamics. [paper] [episode]
- Are We Really Making Much Progress in Text Classification? A Comparative Review This review compares the current state of progress in text classification tasks across various methods. [paper] [episode]
- MONOVAB An Annotated Corpus for Bangla Multi-label Emotion Detection This paper introduces a labeled dataset specifically for multi-label emotion detection in the Bengali language. [paper]
- PRILoRA Pruned and Rank-Increasing Low-Rank Adaptation This technique prunes and ranks low-rank adaptations to improve model performance efficiently. [paper] [episode]
- Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents This paper uses two stages of reinforcement learning agents for dynamic optimization of large language model ensembles. [paper] [episode]
- Entity Linking using LLMs for Automated Product Carbon Footprint Estimation This work uses large language models to link entities and estimate the carbon footprint of products automatically. [paper] [episode]
- Choices Speak Louder than Questions This paper explores how the choices made by an agent can reveal more about its reasoning than its questions. [paper] [episode]
- Predicting Team Performance from Communications in Simulated Search-and-Rescue This study predicts team performance based on communications during simulated search-and-rescue missions. [paper] [episode]
- Thunder-NUBench A Benchmark for LLMs' Sentence-Level Negation Understanding This benchmark tests large language models' ability to understand sentence negation at the sentence level. [paper] [episode]
- CCQA Generating Question from Solution Can Improve Inference-Time Reasoning in SLMs This paper shows that generating a question from a solution can boost reasoning during inference in small language models. [paper] [episode]
- Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs This research causally separates different types of sycophantic behaviors observed in large language models. [paper] [episode]
- TagPR Tag-Guided Process Supervision for Personalization Reasoning in Large Language Models This method uses tags to guide the process supervision for better personalization reasoning in large language models. [paper] [episode]
- Inducing Dyslexia in Vision Language Models This paper shows how to induce dyslexia characteristics within vision language models. [paper] [episode]
- POET Preference Optimization for Enhanced Text-to-Image Generation This paper uses preference optimization to enhance the quality of text-to-image generation. [paper] [episode]
- Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models This research investigates how temporal biases affect information retrieval in transformer and state-space models. [paper] [episode]
- The Collective Turing Test: Large Language Models Can Generate Realistic Multi-User Discussions This paper demonstrates that large language models can generate realistic discussions among multiple users. [paper] [episode]
- Evidence-Guided Schema Normalization for Temporal Tabular Reasoning This method uses evidence to normalize schemas for better temporal tabular reasoning. [paper] [episode]
- Block Sparse Flash Attention This paper proposes a technique called block sparse flash attention to improve the efficiency of attention mechanisms. [paper] [episode]
- Thunder-KoNUBench A Corpus-Aligned Benchmark for Korean Negation Understanding This benchmark tests large language models' negation understanding specifically on Korean text. [paper] [episode]
- Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling This framework proposes a new alignment method using test-time scaling for asymptotic universal alignment. [paper] [episode]
- IESR Efficient MCTS-Based Modular Reasoning for Text-to-SQL with Large Language Models This paper presents an efficient method using Monte Carlo Tree Search for modular reasoning in text-to-SQL tasks with LLMs. [paper] [episode]
- Evaluating Alignment of Behavioral Dispositions in LLMs This research evaluates how well the behavioral dispositions of large language models are aligned. [paper] [episode]
- On Calibration of Large Language Models: From Response To Capability This paper examines the calibration between the model's response and its actual capability. [paper] [episode]
- Vision Wormhole Latent-Space Communication in Heterogeneous Multi-Agent Systems This work explores latent space communication for heterogeneous multi-agent systems using a vision wormhole concept. [paper] [episode]
- Evaluating Test-Time Scaling of General LLM Agents This paper evaluates how test-time scaling affects the performance of general large language model agents. [paper] [episode]
- SiDiaC-v.2.0 Sinhala Diachronic Corpus Version 2.0 This is an updated version of a diachronic corpus specifically for the Sinhala language. [paper] [episode]
- RAWR Reward Assignment Without Rollouts in Verifiable Domains This method assigns rewards without needing to perform full rollouts in verifiable domains. [paper] [episode]
- Screening Is Enough This paper argues that screening methods are sufficient for certain tasks without needing more complex methods. [paper] [episode]
- Revisiting the Capacity Gap in Chain-of-Thought Distillation from a Practical Perspective This paper re-examines the capacity gap when distilling chain-of-thought reasoning practically. [paper] [episode]
- Think Multilingual, Not Harder: A Data-Efficient Framework for Teaching Reasoning Models to Code-Switch This framework provides a data efficient way to teach models to switch between languages for coding.
- GRAVITY Architecture-Agnostic Structured Anchoring for Long-Horizon Conversational Memory This architecture uses structured anchoring for long conversational memory regardless of the underlying architecture. [paper] [episode]
- Relative Kinetic Utility Calibrating Cross-Layer Credit for Global Structured LLM Pruning This method calibrates credit across layers to prune large language models efficiently based on relative kinetic utility. [paper] [episode]
- PowerStep Memory-Efficient Adaptive Optimization via p-Norm Steepest Descent This technique uses p-norm steepest descent for memory efficient adaptive optimization. [paper] [episode]
- PRISM A Geometric Risk Bound for Decomposing Drift into Scale, Shape, and Head This paper provides a geometric risk bound to decompose drift into scale, shape, and head components. [paper] [episode]
- Decomposing and Measuring Evaluation Awareness This research focuses on decomposing and measuring the awareness of evaluation in models. [paper] [episode]
- KSafe-MM A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks This benchmark assesses multimodal safety by localizing cultural risks specifically for Korean contexts. [paper] [episode]
- Diversifying RLVR Rollouts via First-Token Exploration This method diversifies reinforcement learning rollouts by exploring based on the first token. [paper] [episode]
The papers
- Phaedra: Learning High-Fidelity Discrete Tokenization for the Physical Science — As an excellent, fastidious, and diligent researcher, I have meticulously analyzed both provided texts regarding "Phaedra." The first text is a detailed abstract/summary of Phaedra's methodology and results for scientific image tokenization (PDE data), while the second text appea [episode]
- EVolSplat4D: Efficient Volume-based Gaussian Splatting for 4D Urban Scene Synthesis — EVolSplat4D proposes a novel feed-forward 3D Gaussian Splatting framework designed for efficient and consistent reconstruction of both static and dynamic urban scenes in 4D. [episode]
- PhysPlan: Grounded Physical State Reasoning and Graph-Guided Optimization for Physically Plausible Video Generation — Video diffusion models (VDMs) have shown remarkable success in generating high-fidelity video content, but they fundamentally lack an intrinsic understanding of physical laws, often resulting in causally illogical sequences and structural hallucinations. [episode]
- CST-WM: A Causally Structured World Model for Embodied Visual Tracking — Embodied visual tracking requires robots to make predictive decisions over future target observability and apparent scale, especially when dealing with occlusion or distractors. [episode]
- MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models — As a fastidious and diligent AI researcher, I will synthesize these excerpts into a comprehensive, detailed summary of the MonitorBench benchmark paper. Precision is paramount; every detail must be accounted for to ensure no critical nuance is lost. [episode]
- COMiT: Learning Structured Visual Tokens through Sequential Communication — Discrete image tokenizers are crucial for modern vision systems, but existing methods often optimize for reconstruction and compression, yielding tokens that capture local texture rather than object-level semantic structure. [episode]
- Mitigating Cross-Image Information Leakage in Multi-Image Understanding with Large Vision-Language Models — Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks, but their capabilities degrade significantly when handling multi-image inputs due to a phenomenon termed cross-image information leakage. [episode]
- Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees — Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees provides a principled framework for selecting reusable skills into an LLM agent's context window by formulating it as a regularized submodular maximization problem under a hard token budget, yielding the f [episode]
- From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation — This paper introduces Policy Iteration with Human Feedback (PIHF), a framework designed to bring post-training reinforcement learning principles to in-context learning systems, specifically for complex tasks like rare disease diagnosis. [episode]
- SimWAM: A Simple World Action Model for End-to-End Autonomous Driving — SimWAM presents a novel World-Action Model (WAM) designed to improve end-to-end autonomous driving by leveraging future video prediction as a training signal, thereby avoiding costly test-time future imagination. [episode]
- Evaluating Test-Time Scaling of General LLM Agents — LLM agents are increasingly expected to function as general-purpose systems capable of resolving open-ended user requests, necessitating new evaluation methods that move beyond domain-specific settings to assess their performance across multiple skills and tools. [episode]
- Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory — Matrix-Game 3.0 is a memory-augmented interactive world model designed for 720p real-time longform video generation, addressing the critical challenge of simultaneously achieving memory-enabled long-term temporal consistency and high-resolution real-time performance. [episode]
- MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing — MoCA-Video presents a training-free framework for semantic mixing in videos, operating within the latent space of frozen video diffusion models to enable controllable and high-quality video editing under semantic shifts. [episode]
- Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs — Large language models often exhibit sycophantic behaviors, but it is unclear whether these behaviors arise from a single mechanism or multiple distinct processes. [episode]
- TagPR: Tag-Guided Process Supervision for Personalization Reasoning in Large Language Models — TagPR introduces a novel training framework that significantly enhances Large Language Models' intrinsic capacity for personalization reasoning by forcing them to externalize their logic into structured, interpretable steps. [episode]
- Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems — Multi-Agent Systems powered by Large Language Models are currently bottlenecked by discrete text communication, which imposes runtime overhead and information quantization loss. [episode]
- On Calibration of Large Language Models: From Response To Capability — Large language models are widely deployed as general-purpose problem solvers, making accurate confidence estimation critical for reliable use, and this work introduces capability calibration to address its limitations. [episode]
- CPATTA: Conformal Supervision Allocation For Active Test-Time Adaptation — Active Test-Time Adaptation (ATTA) aims to improve model robustness under domain shift by selectively querying human annotations during deployment, but existing methods suffer from low data selection efficiency, wasting human annotation budget. [episode]
- ProCompNav: Proactive Instance Navigation with Comparative Judgment for Ambiguous User Queries — ProCompNav is a two-stage framework that first constructs a candidate pool and then identifies the target through comparative judgment, replacing independent matching strategy with a two-stage collect-then-compare pipeline: candidate pool construction and candidate pool pruning t [episode]
- Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping — Multimodal large language models (MLLMs) often struggle with fine-grained perceptual grounding, frequently missing small details and spatial relationships in cluttered scenes, which leads to errors in complex reasoning tasks. [episode]
- Formalizing the Sampling Design Space of Diffusion-Based Generative Models via Adaptive Solvers and Wasserstein-Bounded Timesteps — Diffusion-based generative models are powerful tools for high-fidelity synthesis, yet their practical deployment is often hindered by high sampling costs due to static heuristics governing solver selection and scheduling. [episode]
- Relative Kinetic Utility: Calibrating Cross-Layer Credit for Global Structured LLM Pruning — Chain-of-Thought (CoT) prompting has significantly improved Large Language Models' reasoning capabilities, but scaling test-time computation introduces severe inference latency and KV cache bottlenecks. [episode]
- NIV: Neural Axis Variations for Variable Font Generation — NIV (Neural Axis Variations) introduces a neural method that automatically converts static fonts into fully functional variable fonts by predicting per-point displacements conditioned on desired design axes. [episode]
- Reasoning with Sampling: Cutting at Decision Points — Frontier reasoning models are produced by posttraining base language models with reinforcement learning, and this work introduces Entropy-Cut Metropolis–Hastings as a training-free method to sample from a sharpened version of the base model’s distribution, showing that it eli [episode]
- Do Proactive Agents Need an LLM to Decide When to Act? — This research addresses a critical architectural challenge in building proactive agents that monitor user activity: how to efficiently decide *when* to invoke a powerful Large Language Model (LLM) for action, without incurring the high cost and latency of constant LLM calls. [episode]
- Certified Adaptive Refresh: Anytime-Valid Monitoring for Federated Conformal RAG — Federated Conformal RAG (FC-RAG) provides distribution-free coverage guarantees for weak language model swarms under bandwidth constraints, but only for a fixed horizon. [episode]
- KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks — Multimodal Large Language Models (MLLMs) introduce significant safety risks by combining language and vision modalities, necessitating evaluation tools that are culturally grounded rather than generic. [episode]
- Optimization Risk Bounds for Kolmogorov-Arnold Networks Trained by DP-SGD with Correlated Noise — As a fastidious and diligent AI researcher, I have meticulously analyzed both provided texts. [episode]
- FactorizedHMR: A Hybrid Framework for Video Human Mesh Recovery — Human Mesh Recovery (HMR) is fundamentally ambiguous, meaning multiple 3D bodies can explain the same visual evidence under occlusion or weak depth cues. [episode]
- Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution — Video world models often fail to maintain evolving states when evidence is unobserved, leading to frozen or implausible resets upon re-observation. [episode]
- From Concept Erasure to Style Purification: Contrastive Eigenbases for Artist Style Protection — The DICE framework is a training-free, inference-time method designed for on-the-fly artist style erasure in diffusion models to combat style mimicry and protect intellectual property. [episode]
- IESR:Efficient MCTS-Based Modular Reasoning for Text-to-SQL with Large Language Models — IESR proposes a modular reasoning framework that integrates information understanding, MCTS-based search, and trajectory verification to achieve state-of-the-art performance on complex Text-to-SQL tasks by decoupling mathematical computation from SQL generation. [episode]
- Persistent Tri-State Message Passing — Persistent Tri-State Message Passing (SpaM) is a framework designed to address the challenges of structural uncertainty, edge noise, and heterophily in semi-supervised learning on real-world graphs. [episode]
- TeD-Loc: Text Distillation for Weakly Supervised Object Localization — Weakly supervised object localization (WSOL) models trained on image-level class labels are limited because traditional methods often focus only on discriminative regions, missing the full spatial extent of objects. [episode]
- Query Lower Bounds for Diffusion Sampling — This work establishes the first information-theoretic lower bounds for diffusion sampling, proving that acquiring a nontrivial sample from high-dimensional distributions using polynomial accuracy score estimates requires an intrinsic barrier on query complexity. [episode]
- Block Sparse Flash Attention — Block Sparse FlashAttention (BSFA) is a training-free method that accelerates long-context prefill inference by computing exact query-key scores before selectively processing value blocks. [episode]
- Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models — In-context learning in Large Language Models (LLMs) is governed by both temporal and semantic relationships, shaping how they retrieve contextual information. [episode]
- AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification — AdvancedMathBench is introduced as a new benchmark suite designed to rigorously evaluate Large Language Models (LLMs) on advanced mathematical proof generation and verification, addressing the current gap where existing benchmarks rely too heavily on final answers or coarse judgm [episode]
- Beacon: Knowing When and How to Perform Agentic Visual Reasoning — Beacon introduces a novel agentic visual reasoning model designed to improve performance on complex tasks by focusing on two critical dimensions of tool use: Mode Adaptiveness and Tool Effect. [episode]
- Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding — As a meticulous AI researcher, I have thoroughly reviewed the provided excerpts from the paper "Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding." Below is a comprehensive and detailed synthesis of the work, designed to capture all critical aspects of [episode]
- FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Self-Verification — Agentic vision-language models (VLMs) have shown great potential for multimodal reasoning by interleaving textual reasoning with explicit tool calls for image manipulation, yet these models frequently use tools unfaithfully, often employing decorative or irrelevant process images [episode]
- UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs — As a fastidious researcher, I must first state that I have analyzed both provided inputs. [episode]
- Diversifying RLVR Rollouts via First-Token Exploration — Reinforcement Learning with Verifiable Rewards (RLVR) often suffers from a bottleneck in rollout diversity, which this paper addresses by introducing an intervention that exploits low-load, high-leverage positions in the reasoning trace. [episode]
- AdvantageFlow: Regularized Advantage-Weighted RL in Flow Models — AdvantageFlow introduces a forward-process reinforcement learning algorithm, AdvantageFlow, specifically designed for rectified flow models. [episode]
- Decomposing and Measuring Evaluation Awareness — The paper moves beyond simple behavioral observation to propose a rigorous, psychology-grounded framework for understanding how models detect when they are being tested and how that detection influences their subsequent actions. [episode]
- Conformal Prediction via Transported Beta Laws — This paper introduces a novel framework for analyzing split conformal prediction by focusing on the law of realized calibration-conditional coverage rather than just its marginal guarantee. [episode]
- BridgeMatch: Conditional Transport Bridges in Matching Matrix Space for 3D Deformable Registration — BridgeMatch is a novel two-stage generative solver designed to estimate reliable, non-rigid point cloud correspondences by maintaining and refining the complete soft matching matrix across different resolutions. [episode]
- Screening Is Enough — A core limitation of standard softmax attention is that it does not provide an independently interpretable measure of query–key relevance, while Multiscreen introduces an architecture built around a mechanism called screening that enables absolute query–key relevance. [episode]
- CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMs — Cycle-Consistency in Question Answering (CCQA) is a novel reasoning method designed to improve inference-time reasoning accuracy for Small Language Models (SLMs) by generating and comparing questions derived from each candidate solution. [episode]
- Fisher-Guided Progressive Parameter Selection for Adaptive Fine-Tuning — FisherAdapTune is a Fisher-guided Adaptive Fine-Tuning framework designed to dynamically select which parameter groups should remain trainable during fine-tuning by tracking the temporal drift of their Fisher geometry. [episode]
- Suppression Is Not Forgetting: Residual Recoverability in Visual Concept Unlearning for VLMs — This paper introduces VLM-UnBench, the first benchmark designed to rigorously evaluate training-free visual concept unlearning in Vision-Language Models (VLMs). [episode]
- ActiveSAM: Fast and Accurate Open-Vocabulary Semantic Segmentation with Frozen SAM 3 — ActiveSAM introduces a training-free, zero-shot inference framework that adapts the frozen SAM 3 model into an active-vocabulary segmenter for open-vocabulary semantic segmentation (OVSS). [episode]
- Empowering Microscopic Traffic Simulators with Realistic Perception using Surrogate Sensor Models — This paper introduces MIDAR, a novel surrogate LiDAR detection model designed to bridge the gap between scalability and perception realism in simulating intelligent transportation systems (ITS). [episode]
- Fast weight programming and linear transformers: from machine learning to neurobiology — As a fastidious researcher, I have meticulously analyzed both provided texts. [episode]
- Robust Promptable Video Object Segmentation — Promptable video object segmentation (PVOS) models suffer substantial performance degradation when faced with input corruptions, which severely limits their deployment in safety-critical domains such as autonomous vehicles and robotics. [episode]
- Inducing Dyslexia in Vision Language Models — Dyslexia, a neurodevelopmental disorder characterized by persistent reading difficulties, has been modeled using large-scale vision-language models (VLMs) to functionally identify and perturb artificial analogues of word processing. [episode]
- Does Anthropomorphic Language Impact Public Perceptions of AI? — As a fastidious and diligent researcher, I have carefully analyzed both provided texts against your request. My primary directive is accuracy, especially when dealing with potentially high-stakes research findings. Analysis of Input: 1. [episode]
- Reward Valuation in Large Language Models: Causal Induction of Anhedonia — Reward valuation in vision-language models is explored through a mechanistic framework inspired by clinical tests for anhedonia, revealing that perturbing specific reward-anticipatory units can induce behavioral effects mirroring human anhedonia. [episode]
- JUMP: Efficient Membership Inference on Fine-Tuned Diffusion Language Models — Public open-weight language models are often fine-tuned on private or domain-specific data before deployment, creating a need to audit whether individual records were used during adaptation. [episode]
- FusionBERT: Multi-View Image--3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder — FusionBERT is a novel multi-view visual fusion framework designed for image–3D multimodal retrieval, addressing limitations in existing methods that focus on single-view alignment and lack structured inter-view feature fusion. [episode]
- Meta-TTL: Meta-Learning Self-Improvement Policies for Language Agents — Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions at inference time, and this paper introduces META-TTL, a framework that formulates discovery of effective adaptation policies as a bi-level optimization problem. [episode]
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model — Matrix-Game 2.0 is an open-source real-time and streaming interactive world model designed to generate long videos on-the-fly via few-step auto-regressive diffusion, addressing the limitations of existing models that suffer from latency and error accumulation when simulating real [episode]
- JointEdit3D: Feed-Forward 3D Scene Editing in a Unified Latent Space — JointEdit3D introduces a feed-forward framework for 3D scene editing that operates within a unified RGB-geometry reconstruction-generation latent space, addressing the limitations of existing methods by coupling appearance synthesis and geometry prediction during editing. [episode]
- PowerStep: Memory-Efficient Adaptive Optimization via p-Norm Steepest Descent — Adaptive optimizers like Adam are standard for training large neural networks but suffer from substantial memory overhead due to storing running estimates of first and second moments. [episode]
- Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge Devices — Running a language model on edge hardware provides private and low-latency reasoning without a network connection, and yet the small models that fit on such devices are unreliable on structured reasoning tasks like arithmetic or formal logic problems. [episode]
- ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis — Diffusion models have revolutionized text-to-image synthesis, yet they still struggle to properly render spatial relationships described in text prompts, which is particularly problematic for applications like urban planning. [episode]
- Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning — Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge to score model outputs according to rubrics as rewards, but policy models may exploit latent biases in this judge, leading to reward hacking and ineffective training outcomes. [episode]
- Gated Spatial Redundancy Projection for Pathology Transformer Attentions — Transformer models are increasingly used for whole-slide image analysis in computational pathology, but they face a fundamental challenge because whole-slide images (WSIs) differ from natural images by having neighboring patches that often contain highly similar tissue types, sta [episode]
- PRILoRA: Pruned and Rank-Increasing Low-Rank Adaptation — PRILoRA introduces a novel parameter-efficient fine-tuning method that combines linearly increasing low-rank allocations with ongoing pruning to improve performance over existing LoRA variants. [episode]
- It Just Takes Two: Scaling Amortized Inference to Large Sets — Neural posterior estimation (NPE) has emerged as a powerful tool for amortized inference, but its effectiveness is often limited by the computational costs associated with training estimators on large sets of observations. [episode]
- Non-Myopic Active Feature Acquisition via Pathwise Policy Gradients — As a meticulous researcher, I have thoroughly analyzed both provided texts concerning the paper "Non-Myopic Active Feature Acquisition via Pathwise Policy Gradients" (NM-PPG). [episode]
- EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies — EventVLA introduces an end-to-end framework designed to solve the critical memory bottleneck in long-horizon Vision-Language-Action (VLA) policies by employing sparse visual evidence memory. [episode]
- ArtifactWorld: Scaling 3D Gaussian Splatting Artifact Restoration via Video Generation Models — ArtifactWorld is a comprehensive framework designed to resolve geometric and photometric degradations in 3D Gaussian Splatting (3DGS) models under sparse-view constraints by systematically expanding training data and employing a homogeneous dual-model paradigm. [episode]
- Evaluating Generative Models via One-Dimensional Code Distributions — This paper introduces a novel framework for evaluating generative models by shifting focus from continuous recognition features to discrete visual tokens, proposing two new metrics that operate purely on token statistics. [episode]
- Cryo-EM as a Stochastic Inverse Problem — Cryo-electron microscopy (Cryo-EM) faces a major challenge in 3D reconstruction due to structural heterogeneity, where biomolecules adopt multiple conformations under identical experimental conditions. [episode]
- HiReFF: High-Resolution Feedforward Human Reconstruction from Uncalibrated Sparse-View Video — HiReFF is a feed-forward method designed for high-resolution (2K) 360° human video reconstruction from uncalibrated, sparse-view videos, addressing the critical need for temporal consistency and computational efficiency in applications like holographic communication and AR/VR. [episode]
- Amortized Optimal Transport from Sliced Potentials — This paper proposes two novel amortized optimal transport (OT) methods, RA-OT and OA-OT, which leverage Kantorovich potentials derived from sliced OT to efficiently predict OT plans across multiple measure pairs. [episode]
- Beyond Pixels: A Vector-to-Graph Framework for Reliable Schematic Auditing — Multimodal Large Language Models (MLLMs) suffer from structural blindness when analyzing engineering schematics because their pixel-driven paradigm discards explicit vector-defined relations necessary for topological reasoning. [episode]
- A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding — This paper introduces MVX-Bench, a new benchmark designed to rigorously evaluate Multimodal Large Language Models (MLLMs) on complex multi-video understanding tasks that go beyond simple event-level comparison. [episode]
- ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge — ARK is introduced as a dual-axis multimodal retrieval benchmark designed to systematically analyze retrieval systems from two complementary perspectives: knowledge domains and reasoning skills. [episode]
- RAWR: Reward Assignment Without Rollouts in Verifiable Domains — Process supervision for chain-of-thought reasoning via Monte Carlo Net Information Gain introduces a novel, efficient method for automatically generating step-level labels to train Process Reward Models (PRMs), significantly improving the reliability and scalability of LLM reason [episode]
- Evidence-Guided Schema Normalization for Temporal Tabular Reasoning — Temporal reasoning over evolving semistructured tables poses a challenge to current QA systems, and this work proposes an SQL-based approach that involves generating a 3NF schema from Wikipedia infoboxes, generating SQL queries, and executing them. [episode]
- Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models — This paper introduces Autoregressive Video Inverse problem Solver (AVIS) and its highly accelerated variant, AVIS Flash, which leverage autoregressive video diffusion models to restore videos in a streaming manner. [episode]
- PRISM: A Geometric Risk Bound for Decomposing Drift into Scale, Shape, and Head — Comparing post-training LLM variants—quantized, LoRA-adapted, distilled—needs a diagnostic that pinpoints how the variant has drifted, not just that it has; existing similarity scores (CKA, SVCCA) flag degradation without linking to risk or mechanism. [episode]
- Are We Really Making Much Progress in Text Classification? A Comparative Review — As a meticulous researcher, I have thoroughly analyzed the provided text excerpts from the paper, "Are We Really Making Much Progress in Text Classification? A Comparative Review," to construct a comprehensive and detailed summary. [episode]
- Training-Free Global Geometric Association for 4D LiDAR Panoptic Segmentation — Geo-4D introduces a novel, training-free framework for 4D LiDAR panoptic segmentation that unifies spatial and temporal reasoning to enable holistic perception over long time horizons. [episode]
- Poly-attention: a general scheme for higher-order self-attention — As a researcher, I must ensure absolute precision. [episode]
- Evaluating Alignment of Behavioral Dispositions in LLMs — Models often fail to appropriately reflect consensus opinion in scenarios with high human consensus and fail to reflect the diversity of opinions in scenarios with low human consensus. [episode]
- GRAVITY: Architecture-Agnostic Structured Anchoring for Long-Horizon Conversational Memory — Long-horizon conversational agents require memory systems that can provide relational, temporal, and thematic structures to ground complex reasoning, which this paper addresses by introducing GRAVITY, a plug-and-play module that injects structured knowledge at generation time. [episode]
- When Does Equivariance Help? Canonical Alignment in Neural Fluid Surrogates — Neural surrogates offer an accelerated alternative to high-fidelity Computational Fluid Dynamics (CFD) simulations, but their success depends critically on how they handle scalability and inductive biases. [episode]
- NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training — Diffusion models have achieved remarkable success across various generative tasks, but their training paradigm largely treats injected noise as uniformly informative. [episode]
- Observation-Aligned Mask Priors for Learning Physical Fields from Authentic Occlusions — Learning physical dynamics directly from incomplete observations is challenging because authentic occlusions are structured, sample-dependent, and often missing not at random (MNAR), unlike generic corruption. [episode]
- Reverberation: Learning the Latencies Before Forecasting Trajectories — As a meticulous AI researcher, I have thoroughly analyzed the provided excerpts from "Reverberation: Learning the Latencies Before Forecasting Trajectories." My synthesis will be comprehensive, precise, and structured to reflect both the core methodology and its critical technica [episode]
- Edge Selection for the Effective use of Piecewise-Constant Distributions as Neural Network Outputs for Event Prediction — Categorical distributions are effective neural network outputs for event prediction across various datasets, demonstrating their utility in modeling both continuous-time and discrete-time event sequences. [episode]
- Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding — Negation understanding in Korean remains an under-explored area, and this research introduces Thunder-KoNUBench, a new sentence-level benchmark designed to systematically evaluate how large language models handle negation phenomena in Korean. [episode]
- Think Multilingual, Not Harder: A Framework for Analyzing and Teaching Code-Switched Reasoning — LLMs are increasingly observed to code-switch during reasoning, and this work introduces a linguistically and behaviorally motivated fine-tuning framework to teach these models beneficial code-switched reasoning behaviors in a data-efficient manner. [episode]
- Guided Trajectory Optimization with Sparse Scaling for Test-Time Diffusion — The paper introduces RTS, a novel Reward-guided Trajectory Scaling method designed to enhance diffusion model generation performance during inference by efficiently navigating the high-dimensional noise space. [episode]
- Measuring (some aspects of) the metacognition of AI — A robust decision-making process must take into account uncertainty, especially when choices involve inherent risks, making it crucial to employ robust methods to measure and regulate the metacognitive capabilities of artificial intelligence systems. [episode]
- Achieving detailed medial temporal lobe segmentation with upsampled isotropic training from implicit neural representation — Imaging biomarkers in magnetic resonance imaging (MRI) are crucial for diagnosing, tracking, and treating Alzheimer's disease (AD), particularly because neurofibrillary tau pathology spreads through brain subregions, beginning in the medial temporal lobe (MTL). [episode]
- Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning — Scaffolding Minds proposes a two-stage training paradigm to optimize latent visual representations for multimodal reasoning, addressing limitations in existing methods where latent targets are suboptimal and RL lacks explicit exploration capabilities. [episode]
- POET: Preference Optimization for Enhanced Text-to-Image Generation — This work introduces an input-side inference-time scaling framework that enhances text-to-image generation through LLM-based prompt rewriting, demonstrating that optimizing only the input text can significantly improve image quality, alignment, and aesthetics across diverse T2I b [episode]
- Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives — This research investigates whether language models (LMs) possess genuine language understanding by evaluating their capacity for entity tracking across naturalistic narratives, comparing their performance against human comprehension. [episode]
- Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics — Exact Flow Linear Attention (EFLA) introduces an exact-flow formulation of delta-rule linear attention by interpreting its update as an explicit Euler discretization of an underlying continuous-time system, thereby eliminating first-order numerical integration errors without intr [episode]
- HARMONI: Aligning Human and Scene Priors for Multi-View 4D Reconstruction — HARMONI is a unified framework designed to jointly reconstruct cameras, scene point clouds, and human meshes from multi-person multi-view videos in a single pass without requiring external modules or preprocessing. [episode]
- Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR — Large language models often struggle with cross-lingual knowledge consistency when querying in low-resource languages, especially Indian languages and their code-mixed variants. [episode]
- AdvMT: Adversarial Motion Transformer for Long-term Human Motion Prediction — The Adversarial Motion Transformer (AdvMT) is a novel model designed to tackle the significant challenge of accurately predicting long-term human motion by integrating a transformer-based motion encoder with a temporal continuity discriminator. [episode]
- Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning — The paper investigates a critical gap in measuring diversity within Large Language Model (LLM) mathematical reasoning: conventional metrics capture only surface-level variation, failing to distinguish between genuine differences in problem-solving strategies. [episode]
- Event-based Scene Synthesis via Inter-Frame Residual Alignment — DESSERT is a Diffusion-based Event-driven Single-frame Synthesis framework that leverages pre-trained Stable Diffusion to predict residual latents between frames, enabling sharp and temporally consistent frame synthesis without relying on traditional optical flow or warping metho [episode]
- Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling — Aligning large language models (LLMs) to serve users with heterogeneous and potentially conflicting preferences is a central challenge for personalized and trustworthy AI, which this paper addresses by formalizing an ideal notion of universal alignment through test-time scaling. [episode]
- Vision Harnessing Agent for Open Ad-hoc Segmentation — The Vision Harnessing Agent for Open Ad-hoc Segmentation (VASA) introduces a novel, training-free framework that enables AI agents to construct visual concepts on the fly by iteratively reasoning over persistent visual states. [episode]
- Quantum Geometry of Data — This paper introduces Quantum Cognition Machine Learning (QCML) as a novel framework for representing data by encoding it as quantum geometry, demonstrating how this approach can capture rich geometric and topological structures in high-dimensional datasets. [episode]
- Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents — Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents introduces RL-Focal, a two-stage reinforcement learning agent framework designed to dynamically optimize ensembles of Large Language Models by learning which ensemble strategies are most effective [episode]
- Data-Efficient Time-Dependent PDE Surrogates: Graph Neural Simulators vs. Neural Operators — Developing accurate, data-efficient surrogate models is central to advancing AI for Science, and this paper argues that Graph Neural Simulators (GNS) offer a principled alternative to traditional Neural Operators (NOs) for time-dependent Partial Differential Equations (PDEs). [episode]
- Leveraging Latent Visual Reasoning in Silence — This paper investigates whether latent visual reasoning, which involves generating continuous latent embeddings before textual generation, retains value when it is not explicitly preserved as an inference-time format. [episode]
- Trustworthiness Costs of Domain Adaptation in Small Language Models:A Cross-Architecture Empirical Study — This paper presents the first systematic cross-domain, cross-architecture empirical study quantifying the trustworthiness cost of domain adaptation in small language models (SLMs). [episode]
- Think, Then Look: Active Spatial Reasoning for House-Scale 3D Scene Understanding — Spatial reasoning in large-scale 3D environments remains challenging for current vision–language models, which are typically constrained to room-scale scenarios. [episode]
- Dropout Universality: Scaling Laws and Optimal Scheduling at the Edge-of-Chaos — This work develops a mean-field theory for dropout as a perturbation of critical signal propagation at the edge of chaos, revealing that front-loaded dropout schedules can cut test loss by 18–35% over constant dropout in MLPs and Vision Transformers at fixed budget. [episode]
- Mirage: a Clean-Label Backdoor against LiDAR 3D Object Detection — MIRAGE introduces proof-of-feasibility for a novel class of backdoor attacks against LiDAR 3D Object Detection (LiDAR 3DOD) models, demonstrating that a stealthy, black-box, clean-label attack can be achieved without requiring access to model architecture or ground truth annotati [episode]
- Predicting Team Performance from Communications in Simulated Search-and-Rescue — Understanding how individual traits influence team performance is valuable, but these traits are not always directly observable. [episode]
- High-Dimensional Partial Least Squares: Spectral Analysis and Fundamental Limitations — This paper provides a comprehensive theoretical analysis of Partial Least Squares (PLS) in high-dimensional settings by employing tools from random matrix theory to characterize its spectral properties and fundamental limitations. [episode]
- TriageRA-CCF: Source-Side Clinical Confidence and Coverage Signals for Adaptive Rank Budgeting in Medical LLMs — Medical large language models are commonly adapted with a fixed low-rank budget, even though medical questions differ substantially in confidence, clinical coverage, and cross-domain difficulty. [episode]
- PrototypeNAS: Rapid Design of Deep Neural Networks for Microcontroller Units — PrototypeNAS is a novel zero-shot Neural Architecture Search (NAS) framework designed to rapidly design and specialize Deep Neural Networks (DNNs) for resource-constrained microcontroller units (MCUs). [episode]
- Revisiting the Capacity Gap in Chain-of-Thought Distillation from a Practical Perspective — Chain-of-thought (CoT) distillation transfers reasoning behaviors from a strong teacher to a smaller student, but prior work reports a capacity gap: distillation may fail when the teacher–student capability mismatch is large. [episode]
- From Seeds to Semantics: Measuring Semantic Accessibility in Deterministic Diffusion Models — This paper investigates how latent structure emerges within deterministic diffusion models by analyzing the relationship between initial noise seeds and generated samples, specifically focusing on how confidence scores can reveal class-relevant organization in the latent space. [episode]
- SelFusion: Self-distillation for Diffusion Language Models — Diffusion language models (DLMs) offer faster inference capabilities compared to autoregressive (AR) large language models (LLMs), making them suitable for real-time applications. [episode]
- Do Gaussian Scenes Contain Enough Structure for Intrinsic Segmentation? — Dynamic 4D Gaussian Splatting reconstructs deforming scenes at high fidelity, but segmenting these scenes for editing or analysis typically relies on costly external 2D masks from foundation models. [episode]
- Generalizing Geometry-Guided Mamba as a Plug-and-Play Context Module for CNN-based Semantic Segmentation — This paper investigates Directional Geometric Mamba (G-Mamba) as a plug-and-play context aggregation module for CNN-based semantic segmentation, addressing limitations in existing context heads that often introduce heavy computation or boundary leakage. [episode]
- The Collective Turing Test: Large Language Models Can Generate Realistic Multi-User Discussions — Large Language Models (LLMs) can generate social media conversations sufficiently realistic to deceive humans when reading them, highlighting both a promising potential for social simulation and a warning message about the potential misuse of LLMs to generate new inauthentic soci [episode]
- Entity Linking using LLMs for Automated Product Carbon Footprint Estimation — Growing concerns about climate change and sustainability are driving manufacturers to take significant steps toward reducing their carbon footprints. [episode]
- Detail Preserving Depth Estimation from a Single Image Using Attention Guided Networks — This paper proposes a novel network architecture for single-image depth estimation that focuses on preserving fine structural details, which is crucial for applications like 3D scene reconstruction and realistic rendering. [episode]
- Choices Speak Louder than Questions — The evaluation of Large Language Models (LLMs) using Multiple-Choice Question Answering (MCQA) is unreliable because model decisions are often more influenced by superficial characteristics of answer options than by genuine comprehension of the question. [episode]
- Segment Anything for Dendrites from Electron Microscopy — This paper introduces DendriteSAM, a vision foundation model based on Segment Anything (SAM), designed for the interactive and automatic segmentation of dendrites in electron microscopy (EM) images. [episode]
- High-dimensional density estimation with tensorizing flow — This paper proposes a novel framework called "tensorizing flow" for estimating high-dimensional probability density functions from observed data, combining tensor-train representation with continuous-time flow models. [episode]
- Fooling Algorithms in Non-Stationary Bandits using Belief Inertia — This paper introduces a fundamentally different approach to establishing worst-case lower bounds for regret in piecewise-stationary multi-armed bandits by leveraging a "belief inertia" argument. [episode]
- Empirical Bayes 1-bit matrix completion — This study develops an empirical Bayes method for 1-bit matrix completion, motivated by the Efron–Morris estimator, which generalizes Stein's estimator to matrices by shrinking singular values toward zero. [episode]
- Cold-Start Active Preference Learning in Socio-Economic Domains — Active preference learning faces a significant hurdle known as the cold-start problem when no initial labeled data are available, which severely limits its performance in socio-economic domains. [episode]
- Moving Object Detection from Moving Camera Using Focus of Expansion Likelihood and Segmentation — This paper proposes Focus of Expansion Likelihood and Segmentation (FoELS), a novel method for separating moving objects from static scenes when viewed by a moving camera. [episode]
- From Laboratory to Real World: A New Benchmark Towards Privacy-Preserved Visible-Infrared Person Re-Identification — Visible-infrared person re-identification (VI-ReID) is a crucial technology for reliable pedestrian identification across varying lighting conditions, but existing methods often rely on centralized training, which raises serious privacy concerns in real-world scenarios where data [episode]
- Hybrid Approach for Enhancing Lesion Segmentation in Fundus Images — Choroidal nevi are common benign pigmented lesions in the eye, but their accurate segmentation from color fundus images remains challenging due to indistinct boundaries and data scarcity. [episode]
- SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0 — SiDiaC-v.2.0 represents the largest comprehensive Sinhala diachronic corpus to date, spanning from 1800 CE to 1955 CE and covering a historical span from the 5th to the 20th century CE, making it a vital resource for Sinhala Natural Language Processing and temporal linguistic stu [episode]
- Boosting Adversarial Robustness and Generalization with Dictionary Structure — This work investigates a novel approach to boost adversarial robustness and generalization by incorporating structural prior into deep learning models, specifically addressing limitations found in existing dictionary learning-inspired convolutional neural networks (CNNs) which pr [episode]
- Hybrid Feature Learning for Handwriting Verification — This paper proposes a Hybrid Deep Learning (HDL) architecture designed to determine the probability that a questioned handwritten word was written by a known writer, addressing the need for more robust handwriting verification methods by combining deep learning with traditional f [episode]
- Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models —
- Hierarchical Utility Calibration for Structured Multiclass Decisions —
- Retrieval Sensitivity to Identity Signals in Queries —
- When Updating Stops Being Learning: Rethinking LLM Self-Evolution via learnable information gain —
- DraftTrace: A Multi-View Analytics Environment for AI-Integrated Writing —
- SCCM: Spherically Consistent Coarse Matching for ERP Dense Feature Correspondence —
- Grounded Revision vs. Prior Injection: Probing Retrieval-Augmented Patent Claim Amendment —
- How Medical VLMs Underutilize Their Vision Encoders: A Dermatology Perspective —
- FM-ReID: Selective Competitive Token Routing for Object Re-Identification —
- ThinkingGuard: Decoding Implicit Hazards via Step-by-Step Risk Attribution in Multimodal Large Language Models —
- AffectReveal: Event-Grounded Emotion Recognition Beyond Visual Appearances —
- Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It —
- SEED: Self-Speculative Decoding via Implicit Encoder-Decoder —
- Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation —
- Scaling Video Generation for Reasoning: At What Cost? —
- Pixel-wise Exposure for Highly Robust In-Vehicle Remote-PPG —
- Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics —
- CrossTimeEdit: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation —
- Generating Edit-Inducing Questions for AI Research Manuscripts —
- Neural Structural Reasoner: A Brain-inspired Architecture for Reasoning over Structured Knowledge —
- Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning —
- What Makes Recurrence Effective in Looped Language Models? —
- PE-OPSD: Internalizing Prompt Enhancement into Flow-matching Models via On-Policy Self-Distillation —
- OCA: ODE-Driven Cross-Attention for Image-to-Point-Cloud Registration —
- VLM4Cluster: Benchmarking Deep Clustering In the Era of Vision-Language Pre-training —
- FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution —
- Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference —
- Not Every Correction Helps: Gain-Guided Continual Test-Time Adaptation —
- You Only Reprogram Once: Rethinking Prolonged Training for Visual Reprogramming —
- G"odel Forest: Balancing Search Depth and Breadth for Data-Centric Recursive Self-Improvement —
- ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking —
- Understanding Private Evolution as Learning-Augmented Clustering —
- Reprogramming Vision-Language Models via Structured Prompt Reparameterization —
- MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization —
- ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context —
- When Semantics Matter: Reliability-Aware Semantic-Rhythm Control for Co-Speech Gesture Generation —
- CHAIN: Calibrated LLM Forecasting via Causal-Temporal Hypergraph Inference —
- Video2Skill: From Streaming Experience to Reusable Embodied Skills —
- AESplat: Advancing Pose-Free Feed-Forward 3D Gaussian Splatting via Decoupled Appearance Modeling —
- Lost in Conversation or Lost in Translation? Diagnosing Multi-Turn Degradation in RAG —
- When Is Coarse Supervision Worth It? Cost-Aware Learning under Unknown Aggregation —
- LAURA: Knowledge Distillation for Interpretable Ambiguous Clause Identification in Legal Contracts —
- ATTUNER: Recomputation-Free KV Cache Reuse via Query-Side Adaptation —
- Can Agents Design Libraries for Agents? —
- Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models —
- Backpropagated Output Momentum: Relocating Optimizer History from Parameters to Task Space —
- SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation —
- Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving —
- QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching —
- Federated Clustering with Unknown Local and Global Cluster Cardinalities —
- Beyond Conditional Independence: Root Cause Analysis with Deep Causal Models —
- Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLMs —
- VAA-CSEC: Vote-guided Advantage Allocation for Chinese Semantic Error Correction —
- Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows —
- CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning —
- On-Policy Visual Evidence Distillation —
- GlassFormer: Learning Real-time Glass Segmentation using Radar-Depth Fusion —
- Rethinking Multimodal Fake News Detection in the Generative AI Era —
- Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents —
- Momentum-Coupled Rubric Adaptation for Detailed Image Captioning —
- RAEGNet: Relation-Aware Evidence Graph Network for Harm-Aware Multimodal Fake News Detection —
- MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation —
- BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR —
- Can Language Models Learn to Forecast Stock Prices —
- Benchmarking Automatic Speech Recognition Tools for Iberian Languages —
- Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation —
- CoEM: Empowering Long-Context Reasoning with Commit-on-Evidence Memory —
- Safe-by-Design Learning via Energy-based Neural Networks —
- Scalable Diffusion SBI for Compositional Inference under Simulator Misspecification —
- ER-JEPA: Experience Replay Improves Joint-Embedding Predictive Learning in Language Models —
- Cool the Sampler, Not the Learner: Sampling Temperature Moves the Staleness Cliff of Importance-Corrected GRPO —
- VStress: Correlation-Aware Auditing and Adaptive Budget Allocation for Repeated Verifiers —
- Chinese-Jev: Bringing System One Model to Chinese-Language Tasks —
- Repetition, Not Length: Isolating the Counting Failure in Neural Text-to-Speech —
- AMU:Admission and Memory Update for Personalized Conversations---Structured Memory with SLM Guided Control —
- SRJudge: Empowering Large Language Models with Selective Reasoning for Fine-Grained Knowledge Concept Tagging —
- CypherTurn: A Multi-Turn Benchmark for Conversational Text-to-Cypher Evaluation and the Autonomy Divergence —
- LatCom: Cross-Agent Latent Compression for Efficient Multi-Agent Collaboration —
- Selecting The Most Informative Tokens in Natural Language Autoencoders —
- Iterative Exact Discrete Guidance for Energy-Based Sampling —
- Learning from Think-Mode Advantage via On-Policy Distillation —
- Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search —
- Identifying ODEs from Unstructured Data with Causal Representation Learning —
- What Does Post-Training Change in Multilingual Reasoning? —
- VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses —
- Interpretable intrinsic dimension estimation through componentwise calibration of distance and angle —
- Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training —
- Cross-Linguistic Effects in Bilingual Phoneme BabyLMs —
- LLM unbranding: Erasing Commercial Identity while Preserving Generic Utility —
- Multimodal Detection of Higher-Order Behavioral Constructs: Self-Compassion in Structured Reflective Interaction —
- When Tools Silently Lie: Evaluating and Mitigating Blind Compliance in Tool-Augmented Data Agents —
- Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories —
- Bridging Semantic Gaps in RAG through Generated Context Knowledge Fusion —
- VLM Fine-Tuning for End-to-End Combinatorial Optimization —
- InsightMap: Structured Spatial Modeling for Embodied Multimodal Reasoning —
- Pointwise or Pairwise: When Do Pairwise Losses Help Reward Learning, Provably? —
- CredWise: A Controlled Agentic Decision-Intelligence Framework for Explainable and Auditable Credit-Risk Assessment —
- Follow the Entities: A Corpus Map for Agentic Search —
- Interacting particle guidance for sampling reward-tilted generative priors —
- Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents —
- V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents —
- AssayRouter: Historical Utility Priors for Frozen Molecular Predictor Routing —
- Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models —
- Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring —
- Solving Without Stopping: On-Policy Distillation at Small Scale —
- A Sharp Transition in Data Reconstruction under Differential Privacy —
- Port-Hamiltonian Latent Deliberation: Mitigating the Deliberation Drift Cliff in Test-Time Compute Scaling —
- SemOPT: Fixing Semantic Errors in LLM-based Optimization Modeling via Reward-Guided Search —
- Compiling Learning Problems into Adaptation Programs for Language Models —
- High-Dimensional Simulation-Based Inference in Latent Spaces —
- Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation —
- Look What You Made Us Cluster: Hate Narrative Extraction from Reddit Discourse —
- Learning to Retrieve Missing Evidence for Long-Term Memory QA —
- Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG —
- Probability Contracts: Accuracy, Coherence, and Decisions Across LLM Interfaces —
- FORUM: Frozen Outputs Reconciled Using Model Agreement for Visual Grounding —
- Regime Boundary Alignment for Evidence-Gated Question Answering —
- Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks —
- Your Benchmark Is Not Saturated: Reviving Multiple-Choice Evaluation with Answer Pooling —
- Larry Caused the Car to Stop, But the Model Didn't Notice: Transformer Blindness to the M-Heuristic —
- The Rashomon Wikipedia: A Data-Perspectivist Analysis of Divergent Historical Narratives —
- Who Warmed the Archives? LLMs Overestimate Historical Warmth —
- Evaluating Bounded Autonomy in Regulated Agentic AI: A Diagnostic Harness with Constitutional Rewards, Escalation Labels, and Runtime Governance —
- From Dissonance to Orchestration: Teacher Intervention in On-Policy Distillation —
- Hierarchical Compression of Vision-Language Model Benchmarks —
- Physical Muon: Orthogonalization as an Equilibrium Computation —
- E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models —
- Why Adaptive Optimizers Underestimate Rare Tokens —
- RunyaNER: Auxiliary Language Selection for Runyankore NER —
- Benchmarking graph-based models for in-silico toxicity prediction in drug discovery —
- Orthogonal Yet Coupled: Decoupling Geometric Components for Model Merging —
- Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models —
- Pair Difficulty Matters: Rethinking Pairwise LLM-as-a-Judge Evaluation and Consistency —
- Ornstein-Uhlenbeck Is Hard to Beat, Yet Superlinear Drift Ships Lower Transport Costs —
- Rational Clarification by Assistive Agents via Value-of-Information Reasoning —
- FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents —
- XU-RS: Explaining Credal Width in Random-Set Language Models —
- Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable —
- Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision —
- RLTL;DR: Self-improvement by Internalizing Self-generated Feedback —
- Co-Linguistics: AI-augmented Theory Construction in Linguistics —
- Flattening the Connectome Spectrum: A Spectral Filter for FC Induces a Pretraining Target for fMRI Encoders —
- Evaluating and Benchmarking the System One Model Jev —
- A Finslerian Approach for Embedding Directed Data —
- Corpus-Guided Dual-Path Propagation for Graph Retrieval-Augmented Generation —
- KUPAS MASTER: Distilling the Tacit Expertise of Master Practitioners into Agent-Ready Experience Corpora —
- LEMON-ZEST: Evolution-Informed Tokenization for Efficient Protein Language Modeling —
- When Models Don't Manipulate Manifolds: The Geometry of a Comparison Task —
- Learning from Shared-Control Overrides: Context-Driven Acceleration Profile Prediction for Personalized Overtaking —
- EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments? —
- Reader Proficiency Shapes Layer-wise Surprisal Profiles —
- Billiger.de Products: A Bilingual Entity Matching Benchmark —
- Predictive Geometry of Hidden Trajectories in Transformers —
- Context Language Models —
- Which papyrus HTR is good enough? Character-error-rate tolerance of four papyrological tasks on Greek texts —
- OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells —
- Selecting What Matters: Semantic Compression-Guided Selective Pooling for Long-Context Embeddings —
- A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses —
- CompOrca: Corpus-Scale Compliance Labelling of Instruction-Tuning Data —
- Thinking in Depth, Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue —
- The Geometry of Inference in Transformer Residual Streams —
- Can a Cacheable Decision Model Follow Rules? —
- Can Vision-Language Models Stay Helpful When Facing Implicit Risks? Intent-Privilege OPSD for Efficient Safety-Helpfulness Alignment —
- Counterfactual Probing for Parallel Unmasking with Hidden Forest Structure —
- AnthroDial: Benchmarking LLM Anthropomorphism in Autonomous Social Interaction —
- Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning —
- One Threshold Does Not Fit All Languages: Language-Conditional Deferral for Reliable and Efficient Low-Resource Text Classification —
- It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them —
- Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR —
- Retrieval Capacity of Self-Attention Under Competition —
- How Many Labels Does a Language Need? Annotation Budgets and Cross-Lingual Pooling for African-Language Text Classification —
- Zero-shot Dependency Parsing with Unsupervised Cross-Lingual Bootstrapping —
- It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs —
- The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment —
- Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning —
- Search Dimension in Unlabeled Projection Pursuit: A Scaling Law for Subspace Restriction —
- Time-Anchored Diffusion Language Models: Latent-Space Caching for Fast Generation —
- Learning What to Remember: Long-horizon Counterfactual Memory Optimization —
- Post-Anomaly Detection Inference for Deep SVDD —
- Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark —
- An Efficient Machine Learning Approach for Degradation Forecasting in AEM Water Electrolysis —
- Identifiability Guarantees for Drivers and Dynamics of Delayed Physical Systems —
- SelfSearch: Reward-Free Search for Self-Improving Agents —
- On Trajectory-Aware Training for Masked Diffusion Language Models —
- S cubed: Spectral Null-Space Swap Makes Reasoning Models Efficient —
- Which Attention Heads are like the Human Head? Not the Ones that Compute —
- BITEM at the NTCIR-19 R2C2 Task: Predicting Confidence from Agentic RAG Pipeline Signals —
- When do data mixtures improve scaling laws? Insights from high-dimensional regression —
- Brain-SAD: A Brain-Inspired Safe Autonomous Driving Control Framework with Dynamic Fear-Oriented Constraint on Dual-Policy —
- Auditing Long-Term Memory Evaluation: Repeated Judging, Reader Variation, and Negative Controls —
- Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models —
- Layer-Informed Fine-Tuning via Three-Stage Functional Segmentation of LLMs —
- Gender bias across LLMs is common and highly heterogeneous —
- EVO-WAM: Evolving World Action Models through Video-Action Verification —
- Latent Inference-Time Guidance of Time Series Foundation Models —
- Traversing the solution space of neural networks with Hessian Null Space Continuation —
- Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces —
- How Local Mixing Encodes Relative Position in Global NoPE Attention —
- From Routing Signals to Selective Review: Visual regrounding in MoE VLMs —
- ReCIRC: Rectified Conformal Risk Control —
- Stochastic World Models for Verifying Vision-Based Neural Feedback Systems —
- Multi-Agent Flow Matching with Decoupled Generative Guidance —
- LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning —
- AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation —
- Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI —
- Pretraining Latent Information Feedback Transformers with Teacher Supervision —
- Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies —
- Rethinking Representations for World-Action Modeling —
- STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization —
- Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering —
- MONOVAB: An Annotated Corpus for Bangla Multi-label Emotion Detection —
- Pixels to Prose: Understanding the art of Image Captioning —
- Estimating the Causal Effects of T Cell Receptors —
- The Hitchhiker's Guide to Agentic AI: From Foundations to Systems —
- Sage: Formalization with Semantic Correction —
- FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech —
- Serverless gossip training of LSTM failure detectors: A matched-protocol comparison with federated, local and centralized learning on NASA C-MAPSS —
- Learning from the Gap Between Pass@K and Pass@1 —
- Sieve and Sage: Efficient Distraction Filtering for Reliable RALM Abstention —
- Developing an OCR model for Extracting Information from Invoices with Korean Language —
- OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing —
- HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization —
- Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models —
- Alignment Forecasting: Predicting Misalignment From Training Data —
- From Lexical Baselines to Agentic Retrieval-Augmented Generation: Structured Skill and Responsibility-Level Extraction with the SFIA Framework —
- Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety —
- When Successful Memories Mislead Embodied Agents:Memory Adaption For Task-Conditioned Execution —
- Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News? —
- TRACE: Deployable Tree-Relational Structure Enhancement for Oncology LLMs —
- Lookahead-R: Budget-Aware Tool Retrieval via Execution-Centric Planning —
- Automated Evaluation of Multi-Turn Dialogues in In-Car Conversational Assistants —
- Constructing Challenging Browser-Use Tasks by Controlled Environment Interventions —
- How to Run Statistics over LLM Judges and Trust the Results: Calibrated Inference for Small-Sample AI Evaluation with evalstats —
- PrimeSeeker: Capability-Oriented Supervision for Deep Search Agents —
- Less Uniform Discrete Diffusion is More Powerful and Scalable —
- tau-Multilingual: Benchmarking Voice Agents Across Languages —
- CrisisFake: Benchmark Validity of AI-Generated Text Detection for Disaster Social Sensing —
- Tracing mechanisms of sycophantic agreement in language models —
- CoVLM-Bench: A Real-World Benchmark for Cooperative Driving Question Answering and Planning —
- Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation —
- Beyond the Context Window: An Adaptive Entropy-Based Routing Framework for Hybrid Retrieval and Long-Context Language Models —
- When Should LLMs Trust Their Own Revisions? A Risk-Aware Study of Intrinsic Self-Correction —
- Hyperspherical Semantic Trajectory Analysis: Mapping Technological Diffusion across Academic Preprints, Patent Signals, and Compute Scaling —
- The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models —
- Resolving the Missing Financial Data Crisis: A Generative AI Pipeline for SEC 10-K Extraction —
- PACT: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions —
- The Price of Token Boundaries: Compression Certificates and Prediction —
- CruxBench: A Benchmark of Information Discovery —
- Wasserstein Causal Forests for Distribution-Valued Outcomes —
- The Decision Value of Perception Compute —
- Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change —
- Question-Specific Knowledge Graphs for Efficient Visual Reasoning —
- HERO: Histology Encoder for Robust Representation in Oncology —
- FluxLite: Inference-Time Proposal Control for Discrete Diffusion Models —
- Intrinsic Associative Memory on Riemannian Manifolds: Curvature, Capacity, and Emergent Modes —
- HEIR: Learning Human-Entity Interactions with Functional Roles —
- Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method —
- Causal and Interpretable Structures in LLM Compositional Tasks —
- Making Cross-Continental Federated Learning Repeatable with FLIP: a Multi-Application Study —
- Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion —
- CoDimRecon: Agentic Reconstruction of Sim-Ready 3D Scenes with Deformable Curves, Surfaces, and Volumes —
- Preferent Compression Bounds Are Tight —
- PowerZooJax: A JAX-based Power System Benchmark for Reinforcement Learning —
- GeoWind2Plan: Mission-Time 3D Urban Wind Prediction for Energy-Efficient UAV Planning —
- Mirror-Score: Calibrated, Inference-only Scoring Exposes the Limits of Sequence-compatibility Ranking in D-peptide Design —
- Mnemon: Raw Records, Fast Judgments, Slow Thoughts —
- AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search —
- A Polyphonic Conception of AI Understanding —
- GeoOutageBench: Benchmarking Ambiguity-aware, Ontology-grounded Geospatiotemporal KGQA for Multimodal Power Outage and Resilience Analysis —
- PADM'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators —
- One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models —
- Why Backdooring Neural Networks is so Easy? —
- A Character-Level Neural Approach to Sinhala Sandhi Splitting —
- Hardware-Aware Functional Kolmogorov-Arnold Networks for Efficient Medical Image Enhancement and Segmentation —
- Xiaomi-OCR-0 Technical Report —
- When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs —
- Language Models Are "Insecure" Reporters —
- From Sharp Eyes to Expert Mind: Internalizing Expert Knowledge in MLLMs for Tampered Text Detection —
- Encoder-Sharing Hierarchical Federated Multi-Task Learning for VANETs —
- Principled Thoughts for Latent Recursive LLM Systems —
- Boosting Metric Depth Completion via Training-Free Adaptive Response Geometry —
- Exploring Learning Models for Topological Relationship Recognition from Image Data —
- Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning —
- FD-AA: A Lightweight Focal-Diffuse And Attenuation-Aware Head for Incidental Abdominal Abnormality Detection in Chest CT —
- Concept Direction Reliability Across Languages with Different Tokenizer Fertility —
- PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents —
- SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety —
- FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models —
- Geometric Representations of African Languages: A Regional Semantic Hub and Cultural Steering —
- The Canonical Order Problem: When Large Language Models Are Unreliable Knowledge Bases for Multi-Valued Relations —
- On the spectral properties of generative denoiser Jacobians —
- Lost in Translation: Measuring the Effect of Non-Native English on End User Performance of Large Language Models —
- Sparse-View Interpretable 3D Animal Behavior Representations for Neural Encoding and Decoding —
- CineSubBench: Evaluating LLMs on Long-Form Narrative and Cultural Understanding from Multilingual Movie Subtitles —
- LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning —
- Mutually Adversarial Self-Training with Evolving Data for Unified Multimodal Models —
- One-Step Next-Latent Prediction Is Not a World Model —
- An Empirical Study and Assessment of EU AI Act Compliance Checkers —
- ChronoSRL: Temporal Geometry for Self-Supervised Reinforcement Learning —
- Cognitive Expert Language Models Better Align with the Corresponding Brain Systems —
- Can Representation Learning Decouple from Loss Minimization? Polar Updates Have an Answer —
- Think Before You Restore: Risk-Aware Manchu Manuscript Restoration with Stroke-Guided Attention —
- Learning from Teacher Continuations at Student States —
- Action Chunking Proximal Policy Optimization with Feedback Correction —
- The Signed Geometry of One-Shot Recourse: On-Path Validity and the Signed-Curvature Criterion —
- Population Fidelity: Evaluating Population Representativeness in LLMs —
- OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models —
- In-Context Learning Amplifies a Latent Symbolic Circuit —
- Stochastic Optimization Under Power-Law Spectra: Tight Bounds and Shuffling Analysis —
- GNA: Granular Neighbor Assembly for Retrieval-Augmented Multivariate Time-Series Forecasting —
- When Trees Are Not Enough: Learning Mixed-Topology Feature Graphs with Adaptive Graph Sparse Autoencoders —
- MoRE: Scaling mixture of experts with hardware-aware low-rank routing —
- HeurEvo: Agentic Evolution of Hybrid Solver-Augmented Heuristics for Time-Critical Mathematical Optimization —
- Cheap and Powerful Tests for Supervised Subspaces: Per-Component Inference for PLS —
- Fractional State Space Transition for Long Sequence Modeling —
- PyroStack: A Multi-Band Spatio-Temporal Sub-Daily Dataset for Wildfires in the United States —
- Training LLMs to Verbalize Evaluation Awareness —
- Adapting Linear-Time Architectures for Tabular In-Context Learning —
- DeepRewind: Predicting and Repairing Premature Commitments in Deep Research Agents —
- Representation by Design in Generation: Cross-View Class-Token Alignment in Diffusion Transformers —
- StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks —
- Explainability from Training with Applications to TCR-Epitope Prediction —
- Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation —
- AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models —
- OTT3R: Multi-View 3D Reconstruction and Fast Dataset Generation at 1% Compute —
- LEGO-Anything: Coding Agents for 3D Scene Reconstruction —
- Stealth Is a Relation, Not a Property: How Event Representations Create Blind Spots for Timing Attacks in Event-Based Perception —
- LOCO-AdaMP: Built-in LOCO Inference for Adaptive Minipatch Ensembles with Enhanced Prediction —
- Calibrated to Whom? Persona and Language Effects on Cultural Values in JEV —
- What Makes High-Magnification Knowledge Transferable? A Study of Cross-Resolution Distillation in Whole-Slide Imaging —
- Eternal Sunshine of the Spotless Mind: Systematically Erasing LLM's Memories —
- Towards Scalable Context-Aware Single-Cell Spatial Transcriptomics Prediction from Histology Images —
- Temporal-Aware Fusion for Robust Outdoor LiDAR Localization —
- RA-CFGCache: From Branch-Level Criteria to Guided-Risk Control under Classifier-Free Guidance —
- MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization —
- Merlin Plus: A Large-Scale, Multi-Cancer, Image-Mask-Report Dataset —
- DARE to Mitigate Hallucination: Dual-path Auto-Regressive-aware Editing —
- Online Versatile Incremental Learning: Towards Class and Domain-Agnostic Adaptation at Any Time —
- Invariant Atoms: Sparse Coordinates of Local Semantic Geometry in Language Model Representations —
- Reliable Parallel Decoding in Masked Diffusion Language Models —
- DynamicHOI: Coupled Dynamics for Physics-aware HOI Reconstruction —
- Memory Consolidation Flattens the Temporal Shape of User Facts —
- Fisher-IRG: Fisher-Induced Local Invariant Representation Geometry across Language and Vision Models —
- FinRT: Distilling Adaptive Red-Teaming Strategies into Reusable Adversarial Generators in Consumer Finance —
- Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models —
- Optimal Multi-Reward Reinforcement Learning —
- Benchmarking Vision-Language Models on Synapse Detection and Proofreading in Connectomics —
- Reimagine Video Dynamics —
- Large-scale factor analysis shows machine intelligence is only partially interpretable —
- Adapting Context Compression for Long-Horizon Agents with Counterfactual Continuations —
- Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling —
Important terms
- Mirror-Score
- This method uses calibrated, inference-only scoring to test sequence-compatibility ranking in D-peptide design, aiming to find better ways to model complex biological interactions.
- Scalable Diffusion SBI
- This technique helps make reliable predictions when the underlying simulation models are imperfect by using scalable diffusion for compositional inference.
- MonitorBench
- This benchmark is designed to assess the chain-of-thought monitorability of large language models, providing a standardized way to probe how these models reason and self-regulate.
- Meta-TTL
- This research explores meta-learning self-improvement policies for language agents, suggesting a path toward autonomous refinement of agent behavior by learning improvement strategies.