AI papers — 2026-10-05
Today's focus is on establishing a new benchmark for generating and understanding RNA dynamics, which is crucial because we need robust models to predict how biological molecules behave over time. We looked at RNADyn, which aims to provide this benchmark by evaluating its ability to generate and interpret RNA dynamics.
A key piece of work involved autoregressive frontier expansion using graph machine learning to grow trees. This suggests a method for structuring complex data relationships. This feeds into the work on neurolens, which focuses on learning latent embeddings of neural semantics from chronic recordings. This provides a way to understand underlying patterns in biological signals.
We also explored broken scale symmetries in undercomplete linear autoencoders. This touches upon how models can capture essential information even when the data space is constrained. This idea connects to stochastic engrams for efficient continual learning, which seeks to improve how models learn new information incrementally without forgetting old knowledge.
Finally, we touched on automatic register identification for the open web using multilingual deep learning and evaluating retrieval robustness of large language models. Both of these explore how these powerful language systems can be made more reliable and context-aware in real-world applications.
The most significant piece of work today involved BioMol-MQA, which attempts to build a multi-modal question answering dataset specifically designed to test how large language models reason about bio-molecular interactions. This matters because it provides a structured way to see if these models can actually grasp the complex relationships between different biological components, moving beyond simple pattern matching.
A related effort focused on agent-to-agent theory of mind, which tests whether large language models can recognize when another model is aware of something they are not. This is important for understanding how sophisticated conversational agents might interact in a real system.
Then there was the Percept-V challenge, which investigates if multimodal large language models can solve straightforward perception problems by combining visual and text inputs. This explores the limits of what these models can perceive when given mixed sensory data.
Another study looked at Enrich-on-Graph, which uses query-graph alignment to help large language models perform complex reasoning by enriching their internal knowledge structures. This suggests a method for giving LLMs better scaffolding for difficult logical tasks.
We also saw work on WAON, which created a large Japanese image-text dataset aimed at improving cultural adaptation in contrastive vision-language models. This is useful for making visual and textual models more nuanced across different cultural contexts.
Finally, there was a unified BERT-CNN-BiLSTM framework used for simultaneously classifying headlines and analyzing sentiment in Bangla news. This shows how to handle multiple text analysis tasks at once on a specific language corpus.
The most significant work today involves developing methods to build interpretable moral contexts from human data because understanding how people make ethical judgments is crucial for aligning large language models. A study explored using probabilistic clustering and large language models to learn these contextual moral frameworks. This approach attempts to capture the nuances of human reasoning by grouping similar ethical scenarios together, which helps in creating a more nuanced understanding of morality than simple rule-based systems allow.
Another important direction is RT-SFT, which focuses on text style transfer from nonparallel corpora using roundtrip translation techniques. This work is relevant because it shows how to adapt language models to adopt specific tones or styles without needing perfectly matched training data. Following this, there is research into sensory-aware sequential recommendation via review-distilled representations. This method aims to improve how systems suggest items by distilling the essence of user reviews, making recommendations more contextually relevant based on what users have actually experienced.
We also looked at trust in large language models when they interact with physical hardware, specifically diving into reasoning ability under nonideality on memristors. This investigation matters because it tests the robustness of LLMs in real-world, imperfect computational environments. Related to this, there is work on GISTBench, which evaluates how well LLM users understand evidence-based interests through interest verification. This provides a way to measure if the model's output aligns with verifiable facts.
Finally, we examined compact portfolios for multi-objective LLM alignment to address the challenge of managing many preferences with few policies.
The work on recursive agent optimization is particularly important because it seeks to improve the efficiency and reliability of complex AI systems by allowing them to self-correct during execution. This approach involves a method where an agent refines its own plan based on intermediate results, which is a key step toward making autonomous agents more robust.
A study investigating how AI agents spend money provides insight into the practical resource demands of these systems. It specifically analyzes and predicts token consumption in agentic coding tasks to see where computational resources are being heavily utilized. This helps us understand the cost associated with deploying these sophisticated tools.
The research on hyperlogic offers a new benchmark for hard, forward-authored Chinese logical reasoning by using execution-derived answers to test the system's reasoning capabilities. This provides a rigorous way to measure complex logical skills in a specific language context, which connects to how well agents can handle intricate instructions.
Furthermore, automated auditing of LLM agent benchmarks addresses the critical issue of who is responsible for setting the standards used to evaluate these agents. This ensures that the benchmarks themselves are not biased or flawed. This work is crucial because it establishes accountability in the evaluation process.
Finally, fastkernels focuses on benchmarking GPU kernel generation in production environments to assess how quickly and effectively hardware-specific code can be generated for high-performance computing tasks. This provides a practical measure of the speed at which specialized AI components can be deployed.
The work on EchoDistill is particularly important because it tackles the robustness of large audio language models by using a self-distillation technique to clean up noisy audio inputs. This directly impacts how reliable these models are when deployed in real-world clinical settings. The researchers tried applying this noisy-to-clean self-distillation method to improve the performance of large audio language models.
One key finding showed that this distillation process significantly enhanced the model's ability to handle corrupted audio, suggesting that cleaning the input before processing is a worthwhile step for these models. This is supported by another study which analyzed how specialized multilingual model adaptation preserves knowledge across different languages when adapting them to new tasks.
Furthermore, there was an investigation into the limits of LLM adaptability by examining how model-internalized priors affect performance on annotation tasks. This suggests that the inherent biases within the model's training data impose constraints on its ability to perform well on specific labeling jobs. This finding connects to work looking at counterfactual evidence audits, which predicts susceptibility to ranked context in LLM agents.
Another piece of research explored how a language model pre-trained exclusively on historical text performs, providing a baseline for understanding the impact of domain specificity during initial training. This contrasts with efforts focused on hint-guided diversified policy optimization, which aims to improve reasoning capabilities through targeted guidance rather than just input cleaning.
The most significant piece of work today involves testing how far we can push lightweight hallucination detection across different tasks like question answering and summarization without needing a powerful GPU. This matters because it directly addresses the reliability of smaller, more accessible models in real-world applications.
We looked at a systematic benchmark to see the limits of this lightweight detection method. The results showed that while performance was good, there are clear boundaries where these models struggle with unsupported content. This finding connects to work on sentence-level context sensitivity, which suggests that understanding immediate surrounding text is a key way to detect unsupported content without needing extensive training.
Another important area explored was understanding the root causes of language model hallucinations by testing their reasoning against prior knowledge. This research investigates how much a model relies on its internal priors when generating incorrect information. This idea builds upon the work examining rule adherence in LLM adjudicators, which assesses whether models follow specific rules in complex scenarios like a tabletop role-playing game.
We also looked at the limits of on-policy self-distillation for continual post-training to see if denser models actually lead to better performance over time. This suggests that simply making the model larger or denser isn't always the answer for sustained improvement. Finally, there is ongoing work developing a multi-timescale recursive self-improvement engine designed to foster open-ended persona growth in language models.
The most significant piece of work today involves finding the move is not winning the game, which means XiangqiBench provides a closed-loop evaluation method for Large Language Model agents. This matters because it moves beyond simple task completion to assess strategic decision-making in complex games.
This XiangqiBench project explored how LLM agents perform in a closed-loop setting by evaluating their ability to find the optimal move, rather than just achieving a final state. It is connected to Hesitation Has a Geometry, which uses entropy-trained hyperbolic probes for sparse activation steering. This suggests that understanding where an agent hesitates might reveal strategic insights into its decision process.
Another important area is decoupling personalization from per-user adaptation through the work Does Every User Need a Private LoRA? This research investigates whether every user requires their own private low-rank adaptation to maintain performance, which has implications for deployment efficiency.
Then there is Lexicographic Multi-Objective On-Policy Distillation, which focuses on distilling knowledge using lexicographic criteria to achieve better performance in complex decision models. This distillation technique aims to preserve the most important decision outcomes while simplifying the underlying model structure.
We also looked at Trained Agentic Context Management, which addresses how agents should manage their context effectively during operation. This work builds upon HakemBench, a Turkish benchmark specifically designed for typed decisions, offering a structured way to test agent reasoning capabilities against predefined decision types.
The work on counterexample generation via per-theorem symbolic verifiers is crucial because it directly addresses the vulnerability of large language models by systematically finding edge cases where imitation fails. This involves using symbolic verifiers to generate specific inputs that expose flaws in the model's reasoning process. This contrasts with FinDialogLens, which focused on event extraction from multi-party financial dialogues to identify missed trades.
FinDialogLens attempts to extract specific events from complex financial chats, a method that provides direct insight into how models handle nuanced conversational data. This is less foundational than the counterexample generation work because it deals with application-specific data rather than core model limitations.
CUEing User Simulators introduces calibrated user embeddings for multi-turn benchmarking. This helps in systematically testing how models respond across different levels of user interaction complexity. This calibrating is important because it provides a standardized way to measure performance over several conversational turns.
Capability scaling-down laws for LLM compression explore methods to reduce the size of large language models while preserving their functional capabilities. This is vital for deploying these powerful tools efficiently. This relates to APDMem, which uses agent-controlled progressive disclosure for query-adaptive long-term memory, showing another path toward making models more efficient and contextually aware.
APDMem focuses on giving agents control over how much long-term memory they reveal based on the specific query. This is a refinement of capability scaling laws because it tailors resource usage to the immediate need. This approach builds upon the idea of providing better control over model access.
From retrieval to typed decisions, calibrated system one models derived from biomedical sentence encoders aim to move beyond simple information retrieval toward making structured decisions based on encoded meaning. This is a significant step in moving toward more reliable reasoning systems. This contrasts with auditing LLM judges for occupational AI measurement, which focuses on validating the fairness of existing measurements rather than creating new decision structures.
Finally, the generative-informed neuro-symbolic framework for syntactic ambiguity resolution using Arabic dependency parses tackles a deep linguistic problem by combining neural and symbolic methods to resolve unclear sentence structures. This work is foundational because resolving ambiguity is a prerequisite for many other complex reasoning tasks that the other papers are trying to solve.
Today's papers
- RNADyn: A Benchmark for Generating and Understanding RNA Dynamics. [paper]
- Autoregressive Frontier Expansion: Growing Trees with Graph Machine Learning. [paper] [episode]
- NeuroLens: Learning Latent Embeddings of Neural Semantics from Chronic Recordings. [paper]
- Broken scale symmetries in undercomplete linear autoencoders. [paper]
- Stochastic Engrams for Efficient Continual Learning. [paper] [episode]
- Automatic register identification for the open web using multilingual deep learning. [paper] [episode]
- The exponential distribution of the order of demonstrative, numeral, adjective and noun. [paper] [episode]
- Evaluating the Retrieval Robustness of Large Language Models. [paper] [episode]
- BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions. [paper] [episode]
- Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models. [paper] [episode]
- The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?. [paper] [episode]
- Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching. [paper] [episode]
- WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models. [paper] [episode]
- Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning. [paper] [episode]
- A Unified BERT-CNN-BiLSTM Framework for Simultaneous Headline Classification and Sentiment Analysis of Bangla News. [paper] [episode]
- Navigating the Reality Gap: On-Device Continual Adaptation of ASR for Clinical Telephony. [paper] [episode]
- Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models. [paper] [episode]
- RT-SFT: Text Style Transfer from Non-Parallel Corpora by Roundtrip Translation. [paper] [episode]
- Sensory-Aware Sequential Recommendation via Review-Distilled Representations. [paper] [episode]
- Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality. [paper] [episode]
- On the Tip of the Tongue: Why LLMs Hallucinate Answers They Can Decode. [paper] [episode]
- GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification. [paper] [episode]
- Many Preferences, Few Policies: Compact Portfolios for Multi-Objective LLM Alignment. [paper] [episode]
- Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity. [paper] [episode]
- Rhetorical Questions in LLM Representations: A Linear Probing Study. [paper] [episode]
- Rank-Turbulence Delta and Interpretable Approaches to Stylometric Delta Metrics. [paper] [episode]
- How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks. [paper] [episode]
- Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks. [paper] [episode]
- Useful Features, Backward Scores: OOD in Language-Model Trajectories. [paper] [episode]
- Recursive Agent Optimization. [paper] [episode]
- HyperLogic: A Hard, Forward-Authored Chinese Logical Reasoning Benchmark with Execution-Derived Answers. [paper] [episode]
- FastKernels: Benchmarking GPU Kernel Generation in Production. [paper] [episode]
- EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation. [paper] [episode]
- Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect. [paper] [episode]
- Specializing Without Forgetting: Analyzing Knowledge Preservation in Multilingual Model Adaptation. [paper] [episode]
- On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance. [paper] [episode]
- Counterfactual Evidence Audits Predict LLM-Agent Susceptibility to Ranked Context. [paper] [episode]
- Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification. [paper] [episode]
- A Language Model from 1913: Pretraining on Historical Text. [paper] [episode]
- Hint-Guided Diversified Policy Optimization for LLM Reasoning. [paper] [episode]
- SANE Schema-aware Natural-language Evaluation of Biological Data. [paper] [episode]
- Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish. [paper] [episode]
- How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation. [paper] [episode]
- Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors. [paper] [episode]
- Denser not equal to Better: Limits of On-Policy Self-Distillation for Continual Post-Training. [paper] [episode]
- Assessing Rule Adherence of LLM Adjudicators in Call of Cthulhu TRPG. [paper] [episode]
- Sentence-Level Context Sensitivity as a Training-Free Detector of Unsupported Content, Evaluated Against Trained Verifiers. [paper] [episode]
- A Multi-Timescale Recursive Self-Improvement Engine for Open-Ended Persona Growth. [paper] [episode]
- Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses. [paper]
- HakemBench: A Turkish Benchmark of Typed Decisions. [paper]
- Does Every User Need a Private LoRA? Decoupling Personalization from Per-User Adaptation. [paper]
- Lexicographic Multi-Objective On-Policy Distillation. [paper]
- Hesitation Has a Geometry: Entropy-Trained Hyperbolic Probes for Sparse Activation Steering. [paper]
- Trained Agentic Context Management. [paper]
- Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents. [paper]
- Are you Synthesizing or Recalling? Evaluating LLMs on Algorithmic Code Retrieval. [paper]
- Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs. [paper]
- FinDialogLens: Event Extraction over Multi-Party Dialogue for Missed-Trade Identification in Financial Chatrooms. [paper]
- CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking. [paper]
- Capability Scaling-Down Laws for LLM Compression. [paper]
The papers
- Recursive Agent Optimization — Recursive Agent Optimization (RAO) introduces a reinforcement learning approach for training recursive agents, which are models capable of spawning and delegating sub-tasks to new instances of themselves. [episode]
- Active Sampling for Ultra-Low-Bit-Rate Video Compression via Conditional Controlled Diffusion — Diffusion models provide a powerful generative prior for perceptual reconstruction at ultralow bitrates, but effective video compression requires controlling the generative process using highly compact conditioning signals. [episode]
- Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks — As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all—they are failures of the benchmark itself: broken specifications, implicit assumptions, and rigid evaluation scripts that penalize valid alternative approaches. [episode]
- EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation — Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. [episode]
- Denser not equal to Better: Limits of On-Policy Self-Distillation for Continual Post-Training — Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities, and this work investigates whether on-policy self-distillation (SDPO) can reliably serve as a stabilizer for continual learning. [episode]
- The Effective Depth Paradox: Topology and Trainability in Deep CNNs — Architectures utilizing identity shortcuts or branching modules maintain optimization stability by decoupling effective depth from nominal depth. [episode]
- Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification — Multimodal Large Language Models (LLMs) are increasingly used for scientific peer review, yet they exhibit a significant performance gap when verifying claims supported by charts compared to tables, even when both modalities represent the same underlying data. [episode]
- Bifidelity Karhunen-Lo`eve Expansion Surrogate with Active Learning for Random Fields — Bifidelity KLEs with Active Learning for Random Fields presents a novel surrogate modeling framework that combines Karhunen–Loève expansions and polynomial chaos expansions with an active learning strategy to efficiently construct accurate, computationally affordable models fo [episode]
- FastKernels: Benchmarking GPU Kernel Generation in Production — As a fastidious and diligent AI researcher, I have meticulously analyzed the provided excerpts (A, B, and C) pertaining to "Fast Kernels: Benchmarking GPU Kernel Generation in Production." The information is fragmented—a high-level abstract/summary (A), a detailed data table ex [episode]
- Rhetorical Questions in LLM Representations: A Linear Probing Study — Rhetorical questions are asked to persuade or signal stance rather than seek information, and understanding how large language models internally represent these questions remains unclear. [episode]
- Sentence-Level Context Sensitivity as a Training-Free Detector of Unsupported Content, Evaluated Against Trained Verifiers — Retrieval-augmented generation (RAG) systems often suffer from hallucination, and this research introduces Grounding-Aware Sensitivity by Perturbation (GASP), a span-level detector that scores each answer sentence by its grounding sensitivity—how strongly its likelihood depends [episode]
- When Does Pooling Pay? Credibility and Resolution under Forgetting in Intermittent-Demand Forecasting — Intermittent demand forecasting presents significant challenges due to sparse observations and cold-start items, and this paper introduces TSB-HB, a hierarchical Bayesian extension that provides a principled generative foundation by enabling partial pooling across items to stabil [episode]
- Attention Sinks in Diffusion Transformers: A Causal Analysis — Attention sinks—tokens that receive disproportionate attention mass—are assumed to be functionally important in autoregressive language models, but their role in diffusion transformers remains unclear. [episode]
- Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish — Turkish is an agglutinative language where meaning resides in morphemes, and current subword tokenizers fail to capture this morphology effectively. [episode]
- Reliability of Probabilistic Emulation of Physical Systems — Two dominant approaches for generating probabilistic forecasts of physical systems are generative models and ensembles of deterministic models trained with continuous ranked probability score (CRPS) loss; this work addresses the reliability gap by developing a framework to evalua [episode]
- World Action Planner: Generalizable Robot Decision-Making with Action-Conditioned World Models — Building generalizable agents for diverse applications remains a fundamental challenge, and this work proposes World Action Planner, a robot planning system that leverages Vision-Language Models (VLMs) and an action-conditioned world model to enable agents to propose, simulate, a [episode]
- Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning — Large Language Models (LLMs) struggle to maintain logic consistency in structured knowledge reasoning tasks like Knowledge Graph Question Answering (KGQA) due to representational differences between unstructured and structured knowledge, leading to "Logic Drift" where LLMs output [episode]
- PatchScene: Patch-based Voxel Diffusion for Large-Scale Scene Completion — PatchScene introduces a novel diffusion framework for large-scale LiDAR scene completion that addresses challenges in geometric fidelity, temporal consistency, and computational scalability. [episode]
- RetiWave-Mamba: A Dual-Stream Network for Retinal Disease Detection based on Multi-scale Context and Feature-Adaptive Mamba Projection — Retinal diseases pose a significant global health challenge, requiring early and accurate diagnosis, which necessitates automated analysis of Optical Coherence Tomography (OCT) images to overcome manual interpretation difficulties. [episode]
- Trade-off Functions for DP-SGD with Subsampling based on Random Allocation: Tight Upper and Lower Bounds — Tight closed-form f-DP analysis for DP-SGD with random shuffling provides transparent and interpretable bounds, establishing that meaningful differential privacy can be guaranteed in specific noise regimes. [episode]
- Evaluating the Retrieval Robustness of Large Language Models — Retrieval-augmented generation (RAG) generally enhances large language models’ (LLMs) ability to solve knowledge-intensive tasks, but it may also lead to performance degradation due to imperfect retrieval and the model’s limited ability to leverage retrieved content. [episode]
- GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification — GISTBench introduces a benchmark for evaluating Large Language Models' (LLMs) ability to understand users from their interaction histories in recommendation systems by proposing novel metrics that verify whether predicted interests are factually grounded in behavioral evidence. [episode]
- Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs — MLLMs currently attend to all visual tokens during generation, leading to diluted focus and unnecessary computational overhead, whereas human visual perception is inherently selective. [episode]
- A Language Model from 1913: Pretraining on Historical Text — We introduce TYPEWRITERLM, a 7.24B History language model (LM) trained exclusively on English text predating 1913, addressing challenges in data quality and temporal leakage to create historically grounded models for NLP research. [episode]
- How Far Does a Shared Linear Map Go? Probing Feature-Space Manipulability for Image Editing — Intermediate feature representations represent the backbone for deep neural networks, and this work investigates their geometric structure by applying various input manipulations to determine if mappings from original to manipulated feature maps can be learned. [episode]
- GB-LSR: Local Spectral Decoding with a Learned Global Bandwidth for Arbitrary-Scale Super-Resolution — GB-LSR presents a fixed-grid local spectral representation that utilizes a single trainable global scalar bandwidth to achieve continuous image reconstruction. [episode]
- An Elastic Shape Variational Autoencoder for Skeleton Pose Trajectories — Deep generative models are applied to skeletal trajectories, but standard Variational Autoencoders (VAEs) often allocate capacity to nuisance factors like camera orientation and speed rather than intrinsic shape dynamics. [episode]
- Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity — Automated fact-checking is a crucial task for responsible information ecosystems, and this work challenges the assumption that incorporating visual evidence universally improves performance in multimodal fact-checking by showing that indiscriminate use can reduce accuracy. [episode]
- Counterfactual Evidence Audits Predict LLM-Agent Susceptibility to Ranked Context — LLM agents are susceptible to manipulation through their ranked information streams, and this research establishes that an upstream ranker can steer an agent's final decision by controlling what content it encounters just before acting. [episode]
- SurGe: Improved Surface Geometry in Point Maps — Recent feedforward 3D reconstruction methods predict point maps and estimate global 3D geometry remarkably well, but these predictions still exhibit inaccurate local surface geometry, which is clearly visible qualitatively but only weakly reflected in common metrics. [episode]
- OpenBox: Annotate Any Bounding Boxes in 3D — OpenBox introduces a novel two-stage automatic annotation pipeline that leverages 2D vision foundation models to generate high-quality, open-vocabulary 3D bounding box annotations for vehicles, pedestrians, and cyclists without requiring self-training. [episode]
- How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks — Agentic coding tasks are uniquely expensive, consuming orders of magnitude more tokens than code reasoning and code chat tasks, and models vary substantially in token efficiency across different frontier LLMs. [episode]
- LightLoc++: Sensor-Robust Representation Learning for Efficient Outdoor LiDAR Localization — Scene coordinate regression (SCR) achieves strong performance in outdoor LiDAR localization, but it usually requires scene-specific training that can take days, limiting its practicality for time-sensitive deployment. [episode]
- Understanding Affective Adaptation in Multimodal Foundation Models: Emergent Functional Specialization — Understanding where and how emotions are represented in large-scale foundation models remains an open problem, particularly in multimodal affective settings. [episode]
- ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation — Unified 3D foundation models aim to bridge 3D understanding and generation within a single backbone, but their text–3D interaction remains largely implicit. [episode]
- Escaping the Capacity Ceiling: Routing on the Stiefel Manifold for Bilinear SPD Layers — Cross-domain EEG decoding remains challenging despite advances in Riemannian deep learning, as covariance matrices from different subjects occupy systematically distinct regions of the SPD manifold. [episode]
- Beyond Log-Concavity and Score Regularity: Improved Convergence Bounds for Score-Based Generative Models in W2-distance — Score-based generative models (SGMs) aim to sample from target distributions by learning score functions, and this work presents a novel framework for analyzing their convergence in W2-distance by relaxing stringent assumptions like log-concavity and score regularity. [episode]
- On the Tip of the Tongue: Why LLMs Hallucinate Answers They Can Decode — Language models hallucinate because they fail to integrate internal signals of uncertainty into their output generation process, rather than due to a lack of knowledge. [episode]
- Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models — As large language models are increasingly integrated into multi-agent and humanAI systems, understanding their awareness of both self-context and conversational partners is essential for ensuring reliable performance and robust safety. [episode]
- Power-SMC: Low-Latency Sequence-Level Power Sampling for Training-Free LLM Reasoning — Power-SMC introduces a training-free Sequential Monte Carlo scheme designed to approximate sequence-level power sampling, which sharpens generation toward high-likelihood trajectories without modifying model weights. [episode]
- Specializing Without Forgetting: Analyzing Knowledge Preservation in Multilingual Model Adaptation — Parameter alignment strategies mitigate catastrophic forgetting when specializing multilingual models into language-family experts by systematically comparing five layer-aware methods against unregularized baselines across diverse languages and tasks. [episode]
- Branch-Centric Tokenization and Test-Time Augmentation for Skeleton Generation — Automatic skeleton generation involves predicting both joint positions and skeletal connectivity, and this work introduces branch-centric tokenization and view-augmented generation to achieve state-of-the-art accuracy across diverse inputs. [episode]
- Automatic register identification for the open web using multilingual deep learning — This research introduces a sophisticated suite of multilingual deep learning models designed to identify diverse text varieties, or "web registers" (such as news reports and discussion forums), across 16 different languages. [episode]
- XClipGS: Exact Half-Space Clipping for Medical Volume Gaussian Splatting — Gaussian-splatting proxies enable interactive rendering of volumetric medical scans, but a clipping plane exposes anatomy not constrained by external-view training and intersects primitives that conventional splatting can only keep or drop whole. [episode]
- UniFLM: United Segmentation and Measurement on Fetal Limb Ultrasonic Image — Prenatal ultrasound examination is crucial for assessing fetal limb development and detecting congenital anomalies, yet existing artificial intelligence models often overlook fetal lethal skeletal dysplasias due to data scarcity and lack of a unified framework. [episode]
- On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance — Large Language Models (LLMs) are increasingly used for zero-shot annotation and LLM-as-a-judge tasks, yet their reliability hinges on how model-internalized priors interact with user instructions. [episode]
- Learning the Language of Histopathology Images reveals Prognostic Subgroups in Invasive Lung Adenocarcinoma Patients — Learning the language of histopathology images reveals prognostic subgroups in invasive lung adenocarcinoma patients by treating tissue as a structured biological language. [episode]
- DuetMoE: Coupling Inter- and Intra-Subgroup Robustness for Fair Medical Image Analysis — Medical image segmentation models often perform unevenly across different patient subgroups, and existing fairness methods frequently fail by treating each subgroup as internally homogeneous, which obscures difficult cases within those groups. [episode]
- Navigating the Reality Gap: On-Device Continual Adaptation of ASR for Clinical Telephony — Automatic Speech Recognition (ASR) can significantly reduce documentation burden in clinical workflows, but standard models degrade sharply in real-world telephony settings where noisy audio, dialectal variation, and strict data residency constraints prevent cloud-based adaptatio [episode]
- BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions — BioMol-MQA is a novel question-answering dataset designed to test and improve Large Language Model (LLM) reasoning capabilities over complex, multi-modal bio-molecular interactions. [episode]
- Many Preferences, Few Policies: Compact Portfolios for Multi-Objective LLM Alignment — A principled method for selecting a small portfolio of Large Language Models (LLMs) that captures representative behaviors across heterogeneous user preferences addresses the impracticality of maintaining a separate LLM per user. [episode]
- Assessing Rule Adherence of LLM Adjudicators in Call of Cthulhu TRPG — As LLMs are increasingly deployed as autonomous adjudicators in semi-open textual game environments, robust rule adherence becomes critical when user intent conflicts with system rules. [episode]
- Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs — Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements, and this paper presents a framework for quantizing VLMs for efficient inference on resource-constrained hardware. [episode]
- HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers — Understanding chart and table images is essential for applying vision-language models (VLMs) to real-world document understanding, and this work introduces HakushoBench, a challenging Japanese chart and table VQA benchmark built from 33 governmental white papers. [episode]
- HyperLogic: A Hard, Forward-Authored Chinese Logical Reasoning Benchmark with Execution-Derived Answers — As a fastidious and diligent AI researcher, I have thoroughly analyzed both provided texts concerning LLMEval-Logic and HyperLogic. [episode]
- Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect — Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases score differently with free-text. [episode]
- FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views — FRUC presents a feed-forward 3D Gaussian Splatting framework designed for dynamic scene reconstruction from uncalibrated collaborative driving views, overcoming the limitations of existing methods that require precise spatial calibration and slow per-scene optimization. [episode]
- SANE Schema-aware Natural-language Evaluation of Biological Data — High-throughput microscopy generates large, structured datasets capturing cellular responses to pharmacological perturbations, but accessing these datasets typically requires SQL expertise. [episode]
- Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors — Large language models often produce hallucinated answers that violate prompt-level constraints, and this study investigates whether these failures stem from missing knowledge or from an incorrect inference path. [episode]
- A Multi-Timescale Recursive Self-Improvement Engine for Open-Ended Persona Growth — Long-term persona agents need more than memory; they require a way to keep living in an environment that does not collapse with them. [episode]
- Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching — Large Language Models (LLMs) struggle with factual errors and hallucinations in knowledge-intensive tasks like Knowledge Graph Question Answering (KGQA) due to a semantic gap between structured knowledge graphs and unstructured queries. [episode]
- WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models — Contrastive vision-language models have achieved remarkable progress through largescale pretraining, but this work investigates whether global pretraining alone is sufficient for culture-specific understanding or if further adaptation with natively sourced data can boost performa [episode]
- Closing the Train-Test Gap in World Models for Gradient-Based Planning — World models paired with model predictive control (MPC) can be trained offline on large-scale datasets of expert trajectories and enable generalization to a wide range of planning tasks at inference time. [episode]
- RT-SFT: Text Style Transfer from Non-Parallel Corpora by Roundtrip Translation — This study proposes a novel method for Text Style Transfer (TST) that adapts Large Language Models (LLMs) to transfer text from an arbitrary domain to a target style using only monolingual corpora and roundtrip translation. [episode]
- Sensory-Aware Sequential Recommendation via Review-Distilled Representations — I have meticulously analyzed both provided text excerpts from the paper "Sensory-Aware Sequential Recommendation via Review-Distilled Representations." The information presented in both sections is highly detailed, focusing on a novel framework that bridges unstructured text data [episode]
- Can AI Understand the Language of Origami? — Building AI systems capable of planning and acting in physical environments requires understanding causal mechanisms governing physical processes, which necessitates internal representations that link observations, actions, and environmental changes. [episode]
- Data Synthesis Improves 3D Myotube Instance Segmentation — Myotubes are crucial model systems for studying muscle physiology and disease, but existing 3D segmentation models fail to generalize due to a lack of large annotated datasets. [episode]
- PlantPlotGAN: A Physics-Informed Generative Adversarial Network for Plant Disease Prediction — Monitoring plantations is crucial for crop management and producing healthy harvests, but limited data from plant disease signals hampers prediction models due to unbalanced datasets. [episode]
- Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models — Moral actions are judged by their context, and this framework models how context shapes the acceptability of ambiguous actions by integrating a probabilistic context learner with LLM-based semantic abstraction and human moral evaluations. [episode]
- VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation — Autoregressive (AR) models for image generation rely critically on visual tokenizers (VT), and this paper introduces VTBench, a comprehensive benchmark designed to systematically evaluate VTs across three core tasks—Image Reconstruction, Detail Preservation, and Text Preservati [episode]
- When Are Concepts Erased From Diffusion Models? — In concept erasure, a model is modified to selectively prevent it from generating a target concept, and this research investigates whether such methods truly remove the target knowledge or merely redirect generation. [episode]
- How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation — Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale, but most accurate methods depend on GPU-intensive inference or proprietary APIs, making them inaccessible to resource-constrained researchers. [episode]
- Useful Features, Backward Scores: OOD in Language-Model Trajectories — Recent white-box out-of-distribution (OOD) detection methods for large language models are structurally confounded by sequence length, leading to near-chance performance when evaluated under length constraints. [episode]
- Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning — Pathology foundation models (PFMs) offer generalizable representations for whole-slide image (WSI) analysis, yet their clinical adoption remains limited because their predictions lack reliable confidence estimates and no single PFM is universally best across tasks. [episode]
- VIDiff: Translating Videos via Multi-Modal Instructions with Diffusion Models — Diffusion models have achieved significant success in image and video generation, motivating research into video editing tasks guided by natural language instructions. [episode]
- Spectral Alignment in Forward-Backward Representations via Temporal Abstraction — Spectral alignment in Forward-Backward representations via temporal abstraction addresses a fundamental mismatch between low-rank factorization and high-rank transition dynamics in continuous environments, demonstrating that temporal abstraction acts as a low-pass filter to suppr [episode]
- Differential Privacy as a Perk: Federated Learning over Multiple-Access Fading Channels with a Multi-Antenna Base Station — Federated Learning (FL) is a distributed learning paradigm that preserves privacy by eliminating raw data exchange, and this work investigates how inherent channel noise in over-the-air federated learning (AirFL) can be leveraged to achieve differential privacy without resorting [episode]
- Autoregressive Frontier Expansion: Growing Trees with Graph Machine Learning — Tree-like branching structures are common in nature and their structural modeling is central to understanding how biological systems function, making realistic generative models valuable for simulation and data augmentation. [episode]
- Low-Frequency Shortcuts in Texture-Driven Visual Learning — Texture-driven domains suffer from low-frequency shortcuts, where a small number of low-frequency components (LFCs) dominate model decisions despite classification information residing in higher frequencies. [episode]
- Asymptotic Performance of Time-Varying Bayesian Optimization — Time-Varying Bayesian Optimization (TVBO) is a framework for optimizing expensive, noisy, time-varying black-box functions, and this paper provides theoretical upper bounds and algorithm-independent lower bounds for its cumulative regret across various temporal kernel classes. [episode]
- ZeBROD: Zero-Retraining Based Recognition and Object Detection Framework — Object detection often suffers from catastrophic forgetting when new products are introduced, necessitating costly and time-consuming model retraining. [episode]
- The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems? — Multimodal Large Language Models (MLLMs) are being tested on simple visual perception problems to determine if they match human capabilities, and this research introduces Percept-V, a dataset designed to isolate these foundational visual skills. [episode]
- Flow Map Denoisers: Traversing the Distortion-Perception Plane for Inverse Problems — Flow map models implicitly define a one-parameter family of denoisers that continuously spans the distortion-perception (DP) frontier, enabling continuous control over image restoration quality in inverse problems. [episode]
- Hint-Guided Diversified Policy Optimization for LLM Reasoning — Recent developments in Large Language Models (LLMs) have showcased impressive reasoning capabilities, with Reinforcement Learning with Verifiable Rewards (RLVR) being a promising enhancement strategy. [episode]
- Rank-Turbulence Delta and Interpretable Approaches to Stylometric Delta Metrics — This article introduces two novel measures for authorship attribution—Rank-Turbulence Delta and Jensen–Shannon Delta—which generalize Burrows’s classical Delta by employing distance functions derived from probabilistic distributions, thereby providing a more interpretable [episode]
- A Sobel-Gradient MLP Baseline for Handwritten Character Recognition — A multilayer perceptron trained exclusively on first-order Sobel edge maps demonstrates strong performance in handwritten character recognition, suggesting that stroke contours alone capture significant class-discriminative information. [episode]
- Stochastic Engrams for Efficient Continual Learning — The ability to learn continuously in artificial neural networks (ANNs) is often limited by catastrophic forgetting, a phenomenon in which new knowledge becomes dominant. [episode]
- Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality — Memristor-based analog compute-in-memory (CIM) architectures offer high energy efficiency for Large Language Models (LLMs), but intrinsic non-idealities introduce noise that significantly impacts reasoning capability. [episode]
- A Unified BERT-CNN-BiLSTM Framework for Simultaneous Headline Classification and Sentiment Analysis of Bangla News — A unified framework for Bangla news headline classification and sentiment analysis has been proposed by combining BERT, CNN, and BiLSTM to simultaneously capture both aspects of news content. [episode]
- Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions — Using behavioral science, health interventions focus on behavior change by providing a framework to help patients acquire and maintain healthy habits that improve medical outcomes. [episode]
- WAM-OPD: Joint Video-Action Supervision for World Action Model Post-Training with On-Policy Distillation — World action models (WAMs) couple visual future prediction with robot action generation, but accelerated students can lose task capabilities during distillation and later encounter states that are poorly represented by offline data. [episode]
- The exponential distribution of the order of demonstrative, numeral, adjective and noun — The frequency of preferred orders for noun phrases formed by demonstrative, numeral, adjective, and noun has been investigated to determine if an exponential or power law distribution better models their actual distribution. [episode]
- Textualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild — Textualization of modalities augments data with emotional cues to help large language models (LLMs) encode interconnections between all modalities in a shared text space, offering an alternative to traditional feature-based models for compound emotion recognition in real-world vi [episode]
- Estimating prevalence with precision and accuracy — Prevalence estimation in classification tasks requires methods to adjust for training data bias and quantify uncertainty, and this paper proposes Precise Quantifier (PQ), a Bayesian method that achieves narrower prediction intervals than existing quantifiers while maintaining wel [episode]
- Error Propagation in Dynamic Programming: From Stochastic Control to American Option Pricing — This paper investigates theoretical and methodological foundations for stochastic optimal control (SOC) in discrete time, developing a framework to rigorously analyze how errors propagate backward through dynamic programming approximations. [episode]
- Query-aware routing for Cross-lingual performance gains in Encoders —
- Seeing, Saying, but Not Using: From Reportable Spatial Facts to Usable States in Multimodal Large Language Models —
- Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty —
- Evaluating VQA in Vision Language Models using Cooperative Principles —
- Found but Not Read: When Extracted Text Closes the Retrieval-Reading Gap in Document Vision-Language Models —
- Misinformation Without Triggers: From Factual Answers to Downstream Decisions —
- Revealing Epistemic Uncertainty in MLLMs via Causal-Invariant Masking —
- ViTok: Improving Dense Semantics in AM-RADIO-Style Multi-Teacher Distillation with PHI-S and Masked Image Modelling —
- Frequency Is Not Sensitivity Identifying Safety-Sensitive Experts in Sparse MoE LLM —
- Probe the Harness: Setup Checks for Stale-Data RL Comparisons in Language Models —
- Custom Forcing: Training-Free Subject Customization for Autoregressive Video Generation —
- Output Language Confusion under Multilingual Prompt Contamination —
- Kinematics-Induced Multimodal 3D Human Pose Estimation with Subject-Level Privacy —
- Continual Graph Memory for Mathematical Research Agents —
- When Predicting Nothing Beats SAM 3: Revisiting Evaluation in Video Object Segmentation —
- Enhancing Biomedical Named Entity Recognition via Multiple Programming Languages Instruction Tuning and Ensemble Method —
- Understanding Trajectory Heterogeneity in Federated World Model Learning —
- TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows —
- Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards —
- A Guideline-Augmented Multi-Agent Framework for Schema-as-Code Biomedical Named Entity Recognition —
- OLMo-Detect: A Multi-Stage, Confounder-Controlled Benchmark for Membership Inference on Large Language Models —
- Sentry: Learning to Recover from LLM Agent Failures at Test Time —
- OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination —
- Recursive Self-Improvement in Unified Multimodal Models —
- Rethinking Fixed Temporal Grids: Frequency-Disentangled Motion Generation —
- RYOPO: Bringing End-to-End Category-Level Object Pose Estimation into Real Time —
- OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection —
- Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations —
- ReSCUE: Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation —
- Verifiable, Articulable, and Tacit Components of Preference —
- Tailoring the Quantization Space for 1-Bit KV Cache Compression —
- CrowdOcc: Monocular Semantic Scene Completion for Quadruped Robots in Crowded Indoor Environments —
- Adaptive Second-Order Solvers for Fast Stochastic Diffusion Sampling —
- WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites —
- HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning —
- Parasitic Co-Denoising: Unlocking 3D Human Motion Generation in a Frozen Video Diffusion Model —
- BeeWhere: Segmenting Bumble Bee Colonies to Quantify Behavioral Effects —
- The Geometry of Knowledge Accessibility in Large Language Models —
- HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation —
- Where to Look Is Not How to Fix: Pre-Denoising Diagnostics and Modality-Dependent Control in Diffusion Composition —
- Unmasking Propaganda: A Comparative Analysis of Masked and Causal Language Models —
- An automated pipeline for standardised speech-unit annotation in spontaneous dialogue —
- NegT2IBench: When Negation Changes the Picture. A Polarity Benchmark for Text-to-Image Models —
- Peer Influence across Heterogeneous AI Models —
- Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning —
- Invariance of Clustering Operations in Causal Effect Identification —
- Ask, Relax, or Act? Evaluating Actionable Indeterminacy in LLM Preference Reasoning —
- A Benchmark for Spatially Grounded Gesture Generation —
- Emergent Structure in the Marginal Attention Space of Language Models —
- Ontological Instability and Statistical Amplification: The Paradox of "Humanizing" LLM-Generated Text —
- Building Interpretable Feature Representations for Resume-Vacancy Matching by Distilling Production LLM Signals —
- How to Find and Reuse Policies for Continuous Adaptation in Lifelong Reinforcement Learning —
- In-Distribution Forcing for Long Video Generation at Test Time —
- Foresight: planning future perception in streaming VLMs without retraining —
- Investigating the Role of Reasoning-Language Alignment in Monolingual Retrieval-Augmented Generation —
- Behavior Pack Optimization for Video MLLM Post-Training —
- CalCErt: Bin-wise Certification of Confidence Calibration in Medical Image Classification —
- Does Physics Live in the Activations? Localizing Physical Quantities in Video Diffusion Models —
- Predicting Steering Vectors and Adapter Weights for Few-Shot Author-Style Transfer —
- Geometry-Aligned Semantic Matching for Cross-Modal Planar Image Registration —
- Hindsight-Guided Rationale Distillation for Rare Disease Diagnosis —
- Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective —
- Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case —
- PocketSplat: Mobile Gaussian Reconstruction via World-Space Latent Allocatio —
- Bridging Research and Practice: A Systematic Evaluation of Generalist and Dermatology-Specific Models in Clinical Skin Lesion Classification —
- Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It —
- KV squared: A Self-Refining KV Cache —
- Predicting and Repairing Merge Collapse in Large Language Models —
- Predictively Oriented Gaussian Process Posteriors —
- Contextual Flow Matching: Adaptive Step Selection in Flow Models for Efficient Visual Generation —
- StanceEval 2026: The Second Stance Detection Shared Task —
- VisionMX: Unlocking Microscaling Post-Training Quantization for Vision Models —
- VDOT++: Unified Few-Step Video Generation via Unbalanced Optimal Transport Distillation —
- AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning —
- Uncertainty as a Proxy for Semantic Correctness in Diffusion-Based Medical Image Synthesis —
- Collective Bias Mitigation via Model Routing and Collaboration —
- EmbPASS: Towards Cross-Embodiment Open Panoramic Segmentation —
- COSMI: COmpositional Synthesis of Multi-object Interactions —
- Consecutive Posterior Fusion for Diffusive Recovery of Unobservable Image Structures —
- Shrome at Touch'e: Soft-Vote Ensembling and Counter-Causal Augmentation for Causality Extraction —
- Moving Forward with Video Saliency: A New Dataset and Benchmark where Motion Matters —
- T3lescope: Arbitrary-Resolution High-Fidelity Generative Surface Reconstruction from Images —
- SDECast: Probabilistic Weather Forecasting in Continuous Time with Neural SDEs —
- DAWIS: Data Assimilation with Windowed Inverse Sampling via Multitask Interpolants —
- To Jev or Not? Evaluating the Accuracy and Efficiency of Structured Decision Models for Hate-Speech Moderation —
- SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models —
- A Fully Automatic Pipeline for 3D Dendrite Instance Segmentation in SBF-SEM —
- Multilingual GSM-Symbolic: What determines capability transfer across languages? —
- LAS-CLIP: A Lightweight Adapter Steering Approach for CLIP's Visual Encoder —
- EVEWorld: Physical Evolution Supervision for Embodied World Models —
- Interpretable Deepfake Detection in Videos via Explicit Forensic Features and Temporal Modeling —
- Benchmarking candidate coverage and rejection policy transfer in typed decision models —
- From Patching to Pruning Visual Computation in Vision Language Models —
- Native Action-Prior Learning from Videos for World Action Models —
- Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning —
- Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning —
- ForestQuery: Boundary-Aware and Spatially Anchored Query Learning for Unified Forest Point Cloud Segmentation —
- Iterating Consistency Models: Stability, Error Bounds and Noise Schedules —
- CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation —
- OuroReward: Sequential Reward Scheduling for Reinforcement Learning in Text-to-3D Generation —
- Depth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth Estimation —
- ChromaGS: Text-Driven Semantic Editing of 4D Gaussian Avatars —
- Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally —
- A Near-Zero Monitor Readout Is Not Evidence of Behavioral Control —
- When Is Accuracy Evidence? A Unified Theory of Generalisation, Validation, and Information Fusion —
- Preserving Anatomical Continuity: Three-Stage Pipeline for Colon Segmentation in 3D Abdominal CT Scans —
- A Vision-Language Model (VLM)-based Pipeline for End-to-End Procedural Modeling of Field-Grown Maize from Point Clouds —
- UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation —
- Fed-ADApt: Federated Anytime Depth Adaptation for Resource-Aware Medical Image Segmentation —
- Single or Multiple Policies for Phase-Structured Reinforcement Learning? —
- Single-Pass Uncertainty Heads for Claim-Level Hallucination Detection in Persian Medical Language Models —
- AREX: Affine-Residual Exponential Integrator for Few-Step Sampling in Flow Matching —
- Below what training size do deep tabular generators stop beating trivial baselines? A preregistered benchmark on a size ladder of clinical and standard datasets —
- Certified Mechanistic Edits: Behavioral Guarantees for Skill Removal and Preservation —
- Weave Forcing: Compositional Memory Routing for Interactive Long Video Generation —
- ProgressNet: Sketching and Prompting with a Frozen Text-to-Image Model —
- Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation —
- Reasoning Models Are Accurate but Unsound on Identification —
- Feedforward Novel View Synthesis for Heterogeneous Cameras —
- Structured Composition of Verifiable Atomic Insights for Table-to-Report Generation —
- Divergence controls entropy in distillation —
- Author Representation Strategies for Zero-Shot Authorship Attribution: A Comparative Study of LLM-Based and Embedding-Based Approaches —
- DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation —
- Writerslogic at PAN 2026: Process over Content for Robust Detection under Domain Shift —
- Writerslogic at the CLEF 2026 SimpleText Track: Multi-Candidate LLM Simplification and Stacked Complexity Spotting —
- Rethinking What to Cache in Few-Step Diffusion Transformers: Solver-Aware Target Selection —
- ManifoldSplat: Language-Guided Semantic Shape Editing of 3D Gaussian Head Avatars —
- DEPICT: Scoring Text-to-Image Alignment by Answer Agreement —
- Low-Cost Video--Time Priors as a Strong Baseline for EEG--fNIRS Emotion Regression on Familiar Videos —
- UniIntervene++: An Adaptive Intervention Agent for Efficient Real-World Reinforcement Learning —
- FALCON: A Model and Dataset Agnostic Framework for Synthetic Data Generation for NL2SQL Pairs —
- World Embedding Benchmark —
- LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation —
- Broken scale symmetries in undercomplete linear autoencoders —
- When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity —
- Amortized Structured Stochastic Variational Inference for Gaussian Process Latent Variable Models —
- On-Board Anomaly Detection for Efficient Marine Environmental Monitoring —
- ProAR: Learning Prospective Reasoning with Autoregressive Video Models —
- Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models —
- Planning to Learn —
- Simulation-Free Learning of Population Dynamics with Wasserstein Lagrangian Residuals —
- SigLIP2 for aerial fire risk classification —
- FlowHMR: Physically Plausible Motion Capture from Video —
- Language Models that Play Chess and Explain Their Moves —
- Decoding the Functional Roles of Register and High-Norm Patch Tokens in Vision Transformers —
- RNADyn: A Benchmark for Generating and Understanding RNA Dynamics —
- What Should World Models Forget? Stratified Retention for Continual Adaptation —
- 4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes —
- MoSE3: Learning World-Space SE(3) at Every Pixel —
- Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis —
- Nearest-neighbour baselines for fingerprint prediction from MS/MS spectra under different assumptions —
- Counterfactual Predictions in Scientific Emulators Without Controlled Experiments —
- TRACE: A Reproducible Benchmark for Electricity Price Forecasting with Official Operational Text —
- Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses —
- From Mathematical to Executable Certificates for Machine Unlearning —
- HakemBench: A Turkish Benchmark of Typed Decisions —
- EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling —
- DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents —
- SCION: Scene Composition with Instanced Neural Primitives —
- SCOPE-4D: Endoscopic 4D Geometry Foundation Models —
- Does Every User Need a Private LoRA? Decoupling Personalization from Per-User Adaptation —
- Conformal Prediction for Time Series with Deep Sequence Models —
- Lexicographic Multi-Objective On-Policy Distillation —
- Confidence-Controlled XAI Auditing for Pedestrian Detection under Domain Shift —
- Co-design Gym: A Unified Benchmark for Embodiment-Policy Co-optimization —
- EviDent-CBCT: Evidence-Bottlenecked Report Generation from Dental CBCT under Non-Exhaustive Report Supervision —
- FactorSplat: Appearance-Controllable Gaussian Proxies for Medical Volume Rendering —
- Octrees as an Explicit 3D Language —
- Hesitation Has a Geometry: Entropy-Trained Hyperbolic Probes for Sparse Activation Steering —
- Trained Agentic Context Management —
- An AI-Based Multi-Stage Approach for Androgenetic Alopecia Assessment from Low-Magnification Scalp Images —
- Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents —
- Learning Style, Forgetting Semantics: A Case Study of SFT and RFT on Classification Tasks —
- Are you Synthesizing or Recalling? Evaluating LLMs on Algorithmic Code Retrieval —
- Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs —
- A Simulation-Grounded Agentic VLM Framework for Wildfire Monitoring and Reporting —
- FinDialogLens: Event Extraction over Multi-Party Dialogue for Missed-Trade Identification in Financial Chatrooms —
- CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking —
- Capability Scaling-Down Laws for LLM Compression —
- APDMem: Agent-Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory —
- From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders —
- Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement —
- DeepStratNet: A Context-Aware Coordinate Regression Framework for Seismic Horizon Tracking under Sparse Labels —
- Multi-Fidelity Policy Gradients Stabilize Data-Scarce Reinforcement Learning —
- MeshQuery: Agentic Seam Planning for UV Parametrization —
- World Action Modeling with Progressive Visual Planning —
- Post-Training Quantization of Autoregressive Weather Models —
- From Fragments to Global Maps: Learning Vectorized Map Aggregation with Large Language Models —
- Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory —
- A generative-informed neuro-symbolic framework for syntactic ambiguity resolution: Evidence from Arabic DPs —
- ENCORE: Exact Non-equilibrium COntrol with Replica Exchange for Diffusion Generation —
- Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks —
- Test-time Multi-agent Coordination by Decomposed Value Gradient Flow —
- Oracle headroom without signal: null-calibrated evaluation of candidate selection for thermal heart rate estimation —
- DAGS: Disentangled Appearance-and-Geometry Steering of a Frozen Image DiT for Temporally Stabilized Generative Rendering —
- High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning —
- Physical AI Smart Spaces: A Large-Scale Benchmark for Multi-Camera 3D Perception in Smart Spaces —
- How Causality Bridges the Semantic Gap —
- GRAFT: Growing Agglomerative Foundation Models via Continual Teacher Distillation —
- Scale-Recursive Rectified Flows for Few-Step Precipitation Ensembles —
- Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation —
- VERSE: Verified Self-Evolving Optimizer for Agent Harnesses —
- Imagine the Future, Internalize the Gist: Efficient VLA Reasoning via Internalized Spatiotemporal Imagination —
- Capturing Dynamics: The 4D Facial Expression Intensity Dataset —
- SpectralCache: Accelerating Diffusion-Based World Models via Spectral Feature Caching —
- Mind the Refinement Gap: When Safe High-Level Robot Plans Produce Unsafe Executions —
- Generalization Bounds for Flow-matching Generative Models for Intrinsically Low-dimensional Data —
- Large Language Continuous Diffusion Models —
- CHASE-VLA: Post-Training Quantization Framework for Vision-Language-Action Models with Chunk-Aware Scale Estimation —
- LEAP: Learning Efficient Action Proposals For LLM Agents —
- Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves —
- Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation —
- Silent Dissent: LLM Agents That Yield to the Majority Still Represent Their Original Premise —
- WakeKV: Reactive, Reversible KV Residency for Heads That Change Their Minds —
- Differential Privacy of Gradient Descent on Perturbed Objectives —
- Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis —
- SymRegFlow: Symmetry-Regularized Flow Matching for Video World Models —
- TPBench: A Turning-Point Benchmark for Dialogue Compression —
- Inner Momentum for Differentially Private Muon —
- Beyond Correctness: Resolving Underspecification in Agentic Text-to-SQL —
- EpiWorld: Grounding LLM Policy Agents in Epidemiological World Models —
- Correcting Guided Diffusion Trajectories with Spectral Alignment —
- FiberGeoText: A Vision-Language Model for Population- Level Organization of Superficial White Matter —
- When History Fails to Become Experience: Action Calibration in Language Agents —
- AptMQL-Bench: From Text-to-SQL to Text-to-MQL via Access-Pattern Schema Design and Data-Preserving Migration —
- Nearly Optimal Fixed-Confidence Best-Arm Identification with 1-Bit Feedback —
- Improving Atomic-Fact Recall via Focused Views in Unstructured Knowledge Editing —
- Automatic Evaluation of Mental Health Stigma in Online Communication —
- TRAC: Trajectory-aware Reuse and Adaptive Correction for Efficient Autoregressive Video Generation —
- OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation —
- Hold-Out Scoring for Efficient Gaussian DAG Learning —
- Muon Learns Facts Better: Understanding the Role of Spectral Orthogonalization —
- FUSEye: Training-Light Fisheye Detection with Overlapping Views and Zero-Initialized Adapters —
- VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning —
- ROUTEAUDIT: Interaction-Aware Identification for Budgeted Multi-Verifier Routing —
- Text-Centric Post-Training for Omni-Modal Reasoning —
- TerrainForge: Physics-Grounded road geometry Editing for Counterfactual Autonomous Driving —
- FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training —
- Clinical Concept Centers in LLMs —
- To Explore The Strange New World Beyond Data Distribution: System Behavior, Causality Tax, and Non-causal Base Model —
- How Robust Is Multimodal Claim Verification to LLM Rewriting? —
- Adaptive Mutual Distillation for Balanced Multi-Task Post-Training of Large Language Models —
- NeuroLens: Learning Latent Embeddings of Neural Semantics from Chronic Recordings —
- ConvoDrift: A Multi-Turn Conversational Dataset for Modeling Stylistic Tone Evolution —
Important terms
- RNADyn
- A new benchmark designed to evaluate a model's ability to generate and correctly interpret complex RNA dynamics over time, helping build better biological prediction models.
- Autoregressive frontier expansion
- A method using graph machine learning to grow trees, suggesting a way to structure and manage very complex data relationships efficiently.
- BioMol-MQA
- A multi-modal question answering dataset specifically built to test if large language models can truly reason about complex interactions between different biological molecules.
- Agent-to-agent theory of mind
- Research testing if LLMs can recognize when another AI model is aware of information they themselves are not, which is key for understanding sophisticated agent interactions.
- Interpretability in moral contexts
- Using probabilistic clustering and LLMs to learn human ethical frameworks, helping us understand the nuances of how people make moral judgments.