AI papers — 2026-10-06
Automating molecular dynamics simulations for proteins promises to speed up the discovery of new drug candidates by allowing more complex simulations than traditional methods permit. NAMD-Agent was explored as a system designed to automate these simulations, and it learns latent protein languages through autoregressive generation, which teaches it the underlying rules of how proteins behave in a sequence of steps.
The core mechanism uses three tiers of computation within transformer architectures and brain models to manage the complexity of this learning process. This hierarchical structure helps in understanding how different levels of abstraction contribute to predicting protein movement. Furthermore, Brain-IT-VQA was looked at, which takes brain signals and translates them into answers, suggesting a pathway for interpreting complex biological data streams.
A key finding relates to how hidden states function as value gradients within recurrent policies, which provides a mathematical framework for controlling the simulation's direction. This is complemented by DynaMiCS, a technique that fine-tunes large language models using performance constraints through dynamic mixtures of data, allowing the model to guide the model toward better simulation results. Finally, studying the soupability of documents in state space models was examined to see if this framework could apply to understanding protein dynamics in a more abstract state space.
The most significant development today involves EulerESG, which automates ESG disclosure analysis using large language models. This is important because it directly addresses the growing need to efficiently process and understand complex environmental, social, and governance reporting across vast amounts of text. The study shows that EulerESG can automate this analysis by leveraging LLMs to extract key information from these disclosures.
Another piece of work that warrants attention is ATOD, which establishes an evaluation framework and benchmark for agentic task-oriented dialogue systems. This matters because it provides a standardized way to measure how well these systems perform when they are tasked with completing specific goals through interaction. The framework allows researchers to compare different approaches fairly against a common set of criteria.
Moving down the list, there is research on Adaptive Information Control for Search-Augmented LLM Reasoning. This work explores how to dynamically adjust information control within search-augmented language models to improve their reasoning capabilities when facing uncertainty in the retrieved data. This helps systems make better decisions when they are relying on external search results.
Then there is MDKeyChunker, which investigates what specific textual elements one large language model considers valuable for markdown retrieval within document chunks. This is a finer point of work because it seeks to understand the granular structure that LLMs prioritize when indexing information. This finding connects to how other models might process similar data structures.
The most significant development concerns how models handle long sequences, specifically showing that incorporating structure-originated reasoning data enhances their ability to maintain coherence over extended contexts. This matters because current large language models often lose track of information as conversations grow longer, and this work suggests a structural approach is key to improving that memory.
A study explored using a method called pi squared which involves structure-originated reasoning data to boost long-context reasoning ability in large language models. The findings indicated that this structured input significantly improved the models' performance when dealing with very long texts compared to standard training methods. This finding builds on previous work where asynchronous on-policy self-distillation under positive rollouts was used to improve general reasoning ability, setting a foundation for how structured data can guide learning within those reinforcement learning frameworks.
Another piece of research focused on narrative forecasting as a metric for tension in large language model storytelling. This work attempted to quantify the emotional weight or conflict present in generated narratives by measuring how well they predicted future narrative tension. This is less directly about reasoning but provides a new way to evaluate the quality of creative text generation, which complements efforts to improve logical coherence.
Furthermore, there is ongoing work addressing the limitations of retrieval-augmented generation systems when they rely solely on factual grounding. Researchers are investigating ways for these systems to represent diverse opinions rather than just retrieving verifiable facts, suggesting a need for richer representation methods beyond simple document fetching. This contrasts with the focus on structural reasoning data, as both aim to deepen model understanding in different ways.
The most critical finding from today's work involves the validation-gated causal interventions used to interpret high-stakes large language model behavior. This matters because it helps us understand when these models might be exhibiting problematic patterns like suicidal ideation. We tested this method on a case study related to suicidality detection, and the results showed that by selectively intervening in specific causal pathways within the model's decision-making process, we could gain clearer insight into its underlying reasoning.
This intervention work builds upon earlier efforts to understand model behavior, showing how targeted manipulation can reveal hidden mechanisms. We also looked at compression techniques and found that fixed retrieval augmentation compression can indeed distort how readers compare different information sources, which is a significant limitation for reliable evaluation.
Another piece of research explored narrative generation for ultra-fine entity typing using narrative-UFET, which attempts to improve how language models categorize specific entities within stories. This is important because better entity typing is crucial for more nuanced content understanding. Following that, we examined character-grounded multi-agent story generation designed for long-form narratives, which moves beyond simple personas to create richer plot structures.
Finally, we touched upon the practical challenges of retrieval at massive scales, specifically investigating whether language models can actually retrieve information effectively when drowning in documents at million token scale. This work suggests that the sheer volume of context presents a real hurdle for reliable in-context learning and retrieval capabilities.
The work on safe inference time alignment is particularly important because it directly addresses the reliability of deploying large language models in real-time decision-making scenarios. Researchers explored using Lagrangian reward augmentation to ensure that the model's outputs align with desired safety constraints during inference. This method involves modifying the loss function to incorporate a penalty related to constraint violation, which helps guide the model toward safer choices without needing extensive human labeling for every possible failure mode.
Co-LMLM investigates continuous query limited memory language models, focusing on how these architectures maintain relevant context over long interactions. The findings suggest that by limiting the memory access during querying, these models can manage complex conversational states more effectively than purely expansive memory systems. This relates to the challenges seen in medical foundation models, where convergence under label supervision appears less robust than previously hoped.
MemArena introduces an ego-centric benchmark specifically designed for evaluating on-device agentic personal memory assistants at scale. By focusing on how these assistants manage self-relevant information, this work provides a necessary testing ground for practical deployment scenarios. This contrasts with the more abstract explorations into form generation, such as FormuEvo, which uses LLM guidance to discover solver-efficient mixed-integer programming formulations.
The study on dimensionality and measurement precision in multiple-choice subsets of human language understanding focuses on how different dimensional representations affect accuracy in specific tasks. This work is foundational because it establishes the necessary metrics for evaluating performance when dealing with structured knowledge retrieval, which is relevant when assessing the outputs from models like Co-LMLM.
Finally, research into fine-grained emotion classification from mobile app reviews using large language models examines how well these models capture subtle affective states in user feedback. This helps bridge the gap between general language understanding and nuanced sentiment analysis in consumer applications.
The work on proxy confidence is particularly important because it provides a way to audit black-box large language model agents by using surrogate log-probabilities. This method helps us understand how these agents are making decisions without needing to open their internal workings.
We saw an investigation into when evidence changes, which looked at evaluating memory repair and re-reading in language model agents. This study explored how these models handle situations where the underlying facts they rely on are updated, essentially testing their ability to correct errors by re-examining information.
Then there was the work on general decision models that benchmarks and insights beyond Jev, which aimed to provide broader perspectives on decision-making capabilities. This piece sought to move past narrow metrics to see how these models perform in more complex scenarios.
This is connected to the training of numerical intelligence via auto-diagnosis and skill discovery, which involved training systems through auto-diagnosis and skill discovery processes. This suggests a path toward developing models that can learn quantitative reasoning directly from experience rather than just pattern matching.
The work on BAIBAICHUCHU is particularly important because it tackles the core question of whether maximum possible profit from investor text can be predicted, which has direct implications for financial modeling. Researchers explored this by developing a method to predict this profit using investor text, and the results showed that the approach performed well. This was built upon earlier work that looked at how language models create synthetic personas, suggesting that understanding textual representation is key to these predictive tasks.
Another significant piece of research involved exploring whether language models require a trainable input embedding table by testing fixed minimal token codes at a one point seven billion parameter scale. The findings suggest that this approach is viable, which simplifies the architecture for smaller models and moves away from needing extensive training data just for basic representation learning. This idea connects to how interaction-aware circuit discovery in language models is being used to understand their internal workings better.
The exploration of interaction-aware circuit discovery in language models aimed to uncover the specific pathways within these large language models that govern their behavior, which is a deeper dive into understanding the model's reasoning process. This contrasts with the work on extending frozen language models beyond their context window, which focuses on increasing the amount of information a model can handle without retraining.
Furthermore, ideas around grounding scientific ideation through orchestrating agents were examined to see how to make language models generate scientifically sound ideas by connecting them to external knowledge bases. This is related to representation-aligned auxiliary supervision for language model adaptation, which seeks to guide the model's learning process more effectively during fine-tuning. Finally, the SEER project focused on self-evolving event reasoning and retrieval for time series forecasting, which represents a different kind of prediction task entirely.
The most significant finding today relates to how long context windows affect the comprehension of online discussion threads. We tested this using LongSocialBench, and the results showed that models with longer context windows demonstrated a better ability to understand complex online conversations compared to shorter ones. This suggests that simply having more memory allows language models to grasp nuanced social dynamics within extended text.
This relates back to work on representation control over self-report and behavior coherence in LLM risk-taking, which explored how controlling the model's internal representations can lead to more consistent and reliable outputs when it makes decisions. A related effort involved PB-GRPO, which focused on learning socially adaptive LLM agents through persona-driven simulations using preference-batched GRPO. This simulation approach seems crucial because it allows the agent to learn appropriate social behaviors in a controlled setting before deployment.
We also looked at how to guide mixture-of-experts training using ExpertMuon-Compass, which aimed at alignment by adjusting step sizes based on expert performance. This is important because it helps ensure that different specialized parts of the model contribute effectively during training. Furthermore, Trajectory-Derived Confidence for Reliable, Resource-Aware Clinical Text-to-SQL Agents investigated how to derive confidence scores from trajectories to make these agents more reliable when translating text into database queries.
Finally, Principled Top-k Selection for Language Models with Hybrid Gradients explored using hybrid gradients alongside principled top k selection to improve the quality of model outputs. This work suggests that combining different selection strategies can lead to better overall performance in language modeling tasks.
Today's papers
- Automating MD simulations for Proteins using Large language Models: NAMD-Agent. [paper] [episode]
- Learning Latent Protein Languages for Autoregressive Generation. [paper]
- Three tiers of computation in transformers and in brain architectures. [paper] [episode]
- Brain-IT-VQA: From Brain Signals to Answers. [paper] [episode]
- The Score Is Not the Structure: Brain Alignment and Cross-Lingual Transfer. [paper]
- Hidden States as Value Gradients: The Pontryagin Structure of Recurrent Policies. [paper] [episode]
- DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures. [paper] [episode]
- Studying the Soupability of Documents in State Space Models. [paper] [episode]
- Towards Probabilistic Question Answering Over Tabular Data. [paper] [episode]
- Cross-Lingual Summarization as a Black-Box Watermark Removal Attack. [paper] [episode]
- EulerESG: Automating ESG Disclosure Analysis with LLMs. [paper] [episode]
- ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems. [paper] [episode]
- Adaptive Information Control for Search-Augmented LLM Reasoning. [paper] [episode]
- LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers. [paper] [episode]
- 3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models. [paper] [episode]
- MDKeyChunker: What Does One LLM Call per Chunk Buy for Markdown Retrieval?. [paper] [episode]
- pi squared: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models. [paper] [episode]
- Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling. [paper] [episode]
- Retrieval-Augmented Generation Must Move Beyond Factual Grounding to Represent Diverse Opinions. [paper] [episode]
- Improving Reasoning Ability via Asynchronous On-Policy Self-Distillation under Positive Rollouts. [paper] [episode]
- K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs. [paper] [episode]
- Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes. [paper] [episode]
- Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter. [paper] [episode]
- Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR. [paper] [episode]
- LEAF: Growing Trees Without Branching for Speech-Aware Large Language Model Post-Training. [paper] [episode]
- Validation-Gated Causal Interventions for Interpreting High-Stakes Large Language Model Behavior: A Case Study in Suicidality Detection. [paper] [episode]
- Compression Is Not Evaluation-Neutral: Fixed RAG Compression Can Distort Reader Comparisons. [paper] [episode]
- Narrative-UFET: Narrative Generation for Ultra-Fine Entity Typing. [paper] [episode]
- From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives. [paper] [episode]
- Office Comprehension Benchmark. [paper] [episode]
- Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale. [paper] [episode]
- Subliminal Clocks: Latent Time Modelling in Diffusion Language Models. [paper] [episode]
- Safe Inference-Time Alignment via Lagrangian Reward Augmentation. [paper] [episode]
- Co-LMLM: Continuous-Query Limited Memory Language Models. [paper] [episode]
- Medical foundation models converge less under label supervision. [paper] [episode]
- Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset. [paper] [episode]
- MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale. [paper] [episode]
- FormuEvo: LLM-Guided Evolution for Discovering Solver-Efficient Mixed-Integer Programming Formulations. [paper] [episode]
- Do Motion Tokenizers for Co-Speech Gesture Generation Encode Gesture Semantics?. [paper]
- Fine-Grained Emotion Classification from Mobile App Reviews: An Empirical Study with Large Language Models. [paper]
- Same Output, Different Gold: Measuring How Reference Choice Moves a Multilingual Benchmark Score. [paper]
- OncoNoteBERT: A Foundation Representation Model for Natural Language Processing of Real-World Outpatient Oncology Notes. [paper]
- SYNLAT: Syntax-Aligned Text-Latent Compression for Chain-of-Thought Reasoning. [paper]
- Training Numerical Intelligence via Auto-Diagnosis and Skill Discovery. [paper]
- Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate's Log-Probabilities. [paper]
- When Evidence Changes: Evaluating Memory Repair and Re-reading in Language-Model Agents. [paper]
- General Decision Models: Benchmarking and Insights Beyond Jev. [paper]
- A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass. [paper]
- BAIBAICHUCHU at the NTCIR-19 FinArg-3 Task: When Is Maximum Possible Profit Predictable from Investor Text?. [paper]
- Behavioral History Outperforms Descriptions of the Person for LLM Synthetic Personas. [paper]
- Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale. [paper]
- Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models. [paper]
- Periscope: Extending Frozen Language Models Beyond Their Context Window. [paper]
- IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation. [paper]
- Representation-Aligned Auxiliary Supervision for Language Model Adaptation. [paper]
- SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting. [paper]
- LongSocialBench: Do Long-Context LLMs Understand Online Discussion Threads?. [paper]
- Copying Before Suppression: What Drives a Below-Chance Dip During Language Model Training?. [paper]
- Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking. [paper]
- InvestigationWorlds: An Agentic Environment for Legal Investigation. [paper]
The papers
- Improving Reasoning Ability via Asynchronous On-Policy Self-Distillation under Positive Rollouts — Reinforcement learning with verifiable rewards (RLVR) is being advanced by Positive-Only Policy Optimization (POPO), a novel framework that enables policy improvement through online positive rollouts without relying on negative rollouts. [episode]
- StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback — StyleTailor presents a novel collaborative agent framework that unifies personalized apparel design, shopping recommendation, virtual try-on, and systematic evaluation into a cohesive workflow. [episode]
- Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter — Large language model (LLM) unlearning aims to remove specific data influences from pre-trained models without costly retraining, but recent studies reveal that unlearned models rapidly recover “forgotten” knowledge through relearning attacks, raising serious security concerns [episode]
- Narrative-UFET: Narrative Generation for Ultra-Fine Entity Typing — Ultra-fine entity typing (UFET) assigns highly specific types to entity mentions, but current approaches struggle with types in the long tail. [episode]
- MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale — As a diligent researcher, I have meticulously analyzed both provided texts regarding "MEMARENA." The goal is to synthesize these two descriptions into a single, comprehensive, and highly detailed summary that captures all critical aspects of the research for maximum clarity and a [episode]
- EvoState: Closed-Loop Visual State Management for Long-Form Video Generation — This work introduces an agentic framework designed to overcome identity drift and compounding inconsistencies in long-form video generation by establishing a closed-loop visual-text-memory synergy process. [episode]
- ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems — Recent advances in task-oriented dialogue (TOD) systems, driven by large language models (LLMs) with extensive API and tool integration, have enabled conversational agents to coordinate interleaved goals, maintain long-horizon context, and act proactively through asynchronous exe [episode]
- Missing-Data-Induced Phase Transitions in Spectral Partial Least Squares — Partial Least Squares (PLS) spectral analysis for multimodal learning is characterized by sharp phase transitions when both data views are subject to independent entry-wise missing-completely-at-random masking. [episode]
- Compression Is Not Evaluation-Neutral: Fixed RAG Compression Can Distort Reader Comparisons — Fixed compression can raise average accuracy while simultaneously corrupting cross-reader comparisons, hiding most of a real reader upgrade and reversing model rankings. [episode]
- Safe Inference-Time Alignment via Lagrangian Reward Augmentation — Inference-time alignment steers frozen language models during decoding using auxiliary reward signals, and this work introduces Lagrangian Reward Augmentation (LARA), a general framework that transfers the constrained trade-off from Safe RLHF to decoding by dualizing the safety c [episode]
- Subliminal Clocks: Latent Time Modelling in Diffusion Language Models — Diffusion Language Models (DLMs) are being investigated to determine if they internally represent denoising progress, and this work shows that DLMs do encode a latent representation related to the diffusion timestep within their residual streams. [episode]
- Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning — Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust deployment in dynamic environments: catastrophic forgetting, wherein sequential [episode]
- Medical foundation models converge less under label supervision — As a fastidious and diligent AI researcher, I have meticulously analyzed the provided excerpts from arXiv to construct a comprehensive summary of the research findings regarding representational convergence in medical imaging encoders. [episode]
- Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR — Reinforcement learning with verifiable rewards (RLVR) has emerged as a scalable paradigm for improving large language model reasoning, but its effectiveness is fundamentally limited by exploration, which this work addresses by proposing NUDGERL. [episode]
- Convergence of Statistical Estimators via Mutual Information Bounds — Mutual information (MI) bounds provide a novel information-theoretic framework for analyzing generalization and convergence rates across various statistical models. [episode]
- The JEPA Predictor: A Transferable Operator for Occluded Feature Completion — Joint-Embedding Predictive Architectures (JEPAs) introduce a trainable predictor head that maps visible-context features and target-position tokens to predicted target features in an encoder’s representation space, and this work demonstrates that this predictor functions as a " [episode]
- Neural Voxel Dynamics: Learning Volumetric Feature Advection for 3D Physics in V-JEPA Latent Space — Neural Voxel Dynamics presents a self-supervised framework for learning implicit 3D physical dynamics directly from video-derived supervisory signals by shifting the predictive bottleneck from 2D image space to a ‘lifted’ 3D Volumetric Latent Space. [episode]
- Deep Learning Reforms Image Matching: A Survey and Outlook — I am an excellent, fastidious, and diligent researcher. My task is to synthesize information from the provided text segments concerning "Deep Learning Reforms Image Matching: A Survey and Outlook" into a long, detailed summary. [episode]
- Point-MF: Stabilizing One-Step Mean Flows for Single-Image Point Cloud Reconstruction — Single-image point cloud reconstruction requires inferring complete 3D geometry, including occluded parts, from only visible content in a single RGB image, and this paper introduces Point-MF, a Mean-Flow-based framework that enables lowNFE single-image point cloud reconstruction [episode]
- LEAF: Growing Trees Without Branching for Speech-Aware Large Language Model Post-Training — State-of-the-art methods for speech-aware large language model posttraining suffer from coarse credit assignment, broadcasting identical terminal rewards to every token in a response, which LEAF addresses by recovering structural information within rollout batches through a retro [episode]
- Validation-Gated Causal Interventions for Interpreting High-Stakes Large Language Model Behavior: A Case Study in Suicidality Detection — Large language models are increasingly proposed for mental-health applications such as detecting suicidal content, raising the question of what they rely on. [episode]
- K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs — Large language models (LLMs) are increasingly deployed in K–12 education, yet existing benchmarks only measure factual recall, leaving a critical gap in understanding curriculum cognition—the structured knowledge of how concepts are organized and visually presented. [episode]
- FormuEvo: LLM-Guided Evolution for Discovering Solver-Efficient Mixed-Integer Programming Formulations — Mixed-integer programming (MIP) formulation design remains an expertise-intensive challenge, as mathematically equivalent formulations can differ by orders of magnitude in computational efficiency for downstream solvers. [episode]
- Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale — Language models can act as corpus-scale retrievers, but their performance degrades significantly when scaling to million-token corpora and generalizing beyond training sizes. [episode]
- RAC: Rectified Flow Auto Coder — Inconsistent generation and reconstruction results in traditional Variational Autoencoders (VAEs) are addressed by proposing a Rectified Flow Auto Coder (RAC), which replaces standard VAE decoding with a continuous-time velocity field integration to create a multi-step, correctab [episode]
- MG-VQA: Manipulation Grounded Visual Question Answering with VLMs — Vision-language models (VLMs) are being pushed beyond static image analysis to handle real-world tasks that require physical interaction, such as answering questions about objects hidden in cluttered environments. [episode]
- GeoFunFlow: Geometric function flow matching for joint probabilistic inference of physical fields and complex geometries — Inverse problems governed by partial differential equations (PDEs) are crucial in science and engineering, particularly when dealing with irregular geometries, which poses significant challenges for classical optimization methods. [episode]
- Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset — Fitting a two-parameter logistic (2PL) IRT model to responses from 29 language models on HLE’s text-only multiple-choice subset reveals that the benchmark measures a single general reasoning factor, and measurement precision concentrates at moderate ability levels. [episode]
- Adaptive Information Control for Search-Augmented LLM Reasoning — As a diligent AI researcher, I have thoroughly reviewed both provided texts. [episode]
- 3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models — Current Large Language Models often fail on elementary spatial tasks like block counting due to a critical “spatial intelligence gap,” where they lack a coherent 3D mental representation from 2D observations. [episode]
- Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes — Trainable input embedding tables are standard in language models, but this work investigates whether they are necessary by replacing them with fixed minimal binary token codes and finding that exact token identity is sufficient for useful language modeling. [episode]
- DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures — DynaMiCS is a novel dynamic mixture optimizer designed to handle multi-domain fine-tuning by framing it as a constrained optimization problem. [episode]
- Quantile Adaptive Temperature Scaling for Confidence Calibration — Deep neural networks often produce poorly calibrated confidence estimates, overstating their certainty even when predictions are incorrect, which makes calibration vital for reliable risk estimation in high-stakes domains. [episode]
- Gradient descent dynamics for deep equilibrium models — Deep equilibrium models (DEQs) are a powerful paradigm for training infinitely deep weight-tied neural networks, and this work rigorously studies their gradient descent dynamics in linear and single-index settings to understand their theoretical behavior. [episode]
- MoFlow: One-Step Flow Matching for Human Trajectory Forecasting via Implicit Maximum Likelihood Estimation based Distillation — MoFlow introduces a novel motion prediction conditional flow matching model that predicts multiple future trajectories for all agents in a scene, utilizing an Implicit Maximum Likelihood Estimation (IMLE) based distillation method to achieve state-of-the-art performance while sig [episode]
- LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers — LLMs consistently generate story continuations with 2–4× lower entropy and substantially higher context-dependence than human-authored ground truth—a gap that widens under post-training alignment and persists across model families and scales. [episode]
- Towards causal effect estimation with learned instrument representations — Instrumental variable (IV) methods are crucial for estimating causal effects from observational data, but their reliance on explicitly available instruments often limits their practical application. [episode]
- Office Comprehension Benchmark — The provided text details a significant new resource in Large Language Model (LLM) evaluation, specifically introducing the Office Comprehension Benchmark (OCB). [episode]
- EulerESG: Automating ESG Disclosure Analysis with LLMs — ESG reports are often published as long, heterogeneous PDF documents, making systematic analysis difficult and labor-intensive. [episode]
- MultiLoc: Look Around As You Localize for Fast And Robust Visual Re-localization — MultiLoc introduces a novel multi-view guided Relative Pose Regression (RPR) framework designed for fast and robust visual re-localization across diverse environments, addressing limitations in existing methods by integrating global spatial and geometric understanding. [episode]
- Brain-IT-VQA: From Brain Signals to Answers — Decoding visual content from fMRI signals while answering questions about those images is a challenging problem, and this work introduces Brain-IT-VQA, a framework that decodes language tokens from brain activity to answer visual questions and provides a new benchmark for studyin [episode]
- Co-LMLM: Continuous-Query Limited Memory Language Models — Limited memory language models (LMLMs) externalize factual knowledge during pre-training to a knowledge base (KB), rather than memorizing it in their weights, and this work introduces Continuous-Query LMLM (CO-LMLM), which pairs continuous keys with textual knowledge values to al [episode]
- Three tiers of computation in transformers and in brain architectures — As a fastidious and diligent researcher, I have meticulously analyzed both provided texts from the arXiv paper concerning computational characterization of human language and its relation to transformer-based language models (LMs). [episode]
- MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models — Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, but realizing this potential requires robust and fine-grained visual perception across diverse modalities, scales, and contexts. [episode]
- pi squared: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models — The study introduces a pipeline called π2 that curates high-quality, structure-originated reasoning data from Wikipedia tables to significantly improve the long-context reasoning ability of large language models. [episode]
- Robotic Ultrasound Makes CBCT Alive — Intraoperative Cone Beam Computed Tomography (CBCT) provides essential 3D anatomical context, but its static nature fails to monitor soft-tissue deformations caused by respiration or surgical manipulation, leading to navigation discrepancies. [episode]
- RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Multi-Trajectory Semantic Retrieval — Referring Atomic Video Action Recognition (RAVAR) aims to recognize fine-grained, atomic-level actions of a specific person conditioned on natural language descriptions, which is critical for interactive human action analysis in complex multi-person scenarios. [episode]
- Hidden States as Value Gradients: The Pontryagin Structure of Recurrent Policies — A key capability of intelligent agents is to act effectively under incomplete state observations, and this work shows that recurrent policies implicitly realize a control-theoretic structure where their hidden state tracks the gradient of the value function, playing the role of a [episode]
- Retrieval-Augmented Generation Must Move Beyond Factual Grounding to Represent Diverse Opinions — Retrieval-Augmented Generation systems exhibit a factual bias by optimizing for epistemic uncertainty reduction while ignoring aleatoric uncertainty inherent in opinion-rich content, necessitating a paradigm shift toward opinion-aware design to prevent echo chambers and manipulat [episode]
- Neural Bayesian Filtering — Neural Bayesian Filtering (NBF) is an algorithm designed for maintaining distributions over hidden states, or beliefs, in partially observable systems by mapping these complex belief states to fixed-length embedding vectors that condition generative models for sampling. [episode]
- Contrastive Time Series Forecasting with Anomalies — Time-series forecasting often struggles to distinguish between short-lived noise and persistent, forecast-relevant anomalies that can significantly alter future predictions. [episode]
- MDKeyChunker: What Does One LLM Call per Chunk Buy for Markdown Retrieval? — MDKeyChunker introduces a three-stage pipeline designed to enhance Retrieval-Augmented Generation (RAG) accuracy for Markdown documents by replacing multi-tool extraction passes with a single, structure-aware LLM call per chunk, augmented by a rolling key dictionary and subsequen [episode]
- Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling — LLMs have so far failed both to generate consistently compelling stories and to recognize this failure, which necessitates introducing narrative tension as a key metric for evaluating human-written fiction. [episode]
- Live Interactive Training for Video Segmentation — Interactive video segmentation often requires many user interventions for robust performance in challenging scenarios, yet state-of-the-art models like SAM2 only use corrections for immediate fixes without learning from this feedback, leading to inefficient, repetitive user effor [episode]
- From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives — A unified framework for long-form narrative generation and verification was introduced to address the challenges of narrative consistency and plot discontinuity in large language models. [episode]
- Cross-Lingual Summarization as a Black-Box Watermark Removal Attack — Cross-lingual summarization attacks (CLSA) represent a qualitatively stronger threat to AI watermarking than prior methods because they systematically destroy token-level statistical biases while preserving semantic fidelity. [episode]
- Towards Probabilistic Question Answering Over Tabular Data — Current approaches for question answering (QA) over tabular data, such as NL2SQL systems, perform well for factual questions where answers are directly retrieved from tables. However, they fall short on probabilistic questions requiring reasoning under uncertainty. [episode]
- A PAC-Bayesian View of Generalisation for Physics-Informed Machine Learning — Physics-informed machine learning (PIML) integrates mechanistic knowledge, typically in the form of partial differential equations (PDEs), into data-driven models to improve performance. [episode]
- Automating MD simulations for Proteins using Large language Models: NAMD-Agent — Molecular dynamics simulations are an essential tool for understanding protein structure, dynamics, and function at the atomic level; this work introduces an automated pipeline that leverages Large Language Models (LLMs) to streamline the generation of Molecular Dynamics (MD) inp [episode]
- Implicit Target Shift in Online Learning: Characterization and Correction — Online learning from data streams often struggles under distributional shift, and this work provides a fundamental framework to analyze and improve online learning by characterizing an effective target shift in kernel regression. [episode]
- Studying the Soupability of Documents in State Space Models — Hidden states from Structured State Space Models (SSMs) can be merged post hoc to support downstream reasoning, and this study investigates whether independently encoded document representations can be pooled while preserving information needed for multidocument reasoning. [episode]
- Transfer Learning for Spatial Autoregressive Models with Application to U.S. Presidential Election Prediction — It is important to incorporate spatial geographic information into U.S. [episode]
- InvestigationWorlds: An Agentic Environment for Legal Investigation —
- PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO —
- From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models —
- ExpertMuon-Compass: Alignment-Guided Step Sizes for Mixture-of-Experts Training —
- Consideration Circuits: Depth Separation and Universality Beyond a Single Softmax —
- Watermarks and Fingerprints as Soft Bindings for Content Provenance: An Open-Licence Benchmark for Images, Audio and Video —
- Kepler4D: Controllable Future Video Generation via 4D Scene State Evolution —
- Trajectory-Derived Confidence for Reliable, Resource-Aware Clinical Text-to-SQL Agents —
- Mitigating Over-squashing without Rewiring: A Sheaf Effective Resistance Perspective —
- How RL Reshapes LLM Reasoning: Transferability, Coverage, and Scaling Laws —
- Diagnosis-Conditioned Spatial Gating and Decoder-Level Supervised Contrastive Learning for Radiology Report Generation —
- Principled Top- k Selection for Language Models with Hybrid Gradients —
- Risk-Calibrated Proposal Transport for Finite-Particle Diffusion Steering —
- Language Model Activations Inhabit Privileged Error-Correcting Basins —
- Referring Multi-Object Tracking in Moving-Camera Videos via Global Motion Compensation —
- CellSplat4D: PSF-Aware 4D Gaussian Splatting for Sparse Robotic Live-Cell Imaging —
- Sparse-GS2Mesh: 3D Gaussian Splatting Guided by Novel Stereo Views and 2DGS for Sparse View Surface Reconstruction —
- Clean: Second-order LLM Training at Linear Memory Cost via Nystr"om Sketching —
- Can LLMs Separate Pasted Artifacts from User Speech? Absorption at Unmarked Prompt Seams —
- FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding —
- Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification —
- Playing social deduction games with reinforcement fine-tuned large language models —
- AI-Enabled Quality Assurance for Multiple-Choice Assessment Items —
- Self-Reflection Fine-Tuning: Enhancing Agent Security against Prompt Injection Attacks from Failure Experience —
- Rethinking Self-Distillation for Multi-Teacher Capability Merging —
- OctMesh: A Unified Octree-Hierarchical Framework for Lossless Triangle Mesh Compression —
- First-Order Steering: Translating Weight Adaptation into Activation Steering —
- Local Fisher Information Enables Sparse Causal Discovery —
- Language-Conditioned Token and Reasoning Efficiency in Large Language Models: A Paired Cross-Lingual Study Protocol —
- Questioning the Questions: Sustaining Self-Evolution in Reasoning Models —
- A Geometric-Transformation Feature-Adaptive Manifold Restoration Method for Open-Vocabulary Semantic Segmentation of Remote Sensing Images —
- Evaluating Modeling Approaches for Experience-Level Classification in Job Description —
- Modeling Deletion Requests in Machine Unlearning —
- From Latent Space to Jacobian Space: Measuring, Evading, and Training Against Safety-Content Accessibility —
- Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency —
- LyapuFlow: Controlling Generative Flows with Lyapunov Feedback for Inverse Problems —
- Trust the View That Sees the Target: Mining Cross-View Conflicts for Reliability-Gated Disaster Damage Assessment —
- Bidirectional Preference Synthesis: Learning Prompt-Conditioned Preferences from Boundary Failures —
- Suppressing Pressure, Amplifying Evidence: Self-Guided Attention Steering to Mitigate Sycophancy and Stubbornness —
- Synthetic-to-Real ViT-Based Pose Estimation of a Noncooperative UAV —
- A differentiable Lagrangian-coupled 3D Gaussian Splatting-SPH model for forward simulation and inverse analysis in solid mechanics —
- ShadowMiner v1 - An Experience Report on Implementing and Measuring a Problem-and-Hypothesis Discovery Engine —
- Hierarchical Credit Assignment for RLVR on Fused Gromov-Wasserstein Geometry —
- Any-scale Object Detection using Arbitrary-scaled Images —
- Evidence and Intervention: A Coupled Active-Inference Extension of Rational Speech Act Models —
- LoCoSplat: Real-Time Feed-Forward 3D Gaussian Splatting with Minimal 3D Reasoning —
- SelectOccFlow: Selective Spatiotemporal Aggregation for 3D Occupancy and Scene Flow Prediction —
- TRIM-ReID: Duplication-Aware Token Reduction and Modality-Aligned Interaction for Multi-Modal Object Re-Identification —
- A Bird's-Eye View of Iterative Reward Design —
- A multi-stage probabilistic framework to estimate gas-fired generator performance during extreme winter weather —
- Boundaries Agree, Labels Do Not: Intra-Annotator Dynamics as a Kind of Training Data —
- Beyond Plausibility: Verifiable Fine-Grained Image Editing on Structured Assets —
- Gaussian Flow Dynamics: Simulation-Free Neural SDE Learning Beyond One-Time Marginals —
- Detecting Defects that Matter: An Application-Driven Benchmark for Anomaly Detection in Manufacturing and Retail Logistics (VAND 4.0 Challenge) —
- GlitchPatch: Repairing Glitch Tokens in Frozen Language Models via Local Retokenization —
- XTurnix: Large-Scale Self-Supervised Turn Control through Two-State Binary Decisions —
- Ideological Stance Detection in a Low-Resource Language: Polarization in Bangladeshi Public vs Private University Discourse on Social Media —
- Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define —
- Understanding and Mitigating Hallucination Escape in Tool-Using LLM Agents —
- On the Trade-off Between Information Loss and Generalization in Sparse Attention —
- AgroGround: Multi-Granularity Grounded Recognition in Agriculture —
- UnAct: Gradient-Free Unlearning via Targeted Activation Intervention —
- Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos —
- RAGrasp: Geometry-Semantic Template Retrieval and Grasp Transfer —
- On the Tightness and Computational Tractability of Higher-Dimensional Confidence Sequences —
- Fractal Cross Product: Theory, Differentiable Implementation and Application to Medical Image Analysis —
- Do Motion Tokenizers for Co-Speech Gesture Generation Encode Gesture Semantics? —
- LoRA Direction Extraction for Controllable Light Toggling in FLUX.1 Kontext —
- Energy Variation in Training Modern Computer Vision Architectures —
- What Do Verifiable Rewards Teach Video-Language Models About Time? A Controlled Multi-Model Study —
- WAMJET: A Harness for World Action Model Acceleration —
- StepCAD: Mesh-to-CAD Code Generation via LLM Policy and Geometry-Guided Search —
- Fine-Grained Emotion Classification from Mobile App Reviews: An Empirical Study with Large Language Models —
- Least Squares for Time Series Forecasting —
- Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction —
- Targeted Active Learning for Preference-Based Treatment Effects on Multivariate Outcomes —
- Same Output, Different Gold: Measuring How Reference Choice Moves a Multilingual Benchmark Score —
- Are We Measuring Anticipation? Auditing Privileged Information in Procedural Video Evaluation —
- The Score Is Not the Structure: Brain Alignment and Cross-Lingual Transfer —
- OncoNoteBERT: A Foundation Representation Model for Natural Language Processing of Real-World Outpatient Oncology Notes —
- SYNLAT: Syntax-Aligned Text-Latent Compression for Chain-of-Thought Reasoning —
- Probability flow ODEs in score-based and reflected diffusion models —
- Training Numerical Intelligence via Auto-Diagnosis and Skill Discovery —
- Streaming Multi-Track Timeline Control for 3D Human Motion Generation —
- Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate's Log-Probabilities —
- When Evidence Changes: Evaluating Memory Repair and Re-reading in Language-Model Agents —
- DABACO: A Multi-Camera Dataset and Benchmark for Screen Localization and Pointing Estimation —
- General Decision Models: Benchmarking and Insights Beyond Jev —
- A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass —
- Dynamic Time Step Prediction in Inverse Heat Dissipation for Blur-Like Image Restoration Tasks —
- Evolving LLM-Generated Features for Interpretable Classification —
- Latent Score-Based Bayesian Cram'er-Rao Bound Estimation for High-Dimensional Imaging Systems —
- BAIBAICHUCHU at the NTCIR-19 FinArg-3 Task: When Is Maximum Possible Profit Predictable from Investor Text? —
- Learning Latent Protein Languages for Autoregressive Generation —
- FADE: Frame-Aware Diffusion-Transformer-based Multi-Concept Erasure for Video Unlearning —
- Masked Privileged-Information Distillation for Multimodal Skin Lesion Classification Under Missing Clinical Metadata —
- Behavioral History Outperforms Descriptions of the Person for LLM Synthetic Personas —
- Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale —
- Selective Backpropagation for Efficient Few-Shot Class-Incremental Learning —
- Verifier-Guided Synthetic Augmentation for 3D Human Shape Generation —
- VolS-GS: Relightable Gaussian Splatting with Volumetric Subsurface Scattering —
- PhysMamba: Selective State Space Models as Learned Articulated Body Simulators —
- Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models —
- ClasSAE: Class-Aligned Sparse Autoencoders via Differentiable Feature-Class Affinity —
- Learning Subject-Specific Anatomical Representations via Manifold Expansion: Application to Accelerated Multi-Contrast MRI —
- MAGEFormer: Learning Metric-Consistent Representations for Anisotropic CT Segmentation —
- Adaptive Partitioning Schemes for Optimistic Optimization —
- Dynamic Quadtree Tokenization and Transformer for Adaptive Mesh PDE Forecasting —
- Periscope: Extending Frozen Language Models Beyond Their Context Window —
- ASD-FEAT: A Multi-Modal Infant Video-Derived Dataset for Early ASD Risk Prediction —
- A Theory of Shape Reconstruction from Heat Conduction and Shading —
- Evaluating Zone-Guided Front Extraction for Glacier Calving-Front Delineation in SAR Imagery —
- Why Convolution Still Matters: Evaluating Inductive Biases in Cryospheric Image Classification —
- IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation —
- VCURF: Virtual Camera-based Uncertainty of Radiance Fields —
- Dependable AI-Assisted Engineering: A Formal Framework for AI Participation and Assurance in Safety-Critical Workflows —
- Towards Safer Autonomous Driving in an Open World: A Dual-Process Approach —
- UniBRep: Learning Unified Geometry and Topology for Image-conditioned B-Rep Generation —
- Scaling 3D Visual Grounding in Abdominal CT —
- Representation-Aligned Auxiliary Supervision for Language Model Adaptation —
- Where Does the Semantic Gain Come From? A Reproduction and Extension of Semantic Knowledge-driven Contrastive Learning for Long-Tailed Recognition —
- SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting —
- LongSocialBench: Do Long-Context LLMs Understand Online Discussion Threads? —
- Copying Before Suppression: What Drives a Below-Chance Dip During Language Model Training? —
- Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking —
Important terms
- NAMD-Agent
- This system automates molecular dynamics simulations for proteins by learning their underlying rules through autoregressive generation, using a hierarchical transformer and brain model structure to manage complexity.
- EulerESG
- This development uses large language models to automate the analysis of Environmental, Social, and Governance disclosures. It leverages LLMs to efficiently extract key information from vast amounts of text for reporting.
- ATOD
- This work establishes a benchmark and evaluation framework specifically for agentic task-oriented dialogue systems. It provides a standardized way to fairly measure how well these systems complete specific goals through interaction.
- LongSocialBench
- This study tested how long context windows affect a model's ability to understand complex online discussion threads. The findings suggest that longer memory allows models to grasp nuanced social dynamics better.