AI papers — 2026-10-02
Today's focus is on developing auditable algebraic counting fields to detect hidden pockets within apo structures, which is crucial for drug design because understanding these regions directly impacts drug design. We explored how to build a system that can reliably identify these pockets using algebraic methods. A key part of this involved analyzing the results from the work on PUMA, which focused on learning a mutation-aware vocabulary of protein units to better map out these structural features. This connects to our efforts in Geometric Stability, where we looked at finding a missing axis in representations that might help stabilize these pocket predictions.
We also touched upon effective resistance and graph neural network reliability when considering tissue-specific interactomes. This provides a framework for how complex biological networks behave under stress. This contrasts with the stochastic optimal control approach used for continuous-time fMRI representation learning, which deals with modeling dynamic brain activity over time. The ideas from Poincar'e meets Bellman, concerning revisable memory and evidence-supported learning in changing environments, offer a way to handle the evolving nature of these biological systems.
Finally, we examined the analysis of quantized and efficiently adapted protein language models to see how well they perform in practice for this type of structural prediction. The most significant work today involved ORBIT-FMIB, which attempts to track order resolved epistatic information using ESM-2. This is important because understanding how gene interactions are ordered is key to deciphering complex biological processes. The researchers tried to use the ESM-2 model to capture this ordering and found that it provides a useful representation of these relationships.
Another piece of work focused on pCoMole, which uses discrete flows for Pareto constrained molecule editing. This method aims to modify molecules in a way that respects certain constraints while keeping them chemically valid, which is crucial for drug discovery. This approach builds upon prior work by focusing on making precise edits within specific chemical spaces.
Then there was the effort to increase the width of layers in self-supervised learning to rival end-to-end backpropagation. This suggests that a deeper, wider network structure can achieve similar performance gains without relying solely on the full end-to-end training paradigm. This idea connects to how other models might learn representations more effectively.
We also saw work on inferring multi timescale neural dynamics using switching linear dynamical systems. This is relevant because biological systems operate across different time scales, and this framework tries to model that complexity dynamically. This contrasts with the static representation learning goals of some graph structure learning methods.
Finally, there was a study on graph structure learning with temporal graph information bottleneck for inductive representation learning. This work seeks to learn useful representations from graphs by incorporating temporal information, which is a step toward understanding dynamic biological networks.
The most significant development centers on UniGuardian, which is a unified defense system designed to detect prompt injection, backdoor attacks, and adversarial attacks in large language models. This matters because securing these systems against malicious input is crucial for maintaining the integrity of deployed AI.
AstroAgentBench provided a new way to evaluate agentic planning capabilities specifically within space mission planning tasks. This work is important because it moves beyond simple task completion to assess complex, multi-step reasoning in an agentic framework.
VISPA introduced pluralistic alignment through automatic value selection and activation, which is significant for understanding how models can be steered toward diverse, non-uniform outcomes. This builds upon the foundational work of Textual Planning with Explicit Latent Transitions, which focuses on using explicit latent transitions to guide textual planning processes.
In Vino Veritas and Vulnerabilities explored LLM safety by examining vulnerabilities induced through drunk language. This offered insight into emergent safety issues when models are exposed to unconventional inputs. This contrasts with LEAD, which focuses on layer-wise expert-aligned decoding for generating faithful radiology reports, a specialized application of model fidelity.
The most significant development this morning concerns the framework for auditing grounding claims, which is crucial because it addresses the reliability of large language models when they make factual assertions. This work proposes a specific framework designed to check whether the information generated by an LLM is actually supported by its source material.
This approach builds upon earlier efforts to understand how models handle uncertainty in their outputs, specifically through distributional uncertainty scoring. This rethinks how we evaluate language model performance by accounting for the inherent ambiguity in the data. Relatedly, there is ongoing research into boosting mathematical problem-solving capabilities within large language models using MDToC, a metacognitive dynamic tree of concepts that helps guide the model's reasoning process.
Further refinement comes from iterative topic taxonomy induction using LLMs. This provides a case study on how these models can be used to build structured knowledge hierarchies in electoral advertising contexts. This contrasts with the work focusing on measuring iterative temporal reasoning through time puzzles, which tests a model's ability to handle sequences of events over time.
Finally, there is the interview-grounded personality simulation framework, InterviewSim. This offers a scalable way to simulate personality traits based on interview data. This connects back to the foundational work of MERGE, which tests minimal expression-replacement generalization for natural language inference.
The most pressing work involves AuditBench, which is testing different alignment auditing techniques on models that have hidden behaviors. This matters because understanding these hidden behaviors is crucial for ensuring that large language models behave reliably in real-world applications.
One key part of this involved comparing how different methods of checking model alignment perform when dealing with these complex latent behaviors. Another study looked at separating production and review sessions to improve the quality of large language model output by using cross-context review. This suggests that having a dedicated check phase helps refine the final result.
Then there was research on detecting and attributing LLM ghostwriters, which is important for understanding provenance in generated text. This builds on work examining verbal tics in frontier language models, which critically reviewed current releases and public discussion around these linguistic quirks.
Finally, there is the work concerning multi-perspective LLM annotations for subjective tasks. This aims to validate analyses by having multiple viewpoints look at the same output. This connects back to evaluating LLMs under challenging patient behaviors in medical consultations, where understanding nuanced human responses is key.
The most critical development today involves VIDA, a new dataset designed to capture the kind of visual ambiguity that plagues multimodal machine translation. This work is important because it directly addresses the reliability issues when translating text accompanied by images, which is a major hurdle for real-world applications. The researchers created this dataset by focusing on scenarios where visual context could drastically alter the meaning of a sentence and found it provides rich examples for training models to handle these visual dependencies better.
This dataset feeds into efforts to build domain-adapted small language models that can perform reliable clinical triage. Specifically, the methods explored how to fine-tune these smaller models using this type of complex visual input data, aiming for more trustworthy outputs in medical contexts. Moving down the list of significance, there was work on reinforcement learning applied directly to language models where they use their own internal states to estimate value. This means the model learns to critique its own decisions during training and is a step toward making language models more self-aware of their performance.
Another piece of research focused on adaptive steering and remasking techniques for diffusion language models, which are used for safe content generation. This technique is essentially a way to guide the generation process so that the output stays within acceptable safety boundaries while still being creative. Following that, there was an investigation into probing persona-dependent preferences within large language models to see how much a model's desired personality influences its generated text. This helps researchers understand the subtle ways in which a model's adopted style affects what it produces.
Finally, we saw work on the fragility of chain-of-thought monitoring when dealing with typologically diverse languages. This showed that simply tracking the reasoning steps breaks down across different language structures. This contrasts with MemGuard, which tackles memory contamination in long-term memory systems by implementing specific safeguards to prevent old or irrelevant information from polluting new learning processes.
The work on SPADER is particularly important because it directly addresses how large language models can improve their performance when answering multi-answer questions by using step-wise peer advantage and diversity-aware exploration rewards. This method tries to make the model explore different answer paths more effectively than standard prompting.
This approach builds on previous ideas where models struggle with reasoning failures, such as those seen in the confidence shortcut observed in masked diffusion models. This suggests a specific failure mode when they are asked to reconstruct missing parts of text. The SPADER framework attempts to guide the model away from these pitfalls by rewarding it for exploring diverse answer sets.
Another piece of research tackles the problem of LLM hallucinations through DECK, which creates a taxonomy based on consistency and confidence levels to better understand where and why models go wrong. This classification helps pinpoint the nature of the error, which is crucial for developing targeted fixes.
This understanding of model failure connects to work exploring how emergent misalignment happens through what is called the piggyback hypothesis, suggesting that generalization might be an issue when models are not specifically aligned with certain partners or tasks. Furthermore, research into multimodal agents shows that they can succeed in reference games without ever forming deep conceptual pacts between the different modalities involved. This suggests success doesn't always require a deep internal understanding of the relationship between inputs.
Finally, TRIAGE is focused on using dialectical reasoning to predict risks in medical time series that are irregularly sampled, offering explainability for those predictions. This contrasts with the generative work by focusing on structured reasoning and risk assessment rather than pure text generation or alignment issues.
The most pressing development involves the work on CHILLGuard, which builds fine-grained safety guardrails for Chinese large language models by using scalable data construction and model-aware preference alignment. This is important because it directly addresses the growing need to make powerful generative models safer in deployment.
This work suggests that vision-language models used for chest radiography do not always require the image input to perform their intended tasks. This finding opens up possibilities for more efficient diagnostic tools, which is a significant shift in how we approach medical imaging analysis.
Another piece of research focuses on ReNikud, which achieves audio-supervised Hebrew grapheme-to-phoneme conversion. This method is valuable because it provides a robust way to handle the complexities of converting written text into spoken sound for language processing applications.
Self-conditioned flow map language models via fixed-point flows explore how these models can be conditioned using fixed-point flows. This is an attempt to improve their efficiency and control during generation. This relates to the broader effort in leveraging instruction tuning and merging for reasoning model adaptation, showing a path toward making these adapted models more capable.
Finally, reading between the dots investigates decoding hidden computation across filler tokens in language models. This work is significant because it seeks to uncover how these models actually process information when they are not explicitly generating output, offering deeper insights into their internal workings.
Today's papers
- Auditable Algebraic Counting Field for Cryptic-Pocket Detection from Apo Structures. [paper]
- Effective Resistance and Graph Neural Network Reliability in Tissue-Specific Interactomes. [paper]
- Stochastic Optimal Control for Continuous-Time fMRI Representation Learning. [paper] [episode]
- Poincar'e Meets Bellman: Revisable Memory, Operational Quotients, and Evidence-Supported Learning in Changing Environments. [paper] [episode]
- REALM: Retrospective Encoder Alignment for LFP Modeling. [paper] [episode]
- PUMA: Learning a Mutation-Aware Vocabulary of Protein Units. [paper] [episode]
- Geometric Stability: The Missing Axis of Representations. [paper] [episode]
- Analysis of Quantized and Efficiently Adapted Protein Language Models. [paper]
- ORBIT-FMIB: Tracking Order-Resolved Epistatic Information Through ESM-2. [paper]
- Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?. [paper]
- pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows. [paper]
- Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning. [paper]
- Inferring Multi-Timescale Neural Dynamics with Switching Linear Dynamical Systems. [paper]
- A foundation for systematic analysis of transformers and RNNs for tractography. [paper]
- Graph Structure Learning with Temporal Graph Information Bottleneck for Inductive Representation Learning. [paper] [episode]
- StoCFL: A Stochastically Clustered Federated Learning Framework for Non-IID Data with Dynamic Client Participation. [paper] [episode]
- Linguistic traces of stochastic empathy in language models. [paper] [episode]
- UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models. [paper] [episode]
- AstroAgentBench: Evaluating Agentic Planning on Space Mission Planning Tasks. [paper] [episode]
- VISPA: Pluralistic Alignment via Automatic Value Selection and Activation. [paper] [episode]
- In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement. [paper] [episode]
- Textual Planning with Explicit Latent Transitions. [paper] [episode]
- LEAD: Layer-wise Expert-aligned Decoding for Faithful Radiology Report Generation. [paper] [episode]
- MMMG: a Comprehensive and Reliable Benchmark for Multitask Multimodal Generation. [paper] [episode]
- LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios. [paper] [episode]
- When Guessing is Rewarded: Rethinking Language Model Evaluation with Distributional Uncertainty Scoring. [paper] [episode]
- Iterative Topic Taxonomy Induction with LLMs: A Case Study of Electoral Advertising. [paper] [episode]
- MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference. [paper] [episode]
- A framework for auditing grounding claims. [paper] [episode]
- MDToC: Metacognitive Dynamic Tree of Concepts for Boosting Mathematical Problem-Solving of Large Language Models. [paper] [episode]
- Measuring Iterative Temporal Reasoning with Time Puzzles. [paper] [episode]
- InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation. [paper] [episode]
- AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors. [paper] [episode]
- From Literature to Hypotheses: An AI Co-Scientist System for Biomarker-Guided Drug Combination Hypothesis Generation. [paper] [episode]
- Cross-Context Review: Improving LLM Output Quality by Separating Production and Review Sessions. [paper] [episode]
- Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders. [paper] [episode]
- Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks. [paper] [episode]
- Who Wrote the Book? Detecting and Attributing LLM Ghostwriters. [paper] [episode]
- Beyond Idealized Patients: Evaluating LLMs under Challenging Patient Behaviors in Medical Consultations. [paper] [episode]
- Verbal tics in frontier language models: A critical review of current releases, research evidence, and public discussion. [paper] [episode]
- Domain-Adapted Small Language Models for Reliable Clinical Triage. [paper] [episode]
- VIDA: A Dataset for Visually Dependent Ambiguity in Multimodal Machine Translation. [paper] [episode]
- Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States. [paper] [episode]
- Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models. [paper] [episode]
- Probing Persona-Dependent Preferences in Language Models. [paper] [episode]
- Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs. [paper] [episode]
- The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages. [paper] [episode]
- MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models. [paper] [episode]
- Text-Preserving Lossy Text Compression: A Study of Strategic Deletion and LLM Reconstruction. [paper] [episode]
- The Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion Models. [paper] [episode]
- SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering. [paper] [episode]
- DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations. [paper] [episode]
- The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment. [paper] [episode]
- Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts. [paper] [episode]
- TRIAGE: Dialectical LLM Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series. [paper] [episode]
- Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings. [paper] [episode]
- What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language. [paper]
- Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yor`ub'a. [paper]
- CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment. [paper] [episode]
- Vision-language models for chest radiography do not always need the image. [paper] [episode]
The papers
- Fault-Tolerant Budget Conservation in Distributed Multi-Agent Delegation — This research introduces a novel, fault-tolerant authorization semantics designed for AI agents that delegate complex work across concurrent and failure-prone workers. [episode]
- DSSR-3D: Decoupled Reasoning for View-Dependent Referring in 3D Gaussians — Recent advances in 3D Gaussian Splatting have enabled open-vocabulary and referring segmentation by distilling semantic knowledge from 2D foundation models into 3D representations. [episode]
- ProtoDCS: Towards Robust and Efficient Open-Set Test-Time Adaptation for Vision-Language Models — Large-scale Vision-Language Models (VLMs) face significant challenges in real-world deployment due to distribution shifts, and existing Test-Time Adaptation (TTA) methods fail in open-set scenarios where test streams contain both covariate-shifted in-distribution (csID) and out-o [episode]
- Can AI Oversight Be Zero Knowledge? — AI systems increasingly produce outputs from confidential data, such as medical assessments or drug candidate properties, necessitating verification methods that ensure correctness without revealing underlying sensitive information. [episode]
- Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null — Across three vision–language model architectures, this work reports a universal negative finding for mid-layer interpretability, demonstrating that while mid-layers often encode ground-truth answers in errors, this signal is not causally active for the final prediction. [episode]
- Probing Persona-Dependent Preferences in Language Models — Large language models exhibit preferences and adopt different personas, and this research investigates how these persona-dependent preferences are implemented internally. [episode]
- MMMG: a Comprehensive and Reliable Benchmark for Multitask Multimodal Generation — As a fastidious and diligent researcher, I have meticulously analyzed both provided texts concerning the MMMG benchmark. [episode]
- UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models — A unified defense mechanism has been proposed to detect prompt injection, backdoor attacks, and adversarial attacks in Large Language Models by treating them collectively as Prompt Trigger Attacks (PTA). [episode]
- ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion — Grapheme-to-phoneme (G2P) conversion for Modern Hebrew is needed for applications like text-to-speech (TTS), but is challenging due to the language’s abjad writing system, which leaves vowels largely unwritten, creating substantial ambiguity. [episode]
- Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States — Reinforcement learning for large reasoning models can be made more stable and efficient by leveraging internal representations already computed during policy training to estimate value functions, a concept introduced in POISE. [episode]
- CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment — Malicious content generated by large language models (LLMs) poses severe safety risks, and existing guardrails often fail to adapt to Chinese cultural context and regulatory nuances. [episode]
- LensVLM: Selective Context Expansion for Compressed Visual Representation of Text — LensVLM introduces an inference framework and post-training recipe that enables Vision Language Models (VLMs) to scan compressed images and selectively expand only the relevant images to their uncompressed form via learned tools, pushing compression further than baselines while m [episode]
- A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model — In an era where large language models are becoming ubiquitous, it is crucial to consider their environmental impact due to their extreme size and resource use. [episode]
- VIDA: A Dataset for Visually Dependent Ambiguity in Multimodal Machine Translation — Ambiguity resolution is a key challenge in multimodal machine translation (MMT), where models must genuinely leverage visual input to map an ambiguous expression to its intended meaning. [episode]
- In CEM, a World Model Is Also a Proposal Mechanism — A world model used for planning determines both which actions receive further consideration and how those candidates are subsequently ranked, making it crucial to evaluate these two roles separately. [episode]
- Reading Between the Dots: Decoding Hidden Computation across Filler Tokens — Frontier Large Language Models can perform multi-step reasoning over content-free filler tokens, and this hidden computation is at least partially accessible to interpretability tools when it is invisible to chain-of-thought monitoring. [episode]
- Beyond Localization: A Comprehensive Benchmark of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images — Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. [episode]
- FACET at WMT 2026 Automated Translation Quality Evaluation Task — FACET introduces a reference-free method for evaluating machine translation quality by decomposing it into three distinct passes—Fluency, Accuracy, and Consistency—each tailored to specific error types. [episode]
- Subtoken Vision Transformer for Fine-grained Recognition — Subtoken Vision Transformer (SubViT) introduces a selective image tokenization method, Attention-based Token Subdivision (ATS), and a two-stage training strategy to enhance fine-grained visual recognition by allocating additional representational capacity only to discriminative p [episode]
- Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks — Large language models are increasingly used to annotate texts, but their outputs reflect some human perspectives better than others. [episode]
- Spectral Tail Auxiliary Learning for AI-Generated Image Detection — As generative image models evolve rapidly, making AI-generated image detection increasingly challenging, this paper introduces Spectral Tail Auxiliary Learning (STAL), a novel frequency-domain auxiliary supervision framework designed to improve the generalization and stability of [episode]
- EgoForge: Goal-Directed Egocentric World Simulator — Given a single egocentric image, a high-level goal instruction, and an optional auxiliary exocentric view, EgoForge generates egocentric rollouts that follow user intent and preserve scene structure without requiring dense supervision. [episode]
- Credal Large Language Models for Semantic Commitment under Uncertainty — Large language models often produce fluent but incorrect answers with unwarranted confidence, and this research introduces Credal Large Language Models (CLLMs) to represent uncertainty through credible sets rather than single predictive distributions. [episode]
- SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding — As a diligent researcher, I have meticulously analyzed the provided excerpts from both sources (A and B) regarding the SONIC-O1 benchmark paper. My synthesis will be comprehensive, detailed, and precise to ensure no critical nuance is lost. [episode]
- On-Device Named-Entity Recognition: A Deployability Study of Accuracy, Cost, Reliability, and Confidence — Named-entity recognition (NER) on-device deployment requires evaluating models based on deployability factors like latency and reliability, rather than just leaderboard accuracy. [episode]
- Cross-Context Review: Improving LLM Output Quality by Separating Production and Review Sessions — Large language models struggle to catch errors in their own outputs when the review happens in the same session that produced them, leading to systematic approval of flawed work. [episode]
- LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation — Language-conditioned goal navigation (LGN) requires agents to locate userspecified targets without step-by-step guidance, and this paper introduces HieraNav, an open-vocabulary LGN task with goals specified at four hierarchical semantic levels: scene, room, region, and instance. [episode]
- Stochastic Optimal Control for Continuous-Time fMRI Representation Learning — A foundational brain dynamics model utilizing stochastic optimal control and self-supervised learning bridges state-space modeling and modern representation learning to create an efficient, scalable framework for decoding fMRI signals. [episode]
- What Drives Compositional Generalization in Visual Generative Models? The Importance of Continuous Training Objectives — Compositional generalization, which is the ability for visual generative models to synthesize novel combinations of known concepts, remains inconsistent across different architectures. [episode]
- Generalized Design Choices for Deepfake Detectors — The effectiveness of deepfake detection methods often depends less on their core design and more on implementation details such as data preprocessing, augmentation strategies, and optimization techniques. [episode]
- DramaAgent: Agentic Storytelling Video Generation — DramaAgent proposes a hierarchical, agentic, and model-agnostic framework for long-form text-to-video and audio generation that addresses challenges like narrative drift and unstable character identity by decomposing generation into story planning, persistent character conditioni [episode]
- Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints — Fully Homomorphic Encryption (FHE) presents significant deployment challenges for intelligent systems by requiring non-linear operations to be replaced with polynomial approximations, which can lead to catastrophic numerical divergence in reinforcement learning due to the Bellman [episode]
- InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation — Simulating real personalities with large language models requires grounding generation in authentic personal data, and this framework addresses the gap by introducing an interview-grounded evaluation protocol for personality simulation at a large scale. [episode]
- LLM assisted writing deserves empirical evaluation — LLM-assisted writing deserves empirical evaluation because it raises questions about clarity, integrity, equity, and evaluation beyond simply being treated as a detection problem. [episode]
- Tangut Word Segmentation under Extreme Resource Scarcity: Integrating Traditional Lexicons and Unlabeled Text — Tangut word segmentation under extreme resource scarcity is addressed by integrating traditional lexicons and unlabeled text within a unified BIES–CRF framework. [episode]
- LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios — As a meticulous researcher, I have carefully analyzed both provided excerpts from the arXiv paper "LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios." My synthesis below aims to construct a comprehensive, detailed summary that captures the scope, methodol [episode]
- Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents — Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. [episode]
- Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings — Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge biases. [episode]
- MDToC: Metacognitive Dynamic Tree of Concepts for Boosting Mathematical Problem-Solving of Large Language Models — MDToC (Metacognitive Dynamic Tree of Concepts) is a novel three-phase prompting technique designed to enhance Large Language Models' mathematical reasoning by transforming abstract thoughts into concepts and evaluable calculations. [episode]
- Geometric Data Perturbation with Noisy-Anchor Alignment for Privacy-Preserving Collaborative Learning — As a diligent researcher, I have meticulously reviewed both provided texts concerning "Geometric Data Perturbation (GDP)" and its variants, specifically focusing on the proposed method: Anchor-Aligned Independent Geometric Data Perturbation with Anchor Noise (AA-I-GDP (anchor noi [episode]
- Real-time Appearance-based Gaze Estimation for Open Domains — Appearance-based gaze estimation (AGE) models often fail in practical, unconstrained scenarios due to limited image diversity and inconsistent label fidelity across different datasets, particularly along the pitch axis. [episode]
- TRIAGE: Dialectical LLM Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series — Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irregularly sampled medical time series (ISMTS), must deliver both calibrated risk scores for patient triage and interpretable rationales that clinicians can verify. [episode]
- When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora — A study was conducted to audit whether an authenticity asymmetry, created when synthetic data uses real corpus text as ground truth but generates distractors, translates into measurable shortcut reliance by trained reinforcement learning policies. [episode]
- Emergent Object Binding Has a Finite Spatial Horizon — Emergent object binding has been shown to be spatially local, decaying with distance and exhibiting a semantic floor that tracks object class, which explains behaviors previously unexplained by aggregate accuracy scores. [episode]
- Fully Online Decentralized Learning in Stochastic Games with Unknown Independent Chains — Fully online decentralized learning in stochastic games with unknown independent chains develops an algorithm that allows agents to learn stationary equilibrium policies without requiring knowledge of underlying transition kernels or synchronization across players. [episode]
- The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment — In this work, researchers propose and validate that emergent misalignment in large language models occurs when narrow finetuning causes them to bind learned behaviors to shared chat-template tokens, which then "piggyback" that behavior onto semantically unrelated test queries. [episode]
- Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation — Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation proposes TRACER, a novel framework that enhances mortality and readmission prediction by integrating severity-grounded medical knowledge graphs, trajectory retrieva [episode]
- Vision-language models for chest radiography do not always need the image — Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image. [episode]
- REVEAL: Robust Evolution of Vision-Language Models for Explainable AI-Video Detection — The LAVID framework proposes a novel agentic Large Vision-Language Model (LVLM) approach for diffusion-generated video detection by leveraging explicit knowledge enhancement and online adaptation to improve reasoning and reduce hallucination. [episode]
- Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs — When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in perception or in action? The core finding is that hidden states reliably encode these premise–perception mismatches even when mode [episode]
- What Do Rationales Communicate? A Message-Intervention Study in Role-Specialized QA — Role-specialized QA pipelines increasingly pass rationales from a reasoner to a verifier, but it is unclear what this message actually buys: better answers, stronger support assessment, or a new failure surface. [episode]
- CASCADE Conformal Prediction: Uncertainty-Adaptive Prediction Intervals for Two-Stage Clinical Decision Support — Effective medication management in Parkinson’s Disease (PD) is challenging due to heterogeneous disease progression, variable patient response, and medication side effects. [episode]
- Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs — When diversity collapse in parallel chain-of-thought sampling motivates inference-time interventions, this research investigates whether grafting high-PRM prefixes from pruned chains into sibling chains moves reasoning trajectories at the operating points where prior work suggest [episode]
- A Framework for Egocentric and Exocentric Procedural Understanding via Temporal Segmentation and Semantic Abstraction — Long-horizon ego/exo data contains rich procedural evidence, but are redundant, noisy, and costly to process or retain. [episode]
- MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation — MedVL-SAM2 introduces a unified 3D medical multimodal model that concurrently supports report generation, VQA, and multiparadigm segmentation, including semantic, referring, and interactive segmentation. [episode]
- Backdoor Containment via Expert Quarantine and Shutdown in LLMs — Backdoor large language models (LLMs) pose a serious security concern because they can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. [episode]
- FourierQK: Filter Shape, Admissibility and the Leakage-Coverage Law — Frequency-collapse attention [Zeris, 2026e] achieves large gains over standard dot-product attention by replacing the Q/K dot product with a bandpass-filtered inner product at a learned frequency. [episode]
- MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation — As a fastidious and diligent researcher, I have meticulously analyzed the provided excerpts from the paper "MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation." My objective is to synthesize this information into a comprehensive, detailed summary sui [episode]
- Reachability Is Not Generalization: Understanding Verb--Noun Decomposition in Assembly Action Recognition — Assembly actions are compositional, combining a manipulation (verb) with a part or tool (noun), yet traditional atomic classifiers assign zero probability to unseen combinations by construction. [episode]
- VISPA: Pluralistic Alignment via Automatic Value Selection and Activation — VISPA introduces a training-free pluralistic alignment framework that enables direct control over value expression by dynamic selection and internal model activation steering. [episode]
- On Reliability of Membership Inference Vulnerability Evaluation — Membership inference attacks (MIAs) are popular methods for empirically assessing data leakage, but reliably estimating their true positive rate (TPR), especially at low false positive rates (FPRs), requires excessive computational resources. [episode]
- The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering — As an AI researcher with a meticulous eye for detail, I have thoroughly analyzed both provided texts concerning "The Exceedance Design Effect." The material presents a sophisticated critique of standard methods used to estimate coverage guarantees (like those in conformal predict [episode]
- Beyond Idealized Patients: Evaluating LLMs under Challenging Patient Behaviors in Medical Consultations — Large language models (LLMs) are increasingly used for medical consultation, but their safety in high-stakes settings depends on how they handle patient inputs that are unclear, inconsistent, or misleading. [episode]
- More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe — A large-scale remote sensing vision-language model (VLM) can achieve competitive performance across diverse benchmarks without requiring specialized architectural modifications, provided it is trained at sufficient scale across varied data and tasks. [episode]
- MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models — Memory-augmented large language models extend reasoning beyond fixed context windows by maintaining long-term memory across interactions, but existing systems often suffer from heterogeneous memory contamination where functionally distinct memories become interchangeable and misl [episode]
- Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation — Reasoning language models (RLMs) have shown impressive capabilities in domains like mathematics and coding, but their performance gains are often limited because training them on tasks lacking reliable verifiers remains challenging. [episode]
- BudgetSchemaBench: A Budget-Swept Diagnostic for Schema Context in Text-to-SQL — BudgetSchemaBench introduces an execution-grounded diagnostic for evaluating how database schemas fit into a model's context window when constrained by token budgets, which matters for agents working with large catalogs. [episode]
- Weighted Data Selection: Sharp Upper-Half and Five-Dimensional Laws — How much risk does a small reweighted training support retain? This work establishes sharp laws governing the worst-case risk inflation when selecting a small subset of training data for weighted least squares regression. [episode]
- Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control — Large language models (LLMs) are increasingly considered for deployment as the control component of robotic health attendants, yet their safety in this context remains poorly characterized. [episode]
- Domain-Adapted Small Language Models for Reliable Clinical Triage — Accurate and consistent Emergency Severity Index (ESI) assignment remains a persistent challenge in emergency departments, where highly variable free-text triage documentation contributes to mistriage and workflow inefficiencies. [episode]
- Partial AUC Maximization from Positive-unlabeled Data — The proposed method addresses the challenge of maximizing partial area under the ROC curve (pAUC) from positive-unlabeled (PU) data, which is crucial in real-world applications like cybersecurity and medical care where labeled negative data are often unavailable. [episode]
- Dyna-DINO: Efficient ViT Distillation Via Adaptive Representation Anchoring — Adaptive representation anchoring is a novel training curriculum designed to accelerate knowledge distillation for Vision Transformers by adaptively selecting intermediate teacher features based on similarity metrics. [episode]
- AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors — As a fastidious researcher, I have meticulously analyzed both provided excerpts from the arXiv paper "AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors." My task is to synthesize these details into a comprehensive, high-fidelity summary suitable [episode]
- VideoSTF: Stress-Testing Output Repetition in Video Large Language Models — Video Large Language Models (VideoLLMs) are increasingly powerful for video understanding, yet they suffer from severe output repetition, a generation failure mode that existing benchmarks overlook. [episode]
- Unleashing Diffusion and State Space Models for Medical Image Segmentation — DSM, a novel framework leveraging diffusion and state space models to segment unseen tumor categories beyond training data, addresses the critical need for robust medical image segmentation by integrating these advanced generative and sequence modeling techniques. [episode]
- Robust Is Salient: An Informed Adversary Moves the Optimal Signal onto the Salience Pole — When an informed adversary shares an audience in a constrained signaling channel, the signal that best protects the truth becomes one that best describes it. [episode]
- From Literature to Hypotheses: An AI Co-Scientist System for Biomarker-Guided Drug Combination Hypothesis Generation — CoDHy, an interactive AI co-scientist system, addresses the challenge of systematically connecting biomarker mechanisms to actionable drug combination hypotheses in cancer research by integrating structured biomedical databases and unstructured literature evidence into a task-spe [episode]
- Multi-User mmWave Beam and Rate Adaptation via Combinatorial Satisficing Bandits — Multi-User mmWave Beam and Rate Adaptation via Combinatorial Satisficing Bandits addresses the challenge of efficiently learning optimal beam and data rate assignments in multi-user millimeter-wave Massive MIMO systems without explicit Channel State Information (CSI). [episode]
- Principled Design of Diffusion-based Optimizers for Inverse Problems — Score-based diffusion models are being developed as powerful priors for inverse problems, but their practical deployment is hindered by long inference times and sensitivity to hyperparameter tuning. [episode]
- Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation — A reward-based policy framework enables zero-shot transfer between environments with completely different observation spaces by conditioning the policy only on rewards and actions, which addresses a significant limitation in deep reinforcement learning where standard policies str [episode]
- A Low Grounding Score Is Not an Ungrounded Judge: Identifying the Perceptibility Confound in Multimodal Oversight — Model judges now supervise multimodal systems at scale, filtering training data, selecting outputs, and supplying the reward that shapes multimodal reasoning models. [episode]
- Query Independent Variable Rate Visual Token Coding — Visual-token compression for vision–language models is posed almost entirely as a selection problem, but this work introduces a method to keep every token and vary its rate, which solves the allocation problem exactly based on measured distortion–rate curves. [episode]
- PUMA: Learning a Mutation-Aware Vocabulary of Protein Units — PUMA introduces a method for discovering evolutionarily meaningful protein sequence units by explicitly incorporating mutational relationships into an iterative merging process, offering an interpretable vocabulary that bridges statistical frequency and evolutionary history. [episode]
- EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning — Local editing of 3D objects remains a long-standing challenge, and EditVerse3D proposes a novel end-to-end framework that enables high-quality object editing under coarse guidance by taking as input a 3D object, a coarse 3D bounding box indicating the target region, and a referen [episode]
- Modeling The Object Representations Underlying Human Physical Reasoning — Humans appear to represent objects for intuitive physics with coarse, volumetric “bodies” that smooth concavities – trading fine visual details for efficient physical predictions – yet their internal structure is largely unknown. [episode]
- MaPa: Text-driven Photorealistic Material Painting for 3D Shapes — This research introduces MaPa, a novel framework designed to generate photorealistic and editable materials for 3D meshes directly from textual descriptions. [episode]
- Augmented Equivariant Mesh Networks for Anatomical Segmentation — Anatomical mesh segmentation requires models that operate directly on irregular surface geometry while remaining robust to arbitrary patient pose and mesh resolution variation. [episode]
- The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating — Reading an LLM judge’s verdict from its first generated token reveals that this cheap readout distorts position bias in one direction, meaning figures obtained this way behave as upper bounds rather than accurate measures of judgment. [episode]
- Probing an Embodied LLM: When Higher Observation Fidelity Hurts Problem Solving — Large Language Models are increasingly proposed as cognitive components for robotic systems, yet their opaque decision processes make it difficult to explain success or failure in closed-loop embodied tasks. [episode]
- SAGA: Stable Acceleration Guidance for Autoregressive Video Generation — Autoregressive video diffusion models face instability due to recursive reuse of generated latents, leading to flickering and structural drift during long-horizon generation. [episode]
- DynGhost: Temporally-Modelled Transformer for Dynamic Ghost Imagings — Ghost imaging reconstructs spatial information from a single-pixel bucket detector by correlating structured illumination patterns with scalar intensity measurements, and this work introduces DynGhost, a transformer architecture that addresses limitations in existing deep learnin [episode]
- AEGIS: Anchor-Enforced Gradient Isolation for Knowledge-Preserving Vision-Language-Action Fine-Tuning — Adapting pre-trained vision-language models (VLMs) for robotic control requires injecting high-magnitude continuous gradients from a flow-matching action expert into a backbone trained exclusively with cross-entropy. [episode]
- Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks — Compute-grounded reasoning (CGR) is introduced as a design paradigm for spatial-aware research agents that resolves every answerable sub-problem through deterministic computation before engaging a language model, which matters because this approach yields more reliable, interpret [episode]
- Who Wrote the Book? Detecting and Attributing LLM Ghostwriters — Out-of-distribution generalization remains insufficiently supported by existing authorship attribution datasets, which often focus on short texts and fail to test for unseen authors or domains. [episode]
- Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems — The gist The framework presents Adaptive Policy-Guided Error Mitigation (APGEM) as a context-aware orchestration layer for Hybrid Quantum Reinforcement Learning on NISQ systems. [episode]
- CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding — Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and answer user questions under tight latency constraints. [episode]
- Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models — Diffusion Language Models (DLMs) present a promising alternative to autoregressive models by generating text through iterative denoising, yet this iterative process introduces unique safety vulnerabilities where harmful tokens can propagate across steps to induce unsafe outputs. [episode]
- Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans — Architect-Ant introduces a framework for furnishing residential floor plans by treating furniture layout synthesis as structured sequence generation over editable geometric objects, combining pseudo-labeled data with deterministic rule scoring and direct preference optimization t [episode]
- Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores — Vision Transformers (ViTs) are challenged by their quadratic computational cost due to dense token-to-token self-attention, which limits their scalability in high-resolution domains. [episode]
- Mapping and Classification of Trees Outside Forests using Deep Learning — Trees Outside Forests (TOF) play an important role in agricultural landscapes by supporting biodiversity, sequestering carbon, and regulating microclimates. The gist: FT-UNetFormer achieved the best mean metrics, with mIoU score of 0.739 and mF1 score of 0.843. [episode]
- DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations — Existing hallucination taxonomies classify LLM errors by what is wrong with the output—memorized misconceptions, reasoning failures, fluent fabrications—but this paper proposes a complementary taxonomy that classifies errors by their detectability signature. [episode]
- Meta-reinforcement learning with minimum attention — Minimum attention applies the least action principle to changes of control concerning state and time, first proposed by Brockett. [episode]
- Geometric Stability: The Missing Axis of Representations — This research introduces Geometric Stability as a critical, distinct axis for analyzing internal geometries of neural network representations, arguing that existing methods focusing solely on alignment (like CKA or Procrustes distance) are insufficient because they only measure s [episode]
- Exploiting Exogenous Structure for Sample-Efficient Reinforcement Learning — As a fastidious and diligent researcher, I have thoroughly reviewed the provided excerpts from Wan et al.'s paper, "Exploiting Exogenous Structure for RL 50." My analysis confirms that this work addresses a critical challenge in reinforcement learning—the sample complexity requ [episode]
- Flow-Transformed Implicit Processes for Function-Space Variational Inference — Implicit-process priors define distributions over functions through flexible generative mechanisms, making them attractive for Bayesian function-space modelling. [episode]
- Efficient Generative Modeling beyond Memoryless Diffusion via Adjoint Schr"odinger Bridge Matching — Diffusion models often yield highly curved trajectories and noisy score targets due to an uninformative, memoryless forward process that induces independent data-noise coupling. [episode]
- Measuring Iterative Temporal Reasoning with Time Puzzles — Tool use, such as web search, has become standard in large language models (LLMs), but existing benchmarks fail to capture how LLMs perform temporal reasoning when augmented by tools. [episode]
- The COTe score: A decomposable framework for evaluating Document Layout Analysis models — Document Layout Analysis (DLA) models typically rely on general object detection metrics like IoU, F1, or mAP, which are ill-suited for printed media because they treat images as 2D projections of 3D space rather than natively 2D tessellations. [episode]
- The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages — Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models, but its reliability remains largely unexplored beyond English and across diverse model families. [episode]
- Text-Preserving Lossy Text Compression: A Study of Strategic Deletion and LLM Reconstruction — Lossy semantic text compression, where an encoder strategically deletes parts of a source text and a large language model (LLM) reconstructs it from the retained skeleton, is studied to find practical strategies for context- and bandwidth-limited processing in LLM pipelines. [episode]
- SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering — SPADER is a reinforcement learning framework designed to improve long-horizon tool use reasoning in Multi-Answer Question Answering by addressing challenges in fine-grained credit assignment and coverage-oriented exploration. [episode]
- Aligned but Not Partner-Specific: How Multimodal LLM Agents Succeed in Reference Games Without Forming Conceptual Pacts — Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific conventions grounded in shared interaction history. [episode]
- VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation — VideoWeaver introduces an agent harness and benchmark designed to evaluate and evolve skills for long video generation, addressing the gap in understanding how general-purpose agent frameworks can handle complex, long-horizon multimodal tasks by enabling agents to compose their o [episode]
- REALM: Retrospective Encoder Alignment for LFP Modeling — Spike activity has been dominant for behavior decoding, but Local Field Potentials (LFPs) offer a stable, low-power alternative that can achieve competitive performance in real-time brain-computer interfaces. [episode]
- Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection — The provided text contains both a detailed technical description (A) and a high-level empirical summary (B). I will synthesize these into a thorough, rigorous overview of the paper's methodology, contributions, and findings. [episode]
- Self-conditioned Flow Map Language Models via Fixed-point Flows — Self-conditioned flow language models implicitly learn a fixed-point iteration that refines its own denoising estimate, leading to novel flow map language models capable of one- and few-step generation. [episode]
- Learning Linear Systems under Heavy-Tailed Noise: A Non-Asymptotic Analysis from A Single Trajectory — Learning linear systems under heavy-tailed noise involves establishing non-asymptotic sample complexity bounds for least-squares estimation of vector autoregressive models using a single observed trajectory, which is crucial for understanding the statistical viability of model-ba [episode]
- TANGO: Treating Tokens as Operators — Token-Aggregated Nonlinear Gating Operators (TANGO) and its variant, Windowed Aggregation of Nonlinear Gating Operators (WANGO), introduce a novel block structure that replaces standard self-attention and position-wise feed-forward networks with a single cross-token gated residua [episode]
- De-GAN - Dynamic Parameter Tuned GAN for 3D Medical Image Segmentation: A Step Towards Generalisation — Brain tumor segmentation remains challenging due to low contrast in enhancing tumors and domain shift caused by scanner and site variations, which DE-GAN addresses by synthesizing slice-adaptive FLAIR images that are then used for 3D U-Net training. [episode]
- Spatial Strategies, Not Actions: Vector-Quantized Geodesics as Tools for LLM-Driven Agents — Large language model (LLM) based agents are often criticized for lacking spatial understanding and mainly exploiting statistical text patterns, but this work investigates their spatial comprehension through an architecture combining geometrical tools with an LLM serving as a high [episode]
- Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM-Based Urban Sensing — Street-view imagery adds most information when an urban attribute is visible and existing records are sparse, providing guidance for image collection and map provenance. [episode]
- LEAD: Layer-wise Expert-aligned Decoding for Faithful Radiology Report Generation — Layer-wise Expert-aligned Decoding (LEAD) is a novel framework designed to mitigate hallucinations in radiology report generation by inherently modifying the Large Vision-Language Model's decoding trajectory. [episode]
- Textual Planning with Explicit Latent Transitions — Planning with large language models is bottlenecked by token-by-token generation and repeated full forward passes, making multi-step lookahead and rollout-based search expensive in latency and compute. [episode]
- In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement — Drunk language, defined as language written under the influence of alcohol, acts as an emergent driver for safety failures in large language models (LLMs) by inducing behaviors analogous to those seen in intoxicated humans. [episode]
- A framework for auditing grounding claims — This research introduces a novel, rigorous framework designed to systematically audit and diagnose the quality of symbolic grounding in artificial agents, addressing the fundamental "symbol grounding problem" (SGP)—how discrete tokens like "cat" acquire meaningful reference to [episode]
- A False Discovery Rate Control Method Using a Fully Connected Hidden Markov Random Field for Neuroimaging Data — False discovery rate (FDR) control methods are essential for voxel-wise multiple testing in neuroimaging data analysis, where hundreds of thousands or even millions of tests are conducted to detect brain regions associated with disease-related changes. [episode]
- Metric-Construction Coupling Inflates Measured Synthetic Dialect Recovery — Synthetic data generation for low-resource languages like Korean presents a critical challenge because existing corpora are often unavailable for redistribution, making it difficult to measure how much real-data gain synthetic supervision can recover. [episode]
- Provable FDR Control for Deep Feature Selection: Deep MLPs and Beyond — A flexible feature selection framework based on deep neural networks is developed that provides a theoretical guarantee for controlling the false discovery rate (FDR), addressing a gap between high-dimensional statistics and explainable AI. [episode]
- StoCFL: A Stochastically Clustered Federated Learning Framework for Non-IID Data with Dynamic Client Participation — Clustered federated learning (CFL) addresses performance degradation caused by Non-IID data in federated learning by grouping clients with similar data distributions, and this paper introduces StoCFL, a novel framework that enhances CFL by incorporating stochastic client clusteri [episode]
- Geometry-Aware Adaptation for Pretrained Models — Machine learning models often struggle to reliably predict new classes when trained on datasets where labels are only a small proportion of a larger label space, and this work proposes LOKI, an approach that exploits the geometric information (metric structure) between labels to [episode]
- RAD: A Dataset and Benchmark for Real-Life Anomaly Detection with Robotic Observations — Anomaly detection is essential for robotic perception and industrial inspection, yet most benchmarks are collected under controlled conditions with fixed viewpoints and stable illumination. [episode]
- MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference — The Minimal Expression-Replacement GEneralization (MERGE) test introduces an automated methodology to evaluate how well current state-of-the-art reasoning models generalize when faced with minimal, non-adversarial variants of existing Natural Language Inference (NLI) datasets. [episode]
- Roto-translated Local Coordinate Frames For Interacting Dynamical Systems — Modelling interactions in complex, non-linear dynamical systems requires accounting for their inherent symmetries to achieve better generalization. [episode]
- Supervise What Decides Success: Criterion-Aligned Auxiliary Losses for Latent World-Model Planning — Latent world models plan by scoring candidate action sequences with distances in latent space, but success is judged by physical quantities, and this work proposes an auxiliary loss that uses these success-criterion quantities as training targets to improve planning performance. [episode]
- CAVE-Mem: Boundary-Aware Experience Validation for Memory Search — Long-term memory agents increasingly rely on iterative search and reusable experience to answer complex questions, but current systems optimize relevance in a way that can introduce validity problems when memory substrates or question intents change. [episode]
- Uncertainty-Aware RL-Controlled Adaptive 3D Mapping — Voxel-based volumetric mapping is fundamental to 3D reconstruction, yet fixed-resolution grids remain inherently inefficient – wasting memory in uniform regions and losing detail in complex ones. [episode]
- Domain generalization and synthetic data in object detection: the enabler, the probe, and the gap — Object detection models often experience performance degradation when deployed under distribution shifts, caused by for example changes in weather type, operational environment, or object appearance. [episode]
- Robust Prior-Guided Segmentation for Editable 3D Gaussian Splatting — 3D Gaussian Splatting (3D-GS) enables real-time 3D scene reconstruction but lacks robust segmentation for editing tasks such as object removal, extraction, and recoloring. [episode]
- Signed Lexical Confidence for Risk-Calibrated Intent Routing — Selective intent routing allows an assistant to act on reliable predictions while deferring uncertain requests, and this work introduces a signed lexical gate that combines sentence classifier logit margin with sparse lexical model support to enhance risk-calibrated intent routin [episode]
- Multi-perspective monitoring of wildlife and human activities from camera traps and drones with deep learning models — Multi-perspective monitoring of wildlife and human activities from camera traps and drones with deep learning models addresses the need for understanding spatial distributions of wildlife and human activities to evaluate interactions and inform conservation planning. [episode]
- SCM-based Fairness and Faithful Explainability for Legal Document Classification — SCM-based fairness regularisation for LegalBERT investigates whether debiasing interventions that change fairness also change how faithfully explanations reflect model reasoning. [episode]
- Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks — Detailed, profession-specific system prompts raise token use and estimated cost per response without a consistent gain in accuracy. [episode]
- STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets — Inferring hidden physical properties from motion remains challenging for foundation models, especially when dealing with opaque, asymmetric rigid bodies where surface cues are unreliable. [episode]
- VITA: A Multi-Source Vicinal Transfer Augmentation Method for Out-of-Distribution Generalization — Invariance to diverse types of image corruption, such as noise, blurring, or colour shifts, is essential to establish robust models in computer vision. [episode]
- Verbal tics in frontier language models: A critical review of current releases, research evidence, and public discussion — Verbal tics in frontier language models are repetitive, formulaic linguistic patterns that emerge due to alignment techniques like RLHF, and this phenomenon highlights a significant "alignment tax" on linguistic diversity and authenticity. [episode]
- Linguistic traces of stochastic empathy in language models — Large language models (LLMs) exhibit a chameleonic ability to adjust their writing style and content when instructed to appear human, revealing underlying linguistic strategies that suggest they rely on implicit representations of humanness rather than genuine empathy. [episode]
- Robust Online Aero-Engine Blade Defect Detection via Dual-Alignment Test-Time Adaptation — Reliable visual inspection is essential for quality assurance in aero-engine blade manufacturing, where defect appearance may vary across production lines, imaging conditions, blade poses, and surface backgrounds. [episode]
- Learned Suppression for 3D Keypoint Detection with a Graph-Transformer Backbone — SAGE3D presents a hybrid Transformer-based model for corner detection in airborne LiDAR point clouds, addressing the challenges of extreme class imbalance and lack of grid structure inherent in 3D scene understanding. [episode]
- Measuring Human-Like Bias in LLMs? A Critique of Human-Derived Bias Constructs in LLM Evaluation — Researchers increasingly use human-derived bias constructs to study Large Language Models (LLMs), including social-cognitive constructs such as implicit bias and stereotype activation, and cognitive biases such as anchoring, framing effects, and confirmation bias. [episode]
- When Guessing is Rewarded: Rethinking Language Model Evaluation with Distributional Uncertainty Scoring — Common evaluation paradigms for language models focus on scoring single responses through accuracy metrics or proper scoring rules, failing to capture the full richness of a model’s belief state. [episode]
- Port-Hamiltonian Neural Networks for Systems with Multiple Asymptotically Stable Equilibria — Stable port-Hamiltonian neural networks certify asymptotic stability by construction, yet their global Lyapunov function structure restricts them to systems with only one attractor. [episode]
- GPEC: Efficient Pre-LLM Gaussian Process Embedding Correction for Cardiac Video Caption Generation — Multimodal large language models (MLLMs) show promise for video understanding, but their performance degrades in specialized medical domains like echocardiography, necessitating methods to improve caption generation accuracy. [episode]
- Graph Structure Learning with Temporal Graph Information Bottleneck for Inductive Representation Learning — Temporal graph learning is crucial for dynamic networks where nodes and edges evolve over time and new nodes continuously join the system, making inductive representation learning in such settings challenging due to ineffective representation of unseen nodes and noisy or redundan [episode]
- Spatial Lifting for Dense Prediction — Spatial Lifting (SL) is a novel methodology for dense prediction tasks that lifts standard inputs into a higher-dimensional space and processes them using networks designed for that higher dimension, offering improved performance while reducing model parameters and inference cost [episode]
- The Benchmarking Epistemology: Validity Theory for Evaluating Machine Learning Models — Predictive benchmarking, which evaluates machine learning models based on predictive performance and competitive ranking, requires explicit conditions of construct validity to support substantive scientific inferences. [episode]
- Evaluating the Robustness of Anti-UAV Detection under Controlled Fog Degradation: Fog-Aware Training and Clear-Sky Tradeoff — Vision-based anti-UAV systems must function in poor visibility, yet most benchmarks use only clear-sky footage, and previous work treats fog as simply present or absent. [episode]
- MonoPhysics: Estimating Geometry, Appearance, and Physical Parameters from Monocular Videos — MonoPhysics is a framework designed for monocular inverse physics estimation of deformable objects by jointly optimizing geometry, appearance, and physical parameters from a single camera view. [episode]
- Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale — The Platonic Representation Hypothesis suggests that neural networks trained on different modalities align and eventually converge toward the same representation of reality as they scale. [episode]
- SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples — As machine learning systems increasingly rely on public, untrusted data sources, data poisoning attacks pose a growing threat by injecting malicious examples into training data to induce misclassification. [episode]
- Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders — Multilingual encoder-based language models are widely adopted for code-mixed analysis tasks, yet we know surprisingly little about how they represent code-mixed inputs internally — or whether those representations meaningfully connect to the constituent languages being mixed. [episode]
- The Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion Models — Masked diffusion language models (MDMs) uniquely support any-order generation, but confidence-based decoding serves as the de facto standard inference policy, which can be fundamentally misaligned with complex reasoning trajectories. [episode]
- BalLOT: Balanced k-means clustering with optimal transport — BalLOT introduces an optimal transport approach to alternating minimization, demonstrating that it provides a fast and effective solution to balanced k-means clustering by reformulating cluster assignment as an optimal transport linear program. [episode]
- EyeTheia: A Lightweight and Accessible Eye-Tracking Toolbox — EyeTheia introduces an open, lightweight deep learning pipeline for webcam-based gaze estimation, designed to be accessible for browser-based experimental platforms and real-world cognitive and clinical research. [episode]
- AstroAgentBench: Evaluating Agentic Planning on Space Mission Planning Tasks — ecent advances in agentic Large Language Models (LLMs) have positioned them as generalist planners capable of reasoning and acting across diverse tasks. [episode]
- Legal text classification in Korean sexual offense cases: from traditional machine learning to large language models with XAI insights — Legal text classification in Korean sexual offense cases: from traditional machine learning to large language models with XAI insights investigates the evolution of AI models for classifying complex legal documents, demonstrating that fine-tuning small language models yields supe [episode]
- FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering — FocusGraph develops a novel framework for keyframe selection in long egocentric videos to enhance question answering, addressing performance degradation and increased inference time associated with using multimodal large language models (MLLMs) on extended video sequences. [episode]
- Iterative Topic Taxonomy Induction with LLMs: A Case Study of Electoral Advertising — By combining unsupervised clustering with iterative prompt-based inference from large language models, this research introduces an end-to-end framework for automatically inducing interpretable topic taxonomies from unlabeled political advertising text corpora. [episode]
- Learning Transferable Skills using Goal-Conditioned Bisimulation — Unsupervised skill discovery methods are being developed to pretrain general-purpose policies using reward-free datasets, but current approaches often fail to transfer effectively across different environments. [episode]
- Compressing History into Memory: Distilling Transformers into Recurrent Transformers — Recurrent Transformers address computational limitations in long-horizon sequential data processing by distilling the compression strategy from full-history transformers into fixed-size memory. [episode]
- Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational Humor — Automated rewards for training language models in conversational humor are investigated by examining reward exploits and designing countermeasures to preserve intended humorous behavior. [episode]
- Meta-Multi-Agent Reinforcement Learning for Fast Adaptation of Interactive Policies with Applications to Autonomous Driving — Meta-Multi-Agent Reinforcement Learning for Fast Adaptation of Interactive Policies with Applications to Autonomous Driving develops a meta-MARL framework to enable rapid adaptation of interactive policies in multi-agent systems by modeling problems as Markov games and defining a [episode]
- Wrivinder: Towards Spatial Intelligence for Geo-locating Ground Images onto Satellite Imagery — Aligning ground-level imagery with geo-registered satellite maps is crucial for mapping, navigation, and situational awareness, yet remains challenging under large viewpoint gaps or when GPS is unreliable. [episode]
- Multi-Resolution Feature Fusion U-Net for Magnetic Resonance Imaging Segmentation — The segmentation of anatomical structures in MRI scans is crucial for clinical diagnosis, but existing Deep Learning architectures often struggle to preserve fine-grained details and global contextual information due to the irregular boundaries and variations in shape characteris [episode]
- Integrating Fairness and Explainability in a Multiple Instance Reinforcement Learning System — Predicting student performance from educational interaction data requires models that are both accurate and sufficiently transparent to support meaningful intervention, while demographic information creates an additional risk of unfair predictions. [episode]
- Design and Implementation of a Kalman Filter-Infused Algorithm for Tilt Estimation — Accurate tilt angle estimation is crucial for robotics and embedded control systems, and this paper presents a single-axis tilt angle estimation system based on an MPU6050 inertial measurement unit implemented on an RP2040 microcontroller platform, utilizing a Kalman filter for s [episode]
- Do Vision Language Models Understand Human Engagement in Games? — Inferring human engagement from gameplay video is important for game design and playerexperience research, yet it remains unclear whether vision–language models (VLMs) can infer such latent psychological states from visual cues alone. [episode]
- Image-Domain Poisson-Perturbation Robustness of NCCT Slice Classification — Image-domain Poisson perturbation may alter normalized NCCT appearance and downstream models, prompting an investigation into model robustness under simulated noise. [episode]
- Poincar'e Meets Bellman: Revisable Memory, Operational Quotients, and Evidence-Supported Learning in Changing Environments — Intelligence is re-framed as a dynamical cycle governed by the Context-Content Uncertainty Principle (CCUP), which seeks to balance the Euler equation by minimizing the topological mismatch between dynamic context flow and static content scaffold. [episode]
- SR-Ground: Image Quality Grounding for Super-Resolved Content — Super-Resolution (SR) models, especially diffusion-based ones, introduce subtle yet perceptually significant visual artifacts that existing Image Quality Assessment (IQA) methods fail to distinguish or localize. [episode]
- Revision-Aware Independent Agent Graphs for Dynamic Reasoning —
- Science Utopia? Closed-Loop LLM Simulation of Academic Research Ecosystems —
- IQS-BO: In-Context Query Selection for Bayesian Optimisation —
- Know When to Hold 'em: Correct-Token Retention in Uniform-State Diffusion Language Models —
- SCOPE-AD: Sequential cost-aware ordinal-belief planning with energy-based models for diagnostic agents —
- PickMoment: Continuous-Time Single-Image-to-Video via Learning Deblurring and Blur-to-Video —
- ShelfChange3D: Object-Level 3D Change Detection for Retail Shelf Monitoring —
- Dyna3: VLM-Guided Training-Free 4D Reconstruction via Depth Foundation Models —
- ODDR: One-Step Deshadow Diffusion via Reward Guidance —
- STAGE: Subspace-Targeted Affine Generative Erasure for Text-to-3D Models —
- DAYJOB: A Benchmark for Long-Horizon Professional Work —
- ARROW: Arbitrary Reconstruction and Tracking of 4D Observations in the Wild —
- What Wins a Vote? Formatting, Length, and Lexical Diversity in the French Compar:IA LLM Arena —
- Clifford Sheaf Neural Networks —
- Evaluating Biomedical Reranking for LLM-Based Question Answering over Longitudinal Clinical Notes —
- CLASP: Continual Low-rank Adapters for Spatially Placed Concepts from One Hypernetwork —
- ARCCS: An Automated Regulatory Compliance Checking System —
- MMVistaReason: Toward Open-Data and Post-Training Recipes for Multimodal Reasoning —
- Does AI-Generated Scientific Text Follow Human Argumentation Patterns? A CARS-Based Comparison of Research Article Introductions —
- Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects —
- Gacha Decoding: Eliciting Diverse Generations Through Instruction Following —
- Supervising Sound Localization by In-the-wild Egomotion —
- AiSearch: Interactive Multi-Modal Search with VLMs —
- LLM-Assisted Discovery of Typed Semantic Links for Ontology Network Construction —
- Smoother Flow Matching via Contrastive Trajectory Repulsion —
- Localisation-Aware Uncertainty for Pretrained Object Detection —
- Optimal Transport Meets Reinforcement Learning: A Survey —
- SHAMS: An Audio-Grounded Pronunciation Benchmark for Levantine Arabic —
- Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs —
- MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs —
- The Impact of Processing Parameters on High-Accuracy Measurements in UAV Photogrammetry —
- Uncertainty-Guided Handshake: Efficient Human-in-the-Loop Refinement for Surgical-Grade Glioma Segmentation —
- Tight Transition Time Bounds for Separable Logistic Regression at the Edge of Stability —
- When Does a Second Model Help? Cross-Model Review in LLM Verification —
- Zero Flux: Flow-Based Comparison of High-Dimensional Discrete Distributions —
- FiVOS: A Fish Segmentation Algorithm Based on Interactive Video Object Segmentation and Filter Enhancement —
- The Persona Is Still There, but Who Is Speaking? Latent Identity Reversion in Persistent AI Agents —
- Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories —
- No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse —
- SALD: Self-Referenced Advantage Learning for Diffusion Models —
- VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation —
- FedCKA: Representation-Guided Layer Personalization for Federated 3D Perception Across Driving Domains —
- GAW-PO: Preference Optimization with Gradient-Aligned Token Weights —
- VoxelSynth3D: Interpretable Volumetric Image-Domain Metal Artifact Reduction with a Paired Synthetic CLINIC-Metal Benchmark —
- How the Audit Rule Shapes Faithful Factor Explanations in LLMs —
- SuperMotion: Source-Preserving Denoising for Text-Driven Human Motion Editing —
- Langevin-Informed Transfer Learning: Replacing Target Samples by Black-Box Feedback —
- Exact Distinguishability in Non-Markovian Decision Processes —
- Towards Reliable Vision-Language Models for Autonomous Driving —
- Synthetic training for long-tail haemorrhagic lesion segmentation in data-scarce settings —
- Revisiting Cross-Reconstruction for Generalizable Deepfake Detection —
- QK-Wanda: Coupling Queries and Keys for Unstructured Pruning —
- AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models —
- The hidden advantage of mask resampling: a theory of masked autoencoders —
- PAGER: Partial-to-global Alignment via Geometric and Relational Distillation —
- Two Routes to the Middle: Placement Search and Brain Readouts Converge on Where Continual Learners Should Specialize —
- Which LLM to pick? Online Active Model Selection for Large Language Models —
- Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs —
- Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning —
- Oneira: From Open-Ended Generation to Open-World Interaction in Video World Models —
- Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness? —
- Beyond Domain-Level Adaptation: Margin-Oriented Semantic-Appearance Interaction Correction for Personalized Federated Vision-Language Models —
- What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language —
- Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yor`ub'a —
- Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering —
- Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference —
- DiVid: Diagnosing Dimension-Specific Diversity Collapse in Video Generation Models —
- pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows —
- Do MLLM Judges Judge the Edit? Auditing Bias in Image Editing Evaluation with Verified Quality Preservation —
- When Text-to-Image Helps Editing: The Effects of Conditioning During Denoising —
- Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models —
- Compound interpretation is based on analogy —
- Acmite: Mitigating Gender Bias in LLMs through Concept-Guided Mutual Information —
- Task-Oriented Rank Adaptation for Continual Learning in Text Classification —
- MEGA: Object-Level Mesh Extraction from 3D Gaussian Splatting via Spatial Visual Distillation —
- CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement —
- In-context Learning of Single-index Targets: Comparing Kernel and Feature Learners —
- Rethinking Memorization Mitigation in Diffusion Models: Reinforcing Text Conditioning —
- ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection —
- 3DROID: A Renderable 3D Gaussian Dataset with Measured Per-Scene Reliability —
- FFBL-Coop: Association-Decoupled Cooperative 3D Multi-Object Tracking —
- Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding —
- GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking —
- PhysDEM: Physics-Defined Energy-Matching Diffusion for Spatiotemporal Field Generation under Scarce Measurements —
- OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction —
- VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding —
- A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering —
- GIFTBench: Diagnosing Generalization in Image Forgery Localization and Informing Model Design —
- VETO: Video Efficient Token Optimization for Vision Language Models —
- Inferring Multi-Timescale Neural Dynamics with Switching Linear Dynamical Systems —
- Continuous Conditioning of VLAs with Augmenting EMG and Visual Task Descriptors —
- A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings —
- PhaseAT: Fourier Phase Adversarial Training for Medical Image Domain Generalization —
- Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models —
- The Asymptotics of Language Model Alignment with Memory —
- LiteReality-Agent: An Agentic System for Interactable 3D Indoor Scene Reconstruction —
- From Pixels to Policy: A Multi-Agent System for Intervention and Geo-Spatial Decision Support —
- EvenSplat: Coupled 2D-3D Decomposition for Gaussian Splatting under Exposure and Illumination Variation —
- Memory-Guided B-Roll Generation from User Video Collections —
- Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2 —
- Unsupervised Domain Adaptation for Enhanced Radiometer Image Precipitation Estimation using Conditional Flow Matching —
- A foundation for systematic analysis of transformers and RNNs for tractography —
- MapLightning: Online Vectorized HD Map Construction with 1D Map Tokens —
- DecomVoxel: Harnessing 3D-Native Priors with Guided In-situ Denoising Optimization for Decompositional Scene Reconstruction —
- MoLE: Mixture of Latent Experts for Complementary Visual Reasoning —
- Cross-Lingual Alignment for Decoder-Only Models using MoE Routers —
- CLoSeR: Closing the Loop for Long-Context Streaming Reconstruction —
- Error-Corrected Inference-Time Scaling for Imperfect Diffusion Models —
- A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined —
- Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens —
- Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models —
- Anti-Persona: Disrupting Unauthorized Identity Binding and Recognition in Personalized Vision--Language Models —
- Latent JEPA: Abstract Future Prediction for Latent Reasoning in Chemistry —
- Sharp Non-Asymptotic Analysis of the Penalized Challenger in beta-EB-TCI for Bernoulli Bandits —
- EndoLive: Real-Time Style Transfer for Endoscopic Endonasal Skull Base Surgical Video —
- SIEVE: Selective attention-value Suppression for Vision-Language Models Unlearning —
- Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage —
- RASteer: Retain-Aware Activation Steering for Concept Erasure in Diffusion Models —
- Token-Level Video Reinforcement Learning —
- The Curvature of Regret in Contextual Linear Optimization —
- Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities —
- Continual Concept Erasure in Diffusion Models by Suppressing Cross-Edit Interference —
- Comparing a gradient boosting algorithm to the GOES FDC for wildfire detection —
- From Reasoning Failures to Composable Video Spatial Intelligence —
- Weather-Aware Domain Adaptation for Street-View Weather Recognition —
- Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks —
- Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents —
- Exploring Weaknesses of Generative Image Watermarks against Latent Frequency Masking —
- Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization —
- Task-Adaptive Grounded 3D-Programmers Using 2D VLMs —
- Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation —
- CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning —
- Typological Alignment of Stack-Based Language Models on Mildly Context-Sensitive Artificial Languages —
- DiDE:Direct Injection with Color-Texture DEcoupling for 3D Stylization —
- Form and Void: Entangled Composition through an Autonomous AI Agent —
- Learning from Failure: Leveraging Unreliable Predictions in Semi-Supervised Real-World Adverse Weather Removal —
- LLM-as-Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them —
- Wasserstein Gradient Flows and Forward-Only Diffusion Are Not Enough for Multimodal Sampling —
- GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning —
- Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic) —
- Surface-volume self-supervised representation learning of brain MRI for genetic discovery —
- A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification —
- Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes —
- Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows —
- Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation —
- MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI —
- Finetuning with Sampling: SFT Learns Better Than You Think —
- Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models —
- Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation —
- From Knowledge Access to Source Learning: Developing Source-Specific Competence —
- MosaiChunk: Compositing Spatio-Temporal Memory for Autoregressive Video Generation —
- Muon meets Tamed Langevin: Momentum Preconditioning beyond Convex and gradient-Lipschitz Potentials —
- 4Director: Controlling Video World Models with Rigid 3D Geometry —
- World Observer: Joint Actor-Observer Generation for Persistent World Modeling —
- AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents —
- Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair —
- Effective Resistance and Graph Neural Network Reliability in Tissue-Specific Interactomes —
- Generative Cinematographer: Composing Camera and Object Motion in 3D —
- OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning —
- DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation —
- Hierarchical Continuous Diffusion Language Models —
- HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation —
- FERPO: Forward Entropy-Regularized Policy Optimization —
- VISTA: A Visual Harness for Reasoning in an Interactive World —
- SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation —
- ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research —
- Embedding Prediction Helps Image Generation —
- PROWBench: Do Video Models Render What the Program Specifies? —
- KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards —
- One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars —
- Sphere Encoder 2 —
- Moore, Escher, Penrose: A Conformal Golden Braid —
- A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications? —
- LENS-GRF: Permutation-Invariant Lesion Evidence Network with Gated Residual Fusion for Acne Severity Grading and Multi-Rater Clinical Oracle Analysis —
- Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning —
- Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping —
- Beyond Pixel Reconstruction: Retrieval-Guided Glyph-Aware Restoration for Low-Resource Manchu Historical Documents —
- DuplexSpeechBench-Document Grounding: Benchmarking Document Grounding and Hallucinations in Voice Agents —
- EgoRefine: Ego-Referenced Predictive Alignment and Trajectory-Conditioned Reliability-Aware Fusion for Asynchronous Collaborative Perception —
- Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning —
- CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters —
- ContractRL: Shielded Group-Relative Policy Optimization for Auditable Tool-Call Repair —
- LEGO-OPD: Factorized Teacher Composition for Multimodal On-Policy Distillation —
- Vmem- phi: Low-Compute Out-of-Distribution Detection in Spiking Neural Networks from Membrane-Potential Statistics —
- Manifold-Constrained Initial Noise Optimization for Efficient Generative Model Alignment —
- M squared Weather: A Benchmark for Joint Multi-Station and Multi-Variable Weather Forecasting —
- When Do Attention-Head Ablations Support Causal Claims? Projection-Level Confounds, Floor Effects, and Matched Controls —
- STCFormer: Adaptive Spatio-Temporal Modeling with Dynamic Cluster Transformer for Station-based Weather Forecasting —
- FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training —
- MatrixReward: Reward from Rubric Matrix for Open-Ended Generation —
- Dissonant ballerinas and crafty carrots: a comparative multi-modal analysis of Italian brain rot —
- RACE: Residual-Aware Test-Time Adaptation for Neighbor-Rich Time-Series Foundation Model Forecasting —
- LLM-as-a-Judge for Low-Resource Languages: Adapting Ragas and Comparative Ranking for Romanian —
- UniBuc at SemEval-2024 Task 2: Tailored Prompting with Solar for Clinical NLI —
- VANDAM: Viewing a nucleotide sequence with DNA molecular priors —
- From Image Latent Space to Fuzzy Rules: Interpretable Analysis of Gastrointestinal Foundation Model —
- Transferable Graph Metanetworks —
- Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification —
- Learning to Cover Locally: Graph Neural Combinatorial Optimization under a Hard Information Horizon —
- Target-Dependent Limits of Causal Repair: A Leading-Log Frontier in a Gaussian Model —
- IrekoGPT: Turning Structured Pruning into Post-Hoc Slimmable LLMs —
- ChainLoRA: Geometry-Preserving Task Vector Merging for Continual Learning in LLMs —
- Exact information accounting for SGD methods —
- Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation —
- PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video —
- PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion —
- EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights —
- Fractional Laplace Neural Operators: Exact Architectures, an Expressivity Frontier at Criticality, and Certified Stability for Memory-Driven Network Dynamics —
- Rules Amortize, Pairings Don't: Linguistic Structure Determines What Latent Task Representations Can Replace In-Context Learning —
- Adaptive Conformal Prediction for Image Regression Models with Application to an Inertial Confinement Fusion Emulator —
- Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models —
- Memorizon: Training World Models Beyond Their Context Window —
- PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop —
- Can LLMs Reason Over Long Horizons? An Empirical Evaluation of Context Strategies for Longitudinal Clinical Reasoning —
- Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness —
- FORTE: Adaptive Scoring and Exact Keyframe Selection for Long-Video Question Answering —
- Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL —
- Gestalt: Large Multimodal Interplay Model —
- Discrete Annotation, Continuous Preference: Rethinking Supervision for Accurate and Generalizable Aesthetic Image Cropping —
- Right In-Place (RiP) Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference —
- Just Align x: Aligning Predictions, Not Representations —
- Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities —
- Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents —
- Explainable Suicide Risk Assessment on Social Media with Multi-Task QLoRA —
- Mixture of Decoders for Diverse Dialog Response Generation —
- HAWK: Rethinking Multimodal Drafting for Speculative Decoding —
- Self-Evolving Coding Rules for AI Coding Agents —
- Lingtai: What Concept Geometry Reveals--and Does Not Reveal--About LLM Inference —
- PhysicsMate: A Curriculum-Grounded Bengali Benchmark for Secondary Physics QA with Small-Model Adaptation —
- Analysis of Quantized and Efficiently Adapted Protein Language Models —
- VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision —
- ORBIT-FMIB: Tracking Order-Resolved Epistatic Information Through ESM-2 —
- Closing the Loop: Practical Training Recipes for Looped Language Models —
- Harnessing Vision-Language Models for Perceptual Quality Assessment and Autonomous Content Adjustment in Augmented Reality —
- Bayesian Fine-tuning Yields Language Models that are as Bayesian as their Beliefs Allow —
- Curvature Under Attack in hZACH-ViT: Gauge Symmetry, Boundary Saturation, and Adversarial Failure —
- Grand Canonical Generators —
- SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation —
- Towards Robust Numerical Claim Verification —
- Soundwich: Video Generation with Layered and Controllable Audio —
- FedMAD: Modulation-Aware Directional Aggregation for Federated Learning in Remote Sensing Image Classification —
- How Divergence Becomes Decision Flips in Compressed Language Models —
- Initialization Improves LLM-Driven Discovery —
- Sequential Functional Structured Tucker Compression for Large Language Model Attentions —
- Reason in Style: Discovering and Controlling Style in Language Models —
- Personalized Image Generation with Reasoning and Reflection —
- What Builds the Scene? Luminance Dominates Geometry Formation in 3D Gaussian Splatting —
- Signal-Noise Factorization Isolates Nuisance Variation into Removable Subspaces —
- Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning —
- Learning to Price Electricity for Optimal Demand Response —
- Video Evidence Indexing: Learning Where to Look from Video Previews for Token-Budgeted Long-Video Question Answering —
- Scalable Multi-Task Inverse Reinforcement Learning —
- Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints —
- Effective Synthetic Data Curation Requires Group-Level Signals —
- VTV-FM: Flow Matching through Variational Terminal-Velocity Closure —
- Inference for stochastic differential equations driven by weighted sub-fractional Brownian motion using neural networks and the Euler approximation —
- Can large language models unlock discrete data in ophthalmic diagnostic reports? —
- Sapien: A Stateful Policy Engine for Autonomous AI Agents —
- Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristics —
- Video Generation Models: A Survey of Post-Training and Alignment —
- Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing —
- Verbalized and Internal Probabilities Are Coupled in Large Language Models —
- VERITYGATE: A Four-Gate Schema-Level Faithfulness Framework and Paired Benchmark for Grounded LLM Narrations over Structured Evidence —
- Contextual trajectory and incremental contextual displacement: Towards using LLMs to understand dynamic, utterance-specific meaning construction —
- Geometric Similarity in VLM Low-Level Vision Representations —
- Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning —
- SmoothOperator: Enhancing Representations for Fine-grained Open-set Recognition via Modulated Label Smoothing —
- Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis —
- Lang3DSeg: Annotation-Free Open-Vocabulary 3D Segmentation with Point Transformers —
- CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight —
- Don't Waste the Noise: Importance-Guided Perturbation Allocation under Joint Global and Local Constraints —
- Machine Translation for Sign Languages —
- DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text —
- Towards Fast and Disentangled Counterfactuals for Visual Foundation Models —
- ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization —
- The Geometry of Contextual Relations: Language Models Address Facts by Order of Mention —
- Block Optimism for Nonstationary Bandits with Latent Linear Dynamics —
- EyeTAG: Eye Trajectory-Aware Gaze Estimation —
- Efficient Task Adaptation in Large Language Models: A Survey of Weight-Based, Prompt-Based, and Embedding-Based Adaptations —
- Platonic Task Arithmetic —
- Joint Branch-Space Transform Coding for Diffusion Activation Quantization with Classifier-Free Guidance —
- ReHoPER: Receding-Horizon Planning for Enhanced Reasoning —
- A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions —
- Two Clocks in Diffusion MLLMs: When Answers Stabilize Before Rationales Unfold —
- Beyond Leaderboards: Tokenomics of Agentic Small Language Model Ensembles —
- Role-aware Heuristic Episodic Attention for Conversational LLMs —
- Video-Index: A Curated Meta-Benchmark for Video Understanding —
- A Citation-Grounded Benchmark for Trustworthy Earnings Call Transcript Analysis with Large Language Models —
- RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation —
- Concept Driven Domain Adaptation: Finding an Abstract Needle in a Haystack —
- The Devil Is in the Reconstruction Loss Scale: Rethinking Optimization in LLM Quantization —
- Auditable Algebraic Counting Field for Cryptic-Pocket Detection from Apo Structures —
- The Price of Correlated Tests: How Strict Should a Model Release Gate Be? —
- VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations —
- Distilling Directional Verification —
- Tolerance-Based Fairness Auditing: Violation Certification and Sensitivity Screening —
- Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning —
- VASC: Value-Aware Sparse Attention with Cross-Layer Memory for Efficient 3D Reconstruction —
- Scaling and Distilling Text Embeddings for Better Diffusibility —
- Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization —
- FutureWorlds: Learning Robotic World Models from Alternative Futures —
- Towards Automatic Video Annotation with ASH: Zero-Shot Open-Vocabulary Multi-Object Tracking and Segmentation —
- It Takes Workflows to Evolve Better Workflows —
- LawCompass: Navigating from Legal QA to Multi-Agent Deep Research with Grounded Evidence —
- Posterior sampling by source-space MCMC via prior-based few-step transport maps —
- Bootstrapping Video Interaction Generation with Synthetic State Transitions —
- Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems —
- Empty Commitments: When Agents Promise What They Cannot Deliver —
- Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study —
- Towards Subject Consistency over Dynamic Subject Sets in Video Generation —
- Capturing In-Context Learning Dynamics with Task Operators —
- HierGF: Hierarchical Gaussian Fields via Geometry-perception Message Passing for Sparse-view 3D Reconstruction —
- JoinGR: Learning to Traverse Join Graphs for Table Retrieval —
- Probe with Participation Trophies: Random-Reward RL as a Probe of LLM Capability —
- Overcoming Kernel Redundancy for Scaling Logic Gate Networks —
- Precision over Scale: A Polish-Silesian Benchmark and a Translation System Outperforming Open-Source and Commercial Models —
- Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation —
- Dataset Identity, Not Novelty: The Source of an Inflated OOD Detection Gain —
- MVDG: Efficient Multi-view 3D Disambiguation on Unconstrained Real-World Images —
- AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines —
- Affine-Aligned Atlas for Canonical Gaussian Construction in Video Representation —
- Madeleine: Learning Involuntary Recall for Conversational Memory from Simulated Lives —
- Counting and Min-Cost Encoding for Tokenization in Large Language Models —
- Open Vocabulary Word Recognition From Transcribed Bangla Texts —
- The RSNA Intracranial Aneurysm (RSNA-ICA) Dataset —
- OptimusMesh: Compact Autoregressive Mesh Generation from Point Clouds via Sparse Latent Pivots —
- BanglaDial-Abuse: A Corpus-Grounded Dataset for Regional Dialect Identification in Abusive Bangla Text —
- My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning —
- PhysicsLENS: Diagnosing Physical Property Blindness in Video Generation Models —
- Persistent Depth Ordering amid Shifting Block-Bypass Responses in Language Model Pretraining —
- CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment —
- HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix —
- Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models —
- Skeleton-and-Strategy Prompting: Training-Free Negation Understanding for Vision-Language Models —
- Color Independent Word Segmentation From Transcribed Bangla Passages —
- FlashBack: Knowing When to Remember in Streaming Vision-Language Models —
- Counterfactual Generation via Flow Matching: Coupling-Sensitive End-to-End Rates —
- iSEE: Object Permanence Through Self-Supervision —
- Semantic RGB--Depth Based Surgical Skill Assessment in Microscopic Stereo Videos —
- Resolving Mixed Single-Photon LiDAR Returns for Foreground-View and Hidden Scene Reconstruction —
- EgoFound3R: End-to-End Egocentric Hand Reconstruction in World Space with Point-Wise Interaction Attributes —
- AutoGUIWorld: Image Generators as Visual World Models for GUI Agent —
- AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation —
- A Compact Explicit 4D Representation for Dynamic Scenes —
- Flow Matching Reinforcement for 3D Mesh Generation via Dynamic Homing Optimization —
- ASCRIBE: Atomic and Significance-Based Reasoning for Thai Clinical SOAP Note Generation —
- Harness Annealing: Learning to Act with Less External Control —
- Evaluating the Robustness of Japanese LLMs to IME-Related and Typographical Errors —
- When the Judge Acts: Auditing VLM-Guided Image Selection on Culturally Situated Prompts —
- Right Answers, Wrong States: Hidden Information Failures in Multi-Agent Collaboration —
Important terms
- Auditable Algebraic Counting Fields
- These are mathematical systems being developed to reliably detect hidden structural pockets within protein shapes using algebraic methods, which is vital for designing effective drugs.
- ORBIT-FMIB
- This work uses the ESM-2 model to track and understand the order of epistatic information in protein language models, which helps decipher complex gene interaction sequences.
- UniGuardian
- This is a unified defense system designed to detect various malicious inputs like prompt injections and adversarial attacks in large language models, ensuring their security.
- VIDA dataset
- This new dataset captures the visual ambiguity found in multimodal machine translation, helping researchers train smaller language models for more reliable clinical triage.