AI papers — 2026-10-08
Today's focus is on how we can better understand and classify ADHD using movie fMRI data. This matters because getting a clearer picture of the underlying neural correlates could lead to more targeted interventions. Researchers looked at MovieSTAGE, which aimed to use scene and transition information along with global encoding methods to classify ADHD in subjects based on their brain activity during movie viewing.
A related piece explored Route-Verify-Vote, which is a procedure conditioned self-consistency method designed for mixed-domain reasoning tasks. This is significant because it suggests a way to make complex inferences more reliable when dealing with different types of information simultaneously. Moving down the list, Child ASR Adaptation with Adult Retention was examined, an empirical study that looked at how well child speech recognition models retain adult language patterns.
Then there's Emo-Jev, which uses probabilistic reasoning for emotion classification incorporating Jev, a concept that helps model emotional responses in a more nuanced way. This connects to the work on CoDR, which introduced training-free confidence-drift remasking specifically for diffusion language models. Finally, simultaneous hyperkinetic movement disorders phenotyping was looked at through a cross-cohort pediatric transfer study that used routine videos and markerless pose estimation alongside a tabular foundation model.
The most significant piece of work from yesterday involved using multi-role reinforcement learning to create more faithful plans by incorporating feedback from a solver. This matters because it moves beyond simple policy learning, aiming for plans that actually work in complex environments. The researchers tried training agents with different roles, and the results showed that this approach improved plan fidelity compared to single-role training.
Another important development was the work on learning perturbation robust policies for large language model agents using stable optimization techniques. This is crucial because it makes these AI agents less brittle when faced with slight changes in their input or environment. They found that by applying metric masking, which essentially lets the agent "forget" certain information strategically during adaptation, they could achieve better robustness.
Then there was the development of Tokka-Bench, which evaluates tokenizers across a hundred natural and twenty programming languages. This helps us understand how different tokenizer designs perform when handling diverse code and text inputs. This foundational work feeds into other systems by providing benchmarks for language model capabilities.
Furthermore, there is research into LLM-guided spatio-temporal graph node generation for forecasting unobserved node states, which attempts to predict future states in complex networks. This is useful for tasks where we need to anticipate what happens next in a dynamic system. This idea connects with the work on ideological LLMs for content moderation, as both explore how models can be guided or adapted based on specific structural inputs.
The most significant development from yesterday was the work on Activation-Informed Pareto-Guided Low-Rank Compression, which directly addresses the efficiency bottleneck in large language models and vision language models. This technique attempts to find a compressed representation of model activations by guiding this compression using a Pareto front approach, meaning it seeks the best possible trade-off between compression size and retained information. The results showed that this method can achieve substantial dimensionality reduction while maintaining high performance on downstream tasks, suggesting a pathway toward making these massive models more practical to deploy.
This efficiency gain is supported by research into attention-mass condensation for sparse decoding, which focuses on how to make the process of generating text faster by condensing the attention mechanism. This work explored methods that can reduce the computational load during token generation without sacrificing semantic accuracy, linking it conceptually to the compression efforts mentioned above. Furthermore, there was progress in continuous semantic caching for low-cost LLM serving, which aims to keep frequently accessed information readily available to speed up inference costs.
On a more foundational level, work on epistemic constitutionalism explored how to prevent coherence bias when training models that rely on human feedback. This addresses the inherent problem where models become overly confident in incorrect information because they are trained on flawed examples, and this concept is relevant to ensuring the reliability of any compressed or optimized model. Finally, research into document optimization for black-box retrieval using reinforcement learning showed how to improve how systems search through large document sets by training a model to select the most relevant parts based on desired outcomes.
The most critical work today centered on APCD, which introduces a method for adaptive path contrastive decoding to improve the reliability of large language model generation. This matters because it directly addresses the instability often seen in complex LLM outputs by guiding the generation process more effectively.
We saw that this APCD approach, when applied to generating user personas using beyond cooperative simulators, provided a more robust way to evaluate agent performance than previous methods. This is significant because it moves past simple simulation and offers a better lens for assessing how well agents can embody realistic user types.
Another key development involved latent performance profiling of large language models, which attempts to map internal model states to external behaviors. This work suggests that understanding these latent states could help diagnose why certain models produce specific kinds of outputs.
The study on auto-interpretation labels showed how far these labels generalize when tested across different languages, scripts, and rewordings. This is important because it tests the robustness of the model's internal interpretation capabilities outside of its initial training context.
Then there was WRIT, which synthesizes write-read intensive trajectories for multi-turn user-facing agents. This work provides a framework for tracking the full interaction history needed to properly assess agent performance over time.
Filtered reasoning score evaluation focused on assessing reasoning quality by looking specifically at the model's most confident traces. This method aims to filter out noisy or less certain outputs to get a clearer picture of the actual logical steps taken by the model.
Finally, there was research into rethinking meeting effectiveness using a benchmark and framework for temporal fine-grained automatic evaluation. This is useful because it tries to give structure and measurable criteria to subjective assessments of how well an AI handles real-world communication tasks.
The most significant work today involved the exploration of how discrete diffusion language models handle parallel sampling, which directly impacts their generation quality. Researchers investigated walk fast but be careful, focusing on understanding parallel sampling within masked diffusion processes. This means they were trying to figure out if the way these models sample during generation affects how coherent the resulting text is.
A related effort looked at steering without breaking, focusing on mechanistically informed interventions for discrete diffusion language models. This work sought to understand how to guide these models without causing them to break their underlying structure. It connects with other efforts because understanding this guidance mechanism is key to controlling the model's output behavior in more complex tasks.
Then there was the work on context-grounded reconstruction for biomedical multimodal continued pretraining, which matters because it addresses a major gap in how large models handle specialized data. This involved moving beyond just captions to reconstruct content with better context grounding during continued pretraining. This contrasts with the mixedpeft research, which combined multiple parameter-efficient fine-tuning methods using mixed objectives for unsupervised domain adaptation.
BehaviorBench provided a crucial benchmark for foundation models when applied to behavioral science tasks, helping us measure their actual performance in those domains. This benchmarking effort is important because it sets a standard for evaluating how well these models can perform in real-world, nuanced applications like behavioral science research.
The most critical work today involved testing how to make large language models better at recalling specific, low-density facts within their knowledge bases because if they can’t retrieve the right piece of information reliably, the entire system breaks down. We looked at FTA-Mem, which is a memory technique designed for long-term dialogue that anchors facts with time and affect, aiming to solve issues where models forget details over extended conversations.
This approach builds upon earlier work concerning what reward models actually memorize; specifically, we explored how those memorized patterns relate to the surprisal theory argument suggesting that simple surprise isn't enough without rational grounding. Furthermore, we investigated using LLM-generated explanations as a method for detecting emotionally rewritten fake news, which is important because it helps us spot manipulation by analyzing the tone of generated text against known falsehoods.
Another piece of research focused on LRCC, which is a method for generalizing low-rank compression using conditional computation to improve efficiency in knowledge retrieval. This work connects to the broader theme of improving knowledge base access, showing how structural compression can aid the retrieval process discussed in FTA-Mem.
The most significant work from yesterday involved exploring how to steer large language model agents toward taking actual actions rather than just generating text. This is crucial because moving models from mere prediction to execution is the next big hurdle for practical AI deployment. We looked at a method called From Uncertainty to Action, which seems to involve training LLM agents using feedback loops so they can navigate uncertainty and make decisions in the real world.
Another important piece of research focused on grounding language models in specific knowledge structures, specifically BEACON-SP, an ontology-grounded GraphRAG framework designed for clinical suicide risk assessment. This means the system doesn't just guess; it pulls structured data from a graph to provide more reliable safety evaluations. This contrasts with other work that focuses more on how models learn from direct language feedback through constraint tree exploration, which is a different way of teaching the model to follow rules based on what it says wrong.
Then there was the comparative analysis of Multi-Label Topic Assignment via LLM Distillation, where researchers pitted generative versus discriminative student models against each other to see which one was better at assigning multiple topics simultaneously. This helps us understand how distillation techniques affect the model's ability to handle complex classification tasks. This is related to how we might fine-tune agents, as understanding topic assignment is a prerequisite for complex reasoning tasks like those explored in CARE, which certifies acceleration for vision-language-action inference.
The work on tiny-scale Chinese BERT pretraining is significant because it directly addresses the challenge of adapting large language models to lower-resource languages by testing different pretraining strategies. Researchers compared masked language modeling, word window modeling, and MacBERT approaches on a small dataset, finding that the WWM strategy yielded better performance than MLM alone. This suggests that contextual information within a local window is more valuable for this type of model than simply predicting missing words randomly across the entire sequence.
Steering follow geometry rather than labels is important because it shows how to control emotional directions in full-duplex speech models without relying on explicit emotion labels during training. This technique manipulates the model's internal representation to guide its output toward a desired emotional state based on geometric relationships within the latent space, which is a more robust method than standard supervised fine-tuning.
On KL-regularized policy optimization provides a framework for improving the stability of reinforcement learning policies by penalizing divergence from an initial distribution. This regularization helps ensure that the learned policy does not stray too far from what was initially expected, which is crucial when training complex decision-making systems.
Localizing safety-critical parameters for sparse fault analysis matters because it helps pinpoint exactly where a language model might fail when deployed on devices due to faults. By focusing analysis on specific parameters, this method offers a targeted approach to understanding the fragility of on-device language models.
Quad-state safety evaluation of open-weight large language models is important because it tests how these models handle inputs that fall outside their expected normal range. This evaluation reveals whether a model can maintain predictable safety performance when encountering non-canonical inputs, which is a key concern for real-world deployment.
U-Space uncovers when and why uncertainty appears in language models by analyzing the distribution of predictions across different contexts. This method helps diagnose the specific conditions that cause a model to become uncertain, offering insight into its failure modes.
Same text, different prediction highlights nondeterminism in text classifiers where the serving context can change the outcome unexpectedly. This finding suggests that simply having a fixed text input is not enough; how that text is presented during inference significantly impacts the resulting classification.
Noise your prompt by noising conditioning tokens in continuous diffusion language models to improve robustness against adversarial attacks. This involves adding controlled noise to specific tokens within the prompt, which helps make the continuous diffusion model less susceptible to malicious inputs while still allowing it to generate coherent output.
Today's papers
- MovieSTAGE: Scene, Transition, and Global Encoding for Movie-fMRI ADHD Classification. [paper]
- Route-Verify-Vote: Procedure-Conditioned Self-Consistency for Mixed-Domain Reasoning. [paper]
- Child ASR Adaptation with Adult Retention: An Empirical Study. [paper]
- Emo-Jev: Probabilistic Reasoning for Emotion Classification with Jev. [paper]
- CoDR: Training-Free Confidence-Drift Remasking for Diffusion Language Models. [paper]
- Simultaneous hyperkinetic movement disorders phenotyping: a cross-cohort pediatric transfer study using routine videos, markerless pose estimation and a tabular foundation model. [paper] [episode]
- Do Generative Priors Align with Human Naturalness Perception?. [paper]
- Demystifying Manifold Constraints in LLM Pre-training. [paper] [episode]
- From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning. [paper] [episode]
- Learning Perturbation Robust Policies for LLM Agents with Stable Optimization. [paper] [episode]
- Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages. [paper]
- Just for FUNS: LLM-Guided Spatio-Temporal Graph Node Generation for Forecasting Unobserved Node States. [paper]
- When Forgetting Looks Like Improvement: Metric Masking in Streaming Diarizer Adaptation and the Price of Rehearsal. [paper]
- APE: Selective Fine-tuning with Acceptance Criteria for Language Model Adaptation. [paper] [episode]
- Classification of Spontaneous and Scripted Speech for Multilingual Audio. [paper] [episode]
- Ideology-Based LLMs for Content Moderation. [paper] [episode]
- Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM. [paper] [episode]
- HealthcareNLP: where are we and what is next?. [paper] [episode]
- Epistemic Constitutionalism Or: how to avoid coherence bias. [paper] [episode]
- Attention-Mass Condensation for Sparse Decoding. [paper] [episode]
- Just on Time: Token-Level Early Stopping for Diffusion Language Models. [paper] [episode]
- Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language. [paper] [episode]
- Document Optimization for Black-Box Retrieval via Reinforcement Learning. [paper] [episode]
- Continuous Semantic Caching for Low-Cost LLM Serving. [paper] [episode]
- APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation. [paper] [episode]
- Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents. [paper] [episode]
- Latent Performance Profiling of Large Language Models. [paper] [episode]
- How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings. [paper] [episode]
- WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents. [paper] [episode]
- Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces. [paper] [episode]
- Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation. [paper] [episode]
- Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World Risk. [paper] [episode]
- Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models. [paper] [episode]
- Handle with CARE: Can LLMs Reproduce How Online Communities React?. [paper] [episode]
- Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining. [paper] [episode]
- MixedPEFT: Combining Multiple PEFT Methods with Mixed Objectives for Unsupervised Domain Adaptation. [paper] [episode]
- BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks. [paper] [episode]
- Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator. [paper] [episode]
- Who Brought Easter Eggs to Eid? Auditing LLM-Generated Cultural Translation of Math Word Problems Across Languages and Regions. [paper] [episode]
- Walk fast but be careful: Understanding Parallel Sampling in Masked Diffusion. [paper] [episode]
- Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation. [paper] [episode]
- When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs. [paper] [episode]
- What do Reward Models Memorize?. [paper] [episode]
- FTA-Mem: Fact-Time-Affect Anchored Memory for Low-Density Long-Term Dialogue. [paper] [episode]
- Surprisal Theory is Tautological (without Rational Grounding). [paper] [episode]
- Leveraging LLM-Generated Explanations for Detecting Emotionally Rewritten Fake News. [paper]
- Beyond Risk Prediction: Evidence Grounding and Psychosocial Factor Verification for Explainable Suicide Risk Assessment. [paper]
- LRCC: Generalizing Low-Rank Compression with Conditional Computation. [paper]
- FinVector-Market-4B: A Controlled Study of LoRA Adaptation for Structured Financial Tasks. [paper]
- CARE: Certifying Acceleration for Vision-Language-Action Inference. [paper]
- BEACON-SP: Ontology-Grounded GraphRAG Framework for Clinical Suicide Risk Assessment. [paper]
- Multi-Label Topic Assignment via LLM Distillation: A Comparative Analysis of Generative vs. Discriminative Student Models. [paper]
- Constraint Tree Exploration for Learning from Language Feedback. [paper]
- From Uncertainty to Action: Learning to Steer LLM Agents. [paper]
- Beyond the Sycophancy Score: How Task, Model, and Pressure Shape LLM Yielding. [paper]
- QuanLing: Cross-Branch Validation of Language Distance Quantification on Western Romance. [paper]
- Tiny-Scale Chinese BERT Pretraining: A Controlled Comparison of MLM, WWM, and MacBERT Strategies. [paper]
- Steering Follows Geometry, Not Labels: Emotion Directions in a Full-Duplex Speech Model. [paper]
- On KL-Regularized Policy Optimization. [paper]
- How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis. [paper]
The papers
- MAGIC: Learning from Visibility Asymmetry for Unsupervised Stereo Matching — Unsupervised stereo matching methods often fail to provide accurate supervision in occluded regions, leading to poor disparity estimates. [episode]
- MotionHiFlow: Text-to-motion via hierarchical flow matching — MotionHiFlow proposes a hierarchical flow matching framework to generate 3D human motions progressively from low to high temporal scales, addressing the limitations of single-scale methods by capturing high-level semantics first and then refining details. [episode]
- Bi-CamoDiffusion: A Boundary-informed Diffusion Approach for Camouflaged Object Detection — Bi-CamoDiffusion introduces an evolution of CamoDiffusion by integrating edge priors into early-stage embeddings via a parameter-free injection process, governed by a unified optimization objective that balances spatial accuracy, structural constraints, and uncertainty supervisio [episode]
- Hand-4DGS: Feed-Forward 3D Gaussian Splatting for 4D Hand Reconstruction from Egocentric Videos — Dynamic 3D hand reconstruction from egocentric videos is essential for next-generation computing platforms such as AR/VR and AI glasses, and this paper introduces Hand-4DGS, the first feed-forward framework for reconstructing dynamic 4D hands directly from egocentric videos, enab [episode]
- Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator — Identifying faithfulness hallucinations in LLM-generated outputs remains challenging due to the scarcity of high-quality annotated data, and this work introduces Hallucination SelfPlay (HSP), a novel framework that enables a detector to bootstrap with an evolved generator. [episode]
- Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models — As a meticulous researcher, I have thoroughly analyzed both provided texts from the arXiv paper "Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models." The information is highly technical, focusing on a novel method for controll [episode]
- Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models — Residualized temporal SAEs provide a useful framework for studying temporally structured diffusion activations by separating linearly predictable components from residual components, allowing sparse latents to capture structure beyond what is linearly predictable. [episode]
- Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models — Large vision-language models (VLMs) are increasingly integrated into healthcare, but their tendency to generate clinically plausible yet incorrect statements raises significant safety concerns. [episode]
- Order-Optimal Sample Complexity of Rectified Flows — Rectified flow models achieve an order-optimal sample complexity of Oe(ε−2) for approximating target distributions in Wasserstein distance, surpassing existing bounds for diffusion models and general flow matching. [episode]
- How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings — Sparse autoencoder (SAE) features are increasingly used to interpret language models, with auto-generated natural-language labels serving as the primary interface for understanding what each feature represents. [episode]
- Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs — This paper investigates a critical failure mode in Video Large Language Models (Video-LLMs): directional motion blindness, which is formally diagnosed as a "direction binding gap." This gap signifies that while the necessary information for motion direction is linearly accessible [episode]
- Learning Perturbation Robust Policies for LLM Agents with Stable Optimization — Reinforcement learning policies trained for long-horizon large language model agents are sensitive to various policy perturbations, such as hidden state noise, pruning, and quantization. [episode]
- MAdam: Metric-Aware Multi-Objective Adam — Multi-objective optimization (MOO) solvers almost universally hand their reconciled directions to Adam, but this coupling introduces two systematic gaps: a weighting mismatch where Adam marginalizes time-varying preferences into a history average, and a geometric mismatch where A [episode]
- LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting — Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurate estimation of intrinsic scene properties such as materials, geometry, and illumination. [episode]
- Geometry-Centered 3D Latent World Models for Growing Surfaces — Physical intelligence—anticipating and shaping the world from partial, multisensory observations—is critical for next-generation world models. [episode]
- StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement — As a fastidious and diligent researcher, I have meticulously reviewed both provided summaries of "STRESSDREAM" from arXiv. [episode]
- Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining — Biomedical figures are explained not by captions alone but by body-text passages that discuss them, and this paper introduces context-grounded reconstruction, a source-grounded framework that converts PubMed Central Open Access (PMC-OA) records into referentially coherent interle [episode]
- Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World Risk — Frontier image generation has moved from artistic synthesis toward synthetic visual evidence, creating significant risks for society because systems now produce artifacts that mimic reliable records. [episode]
- Continuous Semantic Caching for Low-Cost LLM Serving — Continuous Semantic Caching for Low-Cost LLM Serving establishes a rigorous theoretical framework for caching responses in an infinite, continuous query space under uncertainty, bridging discrete optimization with continuous representation spaces. [episode]
- SpanVLA: Learning from Negative-Recovery Samples with Fast Action Bridging for Vision-Language-Action Model — SpanVLA introduces a novel end-to-end autonomous driving framework that integrates an efficient action bridge and learns from real-world negative-recovery samples to enhance performance and robustness. [episode]
- Document Optimization for Black-Box Retrieval via Reinforcement Learning — Document expansion is recast as a document optimization problem where an instruction-tuned language model or vision language model is fine-tuned to transform documents into representations that better align with the expected query distribution under a target retriever, using GRPO [episode]
- Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion — Tri-Prompting introduces a unified video diffusion framework that integrates scene composition, multi-view subject consistency, and disentangled motion control within a single model. [episode]
- Prototype-Based Knowledge Guidance for Fine-Grained Structured Radiology Reporting — Structured radiology reporting promises faster, more consistent communication than free text, but automation remains difficult as models must make many fine-grained, discrete decisions about rare findings and attributes from limited structured supervision. [episode]
- Epistemic Constitutionalism Or: how to avoid coherence bias — Large language models increasingly function as artificial reasoners, and this paper argues for an epistemic constitution for AI—explicit, contestable meta-norms that regulate how systems form and express beliefs. [episode]
- Multi-encoder ConvNeXt Network with Smooth Attentional Feature Fusion for Multispectral Semantic Segmentation — MeCSAFNet is a dual-branch encoder-decoder architecture designed for land cover segmentation in multispectral imagery, leveraging separate processing streams for visible and non-visible channels to achieve superior performance compared to models that process all spectral bands co [episode]
- ReToken: Improving Long-Context VLMs with Visual Retrieval Token — Long visual contexts pose a challenge for vision-language models (VLMs) because performance degrades as distractors grow, and processing all tokens at once becomes computationally infeasible. [episode]
- BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks — Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics, but there remains no systematic understanding of how well they perform across diverse behavioral science tasks, contexts, and populations. [episode]
- U-CECE: A Universal Multi-Resolution Framework for Conceptual Counterfactual Explanations — The paper introduces U-CECE (Universal Multi-Resolution Framework for Conceptual Counterfactual Explanations), a unified, model-agnostic framework designed to generate conceptual counterfactual explanations by adapting its complexity and expressivity to the specific data regime a [episode]
- MixedPEFT: Combining Multiple PEFT Methods with Mixed Objectives for Unsupervised Domain Adaptation — This study presents a novel parameter-efficient strategy for unsupervised domain adaptation that combines custom PEFT architectures with mixed-objective training to simultaneously optimize classification performance on labeled source data and masked language modeling (MLM) on unl [episode]
- Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation — LLM agents increasingly answer questions against structured knowledge bases that they themselves help maintain, and this study tests whether restructuring these knowledge bases for progressive disclosure changes answer quality and cost across different agent access regimes. [episode]
- Ideology-Based LLMs for Content Moderation — Persona conditioning introduces subtle ideological biases into LLM outputs, raising concerns about AI systems that may reinforce partisan perspectives under the guise of neutrality. [episode]
- Demystifying Manifold Constraints in LLM Pre-training — Manifold constraints in LLM pre-training are studied to understand how restricting weights to specific geometric spaces shapes training dynamics, revealing functional overlaps with existing stabilization mechanisms and providing principled alternatives to weight decay. [episode]
- Attention-Mass Condensation for Sparse Decoding — Attention sparsity is a learned property of trained transformers, and this paper demonstrates that attention mass concentrates on a small, identifiable subset of positions which can be dynamically selected to achieve numerically exact equivalence with full attention at O(n) compl [episode]
- Faster-WAM: Do World Action Models Need Deep Action Modules? — World Action Models (WAMs) couple robot action prediction with video world models, and this work introduces Faster-WAM, an instantiation of Dock of Transformer (DoT), which enables lightweight task-specific heads to access representations distributed throughout a central video ba [episode]
- FiRe: Fine-grained Multimodal Reasoning for Enhanced Image Generation — As an excellent, fastidious, and diligent researcher, I have meticulously analyzed the provided text snippets from arXiv regarding "FiRe." Given that no full paper content was supplied in your prompt (only a set of quotes and a reward structure), my analysis must be strictly limi [episode]
- MambaDSF: Multi-Scale SSM with Dilated Feature Fusion for Sonar Small Target Detection — Sonar imaging presents significant challenges for detecting small targets due to insufficient pixel coverage, low acoustic contrast, and scale ambiguity across imaging ranges. [episode]
- Component-Adaptive and Lesion-Level Supervision for Improved Small Structure Segmentation in Brain MRI — Small lesions in brain MRI are sparse, spatially localized, and clinically important targets embedded within a large volume of normal-appearing tissue, leading to an extreme imbalance in voxel distribution where standard learning processes often overlook them. [episode]
- Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks — Open-weight LLM fine-tuning defenses are susceptible to simple attacks, revealing that existing safeguards may fail to eliminate harmful knowledge embedded in pretrained models. [episode]
- The Metagame of Interpretability and Meta-Attributions — We introduce METAGAME, a conceptual framework for quantifying second-order interaction effects of model explanations, which provides a principled method to decompose any first-order attribution into directional interaction terms. [episode]
- Just on Time: Token-Level Early Stopping for Diffusion Language Models — Diffusion language models generate text through iterative refinement, a process that is often computationally inefficient because many tokens reach stability long before the final denoising step. [episode]
- Walk fast but be careful: Understanding Parallel Sampling in Masked Diffusion — In this paper, graph random walks are proposed as a verifiable sandbox to study different parallel sampling strategies in masked diffusion models (MDMs), demonstrating that performance critically depends on the underlying graph structure rather than just local uncertainty scores. [episode]
- StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics — As a fastidious and diligent AI researcher, I have thoroughly analyzed both provided texts (A and B). [episode]
- When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs — Large language models (LLMs) perform strongly on academic benchmarks but show substantial weaknesses when tested on everyday, culturally grounded knowledge. [episode]
- What do Reward Models Memorize? — Discriminative training of reward models (RMs) from human preference data results in biased RMs not yet capable of judging response quality in context-dependent scenarios. [episode]
- BiPO: Bidirectional Partial Occlusion Network for Text-to-Motion Synthesis — Generating natural and expressive human motions from textual descriptions is challenging due to the complexity of coordinating full-body dynamics and capturing nuanced motion patterns over extended sequences that accurately reflect the given text. [episode]
- Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language — The Pashto Common Voice corpus represents the first large-scale, openly licensed speech resource for Pashto, addressing its absence in open speech technology despite having over 60 million native speakers. [episode]
- Causal Posterior Estimation — Causal Posterior Estimation (CPE) introduces a novel simulation-based inference method that enhances posterior distribution approximation by explicitly incorporating the conditional dependence structure from both prior and posterior programs. [episode]
- Open-CHOIR: Open-World Contact-Aware 4D Hand-Object Interaction Reconstruction — Reconstructing 4D hand-object interactions (HOI) from monocular RGB videos in open-world settings remains difficult because existing methods often fail under clutter, occlusion, and unseen object geometries. [episode]
- Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces — Given that correctness alone does not reveal reasoning quality, this paper introduces the Filtered Reasoning Score (FRS), a metric that evaluates reasoning quality specifically on a model’s most-confident outputs to uncover structures hidden by standard accuracy evaluations. [episode]
- Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents — Large Language Model (LLM) agents require evaluation environments that replicate real-world user friction, and this paper introduces Persona Policies (PPol), an evolutionary program search framework designed to generate diverse, human-like user personas for robust agent evaluatio [episode]
- VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification — VideoZeroBench introduces a challenging, hierarchical benchmark designed to rigorously verify whether video multimodal large language models can accurately answer questions while simultaneously identifying and localizing the precise spatio-temporal evidence supporting those answe [episode]
- Modified Loss of Momentum Gradient Descent: Fine-Grained Analysis — We analyze gradient descent with Polyak heavy-ball momentum (HB) to prove that it is exactly plain gradient descent with a modified loss on an exponentially attractive invariant manifold, providing rigorous approximation bounds and deriving continuous modified equations. [episode]
- Dual-Pathway Circuits of Object Hallucination in Vision-Language Models — Vision-language models often produce unreliable outputs through object hallucinations, and this study proposes a framework to mechanistically understand these errors by identifying distinct computational circuits within diverse VLMs. [episode]
- Allocation Stability and Wald Inference under Variance-Aware UCB — Precise asymptotic characterization and refined regret analysis for Variance-Aware Upper Confidence Bound (UCB) algorithms provides deep insights into how incorporating variance estimates affects decision-making stability and performance guarantees in Multi-Armed Bandit problems. [episode]
- From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning — Reliable planning requires converting natural-language instructions into executable symbolic specifications, yet large language models remain brittle without costly PDDL annotations and may exploit solver success in semantically unfaithful ways. [episode]
- Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM — Large language models (LLM) and vision-language models (VLM) present significant memory and computing challenges for deployment, necessitating efficient compression techniques. [episode]
- Empirical Evidence for Simply Connected Decision Regions in Image Classifiers — Decision regions learned by deep neural networks are central to understanding their inner workings, and this study provides empirical evidence that these regions are simply connected. [episode]
- WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents — Multi-turn user-facing agents require synthesizing complex training trajectories that capture both long-horizon execution and evidence-grounded decision making. [episode]
- Simultaneous hyperkinetic movement disorders phenotyping: a cross-cohort pediatric transfer study using routine videos, markerless pose estimation and a tabular foundation model — Accurate recognition of movement disorders (MDs) phenomenology remains a demanding task in clinical neurology, especially in pediatric practice where mixed and evolving motor presentations complicate diagnosis. [episode]
- APE: Selective Fine-tuning with Acceptance Criteria for Language Model Adaptation — Adjacent Possible Exploration (APE) presents a selective fine-tuning method for adapting large language models by systematically exploring parameter modifications while maintaining model stability. [episode]
- Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation — Evaluating meeting effectiveness is crucial for improving organizational productivity, and current post-hoc survey approaches are limited by scalability and fail to capture the dynamic nature of collaborative discussions. [episode]
- Attention at Rest Stays at Rest: Breaking Visual Inertia to Mitigate Relation Hallucinations — Visual attention in multimodal large language models (MLLMs) exhibits pronounced inertia, remaining largely static once settled during early decoding steps and failing to support the compositional understanding required for cognitive inference. [episode]
- HealthcareNLP: where are we and what is next? — This tutorial provides an overview of current achievements and future challenges in Healthcare Natural Language Processing (HealthcareNLP) by structuring the field into three hierarchical layers: data/resource, NLP-Eval, and patients. [episode]
- Auto-Regressive Models Need Structural Registers: Semantic Foresight Improves Coherent Driving Video Continuation — ARCON introduces a scheme that alternates between generating semantic and RGB tokens to explicitly learn high-level structural video information, which significantly improves temporal consistency and physical reasonableness when autoregressively generating long videos in autonomo [episode]
- FTA-Mem: Fact-Time-Affect Anchored Memory for Low-Density Long-Term Dialogue — Long-term emotional-support agents require memory mechanisms for personalized understanding across sessions, and this paper proposes FTA-Mem, a structured memory framework that uses situation-level Fact-Time-Affect memory units to address the challenges of low-density dialogue. [episode]
- Latent Performance Profiling of Large Language Models — Large language models frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabilities. [episode]
- Classification of Spontaneous and Scripted Speech for Multilingual Audio — Distinguishing between scripted and spontaneous speech is an essential tool for better understanding how speech styles influence speech processing research, and this paper addresses this challenge by systematically evaluating models across various formats and languages. [episode]
- HDR Video Generation via Latent Alignment with Logarithmic Encoding — HDR generation from SDR input is challenging for generative models because HDR data's linear space and heavy-tailed distributions mismatch their training data, but this work demonstrates that high-quality HDR video can be achieved by leveraging pretrained video models through a s [episode]
- Who Brought Easter Eggs to Eid? Auditing LLM-Generated Cultural Translation of Math Word Problems Across Languages and Regions — Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether those adaptations are consistent across models, preserve cultural diversity at scale, and reveal which cultural entities models trea [episode]
- BONSAI: Bayesian Optimization with Natural Simplicity and Interpretability — Bayesian optimization (BO) is a popular technique for sample-efficient optimization of black-box functions, and BONSAI introduces a default-aware BO policy that prunes low-impact deviations from a default configuration while explicitly controlling the loss in acquisition value. [episode]
- Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities — This paper investigates how popular text-to-image (T2I) models, DALL-E 3 and Gemini 3 Pro Preview, depict people from 206 nationalities when prompted to generate images of individuals engaging in common everyday activities. [episode]
- A Variational Latent-Space Framework for Uncertainty-Aware Spectral Image Emulation — Synthetic hyperspectral image (HSI) generation remains essential for large-scale simulation, algorithm development, and mission design, yet traditional radiative transfer models are computationally expensive. [episode]
- Evaluating Time Series Foundation Models for Electricity Price Forecasting: Contamination Risk, Distributional Shifts, and Covariate Dependence — Time series foundation models (TSFMs) have shown strong zero-shot forecasting performance, but their generalization in covariate-driven, nonstationary settings is underexplored. [episode]
- Toward Realistic Remote Sensing Dataset Distillation with Discriminative Prototype-guided Diffusion — Dataset distillation into remote sensing image interpretation is addressed by proposing a novel framework that synthesizes compact, representative synthetic samples to reduce storage and computational costs while maintaining high semantic fidelity and discriminative quality. [episode]
- Efficient Dense Crowd Trajectory Prediction Via Dynamic Clustering — Efficient dense crowd trajectory prediction in high-risk environments like transportation hubs requires methods that can handle the challenges of massiveness, noisiness, and inaccuracy inherent in dense crowds. [episode]
- Estimating Model-Level Membership Inference Vulnerability Without Reference Models — Membership inference attacks (MIAs) are standard tools for evaluating AI model privacy risks, but current state-of-the-art attacks require computationally expensive reference models, limiting their practicality. [episode]
- HomID: Benchmarking Intrinsic Dimension Estimators on Homogenous Manifolds with Anisotropic Embeddings — Machine learning models rely on data lying on low-dimensional manifolds, and this study introduces a Quantum-Inspired Intrinsic-dimension Estimation (QuIIEst) benchmark to rigorously test existing methods against complex, topologically non-trivial manifolds. [episode]
- A Parameter-efficient Convolutional Approach for Camouflaged Weed Detection in Multispectral Aerial Imagery — FCBNet introduces an efficient model for weed segmentation that leverages a fully-frozen ConvNeXt backbone combined with Feature Correction Blocks (FCBs) to achieve high accuracy while drastically reducing computational requirements. [episode]
- The Utility and Complexity of in- and out-of-Distribution Machine Unlearning — Machine unlearning, defined as selectively removing data from trained models, is crucial for addressing privacy concerns and knowledge gaps postdeployment. [episode]
- Handle with CARE: Can LLMs Reproduce How Online Communities React? — Large language models (LLMs) are increasingly used to proxy computational social analysis, but they struggle to faithfully represent the dynamic, event-contingent linguistic behaviors of online communities. [episode]
- EscherNet++: Simultaneous Amodal Completion and Scalable View Synthesis through Masked Fine-Tuning and Enhanced Feed-Forward 3D Reconstruction — EscherNet++ proposes a unified diffusion model that simultaneously performs amodal completion and novel view synthesis, addressing limitations in existing methods by integrating input-level and feature-level masking for robustness and enabling seamless integration with fast feed- [episode]
- Environmental Change Detection for Real-World Change Analysis — Environmental Change Detection (ECD) addresses the limitations of conventional Scene Change Detection (SCD) by moving from idealized settings to a practical task that accounts for real-world data scarcity and viewpoint misalignment. [episode]
- Semantics-Aware Hierarchical Consensus Learning for Remote Sensing Image Classification — Deep learning has become increasingly important in remote sensing image classification due to its ability to extract semantic information from complex data, and this paper presents a novel Semantics-Aware Hierarchical Consensus (SAHC) approach that integrates hierarchical-level-s [episode]
- Multiparameter Uncertainty Mapping in Quantitative Molecular MRI using a Physics-Structured Variational Autoencoder (PS-VAE) — Quantitative imaging methods, such as magnetic resonance fingerprinting (MRF), aim to extract interpretable pathology biomarkers by estimating biophysical tissue parameters from signal evolutions. [episode]
- Smart-Insertion-V: Photorealistic Video Insertion via a Closed-Loop Feedback Dual-Stream Framework — Mask-free video object insertion has emerged as a challenging task requiring harmonious integration of reference objects into source videos, especially when those references exhibit severe stylistic domain gaps with the source scene. [episode]
- Reduction of Class Activation Uncertainty with Background Information — Multitask learning and transfer learning are powerful techniques for improving generalization in deep learning, but they often require significant computational resources. [episode]
- APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation — Large language models often suffer from hallucinations due to error accumulation in autoregressive decoding, where suboptimal early token choices misguide subsequent generation. [episode]
- Synthetic Benchmarks Overstate Forward-Forward Scaling: Real-Data Limits of Layer-Local Training — Forward-Forward (FF) learning replaces backpropagation with strictly layer-local goodness updates, and this paper rigorously audits whether layer-local training is a viable alternative to full backpropagation at realistic scales by developing DTG-FF as an instrument. [episode]
- FisheyeDistanceNet: Self-Supervised Scale-Aware Distance Estimation using Monocular Fisheye Camera for Autonomous Driving — FisheyeDistanceNet presents a novel self-supervised framework designed to learn Euclidean distance and ego-motion from raw, unrectified monocular fisheye video sequences for automotive applications. [episode]
- A prism hierarchy of learning regimes in large linear autoencoders — Theoretical studies of machine learning models commonly consider different limiting regimes in which the learning dynamics of gradient descent becomes theoretically tractable, but this work proposes a systematic picture for large weight-tied linear autoencoders characterized by i [episode]
- Batch Augmentation with Unimodal Fine-tuning for Multimodal Fusion of Large Language Models — This research proposes batch augmentation combined with unimodal fine-tuning to improve multimodal learning for detecting fetal organs from ultrasound images and associated clinical textual information, achieving near state-of-the-art performance on datasets like UPMC Food-101. [episode]
- The Role of Initialization in 3D Gaussian Splatting — 3D Gaussian Splatting (3DGS) has become a method of choice for photo-realistic novel view synthesis due to its efficiency and compelling visual quality, and this work systematically studies how initialization affects 3DGS performance and geometric quality across various densifica [episode]
- A Lightweight Vision-Language Fusion Framework for Predicting App Ratings from User Interfaces and Metadata — App ratings are among the most significant indicators of mobile application quality, and this study proposes a lightweight vision–language framework that jointly leverages mobile UI visuals and semantic metadata to predict app ratings. [episode]
- Surprisal Theory is Tautological (without Rational Grounding) — Surprisal theory is critiqued for being tautological without rational grounding, suggesting that its falsifiability requires restricting the language model to one grounded in non-empirically motivated cognitive principles. [episode]
- MovieSTAGE: Scene, Transition, and Global Encoding for Movie-fMRI ADHD Classification —
- Why VLMs Miss Small Objects, and When Zooming In Is Safe —
- SAREO-FM: Decoupled Semantic Supervision for SAR-EO Foundation Models —
- Dialect-Robust Speech Language Models with Synthetic Pseudo-Dialect Augmentation —
- VIS-Ground: Video Interactive Storytelling with Contextual Grounding —
- Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation —
- GRC-Net: Global Representation Consistency Network for Unsupervised Multimodal Anomaly Detection —
- CRT-HMAR: Causal Requirement Tracing-Guided Hierarchical Multi-Agent Regulation for Open-Task-Aware Infrared-Visible Image Fusion —
- Benign Overfitting under Heterogeneous Input Fusion —
- TileSkipper: Region-Adaptive Tile Pruning for 3D Gaussian Splatting —
- Do Image Editors Follow Depth-Dependent Blur and Aperture Response? A Rendered-Ground-Truth Pilot Audit —
- OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models —
- Multimodal LLMs Can Learn to Read Brain Signals: A Vision--Language Model for Unified Multi-Task EEG Decoding —
- Global Exponential Convergence of Two-Layer Linear Network Training —
- TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning —
- Closing the Loop on Contrail Avoidance with Satellite Verification —
- Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling —
- Unified Multi-plane Autoregressive Diffusion for 3D Multi-contrast MRI Synthesis —
- PPCAR-Net: Projection-Refined Parametric 3D Coronary Artery Reconstruction from Sparse X-ray Angiographic Views —
- The Persona Hierarchy Model: Understanding Contextual Generalization in Fine-Tuning LLMs —
- Quantifying Volumetric Risk: Class-Aware Asymmetric Weighted Conformal Prediction for 3D Medical Image Segmentation —
- ARCS: Towards Precise Text-to-SQL via Structured Disambiguation —
- One Frame, Full Heartbeat: ECG-Free Cardiac Cine MRI Synthesis via Phase-Conditioned Flow Matching —
- TiTok: Audio-Visual LLM for Multi-Segment Temporal Grounding —
- Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data —
- trACT: temporal revelation Airborne Camera Trap —
- Spatial Latent Reasoning for Embodied Reference Understanding —
- What a Reporting Convention Hides: A Matched-Budget Audit of Quantum Natural Gradient with an Exactly Computed Metric —
- Event-Aligned Visual Action Reasoning for World Action Models —
- SkillCycle: Co-Evolving Agent Policies and Skill Banks —
- Controllable Crowd Generation through World-Model Planning —
- InscriptionOCR: A Dataset and Method for Understanding Inscriptions —
- Mixture of Layers: Dynamic Layer Routing for Visual Reasoning —
- TIRA: Tumor Immune Representation Adaptation for Zero-Shot Cross-Cancer MSI and TMB Prediction —
- LighTROcc: Lightweight 4D Occupancy Forecasting via Instance-Centric 3D Gaussians —
- Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science —
- From Global Alignment to Local Grounding: Zero-Shot Chinese Character Recognition with Radical Verification —
- Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning —
- RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning —
- DSReg: Provably Recovering Individual World Latents without Reconstruction —
- Right Number, Wrong State? Measuring Cross-Jurisdiction Substitution in LLM Recall of State Policy —
- An Invariant Tangent-Angle Descriptor and a Band U-Net for 2D Fragment Adjacency Prediction —
- BanglaRhet: Benchmarking Classical and Transformer Models for Rhetorical and Persuasion Detection in Bangla Political Speech —
- LLM-Enabled UAV Dispatch: A System-Level Survey and Taxonomy —
- CHASE: Channel-Aligned Structure Exploitation for Geometry-Aware Model Engineering —
- InstanceBench: Diagnosing Referential Reasoning and Target Identity in Referring Expression Segmentation —
- Goldsmith: Gold-Loss-Guided Definition Optimization with an Agentic Annotation Harness —
- Gaussian Material Fields for Volumetric Multi-Energy CT Decomposition —
- SpatialUQ: Post-Hoc Uncertainty Quantification from Spatial Consistency in Black-Box Vision Models —
- RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection —
- Adjoint-Based Calibration and Optimal Control of Stochastic Multiscale Bioprocess Digital Twins —
- TERRA: Learning Transportable Latent Actions through Temporal Effect Representation and Relational Alignment —
- LiG-DETR: Local-in-Global Reassembly in Latent Space for Aerial Object Detection —
- OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning —
- STRIKE: Learning Visual State Transitions for Physical World Modeling —
- It Is Not Seeing the Hazard: A Frozen Vision-Language Safety Score Measures Its Caption Bank —
- ActiveLang: Active Open-Vocabulary 3D Mapping with Semantic-Uncertainty-Guided Exploration —
- Reflected Anchored Langevin Algorithms —
- A Comparative Study of Evaluation Metrics for Long-Document Financial Narrative Summarization with Transformers —
- Unpaired Canonical Correlation Analysis —
- KASALv2: Fully Automatic 3D Rotational Symmetry Classification and Axis Localization —
- WAPR: A Foundation Model for Wide-Angle Refinement in Unseen Object Pose Estimation —
- Certified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration Budgets —
- Visual Evidence Under Cross-Examination: Evaluating and Controlling Decision-Level Evidence Use in Vision-Language Models —
- RELATE: An Evaluation Framework for measuring Relational Orientation of Large Language Models —
- How Do LLMs Change Predictions Under Negation? —
- Gradient-Based Trajectory Optimisation over Continuous Poses for Sparse-View Cone-Beam CT —
- Collaborative Reasoning Distillation via Cross-Feedback and Coherent Curation —
- Learning Situation-Conditioned Thinking Policies for Long-Term LLM Agents —
- Scalable Logistic Gaussian Process Density Regression with Kinetic Langevin Sampling —
- STORK: Spatio-Temporal Observation of uterine contRactions via neural networKs —
- Which Language Should a Skeleton Speak? Language Choices in Multilingual Reasoning —
- Identity-Duplication Auditing in National-Scale Neuroimaging Repositories —
- Coding-Agent Benchmarks Should Match Their Users' Task Flows —
- On-Policy Distillation Teaches New Skills but Not New Knowledge —
- When Rank Rises as LLMs Degrade —
- Closed-Form Noise Calibration Against Membership Inference for Random-Allocation DP-SGD —
- MeshSIPP: Efficient Lattice Planning in Dynamic Environment —
- Rubric Spans are Label Representations: Joint LLM Encoding for Short Answer Scoring —
- Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring —
- SAPD: Step-Aligned Privileged Distillation —
- Quasi-Binarized Autoencoders: An Architecture-Independent Information Bottleneck for Medical Image Anomaly Detection —
- InsClaimBench: Benchmarking Insurance Claim Adjudication Across the Decision Chain —
- Gauss-Newton Accuracy and Indefinite Hessians: Uniform Coexistence in Low-Cost Sets —
- A Multi-Source Ultrasound Benchmark Revealing the Limits of Contemporary Self-Supervised Anomaly Detection Methods —
- From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery —
- What Makes Synthetic Hard Negatives Work in Vision-Language Pretraining? —
- Latent Watermarks under Generative Editing: A Benchmark and Analysis of Detection Survival —
- Enhancing Multi-Region Stylization with Interior-Guided Boundary Repair —
- SpikingVLA: Asynchronous Spiking Vision-Language-Action Models —
- DynStream: Online Streaming 4D Gaussian Reconstruction of Dynamic Worlds from Unposed Video —
- MeshCarve: Artisan Mesh Generation with Flow Matching in Compact Latent Spaces —
- Beyond Group Splits: Specimen-Level Cross-Validation and Visual Attribution for Remaining-Shelf-Life Regression in Climacteric Fruit —
- Bridge Routing Heads: Where Multilingual Multi-hop Reasoning Lives in LLMs —
- SoccerNet-FoulRet: Retrieving Semantically Similar Soccer Foul Videos —
- Flow-of-Thought: A Framework for Visual Reasoning —
- Diffusion-Generated Image Watermarking: A Two-Axis Taxonomy and Three Protocol-Bounded Case Studies —
- Shaer: Controlled Arabic Poetry Generation with Meter Subform and Semantic Conditioning —
- EntroPrefill: Renyi-Guided Context Pruning with Conditional Stability Guarantees for Retrieval-Augmented Generation —
- Leaner Transformers Can Easily Learn to Cluster —
- PARC-Loc: Text-to-Point-Cloud Localization with Partial Assignment and Relational Consistency —
- Beyond Policy Support: Interaction Constrained Offline Reinforcement Learning for Autonomous Driving —
- Fluctuations of Nonlinear Observables in Mean Field Neural Network Training —
- From Expert-Guided Proof Search to Automated Open-Problem Solving —
- Decoupling Logic from Persona: Structural Immunity of Edge LLM Agents to Context Pollution —
- Relational Abstractions for Spatial Reasoning with Diffusion Models —
- AdaPS-LiNGAM: Adaptive Predecessor Selection for Linear Non-Gaussian Acyclic Models under Small-Sample Settings —
- CIRSeg: Coarse-to-Fine Intensity-Robust Liver Segmentation with Source-Free Continual Test-Time Adaptation —
- UltraWorld: Learning Interactive Ultrasound World Models from Untracked Clinical Videos with Acoustic Sampling Map —
- Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression —
- Counterfactual Route Optimization for Gaussian Head Avatar Modeling —
- MIRROR: From Imitation to Internalization in LLM Personalization —
- Beyond Masks and Trajectories: Flow-Guided Latent Action Injection for Stable Surgical Video Generation —
- DisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation Models —
- Concentration, Not Uncertainty: Why Targeted Synthetic Data Doesn't Help Camouflaged Object Detection —
- Efficient 3D Gaussian Head Avatars for Edge Devices —
- UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation —
- MOTIF: Person-of-Interest Deepfake Detection Beyond 3DMM Coefficients —
- A Deafening Silence: Catastrophic Forgetting Lives in the Output Embeddings of Tokens the Data Never Speaks —
- ORCA: Hunting Compositional Failures in Text-to-Image Diffusion —
- For Those Who Believe in Faithfulness: Optimizing the Area Under Insertion and Deletion Curves for Ranking Relative Feature Importance —
- Hard, Yet Reducible: Controlled Forward Transfer for Synthetic Degradation Curation —
- DeltaSplat: Iterative Gaussian Refinement for Pose-Free Feed-Forward 3D Gaussian Splatting —
- Training Advisors for LLM Agents from Task Outcomes —
- DeepTopoClustering: Unsupervised Derivation of Surface Process Taxonomy from 4D Point Clouds for Topographic Monitoring —
- Global Average Precision for Representation Learning —
- LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets —
- Bringing BNNs to Fast Event Processing —
- SANet: Selective Attention Network for Infrared Small Target Detection —
- Sparsifying Stochasticity, Not Capacity: Partial Stochasticity via Deep Weight Factorization of Prior Scales —
- Identifiability of a dissipative knowledge-dynamics model: exact recovery under designed excitation, degeneration on observational data —
- Outperformance Inverse Optimization: Learning Objective Functions that Outperform Agent Decisions —
- FedSSMCoOp: SSM Encoders for light-weight Federated Prompt Learning for Few-shot Classification —
- Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking —
- Do Generative Priors Align with Human Naturalness Perception? —
- Expected Sample Complexity in Multi-Armed Bandits —
- Itgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust, Mixed-Dialect and Code-Switched Arabic ASR —
- Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions —
- MSU Team at the Explainable Deepfake Detection Challenge 2026: Grounded Artifact Evidence for Deepfake Detection —
- Possibilistic Radial Transport for Approximate IM Inference —
- EASE: Entropy-Adaptive Distribution Shaping for Evading AI-generated Text Detectors —
- Extreme Binary Classification: Extreme Value Theory for Extreme Constraint on False Negative —
- Perceptually Aligned Evaluation of Style Transfer —
- Playing with Kruskal: algorithms for flat and hierarchical watershed cuts —
- Scalable Patch-Level Self-Supervised Learning —
- Controlling Dependence in Implicit Generative Models via Spread Mutual Information —
- Gaussian Equivalence for Multi-Head Self-Attention —
- AdSpark: A Large-Scale Dataset and Benchmark for Product-Centric Advertisement Video Generation —
- The Long Road to the Same Answer: Cognitive Bias Under Escalating Reasoning Budgets in Large Language Models —
- Cache the Encoder Within:Compact, Reusable Memory across LLM Queries —
- From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs —
- Efficient Provably Private Classification with a Tabular Foundation Model —
- Towards Calibrated Probabilistic Forecasts for Events of Interest via Outcome-Conditional Recalibration —
- Multi-Agent Coordination via Support-Preserving Distillation —
- SkillSandbox: Skill Verification via Dynamic Scenario Synthesis —
- Temporal Residual Bottleneck for Robust Asynchronous Collaborative Perception —
- ExperienceIndex: Artifact-Grounded Memory —
- I would rather quit NLP than read another paper like this: The rise of antithesis in NLP papers —
- Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position —
- A Probabilistic Perspective on Wasserstein-Based Evidential Uncertainty for Out-of-Distribution Segmentation —
- YANchor-4B: Effective Long-Horizon Reasoning in O(N) Time with O(1) Memory —
- m-Set Adversarial Bandits with Winner Feedback —
- HarnessIR: Harnessing Multimodal Foundation Models for Universal Real-World Image Restoration —
- LLM4Impact: Integrating Heterogeneous Information for Scientific Impact Prediction —
- InterView-C: A Synchronized Multimodal Corpus of VR Avatar-Mediated Survey Interviews —
- HySPE: Positional Encoding via Symplectic Dual Shears —
- HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding —
- BagDINO: Multi-View Baggage Re-Identification with DINOv3 —
- Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering —
- Conformal Prediction for Spatially Dependent Data via Sequential Whitening —
- Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents —
- Argos: Adapt Rich Geometric Priors for Generalizable Online Scene-Change-Detection —
- VideoEvolve: Co-Evolving Memory and Retrieval for Long Video Understanding —
- Pre-training of Bayesian Optimization Algorithm through Bayesian Optimization —
- Kinetic Langevin Meets Split Gibbs: Accelerated Posterior Sampling for Imaging Inverse Problems with Diffusion Priors —
- Universal Local Error and Realized Amplification for the First-Order EDM Predictor —
- Broadly Applicable Approximate MCMC for Switching Stochastic Differential Equations Using Uniformization and Time-Conditioned Factorized Neural Likelihood Estimation —
- HuLiGen: Human LiDAR Generation from Parametric Body Models —
- VolCo: Volumetric Contact for High-Fidelity Human Grasp Generation —
- GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning —
- RoBART: Bayesian Additive Regression Trees with Tree-Specific Rotations —
- Finite-Sample Approximation of Hessian-Guided Perturbed Wasserstein Gradient Flows —
- Masked Feature Encoding for Large-Scale Whole Slide Image Representation —
- From Prompts to Trees: Effective LLM-Guided Tree Generation for Few-Shot Tabular Classification —
- LLM Persuasion Is in the Eye of the Evaluation —
- Geometry-Supervised Visual Representation Learning for Multi-Phenotype Lesion Interpretation in Medical VLMs —
- ProtocolMatch: Protocol-Dependent Model Selection for Scientific Dynamics Forecasting —
- PairAudit: Guiding Human Review with Graph Tokens under Distribution Shift —
- LoomSC: Scalable Deep Subspace Clustering with Projector Factorization and Exact Spectral Reduction —
- Video Prediction Policy 2: Predict Better, Act Better —
- PatchBench: Measuring Collateral Damage in Activation Patching —
- Representation: Geometry Supervised Representation Learning of Phenotypes via Counterfactual Reasoning for Medical VLMs —
- TouchScale: 500 Hours of Human Vision and Touch for Visual-Tactile Learning —
- Physics-Aligned Electronic Ground-State Learning Improves Generalization —
- Shared Gaussianization: What Gaussian Regularizers Certify About Contrastive Learning, and What They Miss —
- Revisiting Explainable AI through Model-Independent Concept Dictionaries —
- On the Necessity of Attention-FFN Split in Vision Transformers —
- SemanticFold: Latent Sequence Compression SeparatesLanguage Modeling, Decodability, and Reasoning —
- One-Shot Adaptive Segmentation For Scientific Images —
- Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts —
- From Digital Human Interactions to Physics-Based Humanoid Skills: Physics-Grounded Post-Training of Interaction Generators —
- Performance at What Cost? A Sustainability-Aware Performance Index for Cell and Nucleus Instance Segmentation —
- Learning to Act with Task Progress: Distilling Small Agents from Compact Teacher Supervision —
- How Private is Private? A Comparative Study for Face De-Identification —
- When to Unpair: Regulating Pairing Dependence in Medical Visual In-Context Learning —
- Data Reuse in Non-Stationary Learning —
- Position Forcing: Self-Conditioning 3D Generation —
- Real-Time Joint Audio-Video Generation by Parallel Adapter Composition —
- Dataset Pruning from First Principles: A Label-Free Linear Programming Approach —
- Koopman Observers for Diffusion Acceleration: Correcting Feature Forecasts with Shallow Measurements —
- Temporally Interpretable Differentiable Decision Trees —
- Input-Blind Controls Produce Substantial Oracle Headroom for Layer Programs in Multiple-Choice Evaluation —
- Document-Level Text Simplification in Estonian Using Large Language Models —
- Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving —
- MOTIP2: Spatial Priors for End-to-End Multi-Object Tracking —
- Safe Meta-Policy Design with Risk Control —
- Gaussian Density Splatting Network —
- Self-correction Optimization for Interleaved Multimodal Generation —
- Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models —
- Rubix: Global Correspondence-Free Point Set Alignment through Assignment Geometry —
- Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds —
- Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL —
- GraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural Networks —
- CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution —
- Derivative Gaussian Processes on a Two-Direction Budget —
- SGF+: Decoupling Gradient Flows for Autoregressive Video Generation —
- Detecting Adversarial Images through Response Profiles of Vision-Language Models —
- Q-Learning with Scalar Adjoint Matching —
- ECHO: Embodied Camera Observations of Human Object Carrying —
- RunningTab: Direct Workspace Interaction with Environment-Side Tabs —
- PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs —
- MORCA: Offline-to-Online Reinforcement Learning for Adaptive Cache Reuse in Video Diffusion Acceleration —
- Label-free cell counting and viability prediction with brightfield imaging and deep learning —
- Two-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion —
- Best Arm Identification for Bandits with Shifting Means —
- Insights from Autoresearch for Solar Panel Segmentation —
- QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation —
- Oracle-Efficient and Parameter-Free Agnostic Smoothed Online Learning —
- Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models —
- Your Prompt Should Do More: Effects of Retrieval Instructions in Embedding Models —
- Video-Conditioned Generative Joint 2D-3D Hand Motion Recovery —
- RoboJEPA: Scaling Robotic Latent World Models —
- Why Forget-Only Unlearning Needs Memorization —
- GRACE: Generation-aware latent compression for efficient video generation —
- EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory —
- Decoupling Exploration from Optimization in RLVR —
- Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos —
- Tetris3D: 3D Scene Generation With Objects That Fit Together —
- Generative AI for Autonomous Driving: Frontiers and Opportunities —
- Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages —
- Pre-training, Reasoning, Benchmarking: X-ray Report Generation on CheXpert Plus Dataset —
- Route-Verify-Vote: Procedure-Conditioned Self-Consistency for Mixed-Domain Reasoning —
- Bounded Autonomy and Verifiable Safety for Agentic AI Enabled Automation —
- Just for FUNS: LLM-Guided Spatio-Temporal Graph Node Generation for Forecasting Unobserved Node States —
- HydroSphere: A Framework for Governed, Self-Healing Wastewater Infrastructure —
- HCPN-GCN: Scaling Hierarchical Prototype Networks with Cone Geometry for Continual Graph Learning —
- Autonomous Driving Research Requires a Community-Driven Data Paradigm —
- PanoPed: Beyond Bounding Boxes for Sim-to-Real Panoramic Pedestrian Tracking —
- Child ASR Adaptation with Adult Retention: An Empirical Study —
- When Forgetting Looks Like Improvement: Metric Masking in Streaming Diarizer Adaptation and the Price of Rehearsal —
- Emo-Jev: Probabilistic Reasoning for Emotion Classification with Jev —
- MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models —
- CoDR: Training-Free Confidence-Drift Remasking for Diffusion Language Models —
- Leveraging LLM-Generated Explanations for Detecting Emotionally Rewritten Fake News —
- Beyond the Sycophancy Score: How Task, Model, and Pressure Shape LLM Yielding —
- Beyond Risk Prediction: Evidence Grounding and Psychosocial Factor Verification for Explainable Suicide Risk Assessment —
- QuanLing: Cross-Branch Validation of Language Distance Quantification on Western Romance —
- LRCC: Generalizing Low-Rank Compression with Conditional Computation —
- Tiny-Scale Chinese BERT Pretraining: A Controlled Comparison of MLM, WWM, and MacBERT Strategies —
- FinVector-Market-4B: A Controlled Study of LoRA Adaptation for Structured Financial Tasks —
- Trust-Region Optimization for Smooth Potential-Interaction Energies in Wasserstein Space —
- Steering Follows Geometry, Not Labels: Emotion Directions in a Full-Duplex Speech Model —
- CARE: Certifying Acceleration for Vision-Language-Action Inference —
- VCR-Bench: A Modular Open-Source Benchmark for Video Classification Robustness —
- SPW-Nav: A Streaming Panoramic World Model for Language-Guided Navigation —
- One-Slide Calibration of Pathology Foundation Models —
- RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding —
- On KL-Regularized Policy Optimization —
- Work While They Sleep: Exploiting Evaluation Latency for Fully Bayesian Optimization —
- The Best Optimizer Depends on Batch Size —
- S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens —
- Zero-Shot Brain MRI Inpainting with 2.5D Unconditional Flow Priors —
- LASER: Latent Space Adjoint Matching for Support-Constrained Entropy-Regularized Offline RL —
- How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis —
- Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models —
- Sigma-Hunter: A Domain-Specific Language Model for Threat Hunting and Detection Engineering —
- Personalize at Test Time: Learning User Preferences for Image Generation —
- BEACON-SP: Ontology-Grounded GraphRAG Framework for Clinical Suicide Risk Assessment —
- Beyond Explanation: Debugging Medical Imaging Models via Concept Intervention —
- Shape-Bayes: Bayesian Inference of Structured Shapes under Visual Ambiguity —
- Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs —
- Shared-Roadmap Generation and Evaluator for Multi-Agent Path Planning Using Heterogeneous Graph Neural Network —
- Careful Judge: Safe and Efficient Human-AI Collaborative Decision Making —
- Quadratic Weak-to-Strong Generalization in Random Feature Networks via Random Matrix Theory —
- Learning Transition Kernels of Jump-Diffusion Processes with Conditional Diffusion Models —
- Multi-Label Topic Assignment via LLM Distillation: A Comparative Analysis of Generative vs. Discriminative Student Models —
- GUARD: Geometric Uncertainty-Aware Point Cloud Denoising and Segmentation for Robotic Hard Disk Drive Disassembly —
- Covariate-dependent Joint Modeling of Multivariate Ordinal Preferences and Its Connections with Comparison Models —
- OverLay++: Dense-Overlap Layout-to-Image Generation Dataset —
- DISRQAD: Diffusion Image Super-Resolution Quality Assessment Dataset and Benchmark —
- Anaximander: Interactively Running Geospatial Deep Learning Models on Any Compute Backend —
- U-Space: Uncovering When and Why Uncertainty Arises in Language Models —
- FedRSPO+: A Heterogeneity-aware Algorithm for Decision-focused Federated Learning —
- Constraint Tree Exploration for Learning from Language Feedback —
- Same Text, Different Prediction: Serving-Context Nondeterminism in Text Classifiers —
- From Uncertainty to Action: Learning to Steer LLM Agents —
- SPLATIFY: Reproduce, Discover, Innovate! From Papers and Ideas to Trainable 3DGS Code —
- TopoCurve: Geometry-Aware Topology Reasoning via B'ezier Curves in Autonomous Driving —
- The Deceptive Bandit Problem: Exploratory Coupling and the Fragility of Multi-Agent Learning —
- Depth-to-RGB: Repurposing a Frozen Depth Estimator for Geometry-Guided Compositing —
- The Impact of Likelihood Tempering on the Limiting Predictive Moments of Variational Bayesian Linear Neural Networks —
- A Cognitive-Aware QML-CRL Framework for Detecting Affinity and Romance-Investment Fraud —
- Noise Your Prompt: Noising Conditioning Tokens in Continuous Diffusion Language Models —
- sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak —
- ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation —
- RDGSplat: Render-Dedicated Geometry for Novel View Synthesis —
- LayerRoPE: Dynamic Depth-wise Magnitude & Angular Superposition —
- One Frame, Full Heartbeat: ECG-Free 4D Cardiac Cine MRI Synthesis via Radial-Decomposed Flow Matching —
- R-CNN-Based Chess Position Recognition —
- Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver —
- StableGrasp: Reconstructing Physically Stable Human Hand Grasps from Single Images —
- Exact Dynamics and Finite-Sample Trajectory Recovery of Linear Recursive Feature Machines —
- StyleFields: Multi-Scale AdaIN-Modulated Implicit SDFs for Coarse-to-Fine 3D Shape Reconstruction and Editing —
- Few Bits, One Law: Toward W2A4KV2 —
- Do Vision Models Learn Physical Constraints or Rendering Shortcuts? A Counterfactual Benchmark for Grounded Physical Consistency —
- An Accuracy-Information Tradeoff for Loss-Difference Conditional Mutual Information —
- Multi-Objective Aligned Small Language Model Framework for SUD Patient Dialogue Generation —
- A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video —
- Consistent Distribution Matching for Data-Free Diffusion Distillation —
- PVSync: A Unified Lip-Sync Expert for Timing and Articulation —
- Trajectory Abstraction for the Science of Language Agent Behavior —
- Pooling Representation Autoencoders for Efficient Diffusion —
- An Informational Curse of Horizon in Goal-Conditioned Policy Learning —
- Efficient Best-of-N policy evaluation for inference-time alignment —
- Adaptive Visual Token Reduction for Accelerated Image Understanding —
- PhyDiCT: Plug-and-Play CT Reconstruction from Sparse X-Rays via Differentiable Rendering and Strong Priors —
- The Symbol of the Surrogate: Measuring Numerical Provenance in Neural PDE Solvers —
- Hardware-aware Calibrated Clustered Attention for Efficient Visual Geometric Transformers —
- EM-SNN: Efficiently Modulated Spiking Neural Network for Remote Sensing Image Dehazing —
- LeCuration: A Tiny World Model as a Data Curation Multi-Tool —
- Kuration SDK: Addressing the Virtual2Real Gap via Data Curation —
Important terms
- Activation-Informed Pareto-Guided Low-Rank Compression
- This technique finds a compressed representation of model activations by seeking the best balance between small size and high performance. It's a key method for making huge language and vision models more practical to use.
- From Uncertainty to Action
- This research focuses on training language models so they can actually take real-world actions instead of just generating text. It uses feedback loops to help agents navigate uncertainty and make decisions.
- APCD (Adaptive Path Contrastive Decoding)
- This method improves the reliability of large language model generation by guiding the output process more effectively. It's crucial for making complex LLM outputs less unstable.
- FTA-Mem
- This memory technique anchors facts with time and emotion to solve models forgetting details in long conversations. It helps ensure reliable recall over extended dialogue.