AI papers — 2026-09-28
Today’s work focuses on exploring mixed neural posterior estimation within simulators that possess both discrete and continuous parameters. The research investigated methods to handle this hybrid parameter space to improve the accuracy of posterior estimations in these complex models. One approach involved examining gradient-momentum coupling as a parameter-space proxy for learning progress, suggesting a way to track how much progress is being made based on the gradient dynamics itself. This ties into broader questions about how learning evolves within these systems.
Furthermore, there is work on distribution-conditioned transport which seems relevant to modeling the flow or transformation of data within these parameter spaces. The implications here suggest new ways to better understand and control simulation processes that rely on mixed parameter types. However, the abstracts do not detail specific numerical results for this particular estimation task yet.
The proof of concept study explored using large language models to generate personalized networks derived from therapy session transcripts. Researchers attempted to leverage these models to create tailored connections based on the content of the sessions. This suggests a potential avenue for personalized therapeutic support or understanding relational dynamics within clinical contexts. The findings indicate that this approach is feasible for generating such networks.
However, the abstracts do not detail specific metrics on network utility or user engagement. This work opens up avenues for applying generative AI to sensitive textual data in mental health settings. It remains an exploratory proof of concept rather than a fully validated clinical tool.
The work on neural ideals and neural codes explored an algebraic framework for classifying and interpreting neural networks, suggesting a way to structure the underlying principles of these systems. This contrasts with the manifold projection and iterative autoencoder refinement approach, which focused on masked language modeling by refining projections to better capture data structure.
Furthermore, research into understanding distribution shifts in machine learning force fields addressed how to mitigate changes in data distributions when applying these models. In parallel, scMEDAL introduced a deep mixed effects autoencoder for single-cell transcriptomics analysis specifically to visualize batch effects. Another effort involved improving molecular-morphology contrastive pretraining by utilizing deep-learning based morphology profiles.
LapDDPM contributed to robust single-cell manifold generation through spectral perturbation diffusion. CaC advanced video reward models using hierarchical spatiotemporal concentrating mechanisms. Finally, a multi-agent LLM framework with specialized analyzers was developed for detecting time series anomalies like an expert.
The work on VLAA-GUI focused on developing a modular framework designed to manage the lifecycle of GUI automation tasks by explicitly defining when to stop, how to recover from errors, and when to initiate a search for alternative strategies. This approach was built upon prior explorations in preference-based opponent shaping within differentiable games.
This suggests that structuring the interaction between an agent and its environment could guide these decision points. The findings indicated that this modular structure provides a more robust method for handling the inherent unpredictability of GUI interactions compared to monolithic automation scripts.
Furthermore, the framework's design seems to draw parallels with efforts in diagnosing compositional binding failures in vision-language models. This implies that understanding when a sequence of visual and textual inputs breaks down is key to effective recovery. What remains open is how this preference-based shaping translates directly into concrete stopping criteria within complex, real-world GUI scenarios.
The work on large language models explored the concept of stepwise intrinsic rewards for reasoning, suggesting a method where rewards are structured incrementally to guide the model through complex tasks. This approach aims to improve reasoning capabilities by providing intermediate feedback rather than just a final score.
In parallel, research into geometric-photometric event-based three dimensional Gaussian ray tracing investigated methods for rendering complex scenes. This focused on how light interacts with surfaces using these event-based representations. Another area of focus involved enhancing image quality through consist-retinex, where one step of noise-emphasized consistency training was shown to accelerate the process of high quality retinex enhancement.
Practical applications were addressed by developing a production scheduling framework for reinforcement learning that incorporates real-world constraints. This suggests a move toward more robust deployment strategies. On the acoustic side, polychirp demonstrated multi-species bird song classification using tinyml on low-power acoustic sensors.
Sage introduced a sampling aware global evaluation benchmark specifically for species distribution modeling. Finally, efforts to improve language model flexibility included achieving tokenizer flexibility through heuristic adaptation and supertoken learning techniques. This also includes exploring counterfactual recoverability in on-policy distillation to ensure that not every divergence in model behavior is suppressed during training.
The work on policy regret for embedding model routing explored the application of contextual bandits with low-rank experts to manage the complexity of routing decisions. This approach aimed to find a balance between exploration and exploitation when selecting which expert model to use based on the current context. The findings suggested that this method provided a tractable framework for mitigating policy regret.
However, specific quantitative results regarding performance gains were not detailed in these abstracts. Furthermore, research into smooth piecewise cutting for neural operators addressed the challenge of handling discontinuities and sharp transitions within these models. This suggests a technique to improve their robustness when dealing with non-smooth functions.
Concurrently, investigations into learning budget-efficient thinking under policy-dependent solvability examined how agents can learn optimal strategies when the solvability of a problem itself depends on the chosen policy. This work points toward developing more adaptive decision-making processes that account for inherent uncertainties in the system's structure.
The work on ReasonAudio focused on establishing a benchmark for evaluating reasoning beyond simple text-audio matching. This suggests that current retrieval methods need to be tested against more complex inferential tasks. This effort builds upon the broader landscape of agent development, as seen in SkillFlow's scalable and efficient system for agent skill retrieval.
This implies a need for robust evaluation metrics when deploying such systems. Simultaneously, the survey of Nigerian low-resource languages by NaijaNLP highlights the critical gap in resources available for these complex evaluations across diverse linguistic contexts.
The work on Human-1 by Josh Talks introduced a full-duplex conversational modeling framework in Hindi using real-world conversations. This provides a model for how to structure complex interaction testing. This contrasts with the strategic overclaiming of LLM reasoning capabilities through evaluation design, which suggests that simply designing tests is not enough; the quality of those tests matters significantly.
The spectral-sphere-constrained hyper-connections research points toward novel architectural approaches in knowledge representation. Jagarin addresses the practical deployment challenge by creating a three-layer architecture for hibernating personal duty agents on mobile devices. Finally, FlyAOC explored evaluating agentic ontology curation using scientific knowledge bases from Drosophila, indicating ongoing efforts to validate how agents manage and utilize specialized domain knowledge.
The work on the Affective Flow Language Model for Emotional Support Conversation explored how to create a model capable of maintaining a sustained, emotionally resonant dialogue. This focused on the flow of conversation rather than just discrete responses. This involved training a language model to adapt its output based on the emotional trajectory of the interaction.
Simultaneously, research into RAPTOR focused on ridge-adaptive logistic probes, suggesting a method for probing complex systems where the sensitivity adjusts dynamically based on local conditions within a network. In parallel, prompt-based continual compositional zero-shot learning demonstrated that models could learn new tasks simply by being given natural language instructions without explicit retraining on new data.
Furthermore, the demo involving generative AI in radiotherapy planning showed how user preferences could be integrated into treatment plans. This suggests a pathway for personalized medical applications. The identification of high performing wearable human activity recognition models without training highlighted the potential for zero-shot transfer in sensor data interpretation.
Finally, efforts to improve motion in image-to-video models through rebalancing reference frame dominance and selective off-policy reference tuning with plan guidance indicate ongoing work toward more stable and controllable generative video outputs.
Today's papers
- Mixed neural posterior estimation for simulators with discrete and continuous parameters. [paper] [episode]
- How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation. [paper]
- Gradient-Momentum Coupling: A Parameter-Space Proxy for Learning Progress. [paper]
- Distribution-Conditioned Transport. [paper]
- GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory. [paper]
- AIWizards at MULTIPRIDE: A Hierarchical Approach to Slur Reclamation Detection. [paper] [episode]
- On the Expressive Power of Transformers for Contextual Relations. [paper]
- Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries. [paper]
- Testing the Utility of Using Large Language Models to Create Personalized Networks From Therapy Session Transcripts: A Proof of Concept Study. [paper]
- Adaptive Dual-Mode Distillation with Incentive Schemes for Scalable, Heterogeneous Federated Learning on Non-IID Data. [paper]
- Formal Abductive Latent Explanations for Prototype-Based Networks. [paper]
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations. [paper] [episode]
- A Dual-Helix Governance Approach Towards Reliable Agentic Artificial Intelligence for WebGIS Development. [paper]
- On the Optimality of the Median-of-Means Estimator under Adversarial Contamination. [paper]
- AI Wizards at CheckThat! 2025: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles. [paper] [episode]
- Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments. [paper]
- Neural Ideals and Neural Codes: An Algebraic Framework for Neural Network Classification and Feature Interpretation. [paper]
- Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling. [paper]
- Understanding and Mitigating Distribution Shifts For Machine Learning Force Fields. [paper]
- scMEDAL for the interpretable analysis of single-cell transcriptomics data with batch effect visualization using a deep mixed effects autoencoder. [paper]
- Improving Molecular-Morphology Contrastive Pretraining using Deep-Learning-based Morphology Profiles. [paper]
- LapDDPM: Spectral Perturbation Diffusion for Robust Single-Cell Manifold Generation. [paper]
- CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating. [paper]
- Detecting Time Series Anomalies Like an Expert: A Multi-Agent LLM Framework with Specialized Analyzers. [paper]
- VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation. [paper]
- Preference-based opponent shaping in differentiable games. [paper]
- MedHal: a Synthetic Dataset for Medical Hallucination Detection. [paper]
- When Bias Meets Trainability: Connecting Theories of Initialization. [paper]
- Attribution Bias in Large Language Models. [paper]
- GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow. [paper]
- Latent Generative Solvers for Generalizable Long-Term Physics Simulation. [paper]
- Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models. [paper]
- Stepwise Intrinsic Rewards for Reasoning in Large Language Models. [paper]
- Geometric-Photometric Event-based 3D Gaussian Ray Tracing. [paper]
- Consist-Retinex: One-Step Noise-Emphasized Consistency Training Accelerates High-Quality Retinex Enhancement. [paper]
- A Production Scheduling Framework for Reinforcement Learning Under Real-World Constraints. [paper]
- PolyChirp: Multi-Species Birdsong Classification Using TinyML on Low-Power Acoustic Sensors. [paper]
- SAGE: A sampling-aware global evaluation benchmark for species distribution modeling. [paper]
- Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning. [paper]
- Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation. [paper]
- Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts. [paper]
- Smooth Piecewise Cutting for Neural Operator to Handle Discontinuities and Sharp Transitions. [paper]
- Nice Fold or Hero Call: Learning Budget-Efficient Thinking under Policy-Dependent Solvability. [paper]
- Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing. [paper]
- Combee: Scaling Prompt Learning for Self-Improving Language Model Agents. [paper]
- On the origin of neural scaling laws: from random graphs to natural language. [paper]
- The Plot Twist: Jailbreaking Unified Multimodal Models with a Three-Act NarrativeAttack. [paper]
- AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment. [paper]
- ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval. [paper]
- NaijaNLP: A Survey of Nigerian Low-Resource Languages. [paper]
- Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations. [paper]
- SkillFlow: Scalable and Efficient Agent Skill Retrieval System. [paper]
- Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design. [paper]
- Spectral-Sphere-Constrained Hyper-Connections. [paper]
- Jagarin: A Three-Layer Architecture for Hibernating Personal Duty Agents on Mobile. [paper]
- FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases. [paper]
- Affective Flow Language Model for Emotional Support Conversation. [paper]
- RAPTOR: Ridge-Adaptive Logistic Probes. [paper]
- Prompt-Based Continual Compositional Zero-Shot Learning. [paper]
- Demo: Generative AI helps Radiotherapy Planning with User Preference. [paper]
The papers
- Unsupervised Methods for Video Quality Improvement: A Survey of Restoration and Enhancement Techniques — As a meticulous AI researcher, I have thoroughly analyzed both provided texts regarding "Unsupervised Methods for Video Quality Improvement: A Survey of Restoration and Enhancement Techniques." My objective is to synthesize this information into a comprehensive, detailed summary [episode]
- MSAlign: Aligning Molecule and Mass Spectra representations for Metabolite Identification — Accurately identifying metabolites i.e. small molecules from mass spectrometry data remains a core challenge in metabolomics, with broad applications in drug discovery, environmental analysis, and clinical research [50, 65]. [episode]
- The Communication Map of a Transformer — The components of a transformer communicate by writing to and reading from a shared residual stream, which all attention heads and individual neurons read from and write back to (Elhage et al., 2021). [episode]
- UL-VIO: Ultra-lightweight Visual-Inertial Odometry with Noise Robust Test-time Adaptation — UL-VIO proposes an ultra-lightweight Visual-Inertial Odometry (VIO) network capable of test-time adaptation (TTA) based on visual-inertial consistency, designed to address challenges in deploying VIOs on resource-constrained devices and handling environmental distribution shifts [episode]
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation — Object-centric Self-improving Preference Optimization (OSPO) is a self-improving framework designed to enhance object-level text–image alignment for Text-to-Image (T2I) generation, specifically targeting object hallucination and fine-grained alignment failures. [episode]
- Mixed neural posterior estimation for simulators with discrete and continuous parameters — Neural Posterior Estimation (NPE) enables rapid parameter inference for complex simulators with intractable likelihoods by training an inference network to estimate a probability density over parameters given data, typically assumed to be continuous. [episode]
- Beyond L 2: Generalizing Abductive Latent Explanations to Diverse Prototype-Based Architectures — Prototype-based neural networks are hailed as interpretable-by-design architectures, and Abductive Latent Explanations (ALE) were introduced to provide formal, mathematically guaranteed explanations that leverage the intrinsic structure of these networks to ensure both predictive [episode]
- Representation Transfer of Foundation Models for Ultra-Widefield Retinal Imaging — Despite the widespread adoption of foundation models as feature extractors for medical imaging, relatively little is understood about how different pretraining strategies influence the transferability of learned representations to weakly supervised ophthalmic imaging tasks. [episode]
- Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs — Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities, focusing on English-Mandarin code-switching as a case study. [episode]
- RefRef: A Dataset and Benchmark for Reconstructing Refractive and Reflective Objects — This work introduces a synthetic dataset and benchmark for reconstructing scenes with refractive and reflective objects from posed images, named RefRef, to address limitations in current 3D reconstruction and novel view synthesis approaches that assume straight light paths for op [episode]
- AIWizards at MULTIPRIDE: A Hierarchical Approach to Slur Reclamation Detection — Detecting reclaimed slurs represents a fundamental challenge for hate speech detection systems, as "the same lexical items can function either as abusive expressions or as in-group affirmations depending on social identity and context." This work addresses Subtask B of the MultiP [episode]
- Supervised Deep Multimodal Matrix Factorization for Interpretable Brain Network Analysis — We present Supervised Deep Multimodal Matrix Factorization (SD3MF), an interpretable framework for integrative brain network analysis that generalizes Symmetric Nonnegative Matrix Tri-Factorization (SNMTF) from unsupervised single-graph clustering to supervised prediction over po [episode]
- Learning with Volterra Neural Networks: A System Theoretic Perspective — This paper presents kVNN, a learnable kernelized Volterra Neural operator designed for compact higher-order filtering by combining order-wise Volterra filtering structure with learnable polynomial-kernel atoms. [episode]
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations — Recent advances in pretraining general foundation models have significantly improved performance across diverse downstream tasks, and while autoregressive (AR) generative models like GPT have revolutionized NLP, most visual generative pretraining methods still rely on BERT-style [episode]
- AI Wizards at CheckThat! 2025: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles — This paper presents "AI Wizards’ participation in CLEF 2025 CheckThat! Lab Task 1: Subjectivity Detection in News Articles, classifying sentences as subjective/objective in monolingual, multilingual, and zero-shot settings." The primary strategy explored was to enhance transfor [episode]
- Adapting Visualization Techniques for Time-Series Anomaly Detection: From Convolutional Neural Networks to Convolutional-Recurrent Neural Networks — Deep neural networks are often perceived as ”black boxes”, limiting their adoption where transparency and explainability are crucial, raising ethical and legal concerns, particularly concerning automated decisions under regulations like GDPR which require justification of dec [episode]
- MedHal: a Synthetic Dataset for Medical Hallucination Detection —
- Learning Operators by Regularized Stochastic Gradient Descent with Operator-valued Kernels —
- What is the Added Value of UDA in the VFM Era? —
- Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning —
- Where You Place the Norm Matters: From Prejudiced to Neutral Initializations —
- When Bias Meets Trainability: Connecting Theories of Initialization —
- ChemMLLM: Chemical Multimodal Large Language Model —
- Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design —
- LapDDPM: Spectral Perturbation Diffusion for Robust Single-Cell Manifold Generation —
- A Production Scheduling Framework for Reinforcement Learning Under Real-World Constraints —
- DeepC4: Deep Conditional Census-Constrained Clustering for Large-scale Multitask Spatial Disaggregation of Urban Morphology —
- Who cuts emissions, who turns up the heat? causal machine learning estimates of energy efficiency interventions —
- Neural Bridge Processes —
- Stochastic Bilevel Optimization with Heavy-Tailed Noise —
- Adaptive Dual-Mode Distillation with Incentive Schemes for Scalable, Heterogeneous Federated Learning on Non-IID Data —
- The Plot Twist: Jailbreaking Unified Multimodal Models with a Three-Act NarrativeAttack —
- On the Optimality of the Median-of-Means Estimator under Adversarial Contamination —
- LLMBisect: Breaking Barriers in Bug Bisection with A Comparative Analysis Pipeline —
- Flow Reconstruction from Sparse Measurements in Urban Drainage Networks: An Application and Evaluation of Data-Driven Sparse Sensing —
- Models Got Talent: Identifying High Performing Wearable Human Activity Recognition Models Without Training —
- Spatial Information Bottleneck for Interpretable Visual Recognition —
- Formal Abductive Latent Explanations for Prototype-Based Networks —
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward —
- Anchor to Expand: Semantic Anchoring for Personalized Text-to-Image Diffusion Models —
- Frequency-Decomposed Avatar Representation for Varying Camera Distances —
- Testing the Utility of Using Large Language Models to Create Personalized Networks From Therapy Session Transcripts: A Proof of Concept Study —
- Consist-Retinex: One-Step Noise-Emphasized Consistency Training Accelerates High-Quality Retinex Enhancement —
- Demo: Generative AI helps Radiotherapy Planning with User Preference —
- Prompt-Based Continual Compositional Zero-Shot Learning —
- FLAME: Flow Enhanced Legendre Memory Models for General Time Series Forecasting —
- Geometric-Photometric Event-based 3D Gaussian Ray Tracing —
- OASI: Objective-Aware Surrogate Initialization for Multi-Objective Bayesian Optimization in TinyML Keyword Spotting —
- Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries —
- On the origin of neural scaling laws: from random graphs to natural language —
- DeepFedNAS: Efficient Hardware-Aware Architecture Adaptation for Heterogeneous IoT Federations via Pareto-Guided Supernet Training —
- Generative Modeling of Discrete Data Using Geometric Latent Subspaces —
- MeshGraphNet-Transformer: Scalable Mesh-based Learned Simulation for Solid Mechanics —
- RAPTOR: Ridge-Adaptive Logistic Probes —
- Stepwise Intrinsic Rewards for Reasoning in Large Language Models —
- Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models —
- Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment —
- Pseudo-Invertible Neural Networks —
- Neural Networks With Dense Weights Are Not Universal Approximators —
- Affective Flow Language Model for Emotional Support Conversation —
- FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases —
- Latent Generative Solvers for Generalizable Long-Term Physics Simulation —
- GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory —
- When to Think Fast and Slow? AMOR: Adaptive Entropy Gate for Hybrid Models —
- A Multi-Stage Framework for Kuzushiji Character Recognition in Japanese Historical Documents —
- Scale-invariant Gaussian derivative residual networks —
- A Dual-Helix Governance Approach Towards Reliable Agentic Artificial Intelligence for WebGIS Development —
- Distribution-Conditioned Transport —
- Jagarin: A Three-Layer Architecture for Hibernating Personal Duty Agents on Mobile —
- NBAvatar: Neural Billboards Avatars with Realistic Hand-Face Interaction —
- Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders —
- Layer-wise Target Propagation: Efficient Component Attribution through Target Centric Propagation —
- Spectral-Sphere-Constrained Hyper-Connections —
- Adaptive Subspace Modeling With Functional Tucker Decomposition —
- On the Expressive Power of Transformers for Contextual Relations —
- GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow —
- Combee: Scaling Prompt Learning for Self-Improving Language Model Agents —
- Attribution Bias in Large Language Models —
- Identifying Causal Effects Using a Single Proxy Variable —
- Below-ground Fungal Biodiversity Can be Monitored Using Self-Supervised Learning Satellite Features —
- At FullTilt: Real-Time Open-Set 3D Macromolecule Detection Directly from Tilted 2D Projections —
- Composite Silhouette: A Subsampling-based Aggregation Strategy —
- Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding —
- Unlocking the Forecasting Economy: A Suite of Datasets for the Full Lifecycle of Prediction Market: [Experiments & Analysis] —
- VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation —
- Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations —
- Introducing WARM-VR: Benchmark Dataset for Multimodal Wearable Affect Recognition in Virtual Reality —
- ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval —
- The Scaling Properties of Implicit Deductive Reasoning in Transformers —
- Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing —
- How Long Does Infinite Width Last? Signal Propagation in Long-Range Linear Recurrences —
- Low-Cost Black-Box Detection of LLM Hallucinations via Dynamical System Prediction —
- Detecting Time Series Anomalies Like an Expert: A Multi-Agent LLM Framework with Specialized Analyzers —
- Gradient-Momentum Coupling: A Parameter-Space Proxy for Learning Progress —
- HaM-World: Soft-Hamiltonian World Models with Selective Memory for Planning —
- Geometry-Aware Simplicial Message Passing —
- UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification —
- How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation —
- StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction —
- Agentick: A Unified Benchmark for General Sequential Decision-Making Agents —
- Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression —
- From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation —
- Complex-Valued Phase-Coherent Transformer —
- AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment —
- Selective Off-Policy Reference Tuning with Plan Guidance —
- Nice Fold or Hero Call: Learning Budget-Efficient Thinking under Policy-Dependent Solvability —
- CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating —
- Tabular Imbalanced Learning: A Survey, Benchmark, and Practical Guide —
- SegRAG: Retrieval Augmented Spatial Prompting for Open Vocabulary Semantic Segmentation —
- Attention Sinks and Outliers in Attention Residuals —
- Federated Martingale Posterior Samping —
- Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models —
- C2P-VAR: Continual and Compositional Personalization in Visual Autoregressive Models —
- Smooth Piecewise Cutting for Neural Operator to Handle Discontinuities and Sharp Transitions —
- Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering —
- Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation —
- Large Language Model Selection with Limited Annotations —
- CODESKILL: Learning Self-Evolving Skills for Coding Agents —
- QASM-Eval: A Dataset to Train and Evaluate LLMs on OpenQASM-3 Beyond Quantum Circuits —
- Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation —
- Statistical Priors for Implicit Preferences: Decoupling Skill Selection as a Local Harness in Personal Agents —
- Proper Scoring Rules for Right-Censored Survival Data —
- Where Flow Matching Leaks: Characterising Membership Signals Along the Interpolation Path —
- SciR: A Controllable Benchmark for Scientific Reasoning in LLMs —
- Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts —
- EmoZone-Talker: Regional Semantic Control of Audio-Driven 3DGS Talking Heads via Facial Action Units —
- Holo-World: Unified Camera, Object and Weather Control for Video World Model —
- Do Location Encoders Capture Spatial Effects? A GeoShapley Benchmark Across Scales —
- Cross-Modality Structural Guidance in 3D Latent Diffusion for Robust FLAIR Super-Resolution —
- TeDiServe: High SLO Attainment Serving for Diffusion Language Models —
- Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark —
- Do Neural Networks Preserve Case Structure? Case-Based Decomposition, Interpretation, and Decision Consistency —
- DreamSat-Pose: Spacecraft Pose Estimation from Single-View 3D Reconstructions and Learned 2D-3D Feature Matching —
- Importance-Aware OBS Pruning for Diffusion Models —
- CT-Merging: Consensus Directions and Task-Specific Scaling for LoRA Adapter Merging —
- Regularizing modality contribution drift in multimodal continual learning —
- Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware —
- Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation —
- A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources —
- A Banach-Space Theory of Markovian Halpern Iteration for Non-Expansive Maps —
- LiD-GLM: Lipschitz-constrained Deep Generalized Linear Models —
- Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair —
- PolyChirp: Multi-Species Birdsong Classification Using TinyML on Low-Power Acoustic Sensors —
- AdapToPASS: Ambiguity-aware Adaptive Spherical Transformer for Panoramic Semantic Segmentation —
- PLSR: Progressive and Localized Super-Resolution of 3D Objects via Localized Latent Voxel Diffusion —
- Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding —
- HybridInfer: Thermal-Aware Reinforcement-Learning Tier Routing for On-Device, Edge, and Cloud LLM Inference —
- When the Preconditioning Exponent Turns Negative: Learning-Rate Coupling and Cross-Environment Generalization —
- ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers —
- Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments —
- Distribution of hitting times for dissipative random dynamical systems on R d, with application to stochastic gradient descent —
- Cosine Similarity Is Not Evidence: Measuring the Noise Floor of Interpretability Transfer Under Quantization —
- Why Clipping Matters in AdaGrad? Toward a High-Probability Theory under Generalized Smoothness —
- Fixed Points Without Fixed Diffusion: Implicit Neural Sheaves for Convergent Test-Time Computation —
- Neural Ideals and Neural Codes: An Algebraic Framework for Neural Network Classification and Feature Interpretation —
- Seasonal and Quantum-inspired Models for Neutron Monitor Time Series Forecasting —
- When Does Advection-Aware Graph Nowcasting Help? A Controlled Study of Distributed Solar Ramp Forecasting with a Self-Supervised Cloud-Motion Estimator —
- A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID —
- Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling —
- Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents —
- Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline —
- Bringing AI to Autonomous Systems -- From Cognition to Collective Intelligence —
- A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models —
- Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents —
- SlideLab: Audience-Centered Scientific Slide Generation and Evaluation —
- SignTrace: Describe a Sign, Find the Word —
- NeuralCert: certified computational discovery of extremal mathematical constructions —
- Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops —
- A Benchmark Framework for Screening Automation in Systematic Reviews —
- Staged Depth Training: A Representation Curriculum for PINNs —
- PALM: Point-in-Time Adaptation for Financial Language Models —
- Adaptive Random Matrices in Gaussian Bandits: Spectral Universality and Selection-Induced Outliers —
- ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure? —
- Guarded Gradient-Based Activation Steering of Shutdown Responses in Qwen3.5-0.8B: A Minimum-Step Policy —
- When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess —
- Parameters vs. Context: TRACE Fine-Tuning for Robust Retrieval-Augmented Generation —
- GAUDI: Geometry-Aware Diffusion for Calibrated Air-Quality Time-Series Imputation —
- Bridging LLM Agents and Data Spaces: An Architectural Mediation Approach using the Model Context Protocol —
- Low-Rank Friction for Memory-Efficient Transformer Pretraining —
- Learning coarse-step dynamics and internal mechanical response with graph networks —
- Adaptive multi-resolution Gaussian processes: Scalable exact inference with naturally data-sparse covariance matrices —
- Strategic Self-Consistency —
- AlphaEarth distinguishes cities but compresses urban variation —
- Cost-Aware Best-LLM Identification using Dueling Feedback —
- DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep Learning —
- Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems —
- From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning —
- LiTe-GS: Oracle-Efficient Next Best View Selection for 3D Gaussian Splatting —
- CSCWD: Cross-Scale Channel-wise Knowledge Distillation for Lightweight Tiny Object Detection on Edge Devices —
- A Synthetic Ground-Truth Framework for the Evaluation of Explainable AI Methods —
- What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study —
- Adaptive Multi-Value Control in LLMs via Causal Activation Steering —
- A Unified Account of Concepts and Chunks —
- All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation —
- Electric Vehicle Charging Station Location Selection using Geospatial Artificial Intelligence (GeoAI) —
- Fake News Theories: Harnessing Disciplinary Insights for Computational Modeling, Detection, and Explanation —
- Improving Molecular-Morphology Contrastive Pretraining using Deep-Learning-based Morphology Profiles —
- ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models —
- Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition —
- Bayesian Uncertainty Quantification for fMRI Functional Connectivity via Simulation-Based Inference —
- Predicting Transmembrane Protein Topology from 3D Structure —
- LensDesigner: A Self-Improving Agent for Optical Lens Design —
- Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model —
- Spectral Feedback for Test-Time Alignment of Protein Diffusion Models —
- RAZOR: Pruning Replaceable Experts in LLMs —
- Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification —
- Pretrained ASR Pseudo-labeling for Noisy Police Audio —
- Reliability-aware Cross-sample Enhancement for Robust Multimodal Sentiment Analysis —
- CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production —
- Moment-guided edge sampling —
- Mentored Decoding: Faster Inference meets Boosting —
- The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation —
- Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework —
- Geometric Feature Learning for Functional Data Valued on the Symmetric Positive Definite Manifold —
- BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering —
- Breaking Homogeneity: Diversifying Persona Sets for Creative LLM Outputs —
- Learning to Bias: Machine Learning-Enhanced Particle Filters —
- Ordinary Nonconvex SGD under Distance-Dependent Moments: Finite-Horizon Stationarity and Nagaev Bounds —
- PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control —
- To Solve Bilevel Optimization with Nonconvex Lower Levels, We Need Second-Order Stationarity —
- Federated Targeted Maximum Likelihood Estimation —
- Benchmarking the Connectomes of Caenorhabditis elegans within the Reservoir Computing Framework —
- Inquesto Score: A reliability Protocol For Voice Agents —
- Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence —
- AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework —
- GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for De Novo Peptide Sequencing —
- REALMS: An AI-Assistant Conversational System for Real-Time Exact Audience Sizing over High-Dimensional Nested Profiles —
- Benchy: towards a universal language for task-oriented AI benchmarks —
- Rank-Reliable Teacher-Guided Fitness Approximation for Expensive Evolutionary Optimization: A TinyML Architecture Search Study —
- Dynamic Regret in Online Convex Optimization with Indicator Switching Costs —
- Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms —
- Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content —
- Atlases Are Already Inside: Recovering Population Templates from Pretrained Diffusion Models —
- Atelier: Learning Local Self-Supervised Features for CryoEM Volumes via Hypernetworks —
- HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases —
- Entropy Regularization: A Free Correction to Cross-Entropy for Verified Demonstrations —
- T-RoPE: Time-Aware Rotary Position Embedding for Sequential Recommendation —
- Reinforcement Learning of Communication in a Mesh of Small Language Models —
- Energy-efficient operation of neural operators for virtual sensing —
- QSV: Quat-Sphere-Vision for Coupled Quaternion Attention on Spherical Lattices —
- Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases —
- The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge —
- Probabilistic Robustness-driven Universal Adversarial Perturbations with Explainability against Deep Reinforcement Learning-based Intrusion Detection System —
- MVAgent: Multi-Agent Video Generation via Consistent Condition Construction and Shot-Level Policy Optimization —
- MedTokenBudget: Lesion-Preserving Token Routing for Dermoscopic Image Classification —
- Audio LLMs Know When They Can't Hear You —
- OpenHail: An Event-Driven Gymnasium Environment for Electric Ride-Hailing Fleet Control —
- Stable initialization without the CLT —
- In-Context Binding Capacity in Language Models —
- MARCEDES: Score-based causal discovery under non-Gaussianity with continuous optimization —
- Conditional Predictive Sufficient Statistics for Visual Representation Learning —
- Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation —
- Recursive Self-Improvement via On-Policy Distillation for Reasoning —
- Population loss in shallow ReLU networks: Bias & families of critical points —
- LLM Parkinsonism: Executive-Control Failure, Token-Inefficient Persistence, and an Uncertainty-Aware Global Executive Control Architecture for Autonomous Language-Model Agents —
- StarWM: Self-Supervised Trained Attention Routing for Robust World Models —
- TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding —
- Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding —
- PixSim: a calibrated open-source simulator of instant-payment fraud, recovery and interdiction under analyst capacity constraints —
- LUMO (Lightweight Unified Multilingual Orchestrator): A Privacy Preserving Offline Voice Assistant —
- MM-VeriAgent: Learning to Use Extensive Tools to Verify Multimodal Misinformation with Reinforcement Learning —
- SAGE: Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion —
- The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading? —
- LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information —
- Combining General and Domain-Specific Pretext Tasks for Brain MR Image Segmentation —
- VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control —
- Parameter Estimation for Unnormalized Discrete Models via Empirically Localized Deformed Bregman Divergence —
- CRC-Router: Risk-Constrained Routing for Medical Agentic AI Systems —
- Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4 —
- NEMSim: Learning Control-Conditioned Multi-Event Physical Dynamics via Executable Event-Mechanism Priors —
- When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification —
- TrafficImag: A Benchmark for Counterfactual Roadside Traffic Video Generation —
- EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection —
- Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents —
- Learning Polarization Image Restoration with General Restoration Priors —
- Amplify What You Gaze At: Target Saliency Boosting in Text-to-Image Generation —
- Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows —
- Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache Eviction —
- SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages —
- From Mono to Stereo: Accelerating Binocular Gaussian Splatting via Reprojection and Selective Patching —
- From S3Q Theory to Implementation: Towards an Architecture for Machine Qualia —
- Mechanism-Aware Ensemble Conditioning for Data-Limited Emulation of Extreme Events —
- ORCA: Evaluating LLMs on Data Science Code Translation —
- Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging —
- Differentiable RNA Secondary Structure Extraction for Deep Learning —
- Training-Free Bottleneck Width Planning for Convolutional Autoencoders —
- Selective Amortization of Full-Budget Counterfactual Reasoning for Visual Token Communication —
- LLPR: Location-aware learning and physics-based reconstruction for raindrop removal from a single image —
- Timo: T aming Mult i modal Diffusion Transformer for Human Mo tion Generation —
- HCOE: Hyperbolic Clinical Ontology Embeddings from Biomedical Language Models —
- Insurance Reserve Intelligence Platform —
- Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More —
- Query-Conditioned Prototype Adaptation for Cross-Domain Few-Shot Learning: Single-Query Inference, Controlled Comparisons, and Failure Modes —
- Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance —
- Missingness-Aware Conformal Prediction Under Cross-Hospital Distribution Shift —
- Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models —
- Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays —
- Towards Universal Representation-Based Process Control —
- Motion Style Slider: Endpoint-Supervised Continuous Style Control for Human Motion Diffusion —
- ConsultMind:Towards Automated Diagnostic Consultation via Uncertainty-Aware Reasoning —
- HasMem: Hard-Origin Adaptively Softened Memory for Long-Term LLM Agents —
- Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes —
- Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment —
- Counterfactual Online Conformal Prediction Under Adaptive Logging —
- A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory —
- Learning Provable Neural Network Observer for Uncertain Dynamical Systems —
- Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness —
- Adaptive Interaction Graphs for Particle Simulation —
- PTC-Decoder: Towards Intelligent SLMs on Offline Resource-Constrained Edge Devices —
- MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation —
- Peer-Grounded Counterfactual Path Planning for Chronic Health Management —
- Attention-Based Adaptive Policies for Simultaneous Speech-to-Text Translation —
- Aligning One-Step Generative Models with Reward-Weighted Transport Distillation —
- Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis —
- I-Parakeet: Integer-Only Conformer ASR on Mobile NPU —
- Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength —
- MDSkin-Net: Multi-Task Skin Lesion Analysis Driven by Pattern Analysis Priors and Spatial Alignment Regularization —
- Learning Chance-Constrained MDPs with Bellman Distributional Certificates —
- SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting —
- Persistent Negatives for Adversarial Black-Box On-Policy Distillation —
- Reliability-Regulated Trajectory Optimization for Progressive COLMAP-Free 3D Gaussian Splatting —
- Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations —
- TISD: On-Policy Self-Distillation with Trajectory Intervention —
- EXAONE Demand 1.0: A Time Series Foundation Model for Demand Forecasting —
- Effects of Transcript Compression on LLM-based Medical Misinformation Detection in Japanese YouTube Videos —
- CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters —
- Conformal Prediction under Exponential-Tilt Joint Shift —
- From Tapping to Hopping: Augmenting Mobile GUI Agents with App-Native Deeplinks —
- Training Graph Foundation Models on The Web Graph —
- From annotation to reasoning: Culture in language models —
- ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning —
- Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations —
- JevSoup: System-One Routing for Training-Free LoRA Composition —
- Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring —
- UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound —
- EPOC: Endpoint-Preserving Online Correction With Compressed Residual State for Multi-Horizon Time Series Forecasting —
- ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos —
- Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models —
- Self-Play Search Distillation for Large Language Model Reasoning —
- MACBT: A Multi-Agent Cognitive Behavioral Therapy Decision Support System with Longitudinal Memory —
- Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms —
- Spackle: Completing Large View Single Image NVS with Adaptive Gaussians —
- LogicTree-RAG: Logic Tree-guided Retrieval-Augmented Generation for Long-form Patent Drafting —
- OneWorld: Learning Consistent Physics Across Actions in World Models —
- DAPEVO: Deep Adaptive Patch Frame-Event Visual Odometry —
- PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem —
- Low-Bit Recurrent States in Hybrid Language Models —
- MVVBench: Benchmarking 4D Reasoning in Vision-Language Models —
- FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models —
- Towards Understanding Momentum Acceleration in River-Valley Loss Landscape —
- IDM-Net: A Lightweight Illumination-Decoupled Modulation Network for Low-Light Image Enhancement —
- Where and When to Force: Routed Forcing for Streaming Avatars —
- Gradient Surgery for Physics-Informed Neural Networks —
- MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens —
- FAVoR: Measuring and Mitigating Author-Style Homogenization in Federated Personalized Generation —
- SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents —
- Factorized axis convolutional gated recurrent unit with dynamic adaptive pooling for remaining useful life prediction of rolling bearings —
- LipSSM: Structurally Lipschitz-Bounded Cascaded State-Space Model via Metric Transfer between Consecutive SSM Layers —
- Coupled Usage-Sense Processes: Temporal and Attributable Lexical Semantic Change —
- Does Uniform Discrete Diffusion Need Time? —
- CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models —
- STORM-Bench: Evaluating Online Video QA under Evolving and Incomplete Evidence —
- FARE: Forensic Acceptance Region Estimation for Catching Bait-and-Switch Image Generators —
- THA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer —
- Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries —
- Self-Supervised Perceptually Interpretable Monocular Depth Estimation —
- PhoenixSR: Generative Heterogeneous Distillation Unleashes Efficient Models for Real-World Super-Resolution —
- PICO: Projection-Informed Consistency Optimisation for 6DoF Surgical Tool Pose Estimation —
- FLIP: Final Layer Inference-Time Probing for Vision-Language Models —
- Learning Hierarchical Causal Representations of the Effects of Forcings on Temperature in Climate Models —
- The Linear Representation Hypothesis for Vision-Language-Action Models —
- ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker —
- TRACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking —
- G squared PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation —
- Same Text, Different Numbers: The Divergence of LLM-Based Measures —
- Robust Successor Features —
- Refining Cytology Predictions with Conditional Random Fields —
- Governed Deduction: Policy-Grounded Premise Authorization Beyond Relevance —
- Metacognitive Selective Ensemble for Mobile Systems —
- Robust Graph Clustering Network for Multiple Missing Data —
- Aurora-X: Built for Extreme Time Series Forecasting —
- Exploiting Spatial Structure for Transductive Few-Shot Classification of Whole-Slide Images —
- Modeling Student Sensemaking with LLMs and Knowledge-Graph-Guided Inference —
- Where Compute Matters: Heterogeneous Attention for Efficient Video Diffusion —
- Cheap, open agents make LLM pollution harder to mitigate —
- Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders —
- Distributed Learning as a Service: The Developer's Perspective —
- CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings —
- Externalized CPDAG Summaries Improve LLM Causal Deduction —
- Band-Selection Stability and Semantic Segmentation Performance: A Study on Hyperspectral City —
- Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents —
- OmouAI: Argumentative Human-AI Policy Deliberation with Simulated Personas —
- SAGE: A sampling-aware global evaluation benchmark for species distribution modeling —
- Block Sparse Attention with Log-Linear Complexity —
- The Residual Stream's Effective Depth —
- DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language Models —
- Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods —
- Double-stream registration with pyramid fusion for HDR video with alternating exposures —
- From Shortcut Learning to Discrete Neural Insertion Sort —
- Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning —
- LocUS: Head Selection and Subspace Projection for Targeted Activation Steering —
- Frame the adversary: a structure-aware attack methodology —
- Do we need to answer that question? Salience and Answerability of Potential Questions in Naturalistic Dialogue —
- AtomWorld-Mem: Memory-Restored World States for Long-Horizon Atomistic Evolution —
- Pocket-STVG: lightweight architecture for Spatio-Temporal Video Grounding —
- Toward AI-Augmented Cooperative Engineering Workflows: Requirements and Architecture the European Rover Challenge —
- Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability? —
- Seeing Semantic Shift: Difference-Aware Sentence-Level Temporal Segmentation of Sign Language Videos —
- CRNDiff: Count-Native Diffusion Framework via Chemical Reaction Networks —
- FedHisto-PAST: Parameter-Efficient Stain-Aware Federated Learning for Cross-Site Lung Histopathology Classification —
- HyperErase: Scale-Calibrated Hypernetwork for Multi-Concept Erasure in Text-to-Image Models —
- Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift —
- Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection —
- Momentum-Guided Federated Split Distillation for Personalized Temporal Edge Intelligence —
- ReG-SAM: Reference Graph-Driven SAM for 2D Foundational Vessel Segmentation —
- I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms? —
- WorldTS: World Modeling for Multimodal Covariate-aware Time Series Forecasting —
- Neural State Prediction: Obstructing Shortcut Learning in EEG Foundation Models —
- Audio emotion recognition for atypical hearing —
- Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting —
- TaskIR: Task-Driven Image Restoration via Degradation Adaptation and Task Feedback —
- Semantic Navigation for Issue Localization in Code Repository —
- SPO: Discovering Adaptive Large Neighborhood Search Operators via Stackelberg Program Optimization —
- Where a Model Sends Its Own Repeated Token —
- Accounting for Bias Enables Sustainable LLM Evaluation —
- Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation —
- Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs —
- ALF: An Active Learning Framework for Scientific Discovery —
- Light Field Primitive for Novel View Synthesis —
- Samples, Sources, Space: Decomposing Data Scale in Spatially Structured Representation Learning of Human Brain Microarchitecture —
- Preserve-and-Compose Training for Composed Image Retrieval —
- Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra —
- Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution —
- DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration —
- Budgeted Quotient-Residual Guidance for Frozen Pocket-Conditioned Molecular Diffusion —
- WeaveAgent: A Two-Stage Tool-Routing Agent for Ultra-High-Resolution Remote Sensing Imagery —
- Purin: A Biology-inspired Mechanism for Artificial Neural Networks —
- RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models —
- Geometric Inconsistency Localization in Multi-View Image Sets —
- Gauss What You Need: Compact Gaussian Splatting Across Scene Scales —
- Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization —
- PIA: A Personal Intelligence Agent Turning Health Conversations into Records and Records into Understanding —
- MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries —
- Identifying Scientists on X —
- MA-WAM: Multi-Agent World-Action Model for Test-Time Planning —
- MoTop: Motion-Topological Model For Micro AU Detection —
- G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies —
- Softmax Reparameterization for Output-Head Quantization —
- UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning —
- Geometric Moment Contraction for Stochastic Nesterov Acceleration —
- Benchmarking Attention for Tabular Foundation Models —
- CytoSPM: Open-Vocabulary Cytopathology Detection with Structured Prompt Bank —
- LUCID: Learning Under Confounding for Inference and Discovery in Time Series —
- More Sensors Only One Field: Rethinking Continual Spatio-Temporal Forecasting —
- CG-HAF: An Interpretable Global-Local Lesion-Burden Fusion Framework for Ordinal Acne Severity Grading in Agentic Skincare Support —
- Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design —
- The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models —
- Stale-Document Poisoning: When Outdated Retrieval Overrides Correct Model Answers —
- DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models —
- Progressive Memory Transformer: Memory-Aware Attention for Time-Series —
- Mutable Transcripts: Mitigating Context Pollution through Editable Conversation State —
- Open Vocabulary Domain Unlearning —
- Programs-of-Layers in LLMs through the Lens of Cortical Areas —
- Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning —
- OpenVAM: Open-World Visual Attention Modeling with VLMs —
- Equation discovery with Bayesian tree-adjoining grammars —
- Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness —
- ContraFM-S2O: Flow Matching-Based One-step SAR-to-Optical Image Translation Model with Contrastive Learning —
- Completed Pairs Hide Capped Failures: A ReVerPi Case Study of Selective Context Projection —
- Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding —
- Differential Attention Unlocks Complementary EEG and Speech Fusion for Emotion Recognition —
- Decodable In-Context State and Model Output Across Training —
- Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers —
- Evaluating the accuracy of KV cache reuse techniques —
- Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents —
- AxonSynth: Domain-Randomized Synthetic Data for Zero-Shot 3D Axon Segmentation in Light-Sheet Microscopy —
- Implicit Neural Representation for Hyperspectral Video Compression —
- ViSTA: A Simple Bridge Extends Visual Alignment to Clinical Time-Series Understanding in Multimodal LLMs —
- From Reward Signal to Visual Utility: A Controlled Audit of Medical VLM Post-Training —
- Different Corruptions, Different Signals: Uncertainty and Loss in Federated Data Quality —
- Diagnosing the Sources of Compositional Failure in Vision-Language Models: A Controlled Analysis —
- Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers —
- Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency —
- KneePreM: Towards 3D Knee MRI Foundation Models via Large-Scale Unlabeled Pretraining and Label-Efficient Fine-Tuning —
- Uncertainty-Aware Federated Learning for Infant Movement Analysis —
- Scaffold: Support Graph Theory Based Sparsification for Graph Neural Networks —
- Beyond Empirical Support: Structured Outlier Generation via Sinkhorn Optimal Transport —
- Game Arena: Strategic LLM Evaluation in Competitive Environments —
- "AI is (not) the new...": A Diagnostic Analogy Framework for Generative AI's Cultural Impacts —
- UQ-LOB: Uncertainty-Aware Limit Order Book Mid-Price Forecasting —
- Prompt Minimization: Reducing Input Redundancy Without Sacrificing Output Fidelity —
- Evaluating Cultural Awareness of LLMs for Haitian Creole —
- SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery —
- ClearGS: Reliability-Aware Gaussian Splatting from Handheld Videos —
- Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge —
- Forensic Twins: Self-Supervised Residual Learning for AI-Generated Image Forensics —
- Structured Reasoning Agentic Framework for Interpretable Critical View of Safety Assessment —
- HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning —
- NEXT: Physics-Informed Neuro-Spectral Exponential Time Differencing Architectures —
- A Flow Matching Framework for Neural Representational Dissimilarity —
- BeatGraph: Self-Supervised Heartbeat Graphs for Infant ECG Representations from the Home Environment —
- MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos —
- Region-Level Black-Box Defense Against Stealthy Embedding-Space Backdoors in CLIP —
- Online Learning via Learned Latent Bayesian Tracking —
- Generalization behavior of OPTQ and the role of regularization —
- Multi-agent Scaling Across Disjunctive and Compensatory Tasks —
- Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights —
- DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education —
- Uncertainty and Explainability in Deep Rough Volatility: A Neural Information-Theoretic Posterior Approach —
- Strategically Diverse Sampling for Self-Training —
- OC-GS: Gaussian Splatting for Irregular Turntable Capture —
- How Far Can INRs Go? Cross-Domain Parameter-efficient INR-Based Semantic Segmentation for Brain MRI —
- Trust Guided Decision Transformer —
- Common-Mode Collapse and Recovery in Direct Feedback Alignment —
- GraphWrit3R: End-to-End 3D Scene Graph Writing —
- New LoRA Skills Should Read but Never Write —
- User Model Extraction via Belief Self-Distillation —
- First-Order Stationarity of Reverse Diffusions —
- Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency —
- Antithetic and Monte Carlo kernel estimators for partial rankings —
- FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders —
- Differentially-Private Decision Trees and Provable Robustness to Data Poisoning —
- Foundations of Reinforcement Learning and Interactive Decision Making —
- Efficient Constrained Graph Search for Post-hoc Error Correction in Binary Classifiers —
- FLex: Joint Pose and Dynamic Radiance Fields Optimization for Stereo Endoscopic Videos —
- Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation —
- Soft-Hard Attention U-Net Model and Benchmark Dataset for Multiscale Image Shadow Removal —
- scMEDAL for the interpretable analysis of single-cell transcriptomics data with batch effect visualization using a deep mixed effects autoencoder —
- Preference-based opponent shaping in differentiable games —
- LEAD: An EEG Foundation Model for Alzheimer's Disease Detection —
- LadderMIL: Multiple Instance Learning with Coarse-to-Fine Self-Distillation —
- NaijaNLP: A Survey of Nigerian Low-Resource Languages —
- Understanding and Mitigating Distribution Shifts For Machine Learning Force Fields —
- VFM-UDA++: Improving Network Architectures and Data Strategies for Unsupervised Domain Adaptive Semantic Segmentation —
- AYLA: Architecting a loss landscape in shallow neural networks to accelerate feature recovery —
- SkillFlow: Scalable and Efficient Agent Skill Retrieval System —
Important terms
- Mixed Neural Posterior Estimation
- This research explores methods for estimating probability distributions when a model has both discrete and continuous parameters, aiming to improve accuracy in complex simulation models.
- Gradient-Momentum Coupling
- This is a technique used as a proxy to track how much progress is being made during learning by examining the dynamics of the gradient and momentum.
- Distribution-Conditioned Transport
- This concept relates to modeling the flow or transformation of data within parameter spaces, suggesting new ways to control simulation processes with mixed parameter types.
- Generative AI for Personalized Networks
- Using large language models, researchers created tailored networks from therapy transcripts to suggest personalized connections, exploring generative AI in mental health.
- Policy Regret for Embedding Model Routing
- This approach uses contextual bandits with low-rank experts to manage the trade-off between exploring new options and exploiting known good ones when choosing which model to use.