AI papers — 2026-08-12

Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment. This paper proposes a hybrid planning architecture for automated driving that... Distillation of Foundation Models for Time-dependent PDEs. Teacher Rollout Extension (TREX): Distilling Foundation Models for... Persistent Recursive Worlds Enable Autonomous Software Evolution. EvoX Genesis (hereafter, Genesis) is a system that reorganizes long-horizon... LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence. LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured Reinforced... Interpretable Causal Discovery via Causal-Effect Constraints. Interpretable Causal Discovery via Causal-Effect Constraints Authors: Cixuan... Disentangling the Expressivity of RoPE. This paper, "Disentangling the Expressivity of RoPE" by Selim Jerad, Anej... Fingerprinting Text-to-Image Diffusion Models via Collapsed Generation. This paper presents a non-invasive model fingerprinting framework for... Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces. Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Authors:... Automated binary classification of hazelnut X-ray images: A deep-learning benchmark for quality assessment. This study presents a benchmark for binary hazelnut quality classification... CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network for End-to-End Image Generation. CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network... Training Under Challenge: Executable Certificates and Challenge-Closed Optimality for Neural Networks. The paper addresses a fundamental limitation in neural network training: "A... Non-Degenerate Risk Certification for Automated Security Decisions: A Decision-Contract Theory with ATT&CK-Aligned Triage as a Worked Instance. Non-Degenerate Risk Certification for Automated Security Decisions: A... Exemplar-based objective classification of gust-induced loads across multiple flight conditions. Objective and Motivation This paper investigates whether an objective... -MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution. ε-MemEvo is a framework for cross-task knowledge transfer in LLM-based program... HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference. HetRoute: Heterogeneous and Cost-aware Collaborative Routing Framework for... Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models. Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language... ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents. ToolHazard is a scalable adversarial environment synthesis framework designed... Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning. This paper studies transfer learning in cooperative multi-agent reinforcement... AVA-Encoder: Towards Agent-Native Video Representation Learning. AVA-Encoder: Towards Agent-Native Video Representation Learning Summary This... HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation. HSTGFormer is a graph-enhanced Transformer framework proposed for efficient... Causal Structure is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library. Causal Structure Is Inducible but Functionally Decoupled: The Routing/Readout... Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges. Graph-Structured Rubrics (GSR) compiles a rubric into a response-independent... Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem. This paper takes a first step toward a unified formal account of agent memory,... Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines. This paper presents a systematic five-step methodology for deploying custom... JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis. The paper introduces Ancient Chinese Character Exegesis (ACCE), a... SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization. SSPO (Structure-Aware Similarity-Weighted Preference Optimization) is a method... VICBench: A Multi-Language Benchmark for Code Vulnerability Detection. VICBench: A Multi-Language Benchmark for Code Vulnerability Detection This... An Efficient Near-Optimal Algorithm for Adversarial-Set Bandits. Problem Setting The paper studies adversarial combinatorial bandits with m-set... Direct Acceleration of Stochastic Root-Finding Without Variance Reduction and Regularization. Problem and Motivation The paper addresses stochastic root-finding problems,... Redistribution-based Cost Inference Improves Sparse Safe Offline RL. The paper introduces the Redistribution-based Cost Inference (RCI) framework,... Draw This First. The paper "Draw This First" presents a method for generating ordered vector... Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning. The paper addresses a central challenge in LLM personalization: task-specific... A Local Sinkhorn Framework for Conditional Distribution Reconstruction of Multidimensional Random Fields. This paper proposes a local Sinkhorn divergence framework for conditional... Continuous-Latent Predictive Modeling with Semantic Alignment for EEG-Language Foundation Models. The paper proposes the Brain Latent Predictive Model (BLPM), an EEG-language... Foresight Without Seeing: Latent Futures for World Action Models. ForeWAM is a dynamics-conditioned direct-policy World Action Model (WAM) that... Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder. Confucius4-TTS is a multilingual zero-shot text-to-speech (TTS) system that... The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance. The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance... Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs. The paper introduces an adversarial reinforcement learning framework to expose... DiG-bench: Discovery in Games. DiG-bench (Discovery in Games) is a new benchmark introduced to measure the... A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions. The paper presents a formal, three-step framework for designing human-aligned... Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision. Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise... Confidence Calibration of Deep Learning Systems. This thesis explores novel methods for improving confidence calibration under... Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization. Most image colorization systems operate in Lab space by predicting chroma (ab)... When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide. This paper investigates when offline evaluation of equal-cost top-k allocation... Fast Length-Squared Sampling for Positive-Semidefinite Matrices. This paper presents a simple rejection-sampling-based algorithm for performing... NAE: Normalizing AutoEncoder. The paper introduces the Normalizing Autoencoder (NAE), a generative framework... How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging. This paper evaluates the complete unsupervised domain adaptation (UDA) pipeline... LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration. LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion... Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams. Diagram-MMU is a multi-modal benchmark designed to assess Multimodal Large... ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening. ScreenShot is a hierarchical transformer-based foundation model for few-shot... Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus. This paper presents the first systematic study of massive activations (MAs) in... Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular Missingness. Paper: "Represent, Then Generate: Multimodal-Conditioned Time-Series Generation... From Visual Widgets to UI Code: Efficient Tool-Grounded Generation. Existing screenshot-to-code systems face a trade-off between flexibility and... AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Core Research Question and Setting This paper, from Salesforce AI Research and... Unifying Physical Backpropagation. Core Contribution This paper develops a unifying theoretical framework for... Certifying What Helps Customer-Return Timing: A Screen-and-Confirm Test for Conditioning Signals, and Why Decay Is Nearly Enough. This paper introduces a screen-and-confirm protocol to certify whether a... Dion3: Full-Stack Orthogonal Updates. Dion3 is a revision of the Muon optimizer that targets overhead at every level... Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation. This paper challenges the standard assumption that LLM model rankings are... SoftWater: Class-Aware Rate Allocation for Softmax Quantization. Based on the paper, here is the summary: The paper introduces SoftWater, a... Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill. Spark-to-Paper is an end-to-end research paper generation system implemented as... Scaling Automatic Research Agents via World Models. The paper "Scaling Automatic Research Agents via World Models" introduces World... Structure-preserving uncertainty quantification for GENERIC dynamics. Structure-Preserving Uncertainty Quantification for GENERIC Dynamics Authors:... From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices. The paper argues that the central research problem for using large language... CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution. CAKE: Compiler–Agent Co-Design for Frontier Kernel Evolution presents a... Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System. This paper introduces UAVQA-Bench, a comprehensive, fully human-annotated... G0.5: One Autoregressive Stream for Robot Reasoning and Action. Galaxea G0.5: One Autoregressive Stream for Robot Reasoning and Action Core... Claim-Level Reliability Assessment for Efficient Test-Time Reasoning. Here is a summary of the paper: The paper introduces Claim-Level Reliability... Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework. This paper presents a structured, benchmark-based comparative assessment of... TELLME: Test-Enhanced Learning for Language Model Enrichment. The paper introduces TELLME (Test-Enhanced Learning for Language Model... How Organizations Use AI: Evidence from ChatGPT. This paper studies how organizations use frontier generative AI by linking... Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence. The paper introduces the “Agentic Self-Improvement” framework, a... Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models. This paper introduces a framework for the automated construction of Dynamic... Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review. This paper reports a single, fully instrumented case study of a large-scale... Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion. The paper introduces GDCE-I (Graph Diffusion Counterfactual Explanation via... Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents. Governed Persistent Memory (GPM) is introduced as an auditable bitemporal... Small-Scale Experiments: Are We There Yet?. Small-Scale Experiments: Are We There Yet? Nicholas Lourie, Kyunghyun Cho,... Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL. Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL... Locating and Controlling Implicit Personalization in Large Language Models. Large language models (LLMs) often shift their outputs in response to implicit... The Advective Fisher-Rao Geometry of Deterministic Measure Transport. The paper introduces a novel Riemannian metric, the advective Fisher–Rao... High-Order Liquid Evidence Encoding for Gradual GNSS Spoofing Detection in Autonomous Driving. The paper proposes a causal high-order liquid evidence framework for detecting... Unifying Generative Models with Path Integrals. This paper formulates generative modeling as a path integral, unifying... AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention. AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive... FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents. FrontierFinance is a fully open benchmark introduced by Samaya AI for... Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment. This paper introduces Group Alignment-induced Sycophancy (GAS), a two-sided... When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation. Core Problem The paper investigates a resource-shared evaluation loop in... Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL. This paper presents a framework for learning loco-manipulation policies using... AmbSentry: Mitigating Sensing Eavesdropping in ISAC Systems by Harnessing Ambient IoT Devices. AmbSentry: Mitigating Sensing Eavesdropping in ISAC Systems by Harnessing... Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling. This paper investigates whether on-policy distillation (OPD) truly expands the... Towards Scalable Fuzzy PSI via Efficient Fuzzy Matching. Based on the paper, here is a detailed summary: This paper presents scalable... One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL. Based on the paper, here is a detailed summary: Summary This paper identifies... Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting. Thyroid ultrasound diagnosis requires coordinated lesion localization,... MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents. MindMemOS is a portable and self-evolving memory operating layer for AI agents,... XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication. XB RIDGE: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication... Trie Automata for Constrained Decoding over Large Finite Sets. Trie Automata for Constrained Decoding over Large Finite Sets Authors: Xingzi... Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads. This paper introduces a unified theoretical framework for understanding... TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement. TailBooster is a dual-layer generative framework designed to address two... Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs. The paper introduces Self-Fix Step-DPO (SFS-DPO), a two-stage reinforcement... Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting. Motion-centric video reasoning is fundamental to interactive applications such... IoT-Enabled Autonomous Maritime Navigation in Smart Ports: A Curriculum-Guided Shared Policy Learning Framework. This paper investigates onboard autonomous navigation for IoT-enabled... Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents. Harness-IF is a benchmark that turns operational instruction following into a... When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use. This paper investigates a failure mode in multilingual API calling called... Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation. The paper studies reference-free post-training for multilingual machine... Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction. Based on the paper, here is the summary: This paper introduces DARC... Robust and Efficient Noisy-Label Time-Series Classification via Dynamic Time Warping Based Granular Ball Computing. Dynamic Time Warping (DTW)-based Nearest-Neighbor (NN) classifiers are... DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation. DexterSQL is a prompting-based (non-fine-tuning) Text-to-SQL system that... M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation. M-Net: Integrating Spectral Features and Physical Field Operators into Deep... Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning. The paper presents a multi-tier, frame-level audio tagging framework for... CT- Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models. CT-∆Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference... Sparse and robust geometric twin support vector machine via asymmetric RoBoSS loss function. This paper proposes a novel asymmetric, robust, bounded, sparse and smooth (aR)... MARCH: Scaling Recurrent Memory with Content-Routed State Anchors. MARCH: Scaling Recurrent Memory with Content-Routed State Anchors Abstract... Continual Learning in Transition. and Motivation This survey paper, authored by Zhiyan Hou, Dan Zhang, Tao Feng,... Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents. Here is a summary of the paper: The paper "Beyond Single-Turn Confidence:... When Do Anchor-Based Pointwise LLM Rerankers Help? Retriever Quality, Statistical Scope, and Anchor Design. and Research Questions This paper investigates when anchor-based pointwise LLM... Poly-Dialectal Neural Machine Translation System for Bangla Regional Dialects. This paper presents a unified Poly-Dialectal Neural Machine Translation System... JAPE: Joint Anomaly Prediction and Intrinsic Explanation in Multivariate Time Series. JAPE: Joint Anomaly Prediction and Intrinsic Explanation in Multivariate Time... HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry. HyperANFIS is a hyperbolic extension of the adaptive neuro-fuzzy inference... Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance. This paper investigates how prompt construction methodology affects the... RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation. RT-SEMamba is a fully causal speech enhancement (SE) model built upon causal... MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques. MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language... Coarsening Latent-Class Probabilities: Directional Distortion and Coverage Loss. Coarsening Latent-Class Probabilities: Directional Distortion and Coverage Loss... Rank-Two Frobenius-Linearized Normal Forms and Orthoderivative Dual Coordinates in Quadratic APN Maps. This paper classifies binary-linear two-term Frobenius-linearized operators of... Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks. This paper addresses the question of how large a random low-dimensional search... SchemaLink: An Intelligent Web Editor for LinkML Schema Curation. SchemaLink is a web-based environment for the graphical construction and... Learning Under Treatment-Induced Label Indeterminacy with Expert Annotations of Counterfactual Outcomes: A Case Study in Neurological Prognostication. This paper addresses the problem of treatment-induced label indeterminacy in... What Makes a Peer? Valuation-Anchored Similarity in Private Markets. The paper introduces a supervised similarity learning framework for identifying... SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries. SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries Abstract... @skills: Attention is all you have. Core Problem The paper identifies a fundamental mismatch in the agent skills... Intensional Anaphora. The paper "Intensional anaphora" by Ezra Keshet and Steven Abney (2024)... DIVE: Unlocking Self-Improvement in Frozen Language Models Through Diversity-Driven Skill Evolution. DIVE: Unlocking Self-Improvement in Frozen Language Models Through... Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts. This paper proposes a personalized scorer modeling framework for generating... Exploring Oversmoothing with Householder Matrices. Core Contribution This paper introduces HouseGNN (Householder Graph Neural... A Hybrid Framework of Vision Transformer and Gated Recurrent Unit for Detection of Mosquito Diseases. This paper introduces a three-step framework for detecting dengue and Zika... Don't Cut Corners: How Training Outside the Prior Makes Simulation-Based Inference More Robust. Large astrophysical simulation campaigns often generate training data by... A Remote Approach to Cashew Orchard Detection: Leveraging Active Learning with Satellite Imagery in Guinea-Bissau. This study presents the first openly accessible, countrywide cashew orchard map... A 12-CNOT Double Qubit Excitation Gate. The paper presents the first reported 12-CNOT decomposition of the double qubit... General Probabilities of Causation with Causal Knowledge. This paper addresses the question of whether additional causal knowledge can... Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues. Problem: Behavioral Relapse of Revoked Constraints This paper studies a failure... Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection. Problem Statement Deep learning detectors for anomalies in dynamic graphs have... Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction. Core Contribution The paper introduces Constraint Saturation Evaluation (CSE),... Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence. The Wiggle Framework is a unified stress test for epistemic stability in LLM... EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory. EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon... Novels generated by language models show compressed formal variation. This study investigates whether large language models (LLMs) can produce the... LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning. This paper introduces EpicStar, a framework for enhancing strategic reasoning... Supervised Mixed-Frequency Learning for Macro-Financial Forecasting When Factors are Weak. This paper, "Supervised Mixed-Frequency Learning for Macro-Financial... LookBack: Where and How to Score LVLM Responses via Visual Reference Usage. L OOK BACK is a training-free LVLM response scoring method that augments token... HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression. HAMP-LIC is a Hessian-aware mixed-precision post-training quantization (PTQ)... Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models. Core Contribution This paper introduces Semantic Lenia, a framework that... Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models. This paper investigates where safety-aligned refusal behavior is encoded in... Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages. This paper examines the structural barriers that disadvantage speakers of... A Conceptual Framework for Enhancing Workforce Readiness for Smart Manufacturing in the AI Era. This paper introduces the Workforce Readiness Level (WRL) framework, a... LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation. LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global... Machine Learning-Based Cyber Defense for Cloud Infrastructure: An Adaptive Deep Q-Network Architecture for Intelligent Intrusion Detection and Automated Threat Mitigation. This paper proposes a reinforcement learning-based dynamic cyber defense... Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability. This paper develops a mathematical and engineering architecture for secure... The energetic cost of mitigating AI attacks in cellular networks. The integration of Artificial Intelligence (AI), generally as Machine Learning... Learning from Online User Feedback for Shopping Agents. LOFA, a framework that enables shopping agents to learn directly from real... Slips: Behavioral Evidence Aggregation for Network Security. Slips is a network intrusion detection system that builds host-centered... Evaluating AlphaEarth Foundations Embeddings for Wildfire Susceptibility Mapping. This paper systematically evaluates AlphaEarth Foundations (AEF) embeddings for... Large Language Model-Driven Small-Capitalization Trading: Integrating Financial News Sentiment, Macroeconomic Indicators, and Technical Signals. This paper studies small-capitalization trading with LLM-derived news... Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation. and Motivation This paper investigates how large language model (LLM) agents... Plaintext Recovery Against Post-Filtering Access Control. This paper, "Plaintext Recovery Against Post-Filtering Access Control" by... A Cascaded Unsupervised-Supervised NLP Pipeline for Detecting Accusatory Language in Public Procurement. This paper proposes a cascaded hybrid learning strategy that integrates... Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems. This paper benchmarks the serving cost of agentic memory systems for... FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees. FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision... Autonomous Telerehabilitation via Skeletal Motion Prediction and Joint-Level Performance Assessment. This paper presents a telerehabilitation pipeline that integrates... Look What the Probes Dragged In! Real-World Chest X-ray Shortcuts in MedCLIP. The paper investigates how real-world shortcuts manifest across different... DCM Bandits: Multiplayer Information Asymmetric Cascading Bandits for Multiple Clicks. DCM Bandits: Multiplayer Information Asymmetric Cascading Bandits for Multiple... Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction. Accurate landmark localization in medical images is a fundamental step for... Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release. This paper presents a fully preregistered, end-to-end stress test of the... Can Vision Models Read the Radar Display? On the Feasibility of Radar Imagery for Air Traffic Complexity Estimation. This paper investigates whether radar display imagery, despite its atypical... Forward and Inverse Virtual Metrology for Phototransistor Gain: A Hierarchical, Uncertainty-Aware Approach for Small Production Datasets. The paper addresses the challenge of predicting and optimizing the current gain... Orientation, not magnitude: the causal structure of task-vector interference in merged language models. This paper investigates the causal structure of task-vector interference in... Kernel Methods for Learning Operators with Multiple Inputs and Outputs. This paper introduces a general kernel-based encoder-decoder framework for... Transferable Above-Ground Biomass (AGB) Estimation Model from Multi-Sensor Data with Sparse Field Calibration. The paper presents an operational framework for wall-to-wall above-ground... Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks. This paper introduces BMAT (Bilevel-Minimax Adversarial Transfer), a unified... Latent variable models for simultaneous EOV identification and removal in population-based SHM. The paper introduces a Gaussian-process latent EOV (GLEOV) model for... DYSANOS Generative Dynamic Smooth Arbitrage-free Non-parametric Option Surfaces. DYSANOS is the first generative market model for smooth SANOS option surfaces... RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning. Training Mixture-of-Experts (MoE) models for reinforcement learning (RL)... GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs. GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs... Clustered Randomized Smoothing for Stochastic Prediction Functions. Clustered Randomized Smoothing for Stochastic Prediction Functions proposes a... Inferential Capability Does Not Determine Legal Scope. The paper argues that the term "inference" performs two distinct legal... Calibration Bets on the Past: Post-Training Quantization for Financial Time-Series Forecasting. This paper presents a systematic study of activation calibration for... Asymptotic Risk Calibration for Selective Question Answering. The paper proposes A-CRC-QA, a post-hoc calibration framework for... LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining. LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for... LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification. LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and... GENADA: efficient generative time series adversarial attack framework. GENADA: efficient generative time series adversarial attack framework This... A Comparison of Malware Image Transformations Using Grad-CAM and Hybrid Learning Models. This paper investigates the use of Gradient-weighted Class Activation Maps... Analysis of Motor Signatures of Social Adaptation in Autism for Efficient Human-Centric Systems. This paper proposes a computational analysis framework to identify potential... Hybrid Gated Attention. The paper proposes a Hybrid Gated Attention (HyGA) framework to extend the... An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS. The paper proposes and evaluates a supervised agentic workflow for modernizing... FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting. FM-LLM (Frequency-Enhanced Mixture-of-Experts for adapting LLMs to Time Series... Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents. Convergent Detour Hijacking: Task-Preserving Resource Amplification in... GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings. GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in... Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence. This paper introduces Mechanist, an agentic framework that uses AI as a... Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection. The paper presents BENCH2ROBUST, a framework that converts failure-free... How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models. This paper addresses the challenge of efficiently using limited oracle budgets... A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench. This paper evaluates VITA, a retrieval-augmented generation (RAG) system... HCGRec: Hint-Conditioned Generative Recommendation with Semantic IDs. Problem Statement Semantic-ID generative recommenders represent each item as a... Accuracy and Order Sensitivity Diverge Under Label-Free Strategies. This paper investigates whether preventing a model from seeing option labels... Causal inference for group-contaminated structured outcomes: observable quotients, lossless reduction and exact randomization inference. This paper studies causal inference for structured outcomes (such as microscopy... Towards Model-based Run-time Cybersecurity: On Control-Flow Anomaly Detection, Attack Identification, and Hardware Monitoring. The paper "Towards Model-based Run-time Cybersecurity: On Control-Flow Anomaly... VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies. VAKRA (eValuating API and Knowledge Retrieval Agents) is a benchmark introduced... HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks. HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold... RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks. RealisticTritonBench is a benchmark for evaluating LLM-based Triton kernel... Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness Classification. SGNet (Spectral-Grouped Network) is a lightweight architecture for... Proportional Analogies on Probability Distributions via Bayesian Updating. The paper introduces a new notion of proportional analogy for probability... Consolidator: Learning Persistent Routed Memory Across Context Boundaries. The paper introduces Consolidator, a shared slot-local operator in a Phasor... Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated Research. The paper presents Nutrition Data Service (NDS), source-preserving... User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling. The paper proposes a collaborative distributed inference system that combines... Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing. The paper introduces HPSE (Hybrid-Policy Self-Editing), a method designed to... Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models. Better Slots, Better Worlds: Representation Quality & Robustness in... CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations. CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic... Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models. Self-Generative-Understanding (SGU): A Semantic Closed-Loop Evaluation... High-dimensional Multi-objective Bayesian Optimization with Learned Variable Interactions. This paper proposes ViaMOBO, a high-dimensional multi-objective Bayesian... Adversarial Resilience of Poisson-Process Submodular Maximization over Matroids: From Robust Offline Optimization to Full-Bandit Learning. This paper studies the problem of nonnegative submodular maximization subject... No One to Blame: A Framework of Constitutive AI Unaccountability. This paper introduces the concept of "constitutive AI unaccountability" to... ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models. ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language... Two-Stage Deformable-Convolutional Inverse Design of Nanophotonic Absorbers from Optical Spectra. This paper presents a two-stage deformable-convolutional framework for the... Instruction Alignment for Binary Code Representation Learning. This paper proposes InsnAlign, a training approach that leverages... Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning. This paper introduces the concept of trait-induced safety variation, a failure... Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents. Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents... Few-Shot Ordinal Learning for Day-Wise Freshness Estimation with Hyperspectral Fish Images. This paper introduces the first few-shot learning framework for hyperspectral... MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning. MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning proposes a... Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations. Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead... Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents. This paper presents a comprehensive empirical study of skill-induced agent... How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment. How China-Origin Vision–Language Models Move from Refusal to Reframing in... The Sleeping Agent: What Gist-Based Context Compression Loses and Why. The paper introduces Salience-Weighted Consolidation (SWC), a... Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation. This paper addresses the challenges of open-vocabulary instance segmentation... CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement. CLAIM: Leading Open-domain Active Clarification of Large Language Models with... Hierarchical Federated Transfer Learning in Digital Twin-Based Vehicular Networks. This paper proposes a novel Hierarchical Federated Transfer Learning (HFTL)... RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving Simulation. RoadWeaver is a coarse-to-fine framework for from-scratch generation of... SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward. SCOUT (Structured Chain-Of-Thought Utilizing Process-Supervised RL Training) is... From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models. MPAR-Bench: Evaluating Multi-Point Associative Reasoning in Large Language... Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse. CUE-Bench is a Chinese Unsaid Emotion benchmark that centers on Affective... Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization. The paper introduces a cost-aware method for evolutionary optimization of LLM... AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection. AgenticTwin is an agentic LLM framework integrated with a digital twin (DT) for... Policy-as-logic for robust reasoning over rules. Policy-as-Logic for Robust Reasoning over Rules proposes a hybrid symbolic... APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference. APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference... TradingMoE: Routing the Right Experts in Evolving Markets. TradingMoE: Routing the Right Experts in Evolving Markets proposes a... Regime-Gated Residual Mixture-of-Experts for Cross-Sectional Volatility Forecasting. This paper studies how regime information should be incorporated into a neural... Chain-of-Thought Shows the Path to a Tree: Realizing Branching Complexity. The paper introduces explicit, depth-bounded Chain-of-Thought (CoT)... Robust Ambiguity Detection (RAD) From Model- and Feature-Space Consistency. Robust Ambiguity Detection (RAD): From Model- and Feature-Space Consistency... Towards Truly Unsupervised Evaluation of Feature Selection. The paper "Towards Truly Unsupervised Evaluation of Feature Selection" by Hafiz... TESLA: Taylor Expansion of Sinusoidal Learnable Activations. TESLA: Taylor Expansion of Sinusoidal Learnable Activations Authors: Daehwa Ko,... Reducing Symmetry Increase in Equivariant Neural Networks. Equivariant Neural Networks (ENNs) have empowered numerous applications in... Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization. Epiplexity, a recently proposed measure of the structural information a... Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning. The paper "Drift and Dependence: Layer-wise Information-Theoretic Bounds for... Robustness of AI-Art Detectors under Generator Shift. This paper investigates the robustness of AI-art detectors under generator... FLARE++: Low-rank attention with dynamic attention routing. FLARE++: Low-Rank Attention with Dynamic Attention Routing Vedant Puri, Yongjie... Air Quality Station Simulation via LSTM and Attention-Based Modelling. The paper presents SATADL (SpAtial-Temporal Attention Dual LSTM), a... When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits. This paper presents a diagnostic protocol for selecting proxy rewards and... Fine-Tuning Generative Models for Extreme Events via CVaR-Penalized Wasserstein Gradient Flows. The paper proposes CVaR-GPA (CVaR-penalized Generative Particle Algorithm), a... Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling. This paper investigates whether Earth observation (EO) foundation model... RECAST: A Machine-Learning Framework for Correction and Super-Resolution of Coarse-Grid PDE Solvers. RECAST: A Machine-Learning Framework for Correction and Super-Resolution of... LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection. LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning... FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier Coupon Allocation. FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier... A Factor Graph Approach to Scalable Multi-Output Gaussian Process Regression. This paper presents a factor graph approach to scalable multi-output Gaussian... Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed. This paper presents a comprehensive evaluation of the trustworthiness of Small... Easper: An Accessible ASR Pipeline for Language Documentation. Easper: An Accessible ASR Pipeline for Language Documentation Abstract Audio... SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges. the Paper "SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic... Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages. Language-Conditional Dequantization (LCD) is a post-hoc method that attaches... SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents. SynWeaver is a website-prior task and trajectory co-synthesis framework... NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation. NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and... SoK: From Generation to Consumption of Privacy Documents in Software Systems. This SoK paper provides a unified, lifecycle-oriented view of privacy documents... Adaptive Bregman Proximal Stochastic Gradient with a Stabilized Barzilai--Borwein Step Size. Adaptive Bregman Proximal Stochastic Gradient with a Stabilized... Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach. This paper investigates cooperative target tracking in multi-AUV ad-hoc... HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting. This paper introduces Hugin, a training framework designed to enhance... Prof-K: Probabilistic One-Pass Filtering for Efficient Top-k Selection. Prof-K is a probabilistic one-pass filtering algorithm for efficient top-k... Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control. Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in... Is this Citation on Point?. This paper studies proposition-level citation support verification in legal... A comparison of CNN architectures for Alzheimer's disease detection in single-view MRI scans. This paper proposes a benchmark that evaluates ten different convolutional... CAM-Guided Saliency Cutout and Image-Based Malware Classification. This study evaluated saliency-guided cutout regularization for image-based... Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones. This paper investigates deep learning-based estimation of the Relative Transfer... Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion. The paper "Toward Meaningful Transparency for AI Chatbots: Disclosing... Do Judges Behave Like Algorithms?. This paper investigates whether judges already behave like algorithms in their... PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping. PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised... Which Site, and When: A Free-Satellite-Data Test of Himalayan Glacial Lake Bursts, Landslides, and Ice Floods. This paper tests whether free satellite data can predict glacial lake outburst... ComBodied Agents: a New Paradigm of Human-Centric Agentic AI. Combodied Agents are introduced as a human-centered paradigm of Agentic AI that... LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training. LazyTrain is an optimization-guided scheduler for limited-resource large... On Weak Bisimilarities in CCSK. In the context of CCSK, a reversible extension of CCS, we study different... QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving. QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG... Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs. The deployment of large language models (LLMs) in mental health contexts raises... A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery. This paper introduces NA-UNETR, a 3D transformer-based segmentation framework... ADEPT: A Unified Framework for Deep Learning Test Adequacy. ADEPT: A Unified Framework for Deep Learning Test Adequacy This paper presents... Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment. Generative Semantic Segmentation via an Observable Semantic-Image Interface and... Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers. This paper investigates the role of municipal algorithm registers in providing... The Boolean Power of ReLU. We prove that, on finite simple undirected graphs equipped with a single... Beyond Local Power: Functional Connectivity Analysis for Subject-Independent Learning Style Recognition. This paper proposes an objective Electroencephalography (EEG) approach... Dual-Model Sentiment Analysis of Consumer Reviews in the Retail Coffee Sector Using Machine Learning and Deep Learning Approaches. This study presents a comprehensive sentiment analysis framework applied to... From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection. This paper presents a closed-loop framework for video reflection removal,... Remote Sensing and Machine Learning-Based Analysis of Land Use and Vegetation Change in Dhaka District, Bangladesh. This study employs remote sensing data and machine learning techniques to... EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval. EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under... Attractor Image-Based Deep Learning of Arterial Pulse Waves for Age Classification. This study investigates whether arterial pulse waveform morphology, which... Structuring the Space of Perspectives. The same event can be reported from different perspectives depending on the... DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation. DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial... CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence. CAS: A Causal Attribution Score for Local and Global Explainable Artificial... A Quantum/Classical Example Oracle Separation for Making Things Up. The paper studies the power of quantum examples compared to classical examples... CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications. CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI... Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection. Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark... Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences. Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences... Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study. Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case... A Local-Linearly Convergent Algorithm for Nonconvex Equality-Constrained Optimization. The paper extends the analysis of the Gradient-Eigenstep Algorithm by Goyens et... HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment. HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar... Policy Convergence and Divergence Across National and Within Regional AI Strategies: A Policy Design Element Analysis. Governments worldwide have responded to the rapid expansion of AI by publishing... Not All Nudges Land: Behavioral Controllability and Elaboration Quality in AI-Supported Journaling. This paper analyzes 369 journal entries from an eight-week passive sensing... Geometric and Behavioral Stratification in Transformer Residual Streams. The paper investigates the geometric and behavioral organization of transformer... GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation. Generating actionable financial advice from business records demands that... ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering. ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in...

The papers