AI papers — 2026-08-13

Governing Agentic AI in FinTech. This paper addresses the governance challenges posed by agentic AI systems in... Keep the Future, Drop the Rollout: RIFT for World Action Models. World action models (WAMs) condition robot actions on predicted futures, but... Localize, Then Reason: Visual Latent Structural Reasoning for Molecular Properties and Edits. The paper introduces Visual Latent Structural Reasoning (VLSR), an end-to-end... OEIS Open: How many conjectures can language models turn into theorems?. OEIS Open: How many conjectures can language models turn into theorems? Tom... CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility. CoMedBench is a reproducible benchmark for evaluating synthetic medical data,... Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians. Hamilton-Zero is a foundation model for computing ground states of arbitrary... Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models. The paper addresses the challenge of predicting answers to interventional "what... Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses. DECAF: Decomposition of Evidence, Contradiction, and Fragility in Perturbation... Beyond Parameter Space: NTK-Guided Personalized Aggregation for Robust Federated Learning. LIGHTYEAR is a federated learning (FL) framework that performs update selection... Equivariant learning of a transferable three-dimensional classical density functional. Equivariant learning of a transferable three-dimensional classical density... A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family. The paper addresses a fundamental weakness in how GPU kernel generation systems... Synthetic Persona Pretraining: Alignment from Token Zero. and Motivation The paper introduces Synthetic Persona Pretraining (SPP), a... Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws. Core Contribution This paper introduces the Neural Quadratic Form (NQF), a... EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding. EEG-PRIME is a two-stage EEG foundation model for cross-dataset multi-task EEG... Demand Transfer Estimation at Scale via Restricted Logit Modeling. Item demand forecasting is an integral component of store assortment... Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents. Core Research Question and Motivation This paper investigates whether... ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models. ContactGuard is a pre-contact execution monitoring system for chunked... TANGCO: Learning Topology-Aware Capacity Allocation for Overload-driven Cascading Failures. TANGCO: Learning Topology-Aware Capacity Allocation for Overload-driven... Homomorphic Aggregation of Continuous-Variable GKP States. Homomorphic aggregation of logical quantum information encoded in... Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations. Class activation mapping (CAM) is one of the most widely used visual... Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services. This paper studies a large language model (LLM) service in which a provider... Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development. This paper presents a systematic evaluation framework for long-horizon AI... Foundations of MT-PDCL: Measure-Theoretic Probabilistic Definite Clause Logic. Measure-Theoretic Probabilistic Definite Clause Logic (MT-PDCL) is a... Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation. Where You Measure Decides What You Measure: Position Selection in... Active-Trace Complexity Bounds for Moreau--Yosida Unadjusted Langevin Sampling. Problem and Setting This paper studies the Moreau–Yosita unadjusted Langevin... CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation. CABS+ is an enhanced model merging framework that extends the Conflict-Aware... From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion. Diffusion models have achieved dominant performance in visual generation but... LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning. Long-horizon Earth observation reasoning requires models to organize... AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1). AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report... DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models. This paper introduces DreOPD (Degraded-reference extrapolative On-Policy... History-informed Lagrangian Neural Networks. History-informed Lagrangian Neural Networks (HiLNN) is introduced to address... RealmEye: Virtual Machine Introspection for Arm CCA Realm VMs. RealmEye is the first Virtual Machine Introspection (VMI) system for Arm CCA... CAPRI: Contract-Aware Proof Repair for Isabelle. CAPRI: Contract-Aware Proof Repair for Isabelle Abstract. Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation. This paper introduces a protocol-level identifiability audit for LLM... Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information. Core Contribution The paper proposes "Rule of Thumb" (ROT) explanations, a new... Exponential quantum advantage for learning signals with a single qubit. The paper demonstrates that coupling a single controllable qubit to a... Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval. The paper "Heterogeneous Vision-Language Ensemble with Disagreement-Aware... ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs. ARAC: Benchmarking Auto-Research’s Alignment and Completeness on End-to-End... TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability. TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer... SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data. SPARED introduces an adversarial reinforcement learning framework for... How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures. SciFigBench is a diagnostic vision-language model (VLM) benchmark for... Vero: Can AI Agents Build Formally Verified Software Repositories?. Vero is the first benchmark to evaluate joint implementation and proof... BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving. BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics... Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity. The paper "Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and... AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models. AaLLM is an open-source, end-to-end multi-agent LLM workflow for analog circuit... Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference. This paper introduces Reduced Matrix Multiplication (RMM), a training-free,... Training AI Scientists to Replicate Research. This paper introduces Replica, a scalable task space for paper replication, and... TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint. TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic... StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems. StateBridge is a training-free latent communication approach for large language... CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport. CoverPrune is a training-free token pruning framework for 3D Vision-Language... Bagging Robustly Learns VC Classes with Linear Sample Complexity. This paper proves that VC classes are adversarially robustly learnable with... Sinkhorn Linearization and the Spectral Proxy: Unifying the Statistical and Algorithmic Theory of Feature-Parameterized Inverse Optimal Transport via a Single Spectral Sandwich. Core Contribution The paper develops the statistical and algorithmic theory of... Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization. SNIPER is a two-stage structured pruning framework for large language models... Momentum as Residual-Driven Multiplier Correction for Deep Learning Optimization. The paper proposes an ADMM-Inspired Momentum (AIM) framework based on... A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings. This paper introduces a cloud–edge collaborative architecture for multimodal... Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents. This paper introduces HARD (Harness-based Autonomous Runtime Defense... Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks. This paper introduces Cocktail, a watermarking scheme for large language models... Symmetry-Breaking De Novo Crystal Generation via Markovian Jump Diffusion. The paper introduces SbCD (Symmetry-breaking Crystal Diffusion), a novel... Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories. The paper identifies a critical bottleneck in agent memory systems: while... PIPES: Securing Agent Perception with Provenance and Priors. PIPES: Securing Agent Perception with Provenance and Priors introduces a... ProME: Prototype-Margin Environments with Repair-Aware Selection for Group-Robust Learning. ProME: Prototype-Margin Environments with Repair-Aware Selection for... Fine-tuned Normalizing Flows for ALICE Zero Degree Calorimeter Fast Simulation. Simulating the ALICE Zero Degree Calorimeter (ZDC) neutron detector responses... Federated Compositional Muon Optimizer for Matrix-Wise Models. Federated Compositional Muon Optimizer for Matrix-Wise Models Authors: Wang... CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation. CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation... QuoteBench: How Matched Scores Can Hide Command-Path Failures. QuoteBench: How Matched Scores Can Hide CommandPath Failures Core Problem and... Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents. Here is a summary of the paper: The paper "Beyond Outcome Rewards: Step-Level... OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways. OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for... FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving. FlashDrive is an algorithm-system co-design framework that targets all four... AQuA: Recursively Self-Improving Quantitative Trading Research Agents. AQuA: Recursively Self-Improving Quantitative Trading Research Agents Summary... Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents. The paper introduces CREST (Hierarchical Credit Assignment via Entropy-Gated... VALG: An Agentic System for ML Theory Research. VALG is an agentic system for machine learning theory research that combines... NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents. This paper presents NaviDC-OCR, a unified document parsing framework designed... ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval. ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval Abstract While... Chance-constrained selection of sequential intervention strategies from counterfactual estimates. Problem and Motivation This paper addresses the challenge of selecting... Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement. This survey paper, "Numeracy in Large Language Models: Fundamental Limitations... On the Expressive Power of Transformers. This survey paper provides an overview of the expressive power of transformers,... EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory. EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal... Into the ORBIT for Time Series: Training Regimes for Foundation Models. This paper introduces ORBIT (Omni-Range Bootstrap Incremental Training), a... Exponential Convex Calibration Dimension for the Multi-Label Jaccard Measure. Exponential Convex Calibration Dimension for the Multi-Label Jaccard Measure... Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research. This paper presents a deployment-oriented two-tier agentic system that... LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation. LYCHEE MEMORY V2 is an efficient long-term memory framework for LLM agents that... RippleMem: From Isolated Retrieval to Associative Recollection for Long-Term Agent Memory. RippleMem is a long-term memory system for LLM-based agents that shifts memory... EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval. EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval... Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing. Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth... Algebraic Decomposition Theory for Transformer Length Generalization. This paper establishes the first complete characterization of which regular... HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark. HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark... Latent On-Policy Self-Distillation. The paper addresses the challenge of enabling agents to learn from experience... Virtual Temperature Sensors in Power Transformers Using Neural Ordinary Differential Equations. This paper presents a physics-aware Neural Ordinary Differential Equations... DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees. DARTree is a training-free speculative decoding method that extends a... High-resolution Calibrated Probabilistic Hourly Precipitation from a Deterministic Forecast. This paper describes an “Attention Residual U-Net” method for probabilistic... Technical Report on Resilient and Secure Large-Scale Energy Internet Systems. This IEEE PES Task Force report examines the security and resilience of... Intern-S2-Preview: Scientific Agentic Foundation Model. Intern-S2-Preview is a series of scientific agentic foundation models designed... OmniScientist: An Omni-Modal Omni-Discipline AI Scientist. OmniScientist: An Omni-Modal Omni-Discipline AI Scientist Core Contribution The... InSPECtor: Improving SLEIGH Processor Specification Veracity via Proxy. InSPECtor: Improving SLEIGH Processor Specification Veracity via Proxy Summary... MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning. MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning Summary... LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection. This paper presents a comprehensive benchmark evaluating whether Low-Rank... Falsehood and Impossibility Are Different Directions in an AI's Representation of Language. This paper reports an exploratory activation study of the multimodal... Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning. Critic-Free Pretraining (CFP) is an efficient paradigm for offline-to-online... CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives. CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical... EEG Decoding Using CNN and LSTM Network. This study introduces a hybrid deep-learning architecture that integrates a... Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy. This paper presents a scalable, context-aware Multi-Agent Framework designed to... DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition. DMDINTEL: Interpreting Large Language Models via Dynamic Mode Decomposition... When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation. Based on the paper, here is a detailed summary: The paper investigates whether... Adversarial Robustness in Smishing Detection: A Comparative Analysis of Adversarial Fragility in Classical vs. Transformer-Based Detection Systems. This study evaluates the adversarial robustness of five model architectures for... RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation. RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in... LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation. LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea... HybridRAG-BN: A Retrieval-Augmented Framework with Fine-Tuned Verification for Bangla KBQA. HybridRAG-BN is a retrieval-augmented framework for Bangla knowledge-base... HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models. HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large... Correct Is Not Governed: Provenance Integrity in Agentic Workflows. Core Thesis: The paper defines "governed execution" as distinct from task... Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering. This paper introduces Syfer, a synthesizer-folding framework for multilingual... Understanding Backdoor Vulnerabilities in Vertical Federated Learning: The Gap Between Research and Practice. This paper, "Understanding Backdoor Vulnerabilities in Vertical Federated... Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors. Dual-Stream Cross-Anchor Correction (DSCC) is a fine-tuning framework proposed... Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language. Static analysis-guided agentic AI translation enables Rust as a full stack... Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence. This paper investigates the validity of the conditional-independence assumption... Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs. The paper introduces the Sensitive Entity Alias Generator (SEAG), a... FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation. FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative... A Probe Direction Is a Property of Its Prompt. Core Claim The paper argues that the standard instrument used to detect whether... Branch and Bound for Relational Verification of Neural Networks. Branch and Bound for Relational Verification of Neural Networks Authors: Kota... Finding the Needle in a Haystack: Test-Time Analog Circuit Representation Adaptation for Bayesian Optimization. Test-Time Analog Circuit Representation Adaptation for Bayesian Optimization... Adaptive Nearest Neighbors Classifier via Granular Ball Computing. The paper proposes an adaptive and efficient k-Nearest Neighbor (KNN) approach... Prompts in the Wild: A Large Analyzed Collection of Transactional Prompts in Code. Here is a detailed summary of the paper "Prompts in the Wild: A Large Analyzed... Huber-Wasserstein barycenters for robust distribution-valued data. Huber–Wasserstein barycenters for robust distribution-valued data Authors:... CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers. CW-BASS v2 is a saturation-aware pseudo-label selection method for... Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging. Mr3D-VL is a dedicated visual-language foundation model for multi-parametric 3D... Sparse Orthogonal Regression Technique: A Spectral Framework for Equation Discovery, Approximation, and Integration. The paper develops the Sparse Orthogonal Regression Technique (SORT), "a sparse... Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry. This paper studies decentralized multi-player reinforcement learning in... DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data. DFM Mimir v1 is a 1-billion-parameter language model based on the Hierarchical... Error-Aware Reverse Auction Mechanism for Large Language Model Routing. Error-Aware Reverse Auction Mechanism for Large Language Model Routing Authors:... Discovering Persistent Behavioural Patterns for Interpretable Blockchain Forensics. This paper proposes a scalable, application-agnostic framework for persistent... High-dimensional networks and mean squared error for possibly misspecified models. The paper "High-dimensional networks and mean squared error for possibly... Thermodynamics of Learning: A Typed Four-Component Accounting of Memory, Fit, and Value. Here is a summary of the paper, constructed from direct quotes and detailed... Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia. PRISM (Perturbation-based Regional Interpretability through Subtraction... Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection. This study investigates whether self-supervised learning (SSL) speech... Physics-informed distribution of relaxation times estimation and latent-space condition monitoring of solid oxide fuel and electrolysis cells from electrochemical impedance spectroscopy. The paper proposes a physics-informed convolutional neural network (CNN)... Generative Universal Multimodal Retrieval with Dual-role Identifiers. The paper proposes DrIG, a novel generative framework for universal multimodal... MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination. MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and... Distribution Steering via Sliced Optimal Transport Control. Distribution steering seeks feedback laws that drive the state law of a... Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings. ParliamentRAG is a Retrieval-Augmented Generation (RAG) system designed for the... A Unifying Perspective on Causal World Models: From Observations to Representations to Structure. This paper studies world models (WMs) from a causal perspective across multiple... Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation. This paper audits a preserved MCP (Model Context Protocol) agent security... TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems. TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death... GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport. GeoCache is a training-free acceleration method for multi-view texture... LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service. LipCache: A Local Inference Proxy with Certified Caching for Edge Image... Operationalizing Cyber Threat Intelligence with GraphRAG. This paper investigates whether using a knowledge-graph-based retrieval system... Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling. This paper introduces a novel proxemics-based reward formulation for deep... ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification. ReflectFact is a novel self-reflective agent framework for multi-hop fact... Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models. Self-referential prompting has been shown to reliably induce large language... AI and Consumer Rights in India Working Paper. This working paper examines whether India's Consumer Protection Act, 2019,... Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors. This paper proposes a controllable method to erase copyrighted animation... TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval. TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval... Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation. Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in... UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations. UniTraffic-Agent is introduced as the MR-CAS solution for Track 3 of the 10th... SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback. SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction... Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds. This paper challenges the default paradigm of aligning large language models... BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs. BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through... MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification. The paper proposes ARMDIL, an Adaptive Router for Multi-Domain Image... Sovereign by necessity? Frontier AI export controls, cyber security, and the limits of national AI capability. The most capable artificial intelligence (AI) frontier models are produced by a... Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence. Spatial Memory Agent (SMA) is an experience-grounded runtime framework that... Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting. This paper proposes a unified evidence reasoning framework that addresses two... Slow and Steady: Preventing MEV with Verifiable Delays. This paper presents a defense mechanism against Maximal Extractable Value (MEV)... Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology. The paper proposes an intervention-aware clinical world model for forecasting... LOB-ID: Evaluating Synthetic Market Data by Inception Distances. The paper introduces LOB-ID (Limit Order Book Inception Distance), an... vToken: Token-Level Virtualization for Reclaimable KV Caches. vToken is a lightweight token-level virtualization layer that decouples logical... Comment on "Modeling rapid language learning by distilling Bayesian priors into artificial neural networks". This comment paper argues that McCoy & Griffiths’ (M&G) method of using... Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining. Core Contribution This paper proposes a task-agnostic measure of training data... From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options. Large language models (LLMs) perform well across a wide range of tasks but... PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs. This paper introduces PROVE-RT, an LLM-assisted framework for generating... Which LLM Is Your Ideal Companion? Evaluating Emotional Companion Capabilities of LLMs Based on Adult Attachment Theory. Based on the paper "Which LLM Is Your Ideal Companion? Evaluating Emotional... FlowLOB: Efficient and Controllable Limit Order Book Generation with Flow Matching. FlowLOB is a conditional flow-matching generative model for limit order book... Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies. Large Language Models (LLMs) are increasingly deployed in discovery domains... A Multispectral Framework for the Detection of Calcium Carbide-Induced Ripening and Shelf-Life Estimation in Climacteric Fruits. The proposed study explores a novel, non-invasive multispectral framework for... PatientAct: Theory-Grounded Mental Health Client Simulation. PATIENTACT: Theory-Grounded Mental Health Client Simulation This paper... On the Structural Limits of Machine Learning Decision Systems: An Information-Theoretic, Interaction-Based, and Stochastic-Dynamical Perspective. This paper examines the structural limits of machine learning decision systems... LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles. LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses... When Should Multi-Round RAG Stop? Structured Stopping Judgments and Retrieval Reduction in Search-R1. Core Contribution and Setup This paper addresses the stopping problem in... Foundations of Independent Component Analysis. This paper presents a comprehensive, self-contained mathematical treatment of... Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks. This paper presents a heterogeneity-aware belief synchronization framework for... MergeOver: Post-Training Token Merging for Recursive Vision Transformers. MergeOver is a post-training approach that integrates Token Merging (ToMe) into... NAS-Driven Hardware Accelerator Exploration for Edge AI and Quantization Effects on the Pareto Space. This paper proposes a three-stage hardware-aware Neural Architecture Search... Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data. The paper "Novel Knowledge-Guided Generative Methods for Synthetic... Making AI-Generated Feedback Matter: A Large-Scale Study of Feedback Workflows and Student Enactment. This study examined how different AI-mediated feedback workflows were... Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization. This paper investigates whether instance-wise allocation of a fixed test-time... SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference. SPADE: Speculative Decoding for Precise and Low-Cost Distributed Edge–Cloud... AutoQuREO: A Framework for Automated Quantum Resource Estimation and Optimization. AutoQuREO is an automated framework for full-stack quantum resource estimation... Revisiting Overestimation Bias Problem of Q-learning: Settling Large Discrete Action Space via Action Intersection. This paper considers the overestimation bias problem of Q-learning in the... Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization. The paper "Doubly Robust Estimation of Causal Effect on CVR with Targeted... The Time Value of Evolution. The paper formalizes the concept of the "time value of evolution," which... On the global feature importance for interpretable and trustworthy heat demand forecasting. The paper introduces an ante-hoc Explainable AI (XAI) methodology to assess the... Smart Contract Invariants Protect Against Cybercriminals. This paper investigates whether smart contract invariants can protect against... Beyond Source: An Empirical Study of Python Bytecode Security Risks. This paper presents an empirical study of Python bytecode as a security... Robust data-driven discovery of fractional differential equations via weak formulations and Pareto-based subset selection. The paper "Robust data-driven discovery of fractional differential equations... The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity. The paper introduces the unmasking growth complexity (UGC) as a path-resolved... InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers. InFactPlanner is a trace-driven decision-support framework for what-if analysis... Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents. This paper introduces the concept of skill misevolution in self-improving LLM... Concept Drift Detection and Adaptive Retraining of Malware Classification Models. Concept drift refers to changes over time in the statistical properties of... Jointly Predicting Courses and Grades Using a Transformer-Based Model. This paper introduces a TRansformer for Academic Course-grade Estimation... Rules or Character? Scaling Laws for AI Safety Design. This paper introduces a stylized comparative-statics model to analyze the... Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity. The paper investigates how instruction tuning affects model confidence and the... Deliberate Practice: Learning Robot Skills under a Budget. The paper "Deliberate Practice: Learning Robot Skills under a Budget" by Shivam... PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR. PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in... It's How You Ask: Gender-Associated Linguistic Bias in LLMs. The paper "It's How You Ask: Gender-Associated Linguistic Bias in LLMs" by... Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety. This paper proposes Wrapper-Based Intent-Form Augmentation (WIFA), an automatic... Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales. Normative datasets are often used to train and align AI systems, but the norms... Task- and dataset-specific information in protein language models. This paper investigates the internal representations of protein language models... SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents. SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in... Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI). This study tested whether a language model's explanatory engagement with a rare... Tight Nonasymptotic Local Convergence of Sinkhorn-Knopp. The paper "Tight Nonasymptotic Local Convergence of Sinkhorn-Knopp" revisits... The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use. This paper investigates why latent world models fail at long-horizon planning,... H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities. This paper presents the first comprehensive adaptation and evaluation of... A Compositional Theory of Curvature in Probabilistic Circuits. Core Contribution This paper introduces a compositional theory of curvature in... Memorization Diagnostics for Code LLMs Should be Scale-Aware. The paper investigates whether standard memorization diagnostics for code LLMs... HybridSB-MoE: Dual-Domain Schrödinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement. HybridSB-MoE is a dual-domain framework for speech enhancement that combines a... Knowledge-guided Pattern Discovery via Coupled Tensor Factorizations. This paper introduces a knowledge-guided approach for pattern discovery that... TeleGapper: On the (un)reliability of Privacy Policies in Telegram Mini apps. TeleGapper: On the (un)reliability of Privacy Policies in Telegram Mini apps... OmniSphinx: Active Mix Networks (Extended Version). OmniSphinx is a novel mix format that applies ideas from active networking to... Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency. and Motivation The paper addresses a critical gap in evaluating Joint-Embedding... Dissecting Software Graphs: Structural Insights for Driver-Guided Fuzzing. This paper presents an empirical study of software structure under multi-driver... Full-Key Recovery and Forgery from One MQOM v2.1 Signature. This paper presents a full-key-recovery attack on MQOM v2.1, a Round-3... SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization. SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via... TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies. TopoIntent is a system that compiles natural-language security intent into... Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks. The paper "Large-scale Testing Global Optimization Methods with Black-box... Distinguishing case-mix from context heterogeneity in prognostic regression model synthesis settings. Distinguishing case-mix from context heterogeneity in prognostic regression... Defensive Boosting for Online Probabilistic Forecasting. Problem Setting and Motivation The paper studies online probabilistic... NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video. NARU is a benchmark designed to evaluate narrative evolution and cultural... I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization. I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization Summary... A Commitment-Based Hybrid Post-Quantum Cryptographic Model for Multi-File Cloud Storage. This paper presents a commitment-based hybrid post-quantum cryptographic model... Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning. the Paper Problem Statement and Motivation The paper addresses the fundamental... Statistical Properties of Robust Learning under Distributional Shifts. and Motivation This paper studies the statistical properties of robust learning... Incremental Evaluation and Training in Relational Deep Learning. Incremental Evaluation and Training in Relational Deep Learning Authors: Jakub... CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation. CardioState-JEPA is a cardiac foundation model that learns a single shared... Difference-of-Convex Regularization for Graph Learning by Differentiable Programming. This paper proposes a Difference-of-Convex Regularizer (DCR) graph learning... ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization. ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization Lixing Li,... The Impact of Temporal Context Length and Encoding Strategies on Self-Supervised ECG Representation Learning. The paper investigates how temporal context length and encoding strategies... Decoupled Contrastive Decoding via Expert-Aligned Drafting. Decoupled Contrastive Decoding (DCD) is a method to accelerate Contrastive... VR-Themis: A Scalable Framework for Virtual Reality Application Clone Detection. VR-Themis: A Scalable Framework for Virtual Reality Application Clone Detection... Balanced Adaptive Prototype Selection for Scalable TabPFN Inference on Large-Scale Tabular Data. This paper introduces Balanced Adaptive Prototype Selection (BAPS), a framework... MBA: Multimodal Benchmark and Agents for Real-World Business Ideation. the Paper "MBA: Multimodal Benchmark and Agents for Real-World Business... UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models. UniTexture is a cross-task universal adversarial texture attack that uses a... Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs. The paper "Beyond Visual Evidence: Revealing and Mitigating Relational Privacy... CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment. CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference... Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge. The Information Abundance Paradox, as proposed in this paper, hypothesizes that... REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation. On-policy distillation (OPD) trains a student on its own generated trajectories... AI Guardrail Survival under Single-Cycle Agentic Self-Summarization. This paper investigates how a standing safety rule is lost during a... Foundation models for movement data: Are they ready for prime-time?. Foundation models (FMs) trained on large-scale accelerometer data have been... Designing AI Pipelines for Decision-Ready ITSM Intelligence. This paper reframes ITSM data use as an IS problem of transformation,... CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation. CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice... On the Importance of Geometric Nonlinearity and Temperature-Dependent Properties in Multi-Material Thermo-Mechanical Topology Optimization. This paper investigates two common simplifying assumptions in thermo-mechanical... Mawqif-XT: An Arabic Benchmark Dataset for Cross-Target Stance Detection. Mawqif-v2 is an extension of the original Mawqif dataset, designed as a... Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models. The paper investigates whether output homogeneity in language models originates... ATOBench: Tracing How Autonomous Penetration-Testing Agents Verify Vulnerabilities When Target Evidence Lies. ATOBench introduces an evaluation framework that makes the verification process... AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design. AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Core... Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes. This paper explores the use of Small Language Models (SLMs) to support... LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure. LittleLearner: Language Models Under Pedagogically Controlled Knowledge... Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence. The paper introduces a Polish-language medical visual question answering (VQA)... The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis. The paper proposes a YOLO- and Contrastive Language-Image Pre-training... Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test. The paper introduces a finite capability sheaf to model failures in AI agent... Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension. The paper presents the organizational framework adopted by the Academic League... TabSOM: A tabular-to-image encoding method based on self-organizing maps. TabSOM: A tabular-to-image encoding method based on self-organizing maps... Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks. and Motivation This paper, authored by Taha Shieenavaz, Shabnam Zareshahraki,... iARCS: Iterative Agentic RL for Controllable 3D Scene Generation. iARCS: Iterative Agentic RL for Controllable 3D Scene Generation Overview The... Multi-perspective Imbalance-Conscious 6G Beamforming Optimization and Performance. The study presents a systematic machine learning (ML) study of 6G-IoT... Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model. Mixture of Training (MoT) is a scaffolded modular pre-training procedure that... Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision. This study presents the first use of egocentric vision for freezing of gait... Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs. This preliminary technical report presents a framework for sign language video... EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction. EGRL: Edge generation-guided relation-aware learning for RNA-protein... RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level. This paper addresses the challenge of assessing the maturity of artificial... Learning the Mathematical Property for Designing Low Mutual Coherence Binary Sensing Matrices. This paper proposes a novel learning-based framework for constructing binary... CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model. CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large... Beyond Simulated Benchmarks: Evaluating Motion Representations for Fall Detection Under Real-World Data Scarcity. Falls are a major health concern for older adults, and wearable sensors have... Moose: Latent concept learning with reasoning-shortcut awareness in. Moose: Latent concept learning with reasoning-shortcut awareness in EL++... Multi-Layer Context Camouflaging: A Semantic Superposition and Contextual Lamination Framework for Malpractice-Resilient Online Assessment. This paper extends the Multi-dimensional Spatio-Temporal Context Camouflaging... Uniform Herding: Exemplar Replay with Representation Refresh. Uniform Herding: Exemplar Replay with Representation Refresh proposes a... Efficient Hessian-Free Methods for Multi-Objective Bilevel Optimization with Nonconvex Lower Level. the Paper Problem Statement This paper addresses multi-objective bilevel...

The papers