AI papers — 2026-08-13
Governing Agentic AI in FinTech. This paper addresses the governance challenges posed by agentic AI systems in... Keep the Future, Drop the Rollout: RIFT for World Action Models. World action models (WAMs) condition robot actions on predicted futures, but... Localize, Then Reason: Visual Latent Structural Reasoning for Molecular Properties and Edits. The paper introduces Visual Latent Structural Reasoning (VLSR), an end-to-end... OEIS Open: How many conjectures can language models turn into theorems?. OEIS Open: How many conjectures can language models turn into theorems? Tom... CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility. CoMedBench is a reproducible benchmark for evaluating synthetic medical data,... Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians. Hamilton-Zero is a foundation model for computing ground states of arbitrary... Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models. The paper addresses the challenge of predicting answers to interventional "what... Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses. DECAF: Decomposition of Evidence, Contradiction, and Fragility in Perturbation... Beyond Parameter Space: NTK-Guided Personalized Aggregation for Robust Federated Learning. LIGHTYEAR is a federated learning (FL) framework that performs update selection... Equivariant learning of a transferable three-dimensional classical density functional. Equivariant learning of a transferable three-dimensional classical density... A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family. The paper addresses a fundamental weakness in how GPU kernel generation systems... Synthetic Persona Pretraining: Alignment from Token Zero. and Motivation The paper introduces Synthetic Persona Pretraining (SPP), a... Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws. Core Contribution This paper introduces the Neural Quadratic Form (NQF), a... EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding. EEG-PRIME is a two-stage EEG foundation model for cross-dataset multi-task EEG... Demand Transfer Estimation at Scale via Restricted Logit Modeling. Item demand forecasting is an integral component of store assortment... Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents. Core Research Question and Motivation This paper investigates whether... ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models. ContactGuard is a pre-contact execution monitoring system for chunked... TANGCO: Learning Topology-Aware Capacity Allocation for Overload-driven Cascading Failures. TANGCO: Learning Topology-Aware Capacity Allocation for Overload-driven... Homomorphic Aggregation of Continuous-Variable GKP States. Homomorphic aggregation of logical quantum information encoded in... Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations. Class activation mapping (CAM) is one of the most widely used visual... Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services. This paper studies a large language model (LLM) service in which a provider... Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development. This paper presents a systematic evaluation framework for long-horizon AI... Foundations of MT-PDCL: Measure-Theoretic Probabilistic Definite Clause Logic. Measure-Theoretic Probabilistic Definite Clause Logic (MT-PDCL) is a... Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation. Where You Measure Decides What You Measure: Position Selection in... Active-Trace Complexity Bounds for Moreau--Yosida Unadjusted Langevin Sampling. Problem and Setting This paper studies the Moreau–Yosita unadjusted Langevin... CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation. CABS+ is an enhanced model merging framework that extends the Conflict-Aware... From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion. Diffusion models have achieved dominant performance in visual generation but... LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning. Long-horizon Earth observation reasoning requires models to organize... AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1). AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report... DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models. This paper introduces DreOPD (Degraded-reference extrapolative On-Policy... History-informed Lagrangian Neural Networks. History-informed Lagrangian Neural Networks (HiLNN) is introduced to address... RealmEye: Virtual Machine Introspection for Arm CCA Realm VMs. RealmEye is the first Virtual Machine Introspection (VMI) system for Arm CCA... CAPRI: Contract-Aware Proof Repair for Isabelle. CAPRI: Contract-Aware Proof Repair for Isabelle Abstract. Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation. This paper introduces a protocol-level identifiability audit for LLM... Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information. Core Contribution The paper proposes "Rule of Thumb" (ROT) explanations, a new... Exponential quantum advantage for learning signals with a single qubit. The paper demonstrates that coupling a single controllable qubit to a... Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval. The paper "Heterogeneous Vision-Language Ensemble with Disagreement-Aware... ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs. ARAC: Benchmarking Auto-Research’s Alignment and Completeness on End-to-End... TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability. TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer... SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data. SPARED introduces an adversarial reinforcement learning framework for... How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures. SciFigBench is a diagnostic vision-language model (VLM) benchmark for... Vero: Can AI Agents Build Formally Verified Software Repositories?. Vero is the first benchmark to evaluate joint implementation and proof... BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving. BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics... Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity. The paper "Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and... AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models. AaLLM is an open-source, end-to-end multi-agent LLM workflow for analog circuit... Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference. This paper introduces Reduced Matrix Multiplication (RMM), a training-free,... Training AI Scientists to Replicate Research. This paper introduces Replica, a scalable task space for paper replication, and... TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint. TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic... StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems. StateBridge is a training-free latent communication approach for large language... CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport. CoverPrune is a training-free token pruning framework for 3D Vision-Language... Bagging Robustly Learns VC Classes with Linear Sample Complexity. This paper proves that VC classes are adversarially robustly learnable with... Sinkhorn Linearization and the Spectral Proxy: Unifying the Statistical and Algorithmic Theory of Feature-Parameterized Inverse Optimal Transport via a Single Spectral Sandwich. Core Contribution The paper develops the statistical and algorithmic theory of... Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization. SNIPER is a two-stage structured pruning framework for large language models... Momentum as Residual-Driven Multiplier Correction for Deep Learning Optimization. The paper proposes an ADMM-Inspired Momentum (AIM) framework based on... A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings. This paper introduces a cloud–edge collaborative architecture for multimodal... Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents. This paper introduces HARD (Harness-based Autonomous Runtime Defense... Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks. This paper introduces Cocktail, a watermarking scheme for large language models... Symmetry-Breaking De Novo Crystal Generation via Markovian Jump Diffusion. The paper introduces SbCD (Symmetry-breaking Crystal Diffusion), a novel... Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories. The paper identifies a critical bottleneck in agent memory systems: while... PIPES: Securing Agent Perception with Provenance and Priors. PIPES: Securing Agent Perception with Provenance and Priors introduces a... ProME: Prototype-Margin Environments with Repair-Aware Selection for Group-Robust Learning. ProME: Prototype-Margin Environments with Repair-Aware Selection for... Fine-tuned Normalizing Flows for ALICE Zero Degree Calorimeter Fast Simulation. Simulating the ALICE Zero Degree Calorimeter (ZDC) neutron detector responses... Federated Compositional Muon Optimizer for Matrix-Wise Models. Federated Compositional Muon Optimizer for Matrix-Wise Models Authors: Wang... CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation. CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation... QuoteBench: How Matched Scores Can Hide Command-Path Failures. QuoteBench: How Matched Scores Can Hide CommandPath Failures Core Problem and... Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents. Here is a summary of the paper: The paper "Beyond Outcome Rewards: Step-Level... OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways. OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for... FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving. FlashDrive is an algorithm-system co-design framework that targets all four... AQuA: Recursively Self-Improving Quantitative Trading Research Agents. AQuA: Recursively Self-Improving Quantitative Trading Research Agents Summary... Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents. The paper introduces CREST (Hierarchical Credit Assignment via Entropy-Gated... VALG: An Agentic System for ML Theory Research. VALG is an agentic system for machine learning theory research that combines... NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents. This paper presents NaviDC-OCR, a unified document parsing framework designed... ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval. ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval Abstract While... Chance-constrained selection of sequential intervention strategies from counterfactual estimates. Problem and Motivation This paper addresses the challenge of selecting... Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement. This survey paper, "Numeracy in Large Language Models: Fundamental Limitations... On the Expressive Power of Transformers. This survey paper provides an overview of the expressive power of transformers,... EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory. EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal... Into the ORBIT for Time Series: Training Regimes for Foundation Models. This paper introduces ORBIT (Omni-Range Bootstrap Incremental Training), a... Exponential Convex Calibration Dimension for the Multi-Label Jaccard Measure. Exponential Convex Calibration Dimension for the Multi-Label Jaccard Measure... Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research. This paper presents a deployment-oriented two-tier agentic system that... LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation. LYCHEE MEMORY V2 is an efficient long-term memory framework for LLM agents that... RippleMem: From Isolated Retrieval to Associative Recollection for Long-Term Agent Memory. RippleMem is a long-term memory system for LLM-based agents that shifts memory... EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval. EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval... Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing. Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth... Algebraic Decomposition Theory for Transformer Length Generalization. This paper establishes the first complete characterization of which regular... HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark. HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark... Latent On-Policy Self-Distillation. The paper addresses the challenge of enabling agents to learn from experience... Virtual Temperature Sensors in Power Transformers Using Neural Ordinary Differential Equations. This paper presents a physics-aware Neural Ordinary Differential Equations... DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees. DARTree is a training-free speculative decoding method that extends a... High-resolution Calibrated Probabilistic Hourly Precipitation from a Deterministic Forecast. This paper describes an “Attention Residual U-Net” method for probabilistic... Technical Report on Resilient and Secure Large-Scale Energy Internet Systems. This IEEE PES Task Force report examines the security and resilience of... Intern-S2-Preview: Scientific Agentic Foundation Model. Intern-S2-Preview is a series of scientific agentic foundation models designed... OmniScientist: An Omni-Modal Omni-Discipline AI Scientist. OmniScientist: An Omni-Modal Omni-Discipline AI Scientist Core Contribution The... InSPECtor: Improving SLEIGH Processor Specification Veracity via Proxy. InSPECtor: Improving SLEIGH Processor Specification Veracity via Proxy Summary... MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning. MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning Summary... LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection. This paper presents a comprehensive benchmark evaluating whether Low-Rank... Falsehood and Impossibility Are Different Directions in an AI's Representation of Language. This paper reports an exploratory activation study of the multimodal... Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning. Critic-Free Pretraining (CFP) is an efficient paradigm for offline-to-online... CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives. CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical... EEG Decoding Using CNN and LSTM Network. This study introduces a hybrid deep-learning architecture that integrates a... Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy. This paper presents a scalable, context-aware Multi-Agent Framework designed to... DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition. DMDINTEL: Interpreting Large Language Models via Dynamic Mode Decomposition... When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation. Based on the paper, here is a detailed summary: The paper investigates whether... Adversarial Robustness in Smishing Detection: A Comparative Analysis of Adversarial Fragility in Classical vs. Transformer-Based Detection Systems. This study evaluates the adversarial robustness of five model architectures for... RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation. RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in... LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation. LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea... HybridRAG-BN: A Retrieval-Augmented Framework with Fine-Tuned Verification for Bangla KBQA. HybridRAG-BN is a retrieval-augmented framework for Bangla knowledge-base... HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models. HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large... Correct Is Not Governed: Provenance Integrity in Agentic Workflows. Core Thesis: The paper defines "governed execution" as distinct from task... Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering. This paper introduces Syfer, a synthesizer-folding framework for multilingual... Understanding Backdoor Vulnerabilities in Vertical Federated Learning: The Gap Between Research and Practice. This paper, "Understanding Backdoor Vulnerabilities in Vertical Federated... Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors. Dual-Stream Cross-Anchor Correction (DSCC) is a fine-tuning framework proposed... Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language. Static analysis-guided agentic AI translation enables Rust as a full stack... Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence. This paper investigates the validity of the conditional-independence assumption... Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs. The paper introduces the Sensitive Entity Alias Generator (SEAG), a... FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation. FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative... A Probe Direction Is a Property of Its Prompt. Core Claim The paper argues that the standard instrument used to detect whether... Branch and Bound for Relational Verification of Neural Networks. Branch and Bound for Relational Verification of Neural Networks Authors: Kota... Finding the Needle in a Haystack: Test-Time Analog Circuit Representation Adaptation for Bayesian Optimization. Test-Time Analog Circuit Representation Adaptation for Bayesian Optimization... Adaptive Nearest Neighbors Classifier via Granular Ball Computing. The paper proposes an adaptive and efficient k-Nearest Neighbor (KNN) approach... Prompts in the Wild: A Large Analyzed Collection of Transactional Prompts in Code. Here is a detailed summary of the paper "Prompts in the Wild: A Large Analyzed... Huber-Wasserstein barycenters for robust distribution-valued data. Huber–Wasserstein barycenters for robust distribution-valued data Authors:... CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers. CW-BASS v2 is a saturation-aware pseudo-label selection method for... Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging. Mr3D-VL is a dedicated visual-language foundation model for multi-parametric 3D... Sparse Orthogonal Regression Technique: A Spectral Framework for Equation Discovery, Approximation, and Integration. The paper develops the Sparse Orthogonal Regression Technique (SORT), "a sparse... Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry. This paper studies decentralized multi-player reinforcement learning in... DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data. DFM Mimir v1 is a 1-billion-parameter language model based on the Hierarchical... Error-Aware Reverse Auction Mechanism for Large Language Model Routing. Error-Aware Reverse Auction Mechanism for Large Language Model Routing Authors:... Discovering Persistent Behavioural Patterns for Interpretable Blockchain Forensics. This paper proposes a scalable, application-agnostic framework for persistent... High-dimensional networks and mean squared error for possibly misspecified models. The paper "High-dimensional networks and mean squared error for possibly... Thermodynamics of Learning: A Typed Four-Component Accounting of Memory, Fit, and Value. Here is a summary of the paper, constructed from direct quotes and detailed... Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia. PRISM (Perturbation-based Regional Interpretability through Subtraction... Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection. This study investigates whether self-supervised learning (SSL) speech... Physics-informed distribution of relaxation times estimation and latent-space condition monitoring of solid oxide fuel and electrolysis cells from electrochemical impedance spectroscopy. The paper proposes a physics-informed convolutional neural network (CNN)... Generative Universal Multimodal Retrieval with Dual-role Identifiers. The paper proposes DrIG, a novel generative framework for universal multimodal... MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination. MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and... Distribution Steering via Sliced Optimal Transport Control. Distribution steering seeks feedback laws that drive the state law of a... Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings. ParliamentRAG is a Retrieval-Augmented Generation (RAG) system designed for the... A Unifying Perspective on Causal World Models: From Observations to Representations to Structure. This paper studies world models (WMs) from a causal perspective across multiple... Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation. This paper audits a preserved MCP (Model Context Protocol) agent security... TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems. TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death... GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport. GeoCache is a training-free acceleration method for multi-view texture... LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service. LipCache: A Local Inference Proxy with Certified Caching for Edge Image... Operationalizing Cyber Threat Intelligence with GraphRAG. This paper investigates whether using a knowledge-graph-based retrieval system... Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling. This paper introduces a novel proxemics-based reward formulation for deep... ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification. ReflectFact is a novel self-reflective agent framework for multi-hop fact... Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models. Self-referential prompting has been shown to reliably induce large language... AI and Consumer Rights in India Working Paper. This working paper examines whether India's Consumer Protection Act, 2019,... Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors. This paper proposes a controllable method to erase copyrighted animation... TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval. TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval... Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation. Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in... UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations. UniTraffic-Agent is introduced as the MR-CAS solution for Track 3 of the 10th... SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback. SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction... Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds. This paper challenges the default paradigm of aligning large language models... BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs. BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through... MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification. The paper proposes ARMDIL, an Adaptive Router for Multi-Domain Image... Sovereign by necessity? Frontier AI export controls, cyber security, and the limits of national AI capability. The most capable artificial intelligence (AI) frontier models are produced by a... Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence. Spatial Memory Agent (SMA) is an experience-grounded runtime framework that... Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting. This paper proposes a unified evidence reasoning framework that addresses two... Slow and Steady: Preventing MEV with Verifiable Delays. This paper presents a defense mechanism against Maximal Extractable Value (MEV)... Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology. The paper proposes an intervention-aware clinical world model for forecasting... LOB-ID: Evaluating Synthetic Market Data by Inception Distances. The paper introduces LOB-ID (Limit Order Book Inception Distance), an... vToken: Token-Level Virtualization for Reclaimable KV Caches. vToken is a lightweight token-level virtualization layer that decouples logical... Comment on "Modeling rapid language learning by distilling Bayesian priors into artificial neural networks". This comment paper argues that McCoy & Griffiths’ (M&G) method of using... Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining. Core Contribution This paper proposes a task-agnostic measure of training data... From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options. Large language models (LLMs) perform well across a wide range of tasks but... PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs. This paper introduces PROVE-RT, an LLM-assisted framework for generating... Which LLM Is Your Ideal Companion? Evaluating Emotional Companion Capabilities of LLMs Based on Adult Attachment Theory. Based on the paper "Which LLM Is Your Ideal Companion? Evaluating Emotional... FlowLOB: Efficient and Controllable Limit Order Book Generation with Flow Matching. FlowLOB is a conditional flow-matching generative model for limit order book... Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies. Large Language Models (LLMs) are increasingly deployed in discovery domains... A Multispectral Framework for the Detection of Calcium Carbide-Induced Ripening and Shelf-Life Estimation in Climacteric Fruits. The proposed study explores a novel, non-invasive multispectral framework for... PatientAct: Theory-Grounded Mental Health Client Simulation. PATIENTACT: Theory-Grounded Mental Health Client Simulation This paper... On the Structural Limits of Machine Learning Decision Systems: An Information-Theoretic, Interaction-Based, and Stochastic-Dynamical Perspective. This paper examines the structural limits of machine learning decision systems... LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles. LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses... When Should Multi-Round RAG Stop? Structured Stopping Judgments and Retrieval Reduction in Search-R1. Core Contribution and Setup This paper addresses the stopping problem in... Foundations of Independent Component Analysis. This paper presents a comprehensive, self-contained mathematical treatment of... Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks. This paper presents a heterogeneity-aware belief synchronization framework for... MergeOver: Post-Training Token Merging for Recursive Vision Transformers. MergeOver is a post-training approach that integrates Token Merging (ToMe) into... NAS-Driven Hardware Accelerator Exploration for Edge AI and Quantization Effects on the Pareto Space. This paper proposes a three-stage hardware-aware Neural Architecture Search... Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data. The paper "Novel Knowledge-Guided Generative Methods for Synthetic... Making AI-Generated Feedback Matter: A Large-Scale Study of Feedback Workflows and Student Enactment. This study examined how different AI-mediated feedback workflows were... Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization. This paper investigates whether instance-wise allocation of a fixed test-time... SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference. SPADE: Speculative Decoding for Precise and Low-Cost Distributed Edge–Cloud... AutoQuREO: A Framework for Automated Quantum Resource Estimation and Optimization. AutoQuREO is an automated framework for full-stack quantum resource estimation... Revisiting Overestimation Bias Problem of Q-learning: Settling Large Discrete Action Space via Action Intersection. This paper considers the overestimation bias problem of Q-learning in the... Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization. The paper "Doubly Robust Estimation of Causal Effect on CVR with Targeted... The Time Value of Evolution. The paper formalizes the concept of the "time value of evolution," which... On the global feature importance for interpretable and trustworthy heat demand forecasting. The paper introduces an ante-hoc Explainable AI (XAI) methodology to assess the... Smart Contract Invariants Protect Against Cybercriminals. This paper investigates whether smart contract invariants can protect against... Beyond Source: An Empirical Study of Python Bytecode Security Risks. This paper presents an empirical study of Python bytecode as a security... Robust data-driven discovery of fractional differential equations via weak formulations and Pareto-based subset selection. The paper "Robust data-driven discovery of fractional differential equations... The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity. The paper introduces the unmasking growth complexity (UGC) as a path-resolved... InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers. InFactPlanner is a trace-driven decision-support framework for what-if analysis... Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents. This paper introduces the concept of skill misevolution in self-improving LLM... Concept Drift Detection and Adaptive Retraining of Malware Classification Models. Concept drift refers to changes over time in the statistical properties of... Jointly Predicting Courses and Grades Using a Transformer-Based Model. This paper introduces a TRansformer for Academic Course-grade Estimation... Rules or Character? Scaling Laws for AI Safety Design. This paper introduces a stylized comparative-statics model to analyze the... Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity. The paper investigates how instruction tuning affects model confidence and the... Deliberate Practice: Learning Robot Skills under a Budget. The paper "Deliberate Practice: Learning Robot Skills under a Budget" by Shivam... PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR. PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in... It's How You Ask: Gender-Associated Linguistic Bias in LLMs. The paper "It's How You Ask: Gender-Associated Linguistic Bias in LLMs" by... Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety. This paper proposes Wrapper-Based Intent-Form Augmentation (WIFA), an automatic... Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales. Normative datasets are often used to train and align AI systems, but the norms... Task- and dataset-specific information in protein language models. This paper investigates the internal representations of protein language models... SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents. SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in... Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI). This study tested whether a language model's explanatory engagement with a rare... Tight Nonasymptotic Local Convergence of Sinkhorn-Knopp. The paper "Tight Nonasymptotic Local Convergence of Sinkhorn-Knopp" revisits... The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use. This paper investigates why latent world models fail at long-horizon planning,... H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities. This paper presents the first comprehensive adaptation and evaluation of... A Compositional Theory of Curvature in Probabilistic Circuits. Core Contribution This paper introduces a compositional theory of curvature in... Memorization Diagnostics for Code LLMs Should be Scale-Aware. The paper investigates whether standard memorization diagnostics for code LLMs... HybridSB-MoE: Dual-Domain Schrödinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement. HybridSB-MoE is a dual-domain framework for speech enhancement that combines a... Knowledge-guided Pattern Discovery via Coupled Tensor Factorizations. This paper introduces a knowledge-guided approach for pattern discovery that... TeleGapper: On the (un)reliability of Privacy Policies in Telegram Mini apps. TeleGapper: On the (un)reliability of Privacy Policies in Telegram Mini apps... OmniSphinx: Active Mix Networks (Extended Version). OmniSphinx is a novel mix format that applies ideas from active networking to... Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency. and Motivation The paper addresses a critical gap in evaluating Joint-Embedding... Dissecting Software Graphs: Structural Insights for Driver-Guided Fuzzing. This paper presents an empirical study of software structure under multi-driver... Full-Key Recovery and Forgery from One MQOM v2.1 Signature. This paper presents a full-key-recovery attack on MQOM v2.1, a Round-3... SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization. SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via... TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies. TopoIntent is a system that compiles natural-language security intent into... Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks. The paper "Large-scale Testing Global Optimization Methods with Black-box... Distinguishing case-mix from context heterogeneity in prognostic regression model synthesis settings. Distinguishing case-mix from context heterogeneity in prognostic regression... Defensive Boosting for Online Probabilistic Forecasting. Problem Setting and Motivation The paper studies online probabilistic... NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video. NARU is a benchmark designed to evaluate narrative evolution and cultural... I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization. I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization Summary... A Commitment-Based Hybrid Post-Quantum Cryptographic Model for Multi-File Cloud Storage. This paper presents a commitment-based hybrid post-quantum cryptographic model... Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning. the Paper Problem Statement and Motivation The paper addresses the fundamental... Statistical Properties of Robust Learning under Distributional Shifts. and Motivation This paper studies the statistical properties of robust learning... Incremental Evaluation and Training in Relational Deep Learning. Incremental Evaluation and Training in Relational Deep Learning Authors: Jakub... CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation. CardioState-JEPA is a cardiac foundation model that learns a single shared... Difference-of-Convex Regularization for Graph Learning by Differentiable Programming. This paper proposes a Difference-of-Convex Regularizer (DCR) graph learning... ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization. ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization Lixing Li,... The Impact of Temporal Context Length and Encoding Strategies on Self-Supervised ECG Representation Learning. The paper investigates how temporal context length and encoding strategies... Decoupled Contrastive Decoding via Expert-Aligned Drafting. Decoupled Contrastive Decoding (DCD) is a method to accelerate Contrastive... VR-Themis: A Scalable Framework for Virtual Reality Application Clone Detection. VR-Themis: A Scalable Framework for Virtual Reality Application Clone Detection... Balanced Adaptive Prototype Selection for Scalable TabPFN Inference on Large-Scale Tabular Data. This paper introduces Balanced Adaptive Prototype Selection (BAPS), a framework... MBA: Multimodal Benchmark and Agents for Real-World Business Ideation. the Paper "MBA: Multimodal Benchmark and Agents for Real-World Business... UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models. UniTexture is a cross-task universal adversarial texture attack that uses a... Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs. The paper "Beyond Visual Evidence: Revealing and Mitigating Relational Privacy... CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment. CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference... Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge. The Information Abundance Paradox, as proposed in this paper, hypothesizes that... REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation. On-policy distillation (OPD) trains a student on its own generated trajectories... AI Guardrail Survival under Single-Cycle Agentic Self-Summarization. This paper investigates how a standing safety rule is lost during a... Foundation models for movement data: Are they ready for prime-time?. Foundation models (FMs) trained on large-scale accelerometer data have been... Designing AI Pipelines for Decision-Ready ITSM Intelligence. This paper reframes ITSM data use as an IS problem of transformation,... CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation. CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice... On the Importance of Geometric Nonlinearity and Temperature-Dependent Properties in Multi-Material Thermo-Mechanical Topology Optimization. This paper investigates two common simplifying assumptions in thermo-mechanical... Mawqif-XT: An Arabic Benchmark Dataset for Cross-Target Stance Detection. Mawqif-v2 is an extension of the original Mawqif dataset, designed as a... Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models. The paper investigates whether output homogeneity in language models originates... ATOBench: Tracing How Autonomous Penetration-Testing Agents Verify Vulnerabilities When Target Evidence Lies. ATOBench introduces an evaluation framework that makes the verification process... AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design. AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Core... Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes. This paper explores the use of Small Language Models (SLMs) to support... LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure. LittleLearner: Language Models Under Pedagogically Controlled Knowledge... Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence. The paper introduces a Polish-language medical visual question answering (VQA)... The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis. The paper proposes a YOLO- and Contrastive Language-Image Pre-training... Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test. The paper introduces a finite capability sheaf to model failures in AI agent... Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension. The paper presents the organizational framework adopted by the Academic League... TabSOM: A tabular-to-image encoding method based on self-organizing maps. TabSOM: A tabular-to-image encoding method based on self-organizing maps... Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks. and Motivation This paper, authored by Taha Shieenavaz, Shabnam Zareshahraki,... iARCS: Iterative Agentic RL for Controllable 3D Scene Generation. iARCS: Iterative Agentic RL for Controllable 3D Scene Generation Overview The... Multi-perspective Imbalance-Conscious 6G Beamforming Optimization and Performance. The study presents a systematic machine learning (ML) study of 6G-IoT... Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model. Mixture of Training (MoT) is a scaffolded modular pre-training procedure that... Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision. This study presents the first use of egocentric vision for freezing of gait... Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs. This preliminary technical report presents a framework for sign language video... EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction. EGRL: Edge generation-guided relation-aware learning for RNA-protein... RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level. This paper addresses the challenge of assessing the maturity of artificial... Learning the Mathematical Property for Designing Low Mutual Coherence Binary Sensing Matrices. This paper proposes a novel learning-based framework for constructing binary... CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model. CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large... Beyond Simulated Benchmarks: Evaluating Motion Representations for Fall Detection Under Real-World Data Scarcity. Falls are a major health concern for older adults, and wearable sensors have... Moose: Latent concept learning with reasoning-shortcut awareness in. Moose: Latent concept learning with reasoning-shortcut awareness in EL++... Multi-Layer Context Camouflaging: A Semantic Superposition and Contextual Lamination Framework for Malpractice-Resilient Online Assessment. This paper extends the Multi-dimensional Spatio-Temporal Context Camouflaging... Uniform Herding: Exemplar Replay with Representation Refresh. Uniform Herding: Exemplar Replay with Representation Refresh proposes a... Efficient Hessian-Free Methods for Multi-Objective Bilevel Optimization with Nonconvex Lower Level. the Paper Problem Statement This paper addresses multi-objective bilevel...
The papers
- Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents — This paper investigates whether tool-using language model agents execute the same action policy for semantically identical tasks across different languages.
- CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation — CABS+ is an enhanced model merging framework that extends the Conflict-Aware and Balanced Sparsification (CABS) method to address its limitations in time complexity, GPU memory consumption, and optimization bias.
- From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion — Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead.
- EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding — EEG-PRIME is a two-stage EEG foundation model for cross-dataset multi-task EEG decoding, proposed to address the poor generalization of EEG decoding models across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology.
- A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family — The paper addresses a fundamental weakness in how GPU kernel generation systems verify correctness. The authors state: "Systems that generate GPU kernels with language models report high correctness rates.
- Equivariant learning of a transferable three-dimensional classical density functional — Equivariant learning of a transferable three-dimensional classical density functional Bingqing Cheng arXiv:2608.13506v1 [cond-mat.stat-mech] 13 Aug 2026 Summary This paper introduces Equi-cDFT, an Energy-first, EQUIvariant framework for learning the excess free-energy functional
- Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians — Hamilton-Zero is a foundation model for computing ground states of arbitrary quadratic qubit Hamiltonians, trained on a dataset of hundreds of thousands of Hamiltonian systems.
- OEIS Open: How many conjectures can language models turn into theorems? — OEIS Open: How many conjectures can language models turn into theorems? Tom Adamczewski, Epoch AI arXiv:2608.11941v1 [cs.AI] 12 Aug 2026 Abstract We construct OEIS OPEN, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al.
- Governing Agentic AI in FinTech — This paper addresses the governance challenges posed by agentic AI systems in financial services.
- AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1) — AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1) presents an improved version of AlayaWorld.
- Beyond Parameter Space: NTK-Guided Personalized Aggregation for Robust Federated Learning — LIGHTYEAR is a federated learning (FL) framework that performs update selection in the function space rather than parameter space.
- TANGCO: Learning Topology-Aware Capacity Allocation for Overload-driven Cascading Failures — TANGCO: Learning Topology-Aware Capacity Allocation for Overload-driven Cascading Failures Abstract Many networked systems, from power grids to traffic networks and cloud clusters, carry loads across nodes with limited capacity.
- Homomorphic Aggregation of Continuous-Variable GKP States — Homomorphic aggregation of logical quantum information encoded in continuous-variable phase space is essential for distributed quantum computing.
- Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations — Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence.
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models — [The summary for "Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models" was not provided in the context. [episode]
- Localize, Then Reason: Visual Latent Structural Reasoning for Molecular Properties and Edits — The paper introduces Visual Latent Structural Reasoning (VLSR), an end-to-end framework for molecular property reasoning from molecular images that follows a "localize-then-reason" strategy.
- Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses — DECAF: Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses Summary This paper introduces DECAF (Decomposition of Evidence, Contradiction, And Fragility), a method for interpreting perturbation-based model explanations by decomposing the response magn
- Synthetic Persona Pretraining: Alignment from Token Zero — The paper introduces Synthetic Persona Pretraining (SPP), a method that installs the desired assistant persona from the very beginning of pretraining ("from token zero"), rather than introducing alignment only during post-training.
- DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models — Summary This paper introduces DreOPD (Degraded-reference extrapolative On-Policy Distillation), a post-training method for flow-matching models used in text-to-image generation.
- Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services — This paper studies a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit.
- Keep the Future, Drop the Rollout: RIFT for World Action Models — World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. This paper asks whether action generation requires the evolving rollout trajectory or only its future representation.
- Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws — This paper introduces the Neural Quadratic Form (NQF), a universal minimal model that unifies sudden learning and neural scaling laws across diverse neural network architectures.
- Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation — Author: Valentin Noël (Devoteam) arXiv: 2608.13337v1 [cs.LG], 13 Aug 2026 --- The paper investigates a hidden methodological flaw in how sparse autoencoder (SAE) latents are causally evaluated.
- LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning — Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences.
- Active-Trace Complexity Bounds for Moreau--Yosida Unadjusted Langevin Sampling — This paper studies the Moreau–Yosita unadjusted Langevin algorithm (MYULA) for sampling from a nonsmooth composite target distribution of the form π(dx) ∝ exp −f(x) − g(x) dx, x ∈ R d, where f is m-strongly convex with L f-Lipschitz gradient and g is convex and G-Lipsc
- CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility — CoMedBench is a reproducible benchmark for evaluating synthetic medical data, designed to address the fragmentation in prior evaluations that often use a single generator, one dataset, or a narrow downstream task.
- Demand Transfer Estimation at Scale via Restricted Logit Modeling — Item demand forecasting is an integral component of store assortment optimization (SAO). Existing literature focuses on learning a suitable customer choice model and using this model to determine the value of an objective function (i.e.
- Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development — This paper presents a systematic evaluation framework for long-horizon AI research and development agents that goes beyond final scores to characterize within-run behavior and experience reuse.
- ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models — ContactGuard is a pre-contact execution monitoring system for chunked visuomotor policies in contact-rich robot manipulation tasks.
- Foundations of MT-PDCL: Measure-Theoretic Probabilistic Definite Clause Logic — Measure-Theoretic Probabilistic Definite Clause Logic (MT-PDCL) is a generalized foundational framework for probabilistic logic programming that eliminates the finite-domain restriction of standard frameworks.
- CAPRI: Contract-Aware Proof Repair for Isabelle — CAPRI: Contract-Aware Proof Repair for Isabelle Abstract. We address the use of large language models (LLMs) to help discover Isabelle proofs. An Isabelle build establishes that the submitted theory is accepted, but not that an LLM changed only what the developer authorised.
- History-informed Lagrangian Neural Networks — History-informed Lagrangian Neural Networks (HiLNN) is introduced to address the challenge of long-horizon forecasting of mechanical systems from position-only observations.
- Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation — This paper introduces a protocol-level identifiability audit for LLM evaluation, addressing the problem that "LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure." The authors formalize a di
- RealmEye: Virtual Machine Introspection for Arm CCA Realm VMs — RealmEye is the first Virtual Machine Introspection (VMI) system for Arm CCA Realm VMs.
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents — Summary This paper presents NaviDC-OCR, a unified document parsing framework designed to handle both digital and camera-captured documents. [episode]
- A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings — This paper introduces a cloud–edge collaborative architecture for multimodal clinical screening in resource-constrained rural settings.
- Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents — The paper introduces CREST (Hierarchical Credit Assignment via Entropy-Gated Self-Teacher), a framework for training multi-turn multi-step LLM agents that addresses the hierarchical credit assignment problem in reinforcement learning with verifiable rewards (RLVR).
- ProME: Prototype-Margin Environments with Repair-Aware Selection for Group-Robust Learning — ProME: Prototype-Margin Environments with Repair-Aware Selection for Group-Robust Learning Summary This paper introduces ProME (Prototype-Margin Environments), a two-stage framework for group-robust learning that operates without training-group labels.
- QuoteBench: How Matched Scores Can Hide Command-Path Failures — The paper addresses a critical blind spot in evaluating LLM coding agents that issue Bash commands. The authors argue that "LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output.
- Algebraic Decomposition Theory for Transformer Length Generalization — This paper establishes the first complete characterization of which regular languages transformers can length-generalize on, and provides a polynomial-time decision algorithm for this membership problem.
- Sinkhorn Linearization and the Spectral Proxy: Unifying the Statistical and Algorithmic Theory of Feature-Parameterized Inverse Optimal Transport via a Single Spectral Sandwich — The paper develops the statistical and algorithmic theory of inverse optimal transport (IOT) under the feature-parameterized cost Cθ(i, j) = −θ⊤φ(i, j), where observations are conditional transition operators of the entropic optimal transport plan.
- Chance-constrained selection of sequential intervention strategies from counterfactual estimates — This paper addresses the challenge of selecting sequential intervention strategies under a cumulative resource budget, where the decision maker must maximize an outcome while constraining the probability that the cumulative cost exceeds a budget.
- CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation — CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation Summary This paper introduces Counterfactual Relevance for On-Policy Distillation (CROP), a method for selective on-policy distillation (OPD) that allocates token-level supervision based on task relevan
- Exponential quantum advantage for learning signals with a single qubit — The paper demonstrates that coupling a single controllable qubit to a conventional quantum sensor can exponentially reduce the number of measurements required to learn classical signals.
- CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport — CoverPrune is a training-free token pruning framework for 3D Vision-Language Models (3D VLMs) that shifts the pruning objective from maximizing token diversity to preserving visual evidence coverage.
- Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research — This paper presents a deployment-oriented two-tier agentic system that separates a maintained, point-in-time knowledge library from report writing.
- Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity — The paper "Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity" investigates why large language models (LLMs) fabricate plausible-sounding details about entities outside their knowledge boundary instead of retreating to safer, more general cla
- High-resolution Calibrated Probabilistic Hourly Precipitation from a Deterministic Forecast — This paper describes an “Attention Residual U-Net” method for probabilistic quantitative precipitation forecasting (PQPF) that predicts the hourly probability of no precipitation plus the distribution of positive precipitation from a weighted mixture of two Gamma distribution
- Symmetry-Breaking De Novo Crystal Generation via Markovian Jump Diffusion — The paper introduces SbCD (Symmetry-breaking Crystal Diffusion), a novel diffusion-based generative framework for crystals that produces complete crystallographic structure specifications, including space groups and Wyckoff positions, rather than relying on empirical distribution
- Technical Report on Resilient and Secure Large-Scale Energy Internet Systems — This IEEE PES Task Force report examines the security and resilience of large-scale Energy Internet (EI) systems, in which electricity, information, and market layers are tightly coupled through pervasive digitalization.
- LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation — LYCHEE MEMORY V2 is an efficient long-term memory framework for LLM agents that replaces turn-level consolidation with semantic segment-level consolidation.
- On the Expressive Power of Transformers — This survey paper provides an overview of the expressive power of transformers, the core component of modern large language models, by comparing them to standard models of computation, particularly circuit complexity classes.
- Momentum as Residual-Driven Multiplier Correction for Deep Learning Optimization — The paper proposes an ADMM-Inspired Momentum (AIM) framework based on residual-penalty variable splitting, which interprets momentum as a multiplier-like correction driven by the splitting residual.
- ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval — ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval Abstract While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governing this memory are rarely treated as evolvable components.
- Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing — Authors: Sheng Ren, Yadong Wang, Naiqiang Tan, Jiangang Kong, Jun Fang, Rui Liu, Jun Wang, Kai Chen, Lipeng Liang, Xiang Chen Affiliation: Nanjing University of Aeronautics and Astronautics; Didichuxing Co.
- Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents — Summary This paper introduces HARD (Harness-based Autonomous Runtime Defense Evolution), a framework for self-evolving runtime defenses for LLM agents.
- TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint — TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint Published: COLM 2026 Core Finding: Vision-Language Models (VLMs) can internally distinguish when abstention is required based on visual evidence, but fail to express this restraint in their outputs.
- Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval — The paper "Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval" presents the GENAI4E team's solution to AI City Challenge 2026 Track 4, addressing the task of text-based person anomaly retrieval (TBAPS).
- VALG: An Agentic System for ML Theory Research — VALG is an agentic system for machine learning theory research that combines multi-level Verification, Adaptive formulation of Learning-theory problems, and Graph-structured proof development. [episode]
- Bagging Robustly Learns VC Classes with Linear Sample Complexity — This paper proves that VC classes are adversarially robustly learnable with sample complexity linear in the VC dimension d, providing an exponential improvement over the previous upper bound of Montasser, Hanneke, and Srebro [2019].
- Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories — The paper identifies a critical bottleneck in agent memory systems: while retrieval can identify a past trajectory that may be relevant, it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed.
- Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information — The paper proposes "Rule of Thumb" (ROT) explanations, a new approach to Explainable Artificial Intelligence (XAI) that "identifies the most relevant features for predicting the behaviour of an AI system, for a particular datapoint." The authors state: "We propose 'Rule-of-Thumb'
- RippleMem: From Isolated Retrieval to Associative Recollection for Long-Term Agent Memory — RippleMem is a long-term memory system for LLM-based agents that shifts memory access "from isolated retrieval to adaptive associative recollection." The paper argues that the main bottleneck in long-horizon agent memory "is not simply storing past experience, but recovering the
- Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference — Summary This paper introduces Reduced Matrix Multiplication (RMM), a training-free, input-adaptive inference method that reduces Transformer matrix products by selecting informative slices along their contraction dimensions, without modifying model weights.
- BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving — BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving Abstract Autonomous driving requires planning under both semantic constraints and predictive dynamics.
- Exponential Convex Calibration Dimension for the Multi-Label Jaccard Measure — Author: Mingyuan Zhang (Independent Researcher) arXiv: 2608.13549v1 [cs.LG] 13 Aug 2026 --- The paper studies the per-instance Jaccard score (intersection over union, IoU) in multi-label classification and binary segmentation.
- Federated Compositional Muon Optimizer for Matrix-Wise Models — Authors: Wang Yan, Feihu Huang Affiliation: College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics, Nanjing, China arXiv: 2608.12710v1 [cs.LG] 13 Aug 2026 --- The paper states: "Muon, a more recently developed optimizer, is useful for matri
- Intern-S2-Preview: Scientific Agentic Foundation Model — Intern-S2-Preview is a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The main model evaluated is Intern-S2-Preview-397B.
- Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks — This paper introduces Cocktail, a watermarking scheme for large language models (LLMs) that simultaneously provides provenance tracing and tamper evidence, addressing the vulnerability of "piggyback spoofing" where an adversary alters critical content while retaining attribution.
- EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory — EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory EgoMonth is introduced as the first month-level egocentric video understanding benchmark for evaluating long-term spatiotemporal memory in real-world daily-life settings.
- SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data — SPARED introduces an adversarial reinforcement learning framework for AI-generated image detection that pits two heterogeneous models against each other: a diffusion image editor that learns to edit real photographs into fake counterparts that fool the current detector, and a rea
- Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement — This survey paper, "Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement" by Aoxin Ni, addresses the persistent failure of large language models (LLMs) to perform elementary numerical tasks despite their high-level mathematical reasoning capabilitie
- TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability — TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability Abstract The paper introduces TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation.
- Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents — The paper "Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents" addresses the challenge of sparse reward supervision in training deep search agents, which operate over trajectories spanning dozens of steps but receive only a single binary
- Virtual Temperature Sensors in Power Transformers Using Neural Ordinary Differential Equations — This paper presents a physics-aware Neural Ordinary Differential Equations (Neural ODEs) framework for virtual temperature sensing in power transformers, applied to real-world time-series data from fifteen transformers in the Norwegian transmission grid.
- Into the ORBIT for Time Series: Training Regimes for Foundation Models — This paper introduces ORBIT (Omni-Range Bootstrap Incremental Training), a training paradigm designed to explicitly control the effective pre-training distribution of heterogeneous time series corpora.
- DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees — DARTree is a training-free speculative decoding method that extends a pretrained causally corrected block-parallel drafter from a single chain to a speculative tree. It is designed to accelerate autoregressive language models by verifying multiple draft tokens in parallel.
- How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures — SciFigBench is a diagnostic vision-language model (VLM) benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral reliability under uncertainty.
- FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving — FlashDrive is an algorithm-system co-design framework that targets all four stages of Vision-Language-Action (VLA) model inference simultaneously to bring end-to-end autonomous driving substantially closer to real-time deployment.
- OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways — OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways Abstract Summary: The paper addresses heterogeneous USV cooperative pursuit in constrained port waterways, which requires evader interceptio
- AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models — AaLLM is an open-source, end-to-end multi-agent LLM workflow for analog circuit design that takes user specifications as input and outputs the appropriate netlist, encompassing both topology generation and circuit sizing.
- HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark — HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark Summary This paper introduces HumanTracker, a large-scale benchmark and a preference-aligned metric for evaluating humanoid motion tracking.
- Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization — SNIPER is a two-stage structured pruning framework for large language models (LLMs) that unifies depth and width pruning via binary knapsack optimization.
- ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs — ARAC: Benchmarking Auto-Research’s Alignment and Completeness on End-to-End Researchs proposes ARAC-Bench, a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes.
- PIPES: Securing Agent Perception with Provenance and Priors — PIPES: Securing Agent Perception with Provenance and Priors introduces a defense mechanism against a novel class of indirect prompt injection attacks called "state-corruption attacks." The paper identifies an "agent perception gap" in tool-using agents: tool responses rarely incl
- Vero: Can AI Agents Build Formally Verified Software Repositories? — Vero is the first benchmark to evaluate joint implementation and proof synthesis at the repository level in Lean 4.
- StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems — StateBridge is a training-free latent communication approach for large language model (LLM) multi-agent systems.
- EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval — EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval Summary This paper introduces EviReform, a method for multi-hop graph retrieval that separates revising the retrieval request from aggregating evidence in the graph.
- AQuA: Recursively Self-Improving Quantitative Trading Research Agents — AQuA: Recursively Self-Improving Quantitative Trading Research Agents Summary This paper introduces AQuA, a system comprising two separate language-model-driven research systems for quantitative investment research: one for symbolic factor discovery (Part I) and one for trainable
- Fine-tuned Normalizing Flows for ALICE Zero Degree Calorimeter Fast Simulation — Simulating the ALICE Zero Degree Calorimeter (ZDC) neutron detector responses at the LHC is computationally expensive, requiring complex Monte Carlo chains. We develop a generative surrogate, focusing on Normalizing Flows (NFs).
- Latent On-Policy Self-Distillation — The paper addresses the challenge of enabling agents to learn from experience and internalize it into their policy for self-evolving AI.
- Training AI Scientists to Replicate Research — This paper introduces Replica, a scalable task space for paper replication, and Faraday, a 27B-parameter "AI Scientist" agent trained to replicate research papers.
- MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning — MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning Summary This paper introduces MAG (MAnifold-Guided semi-supervised in-context demonstration selection), a framework designed to improve few-shot multi-modal in-context learning (ICL) by leveraging abundant unlab
- OmniScientist: An Omni-Modal Omni-Discipline AI Scientist — The paper introduces OmniScientist, described as "an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence." The system is designed to address a critical gap in existing AI scientist systems: "Existing systems are in
- LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection — This paper presents a comprehensive benchmark evaluating whether Low-Rank Adaptation (LoRA) of foundation models is sufficient for robust face presentation attack detection (PAD).
- InSPECtor: Improving SLEIGH Processor Specification Veracity via Proxy — InSPECtor: Improving SLEIGH Processor Specification Veracity via Proxy Summary This paper introduces InSPECtor, a framework for systematically validating SLEIGH processor specifications, which are used by tools like Ghidra for disassembly, decompilation, and emulation.
- Falsehood and Impossibility Are Different Directions in an AI's Representation of Language — This paper reports an exploratory activation study of the multimodal open-weight model Gemma 3 4B IT, investigating whether the model internally distinguishes between false statements (states of affairs that are false) and impossible statements (states of affairs that could not b
- Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning — Critic-Free Pretraining (CFP) is an efficient paradigm for offline-to-online (O2O) reinforcement learning that completely abandons offline critic training, allowing a freshly initialized critic to adapt without inheriting biased estimates.
- Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy — Summary This paper presents a scalable, context-aware Multi-Agent Framework designed to automate the construction of "Lines and Ladders" pricing taxonomies for large-scale retail catalogs.
- CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives — CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives Abstract Summary: The paper addresses the challenge of understanding temporal progression of symptoms in clinical narratives, which is critical for disease monitoring, safety surveillance, and c
- EEG Decoding Using CNN and LSTM Network — This study introduces a hybrid deep-learning architecture that integrates a convolutional neural network (CNN) with a bidirectional long short-term memory (bi-LSTM) network for motor imagery (MI) brain–computer interface (BCI) EEG decoding.
- DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition — DMDINTEL: Interpreting Large Language Models via Dynamic Mode Decomposition Summary This paper introduces DMDINTEL, a novel framework for interpreting the predictions of supervised fine-tuned large language models (LLMs) in classification tasks by leveraging Dynamic Mode Decompos
- When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation — Based on the paper, here is a detailed summary: The paper investigates whether rotation-based post-training quantisation (PTQ) can be improved by respecting the structural decomposition imposed by Rotary Position Embedding (RoPE).
- Adversarial Robustness in Smishing Detection: A Comparative Analysis of Adversarial Fragility in Classical vs. Transformer-Based Detection Systems — This study evaluates the adversarial robustness of five model architectures for smishing (SMS phishing) detection: three classical lexical models (Random Forest, XGBoost, CNN+BiLSTM) and two multilingual transformers (mBERT, XLM-RoBERTa), using a dataset of 27,037 messages combin
- RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation — Retrieval-augmented generation (RAG) treats an external corpus as inference evidence, allowing injected documents to promote attacker-chosen claims.
- LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation — LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation Abstract With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention.
- HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models — HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models Abstract Summary Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks.
- HybridRAG-BN: A Retrieval-Augmented Framework with Fine-Tuned Verification for Bangla KBQA — HybridRAG-BN is a retrieval-augmented framework for Bangla knowledge-base question answering (KBQA) that integrates hybrid retrieval using BM25 and BGE-M3, answer generation using the GGUF version of Gemma-4-31B-Instruct, and a LoRA-fine-tuned Gemma-4-31B-Instruct model for answe
- Sovereign by necessity? Frontier AI export controls, cyber security, and the limits of national AI capability — The most capable artificial intelligence (AI) frontier models are produced by a small number of firms based in two nation states.
- VR-Themis: A Scalable Framework for Virtual Reality Application Clone Detection — VR-Themis: A Scalable Framework for Virtual Reality Application Clone Detection Abstract. Repackaging of mobile applications (aka app cloning) not only threatens the security and privacy of mobile users but also infringes upon the copyright of the original app developers.
- NAS-Driven Hardware Accelerator Exploration for Edge AI and Quantization Effects on the Pareto Space — This paper proposes a three-stage hardware-aware Neural Architecture Search (NAS) pipeline for edge AI deployment on Coarse-Grain Reconfigurable Array (CGRA)-based accelerators, and presents an empirical study of the effects of INT4 Post-Training Quantization (PTQ) on the NAS-Ben
- Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks — The paper "Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks" by Wojciech Zarzecki and Jarosław Arabas argues that existing global optimization benchmark suites are limited and proposes that Black-Box Adversarial Attack (BBAA) problems can serve
- The Time Value of Evolution — The paper formalizes the concept of the "time value of evolution," which captures the delayed utility of evolutionary mutations: "a weak child can be a valuable ancestor that makes high-fitness regions reachable." Immediate-return control is "blind to this delayed utility, penali
- Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety — This paper proposes Wrapper-Based Intent-Form Augmentation (WIFA), an automatic intent-group augmentation method that pairs wrapped harmful examples with structurally matched wrapped benign counterexamples, requiring no external teacher or manual per-wrapper intent labels.
- Physics-informed distribution of relaxation times estimation and latent-space condition monitoring of solid oxide fuel and electrolysis cells from electrochemical impedance spectroscopy — The paper proposes a physics-informed convolutional neural network (CNN) framework for estimating the distribution of relaxation times (DRT) from electrochemical impedance spectroscopy (EIS) data, and for using the learned latent representation for condition monitoring of solid o
- Foundation models for movement data: Are they ready for prime-time? — Foundation models (FMs) trained on large-scale accelerometer data have been proposed as general-purpose feature extractors for health monitoring, but systematic evidence of their advantages is lacking.
- It's How You Ask: Gender-Associated Linguistic Bias in LLMs — The paper "It's How You Ask: Gender-Associated Linguistic Bias in LLMs" by Katherine Van Koevering and Anjalie Field, published as a conference paper at COLM 2026, investigates whether large language models (LLMs) respond differently to prompts containing linguistic features more
- A Probe Direction Is a Property of Its Prompt — The paper argues that the standard instrument used to detect whether a language model senses it is being evaluated—a contrastive activation probe built from prompts that announce an evaluation versus prompts that do not—has a free parameter that its readings do not disclose:
- Rules or Character? Scaling Laws for AI Safety Design — This paper introduces a stylized comparative-statics model to analyze the optimal balance between character shaping (e.g., RLHF, Constitutional AI) and rule enforcement (e.g., output filters, safety classifiers) in AI safety design as deployment scale increases.
- TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies — TopoIntent is a system that compiles natural-language security intent into executable, compliance-checked network topologies.
- TeleGapper: On the (un)reliability of Privacy Policies in Telegram Mini apps — TeleGapper: On the (un)reliability of Privacy Policies in Telegram Mini apps Abstract Telegram Mini Apps are Web applications embedded within the Telegram client, forming an ecosystem of third-party services within one of the world's most widely used messaging platforms.
- Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks — This paper presents a heterogeneity-aware belief synchronization framework for semantic communication in AI-native 6G networks.
- Jointly Predicting Courses and Grades Using a Transformer-Based Model — This paper introduces a TRansformer for Academic Course-grade Estimation (TRACE), a novel model that jointly predicts both the set of courses a student will take and their corresponding grades for an upcoming semester.
- Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings — ParliamentRAG is a Retrieval-Augmented Generation (RAG) system designed for the Italian Chamber of Deputies that addresses three specific risks in applying RAG to parliamentary transcripts: "dominance of the most frequent speakers, inability to weight speakers according to topica
- Deliberate Practice: Learning Robot Skills under a Budget — Summary The paper "Deliberate Practice: Learning Robot Skills under a Budget" by Shivam Vats, Sudarshan Harithas, Mete Tuluhan Akbulut, Arvind Raghunathan, and George Konidaris addresses the problem of autonomously learning robot skills for sequential tasks under a limited practi
- Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning — The paper addresses the fundamental question in robust statistics: "Can we design an accurate estimator and a computationally tractable algorithm to recover the true distribution P* under the Wasserstein-1 metric, given samples subject to Total Variation (TV) contamination?" The
- Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes — This paper explores the use of Small Language Models (SLMs) to support edge-based operation of selected components of the Cognitive Embodied Agent Architecture (CEAA), focusing on the Think and Memory processes.
- Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection — This study investigates whether self-supervised learning (SSL) speech representations capture disease-related characteristics of Parkinson’s disease (PD) or instead exploit dataset-specific confounds, particularly under cross-lingual transfer.
- Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity — The paper investigates how instruction tuning affects model confidence and the lexical diversity of generated answer rationales in question-answering tasks.
- LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles — LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles Objectives: As autonomous vehicles reach public roads, their software becomes safety-critical.
- UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models — UniTexture is a cross-task universal adversarial texture attack that uses a single textured 3D object to induce targeted deviations in VLA action predictions across multiple tasks.
- A Unifying Perspective on Causal World Models: From Observations to Representations to Structure — This paper studies world models (WMs) from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing environment dynamics.
- AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design — The paper addresses the challenge of transforming multimodal sources into condensed and structured media outputs (e.g., posters, slides, webpages, videos), which the authors conceptualize as "a long-horizon agentic process centered on a model-harness system." The key gap identifi
- CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment — CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment Published as a conference paper at COLM 2026 Authors: Bingcan Guo1, Eryue Xu2, Jijie Zhou3, Zhiping Zhang3, Tianshi Li3 1 University of Washington, 2 UIUC, 3 Northeastern University Abstract Ali
- Mawqif-XT: An Arabic Benchmark Dataset for Cross-Target Stance Detection — Mawqif-v2 is an extension of the original Mawqif dataset, designed as a benchmark for cross-target stance detection in Arabic. It consists of 996 manually annotated Arabic tweets collected from three public targets: Women Driving, E-Cars, and Trimester System.
- Full-Key Recovery and Forgery from One MQOM v2.1 Signature — This paper presents a full-key-recovery attack on MQOM v2.1, a Round-3 candidate in the NIST additional-signature process. The attack recovers the complete signing key from one accepted signature and uses it to produce a fresh-message forgery accepted by the reference verifier.
- On the Importance of Geometric Nonlinearity and Temperature-Dependent Properties in Multi-Material Thermo-Mechanical Topology Optimization — This paper investigates two common simplifying assumptions in thermo-mechanical topology optimization: small-strain linear elasticity and temperature-independent material properties.
- Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation — Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network’s receptive field exponentially with the number of decomposition levels while keeping the parameter count linear.
- PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR — PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR Summary This paper introduces PAIR (Pairwise-Aware Inclusion Reweighting), a method for adaptive rollout allocation in reinforcement learning with verifiable rewards (RLVR).
- AI Guardrail Survival under Single-Cycle Agentic Self-Summarization — This paper investigates how a standing safety rule is lost during a single-cycle agentic self-summarization (context compaction), where an agent's interaction history is replaced with a model-generated summary.
- Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models — The paper investigates whether output homogeneity in language models originates during pretraining or is introduced by the alignment process.
- CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation — CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Summary This paper proposes CookVoice, a unified non-autoregressive (NAR) framework for multimodal, multi-style, and multi-task human voice generation.
- MBA: Multimodal Benchmark and Agents for Real-World Business Ideation — The paper introduces MBA-Bench, described as "the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone." The authors automatic
- Making AI-Generated Feedback Matter: From Provision to Student Enactment — This study examined how different AI-mediated feedback workflows were associated with students' behavioural engagement, self-assessment confidence, and submitted-work quality.
- REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation — On-policy distillation (OPD) trains a student on its own generated trajectories under dense token-level supervision from a teacher, providing an effective post-training paradigm for large language models.
- Tight Nonasymptotic Local Convergence of Sinkhorn-Knopp — The paper "Tight Nonasymptotic Local Convergence of Sinkhorn-Knopp" revisits the Sinkhorn-Knopp (SK) algorithm for the matrix scaling problem. Despite extensive literature on global convergence, the local linear convergence behavior was less understood.
- Task- and dataset-specific information in protein language models — This paper investigates the internal representations of protein language models (PLMs) to determine where task-relevant information is stored across their layers.
- Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge — The Information Abundance Paradox, as proposed in this paper, hypothesizes that "when task-relevant information is made available through the training context, the model can reduce loss by using that information directly rather than by encoding it in its parameters.
- Designing AI Pipelines for Decision-Ready ITSM Intelligence — This paper reframes ITSM data use as an IS problem of transformation, abstraction, and decision support.
- Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs — The paper introduces the Sensitive Entity Alias Generator (SEAG), a privacy-preserving framework for Retrieval-Augmented Generation (RAG) systems.
- Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies — Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The usual approach is to present the problem to the model and use its answer as the proposed solution.
- Finding the Needle in a Haystack: Test-Time Analog Circuit Representation Adaptation for Bayesian Optimization — Authors: Fin Amin, Sounak Dutta, and Paul D.
- Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging — Mr3D-VL is a dedicated visual-language foundation model for multi-parametric 3D magnetic resonance imaging (mpMRI), designed to address critical limitations in current AI models for brain tumor diagnosis and treatment.
- The Impact of Temporal Context Length and Encoding Strategies on Self-Supervised ECG Representation Learning — Visualization of Long-Context Representations: "Fig. 2b and 2c visualize 10-minute frozen SSL embeddings using t-SNE [18] for discretized VQ and continuous CNN tokenization, respectively. [episode]
- HybridSB-MoE: Dual-Domain Schr"odinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement — HybridSB-MoE is a dual-domain framework for speech enhancement that combines a heterogeneous spectral Mixture-of-Experts (MoE) pathway with a waveform Schrödinger Bridge (SB) pathway, unified by an asymmetric uncertainty fusion design.
- Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia — PRISM (Perturbation-based Regional Interpretability through Subtraction Mapping) adapts subtraction analysis, the standard framework of human neuroimaging, from biological brains to perturbed transformers, and applies the same logic to both substrates in parallel.
- Error-Aware Reverse Auction Mechanism for Large Language Model Routing — Authors: Haolong Chen, Zhengyuan Xin, Liang Zhang, Lei Xue, Guangxu Zhu arXiv: 2608.12719v1 [cs.GT] 13 Aug 2026 --- The paper addresses the challenge of routing each query to a cost-effective large language model (LLM) to balance quality and cost.
- Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence — Spatial Memory Agent (SMA) is an experience-grounded runtime framework that enables a frozen Vision-Language Model (VLM) agent to improve its spatial reasoning through parameter-update-free self-evolution, without relying on external expert spatial tools at inference time. [episode]
- Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors — Dual-Stream Cross-Anchor Correction (DSCC) is a fine-tuning framework proposed to address object hallucination in multimodal large language models (MLLMs), where models confidently describe objects not present in the image.
- PatientAct: Theory-Grounded Mental Health Client Simulation — PATIENTACT: Theory-Grounded Mental Health Client Simulation This paper introduces PATIENTACT, a framework for simulating mental health clients using large language models (LLMs), designed to address critical shortcomings in existing simulators.
- Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry — This paper studies decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmetry: (A) unobserved actions with common rewards, (B) observed actions with independent rewards, and (C) unobserved acti
- ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization — ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization Lixing Li, Cornell University arXiv:2608.12756v1 [cs.CL] 13 Aug 2026 Abstract Adaptive latent tokenization maps a fine-grained input to a shorter sequence of continuous representations associated with input-dependent s
- Difference-of-Convex Regularization for Graph Learning by Differentiable Programming — This paper proposes a Difference-of-Convex Regularizer (DCR) graph learning framework to address the challenge of dense and ill-conditioned graph Laplacian pseudoinverses in Laplacian-regularized minimization.
- Correct Is Not Governed: Provenance Integrity in Agentic Workflows — Core Thesis: The paper defines "governed execution" as distinct from task correctness.
- PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs — This paper introduces PROVE-RT, an LLM-assisted framework for generating mechanized theorem prover scripts for real-time systems schedulability analysis using the PROSA/ROCQ proof assistant.
- Memorization Diagnostics for Code LLMs Should be Scale-Aware — The paper investigates whether standard memorization diagnostics for code LLMs remain effective as model scale increases, and proposes a new methodology to separate memorization from representational robustness.
- CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers — CW-BASS v2 is a saturation-aware pseudo-label selection method for semi-supervised semantic segmentation (SSSS) under foundation-model teachers.
- Thermodynamics of Learning: A Typed Four-Component Accounting of Memory, Fit, and Value — Summary This paper, "Thermodynamics of Learning: A Typed Four-Component Accounting of Memory, Fit, and Value," develops a formal framework for finite-state learning devices that distinguishes between what a device has memorized and what will hold operational value for it on futur
- Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors — This paper proposes a controllable method to erase copyrighted animation characters from text-to-image diffusion models during generation, while preserving overall image fidelity.
- Distribution Steering via Sliced Optimal Transport Control — Distribution steering seeks feedback laws that drive the state law of a dynamical system between prescribed initial and terminal distributions.
- From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options — Large language models (LLMs) perform well across a wide range of tasks but exhibit persistent weaknesses in logical reasoning.
- FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation — FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation proposes a fairness optimization framework to address a previously overlooked fairness issue in Semantic ID (SID)-based generative recommendation, termed "Token Frequency Bias," where "high-frequen
- Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents — This paper introduces the concept of skill misevolution in self-improving LLM agents, where "self-improving LLM agents convert successful trajectories into persistent cross-task state" and "an unsafe success can thereby become reusable policy after its triggering input disappears
- Beyond Source: An Empirical Study of Python Bytecode Security Risks — This paper presents an empirical study of Python bytecode as a security artifact, motivated by the observation that "Python package security is largely source-centric, yet Python runtimes can execute bytecode directly through.pyc files, compiled-only modules, and marshalled code
- Dissecting Software Graphs: Structural Insights for Driver-Guided Fuzzing — This paper presents an empirical study of software structure under multi-driver fuzzing. The authors propose a structural abstraction that uses a static call graph as a shared backbone and projects driver-specific dynamic coverage onto it to derive driver-induced subgraphs.
- AI and Consumer Rights in India Working Paper — This working paper examines whether India's Consumer Protection Act, 2019, adequately addresses harm caused by defective AI products and services, and whether it proportionately allocates liability across the AI value chain.
- Discovering Persistent Behavioural Patterns for Interpretable Blockchain Forensics — This paper proposes a scalable, application-agnostic framework for persistent behavioural pattern discovery from large-scale blockchain activity.
- A Compositional Theory of Curvature in Probabilistic Circuits — This paper introduces a compositional theory of curvature in Probabilistic Circuits (PCs), showing that a sum node's contribution to the Hessian trace of the negative log-likelihood factorizes exactly into two semantically distinct terms: contextual usage (circuit flow) and local
- ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification — ReflectFact is a novel self-reflective agent framework for multi-hop fact verification, proposed to address two critical limitations in existing agent-based methods: objective conflicts and knowledge conflicts.
- Robust data-driven discovery of fractional differential equations via weak formulations and Pareto-based subset selection — The paper "Robust data-driven discovery of fractional differential equations via weak formulations and Pareto-based subset selection" by Pongpisit Thanasutives and Yoshinobu Kawahara proposes Weak-Pareto, a data-driven discovery framework for fractional partial differential equat
- Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation — This paper audits a preserved MCP (Model Context Protocol) agent security evaluation campaign and identifies a critical construct-validity defect: the historical endpoint exhibited direct treatment leakage, where "treatment metadata gated the ATTACK SUCCESS class, so fixed behavi
- Distinguishing case-mix from context heterogeneity in prognostic regression model synthesis settings — Distinguishing case-mix from context heterogeneity in prognostic regression model synthesis settings This paper introduces a diagnostic framework to distinguish between two sources of heterogeneity in prognostic regression models synthesized across multiple sites: case-mix hetero
- Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence — This paper investigates the validity of the conditional-independence assumption (C5) that underpins compositional reliability bounds for multi-agent systems, which typically multiply component reliabilities to estimate the reliability of the whole.
- Adaptive k Nearest Neighbors Classifier via Granular Ball Computing — The paper proposes an adaptive and efficient k-Nearest Neighbor (KNN) approach via granular-ball computing, called GBKNN, to address two key limitations of traditional KNN: the computational cost in high-dimensional and large-scale datasets, and the sensitivity to the choice of t
- Prompts in the Wild: A Large Analyzed Collection of Transactional Prompts in Code — " Summary This paper argues that prompts, the instructions given to large language models (LLMs), should be treated as "first-class objects for empirical scientific and linguistic investigation" rather than informal artifacts.
- Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs — The paper "Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs" investigates privacy risks in multimodal large language models (MLLMs) used for document understanding, specifically in Key Information Extraction (KIE) tasks.
- Revisiting Overestimation Bias Problem of Q-learning: Settling Large Discrete Action Space via Action Intersection — This paper considers the overestimation bias problem of Q-learning in the setting of a large action space, for the purpose of relieving the bottleneck of existing methods. We find that the large action space increases the randomness in Q-value estimation.
- Decoupled Contrastive Decoding via Expert-Aligned Drafting — Decoupled Contrastive Decoding (DCD) is a method to accelerate Contrastive Decoding (CD) while preserving its output distribution.
- InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers — InFactPlanner is a trace-driven decision-support framework for what-if analysis of sustainable AI data center deployment for LLM inference across single and geo-distributed sites.
- Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling — This paper introduces a novel proxemics-based reward formulation for deep reinforcement learning (DRL) social navigation.
- H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities — This paper presents the first comprehensive adaptation and evaluation of Expected Threat (xT) and Valuing Actions by Estimating Probabilities (VAEP) frameworks for professional team handball, utilizing five seasons of tracking-derived event data from the Handball Bundesliga (2021
- Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence — The paper introduces a Polish-language medical visual question answering (VQA) benchmark built from Polish Board Certification Examination (PES) questions for licensed physicians and dentists pursuing specialist certification.
- AutoQuREO: A Framework for Automated Quantum Resource Estimation and Optimization — AutoQuREO is an automated framework for full-stack quantum resource estimation and optimization, introduced in this paper.
- Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency — The paper addresses a critical gap in evaluating Joint-Embedding Predictive Architectures (JEPAs) for world models. JEPAs learn world models that predict in a compact latent space rather than in pixels, reducing pressure to model nuisance appearance.
- CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation — CardioState-JEPA is a cardiac foundation model that learns a single shared representation jointly across electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG), built on a physiology-aware joint-embedding predictive architecture. [episode]
- I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization — I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization Summary This paper introduces I-SDPO (Instance-Level Adaptive Self-Distillation Policy Optimization), a method for reinforcement learning (RL) post-training of large language models (LLMs) that addresses the "d
- The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use — This paper investigates why latent world models fail at long-horizon planning, using a reproduction of LeWorldModel on the TwoRoom environment. The central finding is that the bottleneck is the planner's objective function, not the predictor's quality.
- Understanding Backdoor Vulnerabilities in Vertical Federated Learning: The Gap Between Research and Practice — This paper, "Understanding Backdoor Vulnerabilities in Vertical Federated Learning: The Gap Between Research and Practice" by Ziqi Zhao, Jialin Lu, Junjie Shan, Junyuan Zhang, Shuya Yang, and Ka-Ho Chow from the School of Computing and Data Science at The University of Hong Kong,
- Comment on "Modeling rapid language learning by distilling Bayesian priors into artificial neural networks" — This comment paper argues that McCoy & Griffiths’ (M&G) method of using Model-Agnostic Meta-Learning (MAML) to distill a Bayesian prior into artificial neural networks does not actually instill a prior under the standard interpretation, and that even under a more permissive int
- Generative Universal Multimodal Retrieval with Dual-role Identifiers — The paper proposes DrIG, a novel generative framework for universal multimodal retrieval featuring dual-role identifiers.
- Balanced Adaptive Prototype Selection for Scalable TabPFN Inference on Large-Scale Tabular Data — This paper introduces Balanced Adaptive Prototype Selection (BAPS), a framework for constructing compact, information-preserving contexts for scalable TabPFN inference on large-scale tabular data.
- ATOBench: Tracing How Autonomous Penetration-Testing Agents Verify Vulnerabilities When Target Evidence Lies — ATOBench introduces an evaluation framework that makes the verification process of autonomous penetration-testing agents observable under deceptive target responses. The paper states: "Autonomous penetration-testing agents rely on target responses.
- OmniSphinx: Active Mix Networks (Extended Version) — OmniSphinx is a novel mix format that applies ideas from active networking to make mix networks flexible. In OmniSphinx, senders embed code (called "mix programs") in their packets that determines how they must be processed at each node.
- Incremental Evaluation and Training in Relational Deep Learning — Authors: Jakub Peleška and Gustav Šír, Czech Technical University in Prague Paper: arXiv:2608.13023v1 [cs.LG], 13 Aug 2026 --- Relational Deep Learning (RDL) models multi-tabular databases as temporal heterogeneous graphs to enable end-to-end representation learning.
- Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language — Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be written in an now-unfamiliar language (e.g.
- UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations — UniTraffic-Agent is introduced as the MR-CAS solution for Track 3 of the 10th AI City Challenge, which includes the main Traffic Anomaly Reasoning (TAR) task and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning.
- On the global feature importance for interpretable and trustworthy heat demand forecasting — The paper introduces an ante-hoc Explainable AI (XAI) methodology to assess the global feature importance of Machine Learning models used for heat demand forecasting in intelligent control of District Heating Systems (DHS), with the motivation to facilitate their interpretability
- BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs — BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs Abstract Summary: Organizational decisions are co-created while evidence, constraints, and human priorities continue to change.
- Operationalizing Cyber Threat Intelligence with GraphRAG — This paper investigates whether using a knowledge-graph-based retrieval system (Microsoft GraphRAG) instead of a standard vector-similarity retrieval system (Naive RAG) produces threat hunting plans that rely more on durable, hard-to-evade clues (TTPs) rather than fragile indicat
- Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) — This study tested whether a language model's explanatory engagement with a rare tool failure rises and then collapses as the failure is made asymptotically rarer.
- Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds — This paper challenges the default paradigm of aligning large language models (LLMs) as passive, sycophantic assistants by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavioral reprogramming.
- A Multispectral Framework for the Detection of Calcium Carbide-Induced Ripening and Shelf-Life Estimation in Climacteric Fruits — The proposed study explores a novel, non-invasive multispectral framework for distinguishing safely ripened fruits (naturally ripened and ethephon-induced) from calcium carbide-ripened samples, while also estimating their ripening progression (in percentage) and remaining shelf l
- SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference — SPADE: Speculative Decoding for Precise and Low-Cost Distributed Edge–Cloud Inference Abstract Summary: The paper addresses the challenge of deploying Large Language Models (LLMs) constrained by high computational demands.
- LOB-ID: Evaluating Synthetic Market Data by Inception Distances — The paper introduces LOB-ID (Limit Order Book Inception Distance), an embedding-based framework for evaluating synthetic limit orderbook (LOB) data.
- Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization — This paper investigates whether instance-wise allocation of a fixed test-time sample budget in neural combinatorial optimization (NCO) solvers provides measurable gains over the conventional uniform allocation, and audits the measurement procedure itself. Research questions.
- FlowLOB: Efficient and Controllable Limit Order Book Generation with Flow Matching — FlowLOB is a conditional flow-matching generative model for limit order book (LOB) trajectories. It is trained on multiple Hong Kong Exchange (HKEX) symbols at three sampling frequencies (0.1s, 1s, 10s) in a tick-relative representation that transfers to unseen instruments.
- Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting — This paper proposes a unified evidence reasoning framework that addresses two persistent challenges in multi-source evidence fusion under Dempster-Shafer theory: existing conflict measures assess inter-evidence inconsistency and intra-evidence uncertainty independently, yielding
- Branch and Bound for Relational Verification of Neural Networks — Authors: Kota Fukuda, Zhenya Zhang, Guanqin Zhang, and Jianjun Zhao Affiliation: Kyushu University, National Institute of Informatics, UNSW Sydney --- The paper addresses the verification of neural networks against relational specifications, particularly global robustness, which
- SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback — The paper addresses the challenge of continual improvement for Agent Skills—portable modules that encapsulate domain knowledge and handling procedures in customer support systems.
- Huber-Wasserstein barycenters for robust distribution-valued data — Authors: Carlos Cardoso-Perelló and Alberto González-Sanz (Department of Statistics, Columbia University) arXiv: 2608.13131v1 [stat.ME] 13 Aug 2026 --- The paper proposes a robust barycenter for distribution-valued data by incorporating the Huber loss directly into the optimal
- Statistical Properties of Robust Learning under Distributional Shifts — This paper studies the statistical properties of robust learning methods—specifically Distributionally Robust Optimization (DRO) and Robust Satisficing (RS)—when there are distributional shifts between the source environment (where training data is generated) and the target e
- A Commitment-Based Hybrid Post-Quantum Cryptographic Model for Multi-File Cloud Storage — This paper presents a commitment-based hybrid post-quantum cryptographic model for authenticated multi-file upload to cloud storage.
- MergeOver: Post-Training Token Merging for Recursive Vision Transformers — MergeOver is a post-training approach that integrates Token Merging (ToMe) into the recursively weight-shared Sliced Recursive Transformer (SReT), without retraining.
- LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service — LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service Abstract Summary: As edge-side vision services continue to expand toward low-latency, high-throughput scenarios, reducing the inference cost of vision models without sacrificing reliabi
- Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering — Summary This paper introduces Syfer, a synthesizer-folding framework for multilingual multi-hop question answering (QA).
- Which LLM Is Your Ideal Companion? Evaluating Emotional Companion Capabilities of LLMs Based on Adult Attachment Theory — Based on the paper "Which LLM Is Your Ideal Companion? Evaluating Emotional Companionship Capabilities of LLMs Based on Adult Attachment Theory," here is a detailed summary: Introduction and Motivation The paper addresses the growing use of large language models (LLMs) in emotion
- High-dimensional networks and mean squared error for possibly misspecified models — The paper "High-dimensional networks and mean squared error for possibly misspecified models" by Lourens Waldorp addresses the challenge of estimating networks (Gaussian graphical models) when the number of parameters (edges) exceeds the number of observations (high-dimensional s
- SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents — SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents Summary This paper introduces SkillShapley, a framework for step-level attribution in LLM agent skills, addressing the open problem of quantifying how each individual step within a skill co
- Smart Contract Invariants Protect Against Cybercriminals — This paper investigates whether smart contract invariants can protect against real-world cybercriminal attacks on blockchain systems, and whether automated tools can discover such attack-stopping invariants.
- NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video — NARU is a benchmark designed to evaluate narrative evolution and cultural understanding in Japanese extreme long-form video. It consists of 1,481 multiple-choice questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions.
- TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems — TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems Abstract Summary The paper introduces TsuGO, a process-level reasoning benchmark for evaluating Search Efficiency in LLM reasoning through Go life-and-death problems.
- Foundations of Independent Component Analysis — This paper presents a comprehensive, self-contained mathematical treatment of the foundations of linear Independent Component Analysis (ICA), aimed at readers with a background in measure-theoretic probability theory.
- Knowledge-guided Pattern Discovery via Coupled Tensor Factorizations — This paper introduces a knowledge-guided approach for pattern discovery that brings together data and computational models by jointly analyzing real data and simulated data (generated using a computational model) using coupled tensor factorizations with linear coupling.
- When Should Multi-Round RAG Stop? Structured Stopping Judgments and Retrieval Reduction in Search-R1 — This paper addresses the stopping problem in multi-round retrieval-augmented generation (RAG): deciding when to stop searching as evidence accumulates.
- Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales — Normative datasets are often used to train and align AI systems, but the norms they contain can function as action-guiding patterns rather than neutral moral knowledge.
- GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport — GeoCache is a training-free acceleration method for multi-view texture diffusion models.
- Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data — The paper "Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data" presents a comparative analysis of generative models for transcriptomic data, investigating strategies to incorporate prior biological knowledge via gene graphs to ensure synthetic data captur
- Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models — Self-referential prompting has been shown to reliably induce large language models to produce first-person reports resembling subjective experience, but no prior work measures how consistent these reports are across repeated, independent trials, or how that consistency compares t
- vToken: Token-Level Virtualization for Reclaimable KV Caches — vToken is a lightweight token-level virtualization layer that decouples logical token liveness from physical block placement in KV cache management for LLM serving systems.
- Slow and Steady: Preventing MEV with Verifiable Delays — This paper presents a defense mechanism against Maximal Extractable Value (MEV) opportunities in distributed ledgers, which relies on enforcing a verifiable delay when generating transactions.
- Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization — The paper "Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization" addresses the challenge of estimating causal effects on post-click conversion rate (CVR), defined as P(Y2 = 1 Y1 = 1), where Y1 indicates whether a user clicks and Y2 indicates whether conve
- MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification — The paper proposes ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs. ARMDIL is an ensemble that uses a multimodal large language model (MLLM) agent to dynamically route each image to the most suitable vision backbone.
- Concept Drift Detection and Adaptive Retraining of Malware Classification Models — Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning model.
- MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination — MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination Abstract Summary: The paper presents Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestratio
- TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval — TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval This paper addresses the challenge of efficiently retrieving relevant clips from large-scale driving logs, which is essential for data curation, model development, and safety analysis.
- Sparse Orthogonal Regression Technique: A Spectral Framework for Equation Discovery, Approximation, and Integration — The paper develops the Sparse Orthogonal Regression Technique (SORT), "a sparse spectral framework for learning orthonormal-basis expansions from noisy and irregularly sampled data." SORT "estimates expansion coefficients directly from observations using L1-regularized regression
- On the Structural Limits of Machine Learning Decision Systems: An Information-Theoretic, Interaction-Based, and Stochastic-Dynamical Perspective — This paper examines the structural limits of machine learning decision systems from information-theoretic, interaction-based, and stochastic-dynamical perspectives.
- Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining — This paper proposes a task-agnostic measure of training data influence for language model pretraining. The key innovation is reformulating training data influence without requiring a downstream task or validation set as the attribution target.
- DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data — DFM Mimir v1 is a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture, trained from scratch and delivering "highly competitive performance for English and sets a new state of the art for Danish using only permissible post-training data.
- Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology — The paper proposes an intervention-aware clinical world model for forecasting post-operative outcomes in cardiology, specifically applied to atrial fibrillation (AF) ablation.
- The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity — The paper introduces the unmasking growth complexity (UGC) as a path-resolved measure of data geometry for masking diffusion, and shows that its local increments directly control Kullback–Leibler (KL) discretization error.
- SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization — SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization Summary This paper introduces SAEVerbalizer, a framework that fine-tunes large language models (LLMs) to generate natural-language explanations of sparse autoencoder (SAE) feat
- LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure — LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure introduces LittleCurriculum, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5.
- Defensive Boosting for Online Probabilistic Forecasting — The paper studies online probabilistic forecasting of binary outcomes chosen by an adaptive adversary.
- The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis — The paper proposes a YOLO- and Contrastive Language-Image Pre-training (CLIP)-based vision-language framework to classify mosquito flight frames of uninfected and Dengue virus serotype 2 (DENV2)-infected mosquitoes.
- Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test — The paper introduces a finite capability sheaf to model failures in AI agent harnesses, where locally successful components disagree on shared state.
- TabSOM: A tabular-to-image encoding method based on self-organizing maps — TabSOM: A tabular-to-image encoding method based on self-organizing maps Abstract Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers.
- Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension — The paper presents the organizational framework adopted by the Academic League of Artificial Intelligence (LIA) at the Federal University of Santa Catarina (UFSC), designed to integrate teaching, research, and university extension through a student-centered, project-based approac
- A Progressive Design Study of Visual Encoders and Value Estimation for Replay-Free Parallelized Q-Learning — This paper, authored by Taha Shieenavaz, Shabnam Zareshahraki, and Loris Nanni from the Department of Information Engineering, University of Padua, Italy, presents a systematic investigation into the architectural design space of Convolutional Neural Networks (CNNs) for the Paral [episode]
- iARCS: Iterative Agentic RL for Controllable 3D Scene Generation — The paper presents iARCS, an "iterative agentic reinforcement learning framework that adapts a pretrained scene generator to natural-language task requirements." The work addresses a core limitation in synthetic 3D scene generation: "existing generators often optimize perceptual
- Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model — Mixture of Training (MoT) is a scaffolded modular pre-training procedure that partitions a target Transformer into contiguous layer blocks, trains each block inside a frozen pretrained aligner scaffold, and then recomposes the trained blocks with an optional short end-to-end adap
- Multi-perspective Imbalance-Conscious 6G Beamforming Optimization and Performance — The study presents a systematic machine learning (ML) study of 6G-IoT beamforming optimization (6GBO) using supervised and unsupervised approaches. We compared the predictive power of network, environmental, device, and vision feature groups for 6GBO.
- CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model — CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model Abstract Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners’ speaking performance.
- RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level — This paper addresses the challenge of assessing the maturity of artificial intelligence technologies, which is essential for investment decisions, project management, and policy monitoring.
- EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction — EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction Abstract RNA-Protein Interactions (RPIs) are critical for regulating cellular functions.
- Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs — This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance communication for individuals with hearing impairments.
- Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision — This study presents the first use of egocentric vision for freezing of gait (FOG) detection in Parkinson's disease (PD) during home-based activities of daily living (ADLs).
- Learning the Mathematical Property for Designing Low Mutual Coherence Binary Sensing Matrices — This paper proposes a novel learning-based framework for constructing binary sensing matrices with low mutual coherence for compressive sensing applications.
- Beyond Simulated Benchmarks: Evaluating Motion Representations for Fall Detection Under Real-World Data Scarcity — Falls are a major health concern for older adults, and wearable sensors have been widely explored for detecting falls and enabling timely intervention.
- Moose: Latent concept learning with reasoning-shortcut awareness in EL++ — Moose: Latent concept learning with reasoning-shortcut awareness in EL++ Abstract. The OWL 2 EL profile is used in some of the largest production ontologies, including the Gene Ontology and SNOMED CT.
- Multi-Layer Context Camouflaging: A Semantic Superposition and Contextual Lamination Framework for Malpractice-Resilient Online Assessment — This paper extends the Multi-dimensional Spatio-Temporal Context Camouflaging Model (MSCCM) introduced in prior work on the MARS assessment-resilience suite, developing a substantially more detailed mathematical treatment of its Context Camouflaging Operator.
- Efficient Hessian-Free Methods for Multi-Objective Bilevel Optimization with Nonconvex Lower Level — This paper addresses multi-objective bilevel optimization (MOBL) problems with nonconvex lower-level objectives. The problem is formulated as: > min F (x):= [fi (x, y ∗ (x))]m i=1, x∈Rdx s.t.
- Uniform Herding: Exemplar Replay with Representation Refresh — Uniform Herding: Exemplar Replay with Representation Refresh proposes a replay-based class-incremental learning method that refreshes exemplars in the current feature representation.