AI papers — 2026-08-12
Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment. This paper proposes a hybrid planning architecture for automated driving that... Distillation of Foundation Models for Time-dependent PDEs. Teacher Rollout Extension (TREX): Distilling Foundation Models for... Persistent Recursive Worlds Enable Autonomous Software Evolution. EvoX Genesis (hereafter, Genesis) is a system that reorganizes long-horizon... LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence. LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured Reinforced... Interpretable Causal Discovery via Causal-Effect Constraints. Interpretable Causal Discovery via Causal-Effect Constraints Authors: Cixuan... Disentangling the Expressivity of RoPE. This paper, "Disentangling the Expressivity of RoPE" by Selim Jerad, Anej... Fingerprinting Text-to-Image Diffusion Models via Collapsed Generation. This paper presents a non-invasive model fingerprinting framework for... Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces. Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Authors:... Automated binary classification of hazelnut X-ray images: A deep-learning benchmark for quality assessment. This study presents a benchmark for binary hazelnut quality classification... CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network for End-to-End Image Generation. CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network... Training Under Challenge: Executable Certificates and Challenge-Closed Optimality for Neural Networks. The paper addresses a fundamental limitation in neural network training: "A... Non-Degenerate Risk Certification for Automated Security Decisions: A Decision-Contract Theory with ATT&CK-Aligned Triage as a Worked Instance. Non-Degenerate Risk Certification for Automated Security Decisions: A... Exemplar-based objective classification of gust-induced loads across multiple flight conditions. Objective and Motivation This paper investigates whether an objective... -MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution. ε-MemEvo is a framework for cross-task knowledge transfer in LLM-based program... HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference. HetRoute: Heterogeneous and Cost-aware Collaborative Routing Framework for... Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models. Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language... ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents. ToolHazard is a scalable adversarial environment synthesis framework designed... Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning. This paper studies transfer learning in cooperative multi-agent reinforcement... AVA-Encoder: Towards Agent-Native Video Representation Learning. AVA-Encoder: Towards Agent-Native Video Representation Learning Summary This... HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation. HSTGFormer is a graph-enhanced Transformer framework proposed for efficient... Causal Structure is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library. Causal Structure Is Inducible but Functionally Decoupled: The Routing/Readout... Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges. Graph-Structured Rubrics (GSR) compiles a rubric into a response-independent... Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem. This paper takes a first step toward a unified formal account of agent memory,... Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines. This paper presents a systematic five-step methodology for deploying custom... JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis. The paper introduces Ancient Chinese Character Exegesis (ACCE), a... SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization. SSPO (Structure-Aware Similarity-Weighted Preference Optimization) is a method... VICBench: A Multi-Language Benchmark for Code Vulnerability Detection. VICBench: A Multi-Language Benchmark for Code Vulnerability Detection This... An Efficient Near-Optimal Algorithm for Adversarial-Set Bandits. Problem Setting The paper studies adversarial combinatorial bandits with m-set... Direct Acceleration of Stochastic Root-Finding Without Variance Reduction and Regularization. Problem and Motivation The paper addresses stochastic root-finding problems,... Redistribution-based Cost Inference Improves Sparse Safe Offline RL. The paper introduces the Redistribution-based Cost Inference (RCI) framework,... Draw This First. The paper "Draw This First" presents a method for generating ordered vector... Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning. The paper addresses a central challenge in LLM personalization: task-specific... A Local Sinkhorn Framework for Conditional Distribution Reconstruction of Multidimensional Random Fields. This paper proposes a local Sinkhorn divergence framework for conditional... Continuous-Latent Predictive Modeling with Semantic Alignment for EEG-Language Foundation Models. The paper proposes the Brain Latent Predictive Model (BLPM), an EEG-language... Foresight Without Seeing: Latent Futures for World Action Models. ForeWAM is a dynamics-conditioned direct-policy World Action Model (WAM) that... Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder. Confucius4-TTS is a multilingual zero-shot text-to-speech (TTS) system that... The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance. The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance... Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs. The paper introduces an adversarial reinforcement learning framework to expose... DiG-bench: Discovery in Games. DiG-bench (Discovery in Games) is a new benchmark introduced to measure the... A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions. The paper presents a formal, three-step framework for designing human-aligned... Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision. Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise... Confidence Calibration of Deep Learning Systems. This thesis explores novel methods for improving confidence calibration under... Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization. Most image colorization systems operate in Lab space by predicting chroma (ab)... When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide. This paper investigates when offline evaluation of equal-cost top-k allocation... Fast Length-Squared Sampling for Positive-Semidefinite Matrices. This paper presents a simple rejection-sampling-based algorithm for performing... NAE: Normalizing AutoEncoder. The paper introduces the Normalizing Autoencoder (NAE), a generative framework... How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging. This paper evaluates the complete unsupervised domain adaptation (UDA) pipeline... LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration. LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion... Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams. Diagram-MMU is a multi-modal benchmark designed to assess Multimodal Large... ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening. ScreenShot is a hierarchical transformer-based foundation model for few-shot... Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus. This paper presents the first systematic study of massive activations (MAs) in... Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular Missingness. Paper: "Represent, Then Generate: Multimodal-Conditioned Time-Series Generation... From Visual Widgets to UI Code: Efficient Tool-Grounded Generation. Existing screenshot-to-code systems face a trade-off between flexibility and... AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Core Research Question and Setting This paper, from Salesforce AI Research and... Unifying Physical Backpropagation. Core Contribution This paper develops a unifying theoretical framework for... Certifying What Helps Customer-Return Timing: A Screen-and-Confirm Test for Conditioning Signals, and Why Decay Is Nearly Enough. This paper introduces a screen-and-confirm protocol to certify whether a... Dion3: Full-Stack Orthogonal Updates. Dion3 is a revision of the Muon optimizer that targets overhead at every level... Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation. This paper challenges the standard assumption that LLM model rankings are... SoftWater: Class-Aware Rate Allocation for Softmax Quantization. Based on the paper, here is the summary: The paper introduces SoftWater, a... Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill. Spark-to-Paper is an end-to-end research paper generation system implemented as... Scaling Automatic Research Agents via World Models. The paper "Scaling Automatic Research Agents via World Models" introduces World... Structure-preserving uncertainty quantification for GENERIC dynamics. Structure-Preserving Uncertainty Quantification for GENERIC Dynamics Authors:... From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices. The paper argues that the central research problem for using large language... CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution. CAKE: Compiler–Agent Co-Design for Frontier Kernel Evolution presents a... Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System. This paper introduces UAVQA-Bench, a comprehensive, fully human-annotated... G0.5: One Autoregressive Stream for Robot Reasoning and Action. Galaxea G0.5: One Autoregressive Stream for Robot Reasoning and Action Core... Claim-Level Reliability Assessment for Efficient Test-Time Reasoning. Here is a summary of the paper: The paper introduces Claim-Level Reliability... Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework. This paper presents a structured, benchmark-based comparative assessment of... TELLME: Test-Enhanced Learning for Language Model Enrichment. The paper introduces TELLME (Test-Enhanced Learning for Language Model... How Organizations Use AI: Evidence from ChatGPT. This paper studies how organizations use frontier generative AI by linking... Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence. The paper introduces the “Agentic Self-Improvement” framework, a... Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models. This paper introduces a framework for the automated construction of Dynamic... Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review. This paper reports a single, fully instrumented case study of a large-scale... Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion. The paper introduces GDCE-I (Graph Diffusion Counterfactual Explanation via... Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents. Governed Persistent Memory (GPM) is introduced as an auditable bitemporal... Small-Scale Experiments: Are We There Yet?. Small-Scale Experiments: Are We There Yet? Nicholas Lourie, Kyunghyun Cho,... Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL. Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL... Locating and Controlling Implicit Personalization in Large Language Models. Large language models (LLMs) often shift their outputs in response to implicit... The Advective Fisher-Rao Geometry of Deterministic Measure Transport. The paper introduces a novel Riemannian metric, the advective Fisher–Rao... High-Order Liquid Evidence Encoding for Gradual GNSS Spoofing Detection in Autonomous Driving. The paper proposes a causal high-order liquid evidence framework for detecting... Unifying Generative Models with Path Integrals. This paper formulates generative modeling as a path integral, unifying... AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention. AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive... FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents. FrontierFinance is a fully open benchmark introduced by Samaya AI for... Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment. This paper introduces Group Alignment-induced Sycophancy (GAS), a two-sided... When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation. Core Problem The paper investigates a resource-shared evaluation loop in... Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL. This paper presents a framework for learning loco-manipulation policies using... AmbSentry: Mitigating Sensing Eavesdropping in ISAC Systems by Harnessing Ambient IoT Devices. AmbSentry: Mitigating Sensing Eavesdropping in ISAC Systems by Harnessing... Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling. This paper investigates whether on-policy distillation (OPD) truly expands the... Towards Scalable Fuzzy PSI via Efficient Fuzzy Matching. Based on the paper, here is a detailed summary: This paper presents scalable... One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL. Based on the paper, here is a detailed summary: Summary This paper identifies... Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting. Thyroid ultrasound diagnosis requires coordinated lesion localization,... MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents. MindMemOS is a portable and self-evolving memory operating layer for AI agents,... XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication. XB RIDGE: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication... Trie Automata for Constrained Decoding over Large Finite Sets. Trie Automata for Constrained Decoding over Large Finite Sets Authors: Xingzi... Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads. This paper introduces a unified theoretical framework for understanding... TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement. TailBooster is a dual-layer generative framework designed to address two... Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs. The paper introduces Self-Fix Step-DPO (SFS-DPO), a two-stage reinforcement... Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting. Motion-centric video reasoning is fundamental to interactive applications such... IoT-Enabled Autonomous Maritime Navigation in Smart Ports: A Curriculum-Guided Shared Policy Learning Framework. This paper investigates onboard autonomous navigation for IoT-enabled... Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents. Harness-IF is a benchmark that turns operational instruction following into a... When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use. This paper investigates a failure mode in multilingual API calling called... Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation. The paper studies reference-free post-training for multilingual machine... Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction. Based on the paper, here is the summary: This paper introduces DARC... Robust and Efficient Noisy-Label Time-Series Classification via Dynamic Time Warping Based Granular Ball Computing. Dynamic Time Warping (DTW)-based Nearest-Neighbor (NN) classifiers are... DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation. DexterSQL is a prompting-based (non-fine-tuning) Text-to-SQL system that... M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation. M-Net: Integrating Spectral Features and Physical Field Operators into Deep... Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning. The paper presents a multi-tier, frame-level audio tagging framework for... CT- Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models. CT-∆Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference... Sparse and robust geometric twin support vector machine via asymmetric RoBoSS loss function. This paper proposes a novel asymmetric, robust, bounded, sparse and smooth (aR)... MARCH: Scaling Recurrent Memory with Content-Routed State Anchors. MARCH: Scaling Recurrent Memory with Content-Routed State Anchors Abstract... Continual Learning in Transition. and Motivation This survey paper, authored by Zhiyan Hou, Dan Zhang, Tao Feng,... Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents. Here is a summary of the paper: The paper "Beyond Single-Turn Confidence:... When Do Anchor-Based Pointwise LLM Rerankers Help? Retriever Quality, Statistical Scope, and Anchor Design. and Research Questions This paper investigates when anchor-based pointwise LLM... Poly-Dialectal Neural Machine Translation System for Bangla Regional Dialects. This paper presents a unified Poly-Dialectal Neural Machine Translation System... JAPE: Joint Anomaly Prediction and Intrinsic Explanation in Multivariate Time Series. JAPE: Joint Anomaly Prediction and Intrinsic Explanation in Multivariate Time... HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry. HyperANFIS is a hyperbolic extension of the adaptive neuro-fuzzy inference... Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance. This paper investigates how prompt construction methodology affects the... RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation. RT-SEMamba is a fully causal speech enhancement (SE) model built upon causal... MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques. MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language... Coarsening Latent-Class Probabilities: Directional Distortion and Coverage Loss. Coarsening Latent-Class Probabilities: Directional Distortion and Coverage Loss... Rank-Two Frobenius-Linearized Normal Forms and Orthoderivative Dual Coordinates in Quadratic APN Maps. This paper classifies binary-linear two-term Frobenius-linearized operators of... Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks. This paper addresses the question of how large a random low-dimensional search... SchemaLink: An Intelligent Web Editor for LinkML Schema Curation. SchemaLink is a web-based environment for the graphical construction and... Learning Under Treatment-Induced Label Indeterminacy with Expert Annotations of Counterfactual Outcomes: A Case Study in Neurological Prognostication. This paper addresses the problem of treatment-induced label indeterminacy in... What Makes a Peer? Valuation-Anchored Similarity in Private Markets. The paper introduces a supervised similarity learning framework for identifying... SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries. SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries Abstract... @skills: Attention is all you have. Core Problem The paper identifies a fundamental mismatch in the agent skills... Intensional Anaphora. The paper "Intensional anaphora" by Ezra Keshet and Steven Abney (2024)... DIVE: Unlocking Self-Improvement in Frozen Language Models Through Diversity-Driven Skill Evolution. DIVE: Unlocking Self-Improvement in Frozen Language Models Through... Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts. This paper proposes a personalized scorer modeling framework for generating... Exploring Oversmoothing with Householder Matrices. Core Contribution This paper introduces HouseGNN (Householder Graph Neural... A Hybrid Framework of Vision Transformer and Gated Recurrent Unit for Detection of Mosquito Diseases. This paper introduces a three-step framework for detecting dengue and Zika... Don't Cut Corners: How Training Outside the Prior Makes Simulation-Based Inference More Robust. Large astrophysical simulation campaigns often generate training data by... A Remote Approach to Cashew Orchard Detection: Leveraging Active Learning with Satellite Imagery in Guinea-Bissau. This study presents the first openly accessible, countrywide cashew orchard map... A 12-CNOT Double Qubit Excitation Gate. The paper presents the first reported 12-CNOT decomposition of the double qubit... General Probabilities of Causation with Causal Knowledge. This paper addresses the question of whether additional causal knowledge can... Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues. Problem: Behavioral Relapse of Revoked Constraints This paper studies a failure... Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection. Problem Statement Deep learning detectors for anomalies in dynamic graphs have... Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction. Core Contribution The paper introduces Constraint Saturation Evaluation (CSE),... Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence. The Wiggle Framework is a unified stress test for epistemic stability in LLM... EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory. EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon... Novels generated by language models show compressed formal variation. This study investigates whether large language models (LLMs) can produce the... LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning. This paper introduces EpicStar, a framework for enhancing strategic reasoning... Supervised Mixed-Frequency Learning for Macro-Financial Forecasting When Factors are Weak. This paper, "Supervised Mixed-Frequency Learning for Macro-Financial... LookBack: Where and How to Score LVLM Responses via Visual Reference Usage. L OOK BACK is a training-free LVLM response scoring method that augments token... HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression. HAMP-LIC is a Hessian-aware mixed-precision post-training quantization (PTQ)... Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models. Core Contribution This paper introduces Semantic Lenia, a framework that... Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models. This paper investigates where safety-aligned refusal behavior is encoded in... Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages. This paper examines the structural barriers that disadvantage speakers of... A Conceptual Framework for Enhancing Workforce Readiness for Smart Manufacturing in the AI Era. This paper introduces the Workforce Readiness Level (WRL) framework, a... LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation. LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global... Machine Learning-Based Cyber Defense for Cloud Infrastructure: An Adaptive Deep Q-Network Architecture for Intelligent Intrusion Detection and Automated Threat Mitigation. This paper proposes a reinforcement learning-based dynamic cyber defense... Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability. This paper develops a mathematical and engineering architecture for secure... The energetic cost of mitigating AI attacks in cellular networks. The integration of Artificial Intelligence (AI), generally as Machine Learning... Learning from Online User Feedback for Shopping Agents. LOFA, a framework that enables shopping agents to learn directly from real... Slips: Behavioral Evidence Aggregation for Network Security. Slips is a network intrusion detection system that builds host-centered... Evaluating AlphaEarth Foundations Embeddings for Wildfire Susceptibility Mapping. This paper systematically evaluates AlphaEarth Foundations (AEF) embeddings for... Large Language Model-Driven Small-Capitalization Trading: Integrating Financial News Sentiment, Macroeconomic Indicators, and Technical Signals. This paper studies small-capitalization trading with LLM-derived news... Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation. and Motivation This paper investigates how large language model (LLM) agents... Plaintext Recovery Against Post-Filtering Access Control. This paper, "Plaintext Recovery Against Post-Filtering Access Control" by... A Cascaded Unsupervised-Supervised NLP Pipeline for Detecting Accusatory Language in Public Procurement. This paper proposes a cascaded hybrid learning strategy that integrates... Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems. This paper benchmarks the serving cost of agentic memory systems for... FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees. FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision... Autonomous Telerehabilitation via Skeletal Motion Prediction and Joint-Level Performance Assessment. This paper presents a telerehabilitation pipeline that integrates... Look What the Probes Dragged In! Real-World Chest X-ray Shortcuts in MedCLIP. The paper investigates how real-world shortcuts manifest across different... DCM Bandits: Multiplayer Information Asymmetric Cascading Bandits for Multiple Clicks. DCM Bandits: Multiplayer Information Asymmetric Cascading Bandits for Multiple... Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction. Accurate landmark localization in medical images is a fundamental step for... Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release. This paper presents a fully preregistered, end-to-end stress test of the... Can Vision Models Read the Radar Display? On the Feasibility of Radar Imagery for Air Traffic Complexity Estimation. This paper investigates whether radar display imagery, despite its atypical... Forward and Inverse Virtual Metrology for Phototransistor Gain: A Hierarchical, Uncertainty-Aware Approach for Small Production Datasets. The paper addresses the challenge of predicting and optimizing the current gain... Orientation, not magnitude: the causal structure of task-vector interference in merged language models. This paper investigates the causal structure of task-vector interference in... Kernel Methods for Learning Operators with Multiple Inputs and Outputs. This paper introduces a general kernel-based encoder-decoder framework for... Transferable Above-Ground Biomass (AGB) Estimation Model from Multi-Sensor Data with Sparse Field Calibration. The paper presents an operational framework for wall-to-wall above-ground... Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks. This paper introduces BMAT (Bilevel-Minimax Adversarial Transfer), a unified... Latent variable models for simultaneous EOV identification and removal in population-based SHM. The paper introduces a Gaussian-process latent EOV (GLEOV) model for... DYSANOS Generative Dynamic Smooth Arbitrage-free Non-parametric Option Surfaces. DYSANOS is the first generative market model for smooth SANOS option surfaces... RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning. Training Mixture-of-Experts (MoE) models for reinforcement learning (RL)... GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs. GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs... Clustered Randomized Smoothing for Stochastic Prediction Functions. Clustered Randomized Smoothing for Stochastic Prediction Functions proposes a... Inferential Capability Does Not Determine Legal Scope. The paper argues that the term "inference" performs two distinct legal... Calibration Bets on the Past: Post-Training Quantization for Financial Time-Series Forecasting. This paper presents a systematic study of activation calibration for... Asymptotic Risk Calibration for Selective Question Answering. The paper proposes A-CRC-QA, a post-hoc calibration framework for... LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining. LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for... LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification. LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and... GENADA: efficient generative time series adversarial attack framework. GENADA: efficient generative time series adversarial attack framework This... A Comparison of Malware Image Transformations Using Grad-CAM and Hybrid Learning Models. This paper investigates the use of Gradient-weighted Class Activation Maps... Analysis of Motor Signatures of Social Adaptation in Autism for Efficient Human-Centric Systems. This paper proposes a computational analysis framework to identify potential... Hybrid Gated Attention. The paper proposes a Hybrid Gated Attention (HyGA) framework to extend the... An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS. The paper proposes and evaluates a supervised agentic workflow for modernizing... FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting. FM-LLM (Frequency-Enhanced Mixture-of-Experts for adapting LLMs to Time Series... Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents. Convergent Detour Hijacking: Task-Preserving Resource Amplification in... GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings. GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in... Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence. This paper introduces Mechanist, an agentic framework that uses AI as a... Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection. The paper presents BENCH2ROBUST, a framework that converts failure-free... How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models. This paper addresses the challenge of efficiently using limited oracle budgets... A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench. This paper evaluates VITA, a retrieval-augmented generation (RAG) system... HCGRec: Hint-Conditioned Generative Recommendation with Semantic IDs. Problem Statement Semantic-ID generative recommenders represent each item as a... Accuracy and Order Sensitivity Diverge Under Label-Free Strategies. This paper investigates whether preventing a model from seeing option labels... Causal inference for group-contaminated structured outcomes: observable quotients, lossless reduction and exact randomization inference. This paper studies causal inference for structured outcomes (such as microscopy... Towards Model-based Run-time Cybersecurity: On Control-Flow Anomaly Detection, Attack Identification, and Hardware Monitoring. The paper "Towards Model-based Run-time Cybersecurity: On Control-Flow Anomaly... VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies. VAKRA (eValuating API and Knowledge Retrieval Agents) is a benchmark introduced... HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks. HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold... RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks. RealisticTritonBench is a benchmark for evaluating LLM-based Triton kernel... Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness Classification. SGNet (Spectral-Grouped Network) is a lightweight architecture for... Proportional Analogies on Probability Distributions via Bayesian Updating. The paper introduces a new notion of proportional analogy for probability... Consolidator: Learning Persistent Routed Memory Across Context Boundaries. The paper introduces Consolidator, a shared slot-local operator in a Phasor... Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated Research. The paper presents Nutrition Data Service (NDS), source-preserving... User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling. The paper proposes a collaborative distributed inference system that combines... Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing. The paper introduces HPSE (Hybrid-Policy Self-Editing), a method designed to... Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models. Better Slots, Better Worlds: Representation Quality & Robustness in... CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations. CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic... Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models. Self-Generative-Understanding (SGU): A Semantic Closed-Loop Evaluation... High-dimensional Multi-objective Bayesian Optimization with Learned Variable Interactions. This paper proposes ViaMOBO, a high-dimensional multi-objective Bayesian... Adversarial Resilience of Poisson-Process Submodular Maximization over Matroids: From Robust Offline Optimization to Full-Bandit Learning. This paper studies the problem of nonnegative submodular maximization subject... No One to Blame: A Framework of Constitutive AI Unaccountability. This paper introduces the concept of "constitutive AI unaccountability" to... ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models. ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language... Two-Stage Deformable-Convolutional Inverse Design of Nanophotonic Absorbers from Optical Spectra. This paper presents a two-stage deformable-convolutional framework for the... Instruction Alignment for Binary Code Representation Learning. This paper proposes InsnAlign, a training approach that leverages... Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning. This paper introduces the concept of trait-induced safety variation, a failure... Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents. Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents... Few-Shot Ordinal Learning for Day-Wise Freshness Estimation with Hyperspectral Fish Images. This paper introduces the first few-shot learning framework for hyperspectral... MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning. MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning proposes a... Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations. Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead... Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents. This paper presents a comprehensive empirical study of skill-induced agent... How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment. How China-Origin Vision–Language Models Move from Refusal to Reframing in... The Sleeping Agent: What Gist-Based Context Compression Loses and Why. The paper introduces Salience-Weighted Consolidation (SWC), a... Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation. This paper addresses the challenges of open-vocabulary instance segmentation... CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement. CLAIM: Leading Open-domain Active Clarification of Large Language Models with... Hierarchical Federated Transfer Learning in Digital Twin-Based Vehicular Networks. This paper proposes a novel Hierarchical Federated Transfer Learning (HFTL)... RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving Simulation. RoadWeaver is a coarse-to-fine framework for from-scratch generation of... SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward. SCOUT (Structured Chain-Of-Thought Utilizing Process-Supervised RL Training) is... From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models. MPAR-Bench: Evaluating Multi-Point Associative Reasoning in Large Language... Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse. CUE-Bench is a Chinese Unsaid Emotion benchmark that centers on Affective... Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization. The paper introduces a cost-aware method for evolutionary optimization of LLM... AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection. AgenticTwin is an agentic LLM framework integrated with a digital twin (DT) for... Policy-as-logic for robust reasoning over rules. Policy-as-Logic for Robust Reasoning over Rules proposes a hybrid symbolic... APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference. APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference... TradingMoE: Routing the Right Experts in Evolving Markets. TradingMoE: Routing the Right Experts in Evolving Markets proposes a... Regime-Gated Residual Mixture-of-Experts for Cross-Sectional Volatility Forecasting. This paper studies how regime information should be incorporated into a neural... Chain-of-Thought Shows the Path to a Tree: Realizing Branching Complexity. The paper introduces explicit, depth-bounded Chain-of-Thought (CoT)... Robust Ambiguity Detection (RAD) From Model- and Feature-Space Consistency. Robust Ambiguity Detection (RAD): From Model- and Feature-Space Consistency... Towards Truly Unsupervised Evaluation of Feature Selection. The paper "Towards Truly Unsupervised Evaluation of Feature Selection" by Hafiz... TESLA: Taylor Expansion of Sinusoidal Learnable Activations. TESLA: Taylor Expansion of Sinusoidal Learnable Activations Authors: Daehwa Ko,... Reducing Symmetry Increase in Equivariant Neural Networks. Equivariant Neural Networks (ENNs) have empowered numerous applications in... Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization. Epiplexity, a recently proposed measure of the structural information a... Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning. The paper "Drift and Dependence: Layer-wise Information-Theoretic Bounds for... Robustness of AI-Art Detectors under Generator Shift. This paper investigates the robustness of AI-art detectors under generator... FLARE++: Low-rank attention with dynamic attention routing. FLARE++: Low-Rank Attention with Dynamic Attention Routing Vedant Puri, Yongjie... Air Quality Station Simulation via LSTM and Attention-Based Modelling. The paper presents SATADL (SpAtial-Temporal Attention Dual LSTM), a... When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits. This paper presents a diagnostic protocol for selecting proxy rewards and... Fine-Tuning Generative Models for Extreme Events via CVaR-Penalized Wasserstein Gradient Flows. The paper proposes CVaR-GPA (CVaR-penalized Generative Particle Algorithm), a... Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling. This paper investigates whether Earth observation (EO) foundation model... RECAST: A Machine-Learning Framework for Correction and Super-Resolution of Coarse-Grid PDE Solvers. RECAST: A Machine-Learning Framework for Correction and Super-Resolution of... LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection. LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning... FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier Coupon Allocation. FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier... A Factor Graph Approach to Scalable Multi-Output Gaussian Process Regression. This paper presents a factor graph approach to scalable multi-output Gaussian... Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed. This paper presents a comprehensive evaluation of the trustworthiness of Small... Easper: An Accessible ASR Pipeline for Language Documentation. Easper: An Accessible ASR Pipeline for Language Documentation Abstract Audio... SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges. the Paper "SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic... Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages. Language-Conditional Dequantization (LCD) is a post-hoc method that attaches... SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents. SynWeaver is a website-prior task and trajectory co-synthesis framework... NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation. NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and... SoK: From Generation to Consumption of Privacy Documents in Software Systems. This SoK paper provides a unified, lifecycle-oriented view of privacy documents... Adaptive Bregman Proximal Stochastic Gradient with a Stabilized Barzilai--Borwein Step Size. Adaptive Bregman Proximal Stochastic Gradient with a Stabilized... Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach. This paper investigates cooperative target tracking in multi-AUV ad-hoc... HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting. This paper introduces Hugin, a training framework designed to enhance... Prof-K: Probabilistic One-Pass Filtering for Efficient Top-k Selection. Prof-K is a probabilistic one-pass filtering algorithm for efficient top-k... Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control. Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in... Is this Citation on Point?. This paper studies proposition-level citation support verification in legal... A comparison of CNN architectures for Alzheimer's disease detection in single-view MRI scans. This paper proposes a benchmark that evaluates ten different convolutional... CAM-Guided Saliency Cutout and Image-Based Malware Classification. This study evaluated saliency-guided cutout regularization for image-based... Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones. This paper investigates deep learning-based estimation of the Relative Transfer... Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion. The paper "Toward Meaningful Transparency for AI Chatbots: Disclosing... Do Judges Behave Like Algorithms?. This paper investigates whether judges already behave like algorithms in their... PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping. PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised... Which Site, and When: A Free-Satellite-Data Test of Himalayan Glacial Lake Bursts, Landslides, and Ice Floods. This paper tests whether free satellite data can predict glacial lake outburst... ComBodied Agents: a New Paradigm of Human-Centric Agentic AI. Combodied Agents are introduced as a human-centered paradigm of Agentic AI that... LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training. LazyTrain is an optimization-guided scheduler for limited-resource large... On Weak Bisimilarities in CCSK. In the context of CCSK, a reversible extension of CCS, we study different... QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving. QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG... Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs. The deployment of large language models (LLMs) in mental health contexts raises... A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery. This paper introduces NA-UNETR, a 3D transformer-based segmentation framework... ADEPT: A Unified Framework for Deep Learning Test Adequacy. ADEPT: A Unified Framework for Deep Learning Test Adequacy This paper presents... Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment. Generative Semantic Segmentation via an Observable Semantic-Image Interface and... Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers. This paper investigates the role of municipal algorithm registers in providing... The Boolean Power of ReLU. We prove that, on finite simple undirected graphs equipped with a single... Beyond Local Power: Functional Connectivity Analysis for Subject-Independent Learning Style Recognition. This paper proposes an objective Electroencephalography (EEG) approach... Dual-Model Sentiment Analysis of Consumer Reviews in the Retail Coffee Sector Using Machine Learning and Deep Learning Approaches. This study presents a comprehensive sentiment analysis framework applied to... From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection. This paper presents a closed-loop framework for video reflection removal,... Remote Sensing and Machine Learning-Based Analysis of Land Use and Vegetation Change in Dhaka District, Bangladesh. This study employs remote sensing data and machine learning techniques to... EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval. EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under... Attractor Image-Based Deep Learning of Arterial Pulse Waves for Age Classification. This study investigates whether arterial pulse waveform morphology, which... Structuring the Space of Perspectives. The same event can be reported from different perspectives depending on the... DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation. DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial... CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence. CAS: A Causal Attribution Score for Local and Global Explainable Artificial... A Quantum/Classical Example Oracle Separation for Making Things Up. The paper studies the power of quantum examples compared to classical examples... CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications. CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI... Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection. Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark... Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences. Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences... Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study. Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case... A Local-Linearly Convergent Algorithm for Nonconvex Equality-Constrained Optimization. The paper extends the analysis of the Gradient-Eigenstep Algorithm by Goyens et... HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment. HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar... Policy Convergence and Divergence Across National and Within Regional AI Strategies: A Policy Design Element Analysis. Governments worldwide have responded to the rapid expansion of AI by publishing... Not All Nudges Land: Behavioral Controllability and Elaboration Quality in AI-Supported Journaling. This paper analyzes 369 journal entries from an eight-week passive sensing... Geometric and Behavioral Stratification in Transformer Residual Streams. The paper investigates the geometric and behavioral organization of transformer... GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation. Generating actionable financial advice from business records demands that... ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering. ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in...
The papers
- Interpretable Causal Discovery via Causal-Effect Constraints — Authors: Cixuan Zhang, Guy Van den Broeck, Benjie Wang Affiliation: Computer Science Dept., Yale University; Computer Science Dept., University of California, Los Angeles --- The paper addresses the challenge of conditional causal discovery — the task of uncovering causal struc
- ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents — ToolHazard is a scalable adversarial environment synthesis framework designed to address the challenge of constructing executable, stateful sandboxes for security evaluation and alignment of LLM-based agents.
- AVA-Encoder: Towards Agent-Native Video Representation Learning — AVA-Encoder: Towards Agent-Native Video Representation Learning Summary This paper introduces the Agentic Video Auto-Encoder (AVA-Encoder), a framework for learning agent-native video representations.
- epsilon-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution — ε-MemEvo is a framework for cross-task knowledge transfer in LLM-based program evolution.
- LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence — LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence Summary This paper introduces LODESTAR (Learned Orientation of Directed Entropy, Steering Trustworthy Answer Retrieval), a m
- Non-Degenerate Risk Certification for Automated Security Decisions: A Decision-Contract Theory with ATT&CK-Aligned Triage as a Worked Instance — Author: Zhenpeng Li (Guangzhou Health Science College) arXiv:2608.12444v1 [cs.CR] 12 Aug 2026 --- The paper addresses a fundamental flaw in unconditional risk bounds for automated security decisions.
- HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference — HetRoute: Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference Summary This paper proposes HetRoute, a heterogeneous and cost-aware collaborative routing framework for distributed edge Mixture-of-Experts (MoE) inference.
- Automated binary classification of hazelnut X-ray images: A deep-learning benchmark for quality assessment — This study presents a benchmark for binary hazelnut quality classification (healthy versus defective) using deep learning on X-ray images. The dataset consisted of 799 segmented single-kernel X-ray images (224 × 224 pixels, grayscale) of *Corylus avellana* var. *pontica*, cv.
- HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation — HSTGFormer is a graph-enhanced Transformer framework proposed for efficient monocular 3D human pose estimation.
- Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning — This paper studies transfer learning in cooperative multi-agent reinforcement learning (MARL), addressing the problem of serving objectives that change dynamically after deployment without retraining policies for each new objective.
- Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models — The paper introduces Ripple-Pivot Search (RPS), a novel training-free parallel decoding method for Diffusion Large Language Models (dLLMs).
- Training Under Challenge: Executable Certificates and Challenge-Closed Optimality for Neural Networks — The paper addresses a fundamental limitation in neural network training: "A flat training curve does not reveal whether a neural network has reached a global optimum, is locally trapped, is representation-limited, or is mismatched to its trainer." The authors introduce Training U
- Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges — Graph-Structured Rubrics (GSR) compiles a rubric into a response-independent typed evaluation graph before observing candidate responses.
- Causal Structure is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library — Causal Structure Is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library Authors: Xining Xun (Tsingjiao Information Science (Beijing) Co., Ltd.) arXiv: 2608.11767v1 [cs.CL], 12 Aug 2026 --- The paper investigates whether exposing discret
- Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces — Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Authors: Congchao Wang, Diwakar Singh, Qiaozi Gao, Spyros Matsoukas, Yang Liu, Mahdi Namazifar (Amazon AGI) Summary This paper introduces Reasoning Jury, a system designed to improve the fidelity of judgments f
- Exemplar-based objective classification of gust-induced loads across multiple flight conditions — This paper investigates whether an objective classification criterion can be found to organize the complexity of gust-induced loads across multiple flight conditions, while remaining as interpretable as labeling based on coarse parameters such as flight attitude.
- Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment — This paper proposes a hybrid planning architecture for automated driving that combines learning-based behavior planning with optimization-based trajectory supervision, aiming to exploit the generalization capabilities of learned behavior generation while retaining deterministic,
- Distillation of Foundation Models for Time-dependent PDEs — Authors: Daniel Musekamp, Boshra Ariguib, Andrei Manolache, Mathias Niepert Affiliation: University of Stuttgart, University of Cologne, Bitdefender arXiv: 2608.11937v1 [cs.LG] 12 Aug 2026 --- Foundation models for time-dependent partial differential equations (PDEs) are trained
- Fingerprinting Text-to-Image Diffusion Models via Collapsed Generation — This paper presents a non-invasive model fingerprinting framework for text-to-image (T2I) diffusion models based on "collapsed generation"—a phenomenon where certain input conditions produce highly consistent images across multiple stochastic seeds.
- Disentangling the Expressivity of RoPE — This paper, "Disentangling the Expressivity of RoPE" by Selim Jerad, Anej Svete, Jiaoda Li, and Ryan Cotterell, published at COLM 2026, reconciles two seemingly incompatible accounts of rotary position embeddings (RoPE) in transformers: the theoretical expressivity view linking p
- Persistent Recursive Worlds Enable Autonomous Software Evolution — EvoX Genesis (hereafter, Genesis) is a system that reorganizes long-horizon software development around a persistent recursive project rather than a persistent agent.
- CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network for End-to-End Image Generation — CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network for End-to-End Image Generation Summary This paper proposes CoQui, a coordinate-conditioned quantum implicit generative adversarial network (GAN) for end-to-end image generation.
- Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs — The paper introduces an adversarial reinforcement learning framework to expose how easily large language models (LLMs) abandon correct beliefs under optimized persuasive pressure.
- Direct Acceleration of Stochastic Root-Finding Without Variance Reduction and Regularization — The paper addresses stochastic root-finding problems, stated as finding x ∈ R d such that F(x) = 0, where F is an operator.
- Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning — The paper addresses a central challenge in LLM personalization: task-specific preference adaptation. The authors note that "universal preference summaries often contain information irrelevant to a particular downstream task.
- AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention — AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention introduces a fine-tuning method for Multimodal Large Language Models (MLLMs) that mitigates catastrophic forgetting by dynamically controlling parameter updates based on activation patterns.
- Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL — This paper presents a framework for learning loco-manipulation policies using Sample-based Model Predictive Control (SMPC) demonstrations with sparse offline-to-online Reinforcement Learning.
- Draw This First — The paper "Draw This First" presents a method for generating ordered vector sketches from text descriptions or images, where the drawing order can be specified via text instructions.
- Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion — The paper introduces GDCE-I (Graph Diffusion Counterfactual Explanation via Inversion), a novel framework for generating counterfactual explanations for Graph Neural Networks (GNNs).
- NAE: Normalizing AutoEncoder — The paper introduces the Normalizing Autoencoder (NAE), a generative framework within the "flow autoencoder" paradigm, which combines autoencoder structure with likelihood-based training of normalizing flows by using separately parameterized encoder and decoder networks as approx
- Structure-preserving uncertainty quantification for GENERIC dynamics — Authors: Zequn He and Celia Reina (University of Pennsylvania) arXiv:2608.12624v1 [cs.LG] --- The paper addresses a critical gap in uncertainty quantification (UQ) for structure-preserving machine learning models.
- Confidence Calibration of Deep Learning Systems — This thesis explores novel methods for improving confidence calibration under challenging conditions such as label noise, domain shifts, and privacy constraints.
- The Advective Fisher-Rao Geometry of Deterministic Measure Transport — The paper introduces a novel Riemannian metric, the advective Fisher–Rao metric, for optimization tasks on paths of probability measures governed by the continuity equation.
- Fast Length-Squared Sampling for Positive-Semidefinite Matrices — Summary This paper presents a simple rejection-sampling-based algorithm for performing length-squared sampling on an n × n positive-semidefinite (psd) matrix, i.e., sampling a column with probability proportional to its squared l2-norm.
- When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide — This paper investigates when offline evaluation of equal-cost top-k allocation policies can be trusted, benchmarking six estimators across five datasets and two known-effect sweeps, with a non-simulated paired reference validation.
- Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines — This paper presents a systematic five-step methodology for deploying custom classification models on NVIDIA Jetson Deep Learning Accelerator (DLA) cores with zero GPU fallback, enabling parallel multi-model inference on a single edge SoC.
- Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization — Most image colorization systems operate in Lab space by predicting chroma (ab) while preserving an input-derived luminance channel (L).
- SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization — SSPO (Structure-Aware Similarity-Weighted Preference Optimization) is a method for neural combinatorial optimization (NCO) training that addresses two failure modes in existing approaches: "gradient signal polarization" and "baseline redundancy." Gradient signal polarization occu
- Towards Scalable Fuzzy PSI via Efficient Fuzzy Matching — Based on the paper, here is a detailed summary: This paper presents scalable fuzzy private set intersection (fuzzy PSI) protocols for general Lp ∈ [1,∞] distance, supporting both low- and high-dimensional sets.
- DiG-bench: Discovery in Games — DiG-bench (Discovery in Games) is a new benchmark introduced to measure the capacity for active discovery—formulating novel generalizations through experimentation—in AI systems.
- The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance — Authors: Shailja Thakur, Sungeun An, Chad DeLuca, Hima Patel (IBM Research India, IBM Research Almaden) Paper: arXiv:2608.11694v1 [cs.CL], 12 Aug 2026 --- The paper addresses a fundamental flaw in how LLM benchmark scores are reported: "A benchmark score comes from a single phras
- Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill — Spark-to-Paper is an end-to-end research paper generation system implemented as thirteen composable skills inside an existing coding assistant, without requiring a separate agent platform or orchestration service.
- Unifying Physical Backpropagation — This paper develops a unifying theoretical framework for formally exact physical backpropagation methods based on the adjoint method from PDE-constrained optimization. The authors state: "The main contribution of this work is twofold.
- Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular Missingness — Paper: "Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular Missingness" by Haochen Zhang, Jiaheng Guo, Yu-Chao Huang, Nicholas Konz, and Tianlong Chen (UNITES Lab, University of North Carolina at Chapel Hill).
- Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment — Summary This paper introduces Group Alignment-induced Sycophancy (GAS), a two-sided evaluation framework for steerable pluralistic alignment.
- Foresight Without Seeing: Latent Futures for World Action Models — ForeWAM is a dynamics-conditioned direct-policy World Action Model (WAM) that provides predictive context for action generation without explicitly generating future videos.
- Claim-Level Reliability Assessment for Efficient Test-Time Reasoning — The paper introduces Claim-Level Reliability Assessment (CLR), a training-free framework for test-time scaling that reallocates compute from additional solution sampling to targeted verification.
- Dion3: Full-Stack Orthogonal Updates — Dion3 is a revision of the Muon optimizer that targets overhead at every level of the stack. The Muon optimizer incurs significant overhead due to its cubic-time Newton-Schulz orthogonalization step, and communication overhead compounds this when weights are sharded.
- Locating and Controlling Implicit Personalization in Large Language Models — Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity.
- Scaling Automatic Research Agents via World Models — The paper "Scaling Automatic Research Agents via World Models" introduces World Model RL (WMRL), a method to scale reinforcement learning (RL) for automatic research (AutoResearch) agents by replacing expensive environment execution with a learned world model.
- Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System — This paper introduces UAVQA-Bench, a comprehensive, fully human-annotated benchmark for UAV aerial image understanding and reasoning, and proposes UAV-MAS, a training-free multi-agent system designed to address the unique challenges of visual perception and reasoning in UAV scena
- G0.5: One Autoregressive Stream for Robot Reasoning and Action — The paper introduces G0.5, a pretrained autoregressive Vision-Language-Action (VLA) model in which "a single transformer decoder emits reasoning and action tokens under a single objective." The authors argue against the prevailing "VLM-as-encoder" recipe that "couples a pretraine
- JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis — The paper introduces Ancient Chinese Character Exegesis (ACCE), a vision-language question answering (VQA) task that models the scholarly exegesis process of ancient Chinese characters.
- A Local Sinkhorn Framework for Conditional Distribution Reconstruction of Multidimensional Random Fields — This paper proposes a local Sinkhorn divergence framework for conditional distribution reconstruction of multidimensional random fields.
- From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices — The paper argues that the central research problem for using large language models (LLMs) in medical-device safety analysis is not safety-text generation, but source-linked safety-knowledge support.
- SoftWater: Class-Aware Rate Allocation for Softmax Quantization — Based on the paper, here is the summary: The paper introduces SoftWater, a method for quantizing the softmax output layer (head) of large language models (LLMs) under a rate-distortion framework using KL divergence, rather than the standard weighted mean-squared error (WMSE) used
- Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision — Authors: Shaojie Zhang and Ke Chen, Department of Computer Science, The University of Manchester The paper addresses a realistic setting in pairwise constrained clustering where supervision is not idealized hard binary labels but rather real-valued probabilistic relations that en
- LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration — LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration Abstract Summary The paper introduces LoSA, a training-free sparse-attention method for accelerating video diffusion transformers.
- How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging — This paper evaluates the complete unsupervised domain adaptation (UDA) pipeline in medical imaging, considering both the adaptation stage and the label-free model selection stage together, under clinical deployment conditions.
- VICBench: A Multi-Language Benchmark for Code Vulnerability Detection — VICBench: A Multi-Language Benchmark for Code Vulnerability Detection This paper presents VICBench, a benchmark of 100 verified vulnerability-inducing commits (VICs) for 100 CVEs across 88 projects in Python, Java, and C++, covering 48 CWE types.
- Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling — Summary This paper investigates whether on-policy distillation (OPD) truly expands the reasoning capability boundary of student language models or primarily improves sampling efficiency within existing capabilities.
- One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL — The paper addresses a critical failure mode in multi-agent reinforcement learning for human-AI interaction. Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. [episode]
- Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams — Diagram-MMU is a multi-modal benchmark designed to assess Multimodal Large Language Models' (MLLMs) ability for scientific diagram parsing and understanding. It features 3.7k curated diagrams and 18.3k human-validated questions across six domains.
- Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads — This paper introduces a unified theoretical framework for understanding multiplicative dual-encoder networks, which compute a real-valued output for a pair of inputs as the inner product of their separately learned encodings.
- Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL — Authors: Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu (Scale AI, University of Arizona, University of Texas at Dallas) arXiv:2608.11669v1 [cs.LG] 12 Aug 2026 --- Reinforcement learning against rubrics—li
- From Visual Widgets to UI Code: Efficient Tool-Grounded Generation — Existing screenshot-to-code systems face a trade-off between flexibility and controllability.
- When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation — The paper investigates a resource-shared evaluation loop in entity-level machine translation of the Seungjeongwon Ilgi, a UNESCO Memory of the World record that is only 37.4% translated.
- Small-Scale Experiments: Are We There Yet? — Small-Scale Experiments: Are We There Yet? Nicholas Lourie, Kyunghyun Cho, Karen Ullrich, Sanae Lotfi Summary This paper addresses the long-standing promise of scaling laws to enable cost-effective machine learning experiments by conducting research at small scales and transferri
- XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication — XB RIDGE: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication Summary This paper introduces XB RIDGE, a decode-free communication protocol designed to enable effective communication between heterogeneous large language models (LLMs) from different model families.
- Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence — The paper introduces the “Agentic Self-Improvement” framework, a closed-loop, goal-directed optimization system designed to address the lack of fine-grained control and reliability in modern black-box Image-to-Video (I2V) models.
- A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions — The paper presents a formal, three-step framework for designing human-aligned reward functions in reinforcement learning, aimed at enabling non-experts to instantiate and iterate on reward functions that adhere to a given preference ordering over trajectories.
- Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models — Summary This paper introduces a framework for the automated construction of Dynamic Master Logic (DML) models from system documentation, representing them as Knowledge Graphs (KG-DML) using Retrieval-Augmented Generation (RAG) and Large Language Models (LLMs).
- Redistribution-based Cost Inference Improves Sparse Safe Offline RL — The paper introduces the Redistribution-based Cost Inference (RCI) framework, which addresses the problem of safe offline reinforcement learning when only sparse trajectory-level stop-feedback is available, rather than dense per-step cost annotations.
- AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses — This paper, from Salesforce AI Research and the University of Illinois Urbana-Champaign, investigates whether capability transfer from strong to weak language models can occur at test time rather than through conventional training-time distillation.
- CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution — CAKE: Compiler–Agent Co-Design for Frontier Kernel Evolution presents a system that co-designs a GPU kernel programming language and an agent-driven compiler harness to enable efficient kernel evolution.
- MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents — MindMemOS is a portable and self-evolving memory operating layer for AI agents, designed to address the rigidity of existing memory systems that remain fixed after development.
- FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents — FrontierFinance is a fully open benchmark introduced by Samaya AI for evaluating AI agents on professional investment research. It consists of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six use cases across the full investor workflow.
- Unifying Generative Models with Path Integrals — This paper formulates generative modeling as a path integral, unifying flow-based, diffusion-based, variational, and adversarial models as different evaluation principles for a single master action.
- Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review — This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no preexisting oracle to validate the target behaviour.
- Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework — This paper presents a structured, benchmark-based comparative assessment of publicly benchmarked Indian foundation models against global frontier and comparable-scale models across eight capability domains: general-purpose reasoning, coding and software engineering, agentic AI an
- Trie Automata for Constrained Decoding over Large Finite Sets — Authors: Xingzi Xu and Karim Bouyarmane (Amazon) Published: Conference paper at COLM 2026 --- Large language models increasingly need to generate structured outputs conforming to predefined schemas, with one common constraint being selection from a finite set of valid strings.
- TELLME: Test-Enhanced Learning for Language Model Enrichment — Summary The paper introduces TELLME (Test-Enhanced Learning for Language Model Enrichment), a novel continual pre-training (CPT) method for large language models (LLMs) that applies the Test-Enhanced Learning (TEL) principle from educational psychology to improve domain-specific
- Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents — Governed Persistent Memory (GPM) is introduced as an auditable bitemporal state-transition model for long-horizon agents, addressing the problem that retrieval alone does not determine whether contradictory, superseded, retracted, deleted, or stale records may support an outgoing
- High-Order Liquid Evidence Encoding for Gradual GNSS Spoofing Detection in Autonomous Driving — The paper proposes a causal high-order liquid evidence framework for detecting gradual and subtle GNSS spoofing attacks in autonomous driving.
- Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting — Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review.
- Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus — This paper presents the first systematic study of massive activations (MAs) in layer-interleaved hybrid linear attention large language models (HLA LLMs). The authors uncover two architecture-aligned morphologies of MAs: pre-attention spikes (PAS) and inter-spike plateaus (ISP).
- Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation — This paper challenges the standard assumption that LLM model rankings are stable across inference conditions by systematically varying the token generation budget (maximum tokens a model may produce) across seven levels (64–4,096 tokens).
- AmbSentry: Mitigating Sensing Eavesdropping in ISAC Systems by Harnessing Ambient IoT Devices — AmbSentry: Mitigating Sensing Eavesdropping in ISAC Systems by Harnessing Ambient IoT Devices Summary This paper proposes AmbSentry, a novel integrated sensing and communication (ISAC) system designed to prevent sensing eavesdropping by leveraging naturally distributed passive am
- Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder — Confucius4-TTS is a multilingual zero-shot text-to-speech (TTS) system that supports 14 languages and performs both intra-lingual and cross-lingual reference cloning without requiring transcripts of audio prompts.
- ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening — ScreenShot is a hierarchical transformer-based foundation model for few-shot prediction in combination drug screening.
- Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem — This paper takes a first step toward a unified formal account of agent memory, addressing the lack of a formal definition, a measure of optimality, and a formulation of how memory is written.
- Continuous-Latent Predictive Modeling with Semantic Alignment for EEG-Language Foundation Models — The paper proposes the Brain Latent Predictive Model (BLPM), an EEG-language foundation model that reformulates heterogeneous EEG decoding tasks as a continuous semantic embedding prediction problem.
- An Efficient Minimax-Optimal Algorithm for Adversarial m-Set Bandits — The paper studies adversarial combinatorial bandits with m-set actions. At each round, the learner selects m out of d items and observes only the aggregate loss of the selected items, not the individual item losses.
- How Organizations Use AI: Evidence from ChatGPT — This paper studies how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifications, and public-company financial data through March 2026.
- Certifying What Helps Customer-Return Timing: A Screen-and-Confirm Test for Conditioning Signals, and Why Decay Is Nearly Enough — This paper introduces a screen-and-confirm protocol to certify whether a candidate signal improves a temporal point process (TPP) model's event-timing likelihood, and a model-free ceiling quantifying how little of customer-return timing is point-predictable.
- TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement — TailBooster is a dual-layer generative framework designed to address two complementary failure modes of conventional deep generative models applied to mixed-type tabular aviation records: the systematic under-representation of distributional tails and the production of operationa
- Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs — The paper introduces Self-Fix Step-DPO (SFS-DPO), a two-stage reinforcement learning framework designed to improve self-correction capabilities in large language models (LLMs) for mathematical reasoning.
- DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation — DexterSQL is a prompting-based (non-fine-tuning) Text-to-SQL system that improves SQL generation accuracy by addressing three key challenges faced by existing LLM-based methods: (i) relying on coarse-grained schema information that fails to reveal fine-grained relationships neede [episode]
- Robust and Efficient Noisy-Label Time-Series Classification via Dynamic Time Warping Based Granular Ball Computing — Dynamic Time Warping (DTW)-based Nearest-Neighbor (NN) classifiers are effective for time-series classification but are vulnerable to mislabeled training samples and require numerous DTW computations during inference.
- When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use — This paper investigates a failure mode in multilingual API calling called Argument Language Mismatch (ALM), where a model selects the correct tool but generates argument values in a language inconsistent with the user input.
- Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction — Based on the paper, here is the summary: This paper introduces DARC (Diagnosis-guided Agent Recovery and Correction), a framework that recasts agent self-correction from indiscriminate context expansion to diagnosis-guided recovery.
- Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation — The paper studies reference-free post-training for multilingual machine translation with open large language models.
- IoT-Enabled Autonomous Maritime Navigation in Smart Ports: A Curriculum-Guided Shared Policy Learning Framework — This paper investigates onboard autonomous navigation for IoT-enabled autonomous surface vehicles operating in congested smart port environments under partial observability and dense traffic conditions.
- CT- Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models — CT-∆Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models This paper introduces CT-∆Bench, a benchmark for longitudinal CT difference reporting, a task where a model takes two temporally separated CT scans from the same patien [episode]
- M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation — M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation Purpose: The paper investigates whether explicit mathematical inductive biases—specifically matrix spectral analysis and vector calculus operators—can enhance m
- Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents — Harness-IF is a benchmark that turns operational instruction following into a rule-level measurement problem.
- Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning — The paper presents a multi-tier, frame-level audio tagging framework for infant-centered audio understanding in naturalistic home recordings.
- Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting — Motion-centric video reasoning is fundamental to interactive applications such as robotic manipulation and autonomous navigation.
- Sparse and robust geometric twin support vector machine via asymmetric RoBoSS loss function — This paper proposes a novel asymmetric, robust, bounded, sparse and smooth (aR) loss function for l1-norm penalized geometric twin support vector machine (aRSGTSVM) to handle classification and regression tasks. The l1-norm penalty can achieve the feature selection.
- MARCH: Scaling Recurrent Memory with Content-Routed State Anchors — MARCH: Scaling Recurrent Memory with Content-Routed State Anchors Abstract Transformers owe much of their strong long-context retrieval capability to a token-level memory that grows with context length.
- Unified Multi-Dialectal Neural Machine Translation for Bangla Using the Dwadash Benchmark Corpus — This paper presents a unified Poly-Dialectal Neural Machine Translation System for Bangla regional dialects, addressing the challenge that contemporary NMT architectures and LLMs assume a homogeneous language distribution, leading to severe performance degradation when translatin
- Continual Learning in Transition — This survey paper, authored by Zhiyan Hou, Dan Zhang, Tao Feng, and colleagues, addresses a fundamental shift underway in the field of continual learning (CL).
- When Do Anchor-Based Pointwise LLM Rerankers Help? Retriever Quality, Statistical Scope, and Anchor Design — This paper investigates when anchor-based pointwise LLM reranking helps, using GCCP/PAGC as a representative method. The study is "reproduction-first," using reproduction as a starting point for controlled component-level stress testing.
- Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents — The paper "Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents" investigates whether uncertainty quantification (UQ) methods developed for single-turn language model outputs transfer to the interactive, multi-turn trajectory setting of LLM
- JAPE: Joint Anomaly Prediction and Intrinsic Explanation in Multivariate Time Series — JAPE: Joint Anomaly Prediction and Intrinsic Explanation in Multivariate Time Series Abstract Multivariate time-series anomaly prediction aims to identify whether and when anomalies will occur over a future horizon from historical observations.
- HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry — HyperANFIS is a hyperbolic extension of the adaptive neuro-fuzzy inference system (ANFIS) proposed to address the limitations of conventional ANFIS models, which "generally construct rule antecedents and perform inference in Euclidean space, limiting their representational capaci
- Machine Learning-Based Cyber Defense for Cloud Infrastructure: An Adaptive Deep Q-Network Architecture for Intelligent Intrusion Detection and Automated Threat Mitigation — Summary This paper proposes a reinforcement learning-based dynamic cyber defense framework for cloud infrastructures, deploying a Deep Q-Network (DQN) to train effective defensive strategies against evolving cyberattacks.
- How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models — This paper addresses the challenge of efficiently using limited oracle budgets when guiding protein structure prediction models.
- HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks — HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks Abstract Kolmogorov-Arnold Networks (KANs) enhance nonlinear function approximation by replacing scalar weights with learnable univariate functions.
- NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation — NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation Abstract Summary: Large Language Models (LLMs) are increasingly used in circuit design workflows, yet their reliability on simulator-facing SPICE netlist recognition and manipulation remains po
- SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward — SCOUT (Structured Chain-Of-Thought Utilizing Process-Supervised RL Training) is a framework proposed to enhance the spatial reasoning capabilities of Vision-Language Models (VLMs).
- Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness Classification — SGNet (Spectral-Grouped Network) is a lightweight architecture for hyperspectral fish freshness classification that separates spectral and spatial feature extraction using grouped convolutions and a depthwise spatial pathway.
- Few-Shot Ordinal Learning for Day-Wise Freshness Estimation with Hyperspectral Fish Images — This paper introduces the first few-shot learning framework for hyperspectral imaging (HSI)-based food quality estimation, specifically for day-wise freshness prediction of fish fillets.
- HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression — HAMP-LIC is a Hessian-aware mixed-precision post-training quantization (PTQ) framework proposed for learned image compression (LIC) models.
- An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS — The paper proposes and evaluates a supervised agentic workflow for modernizing legacy scientific code at production scale, using the conversion of the two-electron-integral core of GAMESS from fixed-form Fortran 77 to free-form Fortran 2008 as a case study.
- Regime-Gated Residual Mixture-of-Experts for Cross-Sectional Volatility Forecasting — This paper studies how regime information should be incorporated into a neural network for cross-sectional volatility forecasting. The authors propose RG-ResMoE, a regime-gated residual mixture-of-experts architecture, and evaluate it against several baselines on a U.S.
- Calibration Bets on the Past: Post-Training Quantization for Financial Time-Series Forecasting — This paper presents a systematic study of activation calibration for post-training quantization (PTQ) in financial time-series forecasting, specifically for cross-sectional volatility forecasting on the S&P 500.
- A Cascaded Unsupervised-Supervised NLP Pipeline for Detecting Accusatory Language in Public Procurement — This paper proposes a cascaded hybrid learning strategy that integrates unsupervised clustering and supervised classification to detect accusatory or whistleblowing-style comments in Ecuador's public procurement system.
- Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling — This paper investigates whether Earth observation (EO) foundation model embeddings can serve as effective sub-grid surface descriptors for probabilistic weather downscaling.
- Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents — The paper introduces Convergent Detour Hijacking (CDH), a text-only, runtime-independent attack against skill-based LLM agents that use progressive disclosure.
- A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery — Summary This paper introduces NA-UNETR, a 3D transformer-based segmentation framework designed to improve the segmentation of the Left Anterior Descending (LAD) artery in free-breathing, non-contrast CT images for cardiac dose sparing in thoracic radiotherapy.
- Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages — This paper examines the structural barriers that disadvantage speakers of underrepresented languages in AI infrastructure, using Bengali as a case study.
- VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies — VAKRA (eValuating API and Knowledge Retrieval Agents) is a benchmark introduced in this paper for evaluating multi-hop reasoning across APIs and retrieval under tool-use policies in executable environments. It is built on the structured API generation pipeline of Elder et al.
- Large Language Model-Driven Small-Capitalization Trading: Integrating Financial News Sentiment, Macroeconomic Indicators, and Technical Signals — This paper studies small-capitalization trading with LLM-derived news sentiment, macroeconomic indicators, technical signals, and uncertainty-aware portfolio construction.
- LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining — LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Large language models (LLMs) have achieved remarkable breakthroughs across various applications.
- Which Site, and When: A Free-Satellite-Data Test of Himalayan Glacial Lake Bursts, Landslides, and Ice Floods — This paper tests whether free satellite data can predict glacial lake outburst floods (GLOFs), rainfall-triggered landslides, and smaller glacial floods in High Mountain Asia, with a focus on Nepal.
- Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction — The paper introduces Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that systematically varies the number of simultaneous constraints (k) from 1 to 12, with every constraint scored by a deterministic, rule-based verifier and zero LLM-judge involvement.
- SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents — SynWeaver is a website-prior task and trajectory co-synthesis framework designed to address two key limitations in existing exploration-based data synthesis methods for web agents: (1) they often fail to cover the full functionality of a website, and (2) without sufficient websit
- The energetic cost of mitigating AI attacks in cellular networks — The integration of Artificial Intelligence (AI), generally as Machine Learning (ML) algorithms, in all levels and aspects of cellular networks demonstrates the success of data-driven algorithms; for example, the Radio Intelligence Controller (RIC) of the O-RAN paradigm bestows th
- Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach — This paper investigates cooperative target tracking in multi-AUV ad-hoc networks and proposes the MDCA hierarchical control architecture together with the VGG-MADiffRL algorithm.
- Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection — Deep learning detectors for anomalies in dynamic graphs have achieved strong accuracy, yet they remain opaque: "when an edge is flagged, the analyst receives a score but no reason." This opacity is "untenable in the cooperative, regulated information systems where such detectors
- Exploring Oversmoothing with Householder Matrices — This paper introduces HouseGNN (Householder Graph Neural Network), a novel architecture designed to mitigate oversmoothing in deep graph neural networks.
- Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction — Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. Existing localization methods have advanced, among which multistage refinement is a superior solution.
- Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability — This paper develops a mathematical and engineering architecture for secure network Electronic Health Record (EHR) interoperability, organized around the concept of a "logit boundary" as the interface between untrusted discovery models and a deterministic judgment substrate.
- Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated Research — Summary The paper presents Nutrition Data Service (NDS), source-preserving infrastructure that operationalizes FAIR (findable, accessible, interoperable, reusable) for automated, agent-mediated nutrition research.
- Do Judges Behave Like Algorithms? — This paper investigates whether judges already behave like algorithms in their decision-making, using data from misdemeanor bail hearings in Harris County, Texas.
- From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models — The paper introduces MPAR-Bench, a bilingual (English–Chinese) benchmark designed to evaluate "reasoning breadth" in large language models (LLMs), as opposed to the extensively studied "reasoning depth." The authors argue that while modern LLMs have achieved remarkable success
- Inferential Capability Does Not Determine Legal Scope — The paper argues that the term "inference" performs two distinct legal functions in EU digital law, and that these functions are not aligned.
- Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization — The paper introduces a cost-aware method for evolutionary optimization of LLM prompts and agentic programs, called "Cost-Aware Cross-Tier Transfer." The core idea is to decouple the three roles an LLM plays in the evolutionary loop—fitness evaluation, variation (mutation), and
- Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse — CUE-Bench is a Chinese Unsaid Emotion benchmark that centers on Affective Stance and covers diverse communicative scenarios.
- ComBodied Agents: a New Paradigm of Human-Centric Agentic AI — Combodied Agents are introduced as a human-centered paradigm of Agentic AI that perceives, models, predicts, and supports individual human-state trajectories over time.
- Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance — This paper investigates how prompt construction methodology affects the measurement of political stance in Large Language Models (LLMs), specifically comparing templated prompts against fully synthetic (LLM-generated) prompts.
- FLARE++: Low-rank attention with dynamic attention routing — FLARE++: Low-Rank Attention with Dynamic Attention Routing Vedant Puri, Yongjie Jessica Zhang & Levent Burak Kara Department of Mechanical Engineering, Carnegie Mellon University arXiv:2608.11519v1 [cs.LG] 12 Aug 2026 Abstract Full self-attention (Vaswani et al., 2017) is a stron
- On Weak Bisimilarities in CCSK — In the context of CCSK, a reversible extension of CCS, we study different notions of bisimilarity (strong/weak, forward-only/reversible) and highlight their differences and commonalities.
- Hierarchical Federated Transfer Learning in Digital Twin-Based Vehicular Networks — This paper proposes a novel Hierarchical Federated Transfer Learning (HFTL) architecture for Digital Twin-based Vehicular Ad hoc Networks (DT-VANET) to address challenges of data heterogeneity and data sparsity among vehicles, which cause suboptimal accuracy in federated learning
- A Conceptual Framework for Enhancing Workforce Readiness for Smart Manufacturing in the AI Era — Summary This paper introduces the Workforce Readiness Level (WRL) framework, a nine-stage, rubric-anchored, individual-level assessment model designed to operationalize the evaluation of AI-era smart manufacturing competencies.
- Robust Ambiguity Detection (RAD) From Model- and Feature-Space Consistency — Authors: Manya Singh, Mark T. Keane, Arjun Pakrashi (School of Computer Science, University College Dublin) The paper addresses the problem of predictive ambiguity in machine learning models.
- Fine-Tuning Generative Models for Extreme Events via CVaR-Penalized Wasserstein Gradient Flows — The paper proposes CVaR-GPA (CVaR-penalized Generative Particle Algorithm), a robust, tail-agnostic algorithm for fine-tuning pre-trained generative models to learn heavy-tailed distributions and capture extreme events, requiring no prior knowledge or estimation of the target's t
- When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits — This paper presents a diagnostic protocol for selecting proxy rewards and contextual bandit policies in delayed-feedback settings, where the business north-star (e.g., a conversion) is observed only after a long delay and cannot train the bandit online.
- RECAST: A Machine-Learning Framework for Correction and Super-Resolution of Coarse-Grid PDE Solvers — RECAST: A Machine-Learning Framework for Correction and Super-Resolution of Coarse-Grid PDE Solvers Summary This paper introduces RECAST (Recurrent Error Correction And Super-resolution of coarse-grid Trajectories), a machine-learning framework designed to restore accuracy lost w
- RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving Simulation — RoadWeaver is a coarse-to-fine framework for from-scratch generation of diverse, large-scale lane-level HD maps for autonomous driving simulation.
- A Hybrid Framework of Vision Transformer and Gated Recurrent Unit for Detection of Mosquito Diseases — This paper introduces a three-step framework for detecting dengue and Zika virus-infected mosquitoes from control mosquitoes by analyzing their locomotion behavior in video.
- Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models — This paper investigates where safety-aligned refusal behavior is encoded in large language models by transplanting weights from aligned models into matched unaligned base models at multiple levels of granularity.
- Learning from Online User Feedback for Shopping Agents — LOFA, a framework that enables shopping agents to learn directly from real online interaction logs without human annotation.
- FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting — FM-LLM (Frequency-Enhanced Mixture-of-Experts for adapting LLMs to Time Series Forecasting) is a framework proposed in this paper for adapting frozen Large Language Models (LLMs) to time-series forecasting.
- Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones — This paper investigates deep learning-based estimation of the Relative Transfer Matrix (ReTM), which was recently introduced as a generalization of the relative transfer function for multiple receivers and sources.
- Easper: An Accessible ASR Pipeline for Language Documentation — Easper: An Accessible ASR Pipeline for Language Documentation Abstract Audio transcription is a critical bottleneck in language documentation.
- CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement — CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement Summary: This paper proposes CLAIM, an uncertainty-driven framework for active clarification learning in open-domain human–LLM interactions.
- Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents — Authors: Jun He (OpenKedge.io), Deying Yu (OpenKedge.io) arXiv: 2608.11632v1 [cs.MA] 12 Aug 2026 --- The paper addresses a fundamental issue in persistent AI agents: "Persistent AI agents accumulate versioned state across long horizons, but storage retention alone does not identi
- CAM-Guided Saliency Cutout and Image-Based Malware Classification — This study evaluated saliency-guided cutout regularization for image-based malware classification. The controlled experiments compared no cutout, standard random cutout, low-saliency cutout, and high-saliency cutout using ResNet18.
- Transferable Above-Ground Biomass (AGB) Estimation Model from Multi-Sensor Data with Sparse Field Calibration — The paper presents an operational framework for wall-to-wall above-ground biomass (AGB) estimation that combines a single globally trained convolutional neural network (CNN) with a lightweight field-calibration workflow.
- Robustness of AI-Art Detectors under Generator Shift — This paper investigates the robustness of AI-art detectors under generator shift, specifically examining whether detectors trained on earlier diffusion generators remain reliable when applied to images produced by newer, architecturally different models.
- Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models — This paper introduces Semantic Lenia, a framework that "reimagines the Large Language Model inference process as a continuous dynamical system within the macroscopic logit simplex." The authors propose transforming LLM inference from a static optimization paradigm into a continuo
- Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing — The paper introduces HPSE (Hybrid-Policy Self-Editing), a method designed to improve composability in Unstructured Knowledge Editing (UKE) for large language models (LLMs).
- GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs — GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs Abstract On-policy rollout methods such as GRPO are central to posttraining of large language models.
- FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier Coupon Allocation — FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier Coupon Allocation Abstract Coupon campaigns aim to lift both conversion and revenue, but gross merchandise value (GMV) inherits a deterministic funnel structure from conversion and conditional order valu
- AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection — AgenticTwin is an agentic LLM framework integrated with a digital twin (DT) for anomaly detection in cyber-physical systems (CPS).
- Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation — This paper addresses the challenges of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS), which aim to recognize both predefined and unseen object categories without exhaustive human annotations.
- APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference — APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference addresses the memory bottleneck in Mixture-of-Experts (MoE) model inference at the edge.
- Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning — The paper "Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning" develops a layer-wise information-theoretic framework to analyze the generalization behavior of replay-based continual learning (CL).
- LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection — LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection Abstract Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning.
- HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting — Summary This paper introduces Hugin, a training framework designed to enhance vision-language planning for autonomous logistics sorting systems (ALSS).
- Consolidator: Learning Persistent Routed Memory Across Context Boundaries — The paper introduces Consolidator, a shared slot-local operator in a Phasor Memory Network (PMNet) that transforms routed short-term memory (STM) before accumulating it into long-term memory (LTM), without replaying source tokens.
- Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning — This paper introduces the concept of trait-induced safety variation, a failure mode in aligned large language models (LLMs) where the same user request elicits different safety decisions depending on the character trait or persona assigned in the system prompt.
- High-dimensional Multi-objective Bayesian Optimization with Learned Variable Interactions — This paper proposes ViaMOBO, a high-dimensional multi-objective Bayesian optimization (MOBO) framework that alleviates the curse of dimensionality by exploiting decision-variable interactions.
- Chain-of-Thought Shows the Path to a Tree: Realizing Branching Complexity — The paper introduces explicit, depth-bounded Chain-of-Thought (CoT) realizations of graph traversal algorithms and branching complexity measures using Transformer decoders with unique hard attention.
- Proportional Analogies on Probability Distributions via Bayesian Updating — The paper introduces a new notion of proportional analogy for probability distributions based on Bayesian updating.
- Plaintext Recovery Against Post-Filtering Access Control — This paper, "Plaintext Recovery Against Post-Filtering Access Control" by Zachary Espiritu and David Cash, demonstrates that fine-grained access control (FGAC) mechanisms in databases are vulnerable to side-channel attacks that can be amplified into full plaintext recovery.
- A 12-CNOT Double Qubit Excitation Gate — The paper presents the first reported 12-CNOT decomposition of the double qubit excitation operator, improving upon state-of-the-art (SOTA) implementations with 13 CNOTs.
- Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization — Epiplexity, a recently proposed measure of the structural information a compute-bounded learner can extract from data, provides a mechanism to reason about this relationship.
- MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning — MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning proposes a structure-aware multi-objective optimization (MOO) method that performs gradient manipulation under spectral–nuclear norm geometry and uses orthonormalized updates for matrix-valued parameters, addr
- LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification — LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification Abstract Summary: The paper investigates whether financial time series are useful as an additional input for classifying sentences from Federal
- MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques — MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques This paper introduces MuseCritic, a semi-scalar reward model for the aesthetic evaluation of complete songs.
- A comparison of CNN architectures for Alzheimer's disease detection in single-view MRI scans — This paper proposes a benchmark that evaluates ten different convolutional neural network (CNN) architectures (including ResNet, DenseNet, MobileNet, EfficientNet, and VGG family models) under the same held-out test split protocol.
- Instruction Alignment for Binary Code Representation Learning — This paper proposes InsnAlign, a training approach that leverages instruction-level alignment knowledge to improve binary code representation learning.
- The Sleeping Agent: What Gist-Based Context Compression Loses and Why — The paper introduces Salience-Weighted Consolidation (SWC), a biologically-inspired compression framework motivated by sleep-based memory consolidation, used as a diagnostic probe to study when gist-based context compression helps and when it hurts.
- Coarsening Latent-Class Probabilities: Directional Distortion and Coverage Loss — Author: Marcell T.
- TradingMoE: Routing the Right Experts in Evolving Markets — TradingMoE: Routing the Right Experts in Evolving Markets proposes a trading-oriented sparse Mixture-of-Experts (MoE) framework that augments a frozen dense LLM with lightweight residual experts for direct trading-decision generation.
- Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages — Language-Conditional Dequantization (LCD) is a post-hoc method that attaches per-language rank-2 LoRA corrections to the linear layers of an already-quantized model, adding 0.12% parameters per language and training in under 20 minutes on a single GPU.
- Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion — The paper "Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion" by Adrian Rauchfleisch and Andreas Jungherr investigates whether different types of disclosure reduce the persuasive influence of AI chatbots.
- Reading the Gate, Not the Interference: Output-Side Interference Measurement Does Not Track Merge Collapse — This paper investigates the causal structure of task-vector interference in merged language models, challenging the prevailing assumption that interference magnitude is the key diagnostic axis.
- Towards Model-based Run-time Cybersecurity: On Control-Flow Anomaly Detection, Attack Identification, and Hardware Monitoring — The paper "Towards Model-based Run-time Cybersecurity: On Control-Flow Anomaly Detection, Attack Identification, and Hardware Monitoring" addresses the challenge of increasing system resilience to cyber-attacks through run-time monitoring.
- Hybrid Gated Attention — The paper proposes a Hybrid Gated Attention (HyGA) framework to extend the effectiveness-efficiency Pareto frontier of gated attention.
- Can Vision Models Read the Radar Display? On the Feasibility of Radar Imagery for Air Traffic Complexity Estimation — This paper investigates whether radar display imagery, despite its atypical visual characteristics, is a viable input format for deep learning vision models in the context of air traffic complexity estimation.
- Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks — Summary This paper introduces BMAT (Bilevel-Minimax Adversarial Transfer), a unified optimization framework for transfer-based adversarial attacks.
- How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment — Authors: Guang Yang1, Fengchen Liu2, Alex Wang3∗, Homa Hosseinmardi1, Amir Ghasemian1 1University of California, Los Angeles; 2University of California, Berkeley; 3Stanford University arXiv:2608.11816v1 [cs.CR] 12 Aug 2026 --- State-aligned distortion has been documented in Chi
- Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release — This paper presents a fully preregistered, end-to-end stress test of the detect–localize–release pipeline for intervening on language model representations, using a 25.7M-parameter transformer (C1) trained on causal-evidence discrimination where a known suppression phenomenon
- Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs — The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost.
- Kernel Methods for Learning Operators with Multiple Inputs and Outputs — This paper introduces a general kernel-based encoder-decoder framework for learning operators with multiple inputs and outputs.
- Air Quality Station Simulation via LSTM and Attention-Based Modelling — The paper presents SATADL (SpAtial-Temporal Attention Dual LSTM), a deep-learning model designed to simulate the measurements of an unresponsive air quality station until its operation is restored.
- User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling — The paper proposes a collaborative distributed inference system that combines dedicated server infrastructure with resources contributed by service users' devices.
- LookBack: Where and How to Score LVLM Responses via Visual Reference Usage — L OOK BACK is a training-free LVLM response scoring method that augments token likelihood with a visual lookback score, a lightweight measure of how strongly each response token refers to image tokens.
- Two-Stage Deformable-Convolutional Inverse Design of Nanophotonic Absorbers from Optical Spectra — This paper presents a two-stage deformable-convolutional framework for the inverse design of metal–insulator–metal (MIM) nanophotonic resonators, reconstructing 64×64 resonator geometry masks from 80-dimensional absorption spectra.
- Forward and Inverse Virtual Metrology for Phototransistor Gain: A Hierarchical, Uncertainty-Aware Approach for Small Production Datasets — The paper addresses the challenge of predicting and optimizing the current gain of silicon bipolar phototransistors from fabrication process parameters in a small-sample, hierarchically structured setting.
- DCM Bandits: Multiplayer Information Asymmetric Cascading Bandits for Multiple Clicks — This paper extends the Dependent Click Model (DCM) Bandits to a multiplayer information-asymmetric setting.
- Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems — Summary This paper benchmarks the serving cost of agentic memory systems for long-running conversational agents.
- Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents — This paper presents a comprehensive empirical study of skill-induced agent failures in LLM agents, where "agent skills" are instruction packages that extend LLM agents with reusable guidance.
- Policy-as-logic for robust reasoning over rules — Policy-as-Logic for Robust Reasoning over Rules proposes a hybrid symbolic approach called Policy-as-logic (PaL) that expresses policies in formal logic and separates fact extraction using language models from reasoning using answer set solvers.
- Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models — As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge.
- A Factor Graph Approach to Scalable Multi-Output Gaussian Process Regression — This paper presents a factor graph approach to scalable multi-output Gaussian process (MOGP) regression, addressing the computational challenges of traditional kernel-matrix methods.
- LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training — LazyTrain is an optimization-guided scheduler for limited-resource large language model (LLM) training. It is built as an optimization layer over a MegaTrain-style layer-streaming executor.
- CCZ-Equivalence and Enumeration of Triprojective APN Functions — This paper classifies binary-linear two-term Frobenius-linearized operators of the form L(Y) = AY σ + BY on K3, where K is a finite extension of F2 and σ is a fixed nontrivial Frobenius automorphism of K with fixed field F2.
- Accuracy and Order Sensitivity Diverge Under Label-Free Strategies — Summary This paper investigates whether preventing a model from seeing option labels while committing to an answer removes positional influence and improves performance on multiple-choice question (MCQ) benchmarks.
- ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models — ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models Abstract Roles provide an interpretable interface for organizing language-model agents, yet most multi-agent systems treat them as hand-written prompt labels disconnected from learned behavior and p
- Causal inference for group-contaminated structured outcomes: observable quotients, lossless reduction and exact randomization inference — This paper studies causal inference for structured outcomes (such as microscopy images) that are observed after an unknown, unit-specific transformation (group action).
- LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation — LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation Abstract Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory.
- TESLA: Taylor Expansion of Sinusoidal Learnable Activations — TESLA: Taylor Expansion of Sinusoidal Learnable Activations Authors: Daehwa Ko, Jaehyeon Kim, Seunghyun Ham (Korea Aerospace University), Jay Hoon Jung Abstract: The parity problem—deciding whether the number of ones in a binary vector is odd or even—remains challenging for s
- Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection — The paper presents BENCH2ROBUST, a framework that converts failure-free tool-use benchmarks into controlled stochastic environments with scenario-controlled solvability, where episodes explicitly require retrying, switching, or stopping after available paths are exhausted.
- Slips: Behavioral Evidence Aggregation for Network Security — Slips is a network intrusion detection system that builds host-centered behavioral profiles and organizes activity into time windows. It uses a modular architecture in which independent modules report evidence rather than generating final alerts directly.
- HCGRec: Hint-Conditioned Generative Recommendation with Semantic IDs — Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autoregressively generating this token sequence.
- Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed — This paper presents a comprehensive evaluation of the trustworthiness of Small Language Models (SLMs), comparing models obtained through two primary approaches: pre-trained small models and compressed larger models.
- Latent variable models for simultaneous EOV identification and removal in population-based SHM — The paper introduces a Gaussian-process latent EOV (GLEOV) model for simultaneous identification and removal of environmental and operational variability (EOV) in population-based structural health monitoring (PBSHM). The key contributions are: 1.
- A Remote Approach to Cashew Orchard Detection: Leveraging Active Learning with Satellite Imagery in Guinea-Bissau — This study presents the first openly accessible, countrywide cashew orchard map for Guinea-Bissau, West Africa, created using a fully remote and cost-effective approach.
- CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations — CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations Summary This paper introduces CTBench, a public benchmark designed to evaluate the troubleshooting capabilities of AI agents in realistic telecom network operations and maintenan
- RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks — RealisticTritonBench is a benchmark for evaluating LLM-based Triton kernel generation in realistic, production-like settings.
- Asymptotic Risk Calibration for Selective Question Answering — The paper proposes A-CRC-QA, a post-hoc calibration framework for uncertainty-aware selective question answering.
- Adaptive Bregman Proximal Stochastic Gradient with a Stabilized Barzilai--Borwein Step Size — The paper addresses the finite-sum composite optimization problem on a finite-dimensional normed vector space E: min F(x):= f(x) + h(x), where f(x):= (1/n) Σ fi(x), with h being a proper convex regularizer whose proximal subproblem is tractable.
- Reducing Symmetry Increase in Equivariant Neural Networks — Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields.
- Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence — This paper introduces Mechanist, an agentic framework that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. The work addresses the growing gap between AI model capabilities and our ability to understand and control them.
- Clustered Randomized Smoothing for Stochastic Prediction Functions — Clustered Randomized Smoothing for Stochastic Prediction Functions proposes a framework to address the mode collapse problem in randomized smoothing for stochastic, multi-modal predictors.
- Towards Truly Unsupervised Evaluation of Feature Selection -- Extended Version — The paper "Towards Truly Unsupervised Evaluation of Feature Selection" by Hafiz Saud Arshad, Muhammad Rajabinasab, and Arthur Zimek addresses the problem that established "unsupervised" evaluation techniques for feature selection are not truly unsupervised.
- Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations — Authors: Moshe Butman, Lior Baruch, Doron Friedman, Kfir Bar (Reichman University) Published: Conference paper at ICLR 2025 --- The paper introduces a novel framework called Preference Tree Optimization (PTO) designed to iteratively improve agent models in goal-oriented dialogue
- A Comparison of Malware Image Transformations Using Grad-CAM and Hybrid Learning Models — This paper investigates the use of Gradient-weighted Class Activation Maps (Grad-CAM) as an eXplainable AI (XAI) tool for analyzing eight distinct image types derived from malware samples. The research builds on prior work by Agrawal et al.
- Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models — Authors: Shukrullo Nazirjonov, Sai Prasanna, Anna Manasyan, Georg Martius (University of Tuebingen, Max Planck Institute for Intelligent Systems) arXiv: 2608.12078v1 [cs.CV], 12 Aug 2026 --- The paper investigates whether object-centric (OC) representations actually deliver for p
- Look What the Probes Dragged In! Real-World Chest X-ray Shortcuts in MedCLIP — The paper investigates how real-world shortcuts manifest across different layers of the medical CLIP-based model MedCLIP and its vision encoder, a frozen ResNet-50.
- RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation — RT-SEMamba is a fully causal speech enhancement (SE) model built upon causal time–frequency Mamba blocks.
- No One to Blame: A Framework of Constitutive AI Unaccountability — This paper introduces the concept of "constitutive AI unaccountability" to describe sociotechnical configurations in which AI accountability is conceptually unachievable, regardless of effort.
- QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving — QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving Abstract Summary Retrieval-Augmented Generation (RAG) repeatedly prefills identical text chunks across queries, incurring redundant computations.
- Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control — Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control, by Josef Chen, addresses the question of when deterministic control transitions between model and tool calls in LLM-agent services can be grouped and executed profitably on a GPU.
- Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation — This paper investigates how large language model (LLM) agents respond to graded similarity signals in strategic interactions, and whether such signals can induce cooperative behavior.
- SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges — Summary of the Paper "SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges" 1. Core Problem and Motivation The paper addresses the limitations of existing retrieval-augmented generation (RAG) systems for multi-hop question answering.
- GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings — GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings This paper introduces GUIDE, a governed multi-agent framework designed to transform heterogeneous enterprise guideline documents into structured, deployment-ready work artifacts.
- Adversarial Resilience of Poisson-Process Submodular Maximization over Matroids, and Full-Bandit Learning — Summary This paper studies the problem of nonnegative submodular maximization subject to a general matroid constraint when the offline algorithm is given an arbitrary controlled value oracle.
- A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench — This paper evaluates VITA, a retrieval-augmented generation (RAG) system purpose-built for context-specific knowledge retrieval in India, against several frontier general-purpose large language models (LLMs) on the HealthBench clinical reasoning benchmark.
- FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees — FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees This paper presents the FQTree algorithm for fine-grained quantization-aware training of boosted decision trees (BDTs), together with the QXGB framework for automatic hardware generation targeting
- Autonomous Telerehabilitation via Skeletal Motion Prediction and Joint-Level Performance Assessment — This paper presents a telerehabilitation pipeline that integrates skeleton-based exercise quality assessment and short-term motion prediction into a two-module system operating on marker-free RGB video.
- RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning — Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems.
- Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts — This paper proposes a personalized scorer modeling framework for generating robust sleep stage labels from multi-scored sleep datasets.
- Don't Cut Corners: How Training Outside the Prior Makes Simulation-Based Inference More Robust — Large astrophysical simulation campaigns often generate training data by sampling parameters across a Uniform prior box. Due to the proposal’s sharp edge, neural posterior estimators struggle to learn accurate approximations near the boundaries.
- Learning Under Treatment-Induced Label Indeterminacy with Expert Annotations of Counterfactual Outcomes: A Case Study in Neurological Prognostication — This paper addresses the problem of treatment-induced label indeterminacy in clinical prediction models, using post-cardiac-arrest neurological prognostication as a case study.
- DIVE: Unlocking Self-Improvement in Frozen Language Models Through Diversity-Driven Skill Evolution — Large language models (LLMs) "cannot retain postdeployment experience without parameter updates." The paper introduces DIVE, "a diversity-driven framework that enables frozen LLMs to improve by evolving persistent natural-language skills from task experience and verifier feedback
- SoK: From Generation to Consumption of Privacy Documents in Software Systems — This SoK paper provides a unified, lifecycle-oriented view of privacy documents from a software engineering perspective.
- SchemaLink: An Intelligent Web Editor for LinkML Schema Curation — SchemaLink is a web-based environment for the graphical construction and enhancement of LinkML schemas, designed to address challenges faced by novice curators in developing and maintaining LinkML schemas.
- GENADA: efficient generative time series adversarial attack framework — GENADA: efficient generative time series adversarial attack framework This paper introduces GENADA (GENerative ADversarial Attack), a generative framework for adversarial attacks on time series classification models.
- Analysis of Motor Signatures of Social Adaptation in Autism for Efficient Human-Centric Systems — This paper proposes a computational analysis framework to identify potential biomarkers of autism-related motor behavior, specifically focusing on how social context modulates motor imitation.
- Is this Citation on Point? — This paper studies proposition-level citation support verification in legal documents, focusing on whether current LLMs can detect when a legal citation points to a real case but does not support the proposition for which it is offered.
- Prof-K: Probabilistic One-Pass Filtering for Efficient Top-k Selection — Prof-K is a probabilistic one-pass filtering algorithm for efficient top-k selection, introduced by Tadeusz Dziarmaga, Witold Sikora, Łukasz Struski, Jacek Tabor, and Marcin Mazur from Jagiellonian University.
- DYSANOS Generative Dynamic Smooth Arbitrage-free Non-parametric Option Surfaces — DYSANOS is the first generative market model for smooth SANOS option surfaces for all strikes and expiries which are free of static arbitrage. The model is designed to generate entire paths of daily spot and option prices for years in the future.
- Supervised Mixed-Frequency Learning for Macro-Financial Forecasting When Factors are Weak — Summary This paper, "Supervised Mixed-Frequency Learning for Macro-Financial Forecasting When Factors are Weak" by Ulrich Hounyo and Zhendong Li, addresses the problem of weak factors in mixed-frequency forecasting.
- What Makes a Peer? Valuation-Anchored Similarity in Private Markets — The paper introduces a supervised similarity learning framework for identifying economically meaningful peer companies in private markets, where limited transparency, sparse disclosures, and infrequent transactions make traditional peer identification challenging.
- Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks — This paper addresses the question of how large a random low-dimensional search space must be to access parameters that achieve low loss when training neural networks through random low-dimensional reparameterization.
- Intensional Anaphora — The paper "Intensional anaphora" by Ezra Keshet and Steven Abney (2024) addresses the problem of anaphora to antecedents in intensional contexts, building on prior work by Stone (1999), Stone & Hardt (1999), and Brasoveanu (2010).
- Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues — This paper studies a failure mode in multi-turn LLM dialogues where constraints that should no longer bind still do.
- PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping — PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping Abstract Summary: The paper addresses the critical challenge of labeled training data scarcity in deploying online HD map construction systems to real-world scenarios.
- @skills: Attention is all you have — The paper identifies a fundamental mismatch in the agent skills ecosystem: 56,804 indexed skills compete for fewer than 100 reliable auto-trigger slots per agent.
- LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning — This paper introduces EpicStar, a framework for enhancing strategic reasoning in Large Language Models (LLMs) within long-horizon environments, using StarCraft II as the testbed.
- EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory — This paper introduces EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval), a long-horizon agentic memory framework for egocentric question answering (QA).
- Novels generated by language models show compressed formal variation — This study investigates whether large language models (LLMs) can produce the same range of formal diversity across complete novels as human-written corpora, rather than focusing on detecting individual AI-generated passages.
- Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence — The Wiggle Framework is a unified stress test for epistemic stability in LLM judges.
- SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries — SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries Abstract Long-running LLM agents act through tools. A single step can send an email, merge a pull request, or wire a payment.
- General Probabilities of Causation with Causal Knowledge — This paper addresses the question of whether additional causal knowledge can further tighten the bounds of probabilities of causation (PoCs) in multivalued settings, where treatments and outcomes can take multiple values.
- Evaluating AlphaEarth Foundations Embeddings for Wildfire Susceptibility Mapping — This paper systematically evaluates AlphaEarth Foundations (AEF) embeddings for wildfire susceptibility mapping, using Victoria, Australia (2017–2025) as a case study.
- Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers — This paper investigates the role of municipal algorithm registers in providing transparency and facilitating accountability for algorithmic systems in public services, using a case study of a Dutch city's register and a decision-support tool called Avola used for welfare benefits
- Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment — The paper presents Semantic Prism, a conditional semantic-image generation-and-refinement framework with deterministic inference for semantic segmentation.
- ADEPT: A Unified Framework for Deep Learning Test Adequacy — ADEPT: A Unified Framework for Deep Learning Test Adequacy This paper presents ADEPT, an extensible framework for running deep learning (DL) test adequacy metrics.
- The Boolean Power of ReLU — We prove that, on finite simple undirected graphs equipped with a single Boolean node feature, the Boolean queries expressible in Σ-MPLang, for any collection Σ of eventually constant activation functions and with arbitrary real coefficients, form a strict subclass of the Boole
- Beyond Local Power: Functional Connectivity Analysis for Subject-Independent Learning Style Recognition — This paper proposes an objective Electroencephalography (EEG) approach evaluating Phase Locking Value (PLV) connectivity against localized features across the Active–Reflective (AR) and Verbal–Visual (VV) Felder-Silverman dimensions.
- From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection — This paper presents a closed-loop framework for video reflection removal, unifying physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation.
- Dual-Model Sentiment Analysis of Consumer Reviews in the Retail Coffee Sector Using Machine Learning and Deep Learning Approaches — This study presents a comprehensive sentiment analysis framework applied to Starbucks customer reviews, leveraging both classical machine learning and advanced deep learning models.
- Remote Sensing and Machine Learning-Based Analysis of Land Use and Vegetation Change in Dhaka District, Bangladesh — This study employs remote sensing data and machine learning techniques to analyze spatiotemporal changes in land cover and vegetation dynamics in Dhaka District, Bangladesh, between 2019 and 2024.
- EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval — The paper identifies a "critical reliability gap" in enterprise RAG (Retrieval-Augmented Generation) deployments.
- Attractor Image-Based Deep Learning of Arterial Pulse Waves for Age Classification — This study investigates whether arterial pulse waveform morphology, which evolves with age and reflects structural and functional changes in the cardiovascular system, can be used to classify chronological age in healthy adults using a deep learning approach.
- A Quantum/Classical Example Oracle Separation for Making Things Up — The paper studies the power of quantum examples compared to classical examples in the Probably Approximately Correct (PAC) learning framework.
- Structuring the Space of Perspectives — The same event can be reported from different perspectives depending on the experiences, background, and beliefs of the writer or speaker. A variety of NLP areas engage with perspectives, spanning from text analysis to algorithm optimization.
- DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation — DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation Abstract Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navig
- Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection — Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection Summary This paper addresses the problem of benchmark contamination in large language models, where a model's performance on a benchmark is inflated because the model has memori
- Policy Convergence and Divergence Across National and Within Regional AI Strategies: A Policy Design Element Analysis — Governments worldwide have responded to the rapid expansion of AI by publishing national and regional AI strategies.
- A Local-Linearly Convergent Algorithm for Nonconvex Equality-Constrained Optimization — The paper extends the analysis of the Gradient-Eigenstep Algorithm by Goyens et al. for solving nonconvex equality-constrained optimization problems. The contributions are twofold.
- CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence — CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence Summary This paper introduces the Causal Attribution Score (CAS), a compact score architecture for causal explanation in explainable artificial intelligence (XAI).
- CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications — Mobile GUI agents "remain brittle when deployed to applications absent from source training." The paper studies "novel-app generalization under a limited target interaction budget and without target demonstrations." The authors note that "an agent may instead encounter an applica
- Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study — Author: Simone Mungari (Revelis s.r.l.) arXiv: 2608.11649v1 [cs.CL] 12 Aug 2026 --- This paper investigates "whether and how LLMs express preferences toward political parties and political leaders." The authors introduce "a systematic and reproducible auditing framework in which
- Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences — Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences Summary This paper introduces Drive-to-Music, a context-aware system that generates music in real time from multimodal driving signals.
- HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment — HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment Abstract Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic promotion, and talent discovery.
- Not All Nudges Land: Behavioral Controllability and Elaboration Quality in AI-Supported Journaling — This paper analyzes 369 journal entries from an eight-week passive sensing study (the MindScape system) to determine which behaviors respond to AI-generated journaling nudges and what signals in users' writing predict follow-through.
- ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering — ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering Abstract Summary: Enterprise question answering is framed as retrieving internal documents and generating grounded answers.
- Geometric and Behavioral Stratification in Transformer Residual Streams — The paper investigates the geometric and behavioral organization of transformer residual streams, proposing that the prediction direction—the unembedding direction of the token a model currently predicts—acts as a "content-defined privileged anchor" around which residual-stre
- GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation — Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business.