AI papers — 2026-08-11
Gromov-Wasserstein Quantization and Clustering: Structure, Rates, and Algorithms. This paper introduces and analyzes the Gromov–Wasserstein (GW) quantization... A Recommendation System Approach for Interference-Robust Sensor Subset Selection. This paper develops a method for sensor-subset selection for tracking. A Single Atom in Front of a Mirror is a Universal Reservoir Computer. Core Claim The paper demonstrates that "the minimal architecture can carry the... DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation. DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation... Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation. Dual-Domain Cross-Modal Decoding (DD-CMD) is proposed for clinical text-guided... Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers. This paper investigates whether token-pruning policies transfer across... Most biomedical publications show signs of LLM-assisted writing. Based on the paper "Most Biomedical Publications Show Signs of LLM-Assisted... Hypothesis Frontier: Verifier Guided LLM and Symbolic Search for First-Order Induction. Hypothesis Frontier: Verifier-Guided LLM–Symbolic Search for First-Order... Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents. Core Thesis The paper makes a distinction that is easy to miss: detecting that... Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence. Ex-Omni-2D is an omni-modal dialogue framework that generates a coordinated... Synthesizing Probabilistic Saturating Counters with Differentially Private Formal Guarantees. This paper presents a formal security analysis of Probabilistic Saturating... PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition. PAC-Bayes theory provides generalization guarantees by controlling the... Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences. Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences Core... A Joint-Distribution Route to Fair Representations with Continuous Sensitive Attributes. This paper addresses fair representation learning when the sensitive attribute... Telemetry and Concealment in Self-Adapting Generative AI: Logging Architecture, Adversarial Model Hiding, and the Limits of Detection. This paper addresses the governance problem posed by continually self-adapting... MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction. MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction... Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation. This paper introduces a new formulation for camera-motion understanding called... ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes. ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled... Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts. This paper introduces MOSAIC (Model Optimization via Systems-Aware TraIning... Probing and steering biology across Boltz-1s trunk-diffusion boundary. This paper investigates how biological information is represented and... ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover. ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover... Pair-Centric Graph Rewiring for Over-Squashing via Optimal Transport-Guided Communication Alignment. This paper proposes PairAlign (PAR), a pair-centric graph rewiring framework... Diffract: Spectral View of LLM Domain Adaptation. The paper "Diffract: Spectral View of LLM Domain Adaptation" studies continual... ChemWorld: Programmable Chemical Worlds for Controlled and Replayable Agent Experimentation. ChemWorld is a programmable chemical environment in which reusable process and... Agent Safety Should Be a Runtime Contract. Position: The paper argues that agent safety should not be treated as a... Defending against Model Extraction for GNNs with Model Reprogramming. Graph Neural Networks (GNNs) serve as the backbone for high-stakes applications... DEFT: Data-Efficient Frequency-domain Top-k Sampling via Inverse Discrete Fourier Transform for Spatiotemporal Dynamical Systems Modeling. DEFT: Data-Efficient Frequency-Domain Top-K Sampling via Inverse Discrete... Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation. Generative AI models are primarily designed to imitate the data distribution,... Posterior contraction rates in Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families. The paper studies posterior contraction in positive-order Sobolev norms and... Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations. Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing... Towards Efficient Reasoning in LLM-Based Recommender Systems via Model Merging. This paper introduces REAM (Reasoning-HEad-Aware Merging), the first model... Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention. This paper studies the expressive power of attention mechanisms by isolating... Causality Sum Rules in Conventional Scattering Matrices. This paper, "Causality Sum Rules in Conventional Scattering Matrices" by Ning... SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation. SkillLens converts heterogeneous GUI experience into Visual Skill Cards (VSCs)... Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory. This paper proves inference-time quantum coordination advantages for specified... ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions. ProbGuard is the first completely probabilistic, architecture-agnostic... Recovering Wasted Compute in Autoresearch Agents. This paper studies the modeling pipeline at the core of autoresearch systems... TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation. TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator... How to Verify Consistency of Probabilistic Claims. Problem Statement This paper addresses the question of whether a probabilistic... EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection. This paper presents EVIL-Detect, a multi-signal ensemble framework with... Self-evolving network verifiers. Self-evolving network verifiers Authors: Ioannis Protogeros, Tibor Schneider,... GARLIC: Graph Attention-based Relational Learning of Multivariate Time Series in Intensive Care. GARLIC: Graph Attention-based Relational Learning of Multivariate Time Series... HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging. HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging introduces... Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence. Here is a detailed summary of the paper "Apodex Discovery: Reality Benchmarks... ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls. ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue... Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration. PCFlow: Perceptually Consistent Flow Matching for Efficient Image Restoration... Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases. This paper introduces an analytically exact framework for the controlled... Optimistic Rates for Multiclass PAC Learning. This paper resolves the open problem of optimistic rates for multiclass PAC... SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning. SPEED TUNING: Speeding Up Policy Execution with Lightweight Reinforcement... Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives. This survey presents a unified review of cross-view feature matching, a... Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data. Workflow Cards: Structured Summaries of Workflow Executions Using Provenance... MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows. MAP-Graph is a provenance-aware memory layer for multi-agent workflows that... TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation. TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal... The Illusion of Cross-Lingual Safety in Low-Resource Languages. This paper investigates whether safety alignment in large language models... From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop. The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located... CosMAP: Contrastive Manifold Approximation and Projection for Dimensionality Reduction of Omics and Genealogical Data. CosMAP: Contrastive Manifold Approximation and Projection for Dimensionality... Bayesian Symbolic Regression with Entropic Reinforcement Learning. Bayesian Symbolic Regression with Entropic Reinforcement Learning Oussama... Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models. DURA is a diffusion-based unrestricted robotic attack that generates visually... Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome. Operationalising Relative Causal Knowledge: Backbone Identifiability from... DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation. DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student... From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents. Problem and Motivation Persistent memory in language-model agents enables... Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition. The paper introduces the Whisper-Aware LLM, a framework designed to address the... MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models. MedUP: Awakening Unified Understanding and Perception in Medical... Hierarchical Compositionality for An Assistive AI Agent. This paper presents an architecture for personalized command disambiguation in... RadFusion: Towards Threshold-Controllable Radiology Report Generation. RadFusion: Towards Threshold-Controllable Radiology Report Generation Summary... Mapping and Measuring the Behavioral Evolution of Large Language Models. The paper "Mapping and Measuring the Behavioral Evolution of Large Language... GitSkills: A Dataset of Agent Skills on GitHub. GitSkills is a dataset of 3,797,117 SKILL.md files collected from 282,200... Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent. The paper introduces DataMaster, an agentic instruction data selection system... A lower bound for stepsize-based acceleration of gradient descent. This paper establishes a new lower bound on the convergence rate of plain... Improving TensorSketch Using Complex Random Variables. Improving TensorSketch Using Complex Random Variables Summary This paper... Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport. Weightless Fine-Tuning (WFT) is a training-free, decoding-time method that... Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models. This paper demonstrates that large language models (LLMs) can serve as... SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features. SQuaT (Student-Aware Quantized Teacher Features) is a label-free... Towards Unified Dynamic Face Landmark Detection. This paper introduces Unified Dynamic Face Landmark Detection, a novel... Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics. This paper presents SAMPLED-BPE, a lightweight token-level auditing pipeline... Uncertainty-Aware Deep Learning for Genomics Applications: Insights from an Empirical Study. This paper presents an empirical analysis of uncertainty quantification (UQ) in... Uncertainty-Aware Compositional Localization and Placement Assessment of Catheters and Tubes in Chest X-Rays. Uncertainty-Aware Compositional Localization and Placement Assessment of... Multi-Granular Rationale-Guided Molecular LLM for Property Prediction. MR-MoL is a multi-granular rationale-guided molecular LLM for property... Tensor-normal maximum likelihood estimation at the operator-norm sample threshold. Let be independent Gaussian tensors in whose covariance is a Kronecker product... Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost. SpeedRunner is a method for programmatic skill learning that reduces agent cost. Generative Learning for Quantum Measurement Design. Generative Learning for Quantum Measurement Design Authors: Jun Dai, Olivier... Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue. The paper introduces a dual-loop self-evolution framework for multi-turn... Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR. This paper presents a rigorous, multi-seed evaluation of Automatic Speech... Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection. The paper proposes CALIBDCD, a calibration framework for feature-based data... Physics-informed Diffusion Generative Model for Time-Series Data Synthesis in Dynamic Systems. PhysDGM is a stepwise physics-embedded diffusion generative model for... Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning. Core Contribution The paper introduces J-Access, an inference-time auditing... Mitigating Context Interference for Reliable and Efficient Search Agents. This paper investigates the issue of context interference in multi-turn search... MoE Proxy Models for Low-Cost Failure Reproduction and Diagnosis in LLM RL Post-Training. This paper proposes a multi-view, frequency-aware expert pruning method for... When Is a General Factor Distinguishable? Non-Proportionality, Stable Structure, and the Bifactor Decision. When Is a General Factor Distinguishable? Non-Proportionality, Stable... Simplex Relaxation for Discrete Diffusion. Core Contribution The paper introduces Simplax, "an exact... MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection. MD-ProTector is an input-only encoder detector for LLM-generated text detection... What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research. This paper synthesizes current knowledge about Responsible AI (RAI) practices... The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark. SRE-Bench is the first realistic, contamination-free reverse engineering (RE)... Partially Observable Learning for Multi-Platform Dispatch Optimization. This paper proposes POLO, a Partially Observable multi-agent reinforcement... Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction. Core Contribution This paper develops a decision-resource view of terminal... Backdoor Decontamination Dynamics in LLM Agents. Backdoor Decontamination Dynamics in LLM Agents Authors: Gabriel Huang, Abhay... Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding. Agentic coding READMEs like CLAUDE.md grow without bound in real repositories,... Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes. The paper introduces AD2-Bench, a large-scale benchmark for evaluating... Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs. This paper introduces the first any-to-any backdoor attack on Vision-Language... BREAD: Baseline-Referenced Explanations for Anomaly Diagnosis. BREAD: Baseline-Referenced Explanations for Anomaly Diagnosis proposes a... Adaptation of Generalist Robot Policies with Minimal Data. Core Problem and Setting The paper introduces minimal-data adaptation (MDA), a... Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter. This paper argues that tokenizer vocabulary size in large language models... Market-Information-Aware Gated-LoRA of Foundation Models for Transferable Day-Ahead Electricity Price Forecasting. This paper proposes a market-information-aware adaptation framework that... Cross-Corpus Evaluation of Generalizable Vulnerability Detection in IoT Firmware. This paper introduces IoTVulBench, a human-verified benchmark for cross-corpus... Entropy-based Code Adversarial Translation for Real-world Repository Migration. This paper introduces Entropy-based Code Adversarial Translation (ECAT), a... Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation. Mixture-of-experts (MoE) models have recently moved beyond routing a fixed... TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs. TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs... Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training. The paper "Unlocking the Power of Medical Tabular Data via Semantic-Aware... Rethinking Text-Based Image Retrieval in Specific Domain. The paper introduces SecMM-TBIR, a multi-match benchmark for surveillance... Conversational Orchestration for Organic 6G. The paper proposes a lightweight, decentralized conversational orchestration... Association-based Privacy Attacks in Wireless Protocols: Formal Modeling and Mitigation. This paper formally investigates the root causes of pairing-based privacy... Principal Trait Analysis: Towards Deriving "Skills" in Human-AI Collaboration. Principal Trait Analysis (PTA) is a novel, data-driven algorithm inspired by... The GenAI Catch-22: Use of Generative Artificial Intelligence in Norwegian Newsrooms During the 2025 Parliamentary Election. The GenAI Catch-22: Use of Generative Artificial Intelligence in Norwegian... Conditional Independence Tests for Constraint-Based Causal Discovery: A Survey. This survey reviews conditional independence (CI) testing with emphasis on... Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task. This paper introduces a novel verbal-only conflict task, the "crayon task," to... Rationale-Guided Learning for Multimodal Emotion Recognition. RATIONALE-GUIDED LEARNING FOR MULTIMODAL EMOTION RECOGNITION The paper proposes... Dueling Deep Q-Learning for Intrusion Detection. This study proposes a novel approach to intrusion detection systems (IDS) by... Scheduling Mixed RL Rollouts Beyond Prefix Locality. MISA-T is a routing-layer admission policy for mixed rollout serving in... Leveraging Human Reading Behavior for Keyphrase Extraction: A Webcam-based Eye-tracking Corpus. This paper investigates whether lightweight webcam-based eye-tracking features... On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models. This systematic literature review, conducted under PRISMA 2020 guidelines,... myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR. This paper presents myMediWhisper, a Burmese medical speech recognition... Entropy-Centric Explainable AI for Remote Sensing Image Segmentation. This paper proposes an entropy-centric explainable AI (XAI) method for semantic... Contextual Information Policy Optimization for Search Agents. The paper, authored by Xingyu Guo, Wei Chen, Linlin Yang, and Baochang Zhang... Dual-Primal Graph VAEs for Noisy Label Aggregation. Dual-Primal Graph VAEs for Noisy Label Aggregation proposes a graph VAE... RLMOpt: Adaptive Prompt Optimization via Recursive Language Models. RLMOpt is a prompt optimizer that makes the search policy itself... Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique. Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic... ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation. ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware... FUSE: Frame-Unified Stress Estimation from Facial Video. FUSE (Frame-Unified Stress Estimation) is a facial-video stress detection... Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects. This paper addresses the problem of reasoning shortcuts in neurosymbolic... Beyond Forecasting: Recasting Volatility Control as a Routing Problem. Beyond Forecasting: Recasting Volatility Control as a Routing Problem Core... Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning. Continuous Interaction Diffusion (CID) is a diffusion-native model–runtime... Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry. Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean... Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning. TPSP introduces a policy-aware scene encoder to capture the interaction between... Basin: Efficient and Extensible Numerical Optimization in Rust. Basin is a numerical optimization library for the Rust programming language... InSight-doc: Agentic Visual Perception for Long-Document Understanding. InSight-doc: Agentic Visual Perception for Long-Document Understanding proposes... Hardware-Aware Deployment of Joint SAR Compression and Despeckling on FPGA. This paper presents the first deployment of a joint SAR Despeckling and Data... Click2Poly: A VLM for vector mapping buildings and walls. Click2Poly is a human-in-the-loop AI assistant designed to speed up the manual... BooST: Bridging Semantics and Motions for Efficient Skill Transfer. BooST: Bridging Semantics and Motions for Efficient Skill Transfer introduces a... Language-Structured Relational Q-Learning for Threat-Aware Control in Safety-Critical Driving. The paper introduces Language-Structured Relational Q-Learning, instantiated... Retrieval-Augmented Vision Foundation Models for Robust Leukemia Cell Classification across Multiple Microscopy Datasets. This work presents a robust framework for leukemia classification across... TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling. TIDE RL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling Abstract... UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations. UniProbe is a lightweight, unified, learnable detector for token-level... IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning. IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization... Forward Trajectory Steering for Hamilton-Jacobi Reachability Analysis. This paper introduces STEER2REACH (S2R), a physics-informed neural network... Diffusion-Based Data-Driven Assortment Optimization. Diffusion-Based Data-Driven Assortment Optimization proposes D3AO, a... RelShap: Relationally Consistent Shapley Explanations. RelShap is a framework that incorporates relational constraints and data... Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness. The paper establishes convergence guarantees for gradient descent applied to... Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling. The paper presents a substantial expansion of the IMPROVE benchmark dataset for... -VAEs as Effective Theories: Tolerance-Dependent Dimension. In a beta-VAE, increasing the regularization strength acts as a spectral cutoff... Generator-Guided Inverse Sampling for Lévy-Driven Generative Models. This paper studies inverse sampling for Lévy-driven generative models from the... Invertible Logits Transformation for Accuracy-Preserving Post-Hoc Uncertainty Calibration. The paper introduces Invertible Logits Transformation (InvLT), a post-hoc... Accelerated Learning of High Dimensional Functions with a Tensor-Featured Training Network. This work presents a method to accelerate the optimization of learning high... Efficient Weak-Entropy PINN for Solving Hyperbolic Conservation Laws. The paper introduces a novel physics-informed neural network framework called... MUSE: A Full-Text Cross-Domain Knowledge Base of Scientific Problems, Solutions, and Rationales. MUSE: A Full-Text Cross-Domain Knowledge Base of Scientific Problems,... VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation. The paper introduces VoxSumm, a multilingual corpus and benchmark for joint... Stigma and Support in Online Sexual Violence Narratives on Reddit. This paper introduces the SCOPE (Stigma and COmmunity Peer Expressions)... Decomposition-Induced Context-Memory Conflict: When Fact-Checking Pipelines Contradict Their Own Source Text. Decomposition-Induced Context-Memory Conflict: When Fact-Checking Pipelines... REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs. The paper presents R EAP (Relation-aware Elicitation And Parsing), a system for... Assessing Reliability of BERT-Based Models on Question Answering Tasks. This study evaluates the reliability of four BERT-based models—RoBERTa,... Evo-Bench: Can Language Models Improve Agent Harness?. Evo-Bench is the first benchmark designed to evaluate large language models'... ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering. ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended... ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS. ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech... Who Gets Heeded? An Obligation-Level Audit of Responsiveness in EPA Rulemaking. This paper introduces obligation-level responsiveness auditing, an auditable,... Stability of Finite-Batch Particle Mean-Field Variational Inference Beyond Strong Convexity. This paper studies the stability of finite-batch particle mean-field... Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution. Emotion2Skill is a framework that extracts LLM-internal emotion vectors and... sLTN: Structural Logic Tensor Networks. sLTN: Structural Logic Tensor Networks introduces an extension of Logic Tensor... A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa. This study presents a comparative evaluation of six object detection... DACRI: Decision-Aware Causal Intervention Ranking for Critical Supply Chains. DACRI: Decision-Aware Causal Intervention Ranking for Critical Supply Chains... Efficient Hypergradient Descent for Inverse Reinforcement Learning. This paper addresses the computational challenges of inverse reinforcement... Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation. This paper introduces a Test-Time Self-Evolving framework for GUI visual... Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders. This paper investigates whether the interpretability of individual sparse... ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization. ReRound (Reconstructive Rounding) is a post-training quantization method that... Spectral Embeddings of Degree- Laplacians in Random Dot Product Graphs. This paper studies a continuum of degree-normalized spectral embeddings for... Path Integral Value Matching for Linear Quadratic Stochastic Optimal Control. The paper addresses the computational challenges in solving Linear Quadratic... Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness. Decoding-Level Taboo is a zero-prompt diagnostic stress test that intervenes... Long-Horizon Forecasting of Complete Financial Statements with Forma. Core Contribution and Problem Statement The paper introduces ProForma-20Q, a... Statistically-Secure Bit Commitment and Coin Flipping Protocols Based on Quantum Hardware Assumptions. This paper introduces the first statistically secure bit commitment and coin... Enhanced Filtering Algorithms for the Euclidean Traveling Salesperson Problem and its variants in Constraint Logic Programming. The paper "Enhanced Filtering Algorithms for the Euclidean Traveling... Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks. S INK F LEX-RL is a modular training system for reinforcement learning (RL) in... Spectral graph clustering with inhomogeneous latent geometry. Spectral graph clustering with inhomogeneous latent geometry Authors:... BPG: Balancing Plasticity and Generalization for Domain Incremental Learning. BPG: Balancing Plasticity and Generalization for Domain Incremental Learning... EvoMem: Memory-Augmented Evolution for Code Optimization. EvoMem is a persistent memory architecture for LLM-based evolutionary program... Iterative Erasure Count Is Not an Affine-Invariant Concept Dimension. Core Claim This paper argues that iterative erasure count is not an... Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy. This paper investigates whether a Large Language Model (LLM) can replace the... When and Where Faults Matter: A Study of Transient Errors in CKKS Multiplication. This paper presents an in-depth analysis of the resilience of homomorphic... Battlefield 5G: Dual-PKI and TPM-Based UE Attestation for Tactical 5G Standalone Networks. The paper presents Battlefield 5G, a pre-authentication framework for tactical... Goodness-of-Fit Tests and Calibration Machine-Learning Algorithms for Logistic Regression with Sparse Data. Based on the paper "Goodness-of-Fit Tests and Calibration Machine-Learning... Improved cross-validated distances for multivariate pattern analysis. The paper "Improved cross-validated distances for multivariate pattern... MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment. MultiModal Code-Switching: Interleaving Visual Objects into Language for... Beyond Detection Accuracy: Measuring Explanation Cost, Stability, and Utility for Resource-Aware IoT Intrusion Detection. This study jointly evaluates predictive effectiveness, explanation cost, local... VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?. VisEditBench is a benchmark for evaluating vision-language models (VLMs) on the... Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation. The paper studies persona conditioning as a diagnostic mechanism for exposing... Quantum Incremental Learning with Mixed State Prototypes. Quantum Incremental Learning with Mixed State Prototypes Abstract Incremental... RevCRN: Reversible Analog Computation using Chemical Reaction Networks. This paper introduces the Reversible Chemical Reaction Network (RevCRN) model... Predicting Space Groups of Double Perovskites by LLM with Dynamic Few-Shot Learning. DyRIS, an LLM-agent-based framework, predicts ranked space-group (SG)... Automatic Field-of-View Adjustment for a View-Expansive Microscope via LSTM-Based Gaze and Pipette Motion Interpretation. Intracytoplasmic sperm injection (ICSI) operators frequently adjust the... Benchmarking Time Series Generation Methods for Privacy-Preserving Forecasting. the paper "Benchmarking Synthetic Time Series Generation Methods for... Self-Knowledge Retrieval Augmented Generation Framework for Patent Matching. This paper proposes a self-knowledge retrieval-augmented generation (RAG)... Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting. Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series... ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation. On-policy distillation (OPD) applies token-level teacher supervision to... Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension. The paper proposes hierarchical empirical-Bayes Naive Bayes (HEB-NB), a... Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift. The paper introduces the Floor Certification Map, a theoretical framework for... 3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment. This paper presents the 3D Sheep Pain Facial Expression System (3D-SPFES), a... GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning. GeoForge is a training-free, self-evolving framework for Earth observation (EO)... CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classification. CLEAR (Class-wise reLiability-aware Expert Aggregation for long-tailed... A Lightweight Fault-Detection Scheme for Barrett Modular Multiplication Using Multiple Conditional Reduction Paths. This paper proposes a lightweight fault-detection scheme for Barrett Modular... Narrative Keyframing for Generative Creative Writing. The paper introduces narrative keyframing, a new interaction technique for... On the Sensitivity to Errors in Homomorphic Computing: Single Transient Bit-flip Client-side Error Characterization. This paper analyzes the sensitivity of Homomorphic Encryption (HE) to bit-level... A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona Problem. The paper reports on a production deployment of a centralized MCP (Model... Trigger the Straggler: Load Hijack on Mixture-of-Experts LLMs. This paper introduces Load Hijack, a supply-chain attack against... A Quantum Roadmap for Softmax Attention: Exact Born-Rule Analogs for Softmax Attention on the Probability Simplex. The paper presents a theoretical construction establishing an exact equivalence... Physics-Informed Implicit Neural Representations for Improved Myocardial Perfusion MRI Quantification. This paper presents a physics-informed neural network (PINN) framework with... Dynamics Models for Offline Hyperparameter Selection in Real-World RL. A key obstacle to deploying reinforcement learning in real-world systems is... Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits. Offline reinforcement learning (RL) offers considerable promise for optimizing... AutoGrable: What Is a Good Graph for a Table?. The paper addresses the fundamental question of graph construction for tabular... TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification. TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing... Gaussian Meta-Space Augmentation for Stacking Ensembles in Multimodal IPMN Risk Stratification. The paper introduces cUPMI (calibrated Upstream Probabilistic Meta-Imputation),... Self-Normalized Inference for Constant-Stepsize Temporal-Difference Learning under Markovian Sampling. This paper develops inferential methods for constant-stepsize... Attention-Path Fragility as an Uncertainty Signal in Large Language Models. The paper proposes that a model's uncertainty about a token is reflected not... Mechanism Design for Generative Engines: From Exploitation toward Win-Win Outcomes. Mechanism Design for Generative Engine: From Exploitation to Win-Win... ODE-Based Transformer Decoders for Iterative Sign Language Translation. Sign language translation has achieved strong results with Transformer... Self-Evolving Embodied Agents via Skill-Harness Evolution. Core Contribution The paper introduces SHAPER, a self-evolving framework for... ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models. ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language... An Empirical Study of Output-to-Input Loops for Black-Box Backdoor Detection in Fine-Tuned Open-Weight LLMs. An Empirical Study of Output-to-Input Loops for Black-Box Backdoor Detection in... Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory. ReTree is a self-correcting tree-structured memory mechanism for LLM-based... MIRA: Medical Image Reflection for Agentic Diagnosis. MIRA (Medical Image Reflection for Agentic Diagnosis) is a medical visual... ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling. ThinkRetrieve is a test-time scaling framework that augments the reasoning... RTSKG: Building a Rail Transit Station Knowledge Graph Dataset. RTSKG is a new rail transit station knowledge graph dataset that explicitly... Towards an approach to multivariate outlier detection for District Heating System data. This paper tests different methods for multivariate detection of outliers in... Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning. Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning... FedCGR: Federated Cross-Domain Generative Recommendation. FedCGR: Federated Cross-Domain Generative Recommendation proposes a federated... Socioduality: A Relational Process Framework for Human-AI Interaction. Socioduality is defined as "a sequential, reciprocal, and history-carrying... IO Factory: Simulating AI-Enabled Influence Campaigns at Scale. IO Factory is an AI-driven framework for simulating information and influence... FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation. The paper introduces FaithformBench, a benchmark for evaluating the... -SUB: A Physics-Informed Synthetic Underwater Benchmark Dataset for Underwater Image Enhancement. This paper presents π-SUB, a physics-informed framework for generating... Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement. This paper presents a pre-registered four-week longitudinal study (N = 72,... Cross-View Sequential Visual Localization with Spatio-Temporal Context Modeling for Autonomous Driving. This paper proposes a temporal-context-enhanced framework for cross-view... Benchmarking Cyberattack Detection in Electric Vehicle Charging Infrastructure with Benign User Updates. This paper develops a leakage-controlled session-level benchmark for... Variational Parameter Calibration with Physics-Aware Latent-Space Surrogates. This paper introduces a physics-aware neural-network-based latent-space... Information Bottleneck under Perfect Privacy. The paper studies the information bottleneck problem under a perfect privacy... Long-Time Trajectory Approximation via SA-NODEs: Model Predictive and Floquet Strategies. This paper studies the approximation of dynamical systems by semi-autonomous... SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training. SCOUT is a unified runtime failure-localization framework for LLM pre-training,... What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model. This paper asks what iterated self-feeding probes of language models measure,... Link-adaptive digital twin for robust physical-layer modeling in hybrid-amplified ultra-wideband optical networks. Link-Adaptive Digital Twin for Robust Physical-Layer Modeling in... CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening. CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and... Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving. Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient... Can Bayesian Optimization Efficiently Find a Strong Single Expert in Neural Thickets?. The paper investigates whether Bayesian optimization (BO) can efficiently find... MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices. MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative... Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement. This paper introduces expert-guided g-computation (egg-computation), a novel... MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model. MazzikaAI is a knowledge-based system for real-time Arabic maqam accompaniment... Strengthening Full Justified Representation: Efficient Verification and Computation. This paper introduces FJR+, a strict strengthening of the full justified... Let it Cook: Learning to Wait in Sequential Decision Making. The paper addresses the question of whether agents in sequential decision... From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate. The paper "From Numbers to Judgment: Specialist LLM Agents and Reinforcement... Contextual Quality-Diversity Evolutionary Reinforcement Learning for HVAC Control in Tropical Commercial Buildings. This paper proposes CQD-ERL, a contextual quality-diversity evolutionary... V-FiLLM: Verified Financial LLM Reasoning Benchmark. V-FiLLM is a framework that generates financial reasoning benchmarks from... Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation. This paper presents the first large-scale empirical study investigating whether... From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation. Traditional offline recommendation evaluation relies heavily on complex,... Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology. Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical... Gaze Target Estimation Anywhere with Concepts. This paper introduces the Promptable Gaze Target Estimation (PGE) task and the... AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations. AlbumentationsX is a data augmentation library that stores the transform list,... Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards. The paper introduces a neuro-symbolic safety guard for end-to-end autonomous... Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval. This paper presents the first direct comparison of natively multimodal... Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits. The paper studies cooperative multi-agent bandits in continuous (Lipschitz)... Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification. This paper proposes an uncertainty-aware and explainable deep learning... VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus. This paper introduces VERDICT (VERification via Disagreement-Informed Coupled... Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models. Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language... Evaluating Rational Contracting in Natural Language. and Motivation The paper addresses the challenge of evaluating how well... A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models. This paper presents a cost-efficient routing pipeline for multilingual... TACTICL: Task-Aware Compression of Tabular ICL Models. TACTICL: Task-Aware Compression of Tabular ICL Models Abstract Summary: The... Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution. The paper addresses a fundamental limitation in automated research ideation... SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information. SPIEVAL is a human-curated benchmark introduced to evaluate large language... Retrieval-Corrected Conformal Prediction for Time Series. Retrieval–Corrected Conformal Prediction (RCCP) is a retrieval-augmented... Measuring Semantic Abstractness of SAE Features via Nonlocality. and Motivation The paper addresses a central challenge in Mechanistic... XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving. XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving... Compositional Benchmark Synthesis for Hierarchical Human Action Recognition. Compositional Benchmark Synthesis for Hierarchical Human Action Recognition... Smart Enough to Go Extinct? An Evolutionary Challenge to the Value of General Intelligence and Its Ethical Implications for AGI. The paper "Smart Enough to Go Extinct? An Evolutionary Challenge to the Value... Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry. The paper studies multi-agent multi-armed bandits (MAB) with heavy-tailed... ReLTEx: Reliable LLM-based Taxonomy Expansion. ReLTEx: Reliable LLM-based Taxonomy Expansion Zeinab Ghamlouch, Mehwish Alam... VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?. VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living... Optimal Stopping of Self-Refining Foundation Models. The paper "Optimal Stopping of Self-Refining Foundation Models" by Kim Hammar,... Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control. Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control Core... DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction. Problem and Motivation Proteolysis-targeting chimeras (PROTACs) induce protein... Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models. This paper introduces a novel denial-of-service (DoS) attack targeting... SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation. SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation... MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph. MEGA (Meta Evaluation-Grounded Adaptation) is presented as a self-evolving... Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models. The paper introduces SALT, a Semantically ALigned action Tokenizer, to address... DuplexWorld: Can voice agents help you get through the day?. DUPLEX WORLD introduces a benchmark for holistically evaluating... Curate Before You Connect: Identity and Ontology Tagging in a Production Knowledge Graph. This paper describes the ingestion and ontology-tagging layer that turns a... Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training. Large-scale neural recommender systems are typically trained with a softmax... Fisher8: Stabilizing Neural Heteroscedastic Regression via Output-Layer Fisher Geometry. Fisher8: Stabilizing Neural Heteroscedastic Regression via Output-Layer Fisher... FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data. FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data Viktoria Schuster,... Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance. This paper proposes a Conversational XAI interface powered by Large Language... Evaluation Resolution Confounds Learning-Rule Comparisons in Model-Brain RSA of Early Visual Cortex. The paper investigates a methodological confound in representational similarity... Predictive Allostatic Organization in Recurrent and Spiking Agents Under Partial Observability. Adaptive behavior under partial observability depends on internal organization... Stay or Stray - A Dynamical Systems Viewpoint of Popularity Bias. The paper "Stay or Stray – A Dynamical Systems Viewpoint of Popularity Bias"... SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning. SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning... Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation. Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic... MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale. MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at... R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video. R4DSG introduces a relative 4D scene graph memory for long egocentric video,... Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization. This paper presents the Reinforcement Learning for Laser Cutting (RL2C)... CARE: Confidence-Aware Reasoning for Reliable Medical VQA. CARE: Confidence-Aware Reasoning for Reliable Medical VQA proposes a framework... Analysis of Federated Aggregation under Model Poisoning and Backdoor Attacks: A Reconstructed Cross-Dataset and Cross-Architecture Benchmark. The paper presents a reconstructed comparative benchmark and evidence audit of... Threshold Structure of Optimal Policies in Restart POMDPs. We study a Restart POMDP (Partially Observable Markov Decision Process) on a... Post-Calibration Reliability Reranking of Relevance Decisions via Label-wise Monotone Projection. Web search, product search, and question-answering retrieval systems often... On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation. The paper introduces LingT2I, a new benchmark designed to evaluate... Clinical Feasibility of Low-Magnification Fluorescence Imaging for Breast Cancer Margin Detection Using Texture Analysis and Deep Learning. This study investigated how 4× and 10× magnifications affect margin-level... StreamFlow: Dynamic Memory Flows for Streaming Video Understanding. StreamFlow introduces an efficient visual memory framework for streaming video... Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?. The paper "Can Released LLM Vocabularies Support Token-Level Estimation of... Benchmarking LLM Judges for Mobile Agent Evaluation. M OBILE J UDGE B ENCH is introduced as "to our knowledge the first benchmark... DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?. DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real... Data Attribution of Emergent Misalignment with Persona Features. Core Research Question This paper investigates emergent misalignment (EM) in... X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction. X2-Turn presents a frame-synchronous turn state prediction method via... When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs. Title: When Self-Consistency Backfires: Majority Vote Hurts the Majority of... A Study of Kernel Telemetry Options for Security-Oriented Provenance. This paper studies kernel telemetry options for building the capture layer of... Knowledge-Graph-Guided Retrieval-Augmented LLMs for Explainable Root Cause Analysis in Automotive HiL Validation. This paper proposes a knowledge-graph-guided retrieval-augmented large language... How Robust Are LLMs to Vietnamese Dialects?. This paper introduces VialectBench, the first end-to-end human-annotated... DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition. DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech... Gloss-Free Representation Learning for Cross-Dataset Sign Spotting. Sign-language research for resource-constrained languages is often limited by... Derivative Computation in PINNs: Automatic Differentiation, Finite Differences and Beyond. The paper systematically investigates finite-difference (FD) derivative... Topological Feasibility Guarantees for Differentiable Predictive Control. This paper establishes deterministic feasibility guarantees for differentiable... Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces. The paper proposes an Inverse Theory of Mind (IToM) pipeline for content... CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation. CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation This... Federated Learning for Distributed CNC Tool Wear Prediction. This paper investigates federated learning for CNC tool wear prediction using... EliSeg: Verified Target Construction for Report-Grounded Abnormality Segmentation. Problem Statement The paper addresses report-grounded abnormality segmentation,... ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment. the Paper Authors: Abdulkadir Küçe, Alihan Esen, Çağla Fikir, Berke Kurt,... XGBoost "is all you need": the case of forecasting transmitted heat energy in District Heating Systems. This paper presents a comparative study of two distinct approaches, XGBoost and... Blast Radius. by M.Y. Strategies to Avoid Illegal Data Access. This study examines technology solutions, personnel training, and policy... Multiclass Sentiment Analysis for Identifying Political Viewpoints. The paper investigates multiclass sentiment analysis of political viewpoints on... Reoptimization Algorithms for Contextual Bandits with Knapsack Constraints. The paper "Reoptimization Algorithms for Contextual Bandits with Knapsack... Decision-Aware Approximation of Belief Functions for Evidential Combinatorial Optimization. The paper introduces a decision-aware approximation method for belief functions... When Agents Talk: Honeytokens under Shared Memory. The paper "When Agents Talk: Honeytokens under Shared Memory" by Joshua S. Modelling Geographic Atrophy Progression using Implicit Neural Representations. Age-related Macular Degeneration (AMD) is the major cause of blindness in the... Every pooling rule has its world: matching probability combination rules to situations and stakes. Every pooling rule has its world: matching probability combination rules to... The Signal Rail: A Deterministic Motion Grammar for Communicating Conversational Agent State in Terminal Interfaces. The Signal Rail: A Deterministic Motion Grammar for Communicating... FITTER: Vocabulary-Agnostic Cross-Domain Inference on Temporal Knowledge Graphs. FITTER: Vocabulary-Agnostic Cross-Domain Inference on Temporal Knowledge Graphs... Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure. The paper "Why Post-Norm Transformers Collapse: Attention Amplification and... SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure. SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by... MVTrack: Ultrafast Appearance-Free Moving Object Tracking from Compressed Bitstreams. MVTrack is an ultrafast tracker for moving objects that operates directly on... REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems. REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent... A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language. This paper introduces a new dataset and baseline models for fine-grained... A Runtime Decentralized Attestation and Coordinated Repair Framework for Securing Automotive ECUs. The paper introduces DACER, a runtime decentralized attestation and coordinated... A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization. SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration) is an... A Systematic Sample Size Analysis of ML-Based Path Loss Prediction for LPWAN. This paper investigates how training-set size affects the accuracy of machine... VIDS-Seg: Towards Reliable Uncertainty Quantification in Pediatric Cardiac Ultrasound Segmentation. VIDS-Seg: Towards Reliable Uncertainty Quantification in Pediatric Cardiac...
The papers
- A Joint-Distribution Route to Fair Representations with Continuous Sensitive Attributes — This paper addresses fair representation learning when the sensitive attribute S is continuous (e.g., age, income, risk score).
- RAISE: Diagnosing Acquisition Collapse in Costly LLM Signals — The paper makes a distinction that is easy to miss: detecting that an auxiliary, model-derived signal helps on average is not the same as learning to act on it per instance. A reward–SNR floor governs when the second is even possible.
- Causality Sum Rules in Conventional Scattering Matrices — Summary This paper, "Causality Sum Rules in Conventional Scattering Matrices" by Ning Han, Rui Zhao, Shuxing Yang, Mingzhu Li, Hongsheng Chen, and Yihao Yang, presents a theoretical framework that formulates causality sum rules directly in the conventional multichannel scattering
- Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation — Dual-Domain Cross-Modal Decoding (DD-CMD) is proposed for clinical text-guided pulmonary infection segmentation, integrating two complementary forms of language guidance during decoding.
- Probing and steering biology across Boltz-1s trunk-diffusion boundary — This paper investigates how biological information is represented and transformed across the architectural boundary between the Pairformer trunk and the diffusion module in Boltz-1, an AlphaFold3-class protein structure predictor.
- PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition — PAC-Bayes theory provides generalization guarantees by controlling the Kullback–Leibler (KL) divergence between posterior and prior distributions over a chosen hypothesis representation.
- Pair-Centric Graph Rewiring for Over-Squashing via Optimal Transport-Guided Communication Alignment — This paper proposes PairAlign (PAR), a pair-centric graph rewiring framework that addresses the over-squashing problem in message-passing neural networks (MPNNs) by treating it as a pair-level communication shortage.
- MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction — MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction Problem Statement Industrial ads ranking systems estimate impression conversion probability by factorizing it into click-through rate (CTR) and post-click conversion rate (CVR) as an exact marginal factori
- Recovering Wasted Compute in Autoresearch Agents — This paper studies the modeling pipeline at the core of autoresearch systems and identifies common failure modes when applied to tabular datasets: "(1) they waste compute resolving the same bugs over and over again; (2) they often fail to tune hyperparameters even when they have
- ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions — ProbGuard is the first completely probabilistic, architecture-agnostic guardrail that leverages LLM early output distributional signals to estimate and calibrate the safety risk of continued generation, enabling early stopping of unsafe ongoing outputs.
- Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences — This paper identifies a foundational bottleneck in generative multimedia evaluation: the "judgment crisis." As the authors state, "While human perception naturally synthesizes the temporal and logical flow of a story, automated evaluation systems remain largely 'blind' to sequent
- Towards Efficient Reasoning in LLM-Based Recommender Systems via Model Merging — This paper introduces REAM (Reasoning-HEad-Aware Merging), the first model merging framework for reasoning compression in LLM-based recommender systems.
- ChemWorld: Programmable Chemical Worlds for Controlled and Replayable Agent Experimentation — ChemWorld is a programmable chemical environment in which reusable process and observation components are compiled into executable worlds. ChemWorld separates the public experimental contract available to an agent from evaluator-owned chemical and material laws.
- Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory — This paper proves inference-time quantum coordination advantages for specified AI state-tracking tasks. A solver compresses semantic history into a future-accessible boundary state and later answers a query.
- Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention — This paper studies the expressive power of attention mechanisms by isolating the basic operation of content-dependent selection, specifically through the Minimum Inner Product (Min-IP) task.
- TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation — TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation This paper introduces TangPoetryBench, the first multi-dimensional, human-grounded benchmark specifically built for poetry-to-image generation, along with PoemAutoEvalua
- Telemetry and Concealment in Self-Adapting Generative AI: Logging Architecture, Adversarial Model Hiding, and the Limits of Detection — This paper addresses the governance problem posed by continually self-adapting generative AI systems—models that update their own weights during production deployment—which fundamentally violate the static model lifecycle assumptions of traditional model risk management (MRM)
- Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers — This paper investigates whether token-pruning policies transfer across different vision tasks (image classification, semantic segmentation, and object detection) when using pretrained Vision Transformers.
- Defending against Model Extraction for GNNs with Model Reprogramming — Graph Neural Networks (GNNs) serve as the backbone for high-stakes applications in Machine-Learning-as-a-Service (MLaaS). Still, their black-box deployment exposes them to Model Extraction (ME) attacks, in which adversaries steal intellectual property by querying APIs.
- SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation — SkillLens converts heterogeneous GUI experience into Visual Skill Cards (VSCs) that provide visual procedural memory to frozen VLM executors.
- ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover — Authors: Minwoo Kim, Soochang Song, Namyoon Lee, Bang Chul Jung, and Yongjune Kim (POSTECH and Ajou University) arXiv: 2608.10545v1 [cs.NI] 11 Aug 2026 --- Edge LLMs must preserve inference continuity when a user hands over between edge nodes, requiring key-value (KV) cache trans
- Posterior contraction rates in Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families — The paper studies posterior contraction in positive-order Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families.
- DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation — DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation introduces a benchmark for evaluating large language models on the open-ended generation of interactive analytic dashboards.
- Gromov-Wasserstein Quantization and Clustering: Structure, Rates, and Algorithms — This paper introduces and analyzes the Gromov–Wasserstein (GW) quantization problem, which extends classical Wasserstein quantization (and thus k-means clustering) to the setting of gauged measure spaces, where the goal is to approximate a probability measure by a discrete one
- ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes — ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes Abstract Text-Attributed Graphs (TAGs), endowed with abundant textual content along with topological structures, have emerged as a versatile backbone for real-world anomaly de
- A Recommendation System Approach for Interference-Robust Sensor Subset Selection — This paper develops a method for sensor-subset selection for tracking.
- Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation — Summary This paper introduces a new formulation for camera-motion understanding called "temporally grounded, compositional recognition," which requires a model to localize motion-consistent intervals within a shot and identify every movement active within each interval.
- A Single Atom in Front of a Mirror is a Universal Reservoir Computer — The paper demonstrates that "the minimal architecture can carry the guarantee: a single atom before a mirror is a provably universal reservoir computer." The authors show that "universality can be associated with a single reservoir, considering a minimal setup of a single atom in
- Synthesizing Probabilistic Saturating Counters with Differentially Private Formal Guarantees — This paper presents a formal security analysis of Probabilistic Saturating Counters (PSCs) under Prime+Probe side-channel attacks, and proposes an enhanced PSC design with provable differential privacy guarantees.
- DEFT: Data-Efficient Frequency-domain Top-k Sampling via Inverse Discrete Fourier Transform for Spatiotemporal Dynamical Systems Modeling — Modeling spatiotemporal dynamical systems governed by partial differential equations (PDEs) poses two major challenges: "it either requires expensive physics-based simulators that entail iterative numerical solving at high computational cost, or it depends on abundant training da
- Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations — Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations Vasundra Srinivasan, Independent Researcher, Stanford School of Engineering. arXiv:2608.11323v1 [cs.AI], 11 Aug 2026.
- Agent Safety Should Be a Runtime Contract — Position: The paper argues that agent safety should not be treated as a property instilled during model training (via RLHF, DPO, Constitutional AI, etc.), but rather as a runtime contract enforced by the harness—the non-model infrastructure connecting a foundation model to the
- Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation — Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself. [episode]
- Most biomedical publications show signs of LLM-assisted writing — Based on the paper "Most Biomedical Publications Show Signs of LLM-Assisted Writing" by Holzwarth, González-Márquez, and Kobak: Summary This paper presents a new method to estimate the prevalence of LLM-assisted writing in biomedical publications and applies it to a large corpu
- Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts — This paper introduces MOSAIC (Model Optimization via Systems-Aware TraIning Co-design), a framework that jointly optimizes model architecture, training-token budget, and distributed parallel execution layout for a fixed cluster and training window, specifically instantiated for s
- Hypothesis Frontier: Verifier Guided LLM and Symbolic Search for First-Order Induction — Hypothesis Frontier: Verifier-Guided LLM–Symbolic Search for First-Order Induction First-order concept synthesis asks a system to infer one formula that classifies labeled objects consistently across several finite relational structures.
- Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence — Ex-Omni-2D is an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video.
- Diffract: Spectral View of LLM Domain Adaptation — The paper "Diffract: Spectral View of LLM Domain Adaptation" studies continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text.
- Generative Learning for Quantum Measurement Design — Authors: Jun Dai, Olivier Nahman-Lévesque, Guillaume Rabusseau, Hong-Ye Hu, and Cunlu Zhou Affiliations: Mila – Québec AI Institute, Université de Montréal, Institut quantique, Université de Sherbrooke, Harvard University, CIFAR AI Chair arXiv: 2608.11396v1 [quant-ph], 11
- Adaptation of Generalist Robot Policies with Minimal Data — The paper introduces minimal-data adaptation (MDA), a regime in which "a pre-trained robot policy must learn a new task from as little as one demonstration followed by autonomous online interaction." The authors note that "fully autonomous learning remains difficult with current
- Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives — This survey presents a unified review of cross-view feature matching, a fundamental computer vision problem concerned with establishing reliable correspondences between images with large viewpoint variations.
- Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding — Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale.
- From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents — Persistent memory in language-model agents enables cross-session reuse of preferences, observations, and experience, but it also makes errors durable: "a poisoned, stale, or misattributed record can alter reasoning, tool use, answers, and subsequent memory writes." Existing defen
- Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases — This paper introduces an analytically exact framework for the controlled behavioral evaluation of Large Language Models (LLMs), shifting from large, unstructured benchmarks to fully crossed factorial experiments.
- EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection — This paper presents EVIL-Detect, a multi-signal ensemble framework with conflict-aware fusion for NLPCC 2026 Shared Task 6, which requires three-class classification of Chinese text into human-written text (HWT), LLM-generated text (LGT), and LLM-refined text (HLT).
- The Illusion of Cross-Lingual Safety in Low-Resource Languages — This paper investigates whether safety alignment in large language models (LLMs), which is primarily developed in English, successfully transfers to low-resource languages. The authors focus on four African languages: Twi, Hausa, Amharic, and Swahili.
- Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter — This paper argues that tokenizer vocabulary size in large language models (LLMs) is not a fixed constant but a deployment-regime-dependent infrastructure parameter that should be optimized based on serving conditions.
- SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features — SQuaT (Student-Aware Quantized Teacher Features) is a label-free Quantization-Aware Training (QAT) framework with Knowledge Distillation (KD) that addresses a fundamental limitation in prior work combining QAT with KD: "during distillation, the range mismatch between the teacher
- RadFusion: Towards Threshold-Controllable Radiology Report Generation — RadFusion: Towards Threshold-Controllable Radiology Report Generation Summary This paper introduces RadFusion, a framework for threshold-controllable radiology report generation.
- From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop — The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic un
- How to Verify Probabilistic Consistency of Predictive Models — This paper addresses the question of whether a probabilistic predictor's answers to many conditional-probability queries are self-consistent, and whether this consistency can be verified in polynomial time.
- A lower bound for stepsize-based acceleration of gradient descent — This paper establishes a new lower bound on the convergence rate of plain gradient descent (GD) when the stepsize schedule is predetermined and nonnegative, showing that such schedules alone cannot achieve the optimal O(T-2) rate for smooth convex optimization.
- What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research — This paper synthesizes current knowledge about Responsible AI (RAI) practices in industry through a literature review of 161 empirical studies published between 2019 and August 2025.
- Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition — The paper introduces the Whisper-Aware LLM, a framework designed to address the challenges of whispered speech recognition in ASR systems.
- HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging — HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging introduces a framework for task vector merging that addresses the limitations of existing methods, which typically require per-subset scalar tuning and are restricted to linear rescaling.
- Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration — Image restoration is fundamentally constrained by the tradeoff between distortion and perception: "minimizing pixel-wise error yields over-smoothed results, whereas optimizing for perceptual realism often introduces structural deviations." Recent approaches attempt to balance thi
- Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation — Mixture-of-experts (MoE) models have recently moved beyond routing a fixed number of complete experts. Shared-expert designs preserve reusable knowledge, fine-grained methods vary computation within experts, and dynamic routers adapt the number of active experts.
- Optimistic Rates for Multiclass PAC Learning — This paper resolves the open problem of optimistic rates for multiclass PAC learning, closing the gap between the known realizable rate and agnostic rate.
- Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models — DURA is a diffusion-based unrestricted robotic attack that generates visually natural adversarial patches for Vision-Language-Action (VLA) models.
- Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent — The paper introduces DataMaster, an agentic instruction data selection system that translates natural-language user requirements into task-specific selection strategies for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs).
- BREAD: Baseline-Referenced Explanations for Anomaly Diagnosis — BREAD: Baseline-Referenced Explanations for Anomaly Diagnosis proposes a scalable, model-agnostic explanation method for AI-based prospective anomaly detection.
- Partially Observable Learning for Multi-Platform Dispatch Optimization — This paper proposes POLO, a Partially Observable multi-agent reinforcement Learning framework for multi-platform dispatch Optimization in instant delivery systems.
- GitSkills: A Dataset of Agent Skills on GitHub — GitSkills is a dataset of 3,797,117 SKILL.md files collected from 282,200 public GitHub repositories owned by 195,841 accounts, gathered in July 2026. [episode]
- MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection — MD-ProTector is an input-only encoder detector for LLM-generated text detection that represents each class (human-written and machine-generated) with multiple trainable reference vectors, called prototypes, in the encoder embedding space.
- Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection — The paper proposes CALIBDCD, a calibration framework for feature-based data contamination detection (DCD) in large language models (LLMs).
- Bayesian Symbolic Regression with Entropic Reinforcement Learning — Bayesian Symbolic Regression with Entropic Reinforcement Learning Oussama Boussif, Mohammed Mahfoud, Xiaoyin Chen, Younesse Kaddar, Moksh Jain, Yoshua Bengio, Sida Li, Damiano Fornasiere, Esmeralda S.
- Physics-informed Diffusion Generative Model for Time-Series Data Synthesis in Dynamic Systems — PhysDGM is a stepwise physics-embedded diffusion generative model for synthesizing time-series data that are consistent with the underlying physical laws of dynamical systems.
- Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes — The paper introduces AD2-Bench, a large-scale benchmark for evaluating Multimodal Large Language Models (MLLMs) in visually adverse and complex urban scenes, and EGVOR (Evidence-Grounded Visual Reasoning), a method to improve trustworthy reasoning in such conditions.
- Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs — This paper introduces the first any-to-any backdoor attack on Vision-Language Models (VLMs), moving beyond the static one-to-one and N-to-N attack paradigm.
- Simplex Relaxation for Discrete Diffusion — The paper introduces Simplax, "an exact Dirichlet–categorical augmentation of uniform discrete diffusion" that "preserves the original uniform diffusion process as its categorical marginal" while introducing "an auxiliary simplex-valued variable" to enrich the training objectiv
- GARLIC: Graph Attention-based Relational Learning of Multivariate Time Series in Intensive Care — The paper introduces GARLIC (Graph Attention-based Relational Learning for Intensive Care), a novel neural network architecture designed to address challenges in analyzing Intensive Care Unit (ICU) data.
- Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue — The paper introduces a dual-loop self-evolution framework for multi-turn empathetic dialogue, driven by verifiable emotion feedback.
- Cross-Corpus Evaluation of Generalizable Vulnerability Detection in IoT Firmware — This paper introduces IoTVulBench, a human-verified benchmark for cross-corpus IoT firmware vulnerability detection, and uses it to systematically evaluate how training-data source, model architecture, tuning method, and curriculum design affect detection accuracy, efficiency, an
- TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation — TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation Summary This paper introduces TimeRoute, a diffusion-based multi-modal recommender system designed to address the problem of "modality time-scale mismatch," where the relevance of different item m
- MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models — MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models Abstract Medical Vision-Language Models (MedVLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging.
- DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation — DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve.
- Multi-Granular Rationale-Guided Molecular LLM for Property Prediction — MR-MoL is a multi-granular rationale-guided molecular LLM for property prediction. It is, to the authors' knowledge, the first method to feed GNN-derived attributions into an LLM's prompt as evidence for property prediction.
- Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data — Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data This paper introduces Workflow Cards, a structured documentation artifact designed to summarize workflow execution provenance data in a format that is readable by both humans and large language mode
- Hierarchical Compositionality for An Assistive AI Agent — This paper presents an architecture for personalized command disambiguation in household environments that combines ASP-based feasibility filtering, compositional concept representation, multi-signal scoring over user-specific interaction history, and uncertainty-aware clarificat
- Mapping and Measuring the Behavioral Evolution of Large Language Models — The paper "Mapping and Measuring the Behavioral Evolution of Large Language Models" by Dong Qiao, Chris Ding, and Jicong Fan presents a label-free framework for characterizing and comparing the output behavior of 32 large language models (LLMs) from six families (GPT, Claude, Gem
- SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning — SPEED TUNING: Speeding Up Policy Execution with Lightweight Reinforcement Learning Abstract Summary: The paper introduces SPEED TUNING, a reinforcement learning framework designed to enhance the speed of manipulation policies.
- Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome — Authors: Fabrizio Russo (Imperial College London) and Mark Somers (Fifty One Degrees Ltd) arXiv: 2608.10664v1 [cs.AI], 11 Aug 2026 --- The paper addresses a foundational gap in the Relativity of Causal Knowledge (RCK) framework.
- Tensor-normal maximum likelihood estimation at the operator-norm sample threshold — Let X 1,, X n be independent Gaussian tensors in R d 1 R d k whose covariance is a Kronecker product of k unknown positive-definite factors, and put D = product a=1 k d a and d = a d a. A recent result of Franks et al.
- The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark — SRE-Bench is the first realistic, contamination-free reverse engineering (RE) benchmark for evaluating AI agents on binary analysis.
- Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning — The paper introduces J-Access, an inference-time auditing method that uses the Jacobian lens to map intermediate representations into vocabulary space and measure how often target concepts remain accessible along a model's output pathway.
- Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR — This paper presents a rigorous, multi-seed evaluation of Automatic Speech Recognition (ASR) for Garhwali, an under-resourced Indo-Aryan language of the central Himalaya, spoken by roughly 2.5 million people.
- Entropy-based Code Adversarial Translation for Real-world Repository Migration — This paper introduces Entropy-based Code Adversarial Translation (ECAT), a multi-agent framework for automated Android-to-HarmonyOS repository migration.
- Uncertainty-Aware Deep Learning for Genomics Applications: Insights from an Empirical Study — This paper presents an empirical analysis of uncertainty quantification (UQ) in deep learning models, focusing on genomics applications.
- Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics — This paper presents SAMPLED-BPE, a lightweight token-level auditing pipeline for web-scale Chinese corpora, motivated by observations of Chinese web pollution surfacing in LLMs—including "spam- and pornography-related Chinese tokens in ChatGPT's vocabulary" and "Chinese gamblin
- MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows — MAP-Graph is a provenance-aware memory layer for multi-agent workflows that addresses the problem where shared memory helps language-model agents reuse information, yet relevant evidence may not be admissible for a particular agent or action because restrictions propagate through
- When Is a General Factor Distinguishable? Non-Proportionality, Stable Structure, and the Bifactor Decision — Author: Jinsong Chen, Faculty of Education, The University of Hong Kong arXiv:2608.10731v1 [stat.ME] 11 Aug 2026 --- The paper addresses whether an additional general dimension is necessary beyond correlated first-order factors in factor analysis models.
- ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls — ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls Summary This paper introduces ConVAWG, a retrieval-grounded framework for generating controlled synthetic chat dialogues in the domain of Violence Against Wome
- CosMAP: Contrastive Manifold Approximation and Projection for Dimensionality Reduction of Omics and Genealogical Data — CosMAP: Contrastive Manifold Approximation and Projection for Dimensionality Reduction of Omics and Genealogical Data Summary This paper introduces CosMAP (Contrastive Manifold Approximation and Projection), a novel graph-based unsupervised dimensionality-reduction method designe
- Towards Unified Dynamic Face Landmark Detection — This paper introduces Unified Dynamic Face Landmark Detection, a novel framework that addresses two major functional limitations in existing face landmark detection (FLD) methods: (1) different network parameters need to be trained independently for each "N-point" benchmark datas
- Market-Information-Aware Gated-LoRA of Foundation Models for Transferable Day-Ahead Electricity Price Forecasting — This paper proposes a market-information-aware adaptation framework that transfers the Chronos-2 time-series foundation model to day-ahead electricity price forecasting.
- Mitigating Context Interference for Reliable and Efficient Search Agents — This paper investigates the issue of context interference in multi-turn search agents powered by Large Language Models (LLMs).
- Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport — Weightless Fine-Tuning (WFT) is a training-free, decoding-time method that approximates the distributional effect of supervised fine-tuning (SFT) for LLM personalization without updating model weights.
- Improving TensorSketch Using Complex Random Variables — Improving TensorSketch Using Complex Random Variables Summary This paper introduces a novel variant of the TensorSketch algorithm, termed Complex-to-Real (CtR) TensorSketch, which leverages complex random variables to achieve improved variance bounds for polynomial kernel approxi
- Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models — This paper demonstrates that large language models (LLMs) can serve as interpretable validators of metal-organic framework (MOF) structures when crystallographic information is transformed into chemically meaningful text.
- Uncertainty-Aware Compositional Localization and Placement Assessment of Catheters and Tubes in Chest X-Rays — Authors: Harshil Lodhiya (Sliced Health) Paper: arXiv:2608.11288v1 [eess.IV], 11 Aug 2026 --- The paper addresses the safety-critical yet tedious and error-prone task of assessing catheter and tube placement on chest X-rays.
- Backdoor Decontamination Dynamics in LLM Agents — Authors: Gabriel Huang, Abhay Puri, Léo Boisvert, Alexandre Drouin, Perouz Taslakian, Spandana Gella, Christopher Pal (ServiceNow Research, Mila – Quebec AI Institute, Polytechnique Montréal, Université Laval, McGill University, Canada CIFAR AI Chair) arXiv:2608.11295v1 [cs.
- Terminal Symmetry as a Carrier of Asymmetric Process Knowledge: Statewise Refinement for Anytime Verified Construction — This paper develops a decision-resource view of terminal symmetry for directed sequential construction tasks.
- MoE Proxy Models for Low-Cost Failure Reproduction and Diagnosis in LLM RL Post-Training — This paper proposes a multi-view, frequency-aware expert pruning method for constructing low-cost proxy models of Mixture-of-Experts (MoE) large language models (LLMs) to enable low-cost failure reproduction and diagnosis during reinforcement learning (RL) post-training.
- Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost — SpeedRunner is a method for programmatic skill learning that reduces agent cost. The paper argues that representing skills as executable code is more cost-effective than natural language, as it offloads reasoning to a cheaper Turing machine.
- Self-evolving network verifiers — Authors: Ioannis Protogeros, Tibor Schneider, Laurent Vanbever (ETH Zürich) This paper proposes a vision and provides early evidence for automatically evolving symbolic network verifiers so that their control-plane models faithfully capture actual network behavior, without requi
- Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence — " Summary This paper introduces Apodex Discovery, a comprehensive framework for building and evaluating "discoverative AI"—AI systems directed at making genuine discoveries rather than solving predefined problems.
- Rethinking Text-Based Image Retrieval in Specific Domain — The paper introduces SecMM-TBIR, a multi-match benchmark for surveillance scenarios, and proposes the Semantic-Aware Fine-Tuning (SAFT) framework to address performance degradation in domain-specific text-based image retrieval (TBIR).
- TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs — TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs Abstract Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists.
- Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training — The paper "Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training" proposes a novel framework called AID (Adaptive Importance-guided Discretized reconstruction) to improve multimodal representation learning for medical image-tabular data.
- Association-based Privacy Attacks in Wireless Protocols: Formal Modeling and Mitigation — This paper formally investigates the root causes of pairing-based privacy threats, specifically "Association Inference (AInf)" attacks, that are exploited using replay and relay techniques in wireless communication protocols.
- Conversational Orchestration for Organic 6G — The paper proposes a lightweight, decentralized conversational orchestration framework for service provisioning in Organic 6G networks, based on Large Language Model (LLM)-driven domain agents.
- Principal Trait Analysis: Towards Deriving "Skills" in Human-AI Collaboration — Principal Trait Analysis (PTA) is a novel, data-driven algorithm inspired by Principal Component Analysis (PCA) that automatically derives interpretable behavioral "traits" from large corpora of human-AI collaborative coding conversations.
- The GenAI Catch-22: Use of Generative Artificial Intelligence in Norwegian Newsrooms During the 2025 Parliamentary Election — The GenAI Catch-22: Use of Generative Artificial Intelligence in Norwegian Newsrooms During the 2025 Parliamentary Election Mari Reisjå IT University of Copenhagen marre@itu.dk Anders Sundnes Løvlie IT University of Copenhagen asun@itu.dk Abstract The increasing use of Generati
- Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task — This paper introduces a novel verbal-only conflict task, the "crayon task," to study congruency effects in large language models (LLMs).
- Conditional Independence Tests for Constraint-Based Causal Discovery: A Survey — This survey reviews conditional independence (CI) testing with emphasis on assumptions, robustness, and scalability in high-dimensional and mixed-type settings common in biomedical domains.
- Rationale-Guided Learning for Multimodal Emotion Recognition — RATIONALE-GUIDED LEARNING FOR MULTIMODAL EMOTION RECOGNITION The paper proposes Rationale-Guided Learning (RGL), a novel framework for multimodal emotion recognition in conversation (MERC).
- Scheduling Mixed RL Rollouts Beyond Prefix Locality — MISA-T is a routing-layer admission policy for mixed rollout serving in reinforcement learning (RL) post-training pipelines for large language models (LLMs).
- Leveraging Human Reading Behavior for Keyphrase Extraction: A Webcam-based Eye-tracking Corpus — Summary This paper investigates whether lightweight webcam-based eye-tracking features can enhance keyphrase extraction (KPE) from Chinese academic abstracts in Library and Information Science (LIS).
- On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models — This systematic literature review, conducted under PRISMA 2020 guidelines, examines security vulnerabilities in agentic Large Language Models (LLMs). The authors screened 743 records across six databases and retained 85 papers published between 2023 and 2025.
- Dueling Deep Q-Learning for Intrusion Detection — This study proposes a novel approach to intrusion detection systems (IDS) by employing a reward-based, dueling Q-learning model, achieving an average accuracy of 99.68% across multiple attack classes.
- myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR — This paper presents myMediWhisper, a Burmese medical speech recognition framework built on a high-quality 28-hour corpus recorded and validated by native speakers.
- Entropy-Centric Explainable AI for Remote Sensing Image Segmentation — This paper proposes an entropy-centric explainable AI (XAI) method for semantic segmentation in remote sensing imagery, addressing the lack of transparency in deep neural network decision-making.
- Contextual Information Policy Optimization for Search Agents — The paper, authored by Xingyu Guo, Wei Chen, Linlin Yang, and Baochang Zhang (Beihang University and Communication University of China), proposes Contextual Information Policy Optimization (CIPO), an evidence-oriented reinforcement learning framework for training search agents th
- Dual-Primal Graph VAEs for Noisy Label Aggregation — Dual-Primal Graph VAEs for Noisy Label Aggregation proposes a graph VAE architecture for inferring ground-truth labels from noisy crowdsourced annotations.
- ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization — ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals.
- V-FiLLM: Verified Financial LLM Reasoning Benchmark — V-FiLLM is a framework that generates financial reasoning benchmarks from executable computation trees grounded in real tables, yielding items whose answers are correct by construction.
- 3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment — This paper presents the 3D Sheep Pain Facial Expression System (3D-SPFES), a novel, monocular depth-aware geometric graph neural network system that integrates each SPFES facial landmark, such as the ears, eyes, and nose, into 3D Euclidean space estimated from a single RGB camera
- Efficient Hypergradient Descent for Inverse Reinforcement Learning — This paper addresses the computational challenges of inverse reinforcement learning (IRL) formulated as a bilevel optimization problem.
- A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa — This study presents a comparative evaluation of six object detection models—YOLOv5, YOLOv8, YOLO11, YOLO26, Faster R-CNN, and RT-DETR—using a real-world dataset, AgriAISeg, collected manually from Nigerian farms.
- Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training — Large-scale neural recommender systems are typically trained with a softmax cross-entropy objective over the full item vocabulary.
- RTSKG: Building a Rail Transit Station Knowledge Graph Dataset — RTSKG is a new rail transit station knowledge graph dataset that explicitly models the spatial and semantic interactions among different kinds of urban entities to support city-level rail transit station related tasks.
- Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting — Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting proposes TORF, a framework that decouples mean forecasting from uncertainty estimation to address the trade-off between distributional flexibility and accurate mean prediction in probabilistic
- AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations — AlbumentationsX is a data augmentation library that stores the transform list, probabilities, annotation settings, and random seed in one Compose object.
- sLTN: Structural Logic Tensor Networks — sLTN: Structural Logic Tensor Networks introduces an extension of Logic Tensor Networks (LTN) designed to handle structured data, such as temporal sequences, graphs, or other positional organizations.
- Attention-Path Fragility as an Uncertainty Signal in Large Language Models — The paper proposes that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is fragile under perturbation of its attention pathways.
- Goodness-of-Fit Tests and Calibration Machine-Learning Algorithms for Logistic Regression with Sparse Data — Based on the paper "Goodness-of-Fit Tests and Calibration Machine-Learning Algorithms for Logistic Regression with Sparse Data," here is a detailed summary.
- When and Where Faults Matter: A Study of Transient Errors in CKKS Multiplication — This paper presents an in-depth analysis of the resilience of homomorphic multiplication in the CKKS (Cheon–Kim–Kim–Song) fully homomorphic encryption (FHE) scheme in the presence of hardware-induced transient single-bit flip errors.
- DACRI: Decision-Aware Causal Intervention Ranking for Critical Supply Chains — DACRI: Decision-Aware Causal Intervention Ranking for Critical Supply Chains Abstract: Detecting or attributing a supply-chain disruption is not the same as selecting the intervention that maximizes recoverable net value.
- On the Sensitivity to Errors in Homomorphic Computing: Single Transient Bit-flip Client-side Error Characterization — This paper analyzes the sensitivity of Homomorphic Encryption (HE) to bit-level faults, focusing on the CKKS (Cheon–Kim–Kim–Song) scheme, which is widely used for approximate arithmetic in AI and machine learning workloads.
- Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension — The paper proposes hierarchical empirical-Bayes Naive Bayes (HEB-NB), a smoothed categorical NB classifier in which "each class-feature conditional probability is smoothed by a Dirichlet prior whose concentration is learned data-adaptively via Type-II maximum likelihood, enabling
- MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment — Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions.
- A Quantum Roadmap for Softmax Attention: Exact Born-Rule Analogs for Softmax Attention on the Probability Simplex — The paper presents a theoretical construction establishing an exact equivalence between a classical single-head transformer attention layer with a residual connection, operating on the probability simplex, and a quantum circuit where every learnable parameter is a rotation-gate a
- Statistically-Secure Bit Commitment and Coin Flipping Protocols Based on Quantum Hardware Assumptions — This paper introduces the first statistically secure bit commitment and coin flipping protocols based on hybrid hardware assumptions, specifically using Hybrid Locked Physical Unclonable Functions (HLPUFs).
- Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation — This paper introduces a Test-Time Self-Evolving framework for GUI visual grounding that enables models to improve after deployment without human-annotated ground truth. The framework constructs a closed-loop of Exploration, Evaluation, Reflection, and Internalization.
- Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders — This paper investigates whether the interpretability of individual sparse autoencoder (SAE) latents extends to the set-level analysis of active latent signatures, testing the "bag-of-features" hypothesis for SAE representations.
- Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning — Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning Summary This paper introduces the Surgical World-Action Model (Surgical WAM), a unified generative model for surgical robot learning that jointly predicts future endoscopic observations and executable s
- Hardware-Aware Deployment of Joint SAR Compression and Despeckling on FPGA — This paper presents the first deployment of a joint SAR Despeckling and Data Compression (DDC) framework on an embedded FPGA-based platform, using a ZCU102 board with AMD's Vitis AI overlay accelerator.
- Knowledge-Graph-Guided Retrieval-Augmented LLMs for Explainable Root Cause Analysis in Automotive HiL Validation — This paper proposes a knowledge-graph-guided retrieval-augmented large language model (KG-guided RAG-LLM) framework for root cause analysis (RCA) and fault localization in automotive Hardware-in-the-Loop (HiL) validation data.
- Basin: Efficient and Extensible Numerical Optimization in Rust — Basin is a numerical optimization library for the Rust programming language (Matsakis & Klock, 2014).
- Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification — This paper proposes an uncertainty-aware and explainable deep learning framework for multi-class skin lesion classification.
- Physics-Informed Implicit Neural Representations for Improved Myocardial Perfusion MRI Quantification — This paper presents a physics-informed neural network (PINN) framework with spatiotemporal implicit neural representations (INRs) for quantitative myocardial perfusion MRI.
- SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation — SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation proposes a novel decision-based black-box sparse attack framework for semantic segmentation.
- Benchmarking Cyberattack Detection in Electric Vehicle Charging Infrastructure with Benign User Updates — This paper develops a leakage-controlled session-level benchmark for cyberattack detection in electric vehicle (EV) charging infrastructure that preserves the ordered inputs of real Adaptive Charging Network (ACN) sessions and models legitimate revisions as normal behavior.
- CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classification — CLEAR (Class-wise reLiability-aware Expert Aggregation for long-tailed Recognition) is a modular ensemble framework for long-tailed classification that addresses the challenge of uneven prediction reliability across frequent and underrepresented classes.
- Battlefield 5G: Dual-PKI and TPM-Based UE Attestation for Tactical 5G Standalone Networks — The paper presents Battlefield 5G, a pre-authentication framework for tactical 5G Standalone (SA) networks that addresses a gap in standard 5G-AKA: "The standardized 5G Authentication and Key Agreement (5G-AKA) authenticates a subscriber credential stored on a Universal Subscribe
- Clinical Feasibility of Low-Magnification Fluorescence Imaging for Breast Cancer Margin Detection Using Texture Analysis and Deep Learning — This study investigated how 4× and 10× magnifications affect margin-level classification of MUSE images using both TA and DL methods, employing a patch-level framework.
- Spectral graph clustering with inhomogeneous latent geometry — Authors: Konstantin Avrachenkov, Lucas S. Sibemberg, Alexander Van Werde arXiv:2608.11321v1 [cs.SI] 11 Aug 2026 --- The paper studies spectral clustering in the presence of a confounding latent geometry.
- Socioduality: A Relational Process Framework for Human-AI Interaction — Socioduality is defined as "a sequential, reciprocal, and history-carrying relational process between two distinguishable parties in which a response from one party becomes part of the observable conditions under which the other party's subsequent contribution, judgement, decisio
- Contextual Quality-Diversity Evolutionary Reinforcement Learning for HVAC Control in Tropical Commercial Buildings — This paper proposes CQD-ERL, a contextual quality-diversity evolutionary reinforcement-learning controller for the supervisory control of a tropical, water-cooled chiller plant and its associated air side.
- Long-Horizon Forecasting of Complete Financial Statements with Forma — The paper introduces ProForma-20Q, a reproducible benchmark for forecasting 78 statement line items 1–20 quarters ahead for anonymized firms, and Forma, a transformer-based architecture tailored to this task.
- Gloss-Free Representation Learning for Cross-Dataset Sign Spotting — Sign-language research for resource-constrained languages is often limited by the cost of dense linguistic labels, including glosses, temporal boundaries, and sign order.
- Evaluation Resolution Confounds Learning-Rule Comparisons in Model-Brain RSA of Early Visual Cortex — The paper investigates a methodological confound in representational similarity analysis (RSA) comparisons of learning rules against brain data at early visual cortex.
- Evo-Bench: Can Language Models Improve Agent Harness? — Evo-Bench is the first benchmark designed to evaluate large language models' intrinsic harness-evolving capability—their ability to autonomously optimize their own executable agent harnesses, rather than merely solving static tasks.
- Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution — Emotion2Skill is a framework that extracts LLM-internal emotion vectors and incorporates them into both skill selection and skill evolution for skill-based LLM agents.
- ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models — ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models Abstract Large language models are increasingly deployed in education as tutors, teaching assistants, content generators, and learning advisors.
- Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness — Decoding-Level Taboo is a zero-prompt diagnostic stress test that intervenes directly in logit space at runtime to force large language models (LLMs) off their nominal decoding paths.
- Who Gets Heeded? An Obligation-Level Audit of Responsiveness in EPA Rulemaking — This paper introduces obligation-level responsiveness auditing, an auditable, AI-assisted framework for measuring whether public-comment engagement co-occurs with changes to specific regulatory duties in U.S. notice-and-comment rulemaking.
- Topological Feasibility Guarantees for Differentiable Predictive Control — This paper establishes deterministic feasibility guarantees for differentiable predictive control (DPC) using a novel topological analysis of the induced reachable safe set, without requiring online safety filters.
- MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale — MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale MERA is a verifier-backed multi-cycle protocol for improving small language models in agentic systems, treating a single model invocation as the unit of adaptation.
- Narrative Keyframing for Generative Creative Writing — The paper introduces narrative keyframing, a new interaction technique for AI-assisted creative writing that lets writers specify different types of narrative constraints at selected moments in a story, then use AI to generate intervening prose.
- Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement — This paper introduces expert-guided g-computation (egg-computation), a novel causal inference framework for estimating the average time saved by candidate hospital Quality Improvement (QI) interventions on length of stay (LOS).
- Beyond Detection Accuracy: Measuring Explanation Cost, Stability, and Utility for Resource-Aware IoT Intrusion Detection — This study jointly evaluates predictive effectiveness, explanation cost, local explanation stability, and selective explanation for binary Internet of Things (IoT) intrusion detection.
- Accelerated Learning of High Dimensional Functions with a Tensor-Featured Training Network — This work presents a method to accelerate the optimization of learning high dimensional functions using deep neural networks (DNNs) by introducing contextual features into the first layer of a DNN.
- Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks — S INK F LEX-RL is a modular training system for reinforcement learning (RL) in dual-control tool-use environments, designed to address the memory challenges of long-horizon agentic tasks.
- VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation — The paper introduces VoxSumm, a multilingual corpus and benchmark for joint speech summarization and translation (JSumT).
- MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model — MazzikaAI is a knowledge-based system for real-time Arabic maqam accompaniment that uses natural language as the actuator of a real-time control loop.
- MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices — MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices Summary This paper presents MemSpec, a prediction-guided, memory-aware runtime system designed to improve adaptive speculative decoding for large language models (LLMs) on memory-c
- DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? — The paper introduces DSAgentBench, described as "the first benchmark to evaluate whether agents can automate full data-science workflows inside real computer environments." The authors argue that while real-world data science "involves long-horizon workflows that span data wrangl
- Invertible Logits Transformation for Accuracy-Preserving Post-Hoc Uncertainty Calibration — The paper introduces Invertible Logits Transformation (InvLT), a post-hoc calibration method that applies a learned scalar MLP f: R to R element-wise to pre-softmax logits.
- Fisher8: Stabilizing Neural Heteroscedastic Regression via Output-Layer Fisher Geometry — Fisher8: Stabilizing Neural Heteroscedastic Regression via Output-Layer Fisher Geometry Summary This paper addresses the instability of training neural networks to jointly predict mean and uncertainty estimates from noisy observations using the Gaussian negative log-likelihood (N
- Beyond Forecasting: Recasting Volatility Control as a Routing Problem — The paper proposes VolRouter, a modular framework that reformulates volatility control as "persistent, state-conditioned routing over a library of estimator–controller policies." The authors identify "policy selection as an explicit layer between risk estimation and portfolio e
- Generator-Guided Inverse Sampling for L'evy-Driven Generative Models — This paper studies inverse sampling for Lévy-driven generative models from the perspective of Markov generators.
- Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation — The paper studies persona conditioning as a diagnostic mechanism for exposing assessor sensitivity in LLM-based information retrieval (IR) evaluation.
- Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving — The paper addresses the challenge of sample-efficient reinforcement learning for autonomous driving, which is "often limited by the trade-off between data efficiency and model bias." The authors note that "while world models reduce the reliance on costly environment interactions,
- Efficient Weak-Entropy PINN for Solving Hyperbolic Conservation Laws — The paper introduces a novel physics-informed neural network framework called Weak-Entropy PINN (WEPINN) for solving hyperbolic conservation laws with discontinuous solutions.
- Improved cross-validated distances for multivariate pattern analysis — The paper "Improved cross-validated distances for multivariate pattern analysis" by Laurent Caplette and Sarah Lippé proposes improvements to existing cross-validated distance measures used in multivariate pattern analysis (MVPA), specifically for Euclidean and Pearson correlati
- ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation — ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation Abstract: Variational autoencoders (VAEs) generate samples from probabilistic latent representations but do not explicitly distinguish uncertainty in the latent location from variability aro
- Automatic Field-of-View Adjustment for a View-Expansive Microscope via LSTM-Based Gaze and Pipette Motion Interpretation — Intracytoplasmic sperm injection (ICSI) operators frequently adjust the field-of-view (FOV) during procedures, which interrupts workflow and increases procedure time.
- TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling — TIDE RL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling Abstract Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with growing contexts, and fin
- Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning — TPSP introduces a policy-aware scene encoder to capture the interaction between policy behaviors and surrounding environments, enabling scene perturbation aligned with the current policy.
- Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models — This paper introduces a novel denial-of-service (DoS) attack targeting end-to-end (E2E) audio large language models (ALLMs).
- Post-Calibration Reliability Reranking of Relevance Decisions via Label-wise Monotone Projection — Web search, product search, and question-answering retrieval systems often assign a relevance label and confidence score to each query-candidate pair.
- VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback? — VisEditBench is a benchmark for evaluating vision-language models (VLMs) on the task of editing visualization code from multimodal feedback, addressing a gap in existing benchmarks that primarily focus on generating visualizations from scratch.
- How Robust Are LLMs to Vietnamese Dialects? — This paper introduces VialectBench, the first end-to-end human-annotated parallel benchmark for evaluating the robustness of Large Language Models (LLMs) to Vietnamese dialect variation.
- Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry — Author: Liangchen Ge arXiv:2608.10416v1 [cs.DS] 11 Aug 2026 --- This paper presents "a theoretical foundation for inverse-distance attention, from its Euclidean prototype (Resolver) to its non-Euclidean realization (Riemann GeoResolver)." The work is explicitly a theoretical stud
- Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects — This paper addresses the problem of reasoning shortcuts in neurosymbolic systems, where a system produces correct predictions through unintended concepts.
- Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique — Authors: Sanidhya Vijayvargiya and Rahul Lokesh (Samsung Research America) Large Language Models (LLMs) deployed as AI agents frequently exhibit "user specification-grounding failures, executing hallucinated, undesired actions to force a resolution rather than expressing uncertai
- Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance — This paper proposes a Conversational XAI interface powered by Large Language Models (LLM) to address interpretability challenges in machine learning-based Intrusion Detection Systems (IDS) for Unmanned Aerial Vehicle (UAV) networks.
- Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning — Continuous Interaction Diffusion (CID) is a diffusion-native model–runtime architecture that integrates tool interaction into iterative denoising, addressing the mismatch between turn-based tool protocols and continuously revisable diffusion generation.
- FUSE: Frame-Unified Stress Estimation from Facial Video — FUSE (Frame-Unified Stress Estimation) is a facial-video stress detection framework that processes complete recordings as a single input without temporal windowing or external segmentation.
- Quantum Incremental Learning with Mixed State Prototypes — Quantum Incremental Learning with Mixed State Prototypes Abstract Incremental learning models are required to learn new classes sequentially without catastrophic forgetting, while operating under parameter and memory constraints.
- RLMOpt: Adaptive Prompt Optimization via Recursive Language Models — RLMOpt is a prompt optimizer that makes the search policy itself language-model-driven through a recursive language model (RLM).
- Stay or Stray - A Dynamical Systems Viewpoint of Popularity Bias — The paper "Stay or Stray – A Dynamical Systems Viewpoint of Popularity Bias" studies the emergence of popularity bias in recommendation systems through a dynamical systems lens.
- Evaluating Rational Contracting in Natural Language — The paper addresses the challenge of evaluating how well language-based AI agents can negotiate and execute natural language contracts in uncertain, multi-step environments.
- Predicting Space Groups of Double Perovskites by LLM with Dynamic Few-Shot Learning — DyRIS, an LLM-agent-based framework, predicts ranked space-group (SG) candidates for double perovskites (DPs) from a given composition.
- Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models — The paper introduces SALT, a Semantically ALigned action Tokenizer, to address the problem that action representations in vision-language-action models (VLAs) are typically optimized for reconstruction under L1/L2 losses in raw action space, where numerical proximity need not ref
- GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning — GeoForge is a training-free, self-evolving framework for Earth observation (EO) agents that transforms completed trajectories into a structured nonparametric execution state.
- Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation — The paper introduces EDPFRL-IM (Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation), a novel framework that addresses a critical gap in Personalized Federated Reinforcement Learning (PFRL).
- MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph — MEGA (Meta Evaluation-Grounded Adaptation) is presented as a self-evolving infrastructure for agent development that addresses three limitations of current approaches: optimization without knowledge accumulation, knowledge accumulation without compositional reasoning, and a lack
- CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening — CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening Summary This paper presents CARB, a characterization-guided framework for predicting CNN inference costs (energy, latency, and peak memory) and enabling efficient deployment scree
- SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning — SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Abstract Summary Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones.
- Link-adaptive digital twin for robust physical-layer modeling in hybrid-amplified ultra-wideband optical networks — Ultra-wideband optical networks represent a practical solution for expanding communication capacity and supporting emerging applications such as artificial intelligence data center interconnection and 6G-oriented intelligent networks.
- Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models — Authors: Yuhang Song, Bor-Jiun Lin, Jiaxu Liu, Te-Chuan Chiu, Anh Nguyen, and Chun-Yi Lee Affiliation: University of Liverpool, National Tsinghua University, Imperial College London, National Taiwan University arXiv: 2608.10525v1 [cs.CV] 11 Aug 2026 --- The paper addresses the fu
- Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits — The paper studies cooperative multi-agent bandits in continuous (Lipschitz) action spaces when the Lipschitz constant is unknown.
- Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry — The paper studies multi-agent multi-armed bandits (MAB) with heavy-tailed rewards under three information-asymmetry regimes.
- Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy — This paper investigates whether a Large Language Model (LLM) can replace the static routing policy in a load balancer (HAProxy) to automatically isolate faulty backend servers that are degraded (returning HTTP 500s) but not down.
- Measuring Semantic Abstractness of SAE Features via Nonlocality — The paper addresses a central challenge in Mechanistic Interpretability (MI): distinguishing surface-level, token-driven features from genuinely high-level, abstract features in Sparse Autoencoders (SAEs).
- Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization — This paper presents the Reinforcement Learning for Laser Cutting (RL2C) algorithm, a Q-learning-based method with an epsilon-greedy policy designed to optimize laser cutting parameters for optical films.
- Retrieval-Corrected Conformal Prediction for Time Series — Retrieval–Corrected Conformal Prediction (RCCP) is a retrieval-augmented calibration method for time series prediction intervals.
- Iterative Erasure Count Is Not an Affine-Invariant Concept Dimension — This paper argues that iterative erasure count is not an affine-invariant concept dimension.
- pi-SUB: A Physics-Informed Synthetic Underwater Benchmark Dataset for Underwater Image Enhancement — This paper presents π-SUB, a physics-informed framework for generating synthetic underwater benchmark datasets that bridges the synthetic-to-real gap for Underwater Image Enhancement (UIE).
- DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction — Proteolysis-targeting chimeras (PROTACs) induce protein degradation by recruiting a target protein to an E3 ubiquitin ligase, making degradation a joint outcome of the degrader molecule and its biological context.
- beta-VAEs as Effective Theories: Tolerance-Dependent Dimension — In a beta-VAE, increasing the regularization strength acts as a spectral cutoff by collapsing low-utility latent coordinates. In the linear Gaussian VAE, the collapse order matches the ranking of reconstruction utilities exactly, because both are set by the PCA spectrum.
- BooST: Bridging Semantics and Motions for Efficient Skill Transfer — BooST: Bridging Semantics and Motions for Efficient Skill Transfer introduces a two-stage framework for learning a unified skill representation that captures both high-level semantic intent (what) and low-level motion dynamics (how) to enable efficient skill transfer to real robo
- ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS — ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners.
- Trigger the Straggler: Load Hijack on Mixture-of-Experts LLMs — Summary This paper introduces Load Hijack, a supply-chain attack against Mixture-of-Experts (MoE) large language models (LLMs) that exploits expert parallelism (EP) serving architectures.
- Decomposition-Induced Context-Memory Conflict: When Fact-Checking Pipelines Contradict Their Own Source Text — Decomposition-Induced Context-Memory Conflict: When Fact-Checking Pipelines Contradict Their Own Source Text This paper introduces and characterizes a failure mode in decompose-then-verify pipelines (including FActScore-style fact-checkers, hallucination detectors, and long-form
- InSight-doc: Agentic Visual Perception for Long-Document Understanding — InSight-doc: Agentic Visual Perception for Long-Document Understanding proposes a novel framework to address the challenges of long-document understanding, which often requires reasoning over many visually rich pages, making inference costly and prone to context rot.
- IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning — This paper addresses confounding bias in model-based reinforcement learning (MBRL).
- Curate Before You Connect: Identity and Ontology Tagging in a Production Knowledge Graph — This paper describes the ingestion and ontology-tagging layer that turns a validated extraction stream into a knowledge graph of 537,157 entities and 2,198,567 relationships drawn from 98,795 government documents.
- Retrieval-Augmented Vision Foundation Models for Robust Leukemia Cell Classification across Multiple Microscopy Datasets — This work presents a robust framework for leukemia classification across multiple heterogeneous datasets using a two-stage pipeline with a pretrained vision foundation model. Stage 1 performs binary classification (leukemia vs.
- Cross-View Sequential Visual Localization with Spatio-Temporal Context Modeling for Autonomous Driving — This paper proposes a temporal-context-enhanced framework for cross-view sequential visual localization in autonomous driving.
- VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus — Summary This paper introduces VERDICT (VERification via Disagreement-Informed Coupled Thresholding), a training-free, domain-agnostic, step-wise verification approach for multimodal large language models (MLLMs).
- Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement — This paper presents a pre-registered four-week longitudinal study (N = 72, 182,451 lines of conversation) examining whether general-purpose AI systems actively foster relational engagement.
- Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory — ReTree is a self-correcting tree-structured memory mechanism for LLM-based search agents. It constructs a bounded per-step reasoning context while preserving source-linked evidence.
- Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora? — The paper "Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?" investigates whether released tokenizer vocabularies can be used to estimate the corpus ratios of individual tokens in hidden pretraining corpora.
- SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information — SPIEVAL is a human-curated benchmark introduced to evaluate large language models (LLMs) as mobile assistants in scenarios where personal information is scattered across multiple applications.
- Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control — This paper introduces a situated behavioral-data (B-data) framework for studying and controlling LLM behavioral personality, moving beyond traditional questionnaire-based self-report methods.
- DuplexWorld: Can voice agents help you get through the day? — DUPLEX WORLD introduces a benchmark for holistically evaluating speech-to-speech (S2S) voice agents across six worlds—banking, insurance, travel, healthcare, logistics, and Pathfinding—spanning 156 authored scenarios and eleven conversation types, with 3,825 scored conversati
- Optimal Stopping of Self-Refining Foundation Models — The paper "Optimal Stopping of Self-Refining Foundation Models" by Kim Hammar, Tansu Alpcan, and Emil C. Lupu addresses the problem of deciding when to stop the self-refinement process of foundation models.
- Smart Enough to Go Extinct? An Evolutionary Challenge to the Value of General Intelligence and Its Ethical Implications for AGI — The paper "Smart Enough to Go Extinct? An Evolutionary Challenge to the Value of General Intelligence and Its Ethical Implications for AGI" by David Klotz subjects the widely held presupposition that general intelligence is extraordinarily valuable to critical scrutiny.
- A Lightweight Fault-Detection Scheme for Barrett Modular Multiplication Using Multiple Conditional Reduction Paths — This paper proposes a lightweight fault-detection scheme for Barrett Modular Multiplication (BMM) using a Statistical Reduction Monitoring (SRM) method.
- Long-Time Trajectory Approximation via SA-NODEs: Model Predictive and Floquet Strategies — This paper studies the approximation of dynamical systems by semi-autonomous neural ordinary differential equations (SA-NODEs) over long time horizons.
- Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution — The paper addresses a fundamental limitation in automated research ideation systems.
- A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona Problem — The paper reports on a production deployment of a centralized MCP (Model Context Protocol) gateway architecture that resolves a governance crisis caused by rapid, uncoordinated adoption of MCP servers within large enterprises.
- Compositional Benchmark Synthesis for Hierarchical Human Action Recognition — Authors: Farnaz Soleimani, Abdelghani Chibani, Yacine Amirat, Ghazaleh Khodabandelou (LISSI Laboratory, University of Paris-Est Créteil) Core contribution: The paper proposes "a benchmark-generation and evaluation framework that synthesizes a four-level hierarchical-intention be
- Path Integral Value Matching for Linear Quadratic Stochastic Optimal Control — The paper addresses the computational challenges in solving Linear Quadratic Stochastic Optimal Control (LQ-SOC) problems.
- EvoMem: Memory-Augmented Evolution for Code Optimization — EvoMem is a persistent memory architecture for LLM-based evolutionary program search that captures and reuses candidate mutation knowledge across runs and tasks.
- BPG: Balancing Plasticity and Generalization for Domain Incremental Learning — BPG: Balancing Plasticity and Generalization for Domain Incremental Learning Summary This paper introduces BPG, a unified framework for Domain Incremental Learning (DIL) that addresses two key limitations of existing parameter-isolation methods: (1) the "plasticity challenge" ari
- Assessing Reliability of BERT-Based Models on Question Answering Tasks — This study evaluates the reliability of four BERT-based models—RoBERTa, BERT-Base, DistilBERT, and ALBERT—on question-answering (QA) tasks using two datasets: SQuAD 2.0 and QuAC.
- MIRA: Medical Image Reflection for Agentic Diagnosis — MIRA (Medical Image Reflection for Agentic Diagnosis) is a medical visual diagnostic framework that enhances Large Vision-Language Models (LVLMs) with active evidence acquisition, tool-grounded verification, and reflection-driven self-correction.
- UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations — UniProbe is a lightweight, unified, learnable detector for token-level hallucination detection in Large Vision-Language Models (LVLMs).
- TACTICL: Task-Aware Compression of Tabular ICL Models — TACTICL: Task-Aware Compression of Tabular ICL Models Abstract Summary: The paper introduces TACTICL, an automated task-aware compression framework for tabular in-context learning (ICL) models.
- Spectral Embeddings of Degree- alpha Laplacians in Random Dot Product Graphs — This paper studies a continuum of degree-normalized spectral embeddings for network data, defined through the family of matrices D(−α)AD(−α) for α ∈ [0, 1], which includes the adjacency matrix (α = 0) and the symmetric Laplacian (α = 1/2) as special cases.
- FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data — FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data Viktoria Schuster, Sana Tonekaboni, Caroline Uhler Summary This paper introduces Fidelity-Guided Rank Optimization (FiGuRO), a novel framework for estimating the intrinsic dimension (ID) of uni- and multi-modal data.
- Can Bayesian Optimization Efficiently Find a Strong Single Expert in Neural Thickets? — The paper investigates whether Bayesian optimization (BO) can efficiently find a strong single expert model through gradient-free post-training of large language models (LLMs), under a modest evaluation budget.
- VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? — VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? Abstract Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments.
- X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction — X2-Turn presents a frame-synchronous turn state prediction method via delayed-stream modeling, extending a pretrained Voxtral Realtime streaming ASR model with a parallel turn state prediction head.
- Enhanced Filtering Algorithms for the Euclidean Traveling Salesperson Problem and its variants in Constraint Logic Programming — The paper "Enhanced Filtering Algorithms for the Euclidean Traveling Salesperson Problem and its variants in Constraint Logic Programming" by Alessandro Bertagnon and Marco Gavanelli addresses the Euclidean Traveling Salesperson Problem (ETSP) and its variants within the framewor
- Benchmarking Time Series Generation Methods for Privacy-Preserving Forecasting — This paper addresses the underexplored question of how well synthetic time series generation methods perform when their output is the sole training source for forecasting models, and how much privacy risk the released series carry.
- Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift — The paper introduces the Floor Certification Map, a theoretical framework for selective predictors that must certify both a selection-conditioned risk level (at most an α-fraction of returned answers wrong) and a hard coverage floor (answering at least a β-fraction of shifted t
- Self-Normalized Inference for Constant-Stepsize Temporal-Difference Learning under Markovian Sampling — This paper develops inferential methods for constant-stepsize temporal-difference (TD) learning under Markovian sampling, addressing two key challenges: serial dependence in the data and the stepsize-dependent stationary target.
- ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation — On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local confidence or teacher–student agreement to weight, filter, or truncate the sampled trajectory.
- FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation — The paper introduces FaithformBench, a benchmark for evaluating the faithfulness of mathematical chain-of-thought (CoT) autoformalisation (AF) systems—systems that translate natural-language reasoning steps into formal statements in proof assistants like Lean.
- IO Factory: Simulating AI-Enabled Influence Campaigns at Scale — IO Factory is an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes.
- ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling — ThinkRetrieve is a test-time scaling framework that augments the reasoning traces of Large Reasoning Models (LRMs) with dynamically retrieved solved examples at each reasoning step.
- FedCGR: Federated Cross-Domain Generative Recommendation — FedCGR: Federated Cross-Domain Generative Recommendation proposes a federated cross-domain recommendation (CDR) framework that formulates federated CDR "as generation over a stable semantic item language." The paper addresses the fundamental tension in federated CDR: "the behavio
- Threshold Structure of Optimal Policies in Restart POMDPs — We study a Restart POMDP (Partially Observable Markov Decision Process) on a general Borel state space, where the controller either lets the hidden state evolve unobserved or restarts the system and observes the new state.
- A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models — This paper presents a cost-efficient routing pipeline for multilingual short-text classification using small language models, evaluated on two benchmarks: a 15-language subset of SIB-200 for seven-way topic classification and a 15-locale subset of MASSIVE for intent classificatio
- StreamFlow: Dynamic Memory Flows for Streaming Video Understanding — StreamFlow introduces an efficient visual memory framework for streaming video understanding that enables dynamic, on-demand access to historical visual information.
- REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs — The paper presents R EAP (Relation-aware Elicitation And Parsing), a system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting.
- CARE: Confidence-Aware Reasoning for Reliable Medical VQA — CARE: Confidence-Aware Reasoning for Reliable Medical VQA proposes a framework to address confidence miscalibration in medical Multimodal Large Language Models (MLLMs) used for visual question answering (VQA).
- ReLTEx: Reliable LLM-based Taxonomy Expansion — ReLTEx: Reliable LLM-based Taxonomy Expansion Zeinab Ghamlouch, Mehwish Alam Télécom Paris, Institut Polytechnique de Paris, France Summary This paper presents ReLTEx, a framework for reliable LLM-based taxonomy expansion.
- MUSE: A Full-Text Cross-Domain Knowledge Base of Scientific Problems, Solutions, and Rationales — MUSE: A Full-Text Cross-Domain Knowledge Base of Scientific Problems, Solutions, and Rationales Abstract Scientific papers contain fine-grained records of problem solving: authors mention technical obstacles and methods that were used to address them, often along with reasoning o
- XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving — XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving Summary This paper introduces XCoT-VLA, a Vision-Language-Action (VLA) model for autonomous driving that replaces verbose natural-language Chain-of-Thought (CoT) reasoning with a compact, executable Chain-of
- What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model — This paper asks what iterated self-feeding probes of language models measure, and answers that they measure "the model and the probe together, in quantities that are not distinguishable by inspection." The construction is a ring of N token cells, each resampled in place from the
- ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering — ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering Summary This paper introduces ConRub-Med, a reinforcement learning (RL) approach for open-ended medical question answering that uses consensus-based rubrics to provide feedback dur
- On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation — The paper introduces LingT2I, a new benchmark designed to evaluate cross-lingual effects in text-to-image (T2I) generation.
- Information Bottleneck under Perfect Privacy — The paper studies the information bottleneck problem under a perfect privacy constraint, with a particular focus on the active-rate regime where the representation-rate constraint is binding.
- R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video — R4DSG introduces a relative 4D scene graph memory for long egocentric video, designed to support object-centric question answering under weak sensing conditions.
- Derivative Computation in PINNs: Automatic Differentiation, Finite Differences and Beyond — The paper systematically investigates finite-difference (FD) derivative computation in Physics-Informed Neural Networks (PINNs) as an alternative to automatic differentiation (AD).
- Data Attribution of Emergent Misalignment with Persona Features — This paper investigates emergent misalignment (EM) in large language models—the phenomenon where "fine-tuning a model on insecure code completions caused it to advocate for enslaving humans and to give malicious advice in domains entirely unrelated to coding." The authors addre
- Self-Knowledge Retrieval Augmented Generation Framework for Patent Matching — This paper proposes a self-knowledge retrieval-augmented generation (RAG) framework for patent matching.
- SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training — SCOUT is a unified runtime failure-localization framework for LLM pre-training, built on the design principle of "identify outliers through strict-majority consensus among equivalent replicas." It addresses the problem that "in LLM pre-training, synchronization propagates rank-lo
- TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification — TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification proposes a weakly supervised hierarchical text classification (HTC) framework enhanced by LLM-based data augmentation.
- When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs — Authors: Utkarsh Bahuguna (Scaler School of Technology) Venue: Accepted at the COLM 2026 Workshop on Efficient Reasoning Core Finding: On the full GPQA Diamond benchmark (198 graduate-level science questions), majority voting via self-consistency reduces per-problem accuracy on a
- Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits — Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Error (MSE) and Fitted Q-Evaluation (FQE) assess only behavioral imitation and cannot detect Toxic Mimicry, a failure mode in which
- A Study of Kernel Telemetry Options for Security-Oriented Provenance — This paper studies kernel telemetry options for building the capture layer of security-oriented provenance systems.
- Diffusion-Based Data-Driven Assortment Optimization — Diffusion-Based Data-Driven Assortment Optimization proposes D3AO, a model-agnostic framework for assortment optimization in the offline setting, based on guided discrete diffusion.
- Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology — Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology Summary This paper introduces Social Chain of Thought (SCoT), a multi-round pipeline for medical differential diagnosis that structures multi-agent interaction as a delibera
- Analysis of Federated Aggregation under Model Poisoning and Backdoor Attacks: A Reconstructed Cross-Dataset and Cross-Architecture Benchmark — The paper presents a reconstructed comparative benchmark and evidence audit of federated aggregation methods under model poisoning and backdoor attacks, rather than a new-algorithm superiority study.
- Click2Poly: A VLM for vector mapping buildings and walls — Click2Poly is a human-in-the-loop AI assistant designed to speed up the manual correction and editing of building and wall vector layers in geospatial mapping. It extends the Florence-2 Vision Language Model (VLM) and is implemented as a QGIS plugin.
- AutoGrable: What Is a Good Graph for a Table? — The paper addresses the fundamental question of graph construction for tabular and relational data: "when the graph is not given, what is a good graph to learn on, and how do we build it?" The authors note that "Graph learning presupposes the existence of a single or multiple gra
- Stigma and Support in Online Sexual Violence Narratives on Reddit — This paper introduces the SCOPE (Stigma and COmmunity Peer Expressions) dataset, which links stigma signals in online sexual violence survivor narratives on Reddit to the types of support offered in corresponding comment threads.
- Benchmarking LLM Judges for Mobile Agent Evaluation — M OBILE J UDGE B ENCH is introduced as "to our knowledge the first benchmark for evaluating LLM-as-judge methods on mobile agent trajectories." The benchmark comprises "931 human-annotated trajectories spanning 6 benchmarks, 4 agents, and 68 apps," collected from "6 established m
- Variational Parameter Calibration with Physics-Aware Latent-Space Surrogates — This paper introduces a physics-aware neural-network-based latent-space framework for reduced-order forward modeling and variational parameter estimation in parametric dynamical systems.
- DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition — DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition Abstract Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages.
- Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling — The paper presents a substantial expansion of the IMPROVE benchmark dataset for drug response prediction (DRP) models, integrating large-scale pharmacogenomic data primarily from PharmacoDB along with additional smaller data sources.
- Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards — The paper introduces a neuro-symbolic safety guard for end-to-end autonomous driving. The guard is a lightweight module that attaches to the final command interface of an already-trained agent.
- Gaussian Meta-Space Augmentation for Stacking Ensembles in Multimodal IPMN Risk Stratification — The paper introduces cUPMI (calibrated Upstream Probabilistic Meta-Imputation), a class-conditional Gaussian augmentation method for regularizing level-1 stacking ensembles in multimodal IPMN (intraductal papillary mucinous neoplasm) risk stratification.
- Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness — The paper establishes convergence guarantees for gradient descent applied to general feedforward neural networks of arbitrary width or depth, without special requirements on initialization or dataset.
- Forward Trajectory Steering for Hamilton-Jacobi Reachability Analysis — This paper introduces STEER2REACH (S2R), a physics-informed neural network (PINNs)-based solver for Hamilton-Jacobi (HJ) reachability analysis.
- Stability of Finite-Batch Particle Mean-Field Variational Inference Beyond Strong Convexity — This paper studies the stability of finite-batch particle mean-field variational inference (MFVI) beyond strong convexity.
- From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation — Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale.
- Language-Structured Relational Q-Learning for Threat-Aware Control in Safety-Critical Driving — The paper introduces Language-Structured Relational Q-Learning, instantiated through an Ego-Centric Relational Q-Network (ERQ-Net), for threat-aware control in safety-critical driving scenarios.
- Strengthening Full Justified Representation: Efficient Verification and Computation — This paper introduces FJR+, a strict strengthening of the full justified representation (FJR) and EJR+ axioms for approval-based committee elections, which can be both verified and satisfied in polynomial time. Key contributions: 1.
- Predictive Allostatic Organization in Recurrent and Spiking Agents Under Partial Observability — Adaptive behavior under partial observability depends on internal organization that carries information beyond the current observation.
- RelShap: Relationally Consistent Shapley Explanations — RelShap is a framework that incorporates relational constraints and data provenance into Shapley value computation, restricting both background data and coalition evaluation to relationally valid configurations.
- Let it Cook: Learning to Wait in Sequential Decision Making — The paper addresses the question of whether agents in sequential decision making need to actively participate at every timestep.
- Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation — This paper presents the first large-scale empirical study investigating whether psychologically inspired influence tactics, when embedded in prompts, affect LLM-generated code in software engineering tasks.
- Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval — This paper presents the first direct comparison of natively multimodal embedding models against frontier LLMs as zero-shot direct visual rankers for text-to-image retrieval.
- An Empirical Study of Output-to-Input Loops for Black-Box Backdoor Detection in Fine-Tuned Open-Weight LLMs — Authors: Md.
- Dynamics Models for Offline Hyperparameter Selection in Real-World RL — A key obstacle to deploying reinforcement learning in real-world systems is hyperparameter selection, particularly when simulators are unavailable and online experimentation is costly.
- Self-Evolving Embodied Agents via Skill-Harness Evolution — The paper introduces SHAPER, a self-evolving framework for train-free embodied adaptation that keeps model parameters frozen and improves the non-parametric agent system by evolving reusable skills and a context-code harness through target-environment rollouts. [episode]
- ODE-Based Transformer Decoders for Iterative Sign Language Translation — Sign language translation has achieved strong results with Transformer architectures, yet recent improvements largely rely on scaling model capacity at the cost of increased computation.
- RevCRN: Reversible Analog Computation using Chemical Reaction Networks — This paper introduces the Reversible Chemical Reaction Network (RevCRN) model and investigates the computability of real numbers using this framework.
- Gaze Target Estimation Anywhere with Concepts — This paper introduces the Promptable Gaze Target Estimation (PGE) task and the GazeAnywhere model, a new end-to-end, concept-driven paradigm for gaze analysis.
- Towards an approach to multivariate outlier detection for District Heating System data — This paper tests different methods for multivariate detection of outliers in the data of transmitted heat energy in the selected substation of local District Heating System, by also considering outside ambient temperature, namely Z-score (univariate, as a benchmark), Mahalanobis
- From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate — The paper "From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate" studies whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization.
- Mechanism Design for Generative Engines: From Exploitation toward Win-Win Outcomes — Authors: Chen Xu (Carnegie Mellon University), Zitian Guo (University of California, San Diego), Chenyan Xiong (Carnegie Mellon University) Paper: arXiv:2608.11390v1 [cs.LG], 11 Aug 2026 --- The paper addresses a strategic tension in the emerging "generative engine ecosystem," wh
- Federated Learning for Distributed CNC Tool Wear Prediction — This paper investigates federated learning for CNC tool wear prediction using the MATWI dataset, which contains sensor recordings and images from 17 cutting-tool sets.
- CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation — CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation This paper introduces CARD, a framework for simulating realistic credit card discussion threads. The paper states: "We present CARD, a framework for simulating realistic credit card discussions.
- Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces — The paper proposes an Inverse Theory of Mind (IToM) pipeline for content recommendation that reasons backward from observed user interactions to infer the beliefs, preferences, and decision-making traits that explain behavior, producing structured natural-language personas withou
- EliSeg: Verified Target Construction for Report-Grounded Abnormality Segmentation — The paper addresses report-grounded abnormality segmentation, where a model receives a chest radiograph and its associated radiology report, and must "identify the findings that warrant spatial localization, and generate a corresponding mask for each." Unlike target-conditioned s
- ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment — Authors: Abdulkadir Küçe, Alihan Esen, Çağla Fikir, Berke Kurt, Kuzey Arar, Gökhan Ercan, and Faik Boray Tek (Istanbul Technical University) arXiv ID: 2608.06110v1 [cs.AI], 6 Aug 2026 The paper presents ECHO (Enhanced Care & Health Observer), described as "a locally-deployab
- XGBoost "is all you need": the case of forecasting transmitted heat energy in District Heating Systems — This paper presents a comparative study of two distinct approaches, XGBoost and Long-Short Term Memory (LSTM), for forecasting transmitted heat energy in District Heating Systems (DHS).
- Blast Radius — "Blast Radius" by M.Y. Pitsane and Hope Mogale (Mankind Research Labs, Sandton; North-West University and University of Pretoria, RSA; arXiv:2608.07440v1 [cs.AI], 7 Aug 2026) addresses the growing problem of affordability and wasted tokens in agentic coding.
- Strategies to Avoid Illegal Data Access — This study examines technology solutions, personnel training, and policy enforcement as methods to prevent unauthorized data access. Data may be protected from illegal access using technological solutions like firewalls, intrusion detection systems, and encryption.
- Multiclass Sentiment Analysis for Identifying Political Viewpoints — The paper investigates multiclass sentiment analysis of political viewpoints on social media, specifically for Tamil language tweets, using two machine-learning approaches: XGBoost and BERT.
- Reoptimization Algorithms for Contextual Bandits with Knapsack Constraints — The paper "Reoptimization Algorithms for Contextual Bandits with Knapsack Constraints" by Zhen Xu (University of Liverpool) studies new algorithms for Contextual Bandits with Knapsack (CBwK) problems.
- Decision-Aware Approximation of Belief Functions for Evidential Combinatorial Optimization — The paper introduces a decision-aware approximation method for belief functions used in evidential combinatorial optimization, where the goal is not to keep the approximation close to the original mass function by an intrinsic distance (such as Jaccard or Jousselme), but to prese
- A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language — Summary This paper introduces a new dataset and baseline models for fine-grained isolated handshape recognition in sign language, grounded in the Hamburg Notation System (HamNoSys).
- Every pooling rule has its world: matching probability combination rules to situations and stakes — Authors: Tanel Tammet, Priit Järv, Dirk Draheim (Tallinn University of Technology) arXiv:2608.11275v1 [stat.ME] --- Systems often need to combine two numerical assessments of the same yes/no question.
- When Agents Talk: Honeytokens under Shared Memory — The paper "When Agents Talk: Honeytokens under Shared Memory" by Joshua S. Gans examines the feasibility of defensive deception—specifically honeytokens—when AI agents share memory and can pool information.
- A Runtime Decentralized Attestation and Coordinated Repair Framework for Securing Automotive ECUs — The paper introduces DACER, a runtime decentralized attestation and coordinated repair framework for securing automotive ECUs.
- REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems — REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems Abstract Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks.
- FITTER: Vocabulary-Agnostic Cross-Domain Inference on Temporal Knowledge Graphs — FITTER: Vocabulary-Agnostic Cross-Domain Inference on Temporal Knowledge Graphs Abstract Temporal knowledge graphs are central to many uses of the Semantic Web, but existing completion methods assume the entities, relation names, and timestamps to be reasoned about are already kn
- Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure — The paper "Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure" by Xingjian Wang, Qingyu Han, Xiaodong Luo, and Yin Zhang provides a two-stage analysis of rank collapse in Post-Norm decoder-only Transformers, using token similarity as a scalar
- SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure — SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Abstract Summary The paper addresses the problem of skill bloat in self-evolving agents.
- Modelling Geographic Atrophy Progression using Implicit Neural Representations — Age-related Macular Degeneration (AMD) is the major cause of blindness in the Western world. Its late dry phase is characterised by irreversible atrophic areas, namely Geographic Atrophy (GA).
- MVTrack: Ultrafast Appearance-Free Moving Object Tracking from Compressed Bitstreams — MVTrack is an ultrafast tracker for moving objects that operates directly on H.264 bitstreams, combining MVDet, a lightweight detector for motion vector fields, with MVLink, a minimalist kinematic association module.
- The Signal Rail: A Deterministic Motion Grammar for Communicating Conversational Agent State in Terminal Interfaces — The Signal Rail: A Deterministic Motion Grammar for Communicating Conversational Agent State in Terminal Interfaces Abstract Terminal interfaces to conversational agents report rich internal state (listening, thinking, executing tools, awaiting input, failing) almost entirely thr
- A Systematic Sample Size Analysis of ML-Based Path Loss Prediction for LPWAN — This paper investigates how training-set size affects the accuracy of machine learning (ML) models for LoRa path loss prediction in urban Low Power Wide Area Network (LPWAN) deployments.
- A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization — SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration) is an open-source framework that employs natural-language orchestration to guide chemical structure optimization.
- VIDS-Seg: Towards Reliable Uncertainty Quantification in Pediatric Cardiac Ultrasound Segmentation — VIDS-Seg: Towards Reliable Uncertainty Quantification in Pediatric Cardiac Ultrasound Segmentation Summary This paper introduces VIDS-Seg, an extension of the Variational Inference under Distribution Shifts (VIDS) framework, designed for dense image segmentation with a focus on r