AI papers — 2026-09-21
Today’s research landscape is defined by a push toward grounding and reliability. This moves the field beyond mere generation toward structural accuracy.
We see this in the development of Fact Grounded Attention, which seeks to eliminate hallucinations by integrating knowledge directly at the attention level. This effort is complemented by the auditing of GPTKB v1.5, which provides a multi-dimensional analysis of how frontier models elicit knowledge from knowledge bases.
This drive for precision extends to coding and logic. BoostAPR utilizes execution-grounded reinforcement learning with dual reward models to improve automated program repair, while Collab-Solver introduces a collaborative solving policy specifically for mixed-integer linear programming.
Even the internal mechanics are being mapped. Researchers are studying how models detect harmful content through internal representations and the specific activation subspaces that enable in-context learning of addition.
Progress is also being made in optimization. Offline reinforcement learning demonstrates an ability to learn effective scheduling even when starting from suboptimality.
The shift toward more nuanced evaluation and specialized modeling continues as researchers tackle the complexities of multimodal and temporal reasoning. The introduction of the CRYSTAL benchmark marks a move beyond simple final answers, aiming instead to evaluate transparent multimodal reasoning processes.
This focus on depth is mirrored in MemeLens, which addresses the cultural and linguistic hurdles of meme interpretation through multilingual, multitask vision-language models. In the realm of representation learning, new evidence suggests that semantic pairs play a critical role in shaping self-supervised outcomes.
These structural advancements extend to logistics and infrastructure. EAGLE utilizes edge-aware graph learning for proactive delivery delay predictions in smart networks, and NetGent introduces agent-based automation for network application workflows.
Even more specialized applications are emerging. Researchers are using reinforced graph-based physics-informed neural networks with dynamic weights to estimate battery health and remaining useful life. There are also calls for a dynamical systems perspective to truly advance time series modeling.
The shift toward agentic autonomy is being met with new frameworks designed to manage the inherent risks of unverified reasoning. To address the danger of models providing plausible but false rationales, researchers have introduced explanation-bound tool execution.
This method uses server-verified action claims to ensure an agent's behavior matches its stated intent without requiring trust in the model's internal logic. This move toward operationalizing agency is further formalized by the AI-GRACE framework, which maps organizational objectives and obligations directly onto deployment architectures.
While these systems manage high-level governance, specialized agents are pushing the boundaries of technical domains. The SpecOpt approach for molecule optimization utilizes contact-diff reasoning to improve binding specificity.
Even in highly sensitive human-centric fields, progress is visible through clinician-grounded quality assurance for psychiatric intake. This ensures that automated assistance remains tethered to professional standards.
The shift toward specialized architectures is evident in the development of attention-aware routing. This seeks to couple routing mechanisms directly with attention within Mixture-of-Experts models.
While this aims for greater efficiency, other researchers are focusing on the broader utility and safety of these systems. For instance, Self-Meta-Evolve addresses the limitations of static prompting by allowing models to evolve prompts for personalized information extraction.
This drive toward specialized application is further seen in PolyBridgeBench, a new benchmark designed to evaluate how well multimodal large language models handle physics-grounded bridge design tasks. However, as these models become more integrated into sensitive environments, security remains a primary concern.
This is reflected in the testing of CESBench for cryptographic engineering security in IoT devices. Additionally, the development of HE-Guardrail utilizes homomorphic encryption to defend against jailbreak attacks during encrypted inference.
The push toward more efficient and reliable intelligence is manifesting through both architectural refinement and rigorous evaluation. Researchers are looking at how to bridge the gap between high-level reasoning and low-level execution.
One approach is a fully differentiable neuro-soft-symbolic framework designed for perceptual task planning. Another uses implicit rule induction via test-time task embeddings to tackle ARC-like challenges.
Efficiency remains a primary driver. This is seen in the development of RBS-Attention, which utilizes radius-bounded sparse prefill to manage long contexts, and TinyCeNN-LM, which employs quality-gated conversion with CeNN-inspired cellular recurrent layers for pretrained attention.
As these models grow more complex, the need for better benchmarks becomes clear. CogGym is emerging as a way to conduct large-scale comparative evaluations of human versus machine cognition.
This evolution continues into specialized domains. Agents are being tested on their ability to design chips using higher-level abstractions, and GT-anchored verifier co-training is used to improve code generation reliability through information-gain rewards.
The focus shifts toward the internal mechanics of reasoning as researchers attempt to move beyond simple accuracy metrics. In an effort to audit how large language models arrive at their conclusions, the LogicTrack framework utilizes formal logic solvers to trace and examine reasoning trajectories.
This pursuit of transparency is mirrored in the study of model failures. Researchers have begun tracing the topological signatures of impaired context sharing to detect hallucinations.
While these methods attempt to pin down where logic breaks, others are looking at how models learn from their own mistakes through DENSE. This method distills agent trajectories into evidence-grounded shortcut trees designed for self-refinement.
These developments suggest a growing movement toward making the black box of model cognition more structured and verifiable through formal and evidentiary constraints. The day concludes with a look at the mechanics of model distillation, specifically whether a teacher model's influence stems from its accuracy or its specific behavioral patterns.
Researchers investigated this by separating correctness from behavior in self-distillation, categorizing teachers as either repulsive or attractive. This distinction helps clarify how a student model inherits knowledge during the distillation process.
Meanwhile, the security of retrieval-augmented generation systems remains a pressing concern. The introduction of micro-collaborative poisoning demonstrates how coordinated, small-scale inputs can compromise the integrity of retrieved information in RAG systems.
These developments suggest that as we refine the nuances of how models learn from one another, we must simultaneously harden them against increasingly sophisticated, distributed attempts to corrupt their reasoning pipelines.
Today's papers
- Nonnegative Matrix Factorization in the Component-Wise L1 Norm for Sparse Data. [paper]
- Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model. [paper]
- Robust Mixture Models for Algorithmic Fairness Under Latent Heterogeneity. [paper]
- Near-Optimal Primal-Dual Algorithm for Learning Linear Mixture CMDPs with Adversarial Rewards. [paper]
- Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30. [paper]
- LLM Safety From Within: Detecting Harmful Content with Internal Representations. [paper]
- ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction. [paper]
- Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency. [paper]
- Trajectory Entropy Reinforcement Learning for Robust Robot Motor Skill Learning. [paper]
- Assessment of Machine Learning-Based Critical Heat Flux Models in the CTF Subchannel Code for Square Rod Bundle Prediction. [paper]
- DEFEAT: Stitching Fragmented File I/O Contexts for Early Ransomware Detection. [paper]
- A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal. [paper]
- Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation. [paper]
- Draft-OPD: On-Policy Distillation for Speculative Draft Models. [paper]
- Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks. [paper]
- Hierarchical attention interpretation: an interpretable speech-level transformer for bi-modal depression detection. [paper]
- Beyond Accuracy: Centroid-Guided Contrastive Loss for Structured Fraudulent Job Posting Detection. [paper]
- One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction. [paper]
- StableAML: Machine Learning for Behavioral Wallet Detection in Stablecoin Anti-Money Laundering on Ethereum. [paper]
- Scaling Novel Graph Generation via Lightweight Structure-Guided Autoregressive Models. [paper]
- How do LLMs Compute Verbal Confidence. [paper]
- On the Limitations of Large Language Models for Conceptual Database Modeling. [paper]
- Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation. [paper]
- Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education. [paper]
- Enhancing Audio Reasoning via Semantic Summary Prediction. [paper]
- Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction. [paper]
- Joint Remaining Useful Life Prediction and Capacity Estimation of Lithium-Ion Batteries Using Partial-Charging Data. [paper]
- MIRCID: Inferred Hub-miRNAs Drive Cross-Task Improvements in Drug Mechanistic Modeling. [paper]
- Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges. [paper]
- Benchmarking World Models for Continual Learning on Compositional Tasks. [paper]
- Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake. [paper]
- Benchmark of stylistic variation in LLM-generated texts. [paper]
- FairLMs: A Turnkey Library for Fairness in Language Models. [paper]
- CodeMidas: Scaling Agentic Coding RL Environments from Code Itself. [paper]
- ASGARD: Action-Space Guard for UAV Resilience via Reinforcement Learning. [paper]
- AI-GRACE: A Use-Case Operationalization Framework for Agentic AI: From Organizational Objectives and Obligations to Deployment Capabilities and Architecture. [paper]
- Learning Cardiac Features: ECG Biometrics Across Time and Exercise. [paper]
- PlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking. [paper]
- Semantic Calibration Prevails Where Token Confidence Fails: Benchmarking Long-Form Scientific QA. [paper]
- Beyond Kinematics: Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning. [paper]
- Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions. [paper]
- Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding. [paper]
- Geometry-Aware Reinforcement Learning for 2D Irregular Nesting. [paper]
- Offline Constrained RLHF with Multiple Preference Oracles. [paper]
- Intervention Granularity Matters: Coherent Treatment Bundles in Counterfactual Simulation with Clinical World Models. [paper]
- lambda-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource. [paper]
- Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining. [paper]
- OverThink: Slowdown Attacks on Reasoning LLMs. [paper]
- Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design. [paper]
- Data-Driven Integration Kernels for Interpretable Nonlocal Operator Learning. [paper]
- Multi-Domain Clustering via Measure Quantization. [paper]
- Offline Multimodal Large Language Models for Decision Support in Air Operations. [paper]
- Position: A Dynamical Systems Perspective is Needed to Advance Time Series Modeling. [paper]
- Deep Reinforcement Learning with Buffered Quantile Objectives. [paper]
- Rhamba: Region-Aware Hybrid Attention-Mamba Framework for Self-Supervised Learning in Resting-State fMRI. [paper]
- Generalizing Beyond Suboptimality: Offline Reinforcement Learning Learns Effective Scheduling through Random Solutions. [paper]
- Near-Universal Multiplicative Updates for Nonnegative Einsum Factorization. [paper]
- Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention. [paper]
- Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER. [paper]
- What Must Survive? Exact Task-Information--State Frontiers for Resource-Sufficient Learning. [paper]
The papers
- Rollout Total Correlation for Deep Reinforcement Learning —
- Reinforcement Learning under External Influence: Guarantees, Algorithms, and Sample Complexity —
- Hierarchical attention interpretation: an interpretable speech-level transformer for bi-modal depression detection —
- Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks —
- OverThink: Slowdown Attacks on Reasoning LLMs —
- Cultural Alignment in Large Language Models Using Soft Prompt Tuning —
- Trajectory Entropy Reinforcement Learning for Robust Robot Motor Skill Learning —
- Understanding In-context Learning of Addition via Activation Subspaces —
- SRAF: Stealthy and Robust Adversarial Fingerprint for Copyright Verification of Large Language Models —
- VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits —
- Toward accurate RUL and SoH estimation using reinforced graph-based physics-informed neural networks enhanced with dynamic weights —
- Collab-Solver: Collaborative Solving Policy Learning for Mixed-Integer Linear Programming —
- NetGent: Agent-Based Automation of Network Application Workflows —
- Benchmark of stylistic variation in LLM-generated texts —
- Generalizing Beyond Suboptimality: Offline Reinforcement Learning Learns Effective Scheduling through Random Solutions —
- Robust Mixture Models for Algorithmic Fairness Under Latent Heterogeneity —
- Fidel-TS: A High-Fidelity Multimodal Benchmark for Time Series Forecasting —
- Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration —
- Transformers Discover Molecular Structure Without Graph Priors —
- Auditing a KB Elicitation of Frontier LLM Knowledge: A Multi-dimensional Analysis of GPTKB v1.5 —
- MIRANDA: short signatures from a leakage-free full-domain-hash scheme —
- The Impact of Semantic Pairs on Self-Supervised Representation Learning —
- Provably Optimal Reinforcement Learning under Safety Filtering —
- PRIVET: PRoximIty leakage detection Via Extreme value Theory —
- Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance —
- SteganoBackdoor: Evading Data-Poisoning Defenses via Steganographic Backdoors —
- A Hybrid Computational Intelligence Framework for scRNA-seq Imputation: Integrating scRecover and Random Forests —
- K2-V2: A 360-Open, Reasoning-Enhanced LLM —
- Boltzmann generators for amorphous particle systems —
- Foundations and Design Principles of Lightweight Cryptography for IoT Systems —
- Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model —
- BEAT-Net: Injecting Biomimetic Spatio-Temporal Priors for Interpretable ECG Diagnosis —
- MemeLens: Multilingual Multitask VLMs for Memes —
- Learning to Advect: A Neural Semi-Lagrangian Architecture for Weather Forecasting —
- CORDS: Continuous Representations of Discrete Structures —
- Semantic Calibration Prevails Where Token Confidence Fails: Benchmarking Long-Form Scientific QA —
- Near-Universal Multiplicative Updates for Nonnegative Einsum Factorization —
- Gradient-Stable Attention Heads Signal LLM Correctness —
- MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks —
- Position: A Dynamical Systems Perspective is Needed to Advance Time Series Modeling —
- StableAML: Machine Learning for Behavioral Wallet Detection in Stablecoin Anti-Money Laundering on Ethereum —
- Data-Driven Integration Kernels for Interpretable Nonlocal Operator Learning —
- Taming the Adversary: A Cost-to-Disturbance Ratio Approach to Adversarial Reinforcement Learning —
- Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation —
- How do LLMs Compute Verbal Confidence —
- Causal Evidence that Language Models use Confidence to Drive Behavior —
- Near-Optimal Primal-Dual Algorithm for Learning Linear Mixture CMDPs with Adversarial Rewards —
- Transferable knowledge graphs with executable learned operators for algorithm design —
- Nonnegative Matrix Factorization in the Component-Wise L1 Norm for Sparse Data —
- Offline Constrained RLHF with Multiple Preference Oracles —
- EAGLE: Edge-Aware Graph Learning for Proactive Delivery Delay Prediction in Smart Logistics Networks —
- Amortized Filtering and Smoothing with Conditional Normalizing Flows —
- Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation —
- Stability Enhanced Gaussian Process Variational Autoencoders —
- LLM Safety From Within: Detecting Harmful Content with Internal Representations —
- JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems —
- Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions —
- Rhamba: Region-Aware Hybrid Attention-Mamba Framework for Self-Supervised Learning in Resting-State fMRI —
- BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models —
- On the Limitations of Large Language Models for Conceptual Database Modeling —
- ISOMORPH: A Supply Chain Digital Twin for Simulation, Dataset Generation, and Forecasting Benchmarks —
- Exemplar Partitioning for Mechanistic Interpretability —
- Sometin Beta Pass Notin: Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation —
- Draft-OPD: On-Policy Distillation for Speculative Draft Models —
- Survival Reinforcement Learning: Toward Scalable Self-Supervised RL —
- Scaling Novel Graph Generation via Lightweight Structure-Guided Autoregressive Models —
- Geometry-Aware Reinforcement Learning for 2D Irregular Nesting —
- Intent-Governed Tool Authorization for AI Agents —
- MemAudit: Auditing Long-Term Agent Memory via Hidden User-State Recovery —
- Chameleon: Recovering Cyber-Physical Systems from Memory Corruption Attacks via ML Surrogates —
- TVGL-CFM:Generating and Forecasting Time-Varying Trajectories of Dynamic Networks with Conditional Flow Matching —
- ANNLib: A Development Framework for Efficient Approximate Nearest Neighbor Search —
- Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning —
- Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales —
- Contrastive Concept Importance: Explaining Pairwise Class Decisions Through Automatically Extracted Concept Representations —
- Convex losses and their applications to SVM, SVR, and Shallow Neural Networks —
- Wiktionary as a Crowdsourced Lexicon for English Dialects —
- Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges —
- Do small language models know what they don't know? —
- HERMES: Contrast-Aware Knowledge Graph Reasoning from Clinical Notes for Patient Outcome Prediction —
- TALON: A Temporally Aware Longitudinal Framework for Radiology Report Generation —
- From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators —
- Beyond WER: Entity and Disfluency Recall in Accented Conversational ASR —
- SAGE: Schema-Guided LLMs for Grant Review —
- Reviser: Revision-Capable Text Generation via Autoregressive Cursor Actions —
- Recursive Language Models Generalize Out of Domain —
- TatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar —
- Transsion's Speaker-Attributed Multilingual ASR System for the MLC-SLM 2026 Challenge —
- Towards Secure Cloud-Native Computing: Unveiling Kubernetes Misconfigurations with Large Language Models —
- A Generative Grammar Underlying the Voynich Manuscript, the Pastiche Hypothesis: Evidence from Large Language Models —
- PhysioBench: A Unified Benchmark for Physiological Signal Question Answering —
- From Generation to Detection: Exploration of Discourse Driven Scenario based LLM Generated Fake News —
- Curriculum-Based Noise Adaptation for Phoneme-to-Text Reconstruction in Visual Speech Recognition —
- COAL-SQL: Coverage-Guided Augmentation and Failure-Driven Learning for Text-to-SQL Post-Training —
- VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering —
- Boosting Deepresearch and LongContext Ability with Self-Generated Deepresearch Rollouts Traces —
- Reading Less While Writing: A Closed-Form Bandwidth Dial for Streaming Multimodal Decoders —
- Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning Models —
- Reading Anxiety or Reading the Label? Comparing Fine-Tuned and Frontier Models for Anxiety Detection on Social Media —
- Enhancing Audio Reasoning via Semantic Summary Prediction —
- MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs —
- Sparse Priors for Efficient Distribution Learning —
- The Right Tool for the Job: On the Selection of Mitigations for GenAI Privacy Threats —
- BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence —
- Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding —
- Extreme classification: beating chance with one training example from each class —
- Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes —
- Bio-MF: Low-Latency and High-Fidelity EEG-to-fNIRS Cross-Modal Generation for Hybrid Motor-Imagery Brain--Computer Interfaces —
- Continuous Delayed-Memory Stochastic Gradient Descent and Continuous-Time Reinforcement Learning from History of Astrophysical Time Series Studies —
- TPM-Attest: Hardware-Rooted Integrity Attestation as a Kernel-Level Anti-Cheat Alternative for Linux —
- Do Quantum Models Scale Like LLMs? —
- When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation —
- mu squared-Bench: A Multilingual Machine Unlearning Benchmark —
- Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications —
- From Switching to Dynamic Regret: A Simple Reduction via Unbiased Random Sequences —
- RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models —
- Complex Problem Solving in Large Language Models: A Statistical Control Survey and Diagnostic Framework —
- Attention-Aware Routing: Coupling Routing and Attention in MoEs —
- Ranking Competing geologic interpretations via foundation-model-assisted generative hydrologic inversion —
- CaLR: Causal Latent Revision for Robust Diffusion Reasoning —
- ASGARD: Action-Space Guard for UAV Resilience via Reinforcement Learning —
- From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities —
- MOSAIC-SR: Transformer-Guided Symbolic Regression for Scientific Equation Recovery —
- On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation —
- A Smoothed Discrepancy Principle for Random Feature Methods and Neural Networks —
- (Don't) Trust, but (Don't) Verify: Developers' Attention to Security in AI-Generated Code —
- Scaling Discovery through Test-Time Communication —
- Stiefel-AdamW: Geometry-Aware AdamW for Linear Factorization Blocks —
- A Lightweight Plug-in Gate for Transformer-Based Time-Series Forecasters —
- FedeRage: Provably Convergent Agnostic Federated Learning under General Client Drift —
- LoRA Enhanced Contrastive Learning with SAS Vision Transformers —
- Toward individual-level calibration in affect recognition with perceptual adjustment queries —
- Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation —
- Loopjacking: Hijacking Human-in-the-Loop Approval —
- Triply-Scalable Equivariant Gaussian Process Modeling —
- Origin Is All You Need: Provenance-Aware Transformers for Structural Trust-Boundary Separation —
- Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models —
- Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing —
- NetInspector: Measuring and Improving LLM Capabilities for Reliable Intent-Based Networking Policy Generation —
- REFINEPPO: Learning Continuous Control Policies by Iterative Action Refinement —
- Talk to Me, Jarvis: An Open-Source Edge-Deployable Voice Assistant Framework for Autonomous Racecars —
- Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models —
- From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost —
- Signal-Centric Remote Sensing via Alternative Preprocessing and Acoustic Processing for ML-Driven Applications —
- Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers —
- TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers —
- Scaling Forced Alignment to End-User Devices —
- Toss If Perishable: An Ethnographic Study on Building Scenario-Based Training for Non-Perishable Skills —
- Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake —
- EnSol: an environment-aware graph neural network for molecular solubility prediction —
- CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop —
- Can Agents Design Better Chips with a Higher Level Abstraction? —
- HMB-GAN: Hybrid Multi-B'ezier GAN for Vector Shape Synthesis —
- M2G-LLM: Enhancing Clinical Prediction via Multimodal Graph Reasoning and LLM Context Injection —
- SpecOpt: Contact-Diff Reasoning for Agentic Molecule Optimization Toward Binding Specificity —
- TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching —
- Not All Irregularity Is Equal: Causally Isolating a Rare Failure Mode in Japanese Morphological Inflection —
- Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks —
- When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success —
- SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs? —
- AI-GRACE: A Use-Case Operationalization Framework for Agentic AI: From Organizational Objectives and Obligations to Deployment Capabilities and Architecture —
- Reliability-Centered Evaluation of Sparse Longitudinal CT Lesion-Size Forecasting with Conformal Interval Calibration and Gompertz-Inspired Regularization —
- Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation —
- Ability-Residual Decoupled Modeling for Affective Cognitive Diagnosis —
- X-SPUR: Explainable Surprisal-Based Protocol-Aware Unsupervised Reasoning for Automotive Ethernet Intrusion Detection —
- A Fully Differentiable Neuro-Soft-Symbolic Framework for Perceptual Task Planning —
- Combining Exploratory Analysis and Automated Analysis for Anomaly Detection in Real-Time Data Streams —
- Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency —
- Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction —
- When Does Reasoning Help in Machine Translation? A Hierarchical Analysis of LRM Reasoning Traces —
- CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition —
- PlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking —
- Efficient Benchmarking in Production: A Study of an Evolving LLM Agent —
- Transcript-Bound Combiners for Downgrade-Resilient Hybrid Post-Quantum Key Establishment: Definition, Proof, and Embedded-Device Cost —
- How Many Humans Are 32 LLM Judges Worth? —
- MIRCID: Inferred Hub-miRNAs Drive Cross-Task Improvements in Drug Mechanistic Modeling —
- Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding —
- GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development —
- FairLMs: A Turnkey Library for Fairness in Language Models —
- Identifying Security Platform Product Abuse with Machine Learning —
- Fast And Accurate Text Content File Type Identification —
- An Introduction to Compression-Based Machine Learning —
- Diagonalized Attention for Individualized Regression: Latent-Row Localization and Prediction —
- Sparse Identification for Automatic Large-Scale Screening: A Constraint-Aware Framework with Ultra Fast Decoding Algorithm —
- LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces —
- Deep Reinforcement Learning with Buffered Quantile Objectives —
- Routine Blood Tests Outperform CRP for Distinguishing Bacterial From Viral Infection in Children —
- Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees —
- CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices —
- IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts —
- From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media Influencers —
- Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining —
- Batched Paillier-Based Hamming-Distance Computation over Binary Embeddings —
- ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL —
- Knowledge-Graph-Augmented Chronos-2 for HEC-RAS Surrogate Forecasting —
- Probabilistic Forecasting of Business Process Executions with Neural Temporal Point Processes —
- Prediction Dynamics in Depth-Recurrent Language Models —
- Consistent Relexicalization of Clinical Documents using Graph-Based Approach —
- Offline Multimodal Large Language Models for Decision Support in Air Operations —
- Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction —
- Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue —
- TrustBOM: A Scalable Architecture for Confidentiality-Preserving SBOMs Across Organizations —
- Brownian Heads for Deep ReLU Representations: Activation Mass and the Cost of Same-Sample Selection —
- DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement —
- Tracing the Evidence Behind Zero-Shot Time-Series Forecasting: A Source-First Taxonomy and Audit Framework —
- DEFEAT: Stitching Fragmented File I/O Contexts for Early Ransomware Detection —
- Decision-Focused Learning for Mean-Variance Portfolio Optimization via KKT-Based Reformulation —
- GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation —
- Optimal Randomized Proper Online Learning —
- Understanding LLM Quantization through Activation-Guided Compensation and Orthogonal Residuals —
- Improving the Predictive Performance of Bootstrap Aggregating by Dirichlet Resampling —
- Efficient Architecture Search under Leave-One-Subject-Out Evaluation —
- Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving —
- HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference —
- Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving —
- Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations —
- LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers —
- PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design —
- The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models —
- ServeGuard: Verifiable, Bounded-Residual Confinement of Operator-Invisible Channels Without Revealing the Certified Read Factor —
- Learning-to-Optimize as the Missing Architectural Layer of AI-Native Networks —
- What Must Survive? Exact Task-Information--State Frontiers for Resource-Sufficient Learning —
- IncentRL: The Trade-Off Between Preference Guidance and Task Performance —
- OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems —
- Critical sets of Latin squares based on autoparatopisms —
- MACE: Memory-Agent Co-Evolution with Adaptive Memory Graphs for Multi-Agent Systems —
- Dual-Interest Sequential Product Recommendation With Multi-Granular SSM —
- OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios —
- MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance —
- On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation —
- Et Tu, MacBook? Unprivileged Keystroke Inference and Context Profiling via the Built-in IMU Side Channel —
- Micro-Collaborative Poisoning: A Distributed Attack on RAG Systems —
- Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction —
- Beyond Accuracy: Centroid-Guided Contrastive Loss for Structured Fraudulent Job Posting Detection —
- Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education —
- Trading Depth for Time in Recurrent Transformers —
- Calibrating Teacher--Student Discrepancy for On-Policy Distillation —
- One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction —
- Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30 —
- Chinese Competitive Debating Dataset and Benchmark —
- Riemannian Neural Hamiltonian Flows: Geodesic Symplectic Transport and Interpretability —
- Analysing the Linearity of Linguistic Relations in Language Model Embedding Spaces —
- Beyond Gaussian Worlds: Latent Geometry Matters for JEPAs —
- When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap —
- Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER —
- Multi-Domain Clustering via Measure Quantization —
- Accelerating Dense LLMs via L0-regularized Mixture-of-Experts —
- PRISM-BN: A Controlled Corpus and Benchmark for Text-to-Parameterized Bayesian Network Extraction —
- GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation —
- Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation —
- CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents —
- Optimization Geometry of Equivalent Brownian RKHS Representations —
- SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference —
- TERMon: Detecting Persistent Behavioral Threats in Edge AI via Hardware-Native Ternary Runtime Monitor —
- A Framework to Quantify the Probability of Future Cyber Loss Events —
- CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords —
- Verifiable Computation with Trusted Execution Environments and On-Chain Digital Rights Tokens —
- GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning —
- World Modeling in Transformers —
- GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills —
- ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction —
- Bilevel Optimization of Topology and Hyperparameters (BOTH) —
- Per-Aetiology Contrastive Severity Embeddings with Phonological Pseudo-Labelling for Multilingual Dysarthric Speech —
- RegKT: Interpretable and Robust Deep Knowledge Tracing With IRT-Regularizer —
- CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation —
- LLM-Generated Feature Pools for Time Series Anomaly Detection —
- MIST: Multimodal Survival Prediction with Genomic-Guided Histology Attention —
- How Many Posterior Samples? Calibrated Stopping for Adaptive Sensing —
- Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods —
- RheoSampling: Resolving the One-Hot Dilemma in Stochastic Dynamic-Tree Speculative Decoding —
- Federated Deep Clustering Networks for High-Dimensional and Heterogeneous Data —
- EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise —
- Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts —
- The Weight Is Over - Interactive Diffusion on Consumer GPUs —
- Do Personality-Tuned LLMs Make Better Social Agents? —
- Watermarkable Multi-Draft Speculative Sampling via Poisson Processes —
- TrialAtlas: Multi-Agent Research Organization for Clinical Trial Design and Optimization —
- AutoRecLab: Describe the Experiment, Get the Code! —
- Neural Cellular Automata Learn General Features in their Hidden Channels —
- SFPF: Spatio-Frequency Polarization Fingerprint for Anomalous Wireless Device Detection —
- Geometric Mean Pooling for Equal-Weight Multiplicative Coarse-Graining —
- Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective —
- LLMs as Feature Engineers for Text-and-Tabular Prediction —
- ExpBoN: Exponential-Noise Best-of- n for Efficient Test-Time LLM Alignment —
- Intervention Granularity Matters: Coherent Treatment Bundles in Counterfactual Simulation with Clinical World Models —
- Beyond Kinematics: Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning —
- Riemannian Simultaneous Inference for Tangent Vector Field Regression —
- What Should We Ask Next? Retrieval-Aware Question Learning for Interactive ReID —
- Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources —
- Joint Remaining Useful Life Prediction and Capacity Estimation of Lithium-Ion Batteries Using Partial-Charging Data —
- AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory —
- End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery —
- Learning to Move Cities: Deep Meta-Models and Reinforcement Policies for Calibration and Control in Urban Networks —
- RACER: Role-Aligned Competence Estimation for Human-AI Routing —
- Provisional Reachability: Containing Agents by Making Every Crossing Revocable —
- Learning Cardiac Features: ECG Biometrics Across Time and Exercise —
- NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities —
- Time series generation with spectrally aligned latent flow matching —
- Moral Entropy: Auditing Bias and Uncertainty in Moral Judgment —
- Assessment of Machine Learning-Based Critical Heat Flux Models in the CTF Subchannel Code for Square Rod Bundle Prediction —
- A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal —
- RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents —
- Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention —
- DiaVLo: Diagnosing Behaviours of Vision-Language Models —
- COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persistence Modules —
- The Supersingular Isogeny Problem in Time and Memory p 1/3+o(1), Unconditionally —
- QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge —
- lambda-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource —
- An Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency —
- Available Guardrails: Certifying Selective Prediction across ML Systems —
- Particle Competition and Cooperation for Robust Graph Convolutional Network Learning Under Label Noise —
- Benchmarking World Models for Continual Learning on Compositional Tasks —
- BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings —
- CodeMidas: Scaling Agentic Coding RL Environments from Code Itself —
- APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport —
- Structuring occupational accident narratives for cross-sector safety analysis: Transferability of accident-process role classification —
- Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design —