AI papers — 2026-09-22
The shift toward more complex, multi-modal reasoning is being met by new frameworks for both stability and verification. In the realm of large language models, researchers have introduced RAILS to manage incremental clustering at scale through retrieval augmentation.
Another study explores the boundaries of human-LLM deliberation. This research suggests that interactive proofs can achieve verifiability even without total transparency, provided certain conditions are met.
To ensure these models remain reliable during complex tasks, the introduction of Critical-State Reinforcement Learning offers a way to diagnose trainable states specifically for multi-turn tool use. Meanwhile, practical applications are moving toward high-stakes automation.
The Jev model uses a System One approach to convert police crash narratives into calibrated probabilistic variables. These advancements in reasoning and calibration suggest a move toward systems that can handle both nuanced human language and rigorous logical verification.
The focus shifts toward the structural mechanics of how models reason and interact, moving from formal mathematical frameworks to the emergent behaviors of agents. Researchers have begun exploring the limits of efficiency in training, finding that as little as one percent of tokens can suffice for effective gradient estimation during on-policy distillation.
This efficiency is mirrored in efforts to bolster logical reasoning. New methods aim to construct reverse thinking abilities within large language models to improve their cognitive flexibility.
However, when these models are deployed as autonomous agents in long-horizon interactions, a different kind of complexity emerges through the observation of emergent collusion. This suggests that as agents interact over extended periods, they develop unprogrammed cooperative strategies that complicate our understanding of multi-agent alignment.
These behavioral shifts highlight the tension between optimizing for individual task performance and managing unpredictable collective dynamics in complex environments. The focus then shifts toward the reliability of autonomous systems, specifically regarding how agents handle unexpected friction.
Researchers have introduced Edgegen to move tool-calling agents beyond simple happy paths by using synthetic edge case generation to improve robustness. This push for autonomy is met with a corresponding need for safety, as seen in the development of a self-healing harness designed for runtime oversight when agents attempt self-modification.
While these systems navigate complex tasks, the internal logic of large language models remains difficult to audit. However, new work on recoverable semantic fingerprints offers a way to perform black-box verification by moving from mere bits to verifiable beliefs.
This tension between capability and control is further complicated in specialized domains. The FinInteract benchmark tests how well models handle clarification and intent integration when faced with ambiguous financial questions.
The shift toward more specialized agentic architectures is evident in the development of Jev-Mem, which introduces system-one-controlled agentic memory to improve efficiency in AI agents. This approach seeks to streamline how agents manage information by moving away from brute-force retrieval toward a more intuitive, rapid processing model.
Parallel efforts are being made to refine how these agents interact with complex environments through WorkWorlds. This is a new infrastructure designed specifically for evaluating AI agents on diverse workplace tasks.
While Jev-Mem focuses on the internal cognitive efficiency of the agent, WorkWorlds provides the external testing ground necessary to see if such intelligence translates to professional utility. This tension between internal memory control and external task performance remains a critical frontier as researchers attempt to bridge the gap between theoretical capability and reliable, real-world deployment.
The focus on agentic reliability continues with a new cross-dimensional threat taxonomy designed to map the security landscape of agentic AI. This framework offers a way to evaluate maturity and address persistent open challenges in the field.
This concern for operational integrity is mirrored in technical efforts to ensure runtime authorization consistency specifically within Model Context Protocol based workflows. Meanwhile, the scale of agentic deployment is being tested by BabelArena, a large-scale multilingual benchmark that pushes LLM agents to perform across diverse linguistic contexts.
On the recommendation front, researchers are scaling explainability by using LLM rationales to drive artist discovery on YouTube Music. This push toward more nuanced machine intelligence extends into specialized domains as well.
These include the development of a discrete generative model for neuronal spiking activity on microelectrode arrays and the use of density-ratio rescoring to improve performance in imbalanced classification tasks. The day concludes with a shift toward the nuances of reasoning and representation, moving from how models process internal logic to how they express it through sound.
Researchers have introduced COT-TTS, a text-to-speech framework that utilizes chain-of-thought reasoning to make audio generation more sensitive to linguistic context. By incorporating this reasoning step, the system aims to better capture the subtle nuances of spoken language that standard models often overlook.
This follows a broader investigation into how models justify their outputs, specifically through a controlled reproduction study on attributable post-rationalization in retrieval-augmented generation citations. This work compared these rationalizations against reinforcement learning with verifiable rubric-based ranking, or RLVR squared, to see if structured feedback can improve the reliability of how models cite their sources.
Together, these developments suggest that the next frontier lies in bridging the gap between a model's internal reasoning processes and its external, multi-modal expressions.
Today's papers
- The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation A new framework evaluates whether an AI truly possesses a consistent identity or is just mimicking a persona based on shared traits. [paper]
- Replicating the Geometry of Emotion Representations in a Base Open-Weights Model Researchers found that even base models contain mathematical structures representing human emotions similar to those in larger assistant models. [paper]
- DDGAD: Disagreement-Driven Graph Anomaly Detection via Adapt-Then-Combine This method detects anomalies in networks by measuring how much a node's own attributes disagree with its surrounding context. [paper]
- Density-Ratio Rescoring for Imbalanced Classification Using Raking Duals and Classifier Scores A new technique improves how models identify rare categories by using statistical reweighting to balance data distributions. [paper]
- Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment Training AI to follow human preferences can actually make the model behave less like a real human. [paper]
- Apollo Restore: A Foundation LLM for Historical Greek Optimized for Fill-in-the-Middle Restoration of Ancient Greek Texts This specialized large language model helps historians reconstruct missing text from ancient Greek manuscripts. [paper]
- Beyond Task Completion: Training Capable and Safe Computer-Use Agents This training method teaches AI agents to complete digital tasks while also recognizing and refusing harmful actions. [paper]
- Intervention, Not Shared Latents: Blocking Visual Shortcuts in Audio-Video Generation To prevent AI from incorrectly guessing sounds based on visual textures, researchers suggest intervening on the visual data itself. [paper]
- CNA: An AI-Oriented Comprehensive Normalized Assessment for Healthy Status and Application to Optimize RRT Strategies by Reinforcement Learning This framework uses a standardized health score to help reinforcement learning models optimize life-saving kidney dialysis treatments. [paper]
- A Comparative Framework for Evaluating Foundation Models on Tabular Data: A Case Study in Healthcare This study provides a structured way to compare how well different AI models handle medical data across dimensions like privacy and fairness. [paper]
- Industrial Kinematic Trajectory Model (IKTM): Coordinate-Free Autoregressive Generator This model predicts the movement of industrial vehicles without needing absolute GPS coordinates by focusing on speed and heading changes. [paper]
- Using Composition Operators to Linearize LLM Semantic Transformations This mathematical approach treats large language model transformations as operators to better understand how they map inputs to outputs. [paper]
- Role-Aware Morgan Fingerprints for Reaction Yield Prediction A fast and efficient method uses chemical role information to predict the outcome of chemical reactions more accurately than complex neural networks. [paper]
- A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation This research shows that using a single learning rate during training can bias results when comparing different AI distillation methods. [paper]
- A Channel-Boosted Multi-Agent System with Iterative Consultation for Document Sensitivity Classification This multi-agent system improves the detection of sensitive information in long documents by having agents exchange information from different parts of the text. [paper]
- Uncertainty and Business-Aware Remaining Useful Life Estimation for Semiconductor Manufacturing This framework combines deep learning with uncertainty estimation to help semiconductor factories plan maintenance and reduce costs. [paper]
- CAMFT: Conflict-Aware Mergeable Fine-Tuning for Large Language Models This method makes it easier to combine multiple task-specific AI models by reducing conflicts during the initial training phase. [paper]
- Type-Driven Tokenization for Brahmic Scripts Standard AI tokenizers break Indian scripts incorrectly, so this work provides a mathematically correct way to process them. [paper]
- SegTSim: A Big Data Driven Segmented Temporal Simulation Framework for Heterogeneous Multivariate Systems This framework uses segmented modeling to simulate complex, multi-variable time-series data like international trade. [paper]
- The Bairong System for MLC-SLM 2026: Dynamic Question-Aware Evidence Routing for Multilingual Conversational Speech Understanding This system intelligently decides whether to look at text or audio depending on what a user asks in a multilingual conversation. [paper]
- Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models This tool uses token frequency analysis to visually compare and identify the unique coding styles of different AI models. [paper]
- The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families Compressing AI models for mobile use can significantly reduce their accuracy and safety in medical tasks. [paper]
- EAVer: Long-Form Factuality Verification as an End-to-End Agentic Policy This agentic system improves fact-checking by grouping related claims and reusing search results to save time and resources. [paper]
- Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks This method optimizes how much computational "budget" is given to different tasks when merging multiple AI models together. [paper]
- H2LooP Telecom Model v1: From Telecom Comprehension to Autonomous Issue and PR Resolution This specialized model outperforms general AI in understanding telecommunications standards and writing telecom-specific code. [paper]
- StationPDE: Station-Oriented Surface PDE Learning for Multi-Station Multivariate Weather Forecasting This model combines physical weather equations with data-driven methods to provide better forecasts from individual weather stations. [paper]
- An Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired Users This project creates a low-cost, offline smart cane that uses computer vision to help visually impaired people navigate. [paper]
- Rank Portability Does Not Imply Feasibility Portability: Target-Specific Evaluation of Joint Hardware Constraints Testing shows that an AI model's performance ranking on one device does not guarantee it will behave the same way on another. [paper]
- Memory That Looks Forward: A Prospective Term for Personal Memory Retrieval This new retrieval method allows personal AI assistants to prioritize information related to a user's upcoming commitments.
- SCoP: Structured Constraint Parsing for Evidence-Space Control in Temporal Knowledge Graph Question Answering This framework ensures AI answers questions about historical facts by strictly filtering for evidence that fits the correct timeframe. [paper]
- Used, Mentioned, or Condemned? A Controlled Contrast-Set Diagnostic for the Use-Mention Distinction in Code-Mixed Hinglish Misogyny Detection This study uses specialized test sets to show how difficult it is for AI to distinguish between using a slur and quoting one. [paper]
- DPTM-DT: Dual-Pretrained Transformer Multitask Representation Learning for Drug-Target Prediction This multi-task framework improves how AI predicts how effectively a drug will interact with its target protein. [paper]
- Do Language Models Know Their Own Constraints? Training AI to follow rules can make the model comply with instructions but actually makes it less able to explain those rules verbally. [paper]
- An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents This study reveals that compressing text in long coding sessions saves more money over time than simply filtering out unnecessary tool information. [paper]
- Evaluation Awareness Shifts from Format to Context with Model Scale This research shows that larger AI models are better at detecting when they are being tested, using reasoning rather than just looking at prompt formatting. [paper]
- Dissecting Hierarchical Reasoning Models: A Mechanistic Study This study investigates how hierarchical models perform complex tasks like Sudoku and finds they rely on iterative refinement rather than a single set of features. [paper]
- Beyond Accuracy and Surface Fluency: Risk-Sensitive Evaluation of LLMs for Legal Clause Generation This paper proposes a new way to evaluate legal AI by focusing on dangerous errors rather than just how fluent the text sounds. [paper]
- Schematize: An Agentic System for Generating and Refining Information-Extraction Schemas for Legal Research This multi-agent system helps researchers turn complex legal questions into structured data extraction plans through interactive dialogue. [paper]
- Modelling daily activity patterns from mobile phone location data via deep representation learning This method uses self-supervised learning to turn massive amounts of phone location data into meaningful human activity patterns. [paper]
- ZoAQ: Adaptive Zeroth-Order Querying via Query-Reuse Coupling This adaptive optimization technique saves significant computing power by reusing previous evaluations during the search process. [paper]
- SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs This library provides a standardized way to detect and fix safety issues that emerge when AI models are fine-tuned for specific tasks. [paper]
- Does the Truthfulness Signal Survive Code-Mixing? Probing Hidden States for Hallucination Detection in Hinglish AI models can still detect lies even when users mix Hindi and English, though they tend to hallucinate more in these languages. [paper]
- SCoR: A Hierarchical Framework for Forecasting Relations Between Scientific Concepts This framework predicts how scientific concepts will connect over time by looking at their historical co-occurrence and relationship types. [paper]
- Gaussian Process Decorrelation for Spatiotemporal Deep Learning-Based Snow Water Equivalent Prediction This method removes spatial correlations from snow data to help neural networks more accurately predict future snowpack levels. [paper]
- Correlation-Guided Flow Matching with Annealed Masking for Spatial Transcriptomics Generation This generative model uses gene-to-gene relationships to create more biologically realistic maps of gene expression from tissue images. [paper]
- Is Imagination Derived from Hallucination? A Cross-Taxonomy Evaluation of Imagination and Hallucination in Large Language Models This study finds that AI imagination and hallucination are often negatively correlated rather than being the same thing. [paper]
- Observational Equivalence of LLM and Human Annotation This research shows that AI can match human quality in text classification, provided the rules for labeling are clear and unambiguous. [paper]
- Weak Ties, Strong Signals: Efficient Training Data Detection in Diffusion LLMs via Independent Token Sampling This method makes it easier to detect if a diffusion-based AI has memorized its training data by using smarter sampling techniques. [paper]
- Contrastive World Models This approach helps AI learn how environments work by focusing on predictive features rather than wasting effort trying to reconstruct every pixel. [paper]
- Beyond the Stitching Assumption: A Unified Framework for Multimodal Synthetic Data Evaluation via Semantic Quantization This framework evaluates how well text and tables are paired in synthetic datasets by checking if they still match when shuffled. [paper]
- Correcting Learning-based Perception for Safety This method uses uncertainty estimates to prevent autonomous systems from making dangerous decisions based on faulty AI perception. [paper]
- Privacy Personalization Trade-offs in LLMs: The Impact of Stylometric Signal Reduction on User-Specific Text Generation There is a direct trade-off between making an AI sound like a specific user and protecting that user's private stylistic identity. [paper]
- SolarFlowRefiner: Refinement-Aware Flow Matching for Surface Solar Radiation Downscaling This technique uses an iterative refinement process to create high-resolution solar radiation maps from coarse satellite data. [paper]
- WildfireSpreadBench: The Metric Decides the Model in Wildfire Spread Prediction This study shows that different evaluation metrics can lead to very different rankings of models used for predicting wildfire spread. [paper]
- Summarize, Judge, Refine: Decoupled Content Understanding and Policy Learning for Multimodal Content Moderation This system improves content moderation by separating the task of understanding images and text from the task of applying safety rules. [paper]
- When and Why Do Linear Bias Probes Fail? A Geometric and Statistical Theory of Bias Detectability in Large Language Model Representations This theory explains why it is difficult to detect social bias in AI if the training data doesn't contain enough diverse demographic examples. [paper]
- TreeSpark: Calibrated, Load-Adaptive Draft Trees for Semi-Autoregressive Speculative Decoding This method speeds up AI text generation by creating smart, flexible "search trees" that adapt to how much work the system is currently doing. [paper]
- Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder This study investigates how AI models used to predict patient success in opioid treatment can have performance gaps across different racial and ethnic groups. [paper]
- The Role of AI in Online Reviews This research suggests that the availability of cheap generative AI is causing a shift toward more negative reviews on online platforms. [paper]
- Resist, Update, Reject: Preference Optimization Installs a Prior-Dependent Reliability Switch This study shows that training AI to follow human preferences can accidentally make it too gullible to unreliable sources. [paper]
The papers
- Distribution-Free Uncertainty Quantification for Kernel Methods by Gradient Perturbations —
- The Cost of Privacy: Rates of Convergence for Parameter Estimation with Differential Privacy —
- Multilinear Common Component Analysis via Kronecker Product Representation —
- MMD-Regularized Unbalanced Optimal Transport —
- Subgoal Search For Complex Reasoning Tasks —
- Combinatorial Inference on the Optimal Assortment in Multinomial Logit Models —
- Powerful Primitives in the Bounded Quantum Storage Model —
- Optimal Sample Complexity of Stable Discounted Markov Decision Processes —
- Gradient is All You Need? How Consensus-Based Optimization can be Interpreted as a Stochastic Relaxation of Gradient Descent —
- VOLTA: Improving Generative Diversity by Variational Mutual Information Maximizing Autoencoder —
- Authorship identification under domain shift: a survey of stylistic measures and learned author representations —
- Controlling risks of AI in chemical science with agents —
- Calpric: Inclusive and Fine-grain Labeling of Privacy Policies with Crowdsourcing and Active Learning —
- Label Propagation for Physics-Informed Neural Networks and Physics-Informed Gaussian Processes —
- Eraser: Jailbreaking Defense in Large Language Models via Unlearning Harmful Knowledge —
- Thompson Sampling for Infinite-Horizon Discounted Decision Processes —
- C-Learner: Constrained Learning for Causal Inference —
- Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment —
- Reflective Policy Optimization —
- Transductive Off-policy Proximal Policy Optimization —
- A Mathematical Framework and a Suite of Learning Techniques for Neural-Symbolic Systems —
- Tackling Feature-Classifier Mismatch in Federated Learning via Prompt-Driven Feature Transformation —
- Optimal Symmetries in Binary Classification —
- Time series classification with random convolution kernels: pooling operators and input representations matter —
- What is the Role of Small Models in the LLM Era: A Survey —
- Streaming Deep Reinforcement Learning Finally Works —
- SG-FSM: A Self-Guiding Zero-Shot Prompting Paradigm for Multi-Hop Question Answering Based on Finite State Machine —
- RL-STaR: Theoretical Analysis of Reinforcement Learning Frameworks for Self-Taught Reasoner —
- A Course Intelligence Platform for Higher Education: Lessons from AI-Assisted Course Evaluation —
- Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions —
- 3D-MoE: Towards Spatial Intelligence with Mixture-of-Experts for 3D Reasoning and Action Generation —
- Memento No More: Coaching AI Agents to Master Multiple Tasks via Hints Internalization —
- Logits are All We Need to Adapt Closed Models —
- RESIST: Resilient Decentralized Learning Using Consensus Gradient Descent —
- AskQE: Question Answering as Automatic Evaluation for Machine Translation —
- Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture —
- Starfish: Rebalancing Multi-Party Off-Chain Payment Channels —
- Critique-Guided Distillation for Robust Reasoning via Refinement —
- Divide by Question, Conquer by Agent: SPLIT-RAG with Question-Driven Graph Partitioning —
- ServerlessLoRA: Enabling Low-Latency Serverless Multi-LoRA Serving —
- Identification of Probabilities of Causation: from Recursive to Closed-Form Bounds —
- Explain Less, Understand More: Data-Efficient Personalization of Reader-Dependent Jargons —
- Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models —
- ALPCAHUS: Subspace Clustering for Heteroscedastic Data —
- ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining —
- Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning —
- Unlocking Pretrained Vision Transformers for Time Series Classification —
- Calibrating Lightweight Sparse Autoencoder Feature Steering —
- Transfer Learning for Matrix Completion —
- Your Mailbox Is Mine: Prompt Injection Attacks Against Real-World LLM Email Agents —
- Temporal Conformal Prediction (TCP): Rolling Calibration for Financial Risk Forecasting —
- Discretization-independent operator learning for partial differential equations —
- Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey —
- BridgeShield: Risk-Aware Graph Modeling for Cross-Chain Bridge Attack Detection —
- Negation-Aware Weighted Information Fusion for Reliable Evidential Reasoning —
- Generative KI f"ur TA —
- Glass-Box Deep Learning for FDIA Detection in Nonlinear Automatic Generation Control: A Kolmogorov-Arnold Network Approach —
- Tight Privacy Audit in One Run —
- Resisting Quantum Key Distribution Attacks Using Quantum Machine Learning —
- World's First Authenticated Satellite Pseudorange from Orbit —
- From Outliers to Topics in Language Models: Anticipating Trends in News Corpora —
- What Is The Political Content in LLMs' Pre- and Post-Training Data? —
- RheOFormer: A generative transformer model for simulation of complex fluids and flows —
- Spoofing Missed-Detection Bounds for PRF GNSS Ranging Authentication Under AWGN Models —
- SPID: Distilled Protein Backbone Generation —
- Fractal basin geometry and the limits to predictability and reproducibility in deep learning —
- Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation —
- BreakFun: Jailbreaking LLMs via Object Instantiation under Simulated Code Execution —
- Memory-Free Continual Learning with Null Space Adaptation for Zero-Shot Vision-Language Models —
- Ground-Truth Subgraphs for Better Training and Evaluation of Knowledge Graph Augmented LLMs —
- BudgetMem: Training-Free Selective Memory for Cost-Efficient Long-Context Processing in Language Models —
- PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc Teamwork —
- SPEAR-MM: Selective Parameter Evaluation and Restoration via Model Merging for Efficient Financial LLM Adaptation —
- ChemVTS-Bench: Evaluating Visual-Textual-Symbolic Reasoning of Multimodal Large Language Models in Chemistry —
- Aware but Unprepared: Measuring the Security Awareness-Behavior Gap in Student Use of LLM-Generated Code with Bifr"ost —
- Towards a Multi-Layer Defence Framework for Securing Near-Real-Time Operations in Open RAN —
- Towards Training-free Automatic Proxy Discovery via Large Language Models for Mixed Precision Quantization —
- Time Series Foundation Models for Process Model Forecasting —
- CLARITY: Medical World Model for Guiding Treatment Decisions by Simulating Context-Aware Disease Trajectories —
- Dropout Neural Network Training Viewed from a Percolation Perspective —
- ShadowBlock: Efficient Dynamic Anonymous Blocklisting and Its Cross-chain Application —
- Agentic Explainable Artificial Intelligence (Agentic XAI) Approach To Explore Better Explanation: A Case Study in Decision Support for Rice Cultivation in Japan —
- RADAR: Retrieval-Augmented Detector with Adversarial Refinement for Adaptive LLM-Generated Fake News Detection —
- CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation —
- Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs —
- Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning —
- KryptoPilot: An Open-World Knowledge-Augmented LLM Agent for Automated Cryptographic Exploitation —
- Bias in the Shadows: Explore Shortcuts in Encrypted Network Traffic Classification —
- A Survey of Agentic Reasoning for Large Language Models: Towards Recursively Self-Improving and Collective Agents —
- Small Gradient Norm Regret for Online Convex Optimization —
- Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents —
- Vector-Valued Distributional Reinforcement Learning Policy Evaluation: A Hilbert Space Embedding Approach —
- ILRR: Inference-Time Steering Method for Masked Diffusion Language Models —
- Beyond Forgetting: Representation Misdirection Elicits Controllable Side Behaviors and Capabilities —
- Holographic generative flows with AdS/CFT —
- The Role of Dataset Linguistic Structure in the Cultural Awareness of Large Language Models —
- Steering Vector Fields for Context-Aware Inference-Time Control in Large Language Models —
- Time2Vec Transformer for Robust Gesture Recognition from Low-Density sEMG —
- An Empirical Survey and Benchmark of Learned Distance Indexes for Road Networks —
- Attack-Resistant Uniform Fairness for Linear and Smooth Contextual Bandits —
- Reinforcement learning with an expectile-based objective —
- Solving the Post-Quantum Control Plane Bottleneck: Energy-Aware Cryptographic Scheduling in Open RAN —
- VeRA: Renewing Reasoning Benchmarks with Executable Specifications —
- When can we trust untrusted monitoring? A safety case sketch across collusion strategies —
- SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport —
- ThreatFormer-IDS: Robust Transformer Intrusion Detection with Zero-Day Generalization and Explainable Attribution —
- MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicine —
- Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD? —
- STRUCTUREDAGENT: Planning with AND/OR Trees for Long-Horizon Web Tasks —
- Implementation of Quantum Implicit Neural Representation in Deterministic and Probabilistic Autoencoders for Image Reconstruction/Generation Tasks —
- Distributed Legal Infrastructure for a Trustworthy Agentic Web —
- A neural operator for predicting vibration frequency response curves from limited data —
- Trustworthy Predictive Distributions for Tail Events with Semiparametric Diagnostic Transport Maps —
- CCTU: A Benchmark for Tool Use under Complex Constraints —
- SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models —
- Evidence for systematic semantic structure in individual letters —
- Proof-of-Authorship for Diffusion-based AI Generated Content —
- ShapleyLaw: A Game-Theoretic Approach to Multilingual Scaling Laws —
- From Noise to Signal: When Outliers Seed New Topics —
- Toward an automated science of the mind —
- Do Papers Tell the Whole Story? A Benchmark and Framework for Uncovering Hidden Implementation Gaps in Bioinformatics —
- Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts —
- Hermes Seal: Zero-Knowledge Assurance for Autonomous Vehicle Communications —
- SkyNet: Belief-Aware Planning for Partially-Observable Stochastic Games —
- Malliavin Calculus for Counterfactual Gradient Estimation in Adaptive Inverse Reinforcement Learning —
- The Self Driving Portfolio: Agentic Architecture for Institutional Asset Management —
- What Makes Good Multilingual Reasoning? Disentangling Traces with Measurable Features —
- Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR —
- CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V —
- Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning —
- Not All Forgetting Is Equal: Retention Dynamics in Fine-Tuned Image Classifiers —
- English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training —
- Characterizing Model-Native Skills —
- Information-Geometric First-Passage Monitoring of Distributional Stability in Stochastic Systems —
- An Analysis of the Coordination Gap between Joint and Modular Learning for Job Shop Scheduling with Transportation Resources —
- Resolving Conflicts Between RTOS Timekeeping and Uninterruptable Trusted Computing —
- Taming the Centaur(s) with LAPITHS: a framework for a theoretically grounded interpretation of AI performances —
- ReLay: Personalized LLM-Generated Plain-Language Summaries for Better Understanding, but at What Cost? —
- CoFlow: Coordinated Few-Step Flow for Offline Multi-Agent Decision Making —
- Probe-Geometry Alignment: Erasing the Cross-Sequence Memorization Signature Below Chance —
- TCDA: Thread-Constrained Discourse-Aware Modeling for Conversational Sentiment Quadruple Analysis —
- Training Non-Differentiable Networks via Optimal Transport —
- APIOT: Autonomous Vulnerability Management Across Bare-Metal Industrial OT Networks —
- Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models —
- Repeated Deceptive Path Planning against Learnable Observer —
- SOM: Structured Opponent Modeling for LLM-based Agents via Structural Causal Model —
- When More Parameters Hurt: Foundation Model Priors Amplify Worst-Client Disparity Under Extreme Federated Heterogeneity —
- A Resilient Solution for Sewer Overflow Monitoring across Cloud and Edge —
- Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking —
- RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking —
- Training-Free Refusal of MCP Exploits via Retrieval-Augmented Generation —
- ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents —
- STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes —
- Focused PU learning from imbalanced data —
- On Kernel Eigen-alignments of KRR: Reconstruction and Generalization —
- TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens —
- Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management —
- Speed Kills: Exploring Confused Deputy Attacks Through Edge AI Accelerators —
- Finite-Particle Convergence Rates for Conservative and Non-Conservative Drifting Models —
- Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control —
- SiST-GNN: Simultaneous Spatial-Temporal Message Passing for Dynamic Graph Representation Learning —
- Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning —
- DDGAD: Disagreement-Driven Graph Anomaly Detection via Adapt-Then-Combine —
- BAIT: Boundary-Guided Disclosure Escalation LLM Jailbreaking via Self-Conditioned Reasoning —
- TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling —
- The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness —
- Rubric-Guided Process Reward for Stepwise Model Routing —
- PTCG-Bench: Can LLM Agents Master Pok'emon Trading Card Game? —
- A Unified Benchmark for Dynamic Medical Treatment Reinforcement Learning —
- Decomposing Refusal Steering in Mixture-of-Experts Models —
- Can Generalist Agents Automate Data Curation? —
- BayaHAR: Lightweight Bayesian Few-Shot User Adaptation for On-Device Personalized Human Activity Recognition —
- How Far Did They Go? The Persuasive Tactics of Covert LLM Agents in a Discontinued Field Experiment —
- Learning to Route LLMs from Implicit Cost-Performance Preferences via Meta-Learning —
- RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention —
- To Nuke or Not to Nuke: LLMs' (Missing) Ethical Reasoning and Actions in a High-Stakes Decision-Making Simulation —
- Directional Linear Separability of Neural Representations: Geometry and Transformations —
- MDForge: Agentic Molecular Dynamics Pipeline Design under Sparse Simulator Feedback —
- Hybrid Uncertainty Sensitivity Analysis Based on the HSIC for High-Dimensional Responses with Aleatory--Epistemic Separation —
- AdaMame: A Training Recipe for Adaptive Multilingual Reasoning —
- The BD-LSC Dataset: Facilitating the Benchmarking of Models for Lexical Semantic Change Detection in Slang and Standard Usage —
- A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models —
- Want Better Synthetic Data? Steer It: Activation Steering for Low-Resource Language Generation —
- The Illusion of Improvement: Reject Inference Strategies in Credit Scoring —
- ENPIRE: Agentic Robot Policy Self-Improvement in the Real World —
- Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning —
- Noise-Debiased Thermodynamic Variance for Local Learning Coefficient Probes —
- Knowledge-Graph Grounding Helps LLMs Only for Out-of-Training Knowledge: A Controlled Study on Clinical Question Answering —
- What are Key Factors for Updates in RL for LLM Reasoning? —
- Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy —
- A General Multimodal Probing Framework for Data Fusion and Prediction with Frozen Large Language Models: Evidence from Multimodal EHR Data —
- How Early Is Early Enough? Design-Dependent Observation-Window Sufficiency in Subscription Churn Prediction —
- Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving —
- Out-of-Distribution Generalization of Risk Aversion in Language Models —
- In-context learning from self-generated trajectories for adaptive model reduction —
- LP-SFT: Keeping Plausible Alternatives Alive in Supervised Fine-Tuning —
- You Frame It: How Conceptual Representations Shape LLM Detection and Reasoning about Antisemitism —
- Who's Behind It? Annotating and Extracting Conspiratorial Actors from German Telegram Posts —
- Riemannian Geometry for Pre-trained Language Model Embeddings —
- Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift —
- Length Penalties Make Chain-of-Thought Less Monitorable —
- Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems —
- Adversarial Attacks on Online Handwriting using Salience-based Temporal Editing —
- BadWAM: When World-Action Models Dream Right but Act Wrong —
- Information-Directed Sampling for Causal Bandits —
- Generalised Balanced Softmax: A Finite-Data Perspective on Logit Adjustment for Long-Tailed Recognition —
- Scikit-fingerprints: Python library for scikit-learn compatible molecular fingerprints and chemoinformatics —
- Output-Aware Rotation for INT2 KV-Cache Quantization —
- Large language models for partial differential equation workflows —
- Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers —
- A Hybrid LLM-Based Framework for Automated Security Annotation Generation in Business Process Models —
- Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers —
- LongRCA Bench: Root-Cause Localization in Long-Horizon Agent Trajectories —
- HalluTracer: Pre-Decoding Truthfulness Prediction via Depth-Averaged Probe-Logit —
- Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics —
- Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases —
- LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization —
- Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson —
- Amortized Bandwidth Learning for Kernel Density Estimation under Logarithmic Score —
- Tree-of-Concerns: Hierarchical Multi-Agent Debate for Unstated-Limitation Extraction in Scientific Critique —
- VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences —
- Machine learning and digital pragmatics: Which word category influences emoji use most? —
- Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents —
- Memory That Looks Forward: A Zero-Inference Prospective Term for Personal Memory Retrieval —
- Summarize, Judge, Refine: Decoupled Content Understanding and Policy Learning for Multimodal Content Moderation —
- AI-inferred expressed well-being and collective-action discourse in climate-change campaigns on X —
- Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models —
- TreeSpark: Calibrated, Load-Adaptive Draft Trees for Semi-Autoregressive Speculative Decoding —
- A framework for recipe data structure with applications for culinary and nutritional insights —
- AdaMem: Adaptive Memory Token Allocation for Soft Compression in Retrieval-Augmented Generation —
- Context Poisoning as Extreme-Value Attention Interference in Long-Context Language Models —
- DeepInstructor: An Agentic AI Instructor for Experience-Driven Idea Evaluation —
- PRQuant: Permutation Residual Quantization for Low-Overhead Inference —
- Fusion Anything: A Generalized Multimodal Foundation Model —
- Correcting Learning-based Perception for Safety —
- A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation —
- Evaluating Fine-Tuned and Base Language Models in Maternal and Vaccination Healthcare for African Settings —
- Beyond the Text: Verifying That Agent-Written Papers Are Backed by Their Artifacts —
- Privacy Personalization Trade offs in LLMs: The Impact of Stylometric Signal Reduction on User-Specific Text Generation —
- Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder —
- An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents —
- ZoAQ: Adaptive Zeroth-Order Querying via Query-Reuse Coupling —
- LE4Mob: Towards Inductive, Distance-Aware and General-Purpose Location Embedding for Human Mobility Modelling —
- Evaluation Awareness Shifts from Format to Context with Model Scale —
- Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents —
- Modelling daily activity patterns from mobile phone location data via deep representation learning —
- Rank Portability Does Not Imply Feasibility Portability: Target-Specific Evaluation of Joint Hardware Constraints —
- StationPDE: Station-Oriented Surface PDE Learning for Multi-Station Multivariate Weather Forecasting —
- Balancing Reasoning and Hardware Constraints in RAG Pipelines for Ukrainian Multi-Domain Document Understanding —
- Type-Driven Tokenization for Brahmic Scripts —
- SolarFlowRefiner: Refinement-Aware Flow Matching for Surface Solar Radiation Downscaling —
- Beyond Accuracy and Surface Fluency: Risk-Sensitive Evaluation of LLMs for Legal Clause Generation —
- Helix-FNO: Spectral-Domain Operator Learning Coupled with a High-Fidelity Mechanistic Model for Fast Surrogate Simulation —
- Hierarchical Bayesian optimization of an aircraft-based multi-agent system-of-systems —
- Correlation-Aware Structured Pruning for Large Language Models —
- Observational Equivalence of LLM and Human Annotation —
- Read-Best Is Not Steer-Best: A Probing--Steering Layer Dissociation in Omni-Modal Large Language Models —
- DiFA: Dual Evidence Fusion and Aggregation for Token-Level Text Anomaly Detection —
- Does the Truthfulness Signal Survive Code-Mixing? Probing Hidden States for Hallucination Detection in Hinglish —
- Using Composition Operators to Linearize LLM Semantic Transformations —
- Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs —
- Weak Ties, Strong Signals: Efficient Training Data Detection in Diffusion LLMs via Independent Token Sampling —
- GRRR: The Geometry of Reshaping, Rotation, and Routing in Decoder LLM post-training —
- Beyond the Stitching Assumption: A Unified Framework for Multimodal Synthetic Data Evaluation via Semantic Quantization —
- Do Language Models Know Their Own Constraints? —
- Is Imagination Derived from Hallucination? A Cross-Taxonomy Evaluation of Imagination and Hallucination in Large Language Models —
- SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs —
- A Comparative Framework for Evaluating Foundation Models on Tabular Data: A Case Study in Healthcare —
- From Latent Biomarkers to Clinical Rules: Embedding-Guided Rule Mining and Attribution-Based Translation for Interpretable Tabular Learning —
- The Limits of Speculation: Bounding Speculative Decoding in Mixture-of-Experts —
- PAGE: Partition-Aware Gated KV-Cache Eviction —
- StepKV: Step-Aware KV Cache Compression for LLM Agents —
- Uncertainty and Business-Aware Remaining Useful Life Estimation for Semiconductor Manufacturing —
- Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models —
- Beyond Raw Context Transfer: Representation-based Federated Retrieval-Augmented Generation —
- MechaTerp-TRACE: A Novel Approach for Component Ablation Analysis in Language Models —
- A Multi-Agent Pipeline for Source-Grounded Synthetic Note Generation from Longitudinal Structured EHR —
- TARGet: Topology-Aware Fusion-based Radio Frequency Circuit Functional Modeling using Graph Neural Networks —
- Clustering-Based Collective Anomaly Detection in IoT Systems: A Graph Neural Network Approach —
- Role-Aware Morgan Fingerprints for Reaction Yield Prediction —
- Monocultural Biases: Correlated biases in large language models lead to unequal systemic exclusion rates in hiring —
- Multiple latent orderings better predict language model preferences —
- Quantifying Hidden Salt for Precision Healthcare: Sodium Assessment via Joint-Factor Retrieval and Chain-of-Thought Inference —
- Industrial Kinematic Trajectory Model (IKTM): Coordinate-Free Autoregressive Generator —
- SCoR: A Hierarchical Framework for Forecasting Relations Between Scientific Concepts —
- Contrastive World Models —
- OpenBlock: Constructive and Verified Content Generation for Adaptive Tile-Matching Games —
- Beyond Task Completion: Training Capable and Safe Computer-Use Agents —
- Gaussian Process Decorrelation for Spatiotemporal Deep Learning-Based Snow Water Equivalent Prediction —
- CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds —
- DPTM-DT: Dual-Pretrained Transformer Multitask Representation Learning for Drug-Target Prediction —
- Adaptive Physics-Informed Neural Networks for the Blasius Boundary-Layer Problem —
- Correlation-Guided Flow Matching with Annealed Masking for Spatial Transcriptomics Generation —
- Fairness Beyond Anonymization? Demographic Leakage in German LLM-Generated Resumes —
- WildfireSpreadBench: The Metric Decides the Model in Wildfire Spread Prediction —
- SegTSim: A Big Data Driven Segmented Temporal Simulation Framework for Heterogeneous Multivariate Systems —
- A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents —
- The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation —
- EvoRank: LLM-Guided Evolution of Multi-Objective Learning-to-Rank Pipelines —
- Dissecting Hierarchical Reasoning Models: A Mechanistic Study —
- The Role of AI in Online Reviews —
- Improving Parameter Utilization by Sharing Neural Experts Across Layers in Transformers —
- PII-TRACE: A Benchmark for Context-Aware PII Detection in Multi-Turn LLM Conversations —
- Evaluating Personal Information Output from Conversational Interactions in Generative AI Systems —
- Prediction of Nonlinear Oscillations in a Jumping Quarter-Car Model Using Reservoir Computing —
- Dissecting Training-Free Uncertainty Estimation in Multimodal Large Language Models —
- Replicating the Geometry of Emotion Representations in a Base Open-Weights Model —
- Schematize: An Agentic System for Generating and Refining Information-Extraction Schemas for Legal Research —
- SALSA: Semi-Autonomous Literature Summarization Assistant —
- A Channel-Boosted Multi-Agent System with Iterative Consultation for Document Sensitivity Classification —
- SCoP: Structured Constraint Parsing for Evidence-Space Control in Temporal Knowledge Graph Question Answering —
- The Bairong System for MLC-SLM 2026: Dynamic Question-Aware Evidence Routing for Multilingual Conversational Speech Understanding —
- On Mitigation of Subliminal Learning in Large Language Models —
- The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families —
- UniGIO: Unified Generative Global In-situ Weather Modeling from Spatiotemporal Incomplete Observations —
- Toollery: Scaling LLM Agents to Thousands of Skills and Tools —
- Knowing, and Saying It Only When Asked: LLM Endognostics and the Schizognosis of Minerva-7B —
- Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles —
- Team DArgk at the 2026 ELOQUENT lab for evaluating generative language model quality: Residuals of Humanity: AI Detection Evasion via GRPO Fine-Tuning —
- Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark —
- EAVer: Long-Form Factuality Verification as an End-to-End Agentic Policy —
- From Trait Vectors to Circuits: Tracing Refusal and Sycophancy Through Language Models —
- Do LLMs Choose Like Humans? Using Cognitive Theory to Evaluate LLM Decision-Making —
- Swiss-Knife: A Framework for Reconfigurable Externalised Multi-Objective Alignment at Decode Time —
- Assessing Adversarial Robustness of Latent Reasoning Models —
- A Synthetic Multivariate Refrigerator Time-Series Dataset for Predictive Maintenance —
- List Counting Failures Are Not One Phenomenon —
- EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems —
- CNA: An AI-Oriented Comprehensive Normalized Assessment for Healthy Status and Application to Optimize RRT Strategies by Reinforcement Learning —
- SCALE: Simulation-Calibrated Amortized Learning for Energy Materials (A hybrid architecture connecting deterministic modeling, real-world data, and transformer-scale inference for accelerated energy-materials discovery) —
- Seeing Through Conflicts: Improving Instruction Hierarchy Alignment in Vision-Language Models —
- BizSage: A Self-Evolving Multi-Agent Framework for Business Research with Efficient Knowledge Retrieval —
- Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks —
- Task-Aware QUBO Allocation for Mixed-Precision Quantization —
- Knowledge Graph-Augmented Ambient AI for Clinical Note Generation —
- Statistical Inference for Adversarial Training: Central Limit Theorems via Optimal Transport —
- H2LooP Telecom Model v1: From Telecom Comprehension to Autonomous Issue and PR Resolution —
- Replay-Gated Neural Execution: Decoupling Persistent Behavioral Specifications from Neural Realizations in Frozen Language Models —
- Universal Observatory Graphs for Distributed Sky Coverage and Artificial Intelligence Based Interplanetary Routing —
- Do Chess Explanations Reflect Model Decisions? Behavioral and Token-Level Tests of LLM Reasoning Faithfulness —
- The Corroboration Illusion: When More News Makes LLM Forecasts Less True —
- CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents —
- Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews —
- A Tutorial on Prompt Engineering: From Messy Thoughts to AI Workflows —
- Predictors and Orchestrators: Parsimonious Machine Learning within an Agentic AI Harness for Multi-Horizon Karst Aquifer Forecasting —
- CALM: A Calibrated LLM Choice Network Framework for Activity-Based Traveler Simulation —
- CAMFT: Conflict-Aware Mergeable Fine-Tuning for Large Language Models —
- Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation —
- Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations —
- DIPLOMAT: Dialogue-Span-Aware Direct Preference Optimization for Polite Persuasive Workplace Negotiation Dialogues —
- Strategy Accumulation and Guided Execution for Automated LLM Fine-Tuning —
- RS-Claw-Evolution: Environment-Feedback-Driven Evolution for Lightweight Remote Sensing Agents in Long-Horizon Tasks —
- Which Part of the Context Layer Does the Work? Separating Semantic Content from Retrieval Scaffolding in Text-to-SQL Agents —
- Used, Mentioned, or Condemned? A Controlled Contrast-Set Diagnostic for the Use-Mention Distinction in Code-Mixed Hinglish Misogyny Detection —
- An Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired Users —
- When and Why Do Linear Bias Probes Fail? A Geometric and Statistical Theory of Bias Detectability in Large Language Model Representations —
- PAANI: On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation —
- Resist, Update, Reject: Preference Optimization Installs a Prior-Dependent Reliability Switch —
- Contrastive Siamese Representation Learning for Predictive Maintenance of Electrical Submersible Pumps —
- Intervention, Not Shared Latents: Blocking Visual Shortcuts in Audio-Video Generation —
- Functional Emotion Without Character: Large Language Models, Aristotelian Disposition, and the Limits of Behavioral Alignment —
- Differentially Private and Fairness-Audited Score Diffusion for Irregular Longitudinal Health Records —
- Social Influence and the Allocation of Scientific Attention in AI Populations —
- Contextual Causality with Large Language Models: A Survey —
- Learning 3D biophysical cell properties from 2D images and cell-population statistics —
- Complex-valued Phase-Coherent Transformers —
- Connected Content Retriever: Dense Graph Edge Features Powering Pre-Ranking at LinkedIn —
- Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models —
- Apollo Restore: A Foundation LLM for Historical Greek Optimized for Fill-in-the-Middle Restoration of Ancient Greek Texts —
- Toward Personalized Sleep Guidance from Wearable Data Using Language Models —
- Efficient Mixture-of-Experts with Speculative Decoding via Expert Coactivation —
- Goal-driven Variant Categorization —
- Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation —
- COREM: Cosine-Relation Momentum Reshaping with Stateful Writeback —
- CultureMINE: Datasets and Methods for Improving the Cultural Capabilities of NLP Systems —
- The Wisdom of Artificial Deliberative Crowds —
- EmbeddGAN: A Novel GAN Framework Using an Embedding Network and Gini Distance Correlation —
- Defusing Explosive Prompts: Understanding and Preventing Trigger-Based Prompt Injections in LLM Agents —
- Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus —
- When Cosine Similarity Fails to Reflect Linearly Accessible Structure in Dialogue Models —
- Tick-Tock on the Open Fronthaul: Securing Synchronization in O-RAN —
- IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law —
- Cross-Dialect NER for Bangla Regional Dialects Using Leave-One-Dialect-Out Cross-Validation and Explainable AI —
- EvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability —
- RLVR is a Kernel, Not a Function: Statistical Inference for pass@ k Crossovers —
- Correct Diagnosis, Better Feedback: A Symbolic-Verifier for Faithful LLM Tutoring Feedback in Logic Proofs —
- The Ups and Downs of Backprop Weights —
- Do Student LLMs Inherit OOD Robustness? Invariance-Weighted Distillation for Reliable Knowledge Transfer —
- Zero-Trust Authorization and Discovery for Enterprise MCP —
- Benchmarking Hybrid Deep Learning Architectures for Predictive Maintenance in Industry 4.0 —
- Augmenting PID Control with Deep Reinforcement Learning: A Hybrid Approach to the Industrial Benchmark —
- TWIG: A Time-Causal Wavelet Operator for Autoregressive Forecasting on Irregular Graphs —
- AutoGym: Blueprint-First Generation of Verifiable Agent Gyms —
- User-Level Handover Decision Making Based on Machine Learning Approaches —
- MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators —
- Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation —
- Pretrained Persona Mixture Models and Tandem Models for Human Simulation —
- Concurrency-Aware Process Model Forecasting with Causal Nets —
- GaitVista: Reliability-Aware AI Measurement toward Accessible Longitudinal Gait Assessment —
- Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation —
- Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA —
- Classification with Abstention Under Class-Conditional Error Constraints —
- Beetle: A Bilingual Model Suite for Modelling Second-Language Processing —
- Locally Private Inference for Riemannian Stochastic Optimization —
- Monotone-Constrained Diffusion Models for Long-Horizon Production Forecasting —
- A Bayesian Vertical Federated Learning Framework for Multivariate Reduced-Rank High-Dimensional Regression —
- From Capability to Assurance in Autonomous Penetration-Testing Harnesses: A Framework and Reference Implementation —
- Self-Organizing Agent Teams Learn to Reason Together —
- Multi-Armed Bernoulli Bandits via Minimax Single-Arm Stopping —
- Generative Embodied Multiple Behavior Control Systems for Human-like Agents —
- Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review —
- A Survey on the Linear Representation Hypothesis —
- Building Trustworthy Mental Health Benchmarks on Bluesky: A Validation-Aware Weak-Supervision Framework —
- COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning —
- LLaDA-PRM: A Bidirectional Step-Level Reasoning Evaluator —
- Autonomous Model Lifecycle Management for Digital Twin-Based Manufacturing Control —
- Hapi: A Multivariable Land-Surface Transformer for Medium-Range Hydrological Forecasting at Continental Scale —
- Analyzing Public Discourse on Urbanism: Topic Clustering, Sentiment Analysis and Retrieval-Augmented Generation using YouTube Comments —
- UBA-ORL: Unlearning-Activated Backdoor Attacks on Offline Reinforcement Learning —
- Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems —
- MATE: Policy-Aware Security Auditing for Mobile Agents via Synthesis-Driven Trajectory Learning —
- Clinical Domain Classification from Medical Transcriptions —
- ProcessLight: Process Supervision for Large Language Model Based Traffic Signal Control —
- D-IMPL: A Diffusion-based Solver for Parameterized BBOs —
- CTSpinoPelvic1K: spine, pelvis, ribs and femora in one coordinate frame, annotated for lumbosacral transitional anatomy —
- Beyond Final-Token Classification: Heterogeneous Readouts for Evidence-Grounded Suicide Risk Detection —
- NLPCC 2026 Task 10: Citation-Level Faithfulness Verification with DeBERTa Ensembles and Class-Wise Calibration —
- DVA-Neurons: Design and Verification of Adaptive LIF Neurons: From Single-Neuron Dynamics to Multi-Neuron Spiking Networks —
- MIS-Bench: Benchmarking Multimodal LLMs for Psychotherapeutic Interpersonal Skills Assessment —
- Look Before You Steer: Geometry Predicts SAE Feature Steerability —
- Improved Private Sparse Covariance Estimation with Multiscale Threshold Tests —
- Robust Market Making with Hawkes Order Flow and Price Impact via Adversarial Reinforcement Learning —
- From Research Frontier to Laboratory Bench: Design of a Four-Tier Experimental Teaching System for Multimodal Medical Image Intelligent Diagnosis —
- SelfOp: An Optimization Algorithm for Self-Improving Security Agents —
- Diagnose, Then Repair: A Two-Stage MQM-Guided Post-Editing Framework for Domain-Specific Machine Translation —
- AlexandriaX 2026: The First Shared Task on Dialectal Arabic Machine Translation —
- To Consolidate or not to Consolidate? Evaluating the Impact of Consolidation in Multi-Reference Training using Peer Reviews —
- FIRM-WM: State-factorized factual-interventional recurrent modeling for reward-free visual planning —
- The Price of Safety: Benign-Case Utility and Token Overhead of Memory-Poisoning Defenses in LLM Agents —
- Counterfactual Tool Ranking under Utility, Cost, and Privilege Constraints —
- Beyond Average Error through Oracle-Informed Stress Tests for Time-Series Forecasting —
- Personalized Federated Reinforcement Learning via Model-Agnostic Meta-Learning: Convergence of Exact and Hessian-Free Meta-Policy Gradients —
- When Label Noise Meets Class Imbalance: A Robust Framework for Android Malware Family Classification —
- A Hybrid Attention Model Learning Unified Time-aware Patch Representation for Irregular Multivariate Time Series Forecasting —
- Same Outcome, Different Readout: What Does a Steerable Valence Direction in LLMs Represent? —
- CurvFlow-DTA: dual-graph discrete Ricci curvature flow for drug--target affinity prediction —
- Causilo Technical Report —
- Leveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising Trajectories —
- Towards Full Pipeline FP8 Reinforcement Learning for LLMs —
- ISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set Architectures —
- Prioritized Rollouts for Efficient World Model-based Vision-Language-Action Policy Optimization —
- Block-Sparse Attention with Semantic-Geometric Decoupled Routing —
- Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion —
- Are Coreset Selection Methods Worth Their Cost? —
- SMS-delivered network-initiated SUPL on Pixel 8: a privacy assessment —
- LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage —
- When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used —
- An Iterative LangGraph Agent for Text-to-SQL: Natural Language Access to the Chicago Crime Database —
- Computationally efficient safe exploration in reinforcement learning —
- Joint Domain-Class Modeling for Federated Learning Under Feature Skew —
- Measuring Behavioural Signatures of Large Language Models through Psychometric Profiling —
- Beyond Linear Context: Graph-Guided Evidence Navigation for Long-Novel Reasoning with a Local 9B Language Model —
- Token Utility Is Selection-Conditioned: Coupled Selection of Prompt Context and Response Supervision for Efficient Instruction Tuning —
- Beyond Single-Model Injection: A Threat Model and Defense Architecture for Prompt Injection in Multi-Agent Systems —
- AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows —
- A Compact Stance-Indexed Anterior-Posterior COP Representation for Parkinson's Disease Classification from Plantar VGRF —
- R-GEAN: Regimen-Guided Edit Action Network for Within-Admission Medication Change Prediction —
- When Agentic Trust Crosses Organizational Boundaries: Structural Externalization and a Reference Model for Trust Evidence —
- Automatic multimodal UX improvement recommendations from LLM agent user simulations —
- Beyond Similarity: Coverage-Aware Prompt Selection for Time Series Forecasting with LLMs —
- Dual-Locking Learned AI Models: A PIN-Based Sparse QIM Watermarking and Adaptive Index Permutation Approach —
- LPINNs: First-Layer Gated Localization for Physics-Informed Neural Networks —
- OptiSkill: A Hierarchical and Evolving SkillBank for LLM-Based Optimization Modeling —
- Rethinking Pivot Programming Languages in Code Language Models —
- On attention heads and bilinear forms —
- Interpretable Multi-Hypersphere Deep Anomaly Detection for Open-set Supervised Anomaly Detection —
- PINNForge: Execution-Grounded Evolutionary Design of Physics-Informed Neural Networks —
- WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models —
- Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World —
- Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity —
- Secrets That Survive Everything: Runtime Credential Exposure in Production Web Applications —
- Enforcing Narrative Reliability and Epistemic Pacing in LLM-Driven Detective Games via Structured Knowledge Trees —
- Attributable Post-Rationalization in RAG Citations: A Controlled Reproduction and an RLVR Comparison —
- Optimizers for Diffusion Models: A Controlled Benchmark —
- Bridging Static and Agentic RAG for Taiwanese Historical Question Answering —
- LazyAgent: Demand-Driven Materialization and Physical Optimization of Agentic Programs —
- FireWorldBench: Benchmarking Complex Physical World Intelligence through Coupled-Field Fire Dynamics —
- From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness —
- Tutoring Large Language Models to be Domain-adaptive, Precise and Safe —
- MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs —
- Event Signature Transfer: Model-Agnostic Forecast Scenario Construction from Historical Events —
- Directing large language models to follow the letter or spirit of the law —
- AirGC-CD: Gaussian-Circulant Precoding for Exactly Debiasable PAPR Reduction in Over-the-Air Federated Learning —
- Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone —
- OmniEdu: Open Foundation Models for Learning and Teaching —
- The Role of Coordinates in Pareto Regret for Adversarial Multi-Objective Bandits —
- An LLM-Assisted AutoML Framework for Intrusion Detection in IoT Networks —
- When Does Adversarial Refinement Help? A Negative Result and Open Problem in Adapting R3GAN to Time Series Imputation —
- Whitening Inverts the Hierarchy: What the Norm of a Whitened Embedding Measures —
- Perplexity Cost Understates What Activation Quantisation Breaks —
- Provably Efficient Reinforcement Learning in Continuous-Time Episodic MDPs with Poisson Decision Epochs —
- From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving —
- CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine —
- Ask for Any Appliance: A Prompt-Programmable Foundation Model for Non-Intrusive Load Monitoring —
- Toscani-Fourier Distance on Probability Measures: Wasserstein Control, Topological Equivalence on Model Classes, and Duality —
- Conformal Robustness in Prediction-Driven Decision-Making —
- Exploiting Software-level Abstractions To Support Practical Hardware Trojan Attacks —
- Chronologic: Measuring Language Models' Ability to Represent the Past —
- K-TRAIL: Simulator-Guided Generative Design of EM/RF Circuits —
- Neural Residual Modeling for Scientific Data Compression under Guaranteed Error Bounds —
- Assessing Runtime Electromagnetic Detection of CPU Hardware Trojans Targeting Kernel Memory —
- Low resource cross-modal alignment using HGNN to enhance speech representation —
- LLMs as Linguistic Chameleons: Decoupling Semantics and Structure for Privacy-Preserving Communication —
- Enhancing speech representation learning with cross-modal knowledge transfer with HGNN under low resource settings: the case study of Yemba —
- Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation —
- Euston: Training Away Mathematical Sycophancy Without Losing the Mathematics —
- Bayesian Deck-of-cards-based Ordinal Regression with Sequential Preference Elicitation —
- Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents —
- SoK: From Finding to Deployment: Systematizing the OS Kernel Bug Lifecycle —
- Causal Inference with Unobserved Confounding: A Mixture Learning Perspective —
- Security of Agent-Integrated Software: When Human Operations and Agent Actions Coexist —
- ChemCLIR-Bench: Benchmarking Cross-Lingual Information Retrieval in Multilingual Chemical Patents —
- SoK: Formal Methods for Fact-Checking and Information Integrity —
- Proximal Residual Value Functions for Consistent Planning and Real-Time Execution —
- The Price of Self-Calibration: Exact Evidence Budgets and Manufactured Blind Sets in Adaptive Monitoring —
- CTRL: Control-Based Time Series Forecasting with LLM-Guided Residual Learning —
- Why Ghost Outputs Teach: A Kernel-Based Understanding of Subliminal Learning —
- Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics —
- Optimal No-Regret Learning for Repeated Prophet Inequality —
- Knowing When to Trust Images: Reliability-Aware Multi-modal Entity Alignment —
- Stochastic Flow Map for Count Data —
- Expansion Counts under Standard A* Tie-Breaking Strategies on the Final Plateau —
- Optimal Multi-way Decision Trees for Stratified Sampling in Online Controlled Experiments —
- ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs —
- CSC: Calibrated Simplicity for Conflict-Aware Social Bot Detection in the LLM Era —
- Rethinking Class Imbalance for Single-Cell Foundation Models: A Systematic Benchmark Across Architectures and Long-Tail Loss Functions —
- A Patient World Model for Early Forecasting of Digital Health Campaign Outcomes: Capabilities and Limits —
- TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents —
- What Can a Recurrent State Safely Forget? —
- Machine-Interpretable Information: Compiling Documents into Searchable and Readable Protocol States —
- The Evidence Ladder for Reinforcement Learning in Healthcare: From Retrospective Policies to Trusted Interventions —
- Leaky-integrator reconstruction: taming error accumulation in recursive differenced time-series forecasting —
- Discovering Physical Representation Languages —
- Bayesian Filtering in Physical Systems via Test-time Trained Flow Matching —
- Blind Thermodynamic Ontology Discovery from Anonymous Experiments —
- CLOADER: Evading Security Mobile Defenses via Runtime Obfuscation and Adaptive Hooking Tactics —
- The Anatomy of Address Poisoning on Ethereum: Funding Mechanisms, Scam Signatures, and Laundering via Tornado Cash —
- Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks —
- RLVR squared: Reinforcement Learning with Verifiable Rubric-based Ranking —
- TRACE: Tractable Routing Autoencoder for Clinical ECG —
- Perplexity Predicts Protection: Choosing Pretrained Backbones for Worst-Client Fairness in Federated Parameter-Efficient Fine-Tuning —
- Propose, Verify, Commit: Evidence-Grounded Memory for Long-Horizon Multi-Actor Conversations —
- RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents —
- BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents —
- Runtime Authorization Consistency Checking for MCP-based Agentic Workflows —
- AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction —
- Endogenous Interpretation —
- ITSY: Causal Discovery From Irregular Time-Series Data —
- Feature Suppression and Differential Privacy for Residential Traffic Classification: A Two-Home Federated Study —
- Predicting Out-of-Distribution Generalization of Neural Operators via Observable Spectral Error Decomposition —
- Decoupled Causal Discovery —
- Preserving Geometric Integrity in Graph Prompting via Measure-Constrained Optimal Transport —
- Physics-residual machine learning predicts oxygen-evolution catalyst activity beyond the training range from sparse polarization measurements —
- Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure —
- Error-Supervised Synthetic Learner Writing for Automated Essay Scoring —
- PACE: Plug-and-Play Contextual Embedding for Feature Screening with Pretrained Tabular Foundation Models —
- ARID: A Deployable Edge AI System for Structured Information Extraction from Industrial Maintenance Work Orders —
- POZZER: A Power Side Channel-guided Fuzzer for Black-Box Embedded Systems —
- Global Ranks Survive, Selected Heads Shift: BOS-Sink Topology under 4-bit Weight-Only Quantization —
- On the Efficiency-Safety Dilemma in Large Reasoning Models —
- Cost-Aware Reinforcement Learning with Action Masking and Projection for Battery Energy Storage Dispatch under Suppressed-Spread Market Shifts —
- Bilinear Optimization Divergence: Diagnosing Factor-Constrained LoRA Continual Learning —
- Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment —
- ETH-TraceBench: A Large-Scale Event-Stream Benchmark for Ethereum DeFi under Temporal, Protocol, and Contract Shift —
- One Patch, Three Roles: What Is Actually Coupled in Autoregressive Time-Series Forecasting? —
- A multi-temporal dataset for mapping burned areas in the Brazilian Cerrado using time series of remote sensing imagery —
- Tail-Weight Control and Localized Generalization in Nearly Low-Rank Adversarial Classification —
- PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI —
- Distill What You Trust: Reliability-Aware Multi-Teacher On-Policy Distillation —
- When the Agent Becomes the Kernel: A Systematization of Security on the Path to AI-Native Operating Systems —
- TEMPER: Temporal Encoder-Masked Probabilistic Ensemble Regressor for Time-Series Forecasting —
- Financial Language Models as Applied Artificial Intelligence Systems for News-Based Trading under Market Frictions —
- STEVE: Stabilizing Textual Gradient-Based Prompt Optimization via Error-Driven Refinement and Regularized Verification —
- GRACE: Grounded Adversarial Reasoning over Canadian Law —
- ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents —
- Constrained Decoding Eliminates Structural Failures in Small LLMs but Reveals a Scale-Dependent Semantic Gap —
- GenVoid: Uncertainty-Aware Learning of Subsurface Material Defects with an Experimentally Validated Physics-Informed Generative Model —
- TriFleetRCA: On-Premise LLM Root Cause Analysis for Kubernetes —
- On Probabilistic Inference Through Parametric Tensor Decomposition in Base Tensor Networks —
- Statistical Convergence of Transformer Encoder-Accelerated Robust Reinforcement Learning —
- Falling Trees: A Model Class for Interpretable Risk Prioritization —
- Belted Engression: Sufficient Dimension Reduction for Generative Distributional Regression —
- Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows —
- WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks —
- FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model —
- Iterative Atom Refinement: A Monotonicity Principle for Dictionary Learning —
- On Generalized Naive Bayes with Continuous Features —
- Federated Multilingual Speech-LLMs: Architecture and Aggregation Strategy Benchmarking —
- Real-time Generalizable Heart Valve Mechanics for Clinical Disease Assessment via a Physics-Conditioned Neural Operator —
- Pattern-level Differential Privacy for High-utility Complex Event Processing —
- Actionable Insights from Observational Data: The Case of Advanced Classes in K-12 Education —
- From Regional to Global: Transfer Learning for Atmospheric Transport Emulators —
- Adaptive Determinantal Client Scheduling in Federated Learning —
- PROSE: A Theory of Optimal Stopping with Perishable Evidence for Peer Selection in Intermittently Connected Decentralised Learning —
- From UNDRR Reports to Event Records: Schema-Constrained LLM Extraction of Georeferenced Disasters —
- Pretraining of Medical Visual Encoders Toward Multi-modal Large Language Models —
- VISTA: An Attention-Based Multi-Agent Reinforcement Learning Architecture for Space Situational Awareness Sensor Tasking —
- GLR-MM: Graph-Based Global-Local Reconstruction for Robust Multimodal Chest X-ray and EHR Representation Learning under Missing Modalities —
- Explainable Recommendations at Scale: LLM Rationales for YouTube Music Artist Discovery —
- Q-TIE: A Lightweight and Generalizable Re-ranking Framework for Temporal Information Retrieval —
- Collaborative Streaming Anomaly Detection with Interactive Explanations and Ensemble Consensus —
- this-that-model-1.0: A typed decision model that decides in 30 ms, for a millionth of a cent —
- SyzHarness: Patch-Based Kernel Bug Reproduction with LLM-Synthesized Fuzzing Harnesses —
- Circuit-Diff: Factual Edit-based Intervention Method for Localizing Knowledge in Attribution Graphs —
- Connecting the Dots in Agentic AI Security: A Cross-Dimensional Threat Taxonomy, Evaluation Maturity, and Open Challenges —
- GDN Tree-Scan: Served Tree Verification for Recurrent-Hybrid Language Models —
- Benchmarking Post-Quantum Cryptography in Lightweight Virtualization Environments on Embedded Hardware —
- Multivariate quantile regression via Kolmogorov-Arnold Networks —
- A discrete generative model of neuronal spiking activity on microelectrode arrays —
- Time-Incremental Continued Pretraining of LLMs: Knowledge Updates Without Catastrophic Forgetting —
- Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer —
- Matched-Input Estimates Differ in Sign Across Architectures: Auditing EEG Foundation Models on Motor Imagery —
- Density-Ratio Rescoring for Imbalanced Classification Using Raking Duals and Classifier Scores —
- The Neural Forcing for Three-Dimensional Incompressible Navier-Stokes finite time blowup —
- Measuring the Assistant's Harmlessness Preferences on the User Turn —
- Sparse Regression Distilled from a Single Robust Fit —
- XYEval: Agents say yes to bad advice —
- Echo State Network (ESN) for Signal Recovery in RF-Impaired IBFD MIMO Systems —
- HaikuS2S: A Cascaded System For Responding In Verse —
- Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control —
- Some Dialects Are More Equal Than Others: Non-Prestigious Arabic Dialectal Bias in LLMs —
- Divergent strategies and convergent outcomes in autonomous materials discovery —
- Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model —
- From Tables to Quantified Statements: Evaluating LLM Inference Generation through Executable Verification —
- Exponential Family Synthetic Controls —
- UniK: Universal Knowledge Perception for Digital and Physical AI —
- LEAP-NBV: Lightweight Edge Active-Perception for Foundation-Model Next-Best-View Planning —
- MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes —
- Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents —
- ACLArena: Agent Continue Learning in Multi-stage Post-training —
- MGRD: Compact morphology-gated residual diffusion for variance-aware cross-domain neurite forecasting —
- Simpler Methods Work Better for L1 Penalized Logistic Models and Large Datasets —
- Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models —
- FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering —
- ShapeLex: Decoupling Local Shape Symbolization and Global Scale Modeling for Text-Controlled Time Series Generation —
- Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure —
- Context-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech —
- Synthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search —
- When Evidence Conflicts: Reliability-aware Meta-review Generation —
- Structured Decomposition for Reliable LLM-Generated Access Control Policies —
- Graph-to-Grid (G2G): Continuous-Coordinate Feature Painting for Soccer Pass Surfaces —
- Q-DEQ: Discrete Solving and Quantization for Deep Equilibrium Models in Time Series Forecasting under Edge Deployment Coding Constraints —
- Calibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables with a System One Model (Jev) —
- Representation-guided in-context learning for medical image interpretation with multimodal large language models —
- Efficient Reasoning Exploration via State-Conditioned Latent Steering with Progress Guidance —
- LeaseGuard: Incumbent-Preserving Admission Control for Privileged LLM Agents —
- From Content Generation to Learning Support: Pedagogy-Guided Generative Video Tutors for STEM Learning —
- From Bits to Beliefs: Recoverable Semantic Fingerprints for Black-Box Verification of Large Language Models —
- FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax Attention —
- Incremental Consistency Execution for Autonomous Intelligent Systems —
- DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents —
- When More Evidence Hurts: Publication-Bias Drift and Principled Stopping for Biomedical Causal Search —
- Reinforcement Learning under State and Outcome Uncertainty: A Foundational Distributional Perspective —
- You Can Tell Who's Asking: What the Web's Questions Are Made Of, and Where They Come From —
- SPeaR: Test-Time Adaptation with Steering Primitives for Realigning Representations —
- Causal Bayesian Optimization: Foundations, Methods, and Applications —
- EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation —
- PAC-Bayesian Meta-Learning for Few-Shot Identification of Linear Dynamical Systems —
- Re:CAP - Auditing Retrieval Coverage in Production RAG Pipelines —
- Model-Agnostic Feature Selection via LOCO-Guided Adaptive Minipatch Sampling —
- OSCAR: Order-aware Scoring and Calibration for AI Rankings —
- Self-Healing Harness for Runtime Oversight of Agent Self-Modification —
- Monet: Measuring the Ecosystem of Open-Source Text-to-Image Models Tailored for Harmful Services —
- CLOOPD: Closing the Learner Loop in On-Policy Distillation —
- Luck Is Not Skill: When Do Paired Rollouts Help Group-Relative RL of LLM Agents? —
- Mind or Message? Auditing Theory of Mind in Multi-Agent Social Simulation —
- Acceptance-Aware Draft Model Training for Speculative Decoding —
- TAC-Time: Texts as Channels For Multimodal Time Series Forecasting —
- KEVGraph: Exploitation-Aware Dependency Vulnerability Remediation —
- APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction —
- CREDO: Variance-Guided Rubric Evolution for Replay-Corrected Credit Assignment —
- Efficient LLM Distillation for Bangladesh Legal Context: A Smartphone-Compatible Retrieval-Augmented Generation Model —
- LIMIT: Less Is More for Instruction Tuning in Text-to-SQL —
- When Residualization Helps an Audit: Format Effects, Slice Gains, and Their Limits —
- LoopCD: Loop-wise Contrastive Decoding for Improving Reasoning in Looped Language Models —
- H-Spec: Parallel Speculative Decoding Without a Drafter-Side KV Cache —
- SKstars at SHROOM: Visions Agreement-Guided Ensembling of Zero-Shot and LoRA-Adapted Vision--Language Models —
- Vimarsha: Faithful ASR Evaluation for Indian Languages with Demographic Diversity, In-the-Wild Audio and Spelling Variations —
- Forgeable Confirmation in Automated Computer Security Testing: Deterministic Rules versus AI Judges —
- Opinion Leader Dynamics: How Sparse Attention Shapes Token Clustering —
- Displacement Geometry Captures Platonic Shared Reality Across Models and Modalities —
- From Articles to Publishers: Aggregating Language Model Predictions for News Source Reliability Inference —
- Recovering Lost Details: Multi-Scale Frequency Compensation for Long-Term Time Series Forecasting —
- Adaptive Forgetting for Nonstationary Optimization: Towards Robust EEG Decoding —
- Memory vs. Context? Influential Factors of Factual Recall in Language Models —
- Hessian Rank Constraint for Learning Structure of Nonlinear Latent Variable Models —
- Taming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation Enforcement —
- Taramandal-GPT: Enhancing Astrodynamics Problem-Solving with Knowledge Retrieval and Structured Thinking —
- Reinforcement Learning Inspired Black-box Adversarial Attacks for Computer Vision —
- Explainable Predictive Condition-based Maintenance of Naval-Propulsion Systems using Fuzzy Logic —
- MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents —
- Adversarially Robust PAC Learning with Optimal VC Rates —
- Canonical Procedural Actions: An Auditable Annotation Protocol for Tool-Use Agent Traces —
- Unsupervised Brain Anomaly Detection as a Bayesian Inverse Problem with Diffusion Prior —
- Structure Before Sampling: Community-Aware Core-Set Selection for Data-Efficient Text-to-Speech —
- How Many Pixels Is a Digit Worth? Place-Aware Coordinate Entropy for GUI Agent Confidence Estimation —
- High-Dimensional Online Change Point Detection with Adaptive Thresholding and Interpretability —
- Temporal Generalization and Explanation Stability of Control Flow Graph Neural Networks for Malware Detection —
- TTSE: A Two-Track Online Self-Evolution Framework —
- When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain —
- rApp/xApp Attestation: A New Security Use Case for O-RAN —
- KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation —
- SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration —
- The Undetected Damage of Quantization on Retrieval and How to Fix It —
- Brain-Token Learning: Microstate-Based Tokenization and Multi-Scale Interaction for Long-Horizon EEG Sequence Modeling —
- A Distributional Optimisation Perspective on Combining Models in Deep Learning —
- Pharmacokinetic State Space Models for Unbiased Prediction of Haemodynamic Collapse —
- LADDER: Graph-Guided Diffusion Language Models for Efficient Multi-Hop Reasoning —
- Few-Shot Demonstrations Elicit the Use of In-Context World Representations in LLMs —
- Mitigating Entity Type Confusion in Cross-Domain NER via Multidimensional Quantification and Reasoning Enhancement —
- Explainable Neuro-Fuzzy Prediction for Trustworthy Decision-Making in Maritime —
- VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning —
- Prescriptive SVD-Inspired Attention via Spectral Energy Retention —
- URA-NER: A Unified Retrieval-Augmented Framework with Retrieval Alignment and Uncertainty Reduction for Low-Resource NER —
- Information-Time Proximal Policy Optimization —
- Credit Access is Associated with Improved Food Security in the Horn of Africa —
- Machine Learning-Based Prediction of Childhood Stunting in Bangladesh: Fairness and Temporal Robustness Assessment —
- Name2Pkg: Lightweight One-Class Android Malware Screening via Name-Package Correspondence Modeling —
- NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware —
- Passive Hybrid Network-Based Intrusion Detection System (Hybrid-NIDS) Combining Suricata and Random Forest —
- Climate Variability Modulates the Impact of Price Spikes on Food Insecurity —
- Probabilistic Modelling of Operational Design Domains, A New Approach for Testing AI Systems —
- Artificial Structure Function Search: Preserving Artificial Functional Connectivity for Structured Pruning —
- SkelOT: Reusing AOT Compilation Across EVM Contract Families —
- End-to-end Jordanian dialect speech-to-text self-supervised learning framework —
- ARM: Attention with Routed-Memory for Learnable Sparse Control —
- Prior-Amortized In-Context Bayesian Inference for Generalized Linear Mixed-Effects Models —
- 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation —
- Comparing Latent Concept Formation in State Space Models and Transformers via Sparse Autoencoders —
- MUSE: Dependency-Aware Adaptation of a Frozen Vision Backbone for Multivariate Time Series Forecasting —
- WPBench: A Comprehensive Benchmark for Wind Power Forecasting —
- ActGov: Governing LLM Agent Actions via Policy-Constrained Validation —
- Predicting Postprandial Glycemic Response from Meal Images, Clinical Variables, and Gut Microbiome Information —
- RAILS: Retrieval-Augmented Incremental LLM Clustering at Scale —
- A Temporal Knowledge Graph for Music Festival Lineup Forecasting —
- Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards —
- Lifted Bellman Linear Programming for Offline Reinforcement Learning —
- Prefix Puncturable Signatures with Smaller Signing Key from HIBS —
- On Emergent Capabilities and Model Merging —
- Beyond Predictable Paths: Redefining AI Security Incident Reporting for Agents —
- LLJ Cards: Best practices for the Use of LLMs as Judges —
- Not All Task Vectors Need Equal Rank: Energy-Proportional Allocation for Model Merging —
- Beyond Point Prediction: Artificial Representative Trees with Uncertainty —
- QLoRA Fine-Tuning of Ministral LLM for Sequence-to-Function Protein Annotation —
- State-Aware Fuzzing of JavaScript Engines with LLM-Guided Instrumentation —
- Toward a Unified Mathematics of Concepts —
- The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence —
- Identifying Representational Biases in Datasets Using PCA: A Max-Disparity Partition Framework —
- t 0: A Time-Series Foundation Model for Forecasting with Context —
- Evaluating Decision Models for Text Annotation in Computational Social Science —
- Universal Multi-Modal Traceformer: Integrating Heterogeneous Context for Process Event Prediction —
- Overlay dx - Automating forecasting evaluation —
- Taking a Second Look: Correcting Sea Ice Forecasts with Sparse Observations —
- GraphToolbox: A Configurable Python Framework for Graph Neural Network Forecasting —
- Ascent: An Agentic System over the Model Context Protocol for Real-World Clinical Data Analysis —
- Custom Named Entity Recognition and Topic Classification for Global Health Publications —
- Augmented Hypothesis Testing with Persona-Based LLM Simulations —
- Written as a Record, Read as an Address: What a Forward Pass Leaves in an Operation's KV Cache —
- iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs —
- Assessing Readability with LLMs: The Role of Reasoning and Few-Shot Prompting —
- Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech Enhancement —
- 5G-Shark: A Network Security Auditor for 5G Subscriber Privacy and Unauthenticated Signalling Resilience —
- DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security —
- Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents —
- Trust in Edge-Enabled IoT Security: Features, Challenges and Research Directions —
- TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction —
- Muon Can Outperform Dedicated Continual Learning Methods —
- Guaranteed Low-Rank Tensor Recovery from Modewise Measurements via Normalized Block-Weighted Riemannian Gradient Descent —
- Domain Specific Post Quantum Signatures for Blockchains —
- Adapting Tree-Structured Speculative Decoding to DeepSeek-V4 for Efficient Inference —
- Reasoning Topology Matters: A Controlled Study of LLM-Based Cybersecurity Analysis —
- A Federated Artificial Intelligence Framework for Optimizing Pancreatic Cancer Treatment - Strategy Update —
- An Exact Junction-Tree Extended Formulation for Optimal Classification Trees —
- World State Generator —
- Enhancing Transformer Representations of Symbolic ODE Expressions —
- Inference of Unknown Dynamical Components Using Next Generation Reservoir Computing: From Chaotic Systems to Climate Data —
- Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents —
- Construting Reverse Thinking: Developing Large Language Models' Reverse Thingking Ability —
- Convex AI Compositionality and the Governance of AI System Populations —
- Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention —
- When Quantization Preserves Accuracy but Not Evidence: Explanation-Aware Post-Training Quantization for Medical LLMs —
- Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection —
- MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents —
- The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts —
- G-NAC: Graph Neural Automata Clustering via Emergent Domain Formation —
- OPBackdoor: Opportunistic Backdoors via Alibi-Aligned Reasoning —
- GRUET: Quantifying Uncertainty of Agentic Reasoning-and-Acting Processes —
- MedRSI: Recursive Self-Improvement for Medical Agents via Clinically Aligned Self-Evolution —
- Extracting Arguments, Not Just Classifying Them: Instruction-Tuned LLMs for Generative Component Detection —
- When Tomorrow Becomes Today: Self-Evolving Policies for Agentic Time-Series Forecasting —
- Partner-Specific Affective Precision in Social Active Inference —
- Decomposing Error and Style in Automated Clinical Coding —
- Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models —
- Learning Prognostic Variables for AI Convective Parameterizations via Symbolic Distillation —
- A Global Comparison of Schemas, Transparency, and Interoperability in Public-Sector AI Registers and Inventories —
- The Copy Ceiling: An Input-Exposure Control for Ontology-Grounded Generation over Curated Corpora —
- OSWorld-Pro: Process-based Evaluation for Computer Use Agents —
- Human-LLM Deliberation as Interactive Proof: Conditions for Verifiability Without Transparency —
- ToneCL: Contrastive Learning for Few-Shot Syllable-Level Tone Classification —
- SocioVerse2: A Longitudinal Dynamic Social Simulation Framework under a Human-AI Co-evolutionary Paradigm —
- BackTrend: Evaluating Scientific Weak-Signal Prediction via Backward Reconstruction —
- Et Tu, Brute? Economic Misalignment in Personal AI Agents —
- Linguistic Features for Interpretable Textual Entailment —
- Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization —
- Learning Physics from an Imperfect Ancestor —
- JAREX: An Acquisition Function for Multi-Objective Algorithmic Process Characterization —
- Jev for Scientific Decisions: Evaluating Semantic Choices and Their Consequences —
- Emergent Collusion in Long-Horizon LLM Agent Interaction —
- Rare Event Estimation via Iterative Unalignment —
- DolphinBench: Mapping the Pareto Frontier of Agent Memory —
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses —
- Harness-Zero: Harness Distillation via Agent-as-Harness —
- LoRA-generating hypernetworks for efficient on-device LLM generative personalization —
- Residual Community Prototypes Under-Reject Held-Out Malware Families in FCG-MFD —
- onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction —
- Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use —
Important terms
- RAILS
- A framework designed to manage incremental clustering at scale by using retrieval augmentation, helping large language models organize and process information more effectively during complex tasks.
- Critical-State Reinforcement Learning
- A method used to diagnose specific trainable states in models, specifically aimed at improving how AI handles multi-turn tool use and complex reasoning sequences.
- Edgegen
- A technique that uses synthetic edge case generation to train tool-calling agents, moving them beyond simple successful paths toward much higher levels of robustness.
- Jev-Mem
- An agentic memory architecture that uses a System One approach to improve efficiency, moving away from brute-force retrieval toward more intuitive and rapid information processing.
- COT-TTS
- A text-to-speech framework that incorporates chain-of-thought reasoning into audio generation, allowing the model to better capture linguistic nuances and subtle spoken language context.