AI papers — 2026-09-23
Today’s briefings begin by addressing how we interpret machine intelligence across scales ranging from human cognition to recursive computation architectures. These studies suggest our current frameworks for reliability remain dangerously incomplete because even when models appear functional, they often fail at deeper levels of reasoning or systemic stability.
First, large language model explanations fall short when tasked with teaching humans active learning strategies, which implies an urgent need for more pedagogically sound alignment methods. Simultaneously, researchers have discovered that type safety does not guarantee error freedom since decision heads frequently prioritize option names over actual rubric bounds, creating hidden logical gaps.
New findings on reproducible AI emphasize that true reproducibility requires strictly controlled randomness. Recent investigations into recursive models reveal critical questions regarding exactly when such systems finish computing effectively, which makes sense alongside evidence that proximal background context can overshadow distant evidence like sirens pulling focus away from truth itself.
Finally, we must confront trace optimized agent hijacking within MCP ecosystems as well as noise sources where internal versus external disturbances fundamentally alter adaptive regulation patterns. These issues leave us much further from truly robust autonomy than previously assumed.
While researchers have long discussed the benefits of AI in cybersecurity, new formal modeling suggests that human-AI collaboration requires a precise balance to avoid diminishing returns. By treating defense-in-depth as a detection cascade where AI augmentation acts multiplicatively across layers, researchers found that AI provides its greatest marginal gains precisely where traditional layering begins to saturate.
However, the study highlights a critical trade-off in triage: attempting to achieve full human review of all AI-flagged alerts can actually lower system-level detection probability. While increasing analyst capacity toward total coverage can reduce false alarms by approximately 20-fold, it simultaneously introduces imperfect human accuracy into every alert rather than just a filtered subset. This suggests that security operations centers should aim for an interior optimum of capacity rather than maximum coverage to maintain effective detection.
The difficulty of ensuring model reliability extends into how we evaluate code generation efficiency via BigO Benchmarking tests. Researchers investigated whether large language models can produce code within controlled time or space complexity constraints alongside concerns regarding inference stability.
Specifically, new findings suggest that greedy decoding fails to maintain precision invariance because cross precision output divergence occurs during inference tasks even when inputs remain identical across different computational settings. Similarly, addressing model behavior, researchers have begun distinguishing between true receptiveness toward user intent versus mere sycophancy to clarify whether an agent is engaging meaningfully or simply deferring to user bias.
This distinction becomes critical when auditing product decisions made by autonomous agents through investigations into delegation blind spots. In these cases, human oversight often fails to catch errors stemming from delegated agency choices, making robust alignment more complex than simple instruction following suggests.
The push toward more reliable automated reasoning continues with developments like VeriSimpl, which utilizes simplification based verification to ensure optimization models derived from natural language remain robustly accurate during translation into formal structures. Nearby, researchers have turned their attention back to evaluation frameworks such as ReasonLab, which offers controlled auditing for prompting techniques used specifically within multiple choice question answering tasks.
At the same time, studies are addressing broader systemic concerns regarding reproducibility by looking beyond simple agent architecture to examine how execution assumptions impact large language model based trading systems. Elsewhere, progress moves into generative chemistry via STAR VAE, which employs scalable latent variable transformers for controllable molecular generation alongside geometric investigations aimed at discovering data manifold geometry by analyzing its inherent properties.
These advancements suggest an increasing focus on precision, whether that involves ensuring fairness against label poisoning attacks in distributed learning or refining reinforcement learning through critic alignment strategies like PACT.
The challenge of managing long-horizon tasks for agents has led researchers to investigate how task state should be communicated to a model. In a study evaluating whether task state should be shown as text, told via directives, or enforced through external gates, researchers found that simply displaying an accurate transcript is often unreliable.
Interestingly, an unverified ledger written by the agent itself can actually outperform an accurate checklist provided in the prompt. While per-turn directives from a state machine improve performance in proportion to a model's obedience, hard enforcement gates are most effective when failures are frequent and state-decidable, though they can decrease performance if the gate makes incorrect judgments.
This effectiveness is highly dependent on the task; for instance, an enforcement gate improved a 235B parameter agent's pass rate from 0.39 to 0.54 on airline policy tasks but had no effect on smaller models that rarely violated those policies. Meanwhile, the problem of whether an agent can trust its own world model has been addressed through Dual-Frontier, a principle designed to solve the failure-attribution problem where it is unclear if a mistake stems from the agent's decision rule or an inaccurate model.
By only admitting decisions when their predicted advantage exceeds a certified error bound, this method improves reliability across tool-use benchmarks. To support these complex, long-horizon interactions, SAGE introduces a self-evolving graph memory engine that moves beyond static retrieval by using a reader-writer framework to incrementally construct and refine structured memory from interaction histories. This approach has shown significant improvements in multi-hop question answering and retrieval efficiency after just two rounds of self-evolution.
The shift toward more efficient model deployment is evident in recent work on modular ensembling and small language models. Modular Norm RandOpt improves upon standard weight-perturbed sampling by using architecture-aware, module-wise natural norms to select ensemble candidates. This approach achieves higher mean accuracy on tasks like Countdown and GSM8K across various Qwen scales while requiring up to twelve times fewer candidates on GSM8K than previous methods.
Similarly, research into small language models as educational assessment designers suggests they can achieve competitive performance in generating questions aligned with Bloom’s taxonomy. While these smaller models offer vital privacy-sensitive deployment advantages, their model-based evaluations still show systematic biases and inconsistencies when compared to expert human ratings. This tension between efficiency and reliability underscores the ongoing necessity for human-in-the-loop workflows even as model scale decreases.
The landscape shifts toward more autonomous reasoning architectures as researchers explore recursive self improvement within AI research agents alongside new control mechanisms like REFLEX. This system utilizes JEV to enable efficient selective control during agentic workflows by employing JEV as a judge to accept confident outputs or escalate uncertain ones when necessary.
However, this push toward autonomy must contend with fundamental stability issues such as double descent patterns where diffusion models risk malign overfitting despite increasing scale. Meanwhile, geometric approaches seek to bridge theory into practice by investigating whether anomaly detection performance can be predicted directly from embedding space geometry, while algorithms like SuperPCA offer faster ways to navigate these high dimensional spaces via subspace analysis. Throughout these evolving computational domains, foundational mathematical tools like Fourier Bessel wavelets continue providing essential frameworks for signal processing.
The final frontier remains bridging high level reasoning with practical deployment efficiency across specialized domains like engineering software development or biological modeling via graph generation using transport coupled bayesian flows or structural learning within stochastic population dynamics without relying on simulations altogether. Next week, we might see how these paradigms scale when applied to large scale agentic engineering benchmarks such as SWE-bench, which evaluates production inference serving performance during complex coding tasks.
We will also explore if modular specialist agents can outperform monolithic context heavy scaffolds by growing task harnesses instead of expanding input windows. More broadly, our focus shifts toward whether adaptive temporal brain connectivity models like BrainAtcl can maintain precision in functional link prediction even as age estimation becomes more computationally demanding under low rank attention residual constraints. Across these diverse architectures, we continue seeking stability between representation enrichment in conversational speech emotion detection and real world utility in ship radiated noise recognition through linear probing on pretrained embeddings.
Today's papers
- Amplifying Randomized Encodings & Applications Researchers show that one-sided randomized encodings can be amplified to achieve negligible privacy and error. [paper] [episode]
- TriHaRd: Higher Resilience for TEE Trusted Time This protocol provides high resilience against clock manipulation attacks in trusted execution environments through Byzantine-resilient updates. [paper] [episode]
- WiP: Towards a Secure SECP256K1 for Crypto Wallets This hardware architecture uses parallel processing and complete addition formulas to protect cryptocurrency wallets from side-channel attacks.
- Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation DigenRL is a framework that speeds up reinforcement learning for diffusion models by allowing flexible resource allocation. [paper]
- Deep Reinforcement Learning on Item-Compatibility Graphs for One-Dimensional Bin Packing This graph reinforcement learning framework generalizes to any problem size to solve the one-dimensional bin packing problem. [paper]
- Conditional Distributional Treatment Effects: Doubly Robust Estimation and Testing This paper introduces a robust estimator and test to capture how treatments affect the entire distribution of outcomes. [paper]
- The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance Alignment training in large language models can cause consensus collapse, reducing the diversity of opinions in simulated populations. [paper]
- LingLan: An Advancing Traditional Chinese Medicine Diagnosis LLM with Multimodal Data LingLan uses multimodal data to help large language models emulate the diagnostic process of traditional Chinese medicine. [paper]
- Toward Responsible AI-Augmented Cyber Defense: Pattern Recognition, Defense-in-Depth, and the Case for Human-AI Collaboration This paper provides a formal model to optimize the balance between human analysts and AI in cybersecurity operations. [paper]
- Prediction Is Not Detection: Evaluating Pre-Recognition Claims in Longitudinal Clinical AI Researchers propose new definitions to prevent overestimating how early AI can detect clinical events. [paper]
- Canonical locks that encode part-whole hierarchies This method uses high-dimensional vectors and phase differences to help neural networks represent hierarchical structures. [paper]
- VideoX-Qwen: Data-Centric Instruction-Based Video Editing This framework provides a scalable way to train models for complex, instruction-driven video editing tasks. [paper]
- Robust Photoplethysmography Signal Denoising via Mamba Networks This deep learning framework uses Mamba networks to remove noise from heart rate monitoring signals while preserving physiological accuracy. [paper]
- A Lightweight Plastic-Memory Framework for Graph Few-Shot Class-Incremental Learning This framework uses an evolving micro-clustering structure to help models learn new graph classes without forgetting old ones. [paper]
- The Dynamics of Quasiregular Neural Learning This study examines how neural networks navigate the competition between dominant regularities and rare exceptions during learning. [paper]
- Reply to comments arXiv:2512.07881 and arXiv:2601.06104 on quantum structure in human and AI-generated language The authors address criticisms regarding their research on quantum-mechanical statistics in human and artificial language. [paper]
- Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application This work introduces a way to quantify the probability of a multimodal model being jailbroken to better optimize defenses. [paper]
- How It's Made: Uncovering Detection Engineering Processes for Network Intrusion Detection Rules This study analyzes how security professionals create and iterate on network intrusion detection rules. [paper]
- Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament Researchers found that political blame in Denmark has increased significantly, driven by ideological hardening on the right. [paper]
- Quantum-Ready Secure WAN: A Risk Assessment and Migration Framework This paper proposes a framework to help enterprises prioritize which services to migrate to post-quantum cryptography. [paper]
- CRT-Decomposed-Protocols for CSIDH This work constructs a zero-knowledge proof for the CSIDH group action by exploiting its mathematical structure. [paper]
- Domain-Adaptive Pretraining Enhances Water Treatment Semantic Representation for Large-Scale Structured Literature Mining WaterBERT is a specialized model that improves the extraction of information from water treatment research papers. [paper]
- CAffNet: Hard Constraint-Affine Neural Networks This framework allows neural networks to satisfy hard mathematical constraints during the training process. [paper]
- FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents This method increases the reliability of AI agents by applying corrective instructions at states that previously led to failure. [paper]
- A Survey on Long-Term Memory Security in LLM Agents A survey exploring the various attacks, defenses, and governance needs related to memory in large language model agents.
- Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models This framework uses a shared scratchpad for human experts to edit a model's reasoning trace during fact-checking. [paper]
- Towards Mitigating Excessive Forgetting in LLM Unlearning via Entanglement-Guidance with Proxy Constraint This method uses entanglement and proxy constraints to help models unlearn specific data without losing general utility. [paper]
- Discovering Data Manifold Geometry through Geometric Properties This approach learns the underlying structure of data by directly optimizing for geometric properties like commutativity and time coherence. [paper]
- Decoding the Legalese: A Scalable and Quantitative Framework for Analyzing Corporate Privacy Policies This system uses large language models to convert complex privacy policies into standardized, measurable data. [paper]
- MAC-RRG: Iterative Multi-Agent Collaboration for X-ray Radiology Report Generation This multi-agent framework improves radiology report generation by iteratively refining reports using medical knowledge graphs. [paper]
- Self-Cleaning and Captured Anyway: One Measured Primitive for Error in a Store an Agent Writes to Itself This study investigates how agents' own written conclusions can lead to unintended bias or drift over time.
- Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation Using specific personas in LLMs actually makes them worse at predicting real human engagement than simple baseline models. [paper]
- How Strongly Should Task State Influence an LLM Agent? This study finds that showing an agent its task history is often less effective than having a system strictly enforce the task state. [paper]
- Dual-Frontier: When Can an Agent Trust Its World Model? This principle helps agents decide when to trust their internal world models and when to perform extra verification. [paper]
- SAGE: A Self-Evolving Agentic Graph-Memory Engine for Structure-Aware Associative Memory SAGE is a dynamic memory system that uses feedback to continuously update and improve its internal knowledge graph. [paper]
- Convergence of gradient descent for deep neural networks This paper provides mathematical conditions under which gradient descent is guaranteed to find a zero-loss solution. [paper]
- Double Descent and Malign Overfitting in Diffusion Models This work shows that overparameterization can cause diffusion models to suffer from catastrophic memorization rather than beneficial generalization. [paper]
- A Spectral Theory of Grokking: Weight Decay induces Feature Learning This theory explains how weight decay helps neural networks transition from simple memorization to learning complex features. [paper]
- Semantic Self-Distillation for Language Model Uncertainty This technique distills complex semantic distributions into lightweight models to provide fast uncertainty estimates for LLMs. [paper]
- From Utterances to Networks: Modelling Slang Adoption and Diffusion Across Subreddits This study uses LLMs to show how social network structure and linguistic context drive the spread of internet slang. [paper]
- Modular Norm RandOpt: Population-Efficient Ensembling through Architecture-Aware Perturbations This method makes model ensembling much more efficient by using architecture-aware sampling. [paper]
- Error Bounds for Statistical Estimators in BTL Model with Parametric Multivariate Utility Functions This paper derives mathematical error bounds for estimating preferences in the Bradley-Terry-Luce model. [paper]
- Small, Private Language Models as Teammates for Educational Assessment Design Small language models can effectively assist in designing educational tests while maintaining student privacy. [paper]
- Explanation-Guided Medical Named Entity Recognition with Stability and Boundary Awareness for Atopic Dermatitis This framework uses stable explanations to improve the accuracy of identifying medical entities in clinical texts. [paper]
- From Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought This study shows that chain-of-thought reasoning is only truly necessary for difficult tasks, while models often bypass it for easy ones. [paper]
- Learning to Fluctuate: Statistical Foundations for Causal Tabular Pretraining This method improves causal effect estimation by training models to predict how responses fluctuate across different samples. [paper]
- Quantum Model Parallelism for MRI-Based Classification of Alzheimer's Disease Stages This architecture uses parallel quantum circuits to improve the accuracy and efficiency of diagnosing Alzheimer's from MRI scans. [paper]
- What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation This research finds that providing too much information to a self-teaching model can actually hinder its learning. [paper]
- SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation This framework generates new test cases on demand to prevent models from gaming static safety benchmarks. [paper]
- Refit the Probe: Single-Direction Ablation Is Not a Necessity Test This study demonstrates that standard methods for testing if a model uses a specific feature are often mathematically flawed. [paper]
- Agent Memory: Characterization of System Implications of Stateful Long-Horizon Workloads This work provides the first comprehensive look at how different memory systems affect the cost and performance of AI agents.
- A Multi-Timestep LSTM Ensemble regressor for Enhanced Short-Term Runoff Prediction This approach uses an ensemble of LSTM models to more accurately predict river runoff patterns. [paper]
- MultiViewDx: Evidence-Linked Multi-View Clinical Diagnosis This framework improves medical diagnosis by linking multiple imaging views with patient context in a structured workflow. [paper]
- UPPRESSO: Untraceable and Unlinkable Privacy-PREserving Single Sign-On Services This privacy scheme allows users to log into multiple services without letting providers track them or link their identities. [paper] [episode]
- Label-Efficient Learning for Ground-Based Sky-Image Classification: A Benchmark of Transfer Learning, Active Learning, and Pseudo-Labeling on GCD Transfer learning is the most efficient way to train models for cloud classification when labeled data is scarce. [paper]
- CHiME-9 ECHI: A Machine Learning Challenge for Enhancing Conversations to Address Hearing Impairment This challenge focuses on improving speech intelligibility in noisy, multi-person environments for people with hearing impairments. [paper]
- Calibrated Confidence Expression for Radiology Report Generation This reinforcement learning framework trains medical models to provide honest and reliable confidence scores with their reports. [paper]
- FuncCode: Compressing Kolmogorov--Arnold Networks in Function Space with Hardware-Aware Quantization This method significantly compresses KAN models by sharing codebooks across different edges in the network. [paper]
- CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents This technique reduces the cost of long-context coding agents by efficiently compacting their history without losing accuracy. [paper]
- REFLEX with Jev for Efficient Selective Control in LLM Agents This architecture uses a fast decision layer to handle simple tasks, calling a powerful LLM only when necessary to save costs. [paper]
The papers
- UPPRESSO: Untraceable and Unlinkable Privacy-PREserving Single Sign-On Services — "Single sign-on (SSO) allows a user to maintain only the credential for an identity provider (IdP) to log into multiple relying parties (RPs). [episode]
- A Standalone FPGA-based Miner for Lyra2REv2 Cryptocurrencies — This work presents "the first hardware implementation of the specific instance of Lyra2 that is used in Lyra2REv2" and "an FPGA-based hardware implementation of a standalone miner for Lyra2REv2 on a Xilinx Multi-Processor System on Chip." The authors state that "several propertie [episode]
- Amplifying Randomized Encodings & Applications — Overview This paper investigates the fundamental connection between the existence of One-Way Functions (OWFs) and the existence of "lossy reductions." The authors bridge fine-grained complexity—specifically under the Exponential Time Hypothesis (ETH)—with foundational cryptog [episode]
- TriHaRd: Higher Resilience for TEE Trusted Time — "TriHaRd: Higher Resilience for TEE Trusted Time" proposes a "TEE-based trusted time protocol with high resilience against attacks manipulating enclave-perceived clock speeds and offsets." The paper identifies that in Trusted Execution Environments (TEEs) such as Intel SGX, "the [episode]
- WiP: Towards a Secure SECP256K1 for Crypto Wallets: Hardware Architecture and Implementation — "The SECP256K1 elliptic curve algorithm is fundamental in cryptocurrency wallets for generating secure public keys from private keys, thereby ensuring the protection and ownership of blockchain-based digital assets. [episode]
- Convergence of gradient descent for deep neural networks —
- Learning to Find Proofs and Theorems by Learning to Refine Search Strategies: The Case of Loop Invariant Synthesis —
- DA-Cramming: Enhancing Cost-Effective Language Model Pretraining with Dependency Agreement Integration —
- Decidable Reasoning About Time in Finite-Domain Situation Calculus Theories —
- On Minimal Depth in Neural Networks —
- DeepSPoC: A Deep Learning Based Sequential Propagation of Chaos —
- MultiViewDx: Evidence-Linked Multi-View Clinical Diagnosis —
- ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy —
- Hierarchical Sparse Bayesian Multitask Learning for Disease Prediction in Pooled Microbiome Studies —
- FedNIA: Noise-Induced Activation Analysis for Mitigating Data Poisoning in Federated Learning —
- Attack Tree Distance: a practical examination of tree difference measurement within cyber security —
- The Challenge of Identifying the Origin of Black-Box Large Language Models —
- Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application —
- BigO(Bench): Can LLMs Generate Code with Controlled Time and Space Complexity? —
- Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors —
- Small Language Models are the Future of Agentic AI —
- Glucose-ML: A collection of longitudinal diabetes datasets for development of robust AI solutions —
- Optimizing Canaries for Privacy Auditing with Metagradient Descent —
- BrainATCL: Adaptive Temporal Brain Connectivity Learning for Functional Link Prediction and Age Estimation —
- EndoCogniAgent: Closed-Loop Agentic Reasoning with Self-Consistency Validation for Endoscopic Diagnosis —
- SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking —
- Towards Mitigating Excessive Forgetting in LLM Unlearning via Entanglement-Guidance with Proxy Constraint —
- Advanced spectral clustering for heterogeneous data in credit risk monitoring systems —
- Ultra Strong Machine Learning: LLM-Generated Explanations Do Not Yet Suffice for Teaching Humans Active Learning Strategy —
- Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs —
- CHUCKLE -- When Humans Teach AI To Learn Emotions The Easy Way —
- Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique —
- Transport-Coupled Bayesian Flows for Molecular Graph Generation —
- Provable Anytime Ensemble Sampling Algorithms in Nonlinear Contextual Bandits —
- Robust Photoplethysmography Signal Denoising via Mamba Networks —
- Simulation-free Structure Learning for Stochastic Population Dynamics —
- POPI: Personalizing LLMs via Optimized Natural Language Preference Inference —
- Fine-Tuning DialoGPT on Common Diseases in Rural Nepal for Medical Conversations —
- STAR-VAE: A Scalable Latent-Variable Transformer for Controllable Molecular Generation —
- IndirectAD: Practical Data Poisoning Attacks against Recommender Systems for Item Promotion —
- Finding Kissing Numbers with Game-theoretic Reinforcement Learning —
- CID: Measuring Feature Importance Through Counterfactual Distributions —
- Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning —
- Robust AI Security and Alignment: A Sisyphean Endeavor? —
- Pretrained battery transformer (PBT): A foundation model for battery life prediction —
- Navigating Taxonomic Expansions of Entity Sets Driven by Knowledge Bases —
- Linear probing enables Ship-Radiated Noise recognition with pretrained audio embeddings —
- CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding —
- Relative Wasserstein Angle and the Problem of the W 2-Nearest Gaussian Distribution —
- Quantum Model Parallelism for MRI-Based Classification of Alzheimer's Disease Stages —
- Discovering Data Manifold Geometry through Geometric Properties —
- Semantic Self-Distillation for Language Model Uncertainty —
- Communication-Efficient Byzantine-Robust Federated Conformal Prediction via Partial Sharing —
- Real Money, Fake Models: Deceptive Model Claims in Shadow APIs —
- Spectral Overfitting in Noisy Linear Probing of Pretrained Representations —
- Conditional Distributional Treatment Effects: Doubly Robust Estimation and Testing —
- Differential Fault Analysis of Lilliput under Random-Location Nibble Faults —
- Calibrated Confidence Expression for Radiology Report Generation —
- Learning Diagnostic Reasoning for Decision Support in Toxicology —
- RepUCB: Representation Learning-Based UCB for Heterogeneous Multi-Task Linear Bandits —
- AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents —
- Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models —
- A Survey on Long-Term Memory Security in LLM Agents: Attacks, Defenses, and Governance Across the Memory Lifecycle —
- Faithful Autoformalization via Roundtrip Verification and Repair —
- SPLICE: Latent Diffusion over JEPA Embeddings for Conformal Time-Series Inpainting —
- LLM Ghostbusters: Surgical Package Hallucination Suppression via Adaptive Unlearning —
- Event-Based Early Warning of Vineyard Disease Risk from Environmental Time Series —
- Flow Matching for Count Data —
- Practical Scaling Laws: Converting Compute into Performance in a Data-Constrained World —
- SAGE: A Self-Evolving Agentic Graph-Memory Engine for Structure-Aware Associative Memory —
- Tight Sample Complexity Bounds for Entropic Best Policy Identification —
- Small, Private Language Models as Teammates for Educational Assessment Design —
- NIMO Controller: a self-driving laboratory orchestrator based on the Model Context Protocol —
- Unbiased Gradients, Moving Stability Boundaries: Exact Mini-Batch Geometry in Linear Self-Attention —
- CAffNet: Hard Constraint-Affine Neural Networks —
- GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning —
- MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training —
- Refit the Probe: Single-Direction Ablation Is Not a Necessity Test —
- Sharp First-Order Lower Bounds for Higher-Order Smooth Nonconvex Optimization —
- Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads —
- Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems —
- Confidence Composition for Multiagent Language Model Systems —
- Learning Urban Access Costs from Origin-Destination Flows via Inverse Optimal Transport —
- When Good Verifiers Go Bad: Silent Negative Transfer in Verifier-Guided VLM Training —
- Learning task-specific subspaces via interventional post-training of speech foundation models —
- PreUnlearn: Auditing Collateral Knowledge Damage Before Large Language Model Unlearning —
- FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming —
- Evidential Fusion Network for Multimodal Survival Prediction under Missing Modalities —
- KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking —
- Explanation-Guided Medical Named Entity Recognition with Stability and Boundary Awareness for Atopic Dermatitis —
- Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation —
- SOLAR: AI-Powered Speed-of-Light Performance Analysis —
- AnySimLite: A Lightweight Few-Shot Similarity Encoder for On-Device Speech-Adjacent Classification —
- Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits —
- Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training —
- Converge to Surprise: Evolutionary Self-supervised Image Clustering —
- Low-Rank Attention Residuals —
- SPINE: Bridging the Cyber-Physical Gap with Agentic AI —
- ReasonLab: A Controlled and Auditable Evaluation of Prompting Techniques for Multiple-Choice QA —
- Simulate to Generalize: Scaling Stateful Supervision for API-calling Agents using LLM World Models —
- Reinforcement Learning for Delivery Drone-Based Participatory Sensing in Dynamic Environments —
- VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based Verification —
- DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making —
- Adaptive Confidence-weighted Expansion for Trustworthy Multi-Omics Multimodal Fusion —
- Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories —
- Orthogonal JEPA: Factorized Predictive States for Latent World Models —
- LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding —
- What Does 99% Accuracy Measure? A Reproducible Audit of Shortcut Learning in a Widely Used Fake News Corpus —
- Training a Language Model End-to-End in Rust: An Experience Report —
- Same Quantity, Different Answer: Numerical Representation Invariance in Language Models —
- Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation —
- A Computational Approach to Measuring Semantic Change in Sanskrit Literature —
- Do Existing Preconditioners Improve Biomedical Tabular Foundation Learning? An Empirical Study on TabPFN Optimization —
- Not All 4-bit Quantizers Are Equal: Deployment-Time Mitigation of PII Leakage in Fine-Tuned Small Language Models —
- "As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It —
- Retrieved-Span Training for Efficient Query-Focused Meeting Summarization on QMSum —
- From Tone to Trajectory: Continuous Sentiment and the Shape of Monetary Policy Communication —
- 4DGS-JEPA: Temporally Compositional Joint-Embedding Prediction for Dynamic Gaussian Splatting —
- What Does Chain-of-Thought Entropy Measure? A Channel Audit of Scaffolding, Routing, and Content —
- Peerify: Benchmarking Peer-Review Claim Verification —
- AIBuildAI-2.5: Efficient Autonomous AI Model Development Through LLM-Guided Tree Search —
- Teaching a Moving Student: Rethinking the Curriculum of On-Policy Distillation —
- Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibration —
- FrontierMath Erd s —
- LLM-Driven Training-free Location-Attribute Synergic Fusion: A Closed-Loop Paradigm for Dual-source Encrypted POIs and LULC Mapping —
- Self-Cleaning and Captured Anyway: One Measured Primitive for Error in a Store an Agent Writes to Itself, and What a Falling Score Actually Measures —
- LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay —
- MoM: Memory of Memory —
- ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains —
- Graph-Based Inference for Feedback-Driven Word Deduction: A Scalable Framework for the Jotto Problem —
- ChainDoRA: Tensor-Train Factorized Weight-Decomposed Low-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning —
- Understanding Reliability in LLM-based Human Behavior Simulation —
- ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch —
- Federating Quantum and Classical Computing: A Privacy-Preserving Hybrid Approach —
- CPyGraph: A Version-Aware Static Analysis Framework for Native CPython Bytecode —
- FREESIA: Covariance-Aware Posterior Transport for Expressive and Scalable Data Assimilation —
- An Accurate and Interpretable Hyper Graph Neural Network for GBM Survival Prediction —
- Impact Is Not Invalidation: Ask About the Claim, Not the Diff —
- Entropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete Flow —
- The Probabilistic Structure of Large Language Models —
- Stable Unsupervised Continual Chunking with Sheaf SyncMap —
- Variational objectives for amortized Bayesian inference in inverse problems: The role of posterior conditioning —
- Brain-Inspired Hierarchical Modularity for General Continual Learning —
- Dual-GNN Multilevel Coarsening for Maximum Independent Set —
- Exposing Blind Spots in Deep Imbalanced Regression Evaluation —
- Benchmarking Neural Defend ARCAS 1B: A Foundational Multimodal Deepfake Detection Model —
- Empirical Auditing of Edge-Private Graph Generators —
- Learning Neural Feedback Linearization for Data-driven Systems via Augmented Lagrangian —
- Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings —
- Mitigating Sequential Reappearance in Diffusion Data-Point Unlearning —
- Attack Success Rate Is Not a Number: On Measurement Validity in Agentic AI Security Evaluation —
- Multi-Term Fourier Graph Neural Network with Sample Relationship Learning for Enhanced Remaining Useful Life Prediction —
- X-Planner: Event-Structured Task Planning for Embodied Intelligence —
- FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability —
- Lean Pool: An AI-Maintained Archive of Formalized Mathematics —
- Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers —
- The AI Neuroscientist: An Interactive Agentic Interface for Neuroimaging Analysis —
- Partition-Matched Evaluation of Community Features under Distribution Shift in Android Malware Function-Call Graphs —
- MedGate-Fusion: Integrating First-Encounter Semantic Narratives and Physiological Biomarkers for Prospective Stroke Risk Stratification —
- When LLM Agents Fail to Read the Room: ReAdapt for Relational Social Reasoning —
- Attention as a Routing Graph: Live Circuit Extraction from a Single Forward Pass —
- Learned Enterprise Data Comprehension: Compression and Routing for Data Agents —
- Controller-Only False Confirmation in Passive RF UAV Link Detection —
- Correcting Within-Group Self-Selection Bias in Prioritized Replay —
- FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing —
- Making Agents More Consistent: Skills Should Form Habits for Repeat Tasks —
- Potential for Enhanced Learning in Machine Learning Classes by Using Wiki LLM Indexing —
- Topological Signal Processing With Unoriented Operators —
- Spatiotemporal Kronecker Covariance Neural Networks —
- MT-ProtBERT: Multi-task Learning ProtBERT for Intrinsically Disordered Proteins Classification with Scarce Data —
- Clarification Is Not Correction: LLMs Fail to Let Go —
- Concept Drift from a Causal Perspective —
- SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation —
- TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks —
- Quantum ROP: Using Quantum Algorithms for ROP Chain Selection in Exploit Construction —
- From Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought —
- Extending FunctionGemma for Practical On-Device Mobile Function Calling —
- Penalized Nonreversible Langevin for Constrained Sampling —
- PICPIs: Prediction-Interval-Conditional Prediction Intervals —
- Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development —
- Deep Reinforcement Learning on Item-Compatibility Graphs for One-Dimensional Bin Packing —
- Robust Failure, Conservative Repair: Textual Knowledge Distillation from Cross-Model Failures —
- Efficient Iterative Retrieval with Heterogeneous Batching —
- From Offline Proxies to Online Decisions: A Layered Engagement Evaluation Framework for Conversational AI —
- Predictive Uncertainty for Neural CAE Surrogates —
- Lightweight Ranking Heads: Accelerating Multi-Task Experimentation in Production Recommender Systems —
- PermuFormer: Multi-Task Pretraining for Permutation Representation in Algebraic Combinatorics —
- Mining Legal Arguments in U.S. Corporate Case Law —
- ZeroGate: Trust-Preserving Fast Paths for Governed AI Agent Runtimes —
- Mean Velocity Matching: Rethinking Generative Dynamics in Diffusion Models —
- Conduct Under Pressure: What Sixty Language Models Do When a User Pushes —
- Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future Directions —
- Real-Time Hand Gesture Recognition for OpenXR Using Transformer-Based Machine Learning —
- ShowTellArena: Evaluating Business Workflow Understanding from Demonstrations —
- RAG-NAROK: Retrieval-Aware Knowledge Corpus Poisoning in RAG with Source-specific Refutation —
- A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization —
- Spectra: A Rules-Driven LLM Pipeline for Automated KYC Document Processing —
- Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining —
- Learning Defensive Policies against Diverse Inference Attacks for Smart Meter Privacy —
- Identifying Suspected Mislabeled Apps in Google Play Application Removal Prediction: An Empirical Comparison of Label Noise Detection Methods —
- Queer inclusion in speech datasets: An audit and taxonomy of practical tensions —
- Towards participatory speech dataset curation: A queer case study and conceptual framework —
- Continuous Optimization for p-adic Models —
- SMTB: Fast Structure-Mapping with Tight Bounds —
- Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training —
- Matryoshka attribution: Learning to attribute language model outputs to representations and weights —
- FASTAR: FRI Accelerator for Scalable Transparent ARguments of Knowledge —
- Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference —
- A JEPA Recipe for Tabular Foundation Models —
- DefaultGNN: A Dual-Perspective GNN Framework for Predicting Corporate Default from Buyer-Seller Transaction Networks —
- Weakly Supervised Quantum Error Mitigation —
- SambaGraph: Action-Reaction Spatio-Temporal Graphs for Soccer Tactical Response Modeling —
- Recovering Agentic Sovereignty: Mitigating the Consensus Paradox via Contrastive Epistemic Decoding —
- A Behavioral Trait Leaks into Preferences: Diagnosing Trait Interference in LLM User Simulators —
- Direct Optimization of Generators for Search in Automated Theorem Proving —
- Scalable Minimum-Volume Simplex Estimation with Non-asymptotic Analysis —
- Rethinking Backdoor Repair Evaluation: Distinguishing Aggregate Clean Utility from Benign Performance Preservation —
- Gaze responses to false-positive computer-aided detection prompts during colonoscopy: a paired-video and real-time eye-tracking study —
- EMGBlend: Heterogeneity-Aware Self-Supervised Pretraining for Gesture and Force Decoding —
- Transformer Heads Looking for Order —
- Evaluating Coding Agents on Kernel Exploit Generation —
- Rewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs —
- Generalized Deep Regression for Repeated Measurements —
- ArticleMiner: Ontology-Guided Knowledge Graph Construction from Scientific Publications —
- Qwen3.8-Omni: Towards Native Omni-Modal Agents —
- Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRA —
- ChatT2: An Adaptive Framework for Developing a Large Language Model-Based Agent for Natural Product Domain Research —
- What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation —
- An Exploratory Replica-Overlap Probe of the Grokking Transition —
- SLED-IFV: Solver-Validated LLM-Guided Decomposition for Scalable Hardware Information-Flow Verification —
- Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces —
- Efficient Cost-Aware LLM Evaluation via Bayesian Bandit Gittins Indices —
- Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark —
- From Experts to Sub-experts: Fine-grained Parameter-Efficient Fine-Tuning for MoE LLMs —
- Targeted Review for AI-Assisted Biodiversity Surveys: Active Continuous-Score Occupancy Modeling —
- When Riemann flows with Wasserstein: Generative Modeling of Probability Distributions on Manifolds —
- From Utterances to Networks: Modelling Slang Adoption and Diffusion Across Subreddits —
- Marginal Log-Likelihood Increments under Dirichlet-Smoothed Markov Estimation —
- Seeing Is Not Perceiving: When Synthetic Consumers Can and Cannot Pretest Visual Marketing —
- Toolcompass: Guiding Tool Trialing, Not Suppressing It —
- C-to-Rust Fallacy: Automatic Refactoring != Memory Security —
- How Strongly Should Task State Influence an LLM Agent? —
- Graph Domain Adaptation Does Not End with Representation Learning —
- Fully Byzantine-Resilient Multi-Agent Reinforcement Learning —
- On the Gradient Heterogeneity Dynamics of Adversarially Robust Federated Regression —
- Optimal Tradeoffs Between Network Size and Parameter Magnitude in Neural Approximation and Minimax Regression —
- TCMaster: Confidence-Aware Querying and Workload-Guided Physical Design for Multi-Source Traditional Chinese Medicine Knowledge Graphs —
- LingLan: An Advancing Traditional Chinese Medicine Diagnosis LLM with Multimodal Data —
- Slow Decay and Silenced Expression: Iterated Subliminal Trait Transfer in Language-Model Lineages —
- Signed Graph Pre-Training and Prompt Learning —
- Self-Supervised Combinatorial Optimization with Constraints via Frank-Wolfe —
- GuidedRay: Diversity-Guided Direction Discovery for Targeted Hard-Label Black-Box Attacks —
- Beyond Class Marginals: Bounding Rehearsal Gaps without Freezing Class Co-occurrence —
- OmniFysics-Nano-V2 Technical Report: Understanding the Physical World Across Modalities —
- Modular Norm RandOpt: Population-Efficient Ensembling through Architecture-Aware Perturbations —
- Syndrome, Synergy, and Safety: Structured Reasoning and Knowledge-Driven Alignment for TCM Prescription Generation —
- Minimal Recurrent Behavioral Memory for Imitation under Partial Observability —
- The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance —
- Neurosymbolic Action Model Learning under Partial Observability —
- Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction —
- Disentangling Heterogeneous Traffic Dynamics for Multi-Step Traffic Forecasting via Adaptive Spectral Decomposition —
- Statistical Gains from Looped Estimation under Parameter Budgets —
- A Lightweight Plastic-Memory Framework for Graph Few-Shot Class-Incremental Learning —
- Adaptive Traffic Camouflage: Causal and Resource-Aware Defense Against IoT Fingerprinting —
- Evaluating Accuracy and Probabilistic Reliability of Zero-Shot Time Series Foundation Models —
- Reply to comments arXiv:2512.07881 and arXiv:2601.06104 on quantum structure in human and AI-generated language —
- Latest Exact Match Attention —
- The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks —
- When Are Aggregate Agent Traces Diagnosable? Traffic-Governed Interpretation and Calibrated Abstention —
- Auditing Proxy-Based Validation Across Text Spans —
- You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs —
- Multi-View Fair Clustering Guided by Cross-View Sensitive Information Discrepancy —
- CacheDyG: Decoupling Temporal Propagation for Efficient Dynamic Graph Learning —
- On the Construction of Trapdoor Claw-Free Functions with Certifiable Key —
- Protocol before progress: leakage-aware evaluation of AIS trajectory prediction —
- ARAFA: An LLM-Generated Arabic Fact-Checking Dataset —
- In-Context Guidance: Learning Inter-Task Synergies via Numerical Foundational Models for Few-Shot Multitask Optimization —
- Gaussian Flow-Matching Schedules: Implications for Sampling and Training —
- Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts —
- Prediction Is Not Detection: Evaluating Pre-Recognition Claims in Longitudinal Clinical AI —
- MemoryAthena: Adaptive Routing over Latent and Generated Memories —
- BELXTR: Biomedical Entity Linking via Contextualized Token Retrieval —
- Isolated Sign Language Recognition for Icelandic Sign Language: Experiments in a Low-resource Setting —
- AgenticSizing: A Large Language Model-based Multi-Agent Framework for Analog Circuit Sizing —
- Neural Approximation by Function Composition: Rigidity and Doubly Exponential Convergence —
- Evaluating the Effectiveness of SechKAN on 1D Data —
- COBRA: A Content-Agnostic Framework for Zero-Day Detection of Suspicious Domains —
- LoRango: It Takes Two LoRAs to Unlock Hidden Behaviors in Diffusion Models —
- Rethinking Length-Based Training: Batch Composition and Loss Normalization in Speech Token Language Models —
- How It's Made: Uncovering Detection Engineering Processes for Network Intrusion Detection Rules —
- AURA: Angular Update Rate Adaptation for training complex-valued neural networks —
- Beyond Scalar Sensitivity: Activation-Aware Mixed-Precision LLM Quantization with Cross-Layer Refinement —
- Toward Responsible AI-Augmented Cyber Defense: Pattern Recognition, Defense-in-Depth, and the Case for Human-AI Collaboration —
- Conditional Tensor Diffusion: Distributional Counterfactual Learning and Inference —
- Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models —
- Certified Against Which Oracle? Execution Labels Set the Reported Risk of Conformal Abstention for Text-to-SQL —
- ClusterFewshot: Improving Few-shot Optimization for LLMs workflow —
- CausalLoss-Fin: Attributing Financial-Agent Loss to Decisions and Infrastructure Faults —
- Exploring Solver-Level Warmstarting for Neural Network Verification —
- GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression —
- Interweaving Marginals into Multivariate Sample Paths: Training-Free Dependence Construction for Probabilistic Time Series Foundation Models —
- Theory for groupoid equivariant neural networks: an approach for steerable CNNs on bounded domains —
- VideoX-Qwen: Data-Centric Instruction-Based Video Editing —
- The Dynamics of Quasiregular Neural Learning —
- BOBA: Dynamic Bayesian Optimization through Bayesian Active Inference —
- MICRO: Multi-Fidelity Active Search for Severe Error Discovery —
- CQ4OE: A benchmark for assessing LLM-assisted ontology generation from competency questions —
- Domain-Adaptive Pretraining Enhances Water Treatment Semantic Representation for Large-Scale Structured Literature Mining —
- Truth for Believable AI: Expressed Doubt, Provenance, and Belief Revision as an Engineerable Stance —
- xWhyL: Causal Interactive Learning —
- Canonical locks that encode part-whole hierarchies —
- FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents —
- Optimizing Denoising Trajectories in dLLMs: A Lightweight Evolutionary Heuristic Approach —
- Adversarial Course-of-Action Generation: Game-Theoretic Multi-Agent Algorithms for COA matching & COA generation —
- ChainUQ: Reasoning Consistency-Aware Uncertainty Quantification for Large Language Models —
- Towards Adaptive Federated Graph Clustering: A Global Community-aware Contrastive Learning-based Approach —
- FuncCode: Compressing Kolmogorov--Arnold Networks in Function Space with Hardware-Aware Quantization —
- RankCert: When Can Simulated Learners Safely Select an AI Tutor? Robust Decision Certification Under Structural Uncertainty —
- Policy-Backed Selective Regeneration under Tainted Inter-Agent Communication —
- Selection-Invariant Communication Compilers for Privacy-Aware Multi-Agent LLM Workflows —
- Fast Matrix Multiplication in fp8: Certified Coefficient Optimization and Measured Error —
- Learning to Link: Automatic Re-identification of BLE Devices Under MAC Address Randomisation —
- Margin-Drop Coordinates for Cross-Budget Robustness Evaluation —
- CoVeR: Coverage-Based Routing of Verifier Calls in Agentic Retrieval —
- The Architect, the Adversary, and the Judge: Closed-Loop Generation of Standards-Aligned Assessment Items at Scale —
- SpecialEduBench: Benchmarking Vision-Language Models on Knowledge, Skill, and Attitude in Language Intervention for Autistic Children —
- From Bilinear to Linear: Differentially Private Federated LoRA via Low-Dimensional Parameterization —
- CoEvo: Oracle-Grounded Self-Evolution of a Single Model for Multi-Step Causal Reasoning —
- FusionMMT: A Unified Multimodal and Multitask Learning Framework for Nuclear Fusion —
- Beyond Classification Accuracy: Quantifying Fingerprint Complexity in Encrypted Darknet Services —
- One Domain, Many Tongues: Composing Domain and Language LoRAs for Cross-Lingual Remote-Sensing MLLMs without Paired Data —
- TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models —
- Neoadjuvant chemotherapy response prediction using pretreatment diffusion and contrast-enhanced magnetic resonance imaging with clinical variables —
- Early Prediction of Pathological Complete Response to Neoadjuvant Chemotherapy Using Temporal Deep Learning on DWI —
- Certified Mechanistic Interpretability: Lifting Single-Input Findings to Bounded Neighbourhoods —
- Differentiable Fuzzy Inference Layer: A Monotone, Compositional Ordinal Reasoning Head for Large Language Models —
- A Cross-Dataset based Zero-Day Intrusion Detection System by Integrating Siamese Network and Reinforcement Learning —
- DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents —
- FairMon: A Tool for Monitoring and Visualizing Algorithmic Fairness —
- MAC-RRG: Iterative Multi-Agent Collaboration for X-ray Radiology Report Generation —
- When Big Data Becomes a Curse: Spatial Heterogeneity and the Limits of Learning from Passive Acoustic Monitoring Data —
- The Cost of Conservation: Coordination-Memory Laws for Exact-Support Generation —
- Zeta-Transform Evaluation for Higher-Order Vanishing Key Recovery —
- VACS: Value-Aligned Compositional Shielding for Multi-Agent Reasoning —
- When Verifiers Vote Backwards under Verdict Substitution: Signed Pivotal Value in Correlated Self-Consistency —
- Unanimity Without Persuasion: A Single Round of Debate Erases the Disagreement That Verification Needs —
- From Risk Scoring to Risk Allocation: A Density-Driven Framework for Diverse Monitoring in Multi-Agent Systems —
- Block-Level Weight-Space Structure Persists Under Post-Training: An Empirical Study Across LLM Families —
- Toward User-Mediated Self-Repair in Ubiquitous Robots Through Goal-Oriented Agentic AI —
- The Free-Recipe Limit: Every Recipe Effect Measures Which Premise of an Idealised Learner Broke —
- Spectral Tail Interventions in Decoder-Only Language Models: Reasoning-Sensitive Weight Structure from Controlled Surgery —
- Activation-Energy Pruning for Spiking Neural Networks: Unsupervised Personalization via Spike-Count Saliency —
- Component Type, Not Reconstruction Error, Predicts Attention Quantization Sensitivity —
- The Uncontrolled Variable: Vision-Language Refusal Is Conditioned on the Image-Attachment Interface, and Not Robust to Irrelevant Image Properties —
- EADC: Evaluation of Advanced and Deep-level Compliance in Large Language Models —
- Refusing Everything Looks Safe: Restoring the Benign Arm to Encoded-Prompt Evaluation —
- Magnitude Profile Pruning: Calibration-Free Structured Attention Head Removal for Transformer Compression —
- Unread or Unenforced? Separating Representation from Enforcement Failure in Content Guards —
- Modality-Gated Deep Adapters: Adding a Modality to a Frozen Embedding Model with Exact Preservation —
- Dynamic Deep Prompt Optimization for Defending Against Jailbreak Attacks on LLMs —
- TREND-10K: A Comprehensive Dataset for Next-Generation Video Quality Assessment Based on Preference-Driven Media —
- Identifying Intelligent Processes via Online Sequential Testing —
- Partially Observed Sparse Graphs: The Unknown Sampling Rate is a Tail Index —
- RCShift: Certifying When Partial Linkage Suffices for Finite-Sample Decisions —
- Beyond Static Charts: Can Language and Vision Language Models Generate Interactive Data Visualization Interfaces? —
- Same Chart, Different Story: Bias in Vision-Language Chart Interpretation —
- Improved Multiplayer Bandit Algorithm for Bernoulli Rewards —
- Bridging the Data Gap: Digital Twin as a New Paradigm for AI-based Radio Sensing —
- Beyond Imitation: Auditing the Recoverability of Reasoning in Distilled Models —
- MSA-CITE: A Co-Adapted LoRA Specialist Ecology for Fixed-Budget Small-Model Inference —
- Who Assures the Verifier? An Executable Assurance-Locus Audit of the European Digital Identity Wallet —
- Quantum-Ready Secure WAN: A Risk Assessment and Migration Framework —
- High-Order Liquid Evidence Modeling for Continuous and Subtle GNSS Spoofing Detection in Autonomous Driving —
- Towards Effective Black-Box Adversarial Attacks on Deep Code Models via Structural and Identifier Perturbations —
- Damage Predicts Recovery: When Calibration Data Matters in Compressing Financial LLMs —
- Can You Delete a Year of Market Data? Machine Unlearning Against Exact Retraining Oracles —
- A Hybrid AI Framework for Academic Advising: Integrating Ensemble-Based Grade Prediction and a Rule-Based Expert System —
- A Multi-Timestep LSTM Ensemble regressor for Enhanced Short-Term Runoff Prediction —
- PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models —
- eBPF Security in the Wild: Structural Concentration, Failure Mechanisms, and Discovery Gaps —
- Information-Theoretic Decoupled Prompt Tuning for Continual Learning —
- CRT-Decomposed-Protocols for CSIDH —
- Coding Agents are Strong Prompt Optimizers —
- Decoupling Is Not Identification: Supervised Evidential Learning in Next-Token Prediction —
- Mode Collapse Is Cheap to Detect: A Ground-Truth-Free Pre-Flight Check for Neural Samplers —
- JAMPR+/L2D: scalable neural heuristic for constrained vehicle routing problems in dynamic environment —
- FISSION: Label Augmentation for Bot Detection —
- On the Effect of Bit-Level Parameter Perturbations in Machine Learning and Deep Learning Models —
- Image-Based Techniques and Ensemble Soft Voting for Malware Classification —
- Neural Fingerprints for Malware Analysis: An Image-Based Metric Learning Approach with Application to Cross-Domain Classification —
- A Throughput-Oriented Analytical Model for Post-Quantum Security Protocols —
- Quantifying Protocol-Induced Uncertainty in Comparative Predictive-Model Evaluation: Evidence from Large-Scale Daily PM10 Forecasting —
- Learning to Fluctuate: Statistical Foundations for Causal Tabular Pretraining —
- Dual-Frontier: When Can an Agent Trust Its World Model? —
- On the security and privacy of LLMs in Mobility —
- CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference —
- Staged Multi-step UTXO Workflows via Recursive Invariants —
- CHiME-9 ECHI: A Machine Learning Challenge for Enhancing Conversations to Address Hearing Impairment —
- PreGS: A Parameter-Transfer-Based Multi-Expert Graph Neural Network for Node Classification —
- Design and Evaluation of a Controlled Post-Alert Incident Orchestration and Response Subsystem Using a Rule Engine and a Local Large Language Model —
- Error Bounds for Statistical Estimators in BTL Model with Parametric Multivariate Utility Functions —
- Disaggregated Quantization: Specializing LLM Prefill and Decode —
- Designing and Analysing Argument Mining Pipelines: Towards a Comprehensive Assessment —
- Geometry-Aware Hyperbolic Residual-Quantized Variational Autoencoders —
- Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament —
- TransBERT: A Framework for Synthetic Translation in Domain-Specific Language Modeling —
- HYDRA: Proactive Android Malware Drift Adaptation via Hierarchical Graph Contrastive Learning —
- PACT: From Credit Assignment to Critic Alignment —
- Formally Modeling the Terrapin Attack on SSH —
- HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing —
- FairMean: Promoting Fairness in Distributed Learning under Label Poisoning Attacks —
- Layout-Guided Masking for GROBID: Lightweight Structural Gains in Large-Scale Scientific PDF Ingestion —
- Learning to Defer with Guidance on Real World Medical Data —
- TimeInteract: Towards Real-Time Interactive Intelligence for Streaming Time Series —
- Double Descent and Malign Overfitting in Diffusion Models —
- Combining Hierarchical Cognitive Process with Process Supervision for Interpretable Scene Safety Understanding —
- OMatG-flash: An All-Atom Flow Map with Reinforce Adjoint Matching for Scalable Materials Discovery —
- SuperPCA: subspace analysis and an efficient algorithm for high-dimensional PCA —
- Reliability Theory for AI Control —
- Enriching Speech Emotion Representations with Conversational Context —
- DeepFEAv2: Deep Learning for Transient Finite Element Analysis Beyond Structured Meshes —
- The Source of Disturbance Matters: External, Internal, and Control-Generated Noise in Adaptive Regulation —
- One-Step Generative Surrogate Models via Block-Triangular Joint Drifting —
- A Practical Guide on Graphical Model Validation —
- Recursive self-improvement of AI research agents —
- Can We Predict Anomaly Detection Performance from Embedding-Space Geometry? —
- Reproducible AI Requires Reproducible Randomness —
- How to Estimate Whether You Have Found Several Needles in a Haystack: Measuring Calibration in Multi-Label Text Classification —
- When Recursive Models Finish Computing —
- Spoken Language Models that Think Aloud —
- Calibration as a First-Class Criterion in LLM Evaluation —
- Gap-Free Streaming PCA Beyond Rank-One Updates: Near-Optimal Rates and Applications to Differential Privacy —
- A Semiotics-Aware Framework for Evaluating Fidelity and Coverage in Natural Language Generation —
- REFLEX with Jev for Efficient Selective Control in LLM Agents —
- Transcribe, Translate, and Optimize: Joint Reward Learning for Speech Translation —
- Notes on Fourier-Bessel wavelets —
- A retrospective analysis on the use of LLMs to study infant syntax learning —
- JEV-as-a-Judge: Accept When Confident, Escalate When Unsure —
- Rouxii: Exploiting Honeypots with Deception-Aware AI Pentesters —
- Neutral-Atom-based Quantum Optimization for Resource Allocation in NOMA Networks —
- Quantum-Aided Active Device Detection in Energy-Harvesting Symbiotic Radio Networks —
- Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models —
- Hierarchical GNNs for power flow: letting physics shape the hierarchy —
- Semantic Abstraction for Natural Language Inference: a Methodological Framework for Discovering and Compensating Semantic Knowledge and Reasoning Gaps in Large Language Models —
- Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference —
- On Basis Function Selection for Sparse Gaussian Process Regression —
- PERSONAWEAVER: Controllable Diversity Beyond Conventional Archetypes in Procedural Character Generation —
- Label-Efficient Learning for Ground-Based Sky-Image Classification: A Benchmark of Transfer Learning, Active Learning, and Pseudo-Labeling on GCD —
- Knowledge Pull Requests for Continual Document Authoring —
- Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models —
- Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding —
- The Delegation Blind Spot: Auditing Product Decisions from Agent Choices —
- MAGIC: Mixed-Granularity Agent Graphs via Incremental Construction with Dense-Reward Reinforcement Learning —
- A Spectral Theory of Grokking: Weight Decay induces Feature Learning —
- Decoding the Legalese: A Scalable and Quantitative Framework for Analyzing Corporate Privacy Policies —
- From Alignment to Access Control: A Framework for GenAI Policy Enforcement —
- Detecting GPT-Assisted Writing Using Interpretable Stylometric Features —
- Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation —
- Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning —
- Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning —
- The Sirens' Song: When Proximal Background Context Overshadows Distant Evidence —
- EquivSVA: A Formally Verified Dataset of Behavioral Assertions Across Equivalent RTL Implementations —
- Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It —
- Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents —
- A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem —
- SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving —
- CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents —
- SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue —
- Agensh: Scaling Organizational Intelligence to 1,024 Agents —
- A Lyra2 FPGA Core for Lyra2REv2-Based Cryptocurrencies —
- Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs —
- Sparse Reduced-Rank Regression for Simultaneous Rank and Variable Selection via Manifold Optimization —
- Stable Marriage Problems with Ties and Incomplete Preferences: An Empirical Comparison of ASP, SAT, ILP, CP, and Local Search Methods —
- Layer Adaptive Node Selection in Bayesian Neural Networks: Statistical Guarantees and Implementation Details —
- Analysis of Regularized Learning in Banach Spaces for Linear-functional Data —
Important terms
- Dual-Frontier
- A principle used to solve failure-attribution problems in AI. It helps determine if a mistake happened because of a bad decision rule or an inaccurate world model by only accepting decisions that exceed a certified error bound.
- SAGE
- A self-evolving graph memory engine designed for long-horizon tasks. It uses a reader-writer framework to build and refine structured memory from interaction histories, improving performance in multi-hop question answering and retrieval efficiency.
- VeriSimpl
- A verification method that uses simplification to ensure optimization models derived from natural language remain accurate when they are translated into formal mathematical or logical structures.
- Modular Norm RandOpt
- An improvement on standard weight-perturbed sampling. It uses architecture-aware, module-wise natural norms to select ensemble candidates more efficiently, achieving higher accuracy while requiring significantly fewer candidates than previous methods.
- REFLEX
- A control mechanism for autonomous reasoning architectures. It uses a judge (JEV) to enable selective control during agentic workflows, deciding whether to accept confident outputs or escalate uncertain ones for further review.