AI papers — 2026-09-18
Today is mostly about the danger of inherited defaults, specifically how tiny, invisible choices made during data preprocessing can quietly sabotage the signals we are trying to study. This is most urgent in legal research, where scholars treat judicial text as data to uncover ideological or doctrinal shifts.
Many researchers still rely on removing stopwords, such as "the" or "and," based on mid-century information retrieval habits that were never tested for accuracy in a legal context. Researchers ran an exhaustive test by removing roughly 18,500 individual words one by one from over 14,000 Supreme Court opinions.
They found that using standard stopword lists actually performs worse than removing nothing at all. Even when they tried to build the best possible custom list of words to delete, the results were statistically indistinguishable from leaving the text untouched.
This suggests that cleaning text might be distorting the legal signals researchers want to recover, making the decision to remove stopwords a fundamental question of measurement validity. The problem of hidden, distorting information also appears in physical systems.
A machine might appear to be working perfectly while actually losing its ability to perform a future, yet-to-be-defined task. This tension is addressed by new crossover benchmarks that debate whether to prompt a frozen large language model or invest in custom training for tabular data.
These benchmarks quantify the exact point where a classical model's learning curve overtakes an LLM's zero-shot performance. In 86% of cases across eighteen datasets, a trained classical model beats a small frozen LLM using no more labeled data than was already on hand, with the median crossover occurring at just 6% of the training set.
This suggests that for most business applications, collecting a few hundred labels to train a gradient-boosted model is more efficient than relying on an LLM's semantic understanding of feature names. Similarly, in the domain of neural prosthetics, researchers found that zero-shot transfer for surface-EMG gesture decoding fails entirely.
However, providing just three labeled repetitions allows a cross-user encoder to exceed the performance of a standard per-user classifier by 0.190 macro F1. This performance gap is further bridged when the training pool is enriched with amputee data rather than just intact subjects.
While these specialized models require specific calibration, other architectures are finding ways to bypass expensive online reasoning entirely. The VisKG-LM framework demonstrates that knowledge graphs can be compiled offline into a visual memory of relation-labeled paths.
By treating these as cached images for a vision-language model to consult, researchers achieved significant gains in question answering over traditional methods that re-encode subgraphs during every inference step. This transition from theoretical modeling to real-world deployment introduces significant risks regarding fairness and reliability.
In educational technology, a cross-architecture audit of deep knowledge tracing models reveals that accuracy often comes at the cost of equity. Researchers evaluated four architectures—DKT, DKVMN, SAKT, and AKT—across two datasets to see how they handle demographic metadata.
They found that bias is a persistent reality, as every architecture showed significant socioeconomic bias on the Eedi dataset with lower AUC scores for economically disadvantaged students. Most strikingly, the most accurate model, AKT, which gains roughly four AUC points through item-level Rasch embeddings, also exhibited the largest socioeconomic bias.
Attempts to mitigate this through reweighting or adversarial training proved unreliable, as these methods failed to change the ABROCA metric in any configuration that maintained accuracy. Similarly, in physical control systems, a new framework for task readiness addresses the danger of dormant dynamics changes.
An actuator might function perfectly for a current task but fail when a new task is assigned because its degradation was never excited by previous operations. By using an intervention-based Bayesian procedure called Evidence-Gated Matched-Pulse Transport, agents can now diagnose local dynamics changes and provide either a recovered policy or an abstention decision.
This effectively trades off performance for safety and readiness certification. We finally have a way to tell if an agent is actually getting smarter at making decisions or if it is just learning how to stumble into better starting positions.
This is vital because when we train agents in closed loops, their own actions shape the environments they encounter, making it hard to know if a performance boost comes from better reasoning or simply reaching easier states. A new protocol called checkpoint handoff solves this by cloning a specific state reached by one policy and handing it directly to another.
By splitting gains into reach, which is how often an agent arrives at a successful state, and solve, which is how often it finishes once it gets there, researchers found that reinforcement learning improves both. On the ALFWorld benchmark, the data shows that an RL-trained solver is consistently more capable than an SFT-trained one when starting from the exact same point.
This ability to refine how an agent interacts with its environment is echoed in work on optimizing external skills through structured graphs. A framework called SkillAA uses an attribution-guided approach to repair failed procedures by routing errors to specific locations in a unified skill graph.
By using local and big gates to validate changes, it achieved high accuracy across SearchQA, LiveMath, and DocVQA. While SkillAA focuses on repairing external tools, another approach turns the verification process itself into a generator of better ideas.
Instead of just using a verifier to pick the best option from a fixed pool, the Verify-Repair-Reselect method uses feedback to build entirely new, improved candidates. This can actually recover correct answers even when every single initial attempt was wrong.
The question of whether AI agents truly understand the systems they manipulate remains a central tension in hardware design. Using a framework called AutoTuring, researchers tested whether agents designing accelerators were actually reasoning about architecture or simply performing a sophisticated search over numerical knobs.
By presenting the same fifteen-dimensional accelerator space twice—once with meaningful architectural names and once as anonymous variables on a scale from zero to one—they could measure the value of semantic meaning. On a nine-kernel FP16 GEMM task, the results showed that meaning does matter, as the architect agent outperformed its blind counterpart by 12.3 percent on average and required 70.1 percent fewer simulator calls.
However, this advantage is not absolute, as a critic loop was able to recover most of the performance gap for the blind agent without providing any additional benefit to the architect. This suggests that structured critique and architectural knowledge may act as substitutes rather than complements.
This ambiguity extends into the financial sector, where the high-stakes nature of autonomous trading agents introduces unique security risks. Through the FARSIGHT framework, which evaluates agents against market turbulence and various attack vectors, researchers analyzed fifteen academic trading schemes.
They found that 80 percent of these schemes failed at least one core robustness metric, and 100 percent exhibited security vulnerabilities. These failures are deeply interconnected, as a minor misjudgment can cascade into a market-wide crash, a vulnerability that an adversary could exploit to trigger a collapse at minimal cost.
The shift toward more sophisticated agentic evaluation is perhaps most clearly seen in the introduction of checkpoint handoff, a protocol designed to disentangle whether reinforcement learning gains stem from an agent's ability to reach better states or its ability to solve problems once it arrives there. By cloning a state reached by a released checkpoint and handing it to another policy without retraining, researchers can separate these two effects into REACH and SOLVE components.
Across two benchmarks and two independent pipelines, the data shows that the interaction between a reacher and a solver is consistently positive, and an RL history provides more value to an RL solver than to an SFT solver. On ALFWorld, RL improves both metrics, and notably, the SFT solver never succeeds in instances where the RL solver fails.
This granular view of agentic performance is complemented by work on the structural components of coding harnesses, which suggests that the effectiveness of an agent is heavily mediated by its environment. In studies across SWE-Bench Verified and Terminal-Bench 2.1, researchers found that context management becomes critical as budgets tighten, primarily by preventing context-overflow failures.
They also found that rule-based elision outperforms LLM-based summarization for efficiency. Furthermore, while planning acts as an accuracy scaffold for weaker models, it serves more as a cost-saving measure for stronger ones, and the utility of predefined tools depends heavily on the model's native bash proficiency.
The push toward more autonomous systems is manifesting in how we handle complex decision-making and data curation. In the realm of navigation, researchers have moved beyond simple heuristic costs by introducing a deep architecture that enables differentiable shortest-path search.
By running a multi-objective Dijkstra algorithm offline to establish a Pareto optimal candidate set, they designed a neural network capable of jointly optimizing cost functions and route ranking. This allows for highly customizable routes based on specific user preferences, outperforming existing methods in both quality and adaptability.
Similarly, the challenge of data engineering is being met by AutoData, an agentic search tool that treats pre-training data selection as a problem of heuristic engineering. Rather than just adjusting weights over fixed domains, AutoData searches a program space of scoring and stratification rules to discover optimal selection algorithms through iterative refinement.
This approach successfully transferred from small proxy models to larger scales, improving the downstream CORE metric. Even in highly specialized domains like legal reasoning, automation is advancing through a new four-stage framework called S4L that enables the translation of natural-language traffic rules into executable Prolog code.
By performing semantic role extraction and scene completion within a single guided prompt, S4L achieved a 75 percent accuracy rate in formalizing rules, significantly surpassing standard natural-language or logical English baselines. The most significant breakthrough involves making the process of steering large language models automatic rather than manual, which is vital for scaling safety and control.
A new framework called Deep Noir uses architectural timing and causal attribution to find the exact points in a model where you should intervene to change its behavior. In tests across various model sizes, this approach improved spam detection by as much as 42 percentage points in larger models and boosted sentiment accuracy by 13.1 percentage points without requiring any code changes.
However, this newfound control comes with a trade-off, as the study found that steering actually creates a predictable vulnerability to prompt-injection attacks that gets worse the more you steer. This tension between control and vulnerability is mirrored in the way we evaluate model safety through fine-grained signals.
Researchers have found that when you strip away a model's general sense of harmfulness to look at specific categories, like hate speech or violence, those specific signals behave differently across different models. Interestingly, even when these category-specific signals are mathematically orthogonal to general harm at one layer, they still end up amplifying the model's overall harmfulness in later layers.
While we struggle to control what models do, we are also struggling to define how we measure their success. A massive mapping of nearly 15,000 papers shows that as benchmarks evolve, there is a growing shift toward testing how models interact and perform professional tasks.
This raises a looming question about whether using AI to help build and score these very same tests might just create a loop that reinforces the existing biases of the models themselves. The difficulty of setting clear goals extends to specialized agents, such as those tasked with fixing broken travel itineraries.
In a comparison of different strategies, researchers found that while full replanning is effective for complex disruptions, hierarchical repair methods are much better at keeping an itinerary stable and preserving the parts a traveler already accepted. This highlights the constant struggle in agent design between finding a perfect new solution and sticking to what already works.
This need for precision is also evident in how we help agents troubleshoot technical problems. A new framework called RAFT moves away from treating support cases as static documents, instead treating them as evolving timelines of events.
By retrieving specific moments in a historical case that match the current problem, it outperforms standard retrieval methods at every stage of a troubleshooting process. Even the way agents learn to code is being refined to avoid wasting expensive computing resources.
A method called SIFT uses a lightweight tree search and an AI judge to quickly rank potential code improvements, reserving the most expensive evaluations only for the most promising candidates. This allows coding agents to improve themselves much faster and more cheaply than previous methods.
In the realm of causal discovery, researchers are trying to prevent models from accidentally skipping over important relationships between variables. A multi-agent framework called MaSCoD attempts to solve this by organizing structural patterns before the model makes its final judgments.
While this approach helps retain more relevant connections in certain datasets, it doesn't work perfectly across the board and can sometimes increase false positives. Finally, there is a growing concern regarding the underlying security of decentralized systems.
A new study mapping the landscape of maximal extractable value attacks shows that almost every production DAG-based BFT protocol is vulnerable to some form of transaction manipulation. The research suggests that whether an attack succeeds depends more on how the protocol was designed from the start than on how much effort an attacker puts in.
Today's papers
- Stop Removing Stopwords: How an Inherited Preprocessing Default Distorts Legal Text-as-Data Removing common stopwords can actually damage the accuracy and interpretability of legal text analysis. [paper]
- Diagnose, Recover, Certify: Task Readiness under Hidden Dynamics Changes This framework helps autonomous systems detect when their physical components are failing and decide if they are still capable of performing new tasks. [paper]
- AI-Driven Real-Time Relay Optimisation in Smart Urban NR-V2X Networks via Learning-to-Optimise Graph Neural Networks A new graph neural network approach enables vehicles in smart cities to select the best communication relays in real time. [paper]
- JointMatch: A Unified Heterogeneous Graph Neural Solver for Large-Scale Ride-Sharing Matching This system uses a single graph neural network to simultaneously pair riders and assign vehicles, making ride-sharing much faster and more profitable. [paper]
- SAGG: Sample-Adaptive Gradient Gating for Robust Multimodal Learning under Heterogeneous Corruption This method improves AI training by selectively ignoring corrupted data on a per-sample basis rather than applying a blanket fix to entire batches. [paper]
- Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes This technique teaches language models to improve their reasoning by turning their own failed attempts into training data for self-correction. [paper]
- The Public Discourse Corpus (PDC): A Speaker-Attributed Dataset for Valence and Epistemic Modality with Target Speaker Participation Researchers have released a new large-scale dataset of public interviews annotated for emotional tone and certainty of speech. [paper]
- Fine-Tuning Models for Biomedical Relation Extraction Small language models can be highly effective at extracting specific biological relationships from medical literature if they are properly fine-tuned. [paper]
- Subliminal Prompting Beyond Static Geometry: Causal Depth and Multi-Token Confounds This study investigates how hidden traits are passed through language model outputs, finding that the mechanism is more complex than simple word associations. [paper]
- SoftTri: Smooth Triangular Membership Functions for Adaptive Fuzzy Inference Systems This paper introduces a new way to make fuzzy logic systems easier to train using standard gradient-based optimization. [paper]
- Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations You can discover the different stylistic dimensions used by large language models simply by sampling their outputs and analyzing them with math. [paper]
- Small Enough to Know Everything: The Fully-Enumerable Transformer as an Instrument for the Science of Delayed Generalization Tiny, fully searchable transformers can be used as perfect scientific tools to study how AI models suddenly learn complex tasks. [paper]
- Improving Generalization and Robustness in Offline Reinforcement Learning via Boundary-Aware Data Augmentation This method improves how AI learns from fixed datasets by creating synthetic data that stays close to the original boundaries of the real world. [paper]
- Lens: Bringing the Right Semantic Perspective into Focus for Training-Free Multimodal Representation Learning This framework allows AI to extract information from images and text based on a specific task's needs without needing extra training. [paper]
- Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry A new testing method shows that even advanced AI struggles with the complex, high-level structural rules of classical Chinese poetry. [paper]
- Enhanced Agriculture-informed Neural Network by Domain Knowledge This hybrid model combines deep learning with actual agricultural science to predict nitrogen emissions more accurately and reliably. [paper]
- Do AI Agents Understand Computer Architecture? This study tests whether AI designers actually understand hardware or are just guessing by seeing if they perform better when the components have meaningful names. [paper]
- FedeRICo: Federated Region-Influenced Coupling for Traffic Flow Prediction This framework allows different cities to collaborate on traffic forecasting without sharing private data while still accounting for how traffic flows across borders. [paper]
- An Analysis of Training-Free Self-Reported Confidence in Language Models While language models can tell you how confident they are, their self-reports are highly sensitive to how you ask the question and can sometimes be wrong. [paper]
- From Parameters to Behaviors: A Survey of Model Fusion for Large Language Models This survey provides a comprehensive map of how researchers combine different AI models into single, more powerful ones. [paper]
- PAA: The Probabilistic Allen Algebra: A Generative and Complete Probabilistic Extension of Allen's Interval Relations This new mathematical system allows computers to handle uncertain timing and duration in language more naturally. [paper]
- SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes An investigation into AI trading agents reveals that most current academic models are highly vulnerable to market crashes and security attacks. [paper]
- CARE-VI: Conservative Adaptive Reliability Estimation for Value Improvement in Off-Policy Actor-Critic Learning This method makes reinforcement learning more stable by carefully checking the reliability of new potential actions before using them. [paper]
- Radio Frequency Detection and Classification of Microplastics in Water This research uses radio frequency technology and machine learning to quickly identify tiny plastic particles in water without using chemicals. [paper]
- Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs This evaluation method separates whether an AI agent improved because it reached better situations or because it actually got better at solving them. [paper]
- SkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback This framework helps AI agents improve their skills by pinpointing exactly where they failed in a structured knowledge graph.
- LLM-as-an-Improver: Turning Verification into Better Candidates Instead of just using a verifier to pick the best answer, this method uses feedback to actually fix and improve the initial guesses. [paper]
- Constraint-Safe Graph-Context Scoring for Stable Point-Feature Labels Under Text-Width and Accessibility-Inspired Profiles This system ensures that map labels are easy to read, don't overlap, and stay stable when you move around a digital map. [paper]
- Delphi Scanner: efficient and interpretable static malware detection via API sequence modeling This lightweight tool detects Windows malware by looking at the patterns of system calls it makes, making it both fast and easy to understand. [paper]
- SCGFM-ART: Amortized Relational Transport for Structure-Centric Graph Foundation Models This framework allows AI to quickly transfer knowledge between very different types of networks by mapping them to a universal structural guide. [paper]
- One Intervention per Component is Enough: Towards Identifiability in Linear Stochastic Dynamics from Steady State This mathematical study shows that you only need one targeted change per system component to fully understand how a complex process works. [paper]
- RISC-V and machine learning: a survey This survey explores how the open-source RISC-V chip architecture is being used to create custom, efficient hardware for artificial intelligence. [paper]
- An Empirical Study of Harness Design for Coding Agents This study breaks down which parts of an AI coding assistant's environment—like planning or memory—are most important for success. [paper]
- Customizable and Jointly Optimized Route Planning: A Deep Architecture Enabling Differentiable Shortest-Path Search This new neural network can learn to plan routes that match specific user preferences while still finding the mathematically best path. [paper]
- Structured Four-Stage Legal Translation: From Natural-Language Traffic Rules to PROLOG This framework turns messy, human-written traffic laws into precise computer code that can be used for automated reasoning. [paper]
- AutoData: Agentic Search for Pre-training Data Selection This AI agent automatically searches for the best ways to select and organize data to train other large language models. [paper]
- Dictionary-Constrained Grapheme-to-Phoneme for Unsegmented Languages from LLM-Annotated Data This method uses large language models to help computers learn how to pronounce words in languages that are difficult to segment, like Japanese. [paper]
- Full-Duplex Speech Models Take the Floor When Asked, Not When Needed This study finds that current "always-listening" speech models are good at responding when spoken to but bad at knowing when to interrupt for safety or corrections. [paper]
- Sharpness-Aware Minimization (SAM) Improves Classification Accuracy of Bacterial Raman Spectral Data Enabling Portable Diagnostics Using a specific optimization technique helps AI more accurately identify bacteria from light signals, aiding portable medical tests. [paper]
- A Learning Algorithm for Threshold Boolean Networks with Prescribed Fixed Points This new algorithm can reconstruct complex biological networks while ensuring the model behaves exactly how scientists expect it to. [paper]
- Evaluating Explanation Methods by the Predictors They Induce This test checks if an AI's explanation is actually truthful by seeing if you can rebuild the model's decisions using only that explanation. [paper]
- FreqCondNorm: Towards Cross-domain Predictive maintenance through a Frequency-Conditioned Transformer Foundation Model This model helps predict when machines will break by using a special layer that handles different data sampling speeds. [paper]
- Radio Frequency Convolutional Neural Networks This research shows how to use existing wireless radio hardware to run AI, making it much more energy-efficient for small devices. [paper]
- Federated Soft Clustering via Generalized Total Variation Minimization This method allows multiple devices to group similar data together through a shared network without ever seeing each other's private information. [paper]
- To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals This system speeds up AI text generation by using internal signals to decide whether the model should repeat itself or think of something new. [paper]
- On the Leakage of Massey Secret Sharing Schemes under Linear Computations This study shows how attackers can exploit mathematical relationships between different pieces of a secret to steal the whole thing. [paper]
- V= a kQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering Researchers have created a new benchmark to test how well AI can answer questions based on spoken Telugu. [paper]
- Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits This new way of running online experiments focuses on the relative difference between options rather than their absolute success rates, making it much more efficient. [paper]
- MAPLE: Metadata Augmented Private Language Evolution This method makes it cheaper and faster to create private training data for AI by using extra information to guide the process. [paper]
- Model-Aware Data Cleaning for Tabular Foundation Models This approach uses reinforcement learning to teach an AI how to clean up messy spreadsheets so they are better suited for large-scale models. [paper]
- Steering the Compass: Aligning Dynamic Psychological Counseling Conversations with Cognitive Behavioral Therapy Strategies This new dataset helps train AI counselors to use specific, proven psychological strategies during therapy sessions. [paper]
- MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs This multi-agent system ensures that AI-generated code is mathematically proven to be safe before it is ever executed. [paper]
- Contrastive Privacy: A Semantic Approach to Measuring Privacy of AI-based Sanitization This new testing method uses semantic similarity to see if "sanitized" images or text still accidentally reveal private information. [paper]
- Weather Data Spoofing Attacks on Rain-Adaptive Millimeter-Wave Frequency Selection in V2X Communication Networks This study shows that an attacker could trick self-driving cars into using the wrong radio frequency just by faking weather data. [paper]
- A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents This research finds that giving AI agents reasoning capabilities doesn't make them immune to digital nudges; it just changes how they are influenced. [paper]
- DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression This highly efficient model uses new compression techniques to handle massive amounts of text while using much less memory and power. [paper]
- Competition, Collusion, and Corruption: The Spectrum of MEV Attacks on DAG-Based BFT Consensus Protocols This paper categorizes different ways attackers can manipulate transaction orders in modern blockchain networks to make a profit. [paper]
- MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation This framework helps AI discover cause-and-effect relationships by organizing structural information before making individual judgments. [paper]
- What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks This study analyzes thousands of papers to show how the goals and methods for testing AI are rapidly evolving toward more interactive tasks. [paper]
- Self Improvement via Fast Tree-search This framework allows coding AIs to improve themselves much faster by using a smart "judge" to quickly filter out bad ideas during their search process.of course! Here is the list of papers with their summaries: [paper]
The papers
- How do I update my model? On the resilience of Predictive Process Monitoring models to change —
- Stochastic tensor space feature theory with applications to robust machine learning —
- Explainable Predictive Process Monitoring: A User Evaluation —
- BlockEmulator: An Emulator Enabling to Test Blockchain Sharding Protocols —
- ResNLS: An Improved Model for Stock Price Forecasting —
- Large Language Models for the Automated Analysis of Optimization Algorithms —
- Estimation of multiple mean vectors in high dimension —
- When fairness metrics fail: A utility-based perspective on epsilon-fairness —
- A Computational Tropical Geometry Framework for Neural Networks —
- On-line Anomaly Detection and Qualification of Random Bit Streams —
- SHIRE: Enhancing Sample Efficiency using Human Intuition in REinforcement Learning —
- Information-Geometric Inverse Distillation for Enhancing Adversarial Transferability —
- A Generative-AI Modeling Framework for Explainable Decision Support in Complex Geosteering Scenarios —
- Combinatorial Optimization for All: Using LLMs to Aid Non-Experts in Improving Optimization Algorithms —
- A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI —
- R2DN: Scalable Parameterization of Contracting and Lipschitz Recurrent Deep Networks —
- Sufficient Decision Proxies for Decision-Focused Learning —
- Out-of-Sample Embedding with Proximity Data: Projection versus Restricted Reconstruction —
- A Network Science Approach to Granular Time Series Segmentation —
- Architectures of Error: A Philosophical Inquiry into AI and Human Code Generation —
- Kolmogorov-Arnold Energy Models: Fast, Interpretable Generative Modeling —
- Spherical Cauchy Variational Autoencoders: Heavy Angular Tails and Exact KL Evaluation —
- GeLaCo: An Evolutionary Approach to Layer Compression —
- Evaluating Out-of-Distribution Robustness in Graph-Based Android Malware Classification: A New Principled Benchmark —
- LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations —
- Text-Trained LLMs Can Zero-Shot Extrapolate PDE Dynamics, Revealing a Three-Stage In-Context Learning Mechanism —
- Compass-v3: Scaling Domain-Specific LLMs for Multilingual E-Commerce in Southeast Asia —
- Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training —
- Watermarking Diffusion Language Models —
- SoK: Kicking CAN Down the Road. Systematizing CAN Security Knowledge —
- oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning —
- TripScore: Aligning LLMs for Real-World Travel Planning via Expert-Calibrated Reward —
- The Environmental Impacts of Language Model Training Keep Rising Now is the Time to Catch Impacts on the Rebound —
- SGM: A Statistical Godel Machine for Risk-Controlled Recursive Self-Modification —
- Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry —
- Pre-train to Gain: Robust Learning Without Clean Labels —
- Effective and Efficient Threat Hunting with Small Language Models —
- Behavioral Coherence: A Method for Sensitive-Domain LLM Evaluation —
- Precision autotuning for linear solvers via contextual bandit-based RL —
- Kolmogorov-Arnold Networks-Based Tolerance-Aware Manufacturability Assessment Integrating Design-for-Manufacturing Principles —
- Ambient Dataloops: Generative Models for Dataset Refinement —
- SynCABEL: Synthetic Contextualized Augmentation for Biomedical Entity Linking —
- L2R: Low-Rank and Lipschitz-Controlled Routing for Mixture-of-Experts —
- Why beta 1 = beta 2 Is Dynamically Special in Adam —
- Elastic Spectral State Space Models for Train-Once Budgeted Inference —
- Model Specific Task Similarity for Vision Language Model Selection via Layer Conductance —
- On the Inherent Privacy Amplification of Missing Data —
- Architectural Design, Not Only Model Intelligence, Governs Multi-Agent LLM Performance —
- Score-based diffusion models for severely ill-posed problems in diffuse optical tomography —
- QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning —
- Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design —
- Transformer-based Parameter Fitting of Models derived from Bloch-McConnell Equations for CEST MRI Analysis —
- Perturbing the Phase: Analyzing Adversarial Robustness of Complex-Valued Neural Networks —
- Exploring Sparsity and Smoothness of Arbitrary Lp Norms in Adversarial Attacks —
- KUDA: Knowledge Unlearning by Deviating Representation for Large Language Models —
- A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities —
- Uncertainty quantification in neural network-based glucose prediction for diabetes —
- Social Simulacra in the Wild: AI Agent Communities on Moltbook —
- MAPLE: Metadata Augmented Private Language Evolution —
- Automated Membership Inference Attacks (AutoMIA): Discovering MIA Signal Computations using LLM Agents —
- Guideline-grounded retrieval-augmented generation for ophthalmic clinical decision support —
- Visuospatial Perspective Taking in Multimodal Language Models —
- When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews —
- Automated Gradient-Driven Parameter Sharing for Low-Resource Multilingual Speech-to-Text Translation —
- When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models —
- CounselReflect: Opportunities and Challenges for Designing Tools to Support Self-Reflection on Mental Health and Well-Being Conversations with AI —
- An Online Machine Learning Multi-resolution Optimization Framework for Energy System Design Limit of Performance Analysis —
- Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation —
- PolyJarvis: An LLM-Orchestrated Agent for Automated All-Atom Molecular Dynamics of Amorphous Homopolymers —
- Semantic Feature Analysis: Improving Agents Without Searching Over Rollouts —
- Green-ELM: Efficient Analytic Learning via High-Dimensional Random Projections —
- Model-Aware Data Cleaning for Tabular Foundation Models —
- Contrastive Privacy: A Semantic Approach to Measuring Privacy of AI-based Sanitization —
- Learning to Theorize the World from Observation —
- Fast Training of Mixture-of-Experts for Time Series Forecasting via Expert Loss Integration —
- Clin-JEPA: A Multi-Phase Co-Training Framework for Joint-Embedding Predictive Pretraining on EHR Patient Trajectories —
- PersonalAI 2.0: Enhancing knowledge graph traversal/retrieval with planning mechanism for Personalized LLM Agents —
- How Loud Rumbles Hit Newsstands: A Data Analysis of Coverage and Spatial Bias in German News about Landslides Around the World —
- SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science —
- Multi-Resolution Attribution from Adaptive Routing State —
- Reinforcement Learning for Graph Generation under a Hard Assortativity Constraint —
- Batch Normalization Amplifies Memorization and Privacy Risks —
- By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode —
- LaSR: Context-Aware Speech Recognition via Latent Reasoning —
- Near-Optimal Machine Unlearning Utility for Smooth Strongly Convex Losses —
- Post-Boundary Bridge: Must Local Attention Go Global Between Global Layers? —
- EssentialGIN: a new approach for gene essentiality prediction based on graph isomorphism neural networks —
- SoK: Reconstruction Attacks on Synthetic Tabular Data (Insights from Winning the NIST CRC) —
- Measurement Under Selection: Decoy-Calibrated Failure Audits for Language Models —
- The Neutral Mask: How Alignment Training Provides Shallow Alignment while Leaving Partisan Structure Intact in a Large Language Model —
- Evaluating Bias in Phoneme-Based Automatic Speech Recognition Systems: An Analysis of IPA Transcription Models —
- Exploring a Layer-Wise Design Space for KV Cache Eviction —
- Redact or Keep? A Fully Local AI Cascade for Educational Dialogue De-Identification —
- Do LLM Attribution Metrics Transfer? Auditing Retrieval-Augmented Generation Evaluation Across Datasets and Constructs —
- When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents —
- When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis —
- Accelerating Q-learning through Efficient Value-Sharing across Actions —
- AI Training Manager: Bounded Closed-Loop Control of Adaptive Training Recipes —
- IHDec: Divergence-Steered Contrastive Decoding for Securing Multi-Turn Instruction Hierarchies —
- CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph —
- A Transferable Learned Temporal Prior for Transmission Reconstruction and Decision-Relevant Uncertainty in Real Outbreak Labels —
- Prompt-Driven Exploration: Language as an Exploration Space for VLA Reinforcement Learning —
- Faithful, Not Corrective: Model Capability Governs Message-Format Effects in Multi-Hop Agent Relays —
- When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation —
- Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation —
- Interpretable and Calibrated Classification of Clinical Data Using Supervised Feature Binarization —
- Building a Neural Network from Scratch: Implementation, Evaluation, and Optimization —
- Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering —
- Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning —
- MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice —
- Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models —
- Attributing Preprocessing Invariance in Spectral Foundation Models —
- Reference-free logged energy-oracle recovery for neural approximations of symmetric coercive variational problems: conforming Riesz reconstruction and archive-level selection —
- When Detection Does Not Guarantee Resistance: Reasoning and Poisoned Context in RAG —
- A Standardized Framework for Machine Learning in Power System Protection —
- AIREP: A Protocol for Per-Decision Evidence in AI Runtime Governance —
- Modality Discrepancy Transformer for Ambivalence and Hesitancy Recognition —
- Subliminal Prompting Beyond Static Geometry: Causal Depth and Multi-Token Confounds —
- Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations —
- What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews —
- FakeSpotter: A content and strategy agnostic Viral Misinformation Detection Tool —
- Stop Removing Stopwords: How an Inherited Preprocessing Default Distorts Legal Text-as-Data —
- Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry —
- Towards Proactive Detection of User-Side Implicit Conflicts in Human-LLM Dialogue —
- Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes —
- VisKG-LM: Compiling Knowledge Graphs into Visual Memory for Multiple-Choice Question Answering —
- Advantage Scale Calibration Imbalance in Group-Relative Optimization under Low-Variance Rewards: Diagnosis and Bounded Recovery —
- To Memories and Beyond: From Remembering to Knowing You across Long-Term Multimodal Personal Archives —
- Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes —
- BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research —
- What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks —
- EvoSherlock: Towards Agentic Lifelong Evolution for Unseen Long-Tailed Security-Critical Events in Videos —
- Federated Soft Clustering via Generalized Total Variation Minimization —
- Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer —
- Generative Query Suggestion via Intent Coverage and Query-Level Credit Assignment —
- What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis —
- Layer-wise Curriculum Learning for Efficient LLM Compression —
- PAPC: Platform Mediation for Privacy-Propagation Externalities in AI-Mediated Workflows —
- YNU-HPCC at SemEval-2025 Task 11: Bridging the Gap in Text-Based Emotion Using Multiple Prediction Headers —
- Robust Conformal Intrusion Detection via Traffic-Aware Calibration and Attack-Orbit Invariance —
- Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training —
- Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices —
- Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses —
- Radio-Frequency Convolutional Neural Networks —
- Learning-Induced Dynamical Transition in Recurrent Neural Networks —
- Why Pretraining Fails to Share Cross-Lingual Knowledge —
- AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment —
- A frontend-backend architecture for tool calls in full-duplex speech models —
- Personalized Federated Hierarchical Gaussian Processes for Privacy-Preserving Modeling of Heterogeneous Distributed Systems —
- Scaling Zero Knowledge UNSAT Verification via Normalized Chaining —
- How to Guide Your Language Flow —
- Learning Submanifolds for Subsequent Inference on Random Dot Product Graphs, Part 1: Theory —
- Smart Insole Human Activity Recognition for Continuous Monitoring in Elderly Care —
- Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not —
- The Role of Fine-grained Harm Signals in LLM Safety —
- Machine-Learning Assessment of the Predictive Value of Inflammatory Biomarkers for Cognitive Impairment in an Older Hispanic Adult Cohort —
- FCx: An algorithm for finding Feasible Counterfactual Explanations —
- Do AI Agents Understand Computer Architecture? —
- MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs —
- Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation —
- Less Is More: Graph-free Multimodal RAG via Multi-signal Late Fusion —
- Closed-World Resolution Against Tool Hallucination in LLM Agents —
- Bayesian Optimization with Rich Auxiliary Information via LLMs —
- The syntax and semantics of goals —
- Sharpness-Aware Minimization (SAM) Improves Classification Accuracy of Bacterial Raman Spectral Data Enabling Portable Diagnostics —
- Beyond Private Training: The New Landscape of AI Privacy —
- Compositional Reasoning in Language Models under Reinforcement Learning Post-Training —
- Enhanced Agriculture-informed Neural Network by Domain Knowledge —
- Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models —
- Search at the Cost of Sampling: Nearly-Instant Latent Space Bayesian Optimization —
- Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling —
- For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances —
- Null importance: Disentangling relevance for interpretable machine learning —
- QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training —
- LLM-as-an-Improver: Turning Verification into Better Candidates —
- An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence —
- LSTM-UT and Recurrent-Depth Transformers on Cellular Automata —
- EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data —
- A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems —
- Self Improvement via Fast Tree-search —
- Next-token functional estimation —
- When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent R'esum'e Screening —
- Agentic AI Networking for Heterogeneous Unmanned Aerial Systems in Low-Altitude Wireless Networks —
- Compressed Active Subspaces for Scalable Bayesian Inference —
- Continual Enterprise World Model Discovery in Dynamic Systems —
- From Parameters to Behaviors: A Survey of Model Fusion for Large Language Models —
- FedFIbOS: Fisher Importance based Optimal Submodelling for Heterogeneous Federated Learning —
- CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR Narratives —
- Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents —
- Form Over Content In Gradient-Based Data Attribution Methods —
- Full-Duplex Speech Models Take the Floor When Asked, Not When Needed —
- Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction —
- SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership —
- Semantic Layer Induction from Raw Telemetry via Hierarchical LLM and RAG Abstraction —
- The Complexity Kink: A Prompt-Side Structural Complexity Index for Code-Generation Reliability —
- From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization —
- Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs —
- A Policy Profile for Croissant: Refusal as a Property of the Dataset —
- ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI —
- Improving Cross-Lingual Transfer for Sequential Sentence Classification in Research Papers via Structural Similarity —
- Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions —
- CoRe: Coherence and Relational Alignment for Multivariate Time Series Forecasting —
- When2Think: Learning When and How Much to Reason —
- Conservation Buys Stability and Factoring Buys Counterfactuals in Physical World Models —
- FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA —
- Opinion Dynamics-based Coalition Formation for Federated Learning in Heterogeneous IoT Systems —
- SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes —
- Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits —
- Learn Your Own Thoughts: Abstract Token Curriculum —
- Reachability, Not Observation: Containing Systems Whose Wiring Changes —
- LearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents —
- ALIBI: Adversarial Legitimacy Injection in Binary Input against LLM Malware Analyzers —
- A Phonemically Comprehensive, ASCII-Only Romanization Scheme for Thai and Lao: Systematic Cross-Lingual Correspondence and Chinese-User-Friendly Design —
- Alliance Beats Isolation: Unifying Heterogeneous Allied Datasets Improves Classifier Performance —
- AutoData: Agentic Search for Pre-training Data Selection —
- Rethinking Multi-Agent Collaboration: When More Is Less —
- OceanMoE: Structured Conditional Sparse Computation for Long-Horizon Multivariate Ocean Forecasting —
- TorchCraft: Unified binder design by inverting an all-atom structure predictor —
- Integrating knowledge from case reports: a medical ontology based multimodal information system with structured summary —
- PhyRestore: Physics-Structured Latent-Factor Restoration —
- Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection —
- Learning-Based Reconstruction of Optical Properties in Bilayered Media from Single-distance Time-Resolved Reflectance Measurements —
- Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems —
- Sybil-TraceGuard: Traceability-enhanced Sybil Guardian for Connected and Autonomous Vehicles Using Dynamic Semi-supervised GNN —
- Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search —
- DeliveryGym: An RL Environment for Long-Horizon Embodied Agent Planning with Adaptive Curriculum —
- Dictionary-Constrained Grapheme-to-Phoneme for Unsegmented Languages from LLM-Annotated Data —
- Steering Equilibrium Selection in Regularized Self-Play via the Reference Policy —
- F squared DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows —
- Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization —
- MetaRTL: Meta-path Attention Enhanced Relational Table Learning —
- Beyond Flattened Tokens: Structure-Preserving EEG Decoding with Reusable TriDim Blocks —
- A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents —
- Trust, but Validate the Instrument: Auditing AI-Generated RTL Verification Plans on Authored Security-Regression Proxies —
- Constraint-Safe Graph-Context Scoring for Stable Point-Feature Labels Under Text-Width and Accessibility-Inspired Profiles —
- Expected Hypervolume Maximization for Multiobjective Optimization under Uncertainties —
- Pretrained Medical Representations for the Practical Screening of Drug Repositioning Candidates —
- Reproducibility is not construct validity: LLM measurement of institutionally situated communication —
- Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference —
- Physical knowledge on historical data matters more than enforcing physical constraints on the forecast —
- AURA: Adaptive Uncertainty-Routed Analysis for Email Threat Detection —
- JustMem: Just-Enough Memory Access for Long-Term Conversations —
- Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning —
- V= a kQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering —
- D-Quant: Driftable Entropy Coding for KV Cache Quantization —
- PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces —
- Evaluating Communicative Success in Machine-Translated Conversation —
- Generalization through Lexical Abstraction in Transformer Models: The Case of Functional Words —
- Online Adaptive Kernel Mixing for Gaussian Process Decision Making —
- ClashBench: Conflicts Leading Agents to Seize and Harm —
- Hopper: Bounded-Memory Collaborative Debiasing for Byzantine-Tolerant Peer Sampling —
- TRACE: Accountable Agentic Retrieval for Source Discovery in Digital Archives —
- Delphi Scanner: efficient and interpretable static malware detection via API sequence modeling —
- REARL: A Closed-loop Autonomous Driving Simulation Enhancement Framework with Real Traffic Data and Large Language Models —
- Digital Twins for Opinion Dynamics: A Generative LLM Framework for Social Networks —
- Amortizing Physics-Informed Neural Solvers via Graph Hypernetworks —
- KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms —
- Mind the Gap: How SBOM Specification Ambiguities Lead to Divergent Software Bills of Materials. An Empirical Tool Study —
- From "Who Is This User?" to "What Does This Purchase Mean?": A Deployed Pipeline for Semantic User Profiling at Bank Scale —
- On the Leakage of Massey Secret Sharing Schemes under Linear Computations —
- Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models —
- Error bounds in Sobolev norms for approximations with norm constrained ReLU neural networks —
- Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute —
- MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation —
- Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics —
- One Intervention per Component is Enough: Towards Identifiability in Linear Stochastic Dynamics from Steady State —
- Graph-Based Stochastic Power-UCT: Monte-Carlo Graph Search with Power Mean Estimation —
- Neuro-Symbolic Agentic AI for Networked Low-Altitude UAVs —
- Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence —
- DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression —
- CellRFT: Reinforcement Fine-Tuning for Single-Cell Perturbation Modeling —
- JANUS: Denial-of-Service Attack Against Beam Hopping in LEO Satellite Networks —
- Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification —
- Benchmarking LLM Compliance with China AI Generated Content Regulations —
- Customizable and Jointly Optimized Route Planning: A Deep Architecture Enabling Differentiable Shortest-Path Search —
- E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews —
- EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning —
- Geopolitical Divisions Across Languages in Large Language Models —
- Dynamic Generalized Gromov-Wasserstein Optimal Transport —
- XIR: A Framework for Interoperability across Cross-Chain Protocols Based on a Verifiable Intermediate Representation —
- FedeRICo: Federated Region-Influenced Coupling for Traffic Flow Prediction —
- Can Data Attribution Filter Out Subliminal Learning? Not Reliably —
- Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression —
- DART: Distillation-Aware Reparameterization for Training-Free LoRA Reuse in Few-Step Video Diffusion Models —
- MAGMA-GEN: Validated Recovery Supervision from Ambiguous Failures via Counterfactual Re-Execution —
- WiCleanData: Guaranteeing the Type Consistency of Wikidata by Taxonomy Refinement and Constraint Enforcement —
- Evaluating Explanation Methods by the Predictors They Induce —
- FCA-Guided Counterfactual Explanations for Multi-Modal Breast Cancer Diagnosis: A Framework Achieving Perfect Validity with Emergent Sparsity —
- Marginal utility, matrix factorization, and the Key-Value (KV) cache: a unified information-economic framework for sovereign geo-mining inference —
- Competition, Collusion, and Corruption: The Spectrum of MEV Attacks on DAG-Based BFT Consensus Protocols —
- Tailored to you: longitudinal effects of personalising language models —
- A Proposal for an Agentic AI Architecture to Support Multi-Domain Decision-Making in the Brazilian Armed Forces —
- Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition —
- MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards —
- SETTer: Sparse-Encoder Transformer for Long-term Multivariate Time Series Forecasting —
- UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning —
- Solving Minimum Span Antibandwidth and Cyclic Antibandwidth Labeling Problems —
- A Scalable Trust Discovery Architecture for the Internet of Agents —
- CARE-VI: Conservative Adaptive Reliability Estimation for Value Improvement in Off-Policy Actor-Critic Learning —
- Design of the IBM Granite 5.0 TurboCTC ASR Model —
- Perception, Layout, and Validation: Calibrated Confidence for Reliable Straight-Through Processing of Financial Documents —
- QoS-Aware Federated Learning for Multimodal In-Cabin Interaction in Smart Vehicles —
- Local Sparsity Enables Unsupervised LLM Safety Detection —
- Fast-varying Natural Frequencies and Damping Ratio Identification for Linear Time-Varying System —
- MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents —
- QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability Synchronization —
- A Noise Optimum in Rehearsal-Free Continual Learning: Isolation, Mechanism, and Scope —
- Small Enough to Know Everything: The Fully-Enumerable Transformer as an Instrument for the Science of Delayed Generalization —
- Fine-Tuning Models for Biomedical Relation Extraction —
- Support Thresholds, Not Algorithms, Limit Rare-Association Recovery in Co-Purchase Networks —
- Robust Federated Q-Learning with Almost No Communication —
- PaGNet: A Panel-Aware GBDT--Neural Network for Multi-Target Corporate Tax Avoidance Proxy Forecasting —
- Sequential Contextual Fit Predicts Human Behavioural and Neural Dynamics Across Domains —
- To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals —
- ResumeShield: Channel Separation and an Open Benchmark for Indirect Prompt Injection in AI Resume Screening —
- When Does Retrieval Help Time-Series Forecasting? —
- SoftTri: Smooth Triangular Membership Functions for Adaptive Fuzzy Inference Systems —
- Evaluating Financial Sentiment in the Age of AI —
- Towards a Unified Modality-Agnostic Multimodal Framework for Cognitive Workload Assessment —
- JointMatch: A Unified Heterogeneous Graph Neural Solver for Large-Scale Ride-Sharing Matching —
- Foundations of Stochastic Lexical Calculus: Semantic Descent and Random Dynamics on Probability Simplices —
- Scene-Conditioned Relation Routing for urban cellular activity forecasting —
- Silence Is Endorsement: Verification-Status Laundering in LLM Agent Pipelines —
- Subdomain-aware representation compression for pretrained image embeddings —
- Explaining spatial information flow in short-term traffic forecasting models using a gated graph attention network —
- Is It Still Worth Training a Classical Model in the Era of LLMs? A Crossover Benchmark on Tabular Data —
- Before the Warning Comes Too Late: Incremental Phone-Scam Detection from Speech —
- The Public Discourse Corpus (PDC): A Speaker-Attributed Dataset for Valence and Epistemic Modality with Target Speaker Participation —
- Accuracy Is Not Enough: A Cross-Architecture Audit of Demographic Bias in Deep Knowledge Tracing —
- How Far Can Sub-3B Open Language Models Go in Zero-Shot Essay Scoring on an 8 GB Consumer GPU? —
- Lens: Bringing the Right Semantic Perspective into Focus for Training-Free Multimodal Representation Learning —
- When AI Agents Commit: Cognitive Serializability Across Data, Evidence, Policy, and Authority —
- Improving Online Reinforcement Learning via Bidirectional Behavior Prior Distillation —
- Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks —
- AI-Driven Real-Time Relay Optimisation in Smart Urban NR-V2X Networks via Learning-to-Optimise Graph Neural Networks —
- JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations —
- Labeled Incidence Structures for Native Transformer Modeling of Text, Knowledge Graphs, and Hypergraphs —
- Personalising a Cross-User Surface Electromyography Encoder Under a Small Calibration Budget —
- Intact-to-Amputee Transfer in Surface-EMG Gesture Decoding: Training Source and Calibration Budget —
- A Learning Algorithm for Threshold Boolean Networks with Prescribed Fixed Points —
- Improving Generalization and Robustness in Offline Reinforcement Learning via Boundary-Aware Data Augmentation —
- AgentPProf: Semantic Profiler for Long Horizon AI Agents —
- SAGG: Sample-Adaptive Gradient Gating for Robust Multimodal Learning under Heterogeneous Corruption —
- Viveka-Insight: a cross-lingual concept graph and citation-grounded retrieval resource over the complete works of Swami Vivekananda in English and Bengali —
- Diagnose, Recover, Certify: Task Readiness under Hidden Dynamics Changes —
- Hypernetwork-Parameterized Spatially Adaptive Neural Operators for PDE Learning —
- ZeroHAT: Behavior-Conditioned Zero-Shot Human Activity Trace Generation —
- EviRec: Continual Evidence Learning for Dual Cold-Start POI Recommendation —
- DDQN-MLP: An Explainable and Adversarially Robust DRL-Guided Adaptive Learning Framework for Ransomware Detection —
- NeuSOGA3D: A Neuro-Symbolic Framework for Explainable 3D Geometric Reconstruction —
- Sharp Reconstruction Bounds for Autoencoders Using the Same Forward Map —
- Structured Four-Stage Legal Translation: From Natural-Language Traffic Rules to PROLOG —
- A Qualitative Model for Reasoning about Path and Support —
- COMPASS: Ordered Clustered Routing at 100K Scale —
- Minimax-Optimal Online Contract Design with Unrestricted Bounded Contracts —
- Generating Heterogeneous 3D Geological Microstructures from 2D Images via a Stable Diffusion-Adversarial Model —
- The More It Says, the More You Pay: A Black-Box Audit of Provider-Side Token Inflation in LLM Services —
- Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning —
- Schema-Anchored Latent Reasoning for Semantic Parsing-Based Knowledge Base Question Answering —
- Learning Principal-Agent Contracts for Equitable Smallholder Carbon Farming under Moral Hazard and Adverse Selection —
- Xeno-Interpretability: Investigating the Alien Minds of LLMs —
- The Bias of Nonlinear Two-Time-scale Stochastic Approximation under Constant Step-Sizes —
- Stress-testing Alignment Midtraining —
- SCGFM-ART: Amortized Relational Transport for Structure-Centric Graph Foundation Models —
- The Organization of Inference: Information, Resource Constraints, and AI Production —
- Seismic Site Response Prediction from Sparse Observations Using Finite-Element-Pretrained Latent Dynamics —
- Online Supervised Dimension Reduction with Random Features: Diagnostics and Computational Trade-offs —
- GraphSkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback —
- Fingerprinting Multimodal Large Language Models —
- Training Neural Networks to Approach the Optimum Bayes Estimator in Dense Multi-Emitter Localization —
- Deep Learning-Based Classification of Cognitive and Resting States Using Electroencephalography Signals —
- How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents —
- Edustories: A Collection of Real-world Case Studies from Classroom Practices —
- Distributionally Robust Federated Learning with Multi-Source Data —
- Radio Frequency Detection and Classification of Microplastics in Water —
- When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation —
- SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness —
- Relational Attention for Data-Efficient Language Modeling —
- Towards TEE-Certified DP: Verifiable Differentially Private Training on Legacy GPUs —
- FreqCondNorm: Towards Cross-domain Predictive Maintenance through a Frequency-Conditioned Transformer Foundation Model —
- Refuse, Decompose, Refresh: A Claim-Safe Protocol for Closed-Loop AI Evaluation —
- Parallelism, critical windows, and separations among diffusion language models —
- An Analysis of Training-Free Self-Reported Confidence in Language Models —
- Language-model groups overstate consensus when replaying human deliberation on a reasoning task —
- Mitigating Retaliatory Algorithmic Collusion in Repeated Games —
- Empirical Analysis of Randomness Quality in Differential Privacy Mechanisms —
- Steering the Compass: Aligning Dynamic Psychological Counseling Conversations with Cognitive Behavioral Therapy Strategies —
- TAP Accuracy Below the Fluctuation Scale and Universal Posterior Geometry in Spherical Linear Models —
- Limits of Confidence in Diffusion —
- SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment —
- CrystalMO-TuRBO: Multi-Objective Trust-Region Bayesian Optimization for High-precision Joint Crystal Structure Refinement —
- WiC is Not WSD: A Study on LLMs and Lexical Ambiguity Resolution —
- Recursive Quantum Long Short-Term Memory for Stable Short-Horizon Temperature Forecasting —
- COIN-GP: Cooperative Online Learning in Networked Distributed Systems with Partial Measurements via Gaussian Process Regression —
- Weather Data Spoofing Attacks on Rain-Adaptive Millimeter-Wave Frequency Selection in V2X Communication Networks —
- What Does Privileged Information Add to On-Policy Self-Distillation? —
- Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape —
- Chronicle: Cut-Point Replay for Regression Testing of LLM Agents —
- UniPolicy: Unified Objective-Specific Policies for Generative Search Advertising —
- PAA: The Probabilistic Allen Algebra: A Generative and Complete Probabilistic Extension of Allen's Interval Relations —
- Multi-center Medical Data Mining with FL-Net - A One-stop Shop for Federated Learning —
- Ownership in AI-Assisted Everyday Tasks —
- Epidemiological Causal Graph Identification: Challenges, Identifiability and Algorithms —
- RISC-V and machine learning: a survey —
- HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication —
- Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models --- A Conceptual Framework and Registered Test Protocol —
- Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL —
- Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models —
- Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure —
- On-Demand Attention: Language Models Know When to Recall —
- Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation —
- dQwen3.5: Hybrid-Attention Diffusion Language Models —
- RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents —
- Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation —
- Calibrated RF-Fingerprinting Under Interference With Heterogeneous Transmission Protocols —
- Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations —
- RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning —
- PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers —
- JEPA-Anything: Learning Predictive Models across Different Worlds —
- An Empirical Study of Harness Design for Coding Agents —
- Score Centering Stabilizes Off-policy Reinforcement Learning —
- Unifying Models of Intergroup Hostility in Online Discourse —
- Embedding Models Measure in Peculiar Ways —
Important terms
- Checkpoint Handoff
- A protocol used to figure out if an AI agent is actually getting smarter at solving problems or just getting better at finding easier starting points. It works by cloning a specific state and handing it to a different policy.
- Deep Knowledge Tracing
- A type of AI model used in education to track student learning. Research shows these models can be biased, often performing worse for students from lower socioeconomic backgrounds despite being highly accurate overall.
- Zero-shot Transfer
- The ability of an AI model to perform a new task without any specific training or labeled examples for that task. In some areas, like decoding gestures from muscle signals, this method often fails compared to specialized training.
- Evidence-Gated Matched-Pulse Transport
- A specialized mathematical procedure used to help physical machines diagnose changes in their own internal dynamics. It allows an agent to decide whether to continue a task or stop for safety when it detects something is wrong.