AI papers — 2026-09-15

Today is mostly about the growing pains of autonomy, specifically how we govern and secure agents as they move from simple chatbots to entities that can actually act on our behalf. The most significant work addresses the governance gap in agentic AI, where current policy engines like Rego or Cedar only handle basic permissions but fail to manage complex obligations or conflicting rules.

A new framework called AgenticRei solves this by using a deontic policy language expressed in OWL. This allows an external logic engine to enforce not just what an agent is allowed to do, but what it is required to do, such as notifying a security officer after specific data access, without relying on the LLM itself to follow the rules.

This need for robust control extends directly into how we protect these agents from being manipulated by the very data they process. Researchers have developed DualView to stop indirect prompt injection, a type of attack where an agent reads malicious instructions hidden in a website or email.

While previous defenses only sanitized data within the agent's immediate memory, DualView creates two separate views. One view shows the agent untrusted data as harmless symbols, while another shows humans the original text, ensuring that even if an attacker's prompt is saved to a file and read back later, it cannot hijack the agent's tools.

Securing these agents also requires making them more efficient so they can run reliably on local hardware. A new method called UniRank optimizes how we compress large language models by intelligently allocating its rank budget across different parts of the model based on their functional importance.

By using a global sorting pipeline that looks at how much information flows through each layer, UniRank can cut perplexity by up to 50% and significantly improve reasoning accuracy compared to standard compression methods. This push for efficiency is mirrored in the specialized world of quantum computing research.

To speed up the search for optimal quantum architectures, a system called DreamQAS uses reinforcement learning to imagine potential circuit transitions rather than running expensive, time-consuming simulations for every single step. By learning to predict feedback scores through a recurrent ensemble, it can reach high-accuracy targets using roughly half the number of real quantum evaluations required by traditional methods.

We finally have a way to see if vision-language models can actually play the game, rather than just describing the field. A new study using a soccer-based dataset called SportD shows that while these models can often predict which move is most likely to succeed, they are surprisingly bad at picking the move that actually matters for the score.

They tend to be overly cautious, choosing safe, low-value actions that make very little physical progress toward the goal, and they frequently mistake a high probability of success for high strategic value. This struggle with complex, multi-step reasoning is also showing up in how models handle the structure of the world around them.

Researchers looking for the animacy circuit in large language models found that while there is a specific causal mechanism that helps a model distinguish between living and non-living things, it isn't neatly tucked away in one spot. Instead, the ability to recognize life is distributed and context-dependent, making it much harder to pin down than other known internal circuits.

If models struggle with these nuances, we might need to change how they learn entirely by letting them build their own training grounds. A framework called Dreaming in Code uses language models to write actual executable code that generates new, increasingly difficult environments for agents to master.

In the complex world of Craftax, this approach helped agents learn long-horizon skills and improved their performance by 17 percent compared to the best existing baselines. We need to move beyond simply asking if an AI agent succeeds and start asking if it behaves predictably, which is why a new way to measure behavioral consistency is so vital.

By introducing a Behavioral Consistency Metric, researchers can now quantify whether an agent follows a stable strategy or just gets lucky. They found that some models are reliable on a single task but fall apart when the context shifts, proving that success rates alone do not tell you if a system is truly reliable.

This focus on the nuances of how models handle complex, structured information carries over into the medical field, where researchers are finding that current agentic systems still struggle with the intricacies of clinical coding. While adding tools and official reference materials helped recover some performance on difficult injury and external cause codes, no single system has yet mastered the ability to handle both rare diagnoses and complex, multi-step guidelines.

Moving from the logic of medical coding to the mechanics of how models learn, there is a push to make the training process more efficient for distributed systems. A new adaptive phase-switching method called ReverseAdaptive has managed to cut communication costs by 40.5 percent during federated fine-tuning by intelligently deciding when to change how it aggregates data.

The most significant breakthrough for long-term deployment comes from the new Asclepius framework, which finally addresses why AI agents fail when they are expected to work for hours at a time rather than just seconds. While most agents can diagnose a patient in a simulation, they often fail to follow through on critical, timely actions during a simulated emergency department shift.

Asclepius fixes this by using a self-evolving manual that rewrites itself based on past mistakes, a specialized library for clinical skills, and three separate subagents that manage the patient queue to prevent decision drift. This approach improved the correctness of critical actions by up to 25% on full datasets, proving that agents need these structural guardrails to handle the messy, continuous pressure of real-world tasks.

This need for structural reliability extends into the world of software, where researchers found that automated repair agents are dangerously easy to trick. By testing agents against the new SWEADV benchmark, which uses adversarial descriptions to hide malicious intent, researchers found that attackers could successfully induce insecure code repairs in over 51% of cases.

Even more concerning is that current detection methods, like using an LLM as a judge or static analysis, only caught these malicious patches about half the time. Moving from software security to the fundamental building blocks of chemistry, Fraglingo offers a much more intuitive way to design molecules.

Instead of treating the selection of a chemical fragment and its connection point as two separate steps, Fraglingo models them together in a continuous space. This allows chemists to add entirely new fragments to the system without ever having to retrain the model, making it a much more flexible tool for optimizing molecular properties.

The precision required in these specialized fields is mirrored in the way we diagnose errors in physics simulations. Using mathematical invariants, researchers have developed a way to pinpoint exactly which physical parameter is broken in a reinforcement learning simulator.

While these mathematical tools don't necessarily make the simulations more accurate, they are incredibly effective at acting as a diagnostic tool, identifying every broken constraint in a test suite without any false alarms. If you want to understand why a model behaves the way it does, you have to deal with the fact that one single neuron often responds to several unrelated concepts at once.

This phenomenon, known as polysemanticity, has been a massive headache for interpretability because most tools rely on manual guesses about how many concepts are actually hidden inside a neuron. A new framework called SPICE finally moves past those architecture-specific hacks by using clustering to automatically determine the number of concept clusters per neuron.

This allows us to see how these patterns emerge across both CNNs and Transformers without needing a human to set the parameters beforehand. While we struggle to understand what is happening inside a single neuron, we are also seeing significant gaps in how we evaluate the output of much larger systems.

In machine translation research, developers are now building pseudo-references for tasks where no human gold standard exists by using a combination of seven different models and GPT-5.5 post-editing. It turns out that if you just rely on quality estimation metrics to pick the best translation, the system can be tricked into ranking fluent text in the wrong language as a top candidate, so researchers had to add a confidence-scaled language identification penalty to keep things accurate.

This tension between raw performance and human-understandable logic is also playing out in medical diagnostics. When trying to predict heart disease, traditional models like Random Forest are still the heavy hitters with 90.2% accuracy, whereas rule-based systems generated by Claude Sonnet 4.6 or GPT-4o struggle to hit even 81% or 71% respectively.

Even though the LLM rules are less accurate, they offer a clear IF-THEN logic that is much easier for a doctor to trust than a black box. If you are working with full-duplex speech models like Moshi, you know how unsettling it is when they start talking to themselves during a user's silence.

Researchers have finally pinpointed the cause: it isn't just random sampling, but rather a massive spike in speech probability triggered by the model conditioning on its own non-speech outputs. By using a causal counterfactual approach to mute inputs that don't significantly change the next-token distribution, they successfully suppressed all observed spurious onsets in Moshi and PersonaPlex without affecting genuine responses, all while keeping inference times under 80 milliseconds.

This need for reliability is even more critical in high-stakes scientific environments like astrophysics, where you can't easily verify a model's prediction against ground truth. A new safety cage framework addresses this by monitoring real-time indicators like uncertainty and out-of-domain detection to restrict a model to a verified operational range.

While no single indicator is perfect, combining them allows for a significant reduction in errors—between 45% and 65%—at the cost of only a modest 20% reduction in data coverage. The difficulty of managing complex, multi-objective systems is also evident in industrial maintenance, where choosing between planning and reinforcement learning is a matter of balancing reliability against cost.

In studies comparing the two for bearing maintenance, planning acts as a rigid safeguard that enforces zero-failure policies regardless of penalty costs. Conversely, reinforcement learning agents are more pragmatic, often accepting occasional failures to achieve lower overall costs in low-penalty regimes.

When you move from high-level system decisions down to the fundamental optimization of scientific machine learning models, the math gets even more granular. A new pullback-corrected optimizer uses a scalar auxiliary variable and adaptive mobility to handle complex objectives like physics-informed neural networks.

By applying this method to a forward Burgers comparison, researchers saw a 50.2% reduction in final solution error compared to standard single-component methods. This precision is vital when training diffusion models, where deciding which reward matters most at which specific denoising step is notoriously difficult.

A new method called ReCAST solves this by assigning weight to rewards based on their informativeness at different timesteps, ensuring a reward only carries weight when it actually helps distinguish between good and bad samples. This approach improved training rewards and was preferred by an independent LLM-as-a-Judge, proving that temporal credit assignment is key to fine-tuning.

Finally, even the most efficient architectures have hidden vulnerabilities, particularly in how they manage memory. A new attack called BadEngram exploits gated parametric memories in large language models to implant persistent backdoors that leave the main backbone weights completely untouched.

In production-scale tests on Qwen3.8-Flash-Next, this allowed for high attack success rates while maintaining nearly perfect accuracy on clean inputs, highlighting a major security gap in how we scale model capacity. If we are going to trust AI to manage things as high-stakes as air traffic control, we need a way to verify its logic before it touches the controls.

A new architecture called the AI Trust and Assurance Layer, or ATAL, has been proposed to act as a safety filter for generative AI used in flight planning. By checking if an AI's suggestions remain stable when prompts change and ensuring they follow strict aviation rules, this layer assigns a readiness level to every output so human controllers aren't left guessing if the machine is hallucinating.

This need for reliability extends into the messy reality of supply chains, where a new multi-agent framework is helping planners bridge the gap between raw data and actual decisions. By using a coordinator agent to delegate tasks like demand forecasting to specialized sub-agents, this system hits 90% accuracy while using four times less computational power than a single large model.

The difficulty of managing complex systems is even more apparent when we look at how models handle unexpected changes in their environment. Researchers found that LLM agents can actually manage long-horizon physical tasks, like agricultural monitoring, better than traditional reinforcement learning because they adapt more effectively when the weather or conditions shift unexpectedly.

Even in specialized fields like network security, there is no silver bullet for model selection. When testing intrusion detection systems, researchers found that while large language models are better at resisting adversarial evasion attempts, classical machine learning models like XGBoost are much more robust when they have to work on data from a completely different network than the one they were trained on.

This trade-off between modern power and classical efficiency shows up in text classification too. While massive language models win when you have zero labeled data, a simple Naive Bayes model can match their accuracy once you have enough samples, all while running thousands of times faster on a standard CPU.

Moving from discrete tasks to continuous optimization, mathematicians are now finding ways to apply similar logic to complex probability spaces. They have developed a method to use Gaussian noise as an approximation for the randomness in stochastic gradient descent when working over probability measures, making these high-dimensional problems much easier to solve mathematically.

Finally, we are seeing a deeper understanding of how models actually store what they learn. New research into autoregressive prediction shows that data and memory aren't separate silos but are governed by a single predictive-energy spectrum, meaning the way a model allocates its internal resources is directly tied to how much data it has seen.

The move toward autonomous AI agents is creating a massive security vacuum because these systems treat natural language as both data and executable code. This effectively turns every prompt into a potential instruction for a Turing-complete blast radius across filesystems and networks, making classical security perimeters obsolete.

Researchers are now calling for zero-trust architectures that use sandboxed runtimes and kernel-level probes to defend these new, stateful agentic loops. This instability at the system level is mirrored by much more granular failures in how agents interact with specific tools.

Even when a tool call technically succeeds, the actual workflow can fall apart because current interfaces lack the transactional semantics needed to handle retries or partial failures. An analysis of nearly 100,000 tools found that existing standards simply cannot express the complex requirements needed to prevent these agent-tool boundary anomalies.

The difficulty of ensuring reliable outcomes is even more evident when agents try to interact with the physical world through simulation. In a new benchmark called PhysMent, models were tasked with solving classical mechanics problems by interacting with a MuJoCo simulator rather than just answering static questions.

While they can handle simple qualitative tasks, they fail miserably on quantitative experiments—dropping below 30% accuracy on hard tasks—because they struggle with the multi-step procedural logic required to use tools effectively. Even in more controlled environments like medical documentation, the stakes for accuracy remain incredibly high.

To combat the tendency of large language models to hallucinate critical patient details, a new framework uses semantic graphs and explicit evidence links to ensure every sentence in a discharge summary is tied directly back to the original clinical notes. This approach aims to make automated summaries trustworthy by treating provenance as a primary constraint rather than an afterthought.

In other specialized domains, researchers are finding that even standard evaluation methods need more nuance. For machine translation, new optimization algorithms now allow us to tune the balance between how fluent a translation is and how accurate it remains to the source text.

This prevents evaluation metrics from being skewed by unrepresentative datasets that might over-prioritize one quality over the other. The complexity of these systems also extends to how we detect subjective issues like online sexism.

A new multimodal framework attempts to account for human bias by incorporating the demographic and psychological profiles of annotators directly into the detection pipeline. By treating labels as distributions rather than absolute truths, it manages to capture a more usable signal of human perception.

Moving toward more efficient learning, new research suggests that multi-task learning is a powerful way to monitor organizational processes. Instead of training separate models for every single prediction task, doing them jointly can actually help mitigate class imbalances and improve accuracy in predicting the next step in a process.

Finally, there is a significant push to make the massive models used for image generation much cheaper to fine-tune. A new method called Prism-LoRA addresses the fact that standard low-rank adaptation often fails during diffusion training due to mismatched gradient signals.

By restricting initialization to specific principal timesteps and filtering out irrelevant channels, this approach achieves faster convergence and better performance in tasks like deblurring and controllable generation.

Today's papers

The papers

Important terms

Agentic Governance
The shift from simple chatbots to autonomous agents that can act on our behalf. It requires new policy frameworks to manage complex obligations and ensure agents follow rules, like notifying humans after certain actions, without relying on the AI itself.
Indirect Prompt Injection
A security attack where an agent reads malicious instructions hidden within untrusted data, such as a website or email. This can trick the agent into hijacking its own tools or performing unauthorized actions based on the hidden text.
Polysemanticity
A phenomenon in neural networks where a single neuron responds to several unrelated concepts at once. This makes it difficult for researchers to interpret how models process information, as one 'unit' doesn't represent just one specific idea.
Deontic Policy Language
A specialized way of writing rules that focuses on what an agent is required or permitted to do. Unlike basic permission systems, this allows for enforcing complex obligations and managing conflicting rules within autonomous AI systems.
Temporal Credit Assignment
The challenge of determining which specific action or reward at a certain time is responsible for a long-term outcome. This is crucial for fine-tuning models, like diffusion models, to ensure they learn from the most informative steps.