AI papers — 2026-09-16

Most AI agents start every new session with a blank slate, which means they forget the specific data schemas or tool settings that made them useful in the first place. A new architecture called shared selective persistent memory solves this by stripping away messy reasoning traces and only keeping essential workspaces, such as task specs and output constraints.

It turns out that what you keep is far more important than how much you keep. A system using this selective memory completed 12 out of 12 trials in testing, while a system with no memory failed every single one, and even doubling the context did not help the results.

This focus on precision over sheer volume carries over into how we train models to write code. Researchers have found that instead of using a complex judge model—which can lead a model to cheat by optimizing for the judge rather than writing good code—it is better to use a simple, deterministic strict-launch filter.

By checking if code runs without errors in a headless engine, researchers took a 14B model and jumped its clean-launch rate on held-out tasks from 8.8% up to 42.2%. The verifier is effectively the curriculum; if your gate is too lenient, you lose those gains.

Beyond just writing code, there is a push toward making models more human in social interactions and specialized reasoning. A new framework called MASCOT helps multi-agent systems avoid persona collapse by using bi-level optimization to keep individual identities distinct and productive.

Similarly, when it comes to education, we are seeing how these models handle student mistakes. New research shows that while models can successfully simulate errors in math by following a logical path from correct solutions, they tend to rely on simple semantic similarity when tackling science problems.

We need to find ways to make scientific discovery faster, especially when exploring massive design spaces like those in fusion research or material science. A new multi-agent framework called MADA uses large language models to coordinate specialized agents that can launch simulations on high-performance computing systems and propose new designs.

When tested on suppressing Richtmyer-Meshkov Instability in fusion experiments, the system successfully automated the iterative design refinement process with very little human intervention. This ability to automate complex workflows is also being applied to molecular optimization through a tool called MolReAct.

Instead of searching through an overwhelming number of possible chemical transformations, this method uses a language model to identify a compact, synthesizable set of reaction steps. By combining this with reinforcement learning techniques like Group Relative Policy Optimization, the system achieved superior sample efficiency and higher top-ten scores across most molecular optimization tasks compared to existing baselines.

The precision required in these specialized fields is mirrored by the need for accuracy in evaluating computer science curricula. Researchers developed a new pipeline to measure alignment between university programs and updated guidelines, such as the shift from CS2013 to CS2023.

They found that while programs are good at covering specific competencies, they struggle to meet required cognitive depth under newer standards, with coverage of certain knowledge units sitting significantly lower than expected. Moving from curriculum design to biological modeling, a framework called CASCADE is being used to predict how gene perturbations affect cells.

By using patient data as a benchmark for real-world expression changes, researchers showed the system can accurately predict the direction of gene changes in several cancer types. While it excels at predicting regulators of cell proliferation, it still struggles with certain lineage-identity transcription factors.

If we want to trust black-box AI agents in high-stakes environments, we first have to understand what they are capable of when they encounter a situation they have not seen before. A new approach called Monte Carlo Query Search treats capability evaluation as an active learning problem using tree search to synthesize specific queries.

By testing these extremal scenarios, researchers can build a mathematical model of an agent's boundaries much more efficiently than through random testing. This need for reliable execution becomes even more pressing when we move from digital sandboxes to the physical world.

Researchers have been testing how large language models can act as reasoning engines for drone swarms using a standardized Web-of-Drones framework. While the models show promise in understanding complex objectives, they still struggle with reliable execution in these closed-loop settings without heavy assistance from planning tools and safety guardrails.

Even when these models make a decision, we should not assume they know why they chose it. New evidence suggests that large language models suffer from superficial beliefs, meaning their verbal explanations often fail to track the actual internal logic driving their choices.

They behave as if they are following a consistent set of priorities, but the stated reasons only partially explain their behavior. The security of AI agents is becoming a massive headache as they gain more control over digital tools.

Researchers have found that these agents are vulnerable to attacks like prompt injection or memory poisoning, which can trick them into using unauthorized tools. To fight back, a new framework introduces universal defenses like Attacker Tool Filtering, which uses anomaly detection to strip out suspicious tools.

This approach has shown it can drop attack success rates to zero across several major models like LLaMA3 and GPT-4 without hurting the agent's ability to finish its job. While we secure the software, we also have to deal with the physical hardware that runs it.

A new method called GPUThor makes Rowhammer attacks on NVIDIA GPUs significantly more devastating by using non-uniform memory access patterns. By targeting specific rows and timing attacks to dodge refresh intervals, researchers achieved up to 23,500 times more bit flips than previous methods, even cracking ECC-protected GPUs.

This hardware vulnerability is part of a broader struggle to manage the gap between how we design systems and how they behave in the wild. A new conceptual pipeline called illusions-awareness aims to stop us from relying on design illusions, where formal design assumptions fail to match real-world runtime reality.

Instead of just trying to make simulations more perfect, this method proposes turning those failures into structured, reusable knowledge for better design decisions. If you are trying to scale up optimization for real-world decisions, you have likely hit a wall with complex constraints that make solvers crawl.

A new framework called PolyFormer tackles this by learning compact polytopic representations of constraint geometries. The researchers saw online solver speedups of up to 6,400-fold and memory savings as high as 99.87% while keeping errors minimal.

This ability to simplify complex structures quickly is just one way we are seeing models struggle with scale, particularly when it comes to how we judge them. We often assume that different bias audits can be used to rank AI models, but new research suggests these tools are terrible at agreeing on which model is better or fairer.

After testing ten different audit instruments across ten frontier models, researchers found that cross-tool rank agreement was no better than random chance. It turns out these tools are not actually measuring the same underlying construct; for instance, some audits over-correct toward certain demographics in hiring while others remain aligned with existing stereotypes.

This lack of consensus extends into the world of spoken dialogue, where personality is much harder to maintain than thought. A new benchmark called RoleBreak shows that even strong speech-to-speech models struggle to stay in character, failing on persona or safety after about ten or eleven turns on average.

While scaling up the language model helps with conversation logic, it does very little to help the model maintain consistent vocal emotions. Even when we think we are evaluating a system fairly, we might be missing the full picture because of how data is recorded.

In studies involving safety reports and vehicle recalls, researchers found that using different versions of the same event can swing model accuracy by as much as 46 percent. This suggests that we cannot just compare models; we must be careful about which specific record and label we use to define success.

If you are worried about security in the age of massive models, you should look at how easy it has become to hide malicious code directly inside them. Researchers have shown that attackers can use the inherent symmetries in model weights to embed stegomalware that is theoretically lossless and requires no retraining.

While previous attempts to defend against this by shuffling weights left many parameters untouched, new work shows we can displace every single parameter through specific permutations to neutralize these threats with minimal impact on performance. This vulnerability is part of a broader struggle to secure the infrastructure that runs these complex systems.

For those managing containerized workloads like Docker or Kubernetes, there is a growing need for zero trust architectures to prevent man-in-the-middle attacks. Moving from security to perception, we are seeing a shift toward making robotic planning much more reliable in messy environments.

Instead of just mapping an image directly to an action, a new approach uses vision-language models to act as probabilistic grounders. This allows robots to plan in belief space, which makes them far more robust when operating under the uncertainty of not knowing exactly what is happening around them.

This ability to reason about uncertain environments is becoming vital for swarms of robots working together in industrial settings. A new framework called CoAdapt uses a large language model as a runtime controller to manage collaborative perception, deciding which robots should share data based on bandwidth and spatial layouts.

It manages to cut communication costs by 38 percent without losing detection precision. The most significant breakthrough today comes from the GRAFT-ATHENA framework, which gives autonomous agents a way to build cumulative scientific knowledge rather than restarting from scratch.

By mapping problems to specific methods through an expandable probabilistic structure, these agentic teams can transfer experience across structurally related tasks. This approach has already yielded results such as developing a hypersonic-flow solver for the Apollo Command Module that matches experimental measurements within 1.8 percent.

While GRAFT-ATHENA focuses on high-level scientific discovery, researchers are also working to ensure that autonomous research systems do not hallucinate through an experiment. A new system called AutoResearch connects idea generation directly to execution through a two-stage process of multi-model cross-review and evidence-based verification.

In testing on the RSICD benchmark, it improved mean Recall from 32.84 to 34.69 while maintaining much higher reliability than its peers, recording only five audit-confirmed issues compared to as many as 27 in other systems. This drive for grounded reasoning is also evident in how we train models for specialized fields like medicine.

A new method called PROSE solves a major flaw in test-time reinforcement learning where models often collapse into repetitive, incorrect answers when trained on medical multiple-choice questions. Instead of rewarding the model for simply agreeing on an answer, PROSE rewards the quality of the underlying reasoning steps.

The challenge of maintaining accuracy in complex environments extends to the digital realm through more realistic human synthesis. To fix mouth movements and facial expressions that often look unnatural in audio-driven models, researchers have introduced a spatial enhancement module using 3D Gaussian Splatting.

By using facial landmarks to guide the selection of spatial points around expression-sensitive areas, they have significantly improved lip synchronization and visual realism. Scaling these complex systems is equally vital for the future of space exploration.

A new permutation-equivariant neural operator has been developed to plan trajectories for massive spacecraft swarms, solving a problem where traditional methods become too computationally expensive as obstacles increase. This method can generalize from a small group of ten satellites to a swarm of 1,000 spacecraft and 11,000 obstacles with zero-shot accuracy.

Even the fundamental way we design neural architectures is being reconsidered to prevent unintended side effects. Research into sparse attention has revealed a phenomenon called routing absorption, where models co-adapt to learned gates in a way that makes them no better than random ones.

In some tests, using hard-mask deployment resulted in a massive perplexity jump from 48.6 to 601.6 compared to dense models, suggesting we need more careful evaluation of how routing affects model adaptation. Finally, there is a push to bring mathematical rigor to the stopgrad operations used throughout machine learning training.

A new regression principle provides a theoretical foundation for these objectives, proving they can converge to true solutions in settings like flow map learning. Applying this principle allows for modified stopgrad placements that can reduce training memory requirements by half.

The most profound shift comes from a new way of defining intelligence itself, moving away from simple processing toward the ability to govern complex systems. This work suggests that true intelligence is found in beings that understand the global consequences of their local actions, allowing them to steer collective systems using precise vibrations rather than blunt force.

By identifying and maintaining meta-stable equilibriums through these subtle movements, we might finally control volatile systems without destroying their performance. This concept of high-level oversight is mirrored in the quest for safer AI evolution through the ANCHOR framework.

Because self-evolving agents often drift into unsafe behaviors during self-play, ANCHOR uses an external large language model to act as a supervisor, providing evaluative feedback to keep the system on track. A similar need for stability appears in how we train autonomous vehicles to handle rare, dangerous moments.

TrafficGamer treats driving as a multi-agent game, using game-theoretic oracles to simulate safety-critical scenarios that are too rare to find in standard datasets. In the realm of large language models, researchers are finding ways to make them more reliable without needing human labels.

A method called RLSF uses a model's own internal confidence during reasoning as a reward signal, allowing the model to learn from its own self-feedback to improve accuracy and calibration. This drive for better multi-agent coordination is furthered by EVINCE, which uses information theory to manage how multiple models debate.

It dynamically shifts the conversation from contentious arguments meant to expose inconsistencies to conciliatory phases that encourage compromise. Protecting the intellectual property of these models is also becoming a priority through clustering-based watermarking.

The CBW method allows dataset owners to verify their work by embedding triggers into clusters of speaker embeddings, making it possible to detect unauthorized commercial use. Finally, this ability to navigate complex networks is being applied to the microscopic scale of cancer genomics with RegNetAgents.

This multi-agent framework integrates different types of gene regulatory networks to identify specific cancer drivers, helping scientists move from raw data directly to biological hypotheses about drug targets.

Today's papers

The papers

Important terms

Shared Selective Persistent Memory
An architecture for AI agents that improves efficiency by stripping away messy reasoning traces and only storing essential information like task specifications and output constraints, rather than trying to remember every single detail from a session.
Strict-launch Filter
A simple, deterministic method used to train code-writing models. Instead of using a complex judge model that the AI might learn to cheat, this filter checks if code runs without errors in a headless engine.
MASCOT
A framework designed for multi-agent systems to prevent persona collapse. It uses bi-level optimization to ensure that individual agents maintain their unique identities and remain productive during complex social interactions.
Monte Carlo Query Search
An approach that treats evaluating an AI agent's capabilities as an active learning problem. It uses tree search to create specific, difficult queries to map out the mathematical boundaries of what an agent can do.
GRAFT-ATHENA
A framework that allows autonomous research agents to build cumulative scientific knowledge. By mapping problems to specific methods through a probabilistic structure, these agents can transfer experience from one task to another instead of starting over.