AI papers — 2026-09-10

A new way to manage overlapping memories offers a breakthrough for long-term agent stability. Instead of asking a language model to rewrite its entire history, which is prone to error, a framework called ROAM organizes information into tiny, distinct atomic units.

The system classifies how these units relate as independent, equivalent, or conflicting. By grouping these atoms into primary observations and supporting evidence, the system fuses details into compact views that prevent outdated data from confusing the agent. This approach boosts answer precision by up to 29.8 percentage points while reducing irrelevant token noise.

Moving from how agents remember to how they act, there is a new call to evolve process mining from retrospective dashboards into active tools. The proposed BlueSky agenda suggests building agents capable of deciding whether an action should be taken based on privacy budgets and organizational authority.

This would require moving toward mineable artifacts like governance contracts and evidence packages. These tools would allow an agent to legitimately refuse or defer a task.

In the world of decentralized finance, neural networks are replacing hand-written formulas for calculating wallet reputation. A new network called zScore-N was trained using original formulas as a teacher to ensure perfect reproduction.

This network handles millions of wallets across massive scales and is resilient to gaps in data. It cuts the error caused by missing features nearly in half.

The focus on reliability extends to robotics, where researchers have made deep Q-learning more cautious about risk. By using mini-batch transition risk mappings, they enabled an underwater robot to navigate complex environments while avoiding destruction.

This method allows the robot to learn from a small number of samples without being misled by environmental randomness. Consequently, its policies work even in environments it never saw during training.

The security of massive models is more fragile than previously thought, especially in mixture-of-experts architectures. A study shows that directional ablation can strip a model of its ability to refuse harmful requests without extra training.

This technique works on the 320B parameter GLM-5.3-Flash, but it behaves unexpectedly. Researchers found that seventy-four percent of the effect occurs only when editing attention, dense writers, and routed experts all at once.

This means following old recipes for finding these directions will result in silent failure because researchers are looking in the wrong place. This vulnerability is a reminder of how little we understand about internal representations, a theme that carries over into mobile data privacy.

Researchers developed CrossLink to show that observers can stitch together different temporary identifiers used by phones for LTE, WiFi, and Bluetooth. By linking these identifiers across protocols and locations, they reconstructed full movement traces for eighty-three percent of users in simulations.

The difficulty of managing complex systems is also evident in how we train models for specialized tasks like translation. We need better ways to model the messy reality of human health, and a new generative transformer called NOAH attempts to do just that.

By training on over 559 million clinical events from nearly 300,000 patients, this model simulates the continuous evolution of a patient's journey. It can process medical images, time-series signals, and unstructured clinical notes to perform tasks like zero-shot classification or simulating responses to medical interventions.

This ability to handle multimodal data is also being applied to make large language models more reliable through better grounding. A new method called Evidence-Aligned Entity Verification addresses hallucinations in retrieval-augmented generation by checking if specific entities align with retrieved evidence.

It uses counterfactual stability analysis to ensure these alignments hold up even when evidence is slightly perturbed. This makes it more robust than methods relying solely on a model's internal knowledge.

While NOAH focuses on the depth of patient data, other researchers are looking at making the infrastructure behind these models more efficient. A new benchmark called Φ-Bench evaluates if large language models can engineer the complex software stacks that power them.

Unlike previous tests that look at small code snippets, this benchmark evaluates long-horizon tasks like end-to-end system optimization. This helps determine if we are getting closer to autonomous AI infrastructure.

Efficiency is also being tackled by improving how models talk to each other during inference. A framework called X-CoSD allows a small model on a device to work with a large model on a server, even with different vocabularies.

By using hybrid resampling, it only exchanges data for the shared parts of the vocabulary. This speeds up token generation without losing the quality of the larger model.

The way we manage model interactions is also evolving to include personal security measures. A system called Bio-Memory uses biometrics like face or palmprint recognition to ensure an AI agent gives retrieved memories to the correct person.

In tests, this biometric layer created significant gaps in retrieval accuracy between owners and non-owners. This provides a practical way to keep personalized data safe in shared environments.

Security remains a theme even in the deployment of specialized models for critical infrastructure. A hybrid architecture combining Spiking Neural Networks with XGBoost has been developed to protect electrical distribution networks from cyber attacks.

By using the spiking network as a fixed feature extractor, the system stays lightweight enough for edge deployment. Even when attackers try to poison data, this model's performance drops by only 0.9% in certain tests.

Finally, there is a growing debate about how agents should use their skills to get things done. Some researchers find it is more effective to treat skills as independent subagents rather than loading instructions into a single long context window.

By spawning fresh context windows for each subtask, these subagents can execute complex tasks more reliably than an agent trying to juggle everything at once. However, the security of agentic AI is a major concern because these systems can act on instructions hidden within images or audio.

A new benchmark called MMPIBench shows that while most multimodal prompt injection attacks are caught during planning, the vulnerability remains significant. In 12.8% of attempts, attackers successfully deliver instructions through QR codes or fake interfaces.

While only 1% of these result in a completed tool call, the risk is much higher when audio is involved. In those cases, attacks complete in up to 75% of instances for certain models.

This vulnerability to unexpected inputs mirrors the need for better efficiency in other complex computational structures. For instance, zero-knowledge machine learning circuits often suffer from massive redundancy that can be stripped away without compromising security.

By using whole-circuit abstract interpretation, this new method ensures every removed check is still logically supported. This reduces prover time by up to 72.8% and cuts constraints by nearly half.

The push for efficiency extends into how we handle long contexts in retrieval-augmented generation systems. Instead of choosing between recomputing expensive caches or risking accuracy, a new hybrid approach fine-tunes models to be aware of cache concatenation.

This method selectively recomputes parts of the cache to slash the time to first token by 80%. It also improves accuracy on long-context benchmarks compared to just recomputing everything.

We must solve how models process massive information without slowing down, and a new framework called ConvMem tackles this by treating long-context reasoning like a hierarchical convolution. Instead of reading text in a slow, linear chain, it uses an LLM as a convolutional kernel to summarize segments into a logarithmic tree structure.

This approach is training-free and highly parallelizable. It allows the model to outperform existing methods on complex reasoning tasks without the risk of overfitting.

While ConvMem helps models understand more, we still have to figure out if they truly understand across different languages. A new benchmark called SWORD reveals that many models are fragile regarding factual consistency.

They struggle significantly more with distorted statements in East Asian languages than in others. Models often rely on how familiar a sentence looks rather than true verification, showing a performance gap of up to 28 percentage points.

This gap in deep understanding is also visible in specialized linguistic tasks like Filipino speech synthesis. Researchers found that fine-tuning a ByT5 model using an LLM-assisted pipeline can significantly improve the accuracy of converting text to phonemes.

This helps the model better predict stress markers and disambiguate words that are spelled the same but pronounced differently. We must also build AI assistants that know when to step in proactively.

A new theoretical framework defines symbiotic agency as a system operating under a standing, revocable mandate from a human. The AI must constantly decide whether to act, monitor, or refrain based on its perception of the user's situation.

By treating individual behavioral episodes as the unit of analysis, this approach links an agent's decision to act directly to its authorized perception and authority containment. However, making these agents reliable is difficult because they often fail to understand their own limitations.

New research shows that when we ask a model how it would behave, it provides a generic theory of how AI agents work rather than actual self-knowledge. Even when shown exact data points, their reports are no more accurate than if they were describing a hypothetical third party.

Using first-person language mostly just makes them sound more flattering and less prone to admitting harmful behaviors. This lack of self-awareness makes it harder to use multi-agent debates to improve reasoning, as the benefits might be an illusion.

While having models argue can shift reported agreement by over 50 percentage points, it does not seem to change what they believe or improve final answers. In many cases, a model's dissent is just a temporary reaction to a hostile instruction.

If we cannot rely on self-correction through debate, we might need to look at how models handle specific constraints during generation. For discrete tasks like Sudoku, standard diffusion models often get stuck with early mistakes.

Simply changing the sampling method to draw directly from clean predictions can jump validity from 31% to 95% without retraining. We can also bypass expensive retraining by using training-free task vectors to steer model behavior.

Instead of fine-tuning, this method uses forward-pass statistics to map activation steering into weight-space edits. It allows us to amplify or suppress behaviors while keeping general problem-solving skills intact.

We must also address privacy leaks in federated learning, as new work shows passive attackers can reconstruct almost entire batches of data. By framing gradient inversion as a problem of erasure-correcting codes, researchers developed a peeling attack that recovers every sample and label in a batch from just one round of updates.

On ImageNet, this method recovered between 94 and 100 percent of batches up to size 128. This vulnerability is mirrored by risks in how we use models to generate information via retrieval-augmented generation.

A study found that poisoning just a few retrieved documents can drop accuracy from nearly 78 percent down to 43.5 percent. Interestingly, the model mostly just stops answering altogether when the context gets messy rather than hallucinating new lies.

To trust these systems in high-stakes areas like law, we must audit every individual claim rather than scoring answers as a single block. A new two-agent system called GANDR uses a Drafter and a Critic to verify claims against their sources.

It hits 70.8 percent strict accuracy on legal benchmarks, outperforming existing baselines by over 11 points. This is largely because it forces the model to commit to citations that actually exist in the retrieved text.

The difficulty of managing complex interactions also appears in how we train models on evolving data. A new benchmark called TTGBench addresses the fact that most temporal graph models are great at predicting structural changes but terrible at tracking semantic shifts.

By testing 17 different methods, researchers found that traditional graph neural networks handle structure well but fail at semantic drift, while large language models do the exact opposite. To make these models usable on phones, we must fix how they handle heat and power.

A new approach called PELM manages computational load by realizing that not every generated word needs to go through the entire neural network. By combining frequency scaling with speculative decoding, it can speed up inference by 23.1% while cutting energy use by over 50%.

This efficiency in hardware-level execution is mirrored by a way to speed up the math inside the transformer itself. A method called EFQ-Softmax maps attention scores directly to low-bit operands using a single affine rule.

On specialized hardware like the A5 vector unit, this cuts latency by about 40% without hurting model quality. Moving from hardware efficiency to training stability, researchers found that scale-invariant optimization in normalized networks is governed by a precise mathematical law.

They discovered that learning rates and weight decay interact through the parameter norm to create a hidden feedback loop. This can lead to unstable behavior if not carefully balanced.

The math behind regularization also becomes clearer when looking at the geometry of the loss landscape. By connecting divergence-based regularization to Sharpness-Aware Minimization, it turns out both methods are essentially penalizing local curvature to find flatter, more robust minima.

In the realm of security verification, a framework called AutoTrans is making it easier to protect new hardware. It uses a specialized signal extractor and template-based prompting to translate security assertions for RISC-V processors with a 78% automatic success rate.

This automation of complex tasks extends into business logic through CARRE, which helps companies predict customer churn. By combining retrieval-augmented generation with counterfactual scoring, it predicts risk reductions roughly 80% better than standard methods.

Finally, protecting the integrity of models remains a moving target. New work on preventative steering shows that defending against malicious fine-tuning requires active, progressive intensity scheduling to keep up with how model parameters adapt over time.

Today's papers

The papers

Important terms

ROAM
A framework for managing long-term AI memory by breaking information into tiny, distinct atomic units. It classifies these units as independent, equivalent, or conflicting to prevent outdated data from confusing the agent.
zScore-N
A neural network designed for decentralized finance that replaces manual formulas for calculating wallet reputation. It is trained using original formulas as a teacher to ensure it can handle massive scales and missing data.
Evidence-Aligned Entity Verification
A method to reduce hallucinations in retrieval-augmented generation by checking if specific entities match the retrieved evidence. It uses counterfactual stability analysis to ensure these links remain accurate even when data is slightly changed.
ConvMem
A framework that handles long-context reasoning by treating it like a hierarchical convolution. It uses an LLM as a kernel to summarize text segments into a logarithmic tree structure, making processing faster and more parallelizable.