AI papers — 2026-09-18

Today is mostly about the danger of inherited defaults, specifically how tiny, invisible choices made during data preprocessing can quietly sabotage the signals we are trying to study. This is most urgent in legal research, where scholars treat judicial text as data to uncover ideological or doctrinal shifts.

Many researchers still rely on removing stopwords, such as "the" or "and," based on mid-century information retrieval habits that were never tested for accuracy in a legal context. Researchers ran an exhaustive test by removing roughly 18,500 individual words one by one from over 14,000 Supreme Court opinions.

They found that using standard stopword lists actually performs worse than removing nothing at all. Even when they tried to build the best possible custom list of words to delete, the results were statistically indistinguishable from leaving the text untouched.

This suggests that cleaning text might be distorting the legal signals researchers want to recover, making the decision to remove stopwords a fundamental question of measurement validity. The problem of hidden, distorting information also appears in physical systems.

A machine might appear to be working perfectly while actually losing its ability to perform a future, yet-to-be-defined task. This tension is addressed by new crossover benchmarks that debate whether to prompt a frozen large language model or invest in custom training for tabular data.

These benchmarks quantify the exact point where a classical model's learning curve overtakes an LLM's zero-shot performance. In 86% of cases across eighteen datasets, a trained classical model beats a small frozen LLM using no more labeled data than was already on hand, with the median crossover occurring at just 6% of the training set.

This suggests that for most business applications, collecting a few hundred labels to train a gradient-boosted model is more efficient than relying on an LLM's semantic understanding of feature names. Similarly, in the domain of neural prosthetics, researchers found that zero-shot transfer for surface-EMG gesture decoding fails entirely.

However, providing just three labeled repetitions allows a cross-user encoder to exceed the performance of a standard per-user classifier by 0.190 macro F1. This performance gap is further bridged when the training pool is enriched with amputee data rather than just intact subjects.

While these specialized models require specific calibration, other architectures are finding ways to bypass expensive online reasoning entirely. The VisKG-LM framework demonstrates that knowledge graphs can be compiled offline into a visual memory of relation-labeled paths.

By treating these as cached images for a vision-language model to consult, researchers achieved significant gains in question answering over traditional methods that re-encode subgraphs during every inference step. This transition from theoretical modeling to real-world deployment introduces significant risks regarding fairness and reliability.

In educational technology, a cross-architecture audit of deep knowledge tracing models reveals that accuracy often comes at the cost of equity. Researchers evaluated four architectures—DKT, DKVMN, SAKT, and AKT—across two datasets to see how they handle demographic metadata.

They found that bias is a persistent reality, as every architecture showed significant socioeconomic bias on the Eedi dataset with lower AUC scores for economically disadvantaged students. Most strikingly, the most accurate model, AKT, which gains roughly four AUC points through item-level Rasch embeddings, also exhibited the largest socioeconomic bias.

Attempts to mitigate this through reweighting or adversarial training proved unreliable, as these methods failed to change the ABROCA metric in any configuration that maintained accuracy. Similarly, in physical control systems, a new framework for task readiness addresses the danger of dormant dynamics changes.

An actuator might function perfectly for a current task but fail when a new task is assigned because its degradation was never excited by previous operations. By using an intervention-based Bayesian procedure called Evidence-Gated Matched-Pulse Transport, agents can now diagnose local dynamics changes and provide either a recovered policy or an abstention decision.

This effectively trades off performance for safety and readiness certification. We finally have a way to tell if an agent is actually getting smarter at making decisions or if it is just learning how to stumble into better starting positions.

This is vital because when we train agents in closed loops, their own actions shape the environments they encounter, making it hard to know if a performance boost comes from better reasoning or simply reaching easier states. A new protocol called checkpoint handoff solves this by cloning a specific state reached by one policy and handing it directly to another.

By splitting gains into reach, which is how often an agent arrives at a successful state, and solve, which is how often it finishes once it gets there, researchers found that reinforcement learning improves both. On the ALFWorld benchmark, the data shows that an RL-trained solver is consistently more capable than an SFT-trained one when starting from the exact same point.

This ability to refine how an agent interacts with its environment is echoed in work on optimizing external skills through structured graphs. A framework called SkillAA uses an attribution-guided approach to repair failed procedures by routing errors to specific locations in a unified skill graph.

By using local and big gates to validate changes, it achieved high accuracy across SearchQA, LiveMath, and DocVQA. While SkillAA focuses on repairing external tools, another approach turns the verification process itself into a generator of better ideas.

Instead of just using a verifier to pick the best option from a fixed pool, the Verify-Repair-Reselect method uses feedback to build entirely new, improved candidates. This can actually recover correct answers even when every single initial attempt was wrong.

The question of whether AI agents truly understand the systems they manipulate remains a central tension in hardware design. Using a framework called AutoTuring, researchers tested whether agents designing accelerators were actually reasoning about architecture or simply performing a sophisticated search over numerical knobs.

By presenting the same fifteen-dimensional accelerator space twice—once with meaningful architectural names and once as anonymous variables on a scale from zero to one—they could measure the value of semantic meaning. On a nine-kernel FP16 GEMM task, the results showed that meaning does matter, as the architect agent outperformed its blind counterpart by 12.3 percent on average and required 70.1 percent fewer simulator calls.

However, this advantage is not absolute, as a critic loop was able to recover most of the performance gap for the blind agent without providing any additional benefit to the architect. This suggests that structured critique and architectural knowledge may act as substitutes rather than complements.

This ambiguity extends into the financial sector, where the high-stakes nature of autonomous trading agents introduces unique security risks. Through the FARSIGHT framework, which evaluates agents against market turbulence and various attack vectors, researchers analyzed fifteen academic trading schemes.

They found that 80 percent of these schemes failed at least one core robustness metric, and 100 percent exhibited security vulnerabilities. These failures are deeply interconnected, as a minor misjudgment can cascade into a market-wide crash, a vulnerability that an adversary could exploit to trigger a collapse at minimal cost.

The shift toward more sophisticated agentic evaluation is perhaps most clearly seen in the introduction of checkpoint handoff, a protocol designed to disentangle whether reinforcement learning gains stem from an agent's ability to reach better states or its ability to solve problems once it arrives there. By cloning a state reached by a released checkpoint and handing it to another policy without retraining, researchers can separate these two effects into REACH and SOLVE components.

Across two benchmarks and two independent pipelines, the data shows that the interaction between a reacher and a solver is consistently positive, and an RL history provides more value to an RL solver than to an SFT solver. On ALFWorld, RL improves both metrics, and notably, the SFT solver never succeeds in instances where the RL solver fails.

This granular view of agentic performance is complemented by work on the structural components of coding harnesses, which suggests that the effectiveness of an agent is heavily mediated by its environment. In studies across SWE-Bench Verified and Terminal-Bench 2.1, researchers found that context management becomes critical as budgets tighten, primarily by preventing context-overflow failures.

They also found that rule-based elision outperforms LLM-based summarization for efficiency. Furthermore, while planning acts as an accuracy scaffold for weaker models, it serves more as a cost-saving measure for stronger ones, and the utility of predefined tools depends heavily on the model's native bash proficiency.

The push toward more autonomous systems is manifesting in how we handle complex decision-making and data curation. In the realm of navigation, researchers have moved beyond simple heuristic costs by introducing a deep architecture that enables differentiable shortest-path search.

By running a multi-objective Dijkstra algorithm offline to establish a Pareto optimal candidate set, they designed a neural network capable of jointly optimizing cost functions and route ranking. This allows for highly customizable routes based on specific user preferences, outperforming existing methods in both quality and adaptability.

Similarly, the challenge of data engineering is being met by AutoData, an agentic search tool that treats pre-training data selection as a problem of heuristic engineering. Rather than just adjusting weights over fixed domains, AutoData searches a program space of scoring and stratification rules to discover optimal selection algorithms through iterative refinement.

This approach successfully transferred from small proxy models to larger scales, improving the downstream CORE metric. Even in highly specialized domains like legal reasoning, automation is advancing through a new four-stage framework called S4L that enables the translation of natural-language traffic rules into executable Prolog code.

By performing semantic role extraction and scene completion within a single guided prompt, S4L achieved a 75 percent accuracy rate in formalizing rules, significantly surpassing standard natural-language or logical English baselines. The most significant breakthrough involves making the process of steering large language models automatic rather than manual, which is vital for scaling safety and control.

A new framework called Deep Noir uses architectural timing and causal attribution to find the exact points in a model where you should intervene to change its behavior. In tests across various model sizes, this approach improved spam detection by as much as 42 percentage points in larger models and boosted sentiment accuracy by 13.1 percentage points without requiring any code changes.

However, this newfound control comes with a trade-off, as the study found that steering actually creates a predictable vulnerability to prompt-injection attacks that gets worse the more you steer. This tension between control and vulnerability is mirrored in the way we evaluate model safety through fine-grained signals.

Researchers have found that when you strip away a model's general sense of harmfulness to look at specific categories, like hate speech or violence, those specific signals behave differently across different models. Interestingly, even when these category-specific signals are mathematically orthogonal to general harm at one layer, they still end up amplifying the model's overall harmfulness in later layers.

While we struggle to control what models do, we are also struggling to define how we measure their success. A massive mapping of nearly 15,000 papers shows that as benchmarks evolve, there is a growing shift toward testing how models interact and perform professional tasks.

This raises a looming question about whether using AI to help build and score these very same tests might just create a loop that reinforces the existing biases of the models themselves. The difficulty of setting clear goals extends to specialized agents, such as those tasked with fixing broken travel itineraries.

In a comparison of different strategies, researchers found that while full replanning is effective for complex disruptions, hierarchical repair methods are much better at keeping an itinerary stable and preserving the parts a traveler already accepted. This highlights the constant struggle in agent design between finding a perfect new solution and sticking to what already works.

This need for precision is also evident in how we help agents troubleshoot technical problems. A new framework called RAFT moves away from treating support cases as static documents, instead treating them as evolving timelines of events.

By retrieving specific moments in a historical case that match the current problem, it outperforms standard retrieval methods at every stage of a troubleshooting process. Even the way agents learn to code is being refined to avoid wasting expensive computing resources.

A method called SIFT uses a lightweight tree search and an AI judge to quickly rank potential code improvements, reserving the most expensive evaluations only for the most promising candidates. This allows coding agents to improve themselves much faster and more cheaply than previous methods.

In the realm of causal discovery, researchers are trying to prevent models from accidentally skipping over important relationships between variables. A multi-agent framework called MaSCoD attempts to solve this by organizing structural patterns before the model makes its final judgments.

While this approach helps retain more relevant connections in certain datasets, it doesn't work perfectly across the board and can sometimes increase false positives. Finally, there is a growing concern regarding the underlying security of decentralized systems.

A new study mapping the landscape of maximal extractable value attacks shows that almost every production DAG-based BFT protocol is vulnerable to some form of transaction manipulation. The research suggests that whether an attack succeeds depends more on how the protocol was designed from the start than on how much effort an attacker puts in.

Today's papers

The papers

Important terms

Checkpoint Handoff
A protocol used to figure out if an AI agent is actually getting smarter at solving problems or just getting better at finding easier starting points. It works by cloning a specific state and handing it to a different policy.
Deep Knowledge Tracing
A type of AI model used in education to track student learning. Research shows these models can be biased, often performing worse for students from lower socioeconomic backgrounds despite being highly accurate overall.
Zero-shot Transfer
The ability of an AI model to perform a new task without any specific training or labeled examples for that task. In some areas, like decoding gestures from muscle signals, this method often fails compared to specialized training.
Evidence-Gated Matched-Pulse Transport
A specialized mathematical procedure used to help physical machines diagnose changes in their own internal dynamics. It allows an agent to decide whether to continue a task or stop for safety when it detects something is wrong.