AI papers — 2026-09-11

Today is mostly about the struggle to make machines understand how things change over time, whether that is the meaning of words or the progression of a disease. We have to start with a major new resource called CHRONOBERG, which finally gives us a way to teach models that language isn't static.

Most current training data lacks a long-term temporal structure, meaning models often fail to grasp how the sentiment or even the definition of a word shifts over decades. By curating 250 years of English books from Project Gutenberg and adding temporal annotations, this work allows us to quantify lexical changes through valence-arousal-dominance analysis.

It turns out that models trained sequentially on this data still struggle to encode these diachronic shifts, proving we desperately need more temporally aware training pipelines. This need for temporal context is just as critical in medicine, where we are trying to predict cancer risk from mammograms.

Usually, a model performs best when it can see a patient's entire history, but in a real clinic, you might only have the current scan available at the moment of decision. A new framework called SEM-HD solves this by using that historical data as privileged information during training to teach a student model how to mimic the insights of a full longitudinal history.

When tested on three major cohorts, this method consistently improved risk prediction even when it only had access to a single exam at deployment. The challenge of managing time and structure extends into how we build intelligent agents through hierarchical reinforcement learning.

This field is trying to help AI discover useful patterns within long streams of experience so they can plan better in complex environments. While there are many ways to do this, from using offline datasets to leveraging large language models, we still don't have a universal definition of what actually constitutes a good temporal structure for an agent to exploit.

We need to figure out if these models are actually following our orders or just mimicking patterns, because the way they handle instructions is far more fragmented than we thought. New research suggests that instruction-following isn't some single, magical ability the model turns on; instead, it looks like a skillful coordination of different linguistic skills that emerge at different stages of the process.

Probes show that models don't have one universal way of checking constraints, but rather rely on task-specific representations that only really become decodable once the model starts generating text. This lack of a unified internal logic is also showing up in much more sensitive areas, like how models handle mental health topics.

While we used to just look at multiple-choice answers to see if a model was biased, looking at the actual reasoning steps reveals much deeper, hidden stigmas that traditional tests miss. It turns out that when you dig into the intermediate logic, you find far more problematic language and flawed reasoning than a simple test score would ever suggest.

The difficulty of trusting these internal processes extends to how we deploy them in high-stakes environments like medicine. A new framework called VeriSim shows that when you move away from perfect textbook cases and introduce the messy, noisy communication typical of real patients, diagnostic accuracy drops by up to 25 percent.

It highlights just how much smaller models struggle compared to their larger counterparts when the conversation gets complicated. Getting agents to work with massive enterprise databases is a nightmare because you can't just shove hundreds of noisy tables into a prompt without breaking everything.

A new framework called TRUST-SQL tackles this by treating the problem as a partially observable process where the agent has to actively hunt for and verify relevant metadata rather than relying on it being pre-loaded. By using a dual-track reinforcement learning strategy that separates exploration rewards from actual execution outcomes, they managed to boost performance by 9.9% relative to standard methods.

It is quite a feat because their 4B and 8B models actually beat out strong baselines that had the luxury of seeing the full schema upfront. This ability to navigate complex, unorganized information is also being applied to how we manage long-context windows in models like DeepSeek-V3.2.

While sparse attention helps scale, the indexer used to find relevant tokens becomes a massive bottleneck as context grows, so a new hierarchical approach called HISA replaces that flat scan with a two-stage process. It first filters out irrelevant blocks of data at a coarse level before doing any fine-grained token work, which speeds things up significantly at 64K context without needing any extra training.

Moving from the architecture of the models themselves to how we actually train them on specific data, there is a push to make active learning more efficient for graphs. A framework called GATTA uses test-time augmentation to help models better estimate their own uncertainty, which turns out to be a much cheaper way to get high performance than trying to engineer incredibly complex acquisition functions.

The most significant breakthrough for scaling scientific discovery comes from a new framework called EvoMaster, which finally addresses the way research agents tend to lose progress or repeat mistakes over long periods. By implementing what they call loop research, the system allows evidence and experience to persist across different stages of an experiment, connecting execution, exploration, and evolution into a single continuous process.

Using GPT-5.4, this framework hit a mean score of 58.02% across ten benchmarks for coding and reasoning, which is a massive jump over the 40.29% achieved by the previous leader, Codex, while actually being 35.6% cheaper to run. This ability to manage complex, multi-stage processes is also being tested in how we evaluate scientific code.

A new benchmark called PETSCAgent-Bench moves beyond simple pass/fail tests to see if AI can actually write code for high-performance computing libraries like PETSc. It turns out that while models can write readable code, they still struggle with the specific API conventions and performance requirements that an expert human would follow.

The difficulty of getting models to behave correctly in specialized environments is mirrored in the messy reality of how humans perceive information. A large-scale study involving over 3,000 people looked at whether AI models actually understand the emotional weight of news headlines regarding geopolitical conflicts.

While top models like GPT-5.2 showed a very high correlation with human sympathy judgments, the researchers found that alignment isn't universal; even when aggregate scores look good, different demographic groups experience these models differently. The most significant breakthrough for autonomous construction comes from a new way to stop large language models from making silly spatial mistakes when building things.

By using a 2.5-D decomposition, researchers have forced the model to plan only in a flat, two-dimensional plane while letting a deterministic system handle the vertical stacking based on column occupancy. This clever trick removes the burden of calculating height from the model entirely, which pushed accuracy up to 94.6 percent on the Build What I Mean benchmark.

It is a massive jump compared to previous systems that were stuck around 76 percent, and it even works beautifully on edge hardware like an NVIDIA Jetson Thor AGX. This same logic of using specialized structures to fix model weaknesses shows up in civil engineering as well.

Instead of letting a model guess at complex physics, a new multi-agent framework uses a closed-loop process of generation and validation to design concrete highway barriers. While a standard model struggled to meet strict safety regulations, this coordinated team of agents achieved a 98.3 percent compliance rate with even the smallest 8B parameter model.

Moving from physical structures to digital information, there is a new way to fight misinformation that actually beats crowdsourced efforts. A system called MUSE uses trust-aware retrieval and multimodal reasoning to identify falsehoods, outperforming highly rated Community Notes by 29 percent.

It doesn't just flag errors; it provides grounded explanations that help people recognize misinformation more effectively. We need to talk about how we protect the very essence of who we are as we move into an era of digital replicas.

A new framework has emerged to address the ethical minefield of cognitive digital twins, which are dynamic computational models that simulate a specific person's cognition using their behavioral and physiological data. These aren't just simple assistants; they are proxies that can act or decide on your behalf, creating massive risks like shadow twins or shifts in epistemic authority.

The authors argue that current governance is insufficient because it focuses on the final decision rather than the cognitive representation itself, proposing a new five-pillar framework to manage these high-risk simulations. This need for deep structural protection extends into how we verify the very data used to train our models.

To combat the risk of people stealing datasets, a new method called CertDW uses conformal calibration to create certified watermarks. It essentially checks if a suspicious model's prediction stability on watermarked samples is significantly higher than its stability on benign ones, allowing owners to prove their data was used even under malicious attacks.

The way we manage these complex systems is also being refined at the most fundamental level of model training. Researchers have found that the instability often seen at the end of large language model pretraining actually stems from the geometry of output embeddings.

By implementing output embedding centering, they can suppress logit divergence more effectively than previous methods like z-loss, making training much more stable without being as sensitive to hyperparameter tuning. Even our ability to trust what a model knows about itself is being scrutinized through new lenses.

While there is a lot of hype around AI sentience, new testing shows that frontier models actually possess limited but measurable metacognitive abilities. They can assess their own confidence and anticipate their own answers, though these skills are qualitatively different from human thought and depend heavily on the specific context provided.

The most significant breakthrough for anyone working on educational content is a new way to fix the messy geometry that often ruins AI-generated animations. While Large Language Models are great at writing code for libraries like Manim, they frequently struggle with spatial logic, resulting in overlapping or illegible objects.

A new module called the Symbolic Geometric Agent solves this by intercepting the code, running it partially to build a symbolic scene graph, and then refining the instructions whenever it detects a collision. This approach led to a 16.1 percent improvement in visual quality scores using a GPT-5.1 pipeline, and in human tests, people preferred these corrected videos over raw outputs 84.4 percent of the time.

Efficiency is also seeing a massive leap in how we serve massive Mixture-of-Experts models. A new system called FluxMoE stops forcing every single expert to live permanently on the GPU, which usually chokes out memory for long conversations.

Instead, it uses an expert paging abstraction to stream weights on demand, keeping the actual computation on the GPU while moving others to host memory. For a large model like GLM-4.5 running on eight H20 GPUs, this boosted throughput by 7.2 times compared to standard vLLM setups without losing any model quality.

Moving from hardware efficiency to the actual mechanics of how models think, researchers are finding that we can track how a model learns by seeing how a single tiny change spreads. By fine-tuning a model on just one adversarial example and measuring how that infection moves to other inputs, they've found that models develop structured linguistic abstractions through experience alone.

This method avoids the geometric assumptions of previous studies and shows that representations are more like conduits for learning than static patterns of activation. This idea of how information is structured leads directly into a deeper look at model safety and refusal behaviors.

There has been a debate about whether a model's refusal to answer is just a single direction in its internal activations, but new comparisons show it is more nuanced. While some methods simply collapse the difference between harmful and harmless states, others can actually flip an activation into the opposite cluster.

This suggests that models encode the absence of a concept in a fundamentally different way than they encode its presence, leaving us with a much richer map of how to steer them. Understanding how machines perceive humor is vital because it requires moving beyond simple pattern matching toward genuine reasoning.

A new framework called Incongruity-Resolution Supervision teaches models to explicitly model the mismatch in a visual scene and then construct a coherent reinterpretation to resolve that tension. By training on these structured reasoning traces, a 72B parameter model achieved a 76.10% ranking score on cartoon captioning, actually outperforming non-expert humans and existing multimodal baselines.

This push for better reasoning extends to how we handle complex constraints in language models. A neuro-symbolic framework named SDDL helps smaller, resource-constrained models solve scheduling problems by translating natural language into formal abstractions that a deterministic solver can handle.

This approach boosted feasibility for the strongest configurations to 55.3%, significantly outperforming direct generation methods which struggled to maintain any consistency at all. The way we train these models also needs refinement to ensure they learn meaningful patterns rather than just memorizing text.

One method uses TF-IDF statistics to weight cross-entropy loss, which de-emphasizes common, low-information tokens and forces the model to focus on semantically rich ones. This simple tweak reduced substring memorization by up to 58% in some fine-tuning scenarios without hurting overall performance.

Even the fundamental way we balance different types of data, like text and video in sentiment analysis, is being questioned. Research shows that current optimization-based balancing methods often fail because they mistake how fast a model learns a modality for how useful that modality actually is.

This suggests we need to move toward measuring utility through held-out performance rather than just looking at gradients. In the realm of generative modeling, there is a growing debate over whether continuous or discrete processes are better for language.

A new model called RePlaid shows that continuous diffusion can scale effectively, establishing a scaling law that rivals discrete models and achieves a state-of-the-art perplexity of 22.1 on OpenWebText. We can also make these generative paths more efficient by making them aware of the model's own errors.

By using fiberwise optimal transport to build schedules that account for prediction risk, researchers achieved a 38.6% relative reduction in FID for flow matching on CIFAR-10. Training efficiency is further improved by automating the most tedious parts of the process, like picking a learning rate.

A tool called ExpTest treats the training loss curve as a signal to perform statistical tests, automatically triggering rate reductions when it detects convergence. This allows models to reach competitive performance across various architectures without any manual tuning.

Finally, we have to confront the reality that chatbots might not be the thinking partners we hope they are. An analysis of metaphorical problem propagation suggests that because LLMs are trained on text that only partially imitates human thought, they lack the cognitive flexibility required for true problem-solving. This implies that simply building larger models may never bridge the gap between imitation and actual understanding.

Today's papers

The papers

Important terms

CHRONOBERG
A major new resource containing 250 years of English books from Project Gutenberg. It uses temporal annotations and valence-arousal-dominance analysis to help teach AI models how word meanings and sentiments change over long periods of time.
SEM-HD
A framework designed for medical imaging that uses a patient's full historical data as privileged information during training. This teaches a student model to make accurate risk predictions even when only a single current scan is available.
TRUST-SQL
A framework that treats interacting with massive enterprise databases as a partially observable process. It uses dual-track reinforcement learning to help AI agents actively hunt for and verify metadata rather than relying on pre-loaded information.
EvoMaster
A scientific discovery framework that uses loop research to prevent AI agents from repeating mistakes. It connects execution, exploration, and evolution into one continuous process, allowing evidence to persist across long-term experiments.
FluxMoE
An efficiency system for Mixture-of-Experts models that uses expert paging. Instead of keeping every expert in GPU memory, it streams weights on demand from host memory, significantly boosting throughput during long conversations.