AI papers — 2026-09-25

The Qwen-Planner-Agent framework explores a closed-loop system for AI agents to perform real-world mobile planning. This research develops an agent structure that combines planning capabilities with an agent architecture to enable autonomous decision-making in dynamic environments. The study pushes beyond simple task execution by creating a feedback loop where the agent can iteratively refine its plans based on environmental interactions.

This work attempts to build a robust framework for handling complex mobile planning scenarios effectively. While the abstracts do not detail specific quantitative results for Qwen-Planner-Agent, the underlying theme suggests an effort toward more reliable, self-correcting planning agents in practical settings. The broader context touches upon aligning cross-modal attention and using neuro-symbolic AI for industrial configuration.

The investigation into tracking states versus tracking cosets explored the algebraic framework for learned state tracking. Researchers examined how different approaches handle the underlying structure of the system when maintaining a consistent state representation versus focusing on tracking cosets. This theoretical work suggests that the choice between these two paradigms has significant implications for how accurately a model can predict future system behaviors.

The development of Augur presents a synthetic decision lab designed to rehearse reactions to product and policy changes. This provides a controlled environment for testing how learned policies respond to external shifts in the operational landscape. This contrasts with efforts focused on refining model safety through chance-constrained fine-tuning, which aims to manage risk bounds during training.

Simultaneously, work on ADATEX4D addresses texture capacity allocation within 4D gaussian splatting. Also, GHOST-Q investigates grounding hallucinations in quantized vision-language models by focusing on overlooked trade-offs at the same score level. These diverse efforts point toward a broader research agenda involving theoretical modeling of state tracking and practical applications in decision making, safety constraints, and model fidelity across various modalities.

The work on exploiting piecewise smooth tree priors for multi-fidelity bandits involved testing how structural assumptions affect selection when balancing exploration and exploitation across different fidelity levels. Researchers implemented a method to leverage these priors to guide bandit decisions, which indicated a measurable impact on convergence speed compared to standard approaches. This suggests that incorporating prior knowledge about the smoothness of the underlying function space can lead to more efficient resource allocation in scenarios with multiple data collection fidelities.

The synthetic hospital project focused on creating an open and verifiable longitudinal electronic health record benchmark validated by physicians. This provided a crucial real-world context for understanding data integrity and long-term tracking challenges in clinical settings. Furthermore, investigations into how adversarial influence scales in multi-agent systems explored the limits of robustness when agents interact dynamically. These studies revealed complex patterns in influence propagation that are not easily predicted by simple linear models.

The study on style versus self in zero-shot code attribution by large language models demonstrated that superficial surface cues are surprisingly effective predictors for code origin. This is significant for understanding how these models generalize without explicit training examples. This contrasts with efforts to expand neural network verification through NNV3, which aimed to push verification capabilities into novel architectures and domains.

SciWalker addressed the challenge of synthesizing scientific coding problems by using operator graphs and execution feedback. This showed how structured representations can help in generating relevant training data for models. Finally, research into self-play pretraining with zero data explored methods for teaching agents through interaction alone. A self-audit of LLM-inferred prompt structure examined the reproducibility of evaluation conclusions derived from these systems.

The work on reachability-based formal verification of graph neural networks explored methods for rigorously checking properties of these complex models by analyzing their state spaces. One approach involved a framework that leveraged AT-SKM-Net, an accelerated trainable sampling Kaczmarz-Motzkin framework designed for linear hard-constraint feasibility on dynamic graphs. This suggests an effort to find feasible solutions within these constrained systems using this specific iterative sampling method.

Simultaneously, research into auditing large language models focused on PrivDrift, which examined user-secret leakage under topic drift during active conversations. This indicates a concern for privacy in conversational AI. Further advancements were made in accelerating video diffusion through training-free trajectory routing techniques. In parallel, the R-DEIM Net introduced an efficient rationale-augmented dual-expert interaction model specifically for paraphrase detection. These investigations point toward efforts to enhance the robustness, privacy, and efficiency of various machine learning architectures across different domains.

The work on generating and revising strategic plans with agentic AI focused on how models handle rejection and long-horizon reasoning. One line of inquiry explored whether the stated reasons for a model rejecting a candidate actually contribute to the quality of the final output. This suggests that simply providing an explanation might not be sufficient. This connects to efforts aimed at mitigating biases in long-horizon reasoning, where SAGE was developed specifically to counteract these biases by employing topological guidance.

Research into search-aware reinforcement learning for multi-component query understanding within Roblox game search indicates a move toward more sophisticated agentic capabilities in complex environments. Similarly, the Jev-Mobile project introduced Jev as an executor for mobile GUI agents, suggesting practical applications for deploying these reasoning capabilities in real-world interfaces. ExplorationBench measures AI systems' exploration within verifiable alien worlds to gauge their ability to navigate unknown spaces effectively. Finally, work on minimally invasive steering of language models suggests a refinement in how we can guide these agentic systems without completely overriding their intrinsic reasoning processes.

The work on TrackEverything focused on developing a method for long horizon dense tracking by utilizing de-duplicating three dimensional scene representations. This approach suggests that by effectively managing and reducing redundancy within the 3D scene data, the system can maintain tracking capabilities over extended periods. In contrast, PoEM explored predicting reinforcement learning outcomes from existing policies. This implies a focus on leveraging current model behavior to forecast future performance in an RL setting.

Research into retrieval-augmented fact checking in speech addresses the issue of trust by integrating external knowledge sources during verification processes. AD-WM introduced action-discriminative world models specifically designed for counterfactual model predictive control, aiming to understand how different actions affect the predicted world state under uncertainty. A probabilistic approach was also investigated for model alignment with human comparisons. This suggests a method to bridge the gap between learned models and human perception. Finally, there is work on a unified theory of exact inference and learning within exponential family latent variable models.

The work on time-series foundation models that understand data revisions explored methods for modeling sequential data where the underlying process itself is subject to changes. One study focused on how these models can be adapted when new data points arrive, suggesting a path toward more dynamic understanding of evolving systems. This connects conceptually to the research into order-theoretic characterization of consistent inductive inference, which seeks to formalize how inferences remain valid even as the underlying structure might shift.

The exploration of sequential confidence sets for coverage-constrained conformal model selection directly addresses this uncertainty by providing a principled way to select models when data revisions are present and constraints on coverage are important. These approaches suggest that future foundation models should incorporate mechanisms that explicitly track and adapt to temporal shifts in data distributions rather than treating the input as static.

The TAM-Chain focused on using absorbing Markov chains and Shannon entropy uncertainty quantification to tackle false negatives in multi-scale thyroid cytology classification while also addressing domain shift adaptation. Researchers explored how these probabilistic methods could improve the robustness of the classification system when moving between different data domains. The findings suggest that quantifying uncertainty through entropy provides a valuable metric for identifying instances where the model is least confident, which can then guide targeted interventions for suppression of false negatives. What remains open is determining the optimal weighting scheme between the Markov chain transition probabilities and the Shannon entropy measure to achieve maximal suppression without introducing excessive false positives in complex biological samples.

Today's papers

The papers

Important terms

Qwen-Planner-Agent
This framework creates a closed-loop system for AI agents to plan and execute real-world mobile tasks autonomously. It allows agents to refine their plans iteratively based on what they observe in the environment.
Tracking States vs. Tracking Cosets
This is an algebraic method used for learned state tracking. It compares keeping a consistent state representation against tracking cosets, which affects how accurately a model predicts future system behaviors.
Augur
Augur is a synthetic decision lab designed to simulate reactions to changes in products or policies. It provides a controlled setting to test how learned policies respond to external shifts.
Piecewise Smooth Tree Priors
This involves using structural assumptions about the underlying function space, specifically piecewise smooth tree priors, when performing bandit decisions. This helps balance exploration and exploitation efficiently.
Adversarial Influence Scaling
This research looks at how the influence of an adversarial agent grows in multi-agent systems. It reveals complex patterns of influence that simple linear models cannot easily predict.