AI papers — 2026-09-11
Today is mostly about the struggle to make machines understand how things change over time, whether that is the meaning of words or the progression of a disease. We have to start with a major new resource called CHRONOBERG, which finally gives us a way to teach models that language isn't static.
Most current training data lacks a long-term temporal structure, meaning models often fail to grasp how the sentiment or even the definition of a word shifts over decades. By curating 250 years of English books from Project Gutenberg and adding temporal annotations, this work allows us to quantify lexical changes through valence-arousal-dominance analysis.
It turns out that models trained sequentially on this data still struggle to encode these diachronic shifts, proving we desperately need more temporally aware training pipelines. This need for temporal context is just as critical in medicine, where we are trying to predict cancer risk from mammograms.
Usually, a model performs best when it can see a patient's entire history, but in a real clinic, you might only have the current scan available at the moment of decision. A new framework called SEM-HD solves this by using that historical data as privileged information during training to teach a student model how to mimic the insights of a full longitudinal history.
When tested on three major cohorts, this method consistently improved risk prediction even when it only had access to a single exam at deployment. The challenge of managing time and structure extends into how we build intelligent agents through hierarchical reinforcement learning.
This field is trying to help AI discover useful patterns within long streams of experience so they can plan better in complex environments. While there are many ways to do this, from using offline datasets to leveraging large language models, we still don't have a universal definition of what actually constitutes a good temporal structure for an agent to exploit.
We need to figure out if these models are actually following our orders or just mimicking patterns, because the way they handle instructions is far more fragmented than we thought. New research suggests that instruction-following isn't some single, magical ability the model turns on; instead, it looks like a skillful coordination of different linguistic skills that emerge at different stages of the process.
Probes show that models don't have one universal way of checking constraints, but rather rely on task-specific representations that only really become decodable once the model starts generating text. This lack of a unified internal logic is also showing up in much more sensitive areas, like how models handle mental health topics.
While we used to just look at multiple-choice answers to see if a model was biased, looking at the actual reasoning steps reveals much deeper, hidden stigmas that traditional tests miss. It turns out that when you dig into the intermediate logic, you find far more problematic language and flawed reasoning than a simple test score would ever suggest.
The difficulty of trusting these internal processes extends to how we deploy them in high-stakes environments like medicine. A new framework called VeriSim shows that when you move away from perfect textbook cases and introduce the messy, noisy communication typical of real patients, diagnostic accuracy drops by up to 25 percent.
It highlights just how much smaller models struggle compared to their larger counterparts when the conversation gets complicated. Getting agents to work with massive enterprise databases is a nightmare because you can't just shove hundreds of noisy tables into a prompt without breaking everything.
A new framework called TRUST-SQL tackles this by treating the problem as a partially observable process where the agent has to actively hunt for and verify relevant metadata rather than relying on it being pre-loaded. By using a dual-track reinforcement learning strategy that separates exploration rewards from actual execution outcomes, they managed to boost performance by 9.9% relative to standard methods.
It is quite a feat because their 4B and 8B models actually beat out strong baselines that had the luxury of seeing the full schema upfront. This ability to navigate complex, unorganized information is also being applied to how we manage long-context windows in models like DeepSeek-V3.2.
While sparse attention helps scale, the indexer used to find relevant tokens becomes a massive bottleneck as context grows, so a new hierarchical approach called HISA replaces that flat scan with a two-stage process. It first filters out irrelevant blocks of data at a coarse level before doing any fine-grained token work, which speeds things up significantly at 64K context without needing any extra training.
Moving from the architecture of the models themselves to how we actually train them on specific data, there is a push to make active learning more efficient for graphs. A framework called GATTA uses test-time augmentation to help models better estimate their own uncertainty, which turns out to be a much cheaper way to get high performance than trying to engineer incredibly complex acquisition functions.
The most significant breakthrough for scaling scientific discovery comes from a new framework called EvoMaster, which finally addresses the way research agents tend to lose progress or repeat mistakes over long periods. By implementing what they call loop research, the system allows evidence and experience to persist across different stages of an experiment, connecting execution, exploration, and evolution into a single continuous process.
Using GPT-5.4, this framework hit a mean score of 58.02% across ten benchmarks for coding and reasoning, which is a massive jump over the 40.29% achieved by the previous leader, Codex, while actually being 35.6% cheaper to run. This ability to manage complex, multi-stage processes is also being tested in how we evaluate scientific code.
A new benchmark called PETSCAgent-Bench moves beyond simple pass/fail tests to see if AI can actually write code for high-performance computing libraries like PETSc. It turns out that while models can write readable code, they still struggle with the specific API conventions and performance requirements that an expert human would follow.
The difficulty of getting models to behave correctly in specialized environments is mirrored in the messy reality of how humans perceive information. A large-scale study involving over 3,000 people looked at whether AI models actually understand the emotional weight of news headlines regarding geopolitical conflicts.
While top models like GPT-5.2 showed a very high correlation with human sympathy judgments, the researchers found that alignment isn't universal; even when aggregate scores look good, different demographic groups experience these models differently. The most significant breakthrough for autonomous construction comes from a new way to stop large language models from making silly spatial mistakes when building things.
By using a 2.5-D decomposition, researchers have forced the model to plan only in a flat, two-dimensional plane while letting a deterministic system handle the vertical stacking based on column occupancy. This clever trick removes the burden of calculating height from the model entirely, which pushed accuracy up to 94.6 percent on the Build What I Mean benchmark.
It is a massive jump compared to previous systems that were stuck around 76 percent, and it even works beautifully on edge hardware like an NVIDIA Jetson Thor AGX. This same logic of using specialized structures to fix model weaknesses shows up in civil engineering as well.
Instead of letting a model guess at complex physics, a new multi-agent framework uses a closed-loop process of generation and validation to design concrete highway barriers. While a standard model struggled to meet strict safety regulations, this coordinated team of agents achieved a 98.3 percent compliance rate with even the smallest 8B parameter model.
Moving from physical structures to digital information, there is a new way to fight misinformation that actually beats crowdsourced efforts. A system called MUSE uses trust-aware retrieval and multimodal reasoning to identify falsehoods, outperforming highly rated Community Notes by 29 percent.
It doesn't just flag errors; it provides grounded explanations that help people recognize misinformation more effectively. We need to talk about how we protect the very essence of who we are as we move into an era of digital replicas.
A new framework has emerged to address the ethical minefield of cognitive digital twins, which are dynamic computational models that simulate a specific person's cognition using their behavioral and physiological data. These aren't just simple assistants; they are proxies that can act or decide on your behalf, creating massive risks like shadow twins or shifts in epistemic authority.
The authors argue that current governance is insufficient because it focuses on the final decision rather than the cognitive representation itself, proposing a new five-pillar framework to manage these high-risk simulations. This need for deep structural protection extends into how we verify the very data used to train our models.
To combat the risk of people stealing datasets, a new method called CertDW uses conformal calibration to create certified watermarks. It essentially checks if a suspicious model's prediction stability on watermarked samples is significantly higher than its stability on benign ones, allowing owners to prove their data was used even under malicious attacks.
The way we manage these complex systems is also being refined at the most fundamental level of model training. Researchers have found that the instability often seen at the end of large language model pretraining actually stems from the geometry of output embeddings.
By implementing output embedding centering, they can suppress logit divergence more effectively than previous methods like z-loss, making training much more stable without being as sensitive to hyperparameter tuning. Even our ability to trust what a model knows about itself is being scrutinized through new lenses.
While there is a lot of hype around AI sentience, new testing shows that frontier models actually possess limited but measurable metacognitive abilities. They can assess their own confidence and anticipate their own answers, though these skills are qualitatively different from human thought and depend heavily on the specific context provided.
The most significant breakthrough for anyone working on educational content is a new way to fix the messy geometry that often ruins AI-generated animations. While Large Language Models are great at writing code for libraries like Manim, they frequently struggle with spatial logic, resulting in overlapping or illegible objects.
A new module called the Symbolic Geometric Agent solves this by intercepting the code, running it partially to build a symbolic scene graph, and then refining the instructions whenever it detects a collision. This approach led to a 16.1 percent improvement in visual quality scores using a GPT-5.1 pipeline, and in human tests, people preferred these corrected videos over raw outputs 84.4 percent of the time.
Efficiency is also seeing a massive leap in how we serve massive Mixture-of-Experts models. A new system called FluxMoE stops forcing every single expert to live permanently on the GPU, which usually chokes out memory for long conversations.
Instead, it uses an expert paging abstraction to stream weights on demand, keeping the actual computation on the GPU while moving others to host memory. For a large model like GLM-4.5 running on eight H20 GPUs, this boosted throughput by 7.2 times compared to standard vLLM setups without losing any model quality.
Moving from hardware efficiency to the actual mechanics of how models think, researchers are finding that we can track how a model learns by seeing how a single tiny change spreads. By fine-tuning a model on just one adversarial example and measuring how that infection moves to other inputs, they've found that models develop structured linguistic abstractions through experience alone.
This method avoids the geometric assumptions of previous studies and shows that representations are more like conduits for learning than static patterns of activation. This idea of how information is structured leads directly into a deeper look at model safety and refusal behaviors.
There has been a debate about whether a model's refusal to answer is just a single direction in its internal activations, but new comparisons show it is more nuanced. While some methods simply collapse the difference between harmful and harmless states, others can actually flip an activation into the opposite cluster.
This suggests that models encode the absence of a concept in a fundamentally different way than they encode its presence, leaving us with a much richer map of how to steer them. Understanding how machines perceive humor is vital because it requires moving beyond simple pattern matching toward genuine reasoning.
A new framework called Incongruity-Resolution Supervision teaches models to explicitly model the mismatch in a visual scene and then construct a coherent reinterpretation to resolve that tension. By training on these structured reasoning traces, a 72B parameter model achieved a 76.10% ranking score on cartoon captioning, actually outperforming non-expert humans and existing multimodal baselines.
This push for better reasoning extends to how we handle complex constraints in language models. A neuro-symbolic framework named SDDL helps smaller, resource-constrained models solve scheduling problems by translating natural language into formal abstractions that a deterministic solver can handle.
This approach boosted feasibility for the strongest configurations to 55.3%, significantly outperforming direct generation methods which struggled to maintain any consistency at all. The way we train these models also needs refinement to ensure they learn meaningful patterns rather than just memorizing text.
One method uses TF-IDF statistics to weight cross-entropy loss, which de-emphasizes common, low-information tokens and forces the model to focus on semantically rich ones. This simple tweak reduced substring memorization by up to 58% in some fine-tuning scenarios without hurting overall performance.
Even the fundamental way we balance different types of data, like text and video in sentiment analysis, is being questioned. Research shows that current optimization-based balancing methods often fail because they mistake how fast a model learns a modality for how useful that modality actually is.
This suggests we need to move toward measuring utility through held-out performance rather than just looking at gradients. In the realm of generative modeling, there is a growing debate over whether continuous or discrete processes are better for language.
A new model called RePlaid shows that continuous diffusion can scale effectively, establishing a scaling law that rivals discrete models and achieves a state-of-the-art perplexity of 22.1 on OpenWebText. We can also make these generative paths more efficient by making them aware of the model's own errors.
By using fiberwise optimal transport to build schedules that account for prediction risk, researchers achieved a 38.6% relative reduction in FID for flow matching on CIFAR-10. Training efficiency is further improved by automating the most tedious parts of the process, like picking a learning rate.
A tool called ExpTest treats the training loss curve as a signal to perform statistical tests, automatically triggering rate reductions when it detects convergence. This allows models to reach competitive performance across various architectures without any manual tuning.
Finally, we have to confront the reality that chatbots might not be the thinking partners we hope they are. An analysis of metaphorical problem propagation suggests that because LLMs are trained on text that only partially imitates human thought, they lack the cognitive flexibility required for true problem-solving. This implies that simply building larger models may never bridge the gap between imitation and actual understanding.
Today's papers
- CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models Researchers introduce a new temporal corpus to help language models better understand how word meanings change over centuries. [paper] [episode]
- Longitudinal Risk Prediction in Mammography with Privileged History Distillation This framework uses past medical images during training to improve breast cancer risk prediction even when only the current scan is available during actual use. [paper] [episode]
- Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning This survey examines how hierarchical reinforcement learning helps AI agents find useful patterns in complex sequences of actions. [paper] [episode]
- A Constraint Programming Approach for n-Day Lookahead Playoff Clinching in the NHL A new algorithm uses mathematical constraints to predict exactly when a hockey team will secure a spot in the playoffs. [paper] [episode]
- CQD-SHAP: Explainable Complex Query Answering via Shapley Values This method uses game theory to explain which parts of a complex question are most important for finding an answer in a knowledge graph. [paper] [episode]
- Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records This model combines medical lab results with transformer architectures to provide interpretable predictions from patient records. [paper] [episode]
- Relevance Is Not Permission: Localizing and Controlling Metric-Facing Attention Contributions This paper introduces a method to ensure that the parts of data an AI focuses on actually contribute to its final prediction.
- Smoothing the Score Function to Enhance Generalization in Diffusion Models New mathematical techniques are proposed to prevent image generation models from simply memorizing their training data. [paper] [episode]
- RDQ: Residual Distribution Quantization for Large Language Models This quantization method reduces the memory footprint of large language models by accounting for how data distributions shift through different layers. [paper] [episode]
- From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents This survey explores how tracking the step-by-step actions of AI agents can make them more reliable and easier to audit. [paper] [episode]
- Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC) This technique allows speech models to recognize new or specific words without needing long, slow prompts.
- HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention A new indexing method speeds up sparse attention mechanisms by using a two-stage search process to find relevant data faster. [paper] [episode]
- A Density-Matrix Framework for Electronic-Structure Analysis of Electrolytes for Lithium Batteries This AI platform uses density matrices to help scientists design better battery electrolytes by predicting molecular behavior. [paper] [episode]
- Analyzing LLM Reasoning to Uncover Mental Health Stigma Researchers analyzed the internal logic of language models to reveal hidden biases and stigmas regarding mental health. [paper] [episode]
- VeriSim: A Configurable Framework for Stress-Testing Medical AI Under Patient Communication Noise This framework simulates noisy, realistic patient speech to test how well medical AI handles real-world communication challenges. [paper] [episode]
- How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism This study suggests that following instructions is a result of coordinating various skills rather than one single universal mechanism. [paper] [episode]
- DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum This method improves privacy in machine learning training by using specialized mathematical updates to reduce noise distortion. [paper] [episode]
- SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations This approach allows neural networks to adaptively blend different activation functions for better performance across various tasks. [paper] [episode]
- Amulet: a Python Library for Assessing Interactions Among ML Defenses and Risks This library helps researchers see how different security defenses might accidentally create new vulnerabilities in machine learning models. [paper] [episode]
- TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas This agent uses a structured protocol to navigate unknown database structures to accurately convert natural language into SQL queries. [paper] [episode]
- GATTA: Graph Active Learning with Test-Time Augmentation This framework improves how AI learns from graph data by using multiple augmented views to get better estimates of uncertainty. [paper] [episode]
- Monotone Neural Policy Iteration for High-Dimensional Hamilton--Jacobi--Bellman Equations This paper presents a new way to solve complex control problems in high dimensions using neural networks and monotone operators.
- Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks This study provides a standardized way to compare how different AI routers select the best model for a specific task. [paper] [episode]
- Semidefinite Programming for Quantum Channel Learning This research uses convex optimization to efficiently reconstruct quantum channels from classical data samples. [paper] [episode]
- Optimizing Three Critical Factors for Practical and Effective OOD Detection Fine-Tuning This method improves how AI detects unfamiliar data by balancing accuracy with the ability to recognize outliers. [paper] [episode]
- A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems This review examines the evolving landscape of attacks targeting voice recognition systems, such as deepfakes and adversarial audio. [paper] [episode]
- Aslema at NADI 2026: Data Augmentation for Intent Recognition and Slot Filling This system uses synthetic data generation to improve how AI understands specific dialects like Tunisian Derja. [paper] [episode]
- EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale This framework allows AI agents to conduct long-term scientific research by maintaining a continuous memory of their experiments. [paper] [episode]
- Customized large language models can outperform Community Notes in correcting misinformation This research introduces a model that uses real-time evidence to correct misinformation more effectively than crowdsourced systems. [paper] [episode]
- A Lightweight Multi-Agent Framework for Automated Concrete Barrier Design This multi-agent system automates the design of highway barriers by coordinating specialized agents to follow strict engineering rules. [paper] [episode]
- An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc This benchmark evaluates whether AI can write high-performance scientific code that follows professional library standards. [paper] [episode]
- Prediction--Loss Alignment for Sampler--Robust Flow Matching Training This work shows how aligning prediction and loss functions makes diffusion models more stable during training. [paper] [episode]
- Assessing Predictive Models for Fairness Based on Activity-Space Patterns This approach measures fairness by looking at whether models treat people differently based on where they spend their time. [paper] [episode]
- Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups This study tests if language models understand the emotional nuances in news headlines across different demographic groups. [paper] [episode]
- 2.5-D Decomposition for LLM-Based Spatial Construction This method improves how AI builds 3D structures by letting a language model plan in 2D and a math engine handle the vertical placement. [paper] [episode]
- Domain Elastic Transform: Bayesian Function Registration for High-Dimensional Scientific Data This framework allows researchers to align complex scientific data, like gene expression, without needing to use a fixed grid. [paper] [episode]
- Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents This method helps AI agents learn how to actively explore their environments to find useful information for future tasks. [paper] [episode]
- Timely Clinical Diagnosis through Active Test Selection This framework uses Bayesian design to help clinicians choose the most informative medical tests at each step of a diagnosis. [paper] [episode]
- Generalization in VAE and Diffusion Models: A Unified Information-Theoretic Analysis This paper provides a single mathematical theory to explain how both VAEs and diffusion models generalize to new data. [paper] [episode]
- On the Societal Impact of Machine Learning This thesis explores how machine learning affects society and proposes ways to reduce algorithmic discrimination. [paper] [episode]
- mmWave Sensing with End-to-End Fully Homomorphic Encryption This system allows for secure radar sensing by performing all signal processing on encrypted data in the cloud.
- Evidence for Limited Metacognition in LLMs This research uses new methods to show that while large language models have some self-awareness abilities, they are limited compared to humans. [paper] [episode]
- Learning Intrinsic Water-Quality Dynamics with Rainfall for Data-Driven Forecasting This model uses rainfall data to help predict how water quality changes over time across different locations. [paper] [episode]
- CertDW: Towards Certified Dataset Ownership Verification via Conformal Calibration This method provides a way to mathematically prove if a model was trained on a specific protected dataset. [paper] [episode]
- Output Embedding Centering for Stable LLM Pretraining This technique improves the stability of training large language models by centering their output representations. [paper] [episode]
- Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu This study shows that fine-tuning small parts of a model is highly effective for detecting hate speech in languages like Roman Urdu. [paper] [episode]
- Using Seismic Statistical Features and VQ-VAE to Improve Spatiotemporal Seismicity Predictability This method combines earthquake statistics with deep learning to better predict where earthquakes will occur. [paper] [episode]
- Cognitive Digital Twins: Ethical Risks and Governance for AI Systems That Model the Mind This paper discusses the unique ethical risks of creating digital simulations of human cognition and how to govern them. [paper] [episode]
- SGA: Plug&Play Geometric Verification for Educational Video Synthesis This module checks AI-generated animation code to ensure that objects are placed correctly in 3D space without overlapping. [paper] [episode]
- FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving This framework speeds up large mixture-of-experts models by intelligently managing how expert weights are stored in memory. [paper] [episode]
- Verification of Adaptive Agentic Controllers through Finite Rule Revision This study proposes a way to test and fix errors in autonomous AI controllers using symbolic rules. [paper] [episode]
- SAC-Copula: Quality-Preserving Watermarking for Diffusion Language Models via Smooth Correlated Gumbel Fields This method adds watermarks to generated text that are harder to remove and do not ruin the quality of the writing. [paper] [episode]
- Perturbation: A simple and efficient adversarial tracer for representation learning in language models This approach uses small changes to input data to reveal how language models organize information internally. [paper] [episode]
- Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP This study compares different ways to modify AI behavior, finding that some methods are better at controlling how a model refuses harmful requests. [paper] [episode]
- Gradient-based Learning in State-based Potential Games for Self-Learning Production Systems This research introduces faster ways for multi-agent systems to learn optimal behaviors using gradient descent. [paper] [episode]
- Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage This framework uses AI to analyze police footage and identify patterns of interaction between officers and citizens. [paper] [episode]
- Some hypotheses on how chatbots work in problem-solution-driven conversations: Large Language Models as confirmation of the Innovation Illusion This paper argues that chatbots lack the cognitive flexibility of humans because they rely on imitating text patterns rather than true understanding. [paper] [episode]
- ExpTest: Loss-Curve Hypothesis Testing for Autonomous Learning-Rate Selection in Deep Neural Networks This tool automatically adjusts a model's learning rate by statistically analyzing its training loss curve. [paper] [episode]
- Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions This framework helps small language models solve complex scheduling problems by translating them into a format that math solvers can understand. [paper] [episode]
- Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding This method teaches AI to understand humor by training it to recognize the mismatch between an image and its funny explanation. [paper] [episode]
The papers
- CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models — CHRONOBERG is a specialized corpus designed to capture "Language Evolution and Temporal Awareness in Foundation Models." This resource is critical for advancing Natural Language Processing research by providing researchers with a mechanism to study how language changes over time, [episode]
- Longitudinal Risk Prediction in Mammography with Privileged History Distillation — Longitudinal risk prediction in mammography is a critical area of research aimed at improving cancer screening by moving beyond single-time point assessments. [episode]
- A Constraint Programming Approach for n-Day Lookahead Playoff Clinching in the NHL — As a diligent researcher who understands that any mistake could cost millions, my primary directive is accuracy. [episode]
- Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning — Hierarchical Reinforcement Learning (HRL) is critical for enabling agents to solve complex, long-horizon tasks by decomposing them into manageable subgoals. [episode]
- CQD-SHAP: Explainable Complex Query Answering via Shapley Values — The paper introduces "CQD-SHAP: Explainable Complex Query Answering via Shapley Values," a methodology designed to provide deep interpretability into complex question answering systems. [episode]
- Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records — The paper addresses the critical need for robust and interpretable predictive models when analyzing complex, longitudinal Electronic Health Records (EHRs) for clinical tasks. [episode]
- Smoothing the Score Function to Enhance Generalization in Diffusion Models — This paper introduces a novel geometric and optimization-based framework for understanding diffusion models, arguing that generalization failures stem from the model's inability to properly smooth the empirical score function. [episode]
- Warranted Attention: Learning What to Pass from Attention to Prediction — The paper investigates the complex relationship between attention mechanisms and semantic relevance in large language models, arguing that "Relevance Is Not Permission." It provides a rigorous framework for localizing and controlling specific metric-facing contributions within tr [episode]
- From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents — This survey paper addresses the "process-level accountability gap" emerging in Large Language Model (LLM)-based agents. [episode]
- RDQ: Residual Distribution Quantization for Large Language Models — This paper introduces RDQ (Residual Distribution Quantization), a post-training quantization (PTQ) framework designed to mitigate the sharp performance degradation observed in large language models (LLMs) when using sub-4-bit precision. [episode]
- LOGIC: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration — Contextual biasing is crucial for grounding large language models (LLMs) in specific domain knowledge or user context, addressing the inherent limitations of general pre-training. [episode]
- How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism — The ability of Large Language Models (LLMs) to adhere precisely to complex instructions represents a critical frontier in AI research, moving beyond simple text generation toward demonstrable reasoning and constraint satisfaction. [episode]
- Analyzing LLM Reasoning to Uncover Mental Health Stigma — The paper analyzes how Large Language Models (LLMs) generate reasoning traces when applied to mental health vignettes, focusing specifically on the presence and severity of stigmatizing language. [episode]
- A Density-Matrix Framework for Electronic-Structure Analysis of Electrolytes for Lithium Batteries — Please provide the full body text of the paper, "A Density-Matrix Framework for Electronic-Structure Analysis of Electrolytes for Lithium Batteries." The material you have provided consists only of supplementary data (Tables S12, S13, and S14) and a list of references. [episode]
- DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum — The paper details advanced methodologies for performing large-scale language model training while adhering to strict privacy constraints. [episode]
- HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention — The paper introduces HISA (Hierarchical Indexed Sparse Attention), a novel and efficient method designed to enhance sparse attention mechanisms for processing extremely long context sequences. [episode]
- VeriSim: A Configurable Framework for Stress-Testing Medical AI Under Patient Communication Noise — VeriSim introduces a novel and comprehensive computational framework designed to rigorously stress-test diagnostic Artificial Intelligence models when deployed in real-world clinical settings. [episode]
- TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas — This paper introduces TRUST-SQL, a novel framework utilizing "Tool-Integrated Multi-Turn Reinforcement Learning" to tackle the challenging task of Text-to-SQL generation when facing unknown database schemas. [episode]
- GATTA: Graph Active Learning with Test-Time Augmentation — GATTA (Graph Active Learning with Test-Time Augmentation) is a framework designed to enhance active learning for graph-structured data by improving the reliability of uncertainty estimates. [episode]
- SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations — The paper introduces SG-Blend, a novel activation function designed by interpolating between the improved Swish variant (SSwish) and the established Gaussian Error Linear Unit (GELU). [episode]
- Amulet: a Python Library for Assessing Interactions Among ML Defenses and Risks — I apologize, but the provided text appears to be an excerpt from a bibliography page (Page 13) containing citations for various machine learning security papers. [episode]
- Monotone Neural Policy Iteration for High-Dimensional First-Order Hamilton--Jacobi--Bellman Equations — The paper addresses the challenge of solving high-dimensional first-order Hamilton–Jacobi–Bellman (HJB) equations, which are fundamental in optimal control theory and reinforcement learning. [episode]
- Semidefinite Programming for Quantum Channel Learning — I apologize, but you have provided a bibliography section and command-line syntax, but not the actual text content of the arXiv paper titled "Semidefinite Programming for Quantum Channel Learning." As an AI researcher where any mistake could cost millions of dollars, I cannot gen [episode]
- Aslema at NADI 2026: Data Augmentation for Intent Recognition and Slot Filling — Please provide the scientific paper ("Aslema at NADI 2026: Data Augmentation for Intent Recognition and Slot Filling") you would like me to summarize. Once you provide the text, I will adhere strictly to your detailed requirements: 1. [episode]
- Optimizing Three Critical Factors for Practical and Effective OOD Detection Fine-Tuning — As a diligent researcher, I have reviewed the provided text. [episode]
- Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks — The paper presents a comprehensive and rigorous evaluation of four prominent open-source model routing systems—RouteLLM, Aurelio Semantic Router, LiteLLM Router, and vLLM Semantic Router—across a suite of complex benchmarks. [episode]
- A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems — I apologize, but you have provided only a section of a bibliography and not the actual content of the paper, "A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems." To fulfill your request—which requires extracting specific details, quoting key phrases, an [episode]
- EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale — I am prepared to execute this summary with maximum diligence and precision. [episode]
- A Lightweight Multi-Agent Framework for Automated Concrete Barrier Design — The paper introduces a novel methodology, a lightweight Multi-Agent Framework (MAF), designed to automate and enhance the process of concrete barrier design. [episode]
- Prediction--Loss Alignment for Sampler--Robust Flow Matching Training — The paper details advanced methodologies for training generative models using flow matching, specifically addressing stability and robustness in challenging domains like communication systems (MIMO detection) and binarized image processing. [episode]
- 2.5-D Decomposition for LLM-Based Spatial Construction — The paper "2.5-D Decomposition for LLM-Based Spatial Construction" addresses the critical challenge of enabling Large Language Models (LLMs) to perform complex, geometrically consistent spatial reasoning and construction tasks. [episode]
- An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc — I am unable to provide the summary for "An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc" because you have provided a bibliography of related papers rather than the full text of the article itself. [episode]
- Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups — The paper, "Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups," provides a rigorous evaluation of how large language models' alignment scores vary when assessed against human responses across diverse sociodemographic groups and complex geopolitical topic [episode]
- Customized large language models can outperform Community Notes in correcting misinformation — This paper introduces M USE, a scalable approach for "multimodal misinformation correction" designed to address the rapid spread of false or misleading content on social media. [episode]
- Assessing Predictive Models for Fairness Based on Activity-Space Patterns — This paper introduces a comprehensive methodology for assessing fairness in predictive models by analyzing "activity-space patterns," particularly focusing on movement data. [episode]
- On the Societal Impact of Machine Learning — This paper, titled *On the Societal Impact of Machine Learning*, provides an exceptionally comprehensive and multi-faceted analysis, bridging foundational theoretical concepts of algorithmic fairness with practical, actionable interventions designed to mitigate systemic societal [episode]
- Timely Clinical Diagnosis through Active Test Selection — This paper introduces ACTMED (Adaptive Clinical Test selection via Model-based Experimental Design), a diagnostic framework designed to emulate "real-world diagnostic reasoning." It addresses the critical need for "timely and cost-effective decisions" in clinical settings where d [episode]
- Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents — This paper introduces a novel framework designed to enhance Large Language Model (LLM) agents by equipping them with proactive exploration capabilities, addressing the inherent limitations of reactive decision-making in complex environments. [episode]
- Generalization in VAE and Diffusion Models: A Unified Information-Theoretic Analysis — This paper presents "Generalization in VAE and Diffusion Models: A Unified Information-Theoretic Analysis," offering a rigorous theoretical framework for understanding the generalization capabilities of two major classes of generative models: Variational Autoencoders (VAEs) and D [episode]
- Domain Elastic Transform: Bayesian Function Registration for High-Dimensional Scientific Data — As a diligent AI researcher, I must inform you that while I have fully internalized your required structure—including the opening orienting paragraph, 3-5 sections with bold headers, bulleted lists, and the strict word count and quoting requirements—the actual content of the [episode]
- Learning Intrinsic Water-Quality Dynamics with Rainfall for Data-Driven Forecasting — The paper, "Learning Intrinsic Water-Quality Dynamics with Rainfall for Data-Driven Forecasting," addresses the critical challenge of predicting water body quality parameters by explicitly modeling the complex, non-linear interactions between pollutant concentrations and meteorol [episode]
- Evidence for Limited Metacognition in LLMs — This paper investigates the extent of metacognitive abilities in Large Language Models (LLMs) by designing complex game-theoretic tasks that require models to introspect on their own certainty and decision-making processes. [episode]
- mmFHE: mmWave Sensing with End-to-End Fully Homomorphic Encryption — The paper "mmFHE: mmWave Sensing with End-to-End Fully Homomorphic Encryption" presents a novel framework for performing complex radar signal processing entirely within an encrypted environment. [episode]
- CertDW: Towards Certified Dataset Ownership Verification via Conformal Calibration — The paper addresses "Dataset Ownership Verification," aiming to establish certified robustness for verifying dataset ownership even when perturbations are applied. [episode]
- Output Embedding Centering for Stable LLM Pretraining — The paper addresses critical stability issues encountered during Large Language Model (LLM) pretraining, proposing novel techniques to enhance model robustness and performance. [episode]
- Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu — As a diligent AI researcher where accuracy is paramount, I am fully prepared to conduct this detailed extraction. [episode]
- Using Seismic Statistical Features and VQ-VAE to Improve Spatiotemporal Seismicity Predictability — As a diligent researcher, I recognize that extracting a detailed summary requires access to the full body of text from "Using Seismic Statistical Features and VQ-VAE to Improve Spatiotemporal Seismicity Predictability." The material provided contains only the bibliography section [episode]
- Cognitive Digital Twins: Ethical Risks and Governance for AI Systems That Model the Mind — The paper, "Cognitive Digital Twins: Ethical Risks and Governance for AI Systems That Model the Mind," provides a critical examination of Cognitive Digital Twins (CDTs)—advanced AI systems designed to model, predict, and simulate complex aspects of human thought, behavior, and [episode]
- SGA: Plug&Play Geometric Verification for Educational Video Synthesis — The paper introduces SGA, a "Plug&Play Geometric Verification" system designed to enhance educational video synthesis by rigorously detecting and correcting geometric inconsistencies in generated scripts. [episode]
- FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving — As a diligent researcher, I understand that accuracy and adherence to structure are paramount, especially when dealing with high-stakes technical documentation. [episode]
- Verification of Adaptive Agentic Controllers through Finite Rule Revision — The paper, "Verification of Adaptive Agentic Controllers through Finite Rule Revision," addresses the critical challenge of ensuring that self-modifying or adaptive automated controllers—particularly those operating in complex, dynamic environments like supply chains—maintain [episode]
- SAC-Copula: Quality-Preserving Watermarking for Diffusion Language Models via Smooth Correlated Gumbel Fields — The analysis presented rigorously evaluates the robustness and transferability of detection mechanisms across various language model sources, tasks, and backbones. [episode]
- Perturbation: A simple and efficient adversarial tracer for representation learning in language models — The paper "Perturbation: A simple and efficient adversarial tracer for representation learning in language models" introduces a novel method designed to enhance understanding of how language models encode meaning by systematically perturbing inputs. [episode]
- Gradient-based Learning in State-based Potential Games for Self-Learning Production Systems — This paper introduces novel gradient-based optimization methods for state-based potential games (SbPGs) within self-learning distributed production systems. [episode]
- Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP — The provided material consists entirely of quantitative metrics, ablation study results, and comparative numerical scores across various model configurations and modifications. [episode]
- Some hypotheses on how chatbots work in problem-solution-driven conversations: Large Language Models as confirmation of the Innovation Illusion — The paper investigates the underlying cognitive mechanisms by which Large Language Models (LLMs) operate when engaged in structured problem-solution dialogues, framing their performance through the lens of psychological biases. [episode]
- ExpTest: Loss-Curve Hypothesis Testing for Autonomous Learning-Rate Selection in Deep Neural Networks — I apologize, but you have provided a list of references and citations rather than the full text or abstract for the paper titled "ExpTest: Loss-Curve Hypothesis Testing for Autonomous Learning-Rate Selection in Deep Neural Networks." To generate the detailed, structured summary y [episode]
- Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage — This paper presents OpenBWC, a novel interdisciplinary framework designed to analyze police body-worn camera (BWC) footage using advanced artificial intelligence and statistical machine learning. [episode]
- Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions — The paper addresses the critical challenge of evaluating large language models (LLMs) on complex combinatorial optimization tasks, specifically within resource-constrained scheduling problems. [episode]
- Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding — The paper, "Learning to Think Like a Cartoon Captionist," addresses the complex challenge of multimodal humor understanding by proposing a novel training paradigm centered on incongruity resolution. [episode]
- Continuous Diffusion Scales Competitively with Discrete Diffusion for Language — The provided text segments cover unrelated topics including university collaborations (Noogame-UPS), local crime reports, and political healthcare legislation. [episode]
- XAI-Arena: Can LLMs Assess the Quality of XAI Explanations? —
- Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations —
- ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance —
- Multi-Agent Agentic Graph Learning via Structural Signatures —
- A Function-Space Approach to the Statistical Mechanics of Learning Dynamics —
- From State Synchronization to Cognitive Self-Evolution: An Operational Architecture for Cognitive Digital Twins —
- Seven Sources of Physical AI Capability Formation —
- RobustSGPO: Search-Space Control for Agent Harness Evolution —
- Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery —
- RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems —
- PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations —
- Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents —
- Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher Distillation —
- Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning —
- LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents —
- Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks —
- Online Inverse Integer Linear Optimization via Small-Gradient Skipping: Constant Regret and Finite Mistakes —
- In Medical Claims Data, Enhancing Predictive Performance for Major Adverse Cardiovascular Events Using Cross Attention —
- TempTPI: Informer-Based trajectory prediction for maritime vessels —
- The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents —
- Exact Degeneracy Under Balanced k-Shot Sampling:Consequences for Small-Sample Discriminant Analysis on LLM Embeddings —
- Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction Fields —
- AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents —
- Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format —
- Forward-Free LLM Depth Pruning via Weight Redundancy —
- Decision Transformer for UAV-Mounted RIS-Assisted Dynamic D2D Communications —
- ProMeta: Few-shot PROTAC-targeted degradation prediction across E3 ligases —
- Grounded Evaluation and Repair for NL-to-PDDL Problem Generation —
- Meta-LinEXP3: Online-within-Online Learning for Adversarial Linear Contextual Bandits —
- A Kernel-Based Modular Discriminant Analysis Framework for Small-Sample Learning —
- Development and Validation of a Physics-Guided Machine Learning Extrapolation Framework Using a Classical Transient Diffusion Benchmark —
- Multi-Pass, Multi-View Blended Learning for High-Fidelity Volumetric CT Synthesis from Chest X-Rays —
- Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models —
- Structural Process Supervision for Latent Chain-of-Thought Reasoning —
- Deep Neural Networks for Learning Intent from sEMG Signals to Support Hardware Devices for Post-Stroke Neurorehabilitation —
- An Explainable Machine Learning Framework for Predicting Blood-Brain Barrier Permeability Using Molecular Descriptors —
- Structure-Aware Unsupervised Anomaly Detection for Spacecraft Telemetry with Adaptive EVT Thresholding —
- Beyond Contact Sensors: Deep learning with Pseudo-Labeling for remote Photoplethysmography —
- Field-level prediction of mid-plane stress tensor fields in concrete target penetration: a cross-velocity graph neural operator surrogate —
- Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability —
- Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States —
- Hybrid Quantum-Classical NLP Classification with Compact Semantic Representations: An Experimental Analysis of Representation Compression —
- A Systematic Evaluation of Molecule Generation Models for De Novo Drug Design: From Benchmarks to Practical Insights —
- A Trust-Network-Based Federated Learning Framework for Multi-Center Aging Clock Prediction —
- Storage-Scalable Progressive Semantic Communication via Knowledge-Base Reuse —
- Agent-Based ML-LLM Fusion with Self-Optimizing Prompts for Plateau Weather Alerts —
- Kernel-Managed Shared Memory for System-Wide Personalization —
- CompassOPD: Cross-Family On-Policy Distillation via Within-Family Likelihood Shifts —
- CoGe-GCD: Reframing Generalized Category Discovery with Compositional Generalization —
- Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning —
- Robust Beam Prediction for V2X Networks with Multi-Modal Sensing —
- Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection —
- Hierarchical and Permutation-Invariant Feature Transformation Learning via Policy-Guided Embedding Search —
- What Should an Agent Forget? Separating What Is Stored from What Is Used —
- Training Trajectories Determine Circuit Removability in Annealable Soft-Prior Transformers —
- A Dominant Diffuse Phase in the Sparse Autoencoder Phase Diagram —
- View-Structured Conformal Prediction for 3D Gaussian Splatting —
- One Loop, Two Gains: Can Active Learning win the Lottery for Free? —
- TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards —
- Cyber-Financial Contagion: Modeling the Propagation of an AI Vendor Compromise Through the Banking System —
- A Later Test Set Is Not a New Domain: Pretraining Familiarity Survives a Contamination-Free Hold-Out —
- OmniMed-FL: A Robust Multimodal Federated Learning Framework for Clinical Diagnosis —
- Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs —
- Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs —
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition —
- Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization —
- Nonmaximal sums of maximally monotone operators under Rockafellar's constraint qualification —
- Learning with Covariance Matrices: Principal Component Analysis Meets Learning with Graphs —
- Quantum Feature Engineering for Credit Default Prediction: When and Why IQP Circuits Help Linear Classifiers —
- M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction —
- Probabilistic Focal Search: Accelerating Bounded-Suboptimal Search via Lower-Bound Advancement —
- Halo: Improving forecast accuracy through heteroscedastic estimation —
- An Empirical Measurement of Jailbreaking Evaluators —
- Threshold Choice, Not Sample Size, Bounds Trustless Verification of Nondeterministic Compound AI Workflows —
- Adaptive Diffusion Freezing: Privacy-preserving Diffusion Models Against Membership Inference Attacks —
- Black-Box Membership Inference via Word-Level Probability Estimation —
- PEARL: A Task-Aware Framework for Evaluating Differentially Private Synthetic Educational Data —
- Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting —
- SoK: Privacy Attacks on Machine Learning via Explainable AI —
- QuantumQUBO Agent: Automating Quadratic Unconstrained Binary Optimization (QUBO) Formulation Generation from Natural Language —
- From Cycle Space to Cycle Manifold: Limits and Achievability of Blind False Data Injection Attacks —
- HermiCache: Enclave-Aware Cache Replacement for Trusted Execution Environments —
- Zero-shot rib design: merging training-free generative prior with topology optimization —
- Byzantine-Robust Federated Fire Detection with a Rotating Coordinator —
- Artificial Intelligence Algorithms for the Detection of Pathologies Related to Lung Cancer through Image Analysis using Convolutional Neural Networks and Data Augmentation: a systematic mapping of the literature —
- A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive Reasoning —
- Understanding LoRA Rank Trade-offs in Diffusion Model Fine-Tuning —
- Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking —
- GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models —
- Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement —
- Architecting the Secure AI-SOC: A Neurosymbolic Framework for Pipeline Integrity and Threat Mitigation —
- An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics —
- NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction —
- CMNIE: An Information Extraction Benchmark for Chinese Military News —
- Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge —
- Towards a Deterministic Math Solver for Clinical Language Models —
- Conformal Calibration Transfer —
- The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes —
- CARTS: Contextual Autoregressive Rank Transcoding Steganography for Full-Capacity Keyed Text Encoding —
- Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking —
- Temporal and Multimodal Deep Learning for Cyberattack Detection in LEO Satellite Systems —
- Adaptive Margin Ordinal Loss: Penalizing Center-Class Hedging in Ordinal Classification —
- Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu —
- Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitive and LLM-Generated Python Code —
- Weighted Empirical Risk Minimization for Machine Learning under Long-Range Dependence: Exact Pathwise Rates and Learning-Error Geometry —
- A Bellman Optimality Equation for Plasticity —
- Counterfactual Marginalisation: Framework for Evaluating Robustness to Nuisance Variables —
- Big Enough to Break Out: Tracking the Rising Capability of LLM Penetration-Testing Agents —
- From Connectivity to Rewards: Dense Reward Learning with Directed State Graphs —
- Analyzing Traditional and Neural Approaches to Multilingual Readability Assessment —
- DR-LabStack: Design and Implementation of a Clinician-Facing Web System for Diabetic Retinopathy Prediction —
- RiVaT-Fuse: Reliability-Calibrated Variational Tensor Fusion with Matrix-Valued Trust for Multimodal Prediction under Modality Uncertainty —
- Larger Context Window, Fewer Overcorrections: Optimizing Prompts and Batching for Minimal-Edit Grammatical Error Correction —
- Studying Without a Syllabus: Task-Agnostic Environment Preprocessing —
- Processing and classifying bird songs using wavelet techniques and supervised learning —
- Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models —
- No-Box Vulnerability Analysis: Description-only Detection of Indirect Prompt Injection Vulnerabilities in MCP Servers —
- Flow Duality and Source Geometry for Categorical Generation —
- Certifying Lower Bounds for Risk-Sensitive Reinforcement Learning under Adversarial State Perturbations —
- A2ABreak: Systematic Security Analysis of the A2A Protocol —
- When Validation Stops Learning: Auditing Update Admission for Continual Embodied Agents —
- Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry —
- AspisAI: A Canonical, Machine-Interpretable Governance Framework for Automated Multi-Standard Compliance Monitoring —
- Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble —
- Relatively Smart II: Tractable or Semi-Supervised Instance-Optimal Learning —
- DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents —
- Does Linguistic Structure Enrichment Enhance Coherence Assessment? Not With Current Architectures —
- LLM-Anchored Paralinguistic Enrichment for Alzheimer's Disease Detection —
- SearchAtlas: Analyzing Agentic Search Strategies via Evidential Query Graphs —
- Auto-RecSys: Harnessing Autonomous Research Agents for Industry-Scale Recommender System —
- Structurally Speaking: Motif-Oriented Graph Captioning through Bidirectional Graph-Text Translation —
- AUC Maximization from Biased Positive-unlabeled Data with Confidence —
- Using Semantic Uncertainty to Estimate Transition Relevance in Turn-taking —
- Empirical Evaluation of Membership Inference Attacks on NLP Text Classifiers: A Baseline Study on SST-2 —
- Robust Multimodal Sentiment Analysis with Incomplete Modalities via Semantic-aware Completeness based Reconstruction —
- Empirical Evaluation of Data Poisoning Attacks in Supervised Learning —
- Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation —
- When More Is Not Better: Component Anti-Synergy in a P300 Speller —
- Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows —
- Phases in a class of associative memories via hidden neurons —
- EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale —
- Thompson Sampling for Non-Monotone Convex Ridge Bandits: Monotonicity Is Not Needed for Polynomial Regret —
- Demystifying the Privacy-Utility Trade-off in LLM Interactions —
- Distribution-aware Language Neuron Identification in Multilingual Large Language Models —
- Importance Weighting for Unlabeled-unlabeled Learning under Distribution Shift —
- Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models —
- DeFiFusion: Combining Transaction Events with Smart Contracts to Detect Price Manipulation Attacks —
- Topological Necessities: Mechanism-Invariant Strategic Subgoals for Cross-Embodiment Goal-Conditioned Control —
- Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks —
- K/V-Cache Interventions Dissociate Representation Alignment from Persona Expression in Decoder-Only Language Models —
- The Missing Boundary: How Autonomous Agents Lose Control —
- BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure —
- Rebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss —
- The Agent Incident Registry: Toward Preventing Repeated AI Agent Failures —
- T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks —
- EMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression —
- Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents —
- Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning —
- The information geometry of large language models is shared, learned, and controllable —
- MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG —
- When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text —
- ToxicRAG: Compromising Retrieval-Augmented Generation Systems via Single-Shot Knowledge Poisoning Attacks —
- Beyond Solver Verdicts: Generative Reward Models for Autoformalization —
- ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute Mediation —
- Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation —
- Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers —
- HERALD: High-Fidelity Exemplar Retrieval with Adaptive Landmark Distillation for Heterophily-Aware Graph Condensation —
- KuaiRP Series Role-playing Models Technical Report —
- From Repetition to Recognition: Inductive Discovery of Disinformation Narratives —
- Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting —
- How Wrong Can a Good Predictor Be? Diverging Updates with Vanishing Predictive KL —
- Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving —
- Bidirectional Multimodal Fusion of Sky Images and Time-Series for Solar Forecasting with Large Language Models —
- The Machines Are Calling: Measuring Automated and Synthetic Voices in Unwanted Inbound Calls —
- Can LLMs Normalize Databases? A Benchmark and Multi-Agent Framework for Schema Normalization —
- Human Agreement and Return Association Are Not Interchangeable Criteria —
- A Dominant Supplier Slows Recursive Drift More Than It Steers It —
- Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation —
- Break Step: Recursive Training Resonates with Replayed Sampling Noise —
- DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat —
- LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry —
- When does a spectral prior help graph learning? Connectivity-loss estimation under road-network disruptions —
- Semi-Tensor Product-Based Multi-Term Randomized T-SVD and Its Visual Applications —
- From Explanations to Interventions: Execution-Guided Counterfactual Synthesis in Temporal Graphs —
- Hierarchical Clustering Can Jointly Satisfy Richness, Consistency, and Scale Invariance —
- Debate-to-Skill: Capability-Bound Process Supervision for Industrial Query-to-Agent Annotation —
- SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics —
- Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment —
- Agentic Share-of-Search: A Multi-Agent AI System for Competitive Decision-Making in LLM-Mediated E-Commerce —
- FlexComp: One Model for Every Ratio in Context Compression —
- An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning —
- Automated Identification of Competing Narratives in Political Discourse on Social Media —
- CryptoL: Towards Scale Dominance and Physics Constraints Mitigation in Financial Multivariate Time Series Forecasting —
- Convex Optimization with Nested Evolving Feasible Sets (CONES) under Time-Varying Loss Functions —
- REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving —
- Legible Failures: Detecting and Repairing In-Context Binding Errors —
- You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements —
- Polyhedral Geometry of Time-to-First-Spike Neural Networks —
- Solving Few-Shot Multiobjective Multitask Optimization via Iterative Sequential Transfer —
- A Voice-Interactive Multi-Agent System for Smart Operating Rooms: Architecture Design and Key Technologies —
- NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment —
- Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents —
- OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models —
- Assessing the Reusability of Public Speech Resources for Low-Resource Languages: A Central Kurdish Case Study —
- The Illusion of Balanced Multimodal Sentiment Analysis: Beyond the Limits of Optimization-Based Methods —
- You've Got a BUD in Me: Authenticated Reads from Per-Block Write Logs —
- MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions —
- AI-Powered Flare Combustion Efficiency Estimation —
- Few-Shot Learning for Network Intrusion Detection: Methods, Datasets, and Performance —
- Predicting Train Delays in Finland Using Machine Learning and Weather Data —
- Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 1 —
- When Does Text Inform? Benchmarking Information-Theoretic Metrics for Multimodal Time-Series Forecasting —
- Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data —
- Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model —
- Memory Compression for High-Fanout Agent Sandboxes —
- A Hilbert-Valued Functional Decomposition Framework for Explaining Time-Dependent Outputs —
- Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation —
- A Dynamic Fusion Large Language Model for Traffic Flow Prediction —
- Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models —
- Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents —
- Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification —
- AI Exposure and AI Resilience: A Two-Dimensional Assessment Framework for Software and Software-Based Business Model —
- MultiHuSE: A Multimodal Dataset for Humour Styles and Emotions —
- Estimating Inconsistency Response Surfaces under Uncertainty in Cyber-Physical System Development —
- E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets —
- On the Impact of Anonymization on the Performance of Large Language Models —
- Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding —
- Reification as a Transferable Vocabulary: Zero-Shot Link Prediction with Vanilla GNNs —
- SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ —
- Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells —
- Local Robustness Quantification for Naive Bayes Classifiers and Generative Forests: a General Approach —
- RAMamba-Net: A Reliability-Aware and Mamba-Based Multimodal Fusion Network for Auditory Attention Detection —
- Beyond Confidence: Stability-Aware Test-Time Adaptation for LLM Reasoning —
- TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs —
- From Queries to Narratives: Cultural Heritage Data Stories for Knowledge Graph Exploration and Quality Assessment —
- Deep-Fake CAPTCHA: Mitigating Next-Generation Social Engineering Attacks —
- "They don't care about this": A Systematic Study of TEE Build Reproducibility in the Wild —
- Heterogeneous Cross-Chain Transaction Tracing for Solana Bridges via Candidate-Set Selective Decision —
- SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations —
- LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician-Evaluated Case Study —
- Chypothermia: Clock Freezing for Static Side-channel Attacks —
- Calibration-Aware Uncertainty Cascades for Efficient Heterogeneous Model Collaboration —
- Prevalence Determines Precision:Silent Contamination in Detector-Defined Datasets —
- Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study —
- RouteRepair: Instance-Level Failure Diagnosis and Targeted Repair in LLM-Based Automated Heuristic Design for Routing Optimization —
- Flexible and Interpretable Accent Distance Measurements —
- ReGround: Grounding Reviewer Comments in Multimodal Evidence —
- The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation —
- Published Unlearning Numbers Move Per Checkpoint, and Not Because the Removed Data Survives: An Audit of 263 Released Batch-Normalized Checkpoints —
- From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development —
- Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study —
- ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps —
- DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis —
- Structural priors for data-efficient language learning —
- Extending SMT Solving with Non-Ground Clause Learning —
- Generalised Score Matching on Convex Domains —
- Risk-Averse Decision Making with Multi-Level Reliability Guarantees —
- Lightweight LiDAR-Based Cone Detection Framework Using Random Forest for Formula Student Driverless —
- Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems —
- On Identifying Sound Conditions for Frontrunning Resistance —
- Particle GFlowNets: Rethinking Generative Marginalization Models —
- Characterizing Job Power Elasticity for Power-Flexible AI Training —
- Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech —
- Accountability in Certificate Transparency and Variants —
- Enabling Knowledge Graph Understanding at Scale with the EXplore Your Graphs ENgine (EXYGEN) —
- A Dataset and Model for Imputing Water Surface Elevation on a Large and Extremely Sparse Spatiotemporal Graph —
- CHOIR: heterogeneity-aware conformal prediction for crash injury severity across driver safety strata —
- From Intent to Execution Grant: An Execution-Boundary Conformance Profile for High-Risk AI Actions —
- Identifiability of Nonnegative Tensor Decompositions via Positive Scattering —
- Making Alternative Data Work: Context-Augmented LLMs for Financial Forecasting —
- PHAT: PHotonic Accelerator for TFHE —
- Distributed Optimization of Modular Production Systems using Model-based Reinforcement Learning with Inverse Models —
- A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC Filings —
- MAPLE: Memory-Augmented Planning with Language and Evolution —
- LoaDiff: Conditional Generation of Electricity Consumption Time Series for Energy Analytics —
- RDDMPI: Residual Denoising Diffusion Model for Probabilistic Multivariate Time Series Imputation —
- Musec: MomentUm SpEctral Clipping for Stable Muon-type Training —
- Learnware and AI Model Management System —
- Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents —
- Geospatial AI, Dataverse Metadata, and the Study of Place-Based Government —
- COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization —
- Structured Transforms for Low-Overhead Quantization of Language Models —
- Negative Self-Distillation: Learning to Reason by Avoiding Flaws —
- When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making —
- Generalization Analysis of Distributed Kernel-based Robust Gradient Descent Algorithms —
- Why Does Post-Training Quantization Work? —
- The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge —
- LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation —
- SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control —
- Signing the Transaction but Not the Decision: Whisper Attacks and a Binding Defense for AP2 —
- RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety —
- Component-Aware Differential Privacy for Federated Multilingual Speech-LLMs —
- A Unified Per-Token Gating Family for On-Policy Distillation: FKL/RKL Mixing with Multi-Channel and Bias Coefficients —
- Recognizing Is Not Reversing: A Controlled Inversion Test of Fact-Preserving News Framing —
- The widening evaluation gap in medical large language model research 2023 to 2026 —
- Differentially Private EEG Feature Anonymization: A Privacy-Utility Case Study in Clinical Neurophysiology —
- Predicting Privacy Leakage from Weight Spectral Density —
- Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech —
- Dynamic language model representations for multi-objective reaction optimisation —
- SpecGuard: Inference-Time Backdoor Detection For Free —
- Thinking with Looped Flows —
- Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead —
- Don't Trust the Super-App: A Case Study of Russia's Max —
- Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models —
- Atlas: Efficient Verifiable Semantic Search —
- Model-Aware Schedules Improve Generation via Fiberwise Optimal Transport —
- IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing —
- BlueSTAR: Tiered Agentic Architecture for Autonomous Cyber Defense —
- From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge —
- Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption Models —
- Epistemic orientation predicts legislative effectiveness among members of the US Congress —
- AdamX: Cosine similarity meets gradient descent —
- Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model —
- Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM Forecasting —
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement —
- On the Regularization Landscape for the Linear Recommendation Models —
- Domain-Specific Hallucination Detection in Large Language Models —
- From Specs to Apps: Verifying and Monitoring Models of Signal and WhatsApp —
- CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search —
- Nuha-Speech: Building General-Purpose Arabic Speech-LLMs —
- CausalArena: Benchmarking Causal Discovery in the Foundation Model Era —
- MindTopo: Can Foundation Models Reason in Topological Space? —
- TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription —
- From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good —
- Artificial Id: Drive and Persistent Alignment in Agentic AI —
- Distance generalization in transformers: why bother with positional encoding? —
- Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact —
- Can Edge-Deployable Vision-Language Models Identify Species? —
- Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data —
- A Short Survey of Viewing Large Language Models in Legal Aspect —
- General Quantification of Covariate and Concept Shifts —
- Label Differential Privacy via Aggregation —
- DNA: Differentially private Neural Augmentation for contact tracing —
- Self-optimization in distributed manufacturing systems using Modular State-based Stackelberg Games —
- Explainable few-shot learning workflow for detecting invasive and exotic tree species —
- Mapping Seven Decades of Philosophy in Colombia: Dynamic Topic Modelling of Ideas y Valores —
- Sublinear Variational Optimization of Gaussian Mixture Models with Millions to Billions of Parameters —
- The observational partial order of causal structures with latent variables —
- Bayesian Optimization of a Lightweight and Accurate Neural Network for Aerodynamic Performance Prediction —
- Near-optimal estimates for the p-Lipschitz constants of deep random ReLU neural networks —
- "Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated —
- The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies —
- BiHDTrans: binary hyperdimensional transformer for efficient multivariate time series classification —
- Test time training enhances in-context learning of nonlinear functions —
- Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection —
- Autonomous-Flow-Based Generation —
- Do Vision-Language Models Understand Visual Persuasiveness? A Diagnosis via Visual Persuasive Factors —
- Statistical analysis of Inverse Entropy-regularized Reinforcement Learning —
- When do cheap embeddings beat protein language models? A theoretically-grounded hashing sketch for biological sequence classification —
- UBCL: A Reinforcement Learning Framework for Controllable and Diverse Player Behaviors —
- DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports —
- engGNN: A Dual-Graph Neural Network for Omics-Based Disease Classification and Feature Selection —
- Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation —
- Evaluating Memory Structure in LLM Agents —
- Partial GFlowNet: Accelerating Convergence in Large State Spaces via Strategic Partitioning —
- What Language is This? Ask Your Tokenizer —
- Probing for Knowledge Attribution in Large Language Models —
- Bilateral Trade Under Heavy-Tailed Valuations: Minimax Regret without a Variance Bound —
- Rescaling Confidence: What Scale Design Reveals About LLM Metacognition —
- Streaming Translation and Transcription Through Speech-to-Text Causal Alignment —
- Surprised by Attention: Predictable Query Dynamics for Time Series Anomaly Detection —
- Measuring Progress in Reasoning Toward Mathematical Discovery with Automatic Verification —
- Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction —
- Alignment Reduces Expressed but Not Encoded Gender Bias: A Unified Framework and Study —
- Generalization Guarantees on Data-Driven Tuning of Gradient Descent with Langevin Updates —
- Timing is Everything: Temporal Scaffolding of Semantic Surprise in Humor —
- A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation —
- Privacy Auditing with Zero (0) Training Run —
- Goal-Oriented Lower-Tail Calibration of Gaussian Processes for Bayesian Optimization —
- Cross-lingual brain-language model alignment is robust but challenges hierarchical and computational accounts —
- Benchmarking non-conformity score functions in conformal prediction —
- MERIT: Matching Expertise via Rubric-Informed Training for Reviewer Assignment —
- Activation-Based Active Learning for In-Context Learning: Challenges and Insights —
- Attention by Synchronization in Coupled Oscillator Networks —
- Characterizing Narrative Content in Web-scale LLM Pretraining Data —
- Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue —
- SafeImpute: Reliable Clinical Data Imputation via Conformal Selection —
- Lower Bounds for PIR with Preprocessing from Blackbox Cryptography —
- Ceci n'est pas une pipe: AI systems as semantic abstractions —
- A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books —
- Transfer Learning of Keystroke Dynamics for Cross-Device User Authentication —
- Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis —
- Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment —
- Toward a First-Principles Update Geometry for the Language-Model Head —
- DelistBench: Evaluating Search-Enabled LLMs for Auditable Corporate-Event Database Completion —
- Adaptive Entangled Game Modules in Artificial General Intelligence —
- Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions —
- An Autonomous GeoAI Agent for Arctic Eco-Navigation —
- The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents —
- Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery —
- Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration —
Important terms
- CHRONOBERG
- A major new resource containing 250 years of English books from Project Gutenberg. It uses temporal annotations and valence-arousal-dominance analysis to help teach AI models how word meanings and sentiments change over long periods of time.
- SEM-HD
- A framework designed for medical imaging that uses a patient's full historical data as privileged information during training. This teaches a student model to make accurate risk predictions even when only a single current scan is available.
- TRUST-SQL
- A framework that treats interacting with massive enterprise databases as a partially observable process. It uses dual-track reinforcement learning to help AI agents actively hunt for and verify metadata rather than relying on pre-loaded information.
- EvoMaster
- A scientific discovery framework that uses loop research to prevent AI agents from repeating mistakes. It connects execution, exploration, and evolution into one continuous process, allowing evidence to persist across long-term experiments.
- FluxMoE
- An efficiency system for Mixture-of-Experts models that uses expert paging. Instead of keeping every expert in GPU memory, it streams weights on demand from host memory, significantly boosting throughput during long conversations.