AI papers — 2026-09-07
Today’s briefing focuses on how we can make AI systems more reliable and efficient by understanding their internal reasoning and decision-making processes. We start with a significant breakthrough in training large language models, where researchers discovered that solution divergence—the variety of different ways a model can solve the same problem—is actually a predictor of intelligence.
By using this diversity as a metric to guide training, they were able to consistently boost success rates across multiple domains, essentially teaching models that there isn't just one way to get the right answer. This theme of improving reliability continues with new ways to manage how agents use their skills.
Instead of relying on imperfect, one-shot instructions, a new framework called SkillRevise allows agents to iteratively fix their own procedural errors by looking at execution evidence and pulling repair principles from memory. This approach nearly doubled success rates on certain benchmarks and produced skills that actually work across different environments.
Moving into the technical infrastructure of these models, we see efforts to make their reasoning more honest through better uncertainty estimation. A large-scale study across twenty-two languages found that if you ask a model to reason in English even when the question is in a low-resource language, its ability to know when it is wrong improves significantly.
This suggests the bottleneck for multilingual models isn't understanding the question, but rather the generation process itself. Finally, we have a look at how these models handle specialized data and complex math.
New protocols for fuzzy private set intersection are making it much faster and cheaper to find matching elements in encrypted datasets, outperforming previous methods by thousands of times in some cases. At the same time, new benchmarks like TeleTables show that while models are getting better at reading technical tables, they still struggle deeply with the complex reasoning required for highly specialized fields like telecommunications.
If you want to understand why deep learning actually works, you have to look past simple curve fitting toward how features are built layer by layer. A new theoretical framework called Neural Low-Degree Filtering suggests that training is essentially a spectral procedure where each subsequent layer selects directions that have the highest correlation with the labels.
This provides a mathematical way to see how concepts emerge from raw data through compositionality, moving us beyond the "lazy" regime where weights barely move. While theory explains the mechanics, practical breakthroughs are happening in how we interpret messy biological signals.
Researchers have successfully turned noisy nanopore sensing data into image-classification tasks by using continuous wavelet transforms to create scaleograms. This approach reached 82% accuracy in identifying peptides and proved robust enough to work on embedded hardware even when half the model weights were zeroed out.
This trend of making complex environments more manageable extends to how agents interact with computers. The EvoCUA-1.5 framework moves away from static imitation learning by using online reinforcement learning to let agents learn through trial and error in sandbox environments.
By using a specialized step-level optimization, these agents achieved 63.2% success on the OSWorld benchmark, proving that they can learn to navigate complex desktop tasks through direct interaction rather than just following pre-recorded traces. If we want to make agentic workflows actually practical, we have to solve the massive bottleneck of evaluating them.
Right now, finding the best workflow requires running it over and over again, which is incredibly slow and expensive. A new framework called GLOW changes that by predicting how a workflow will perform before you even run it.
It does this by combining graph neural networks to understand the structure with large language models to grasp the semantic meaning of the agents involved. When researchers plugged GLOW into an automatic generation framework, they saw optimization times drop by a staggering 98.7% while barely sacrificing any accuracy.
This need for efficiency isn't just about workflow design; it is becoming a central theme in how we manage large language models more broadly. A new survey on the reasoning economy highlights this tension, noting that while "System 2" deep thinking makes models smarter, it creates a massive computational burden.
The researchers argue that finding the sweet spot between high-performance reasoning and limited budgets is one of the most critical challenges facing the field today. Safety is another area where we are seeing a push toward making complex processes more efficient.
Instead of having a model write out its entire thought process to explain why it flagged something as unsafe, a new method called COLAGUARD moves that reasoning into a continuous latent space. This allows the guardrail to be 12.9 times faster and use significantly fewer tokens while still matching the accuracy of models that explain their reasoning explicitly.
While we work on making these models safer and faster, we also have to deal with the growing problem of detecting what is actually human-written. A system called NotAI.AI tackles this by moving beyond simple "yes or no" labels to provide actual explanations for its detections.
By using a mix of sentence curvature and stylometric features, it can tell you exactly which parts of a text look machine-generated, achieving an F1 score of 0.9685 on test data. We need to rethink how we evaluate whether a model actually "knows" something or is just getting lucky, especially when it comes to complex reasoning.
Researchers have introduced KCSAT-ML, which uses a decade of Korean math exam data—complete with real error rates from hundreds of own thousands of students—to see if models make the same mistakes humans do. They found that even when models have similar accuracy scores, they behave very differently; some struggle with the hard stuff like humans, while others fail on easy problems that humans find trivial.
This distinction is vital because it exposes how scaling up computation doesn't always lead to smarter reasoning. In fact, researchers noticed a strange "overthinking" effect where increasing test-time scaling actually makes models perform worse on easier items within the same family.
The struggle to capture true capability is also evident in how we route tasks between different models. Since LLM responses are stochastic and can change even with slight phrasing tweaks, relying on a single answer to decide which model should handle a prompt is incredibly unstable.
To fix this, the DARS framework looks at the whole distribution of possible outcomes—including cost and variability—rather than just one sample, which makes routing much more reliable. This need for precision extends into specialized fields like medicine, where we want high performance without the massive cost of retraining models from scratch.
The OpenMedLM platform shows that clever prompt engineering can actually beat out expensive fine-tuning for open-source models. By using strategies like chain-of-thought and ensemble voting, they pushed an open-source model to 72.6% accuracy on MedQA, proving that we can get state-of-the-art medical expertise out of accessible models without needing massive computational budgets.
If we want to know if AI is becoming a social liability, we have to ask if it can actually think like a conspiracy theorist. Researchers found that large language models show partial agreement with conspiratorial mindsets and are surprisingly easy to manipulate into adopting these views through targeted prompting.
Even more concerning is that conditioning these models with certain socio-demographic attributes reveals latent biases in how they process misinformation, which could make them dangerous tools for reinforcing harmful narratives if deployed in sensitive contexts. This susceptibility to manipulation isn't just a matter of text; it extends to how agents behave when they have to work together under pressure.
A new benchmark called GPTNT uses the cooperative game Keep Talking and Nobody Explodes to test multimodal agents who must communicate asynchronously to defuse bombs. Despite their individual intelligence, not a single model tested was able to defuse a bomb in real time, revealing massive failures in state tracking and error recovery when these models are forced to act outside of simple turn-taking.
The difficulty of coordinating complex tasks is mirrored by the difficulty of writing the code that powers them. A new evaluation framework called KernelGenBench shows that while agentic systems can write specialized hardware kernels, their success is incredibly inconsistent across different chips and software sources.
For instance, an agent's accuracy dropped from 87 percent on NVIDIA hardware to just 25 percent on Iluvatar CoreX, proving that being good at coding for one platform doesn't mean a model is ready for real-world deployment elsewhere. Efficiency is the next logical hurdle, especially when these models are used to retrieve information.
A method called CacheWeaver tries to solve the high cost of retrieval-augmented generation by reordering evidence so that overlapping information can be reused in a cache. By using a simple greedy policy to place the most reusable prefixes first, they managed to cut the time it takes for a model to start speaking by up to 33 percent without losing any answer quality.
We need to move past simply checking if an AI's answer is correct and start measuring how it actually feels to use these models in real time. A new framework called QoNext finally addresses this by bringing networking principles like latency and generation velocity into the evaluation of foundation models.
By building a dedicated database and a neural predictor, researchers can now estimate human satisfaction based on measurable system parameters rather than just static text quality. This shift toward more nuanced evaluation is mirrored in how we understand model limitations, specifically regarding how specific an instruction actually is.
Standard vector embeddings often fail to distinguish between a broad prompt and a highly specific one because they focus too much on topic rather than detail. A new method called Prompt2Box uses box embeddings to capture these specificity relations, which helped identify 13.5% more weaknesses in seventeen different LLMs compared to traditional methods.
While we work on understanding model behavior, we also have to grapple with the legal reality of how these models are built and deployed. There is a growing problem with "models in the dark," which are downstream versions of AI created without enough transparency to honor GDPR rights like rectification or erasure.
Because machine learning exists in complex supply chains, enforcing these privacy rights remains a massive technical and legal hurdle. The complexity of these systems also extends to how they generate content, whether it is a storyline for a game or an entire terrain.
Generative AI has become a powerhouse for procedural content generation, but the field is still struggling to find enough high-quality, diverse training data to keep these models performing well. Even when we use math to explain the world through symbolic regression, we are finding that our current tools are too blunt.
A new approach called Deep Divide-and-Reduce in Symbolic Regression moves away from brute-force searches for sub-expressions and instead uses formal decomposition to find more complex mathematical patterns. We see similar struggles with structural integrity in other data formats, such as the massive taxonomic hierarchies found in Wikidata.
Researchers have developed a new validation method to hunt down classification errors and redundant links, creating a system that lets users inspect these relationships to clean up the graph. In the realm of specialized engineering, like carbon capture and storage, we are finding better ways to optimize expensive simulations by respecting physical symmetries.
A new Gaussian Process kernel called GP-Perm allows Bayesian Optimization to handle unordered sets of data more efficiently, which is vital when managing groups of wells where the specific order doesn't change the underlying physics. Finally, there is a way to make graph neural networks more stable and capable of seeing long-range connections without losing detail.
By integrating reservoir computing with structured convolutions in a model called RGC-Net, researchers have managed to stop node embeddings from becoming indistinguishable, leading to better performance in tasks like modeling how brain connectivity evolves.
Today's papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving Higher solution divergence in an LLM's outputs is linked to better overall problem-solving abilities. [paper] [episode]
- SoK: AI-Augmented Binary Reversing This paper provides a comprehensive overview and taxonomy of how machine learning is used to assist in analyzing binary code. [paper] [episode]
- Efficient Fuzzy PSI under One-Sided Assumptions A new method makes private set intersection more efficient by using lightweight cryptographic primitives under specific assumptions. [paper] [episode]
- SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision This framework helps AI agents iteratively improve their procedural skills by diagnosing and fixing errors found during execution. [paper] [episode]
- Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs This study finds that prompting models to reason in English significantly improves their ability to express uncertainty across many different languages. [paper] [episode]
- TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation This new benchmark tests how well language models can interpret the complex technical tables found in telecommunications standards. [paper] [episode]
- ConfRAG: Confidence-Guided Retrieval-Augmented Generation This method reduces hallucinations and saves computation by only using external retrieval when the model expresses uncertainty.
- Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes This research analyzes how relaxing verification rules during fast text generation can unintentionally degrade the quality of the output. [paper] [episode]
- Consensus Group Relative Policy Optimization for Text Generation This training method allows models to learn from consensus across multiple samples without needing expensive human-labeled data. [paper] [episode]
- Longitudinal Adoption and Deprecation of the Privacy Sandbox Web APIs Researchers analyzed seven years of data to see how web actors actually adopted Google's privacy-enhancing technologies before they were canceled. [paper] [episode]
- Improving Weak World Models Behind Strong Agents in Atari Pong This study shows that even high-performing AI agents often rely on flawed internal models of their environment and proposes a way to fix them. [paper] [episode]
- RL-VLA cubed: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training This framework speeds up training for robotic agents by allowing different parts of the system to work asynchronously. [paper] [episode]
- Unified Deployment-Aware Evaluation of Open Reasoning Language Models This paper argues that evaluating reasoning models should focus on practical deployment trade-offs like speed and memory rather than just accuracy scores. [paper] [episode]
- Advancing Subseasonal Forecasting with Machine Learning A new machine learning framework significantly improves weather forecasts for the two-to-six week window by correcting systematic biases. [paper] [episode]
- Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism This study shows that even when configured to be consistent, GPUs cause small variations in the probability of specific words being chosen. [paper] [episode]
- A Survey on Semantic Modeling for Building Energy Management This survey reviews how structured data models can help make building energy management more automated and interoperable. [paper] [episode]
- Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation This new architecture improves how AI recognizes emotions in conversations by separately modeling individual speakers and the overall dialogue context. [paper] [episode]
- Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning This theoretical work explains how deep networks learn complex features by treating training as a mathematical filtering process. [paper] [episode]
- Deep Learning-Driven Peptide Classification in Biological Nanopores This method turns biological sensor signals into images so that powerful deep learning models can identify peptides more accurately. [paper] [episode]
- EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents This framework enables AI agents to learn how to use computers more effectively by interacting directly with live software environments. [paper] [episode]
- Multi-Modal Time Series Prediction via Mixture of Modulated Experts This approach uses a specialized architecture to combine time series data with text information for better forecasting. [paper] [episode]
- From Architecture to Output: Structural Origins of Hallucination in Large Language Models This research investigates which specific parts of a model's architecture are responsible for generating factual errors.
- Trait-Aware Policy Optimization for Autoregressive Multi-Trait Essay Scoring This training method helps models score essays more accurately across multiple different writing qualities simultaneously. [paper] [episode]
- Boosting Data Augmentation with Stochastic Weight Averaging This study shows that averaging model weights can provide the benefits of data augmentation without the cost of extra training. [paper] [episode]
- Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings A new watermarking technique uses both word and context information to make it harder for people to remove digital signatures from text. [paper] [episode]
- Partial Inverse Design of High-Performance Concrete Using Cooperative Neural Networks for Constraint-Aware Mix Generation This AI framework helps engineers design specific concrete mixtures that meet strict performance requirements in a single step. [paper] [episode]
- ClosureBench: A Constructive Benchmark for Compositional Graph Reasoning This benchmark generates new, unique math problems on the fly to prevent models from simply memorizing old test questions. [paper] [episode]
- The Sample Complexity of Learning Lipschitz Operators with respect to Gaussian Measures This mathematical study establishes the fundamental limits on how many data samples are needed to learn certain types of complex functions. [paper] [episode]
- Cross-Preference Learning for Sentence-Level and Contextual Machine Translation This training framework helps translation models better decide when to use surrounding context and when to focus on individual sentences.
- Towards Efficient Parametric State Estimation in Circulating Fuel Reactors with Shallow Recurrent Decoder Networks This method uses neural networks to monitor the real-time state of nuclear reactors using only limited sensor data. [paper] [episode]
- Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models This method improves image generation models by training them on entire groups of images at once rather than just comparing two at a time. [paper] [episode]
- Multilingual Models for Check-Worthy Social Media Posts Detection This study explores how multilingual AI models can effectively identify both harmful and factually verifiable posts across different languages. [paper] [episode]
- GLOW: Graph-Language Co-Encoding for Agentic Workflow Performance Prediction This framework predicts how well an AI workflow will perform by combining structural graph analysis with language understanding. [paper] [episode]
- NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution This system detects AI-written text and provides human-readable explanations for why it thinks the text is not human. [paper] [episode]
- Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models This survey explores how to balance high-quality reasoning with the high computational costs required to run it. [paper] [episode]
- Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning This multi-stage training method turns mid-sized open models into highly effective educational tools that can outperform much larger systems. [paper] [episode]
- Robust and Efficient Guardrails with Latent Reasoning This method makes AI safety filters much faster by performing reasoning in a hidden mathematical space rather than generating text. [paper] [episode]
- Explainable Clustering of Mixture Models This research provides new mathematical bounds to help explain how well simple decision trees can represent complex data clusters. [paper] [episode]
- ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control This framework uses symbolic logic and memory to help AI write long stories that remain consistent from start to finish. [paper] [episode]
- Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases This study suggests that repeating a small dataset can actually be more efficient than training on a larger one due to specific mathematical biases. [paper] [episode]
- OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models This work shows that clever prompting can make open-source models perform as well as much larger, specialized medical models. [paper] [episode]
- WaveletDiff: Multilevel Wavelet Diffusion For Time Series Generation This new diffusion model uses wavelets to capture the multi-scale patterns found in time series data like finance or weather. [paper] [episode]
- KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty This benchmark uses real human exam error rates to see if AI models actually struggle with the same hard problems as people. [paper] [episode]
- Role-Aware Artificial Intelligence Across Augmentation and Automation in Human-Machine Symbiosis This study proposes a way to identify whether an AI was acting as a helper or a creator when it generates text. [paper] [episode]
- From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing This method improves how we direct tasks to different models by looking at their overall performance range rather than just one single answer. [paper] [episode]
- Compiler-Guided Adaptive Proof Search with Cross-Model Synergy on Context-Dependent Theorem Proving This framework uses software compiler errors to guide AI in finding mathematical proofs more efficiently. [paper] [episode]
- The Geometry of Polynomial Group Convolutional Neural Networks This mathematical study explores the underlying geometric structures of a specific type of neural network architecture. [paper] [episode]
- Fractal and Chaotic Activation Functions in Echo State Networks: Preprocessing Topology Governs the Echo State Property This research shows that using non-smooth, fractal-like functions can actually make certain types of neural networks more stable and faster. [paper] [episode]
- Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Tendencies in Large Language Models This study investigates whether AI models can be easily manipulated into adopting conspiracy theories or showing social biases. [paper] [episode]
- AnomalyMatch: Discovering Rare Objects of Interest with Semi-supervised and Active Learning This framework combines semi-supervised learning with human feedback to find rare objects in massive datasets like astronomical images. [paper] [episode]
- Forecast Skill Is Not Decision Skill: Evidence from Weather-Dependent Decision Tasks This study demonstrates that a weather model that is statistically accurate might still be poor at helping people make real-world decisions. [paper] [episode]
- KernelGenBench: Can LLMs and Agents Write Efficient Kernels Across Operator Sources and Hardware Platforms? This benchmark tests how well AI can write high-performance code for different computer chips and software libraries. [paper] [episode]
- Small Molecule Optimization with Large Language Models This method uses the generative power of language models to help discover new drug molecules with better properties. [paper] [episode]
- CacheWeaver: Cache-Aware Evidence Ordering for Efficient Grounded RAG Inference This technique reorders retrieved information to maximize the reuse of data in computer memory, making AI responses faster. [paper] [episode]
- Relocation of compact sets in R n by diffeomorphisms and linear separability of datasets in R n This mathematical theory shows how complex data can be transformed so that it is easier for neural networks to classify. [paper] [episode]
- GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes This benchmark tests whether AI agents can successfully communicate and cooperate in a high-pressure, real-time video game. [paper] [episode]
- PROMPT2Box: Improving LLM Weakness Discovery and Specificity Estimation by Uncovering Entailment Structure among Prompts This method uses specialized embeddings to help researchers see exactly how specific or vague an AI's prompt is. [paper] [episode]
- Deep Divide-and-Reduce in Symbolic Regression This new approach improves the ability of AI to discover mathematical formulas from data by breaking complex equations into simpler parts. [paper] [episode]
- Short paper: Models in the dark -- Rectification and erasure under GDPR in ML supply chains This paper examines the legal and technical difficulties of fulfilling privacy rights like "the right to be forgotten" within complex AI development chains. [paper] [episode]
- Procedural Content Generation via Generative Artificial Intelligence This survey explores how generative AI is being used to automatically create game content like terrain, items, and stories. [paper] [episode]
The papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving — The paper investigates a novel approach to enhancing Large Language Model (LLM) problem-solving capabilities by focusing on solution divergence—the presence of multiple viable solutions for a single problem. [episode]
- SoK: AI-Augmented Binary Reversing — The paper presents "the first comprehensive systematization of knowledge on AI-augmented binary reversing," addressing a field that has become "increasingly fragmented" despite its critical role in software understanding, vulnerability discovery, and malware investigation. [episode]
- Efficient Fuzzy PSI under One-Sided Assumptions — Fuzzy private set intersection (PSI) allows two parties to identify approximately matching elements between their input sets, where a match occurs if their distance is at most a threshold delta. [episode]
- Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs — This study presents a large-scale evaluation of Uncertainty Estimation (UE) methods, addressing a critical gap in prior research that was "predominantly focused on English." By utilizing two human-curated Multiple-Choice Question Answer (MCQA) datasets and eliciting longform, lan [episode]
- SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision — Agent skills are crucial procedural artifacts that allow LLM agents to manage complex workflows, verify constraints, and recover from failures in dynamic environments. [episode]
- ConfRAG: Confidence-Guided Retrieval-Augmenting Generation — The paper "ConfRAG: Confidence-Guided Retrieval-Augmenting Generation" addresses critical limitations in standard Retrieval-Augmented Generation (RAG) systems, specifically concerning factual accuracy and susceptibility to hallucination. [episode]
- TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation — The paper introduces TeleTables, a novel benchmark designed to rigorously evaluate Large Language Models (LLMs) in their ability to interpret and reason over complex tabular data found within telecommunication standards. [episode]
- Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes — Speculative Decoding (SD) accelerates large language model inference by using a lightweight draft model to propose tokens, which are then verified by the larger target model. [episode]
- Consensus Group Relative Policy Optimization for Text Generation — Consensus decoding methods, such as Minimum Bayes Risk (MBR) decoding, are highly effective for text generation by sampling multiple candidates and selecting the one with the highest consensus score. [episode]
- Improving Weak World Models Behind Strong Agents in Atari Pong — This paper addresses the critical gap in how visual world models—which are typically evaluated only as components of model-based reinforcement learning (MBRL) systems—are assessed for their standalone reliability. [episode]
- Advancing Subseasonal Forecasting with Machine Learning — Subseasonal forecasting—weather predictions two to six weeks ahead—is crucial for agricultural planning and disaster preparedness, yet it remains a "predictability desert" due to compounding model errors and the chaotic nature of the atmosphere. [episode]
- Longitudinal Adoption and Deprecation of the Privacy Sandbox Web APIs — This paper presents a longitudinal measurement and analysis study of the Privacy Sandbox APIs, providing a comprehensive look at their adoption and deprecation over their entire lifespan in Chrome. [episode]
- Unified Deployment-Aware Evaluation of Open Reasoning Language Models — The existing landscape of large language model (LLM) evaluation often suffers from methodological inconsistencies, such as "mixed sample sizes" and "accuracy-centered summaries," making practical model selection ambiguous. [episode]
- RL-VLA cubed: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training — The paper introduces RL-VLA3, a fully asynchronous distributed reinforcement learning framework designed specifically to address the system-level challenges inherent in training Vision-Language-Action (VLA) models. [episode]
- Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation — Addressing the complex task of emotion recognition within natural dialogue requires robust modeling that captures both fine-grained acoustic details and long-range conversational context. [episode]
- Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning — Understanding how deep neural networks learn useful internal representations is a central open problem in theory, and this paper addresses that challenge by proposing a mathematically tractable surrogate for feature learning called Neural Low-Degree Filtering (Neural LoFi). [episode]
- Multi-Modal Time Series Prediction via Mixture of Modulated Experts — Multi-Modal Time Series Prediction via Mixture of Modulated Experts addresses the limitations of conventional token-level fusion in multi-modal time series prediction (MMTSP), where traditional methods struggle to exploit complementary information between textual signals and temp [episode]
- Deep Learning-Driven Peptide Classification in Biological Nanopores — This study addresses the critical need for rapid, low-cost, and accurate methods for identifying proteins and peptides in clinical settings using nanopore devices. [episode]
- A Survey on Semantic Modeling for Building Energy Management — Building Energy Management (BEM) is critical for reducing carbon emissions and energy consumption in the building sector, yet its development is hindered by "heterogeneous data models" and "semantic interoperability" issues arising from IoT technologies. [episode]
- Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism — This paper advances beyond conventional notions of model reproducibility by proposing that analyzing the underlying token probability distributions is necessary to accurately quantify Large Language Model (LLM) nondeterminism. [episode]
- EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents — EvoCUA-1.5 is a novel online reinforcement learning framework designed to address the unique challenges of training multi-turn computer-use agents, which must operate within partially observable, multimodal desktop environments. [episode]
- From Architecture to Output: Structural Origins of Hallucination in Large Language Models and the Amplifying Role of Data — The persistence of hallucination—the production of "fluent, confident, factually wrong outputs"—remains a critical challenge in large language models (LLMs). [episode]
- Trait-Aware Policy Optimization for Autoregressive Multi-Trait Essay Scoring — Multi-trait essay scoring aims to provide "finegrained evaluation of writing quality across multiple dimensions," but traditional reinforcement learning methods struggle with this task because they rely on a "sequence-level scalar reward." This approach fails to exploit the compl [episode]
- The Sample Complexity of Learning Lipschitz Operators with respect to Gaussian Measures — Operator learning, which is the approximation of mappings between infinite-dimensional function spaces using data, has seen significant empirical success in computational science and engineering. [episode]
- Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings — The paper addresses the persistent challenges in attributing Large Language Model (LLM)-generated content, which is often rewritten or translated to meet specific demands. [episode]
- Partial Inverse Design of High-Performance Concrete Using Cooperative Neural Networks for Constraint-Aware Mix Generation — Partial inverse design addresses a significant limitation in traditional concrete mix design: the inability to flexibly handle scenarios where only a subset of variables are unknown or must be optimized while others remain fixed by practical constraints. [episode]
- ClosureBench: A Constructive Benchmark for Compositional Graph Reasoning — ClosureBench introduces a novel, constructive benchmark designed for evaluating compositional graph reasoning capabilities in large language models. [episode]
- Boosting Data Augmentation with Stochastic Weight Averaging — The provided material details advanced theoretical results concerning how group actions, denoted by S g, affect optimization processes, specifically focusing on establishing conditions for invariance and equivariance within loss functions and associated differential operators. [episode]
- Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning — This paper introduces a comprehensive methodology for enhancing large language models (LLMs) specifically for pedagogical tutoring by leveraging advanced optimization techniques on open-source architectures. [episode]
- Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models — * Preference optimization has emerged as an efficient alternative to online reinforcement learning from human feedback (RLHF) for aligning text-to-image diffusion models. [episode]
- Multilingual Models for Check-Worthy Social Media Posts Detection — The research analyzes the performance of multi-label XLMRoBERTa-base models for detecting verifiable factual claims and harmful claims within social media posts. [episode]
- Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models — The paper provides a comprehensive analysis of "reasoning economy," which is the critical balance between performance (benefits) and computational costs (budgets) in Large Language Models (LLMs). [episode]
- Towards Efficient Parametric State Estimation in Circulating Fuel Reactors with Shallow Recurrent Decoder Networks — The development of accurate, real-time state estimation for complex engineering systems is crucial for creating reliable digital twins of operating equipment. [episode]
- GLOW: Graph-Language Co-Encoding for Agentic Workflow Performance Prediction — Agentic Workflows (AWs) represent a promising paradigm for solving complex tasks by coordinating multiple specialized agents through structured collaboration topologies. [episode]
- NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution — The NOTAI.AI framework addresses a critical limitation in machine-generated text detection: while current models achieve high classification accuracy, they often operate as "black boxes," failing to provide transparent evidence for their claims. [episode]
- Cross-Preference Learning for Sentence-Level and Context-Aware Machine Translation — Cross-Preference Learning (CPL) introduces a novel preference-based training framework designed to address a critical limitation in current machine translation (MT) systems: the assumption that contextual information is always beneficial. [episode]
- ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control — " * Abstract and Problem Formulation Long stories inherently suffer from accumulated inconsistency, where existing prompting-based methods often fail to maintain coherence. [episode]
- OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models — OpenMedLM presents a novel prompting platform designed to achieve state-of-the-art (SOTA) performance for open-source (OS) large language models in medical question answering, demonstrating that robust prompt engineering can outperform computationally intensive fine-tuning. [episode]
- Role-Aware Artificial Intelligence Across Augmentation and Automation in Human-Machine Symbiosis — The evolution of artificial intelligence has created an ambiguous boundary between human input and computational output, leading to a complex state of "human-machine symbiosis." When content is generated through this collaboration, the functional role AI plays—whether it acts a [episode]
- WaveletDiff: Multilevel Wavelet Diffusion For Time Series Generation — The paper introduces WaveletDiff, a novel diffusion model designed for time series generation, and critically examines the phenomenon of reproducibility within this domain. [episode]
- Explainable Clustering of Mixture Models — Explainable machine learning aims to provide transparency into complex black-box algorithms, yet much of existing research focused on worst-case guarantees for explainable clustering is "distribution-agnostic." This paper addresses this limitation by introducing a statistical set [episode]
- KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty — Mathematical reasoning is a central axis for evaluating language and vision-language models, yet most existing benchmarks lack a per-item difficulty signal grounded in actual human performance. [episode]
- Robust and Efficient Guardrails with Latent Reasoning — Existing safety guardrails are crucial for deploying Large Language Models (LLMs) in real-world applications, yet current reasoning-based approaches suffer from "steep computational cost" and high latency due to their reliance on generating explicit chain-of-thought (CoT) rationa [episode]
- From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing — From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing The paper identifies a fundamental limitation in existing Large Language Model (LLM) routing methods: the reliance on single-shot supervision. [episode]
- Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases — The paper investigates the "small-vs-large gap," a counterintuitive phenomenon where training on fewer, repeated samples can lead to reduced training compute for a given model compared to using a larger dataset. [episode]
- Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Tendencies in Large Language Models — The paper "Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Mindset in Large Language Models" investigates whether Large Language Models (LLMs) can reproduce complex human psychological constructs, specifically a generalized tendency to endorse conspiracy theories [episode]
- AnomalyMatch: Discovering Rare Objects of Interest with Semi-supervised and Active Learning — AnomalyMatch addresses the critical challenge of identifying rare objects—or outliers—within massive datasets where labeled examples are scarce, a common occurrence in fields like astronomy and computer vision. [episode]
- Fractal and Chaotic Activation Functions in Echo State Networks: Preprocessing Topology Governs the Echo State Property — This paper systematically investigates non-smooth activation functions—including chaotic, stochastic, and fractal variants—within Echo State Networks (ESNs) to challenge the prevailing assumption that smooth, globally Lipschitz continuous functions are required for stable ope [episode]
- Forecast Skill Is Not Decision Skill: Evidence from Weather-Dependent Decision Tasks — Weather forecasts are increasingly critical for guiding decisions in industries ranging from agriculture to renewable energy, yet traditional evaluation metrics—such as CRPS or PIT histograms—often fail to capture how these forecasts truly impact end-users. [episode]
- The Geometry of Polynomial Group Convolutional Neural Networks — The paper rigorously analyzes the Jacobian structure of polynomial group convolutional neural networks, establishing key mathematical identities and proving a recursive relationship for their derivatives. [episode]
- CacheWeaver: Cache-Aware Evidence Ordering for Efficient Grounded RAG Inference — In Retrieval-Augmented Generation (RAG), a primary bottleneck for interactive use is serving latency, driven by the large input length of retrieved context. [episode]
- KernelGenBench: Can LLMs and Agents Write Efficient Kernels Across Operator Sources and Hardware Platforms? — KernelGenBench is a comprehensive, unified benchmark designed to rigorously evaluate LLM- and agent-generated Triton kernels, addressing the critical gap in current research where kernel development remains "a highly specialized and labor-intensive task." Because existing benchma [episode]
- Compiler-Guided Adaptive Proof Search with Cross-Model Synergy on Context-Dependent Theorem Proving — The paper introduces a novel framework for context-dependent theorem proving in real-world Lean 4 projects, addressing the limitations of traditional independent sampling methods. [episode]
- Small Molecule Optimization with Large Language Models — Molecular optimization is a cornerstone of drug discovery, yet traditional methods are time-consuming and costly due to the vast and discrete nature of chemical space. [episode]
- Relocation of compact sets in R n by diffeomorphisms and linear separability of datasets in R n — The paper investigates advanced techniques for manipulating and separating complex topological structures embedded in Euclidean space (R n). [episode]
- Diagnosing and Mitigating Semantic Inconsistencies in Wikidata's Classification Hierarchy — Wikidata, despite being "the largest open knowledge graph on the web," suffers from a degree of taxonomic inconsistency due to its "relatively loose editorial policy." This paper addresses the fundamental problem of misuse and conflation of P31 (instance of) and P279 (subclass of [episode]
- PROMPT2BOX:Improving LLM Weakness Discovery and Specificity Estimation by Uncovering Entailment Structure among Prompts — The paper "PROMPT2BOX: Improving LLM Weakness Discovery and Specificity Estimation by Uncovering Entailment Structure among Prompts" introduces a novel framework designed to enhance the rigorous evaluation of Large Language Models (LLMs). [episode]
- Procedural Content Generation via Generative Artificial Intelligence — Generative Artificial Intelligence (AI) has become a transformative force in procedural content generation (PCG), moving beyond simple rule sets to create complex, dynamic assets. [episode]
- Deep Divide-and-Reduce in Symbolic Regression — Symbolic Regression (SR) aims to discover the underlying mathematical relationship or equation that best explains a set of input-output data, moving beyond mere prediction to provide interpretable scientific models. [episode]
- Short paper: Models in the dark -- Rectification and erasure under GDPR in ML supply chains — The paper presents a holistic survey of challenges in implementing the rights to rectification and erasure under the General Data Protection Regulation (GDPR) within machine learning systems, specifically addressing issues arising from complex ML supply chains. [episode]
- Inducing Permutation Invariant Priors in Bayesian Optimization for Carbon Capture and Storage Applications — The following is a detailed, comprehensive summary of the scientific paper, "Inducing Permutation Invariant Priors in Bayesian Optimization for Carbon Capture and Storage Applications," utilizing only content extracted from the text. [episode]
- GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes — GPTNT is a novel benchmark designed to rigorously evaluate real-time, multimodal collaboration between artificial agents, addressing a significant gap in existing benchmarks that traditionally studied time pressure, information asymmetry, and imperfect communication in isolation. [episode]
- Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills — LLM-powered coding agents are increasingly integrated into software development workflows, making them a critical component of the modern software supply chain. [episode]
- QoNext: Towards Next-generation QoE for Foundation Models — Existing evaluations of foundation models fail to capture "user’s experience during interaction," often treating evaluation as a matter of output correctness alone. [episode]
- Hyperedge Anomaly Detection with Hypergraph Neural Network — Hypergraphs provide a powerful data structure for modeling higher-order associations, allowing researchers to capture complex relationships that conventional graph structures fail to represent. [episode]
- Reservoir-Based Graph Convolutional Networks — The evaluation of advanced graph neural networks necessitates rigorous testing across diverse and complex graph structures, ranging from molecular chemistry to functional brain connectivity. [episode]
- HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning — HLS-Seek is a novel framework designed to address a critical gap in existing Large Language Model (LLM) approaches to High-Level Synthesis (HLS). [episode]
- Predicting California Bearing Ratio with Ensemble and Neural Network Models: A Case Study from Turkiye — The California Bearing Ratio (CBR) serves as a critical geotechnical indicator for assessing the load-bearing capacity of subgrade soils, particularly in transportation infrastructure and foundation design. [episode]
- "**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems — The paper investigates prompt injection (PI) attacks against LLM-based automatic grading (AG) systems, demonstrating that these emerging AI-powered assessment tools are highly vulnerable. [episode]
- The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs — The paper provides a rigorous mechanistic analysis of a specific vulnerability in large language models (LLMs) known as the continuation-triggered jailbreak. [episode]
- Detect, Remask, Repair: Diffusion Editing for Faithful Summarization of Evolving Contexts — Summaries of real-world events are often generated from partial or early contexts; however, as information evolves and new evidence arrives, these initial drafts can become "stale" or contain claims that are no longer supported by the full context. [episode]
- MemCoRe: Recovering Evidence from Progressively Compressed Factual Knowledge for Agent Memory — The paper introduces MemFly, a novel framework designed to address the fundamental dilemma in large language model (LLM) agents where efficient compression of long-term memory conflicts with maintaining the precise fidelity required for complex reasoning. [episode]
- Towards AI epidemiology: a measurement standardisation framework for prospective risk detection — This paper introduces a measurement standardisation framework designed to address the critical governance challenge posed by opaque, deployed AI systems. [episode]
- An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders — " The study addresses the question, "Can pretrained models generalize to new datasets without any retraining?" It investigates whether embeddings from self-supervised learning (SSL) encoders can form meaningful clusters when applied to image datasets that were "not seen during tr [episode]
- Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation — Block attention offers a promising pathway for improving Key-Value (KV) cache reuse in long-context scenarios, such as Retrieval-Augmented Generation (RAG), by processing input text in independent blocks that do not attend to one another. [episode]
- Improving Energy Efficiency of Oil Platforms Through Optimal Loading of Diesel Generators Using Machine Learning and Search Algorithms — The escalating global demand for energy and the environmental impact of fossil fuel use necessitate efficient energy production, a challenge that has often overlooked the internal efficiency of industrial systems like offshore oil platforms. [episode]
- TSMini: A Simple Yet Highly Effective Trajectory Similarity Learning Model — Trajectory similarity is a foundational concept in spatio-temporal data mining tasks such as trajectory clustering and k-nearest neighbor (kNN) queries. [episode]
- Don' t Box Me In: Dynamic Cultural Adaptation and Cognitive Tracking for Social Understanding — The paper "Don’t Box Me In: Dynamic Cultural Adaptation and Cognitive Tracking for Social Understanding" introduces DyCAC, a novel training-free framework designed to overcome the limitations of existing Large Language Models (LLMs) that treat culture as a static demographic at [episode]
- Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective — Data acquisition efficiency is a central challenge in deploying reinforcement learning in business and healthcare operations, where interactions are costly, slow, and often involve humans in the loop. [episode]
- Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents — Modern LLM agents require more than just long context; they need "decision-relevant evidence at the moment of action." This paper addresses the limitations of standard retrieval methods, which are based on semantic similarity rather than impact on decision-making. [episode]
- MemMA: Coordinating the Memory Cycle through Multi-Agent Reasoning and In-Situ Self-Evolution — Memory-augmented LLM agents are crucial for maintaining coherence over long-horizon interactions in conversational settings; however, existing systems often treat the memory cycle—construction, retrieval, and utilization—as isolated subroutines. [episode]
- OR-Agent: Bridging Evolutionary Search and Structured Research for Automated Heuristic Design — Automated heuristic design in complex, experiment-driven domains requires a systematic approach that moves beyond simple iterative mutation loops and stochastic search strategies. [episode]
- Not All LLM Reasoning is Visible in the Chain-of-Thought — The paper details advanced methodologies for improving Large Language Model (LLM) reasoning capabilities, specifically through Reinforcement Learning (RL) fine-tuning on complex arithmetic tasks. [episode]
- Active Inference for an Intelligent Agent in Autonomous Reconnaissance Missions — This paper introduces a novel framework for autonomous reconnaissance missions, applying Active Inference to solve the persistent surveillance problem in a dynamic 2D environment. [episode]
- From Plausible to Actionable: A Position on LLM Self-Explanations — The emergence of Large Language Models (LLMs) has created new opportunities in explainable artificial intelligence (XAI), particularly through self-explanations—natural language explanations generated by the model for its own decisions. [episode]
- MetaCaster: Meta-Harness-Optimized Agent for End-to-End Few-Shot Learning of Lightweight Time Series Forecasters — Lightweight time series forecasting (TSF) is critical for resource-constrained environments, yet these specialized models typically require "substantial training data," which severely limits their applicability in domains characterized by "scarce, slowly accumulated, or privacy-s [episode]
- Collaboration between parallel connected neural networks -- A possible criterion for distinguishing artificial neural networks from natural organs — The paper investigates Parallel Connected Neural Networks (PNNs) using the MNIST dataset to establish a quantitative criterion for measuring the "bionic level" of artificial neural networks. [episode]
- Gradient-based Model Shortcut Detection for Time Series Classification — Deep learning models have become state-of-the-art in Time Series Classification (TSC), but they are prone to relying on "spurious correlations" or shortcut learning—a phenomenon where phenomena not causally related to the true task mislead model performance. [episode]
- Brain4FMs: A Benchmark of Foundation Models for Electrical Brain Signal — The paper "Brain4FMs: A Benchmark of Foundation Models for Electrical Brain Signal" establishes a comprehensive benchmark designed to evaluate the performance and predictive capabilities of various Foundation Models (FMs) when applied to complex electrical brain signal analysis. [episode]
- Quality-diversity in dissimilarity spaces — This paper presents a generalized Quality-diversity (QD) framework known as Go-Explore, designed specifically for optimizing objectives that are "hard to optimize" or "computationally expensive." It addresses the challenge of finding a diverse set of inputs—rather than just loc [episode]
- Inductive Venn-Abers and related regressors — This paper introduces a novel framework for regression estimation, generalizing concepts from Venn–Abers predictors—traditionally used in binary classification—to the challenging domain of unbounded regression. [episode]
- Tracing Audio Grounding and Answer Selection in Audio LLMs —
- SMILE: Bridging Continuous Optimization and Discrete Symbolic Recovery —
- A Cost-Aware Agentic Architecture for NL-to-SQL over Nested Enterprise Schemas, with a New Benchmark —
- CAGE: Coherence-Aware Graph Encoding for Retrieval-Augmented Generation —
- ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying —
- Continual Graph Memory for Adaptive Recommendation under Intent Drift —
- Choosing the Right Language Mode at Inference Time for Multilingual Reliability —
- Interpretability for Turing Machines —
- Harness-agnostic detection and immunization of reward hacking in self-evolving language models —
- ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies —
- WEECFP-SuRGE: A Position-Aware Substructure Encoding Method for Molecular Property Prediction —
- Controlling and Assessing Appropriate Persona Use in LLM-based Dialogue Generation —
- Train What You Deploy:Token-Faithful Post-Training of a Production Coding —
- Predicting Spatiotemporal Mobile Sensing-Based PM2.5 Concentrations Using Low-Rank Adapted Spatially Attentive Graph Neural Network —
- SQL-Zero: Self-Evolving Text-to-SQL —
- Model Retirement Creates Reproducibility Risk in Biomedical AI Publications —
- FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality —
- How Do Language Models Represent and Use Phonological Information for Allomorph Selection? —
- Simulation-free Unbalanced Dynamic Optimal Transport with General Growth Penalty —
- Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models —
- PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces —
- Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models —
- Locating and Steering Refusal Beyond Attention —
- Training Large Language Models for Small-Molecule Design with Synthetic Task Scaling —
- Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM —
- DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems —
- Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs —
- A Fairness Audit of the Duckworth-Lewis-Stern Method: Format-Specific and Gender-Differential Bias, with an Interpretable Calibration Layer for Cricket Target Revision —
- Vectorizing Classical Tamil: Representation Learning for Verse-Commentary Pairs —
- Resilience Beyond Stationary Client Unavailability: Unlocking Efficient and Unbiased Federated Learning —
- Shadow Queries for Private Retrieval in Vector Databases —
- Memory-Efficient Designs for Word-Wise Universal Fully Homomorphic Encryption —
- A Robust Watermark-based Fingerprint Framework for GNNs Ownership Verification —
- Persistent Teacher Anchoring for Tool-Using Agents —
- Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges —
- Dynamic Heterogeneous Graph Representation Learning: A Survey —
- DODR: Deterministic Operator-Driven Reasoning in Latent Space —
- Injected and Leaked: Actively Inducing Side-Channel Leakage Using Electromagnetic Injection and Hardware Nonlinearity —
- Learning-Augmented Algorithms: Guarantees, Construction Mechanisms, and System-Level Implications —
- Can Activation Steering Capture Multidimensional Authorship Style? —
- ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing —
- How Faithful Is Attribution for Sales Forecasting? A Counterfactual Study —
- Whose record is this? Diagnosing and authorizing record use in personalized multimodal models —
- Hierarchical Possession-Aware Graph Pointer Network for Pass Receiver Selection —
- MedFlow: Class-Aware Multi-Scale Generation for Medical Time-Series Synthesis —
- When Financial Fine-tuning Fails: A Three-Level Detectability Analysis of Numerical Hallucination in Domain-Adapted Language Models —
- Recurrence Is Not Enough: Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3 —
- CPR-IE:A Compression-Prediction-Resource Intelligence Efficiency Metric —
- Federated Attack Campaign Detection via Contrastive Encoding of Threat Indicators in Gradient Updates —
- A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures —
- Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution under Analysis Budgets —
- Minimax Lower Bound for Estimating Diffusion-based Local Intrinsic Dimension —
- Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities —
- Generating Constructive Feedback on Stories via Reinforcement Learning —
- Communication-Efficient Personalized Federated Learning via Layer-Wise Multi-Threshold Random Sketching —
- PACE: Propagation-Aware Collaborative Correction for One-Shot Personalized Federated Graph Learning —
- On Epistemic Diversity in Large Language Models —
- Long Horizon Transformer Quantile Fault Prediction for Multi Site Industrial Predictive Maintenance —
- MABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate —
- MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain —
- ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults —
- KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU —
- CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation —
- MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models —
- When Genomic Masking Priors Fail to Transfer: Strong Variant Prediction, Weak Functional Generation —
- MZ-Rain: Moisture-Budget-Guided Zero-Inflated Model for Station-Level Precipitation Nowcasting —
- CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution —
- LLM-Assisted Behavioural and Scenario Augmentation for Agent-Based Energy Adoption Models —
- From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents —
- CHAMP: Cross-domain Hybrid Architecture for Matchmaking and Prediction in Online Multi-Player Games —
- AutoLR: Automating the Path from Research to Launch Review in Industrial Recommender Systems —
- Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents —
- MARLA: A Conceptual Scaffold for Regulatory Learning under the EU AI Act —
- ReCAST: Restoration-aware Cascaded Stage-wise Training for Obfuscated SMS Risk Classification —
- Reinforcement Learning for Sequential Solar PV Policy Design under Uncertainty: An Agent-Based Approach —
- From Deep to Shallow: Unconstrained and Efficient Layer Merging Strategy —
- From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments —
- Cache-Aware Joint Router Adaptation for Memory-Efficient MoE Inference —
- RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents —
- The Security Feature Location Problem —
- Adaptation Interfaces for In-Context Tabular Foundation Models in Time-to-Event Prediction —
- Fast Gauss Sums via Flash Attention —
- Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing —
- Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing —
- Solving Hard XAI Queries Based on a Compiled Dual-Rail Encoding —
- Physics-Aware Random Walk Fingerprints for Scalable Power Grid Graph Classification —
- Discourse Dependency: A Continuous Criterion for Translation Difficulty —
- Why We Care About Understanding: Competence through Predictive Compression —
- Fractal basins trap latent reasoning —
- Robust Coverless Linguistic Steganography via Sentence Embedding Space with Global Resynchronization —
- BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference —
- Global to Local: Topology-Preserving Adaptive Graph Pooling via Granular-Ball —
- A Tree-based RAG Framework for Evidence-Intensive QA via Adaptive Planning and Topology-Aware Evidence Gathering —
- Beyond Homoscedasticity: Decoupled Uncertainty Optimization for Deep Imbalanced Regression —
- BIT.UA at BioASQ 14B: Modular Retrieval with pg textsearch and Qdrant, and Agent-Based Answer Generation —
- Language models judge war differently when tested for alignment —
- TPMSpy: Validation of Measured Boot Systems by Low-Level Tracing of TPM Usage —
- Solution-space heterogeneity shapes federated learning dynamics across partial differential equations —
- Has MIMO decoding been proved hard from lattice problems? —
- Amortizing Scaling Law Construction Costs —
- TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing —
- MoirfEolas and Cr'iochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology —
- Leveraging Low-Level Symbolic Competences for Unsupervised Grounding in Hallucination Detection —
- An Analysis of Self-supervised Pre-training with Dependent Samples —
- Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment —
- How do LLMs Evaluate Perceived Moral Agency? Investigating Moral Decision-Making in Human-Artificial Agents Interactions —
- Towards Efficient Evaluation of Evolutionary Transfer Optimization: Case Studies on Task-Parameterized Applications —
- EuroAlpaca: Task-Preserving Localisation of Instruction Data for European Languages —
- A Structured Debate-Mixture-of-Agents Framework for Complex Clinical Diagnostic Decision Support —
- Confounding-Valid Conformal Inference for Counterfactual KPIs in Wireless Networks —
- Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time —
- MePo++: Unifying Representation Refinement and Reconciliation for General Continual Learning —
- TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents —
- Deep Microcompression: Structured Pruning and Bit-packed Quantization for Microcontrollers —
- Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny? —
- Constructing and Evaluating Clinical Reasoning Trajectories for Medical Agent —
- Operational Roles of QRNG-Derived Quantum Entropy in Bitcoin Proof-of-Work Architectures —
- LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28 —
- ProCA: Progressive Contrastive Alignment for Robust EEG Visual Decoding —
- NEAT-POCKET: Pocket-Conditioned Autoregressive 3D Molecular Generation with a Neighborhood-Guided Set Transformer —
- Improving Language Identification for Code-Switched Utterances with Integer Linear Programming —
- Compact Bellman-Grounded Cognitive Maps for Cost-Aware Navigation —
- Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens —
- A Comparative Study of Counterfactual Explainers for Graph Neural Networks Enabling Multiple Types of Graph Edit —
- TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors —
- Understanding the Privacy-Preserving Potential of HTTP/2 Against Webpage Fingerprinting —
- Single-Query Black-Box Calibration Auditing via Logit Bias —
- Coarse-Graining Hidden Representations: Unsupervised Neuron Selection via Mapping Entropy —
- MomentQuant: an even more minimalist interval method with linear time complexity for time series classification —
- From 80x to 385x: A Best-Matching-Unit Search at the L2 Roof, Measured Against a Symmetrically Tuned Baseline —
- NS-ST-GraphRAG: Neuro-Symbolic Spatio-Temporal GraphRAG for Literary Knowledge Processing —
- SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding —
- A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment —
- A Hybrid Predictive Ensemble of Machine Learning and Deep Neural Networks for Early Cardiovascular Disease Risk Assessment —
- From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making —
- Beyond Stationarity in Time Series: Discovering Causal Structures and Latent Regimes via Markov Blankets —
- Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG —
- Conformal Prediction for Offensive Security —
- Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers —
- The Mirror Agent Model: a Bayesian Architecture for Interpretable Agent Behavior —
- Phase Transition Frequency as a Training Time Predictor of Test Accuracy in ResNets —
- What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection —
- FluxDisco: Symbolic Regression for Stoichiometric Dynamical Systems via Monte Carlo Graph Search —
- PAC-Bayesian Reconstruction Guarantees for Time Series Variational Autoencoders —
- Dimension-Adaptive Batched Lipschitz Narrowing Without Knowing the Zooming Dimension —
- A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR —
- FedDRAW: Federated Dual Reputation Annealing Weighting for Heterogeneous Multi-Institutional Chest Radiograph Classification —
- CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review —
- ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs —
- Substrate-Aware AI Agents: Execution Context as a First-Class Input —
- Hessian-based molecular conformation augmentation for a scalable and efficient strategy of machine learning interatomic potentials —
- PRICE: A Systematic Study of LLM Adaptation Choices for Bitcoin Price Forecasting —
- Governing Bring Your Own AI: A Parameterized Maturity Model —
- Uncensored Open-weight Models: Redistribution as the Persistence Layer —
- Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory —
- A Unified Physics-Aware Quantum Machine Learning Framework across Power GaN HEMTs and Logic Nanowire FETs: Predicting Unseen Process Splits and Held-Out Geometry Combinations with Lower Error and Tighter Split-to-Split Variability —
- GLASS: Graph-Language Alignment with Spherical Scoring for Transferable Graph-Level Anomaly Detection —
- Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions —
- Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents —
- Self-Supervised Lexical Representation Learning for Fast, Large-Scale Phylogenetic Inference —
- CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls —
- AI for Computational Design Science: A Responsible Human-AI Framework and Case Study on Short-Form Video Safety Surveillance —
- How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method —
- Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference —
- Testing Interchangeability in LLM Agent Teams —
- GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity —
- Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods —
- Learning from VAE Errors to support ECG-based Differential Diagnosis of Myocardial Scar —
- RISE: Recursive Improvement via Self-Extrapolating Policy Distillation —
- LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics —
- How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing —
- Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness —
- Optimal Rates for Agentic Networked Information Aggregation —
- LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision Diagrams —
- Embedded Graph Flows for Categorical Graph Generation —
- Machine Unlearning as Private Retroactive Algorithms —
- Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models —
- The History Is the Detector: Executing CVE Patch History, End-to-End —
- Variational Continuation for Double Pendulum Periodic Orbits —
- Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability —
- Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education —
- Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation —
- When LLM Decompilers Recompile More and Preserve Less —
- CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents —
- Propagation Model for SSC attacks: Why SBOM (tools) don't tell the whole truth —
- Molecular D'ej`a Vu: Digit-Level Retrieval of Published Values in Frontier Language Models —
- Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence —
- Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe —
- A Deep Generative Model for Synthesizing Labeled Wireless Signals —
- RegionFed: Federated Learning for Personalized Query Understanding in Heterogeneous Retail Environments —
- WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data —
- How Much Does Corpus Choice Change Dependency-Distance Estimates? —
- EXAONE Finance 1.0: An Attention-free Time Series Foundation Model for Financial Time Series —
- Spectral-Target Physical Latent Structuring for JEPA-Style World Models —
- ProToMEx: Rapid, Interpretable Explanations via Structured Representations —
- A Data Fusion Framework for Grounding Aerospace Surrogate Model via Experimental Wind-Tunnel Observations —
- Quantum-Assisted Memory-Efficient Training for Parameter-Intensive Wi-Fi-Based Human Activity Recognition —
- Evaluating Large Language Models for Forced Outage Risk Prediction: Benefits and Comparison to Machine Learning —
- From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance —
- Memory as transformation: LETHE, a self-referential gan-inspired architecture —
- Evidence Integration in Large Language Models —
- BER-PEF: Unified Human Mobility Predictability Evaluation via Bayes Error Rate Estimation —
- Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation —
- Data-Optimized Contingency Screening: A Machine Learning Approach to Power System Security —
- Iris: Climbing to the Search Frontier —
- Data-Driven Learning of Unknown Nonlinear Differential Equations Using Functional Analysis —
- MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering —
- Modular Deep Recurrent Neural Network: Application to Quadrotors —
- A Removal Based Approach to Improve LLM Faithfulness at Test-Time —
- SharedSAE: One Feature Dictionary Across Language Models —
- Adapting from Downturns: Prediction of Long-Term Conversational-Skill Development in Mental-Health Crisis Counselors —
- A Quantum Variational Approach to Prototypical Recurrent Unit —
- Blockchain-Enabled Secure Logging for Fiscal Electronic Mechanisms: Evaluation of the Greek eSEND and myDATA Tax Systems —
- VERGE: Verification-Enhanced Refinement for Grounded Extraction of Early-Onset Colorectal Cancer Symptoms in Clinical Notes —
- Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets —
- Corporate Language Model (CLM): Transforming Tacit and Fragmented Enterprise Knowledge into a Sovereign, Auditable, and Executable Corporate Intelligence Layer —
- On the Abundance of Critical Points of the t-SNE Energy —
- Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys —
- You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments —
- Candidate Comparability Before Promotion: Conditional Validation in Adaptive Network Intrusion Detection —
- Evaluation of Phonetic Encoding Algorithms on Transcription Datasets —
- The Anatomy of an ASR Hallucination —
- Disentangling Attention in Deep Operator Learning: A Controlled Study of Data-Driven and Physics-Informed Architectures —
- A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models —
- Engineered Persuasion: Evaluating Personalized Pretexts in LLM-Generated Spear Phishing —
- REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation —
- Beyond a Universal Forecasting Selector: Demand-Conditioned Model Selection across Demand Patterns and Horizons —
- A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models —
- What Attention Recalls and Recurrence Controls in Hybrid Language Models —
- GRACE: Graph-Grounded Reflective Agent Copilot Engine for Expert-in-the-Loop Knowledge Expansion —
- HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals —
- Conformity Breaks Conformal Prediction —
- TRILOGUE: A Trilingual Spoken Dialogue Fact-Checking Benchmark with Evidence and Paired Audio —
- When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models —
- On-board ML for Trace Gas detection in Imaging Spectroscopy data —
- Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning —
- Nested Inductive Bias Framework for SPD Manifold Learning —
- Nebulon Enterprise Simulated Threats for Phishing Research (NEST-Phish): A Synthetic Enterprise Phishing Email Dataset for Behavioral and Machine-Learning Research —
- Client-Side Probing of Deleted Ridge Statistics in Federated Unlearning —
- PerfReasoning: How Well Do LLMs Reason on Hardware Performance? —
- Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal —
- Patterns of Priming in Production: Lexical, Semantic and Structural Alignment in Language Model Generation —
- Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning —
- Towards Understanding Pause Token Fine-Tuning Dynamics: A Mode Retention Perspective —
- When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference —
- ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended Reality —
- Hakken: Predicting future discoveries to fill the gaps in today's knowledge —
- Rethinking Indirect Prompt Injection as a Test-Time Search Problem —
- BioSync: Transformer-Based Cross-Modal Fusion for a Multimodal Physiological Digital Biomarker —
- LentEx: Generalizable Latent Entity Extraction via Synthetic Data and Instruction-Tuned LLMs —
- What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents —
- Hoss: Fast Oblivious Semantic Search with Heterogeneous GPU-CPU-TEE Architecture —
- MaxKernel: Agentic Kernel Generation for TPUs —
- Scale-QLoRA: Code-Invariant Adapter Merging for Native 4-bit Microscaling LLMs —
- Towards a universal language of concepts: A survey —
- An Energy-Based Conservative-Dissipative Latent Neural Evolution Operator for Magnetization Dynamics —
- Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One —
- Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection —
- A Calibrated Reflection Approach for Enhancing Confidence Estimation in LLMs —
- Mitra-v2 Technical Report —
- Data-Driven Discovery of Composition-Dependent Constitutive Models for Hyperelasticity and Viscoelasticity of Digital Materials —
- From Answers to Interpretations: Rethinking Ambiguity-Induced Aleatoric Uncertainty Estimation in LLMs —
- Fast Surrogate Modeling of Excitable and Oscillatory FitzHugh-Nagumo Dynamics with Parametric Neural Operators —
- Rhythms of Work: Multi-Scale Interpretation of Human Behavioral Traces for Workplace Agents —
- IPGeoAI: Transformer-Based Geolocation with LLM Semantic Fusion —
- Reducing Hallucinated Transcripts in Whisper via Hallucination Space Projection —
- La Agente 'Optima: Towards Agentic Self-Driving Laboratories —
- Extremely Sparse Supervision Incentivizes Reasoning Ability —
- Optimizing Credential Blast Radius Through Trust Boundaries and Delegation Under Post-Quantum Authentication Costs —
- Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models —
- Optimizer Memory Schedules for Outscaling the Overtraining Axis —
- Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines —
- When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models —
- Representation Redundancy and Structural Complexity in Finite-Field Inversion —
- GNN-Guided Graph Coarsening and Adaptive QUBO Penalties for the Capacitated Vehicle Routing Problem with Time Windows on a Quantum Annealer —
- PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning —
- tau tau-Bench: An Environment for End-To-End, Realistic Agent Construction —
- Why Is SHAP Not a Reliable Standalone Explanation Framework for Malware Detection? —
- Leveraging Imperfect Restoration for Data Availability Attack —
- SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents —
- Too Rare to Learn: Prescribed Cyclone Tracks Degrade a Bay of Bengal Ocean Emulator —
Important terms
- Solution Divergence
- A metric used to measure the variety of different ways a model can solve a single problem. Researchers found that higher diversity in solutions is actually a strong predictor of how intelligent an AI model is.
- SkillRevise
- A framework that helps AI agents fix their own mistakes. Instead of following one-shot instructions, agents look at evidence from their previous attempts and use stored principles to iteratively repair their procedural errors.
- Neural Low-Degree Filtering
- A theoretical framework explaining how deep learning works. It suggests that training is a process where each layer of a model selects specific directions in the data that have the highest correlation with the target labels.
- GLOW
- A framework designed to make agentic workflows more efficient. It uses graph neural networks and large language models to predict how a workflow will perform before it is even run, drastically reducing optimization time.
- COLAGUARD
- A safety method that moves AI reasoning into a continuous latent space. This allows the model to check if its responses are safe much faster and using fewer tokens than methods that require explicit text explanations.