AI papers — 2026-09-16
Most AI agents start every new session with a blank slate, which means they forget the specific data schemas or tool settings that made them useful in the first place. A new architecture called shared selective persistent memory solves this by stripping away messy reasoning traces and only keeping essential workspaces, such as task specs and output constraints.
It turns out that what you keep is far more important than how much you keep. A system using this selective memory completed 12 out of 12 trials in testing, while a system with no memory failed every single one, and even doubling the context did not help the results.
This focus on precision over sheer volume carries over into how we train models to write code. Researchers have found that instead of using a complex judge model—which can lead a model to cheat by optimizing for the judge rather than writing good code—it is better to use a simple, deterministic strict-launch filter.
By checking if code runs without errors in a headless engine, researchers took a 14B model and jumped its clean-launch rate on held-out tasks from 8.8% up to 42.2%. The verifier is effectively the curriculum; if your gate is too lenient, you lose those gains.
Beyond just writing code, there is a push toward making models more human in social interactions and specialized reasoning. A new framework called MASCOT helps multi-agent systems avoid persona collapse by using bi-level optimization to keep individual identities distinct and productive.
Similarly, when it comes to education, we are seeing how these models handle student mistakes. New research shows that while models can successfully simulate errors in math by following a logical path from correct solutions, they tend to rely on simple semantic similarity when tackling science problems.
We need to find ways to make scientific discovery faster, especially when exploring massive design spaces like those in fusion research or material science. A new multi-agent framework called MADA uses large language models to coordinate specialized agents that can launch simulations on high-performance computing systems and propose new designs.
When tested on suppressing Richtmyer-Meshkov Instability in fusion experiments, the system successfully automated the iterative design refinement process with very little human intervention. This ability to automate complex workflows is also being applied to molecular optimization through a tool called MolReAct.
Instead of searching through an overwhelming number of possible chemical transformations, this method uses a language model to identify a compact, synthesizable set of reaction steps. By combining this with reinforcement learning techniques like Group Relative Policy Optimization, the system achieved superior sample efficiency and higher top-ten scores across most molecular optimization tasks compared to existing baselines.
The precision required in these specialized fields is mirrored by the need for accuracy in evaluating computer science curricula. Researchers developed a new pipeline to measure alignment between university programs and updated guidelines, such as the shift from CS2013 to CS2023.
They found that while programs are good at covering specific competencies, they struggle to meet required cognitive depth under newer standards, with coverage of certain knowledge units sitting significantly lower than expected. Moving from curriculum design to biological modeling, a framework called CASCADE is being used to predict how gene perturbations affect cells.
By using patient data as a benchmark for real-world expression changes, researchers showed the system can accurately predict the direction of gene changes in several cancer types. While it excels at predicting regulators of cell proliferation, it still struggles with certain lineage-identity transcription factors.
If we want to trust black-box AI agents in high-stakes environments, we first have to understand what they are capable of when they encounter a situation they have not seen before. A new approach called Monte Carlo Query Search treats capability evaluation as an active learning problem using tree search to synthesize specific queries.
By testing these extremal scenarios, researchers can build a mathematical model of an agent's boundaries much more efficiently than through random testing. This need for reliable execution becomes even more pressing when we move from digital sandboxes to the physical world.
Researchers have been testing how large language models can act as reasoning engines for drone swarms using a standardized Web-of-Drones framework. While the models show promise in understanding complex objectives, they still struggle with reliable execution in these closed-loop settings without heavy assistance from planning tools and safety guardrails.
Even when these models make a decision, we should not assume they know why they chose it. New evidence suggests that large language models suffer from superficial beliefs, meaning their verbal explanations often fail to track the actual internal logic driving their choices.
They behave as if they are following a consistent set of priorities, but the stated reasons only partially explain their behavior. The security of AI agents is becoming a massive headache as they gain more control over digital tools.
Researchers have found that these agents are vulnerable to attacks like prompt injection or memory poisoning, which can trick them into using unauthorized tools. To fight back, a new framework introduces universal defenses like Attacker Tool Filtering, which uses anomaly detection to strip out suspicious tools.
This approach has shown it can drop attack success rates to zero across several major models like LLaMA3 and GPT-4 without hurting the agent's ability to finish its job. While we secure the software, we also have to deal with the physical hardware that runs it.
A new method called GPUThor makes Rowhammer attacks on NVIDIA GPUs significantly more devastating by using non-uniform memory access patterns. By targeting specific rows and timing attacks to dodge refresh intervals, researchers achieved up to 23,500 times more bit flips than previous methods, even cracking ECC-protected GPUs.
This hardware vulnerability is part of a broader struggle to manage the gap between how we design systems and how they behave in the wild. A new conceptual pipeline called illusions-awareness aims to stop us from relying on design illusions, where formal design assumptions fail to match real-world runtime reality.
Instead of just trying to make simulations more perfect, this method proposes turning those failures into structured, reusable knowledge for better design decisions. If you are trying to scale up optimization for real-world decisions, you have likely hit a wall with complex constraints that make solvers crawl.
A new framework called PolyFormer tackles this by learning compact polytopic representations of constraint geometries. The researchers saw online solver speedups of up to 6,400-fold and memory savings as high as 99.87% while keeping errors minimal.
This ability to simplify complex structures quickly is just one way we are seeing models struggle with scale, particularly when it comes to how we judge them. We often assume that different bias audits can be used to rank AI models, but new research suggests these tools are terrible at agreeing on which model is better or fairer.
After testing ten different audit instruments across ten frontier models, researchers found that cross-tool rank agreement was no better than random chance. It turns out these tools are not actually measuring the same underlying construct; for instance, some audits over-correct toward certain demographics in hiring while others remain aligned with existing stereotypes.
This lack of consensus extends into the world of spoken dialogue, where personality is much harder to maintain than thought. A new benchmark called RoleBreak shows that even strong speech-to-speech models struggle to stay in character, failing on persona or safety after about ten or eleven turns on average.
While scaling up the language model helps with conversation logic, it does very little to help the model maintain consistent vocal emotions. Even when we think we are evaluating a system fairly, we might be missing the full picture because of how data is recorded.
In studies involving safety reports and vehicle recalls, researchers found that using different versions of the same event can swing model accuracy by as much as 46 percent. This suggests that we cannot just compare models; we must be careful about which specific record and label we use to define success.
If you are worried about security in the age of massive models, you should look at how easy it has become to hide malicious code directly inside them. Researchers have shown that attackers can use the inherent symmetries in model weights to embed stegomalware that is theoretically lossless and requires no retraining.
While previous attempts to defend against this by shuffling weights left many parameters untouched, new work shows we can displace every single parameter through specific permutations to neutralize these threats with minimal impact on performance. This vulnerability is part of a broader struggle to secure the infrastructure that runs these complex systems.
For those managing containerized workloads like Docker or Kubernetes, there is a growing need for zero trust architectures to prevent man-in-the-middle attacks. Moving from security to perception, we are seeing a shift toward making robotic planning much more reliable in messy environments.
Instead of just mapping an image directly to an action, a new approach uses vision-language models to act as probabilistic grounders. This allows robots to plan in belief space, which makes them far more robust when operating under the uncertainty of not knowing exactly what is happening around them.
This ability to reason about uncertain environments is becoming vital for swarms of robots working together in industrial settings. A new framework called CoAdapt uses a large language model as a runtime controller to manage collaborative perception, deciding which robots should share data based on bandwidth and spatial layouts.
It manages to cut communication costs by 38 percent without losing detection precision. The most significant breakthrough today comes from the GRAFT-ATHENA framework, which gives autonomous agents a way to build cumulative scientific knowledge rather than restarting from scratch.
By mapping problems to specific methods through an expandable probabilistic structure, these agentic teams can transfer experience across structurally related tasks. This approach has already yielded results such as developing a hypersonic-flow solver for the Apollo Command Module that matches experimental measurements within 1.8 percent.
While GRAFT-ATHENA focuses on high-level scientific discovery, researchers are also working to ensure that autonomous research systems do not hallucinate through an experiment. A new system called AutoResearch connects idea generation directly to execution through a two-stage process of multi-model cross-review and evidence-based verification.
In testing on the RSICD benchmark, it improved mean Recall from 32.84 to 34.69 while maintaining much higher reliability than its peers, recording only five audit-confirmed issues compared to as many as 27 in other systems. This drive for grounded reasoning is also evident in how we train models for specialized fields like medicine.
A new method called PROSE solves a major flaw in test-time reinforcement learning where models often collapse into repetitive, incorrect answers when trained on medical multiple-choice questions. Instead of rewarding the model for simply agreeing on an answer, PROSE rewards the quality of the underlying reasoning steps.
The challenge of maintaining accuracy in complex environments extends to the digital realm through more realistic human synthesis. To fix mouth movements and facial expressions that often look unnatural in audio-driven models, researchers have introduced a spatial enhancement module using 3D Gaussian Splatting.
By using facial landmarks to guide the selection of spatial points around expression-sensitive areas, they have significantly improved lip synchronization and visual realism. Scaling these complex systems is equally vital for the future of space exploration.
A new permutation-equivariant neural operator has been developed to plan trajectories for massive spacecraft swarms, solving a problem where traditional methods become too computationally expensive as obstacles increase. This method can generalize from a small group of ten satellites to a swarm of 1,000 spacecraft and 11,000 obstacles with zero-shot accuracy.
Even the fundamental way we design neural architectures is being reconsidered to prevent unintended side effects. Research into sparse attention has revealed a phenomenon called routing absorption, where models co-adapt to learned gates in a way that makes them no better than random ones.
In some tests, using hard-mask deployment resulted in a massive perplexity jump from 48.6 to 601.6 compared to dense models, suggesting we need more careful evaluation of how routing affects model adaptation. Finally, there is a push to bring mathematical rigor to the stopgrad operations used throughout machine learning training.
A new regression principle provides a theoretical foundation for these objectives, proving they can converge to true solutions in settings like flow map learning. Applying this principle allows for modified stopgrad placements that can reduce training memory requirements by half.
The most profound shift comes from a new way of defining intelligence itself, moving away from simple processing toward the ability to govern complex systems. This work suggests that true intelligence is found in beings that understand the global consequences of their local actions, allowing them to steer collective systems using precise vibrations rather than blunt force.
By identifying and maintaining meta-stable equilibriums through these subtle movements, we might finally control volatile systems without destroying their performance. This concept of high-level oversight is mirrored in the quest for safer AI evolution through the ANCHOR framework.
Because self-evolving agents often drift into unsafe behaviors during self-play, ANCHOR uses an external large language model to act as a supervisor, providing evaluative feedback to keep the system on track. A similar need for stability appears in how we train autonomous vehicles to handle rare, dangerous moments.
TrafficGamer treats driving as a multi-agent game, using game-theoretic oracles to simulate safety-critical scenarios that are too rare to find in standard datasets. In the realm of large language models, researchers are finding ways to make them more reliable without needing human labels.
A method called RLSF uses a model's own internal confidence during reasoning as a reward signal, allowing the model to learn from its own self-feedback to improve accuracy and calibration. This drive for better multi-agent coordination is furthered by EVINCE, which uses information theory to manage how multiple models debate.
It dynamically shifts the conversation from contentious arguments meant to expose inconsistencies to conciliatory phases that encourage compromise. Protecting the intellectual property of these models is also becoming a priority through clustering-based watermarking.
The CBW method allows dataset owners to verify their work by embedding triggers into clusters of speaker embeddings, making it possible to detect unauthorized commercial use. Finally, this ability to navigate complex networks is being applied to the microscopic scale of cancer genomics with RegNetAgents.
This multi-agent framework integrates different types of gene regulatory networks to identify specific cancer drivers, helping scientists move from raw data directly to biological hypotheses about drug targets.
Today's papers
- Shared Selective Persistent Memory for Agentic LLM Systems This architecture stores reusable context like tool configs while discarding session-specific reasoning to improve efficiency. [paper] [episode]
- Rhythm of the Deep: Two-tier acoustic organization of sperm-whale codas from click waveforms to second-order sequence dependence Sperm whale communication uses both click waveform details and timing patterns to organize complex sequences. [paper] [episode]
- The Verifier is the Curriculum: Precision Sets the Return on Search in Code Self-Distillation Using a strict code execution check as a training signal significantly improves model generalization without needing a complex reward model. [paper] [episode]
- Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation Large language models can simulate student misconceptions to create plausible wrong answers for multiple-choice questions. [paper] [episode]
- FactorEngine: A Program-level Knowledge-Infused Factor Mining Framework for Quantitative Investment This framework uses LLMs and executable code to discover stable, interpretable predictive signals for financial markets. [paper] [episode]
- MASCOT: Multi-Agent Socio-Collaborative Companion Systems This multi-agent framework prevents generic behavior by optimizing individual personas and collaborative dialogue. [paper] [episode]
- Liberating LLM Capabilities in Full-Duplex Speech Models This method allows speech models to use text as a primary output channel for complex reasoning during real-time interaction. [paper] [episode]
- Multi-Agent Collaboration for Automated Design Exploration on High Performance Computing Systems This multi-agent system automates complex scientific design workflows by coordinating specialized agents for simulation and optimization. [paper] [episode]
- MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning This framework uses mixture modeling to capture diverse human preferences for more personalized AI alignment. [paper] [episode]
- Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023 This study uses an LLM-based pipeline to evaluate how computer science curricula align with evolving educational standards. [paper] [episode]
- LLM-Guided Dynamic Action Spaces for Synthesizable Molecular Optimization This tool uses LLMs to identify feasible chemical reaction paths, making multi-step molecular optimization more efficient. [paper] [episode]
- Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering Visual question answering systems tend to favor information at the start of a prompt rather than the end. [paper] [episode]
- Robustness as an Emergent Property of Task Performance Model robustness appears to emerge naturally as a model masters increasingly difficult tasks. [paper] [episode]
- Autonomous Assessment of Generalizability of AI Agent Capabilities This method uses active learning and Monte Carlo tree search to automatically map out the capability boundaries of AI agents. [paper] [episode]
- Dual Randomized Smoothing: Beyond Global Noise Variance This technique uses input-dependent noise levels to provide better certified robustness for neural networks at various scales. [paper] [episode]
- CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction This framework predicts how gene changes affect biology by validating LLM predictions against real patient data. [paper] [episode]
- Hierarchical Modeling of ICD Codes in EHR Foundation Models Encoding the hierarchical structure of medical codes improves how AI models represent clinical data. [paper] [episode]
- Say the Mission, Execute the Swarm: Agent-Enhanced LLM Reasoning in the Web-of-Drones This framework allows users to control drone swarms using natural language by grounding LLM reasoning in standardized web protocols. [paper] [episode]
- Superficial Beliefs in LLM Decision-Making Large language models often make decisions based on patterns that do not match the specific reasons they provide. [paper] [episode]
- Generating Individual Travel Diaries Using Large Language Models Informed by Census and Land-Use Data This method uses census data to prompt LLMs into generating realistic synthetic travel patterns for transportation modeling. [paper] [episode]
- Permutation-Based Stegomalware in Large Language Models: Threats and Countermeasures Researchers have found that malicious code can be hidden in model weights using symmetries that are difficult to detect. [paper]
- CoAdapt: An LLM-based Framework for Adaptive Collaborative Perception in IIoT Robotic Swarms This framework uses an LLM to manage how robots share data, saving bandwidth while maintaining perception accuracy. [paper]
- Sample-Conditioned Representation Selection for Audio Few-Shot Learning This method improves audio classification by using a learned mask to select only the most relevant features from an input. [paper]
- Channel-Informed Neural Network for Physical Layer Key Generation This approach uses wireless channel measurements to help devices establish secure, shared cryptographic keys. [paper]
- Same Flow, Different Paths: Variance Reduction in Flow Matching Optimizing the mathematical path taken during flow matching can significantly speed up model convergence. [paper]
- The record is part of the task: matched-record evaluation of text classifiers across maintenance, safety and recall reporting Model evaluations should use consistent data records to avoid bias caused by different documentation stages. [paper]
- RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue This benchmark tests whether spoken dialogue models can maintain a consistent persona over long conversations. [paper]
- Deconstructing Stereotypes: Scope-Conditioned Generation for Effective Multilingual Counterspeech This method improves automated counterspeech by explicitly teaching models to address the underlying stereotypes in hate speech. [paper]
- Are We Grading Properly? Understanding Failure Modes in Medical Benchmarks Flawed grading rubrics can lead to inaccurate evaluations of AI models in medical benchmarks. [paper]
- Autonomous Droplet Navigation via Model-Based Reinforcement Learning An AI agent can learn to navigate liquid droplets through complex microfluidic paths using reinforcement learning. [paper]
- HUMAID-NER: A Disaster Tweet Dataset for Joint Named Entity Recognition and Event Classification via Uncertainty-Weighted Multitask Learning This dataset provides annotated disaster tweets to help models better identify entities and events during humanitarian crises. [paper]
- Skill-based Agentic Evaluation for Real-time Data Science Tasks This framework evaluates data science agents by running their code against live, ever-changing datasets. [paper]
- Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks These modular defenses protect AI agents from being manipulated through malicious tool calls or prompt injections. [paper]
- Towards Illusions Awareness in Cyber-Physical System's Design This research proposes a way to turn errors caused by unrealistic design assumptions into useful knowledge for engineers. [paper]
- Can We Stop The Ads? Taxonomy and Characterization of Smartphone Splash Ads and Existing Countermeasures This study analyzes the difficulty of preventing full-screen pop-up ads that interfere with mobile app usage. [paper]
- Bounded Adjustment with Reliability-Guided Embedding for Imbalanced Learning with Noisy Labels This method improves machine learning performance on imbalanced datasets by accounting for noisy or incorrect labels. [paper]
- Toward Governance-Aware Autonomous GIS: A Narrative Review of Ethical and Privacy Risks in LLM-Enabled GeoAI This review identifies privacy and bias risks when using large language models to automate geographic information systems. [paper]
- Structural Negative Transfer in Federated Graph Neural Networks: Diagnosis, Causal Investigation, and the Limits of Divergence-Aware Mitigation Joining a federated learning network can actually hurt a model's accuracy if the data structures are too different. [paper]
- FlashVector: Agent for Hierarchical Model Serving Stack Optimization This agentic system optimizes every layer of a recommendation system to increase throughput and reduce latency. [paper]
- Psychological Effects of Cultural Upheavals from Millions of Song Lyrics Over 100 Years Analysis of song lyrics shows that major historical events change how artists express distress and social connection. [paper]
- GPUThor: Amplifying Rowhammer Attacks via Non-Uniform Patterns to Exploit ECC-Protected GPUs This attack technique significantly increases the effectiveness of memory-based exploits on modern GPUs. [paper]
- JustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management This runtime enables large language models to handle massive amounts of text locally on standard consumer laptops.
- Neuro-Symbolic Synergy for World Modeling This framework combines the reasoning of symbolic logic with the linguistic power of LLMs to create more accurate world models. [paper]
- Learning efficient representations of complex constraints for scalable optimization This method turns complex geometric constraints into simple mathematical forms to speed up optimization problems. [paper]
- Reduplicative constructions in Mandarin: Socio-emotional profiling through distributional semantics Word embeddings can effectively capture the subtle grammatical and semantic meanings of repeated words in Mandarin. [paper]
- ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals Automated grading rubrics can be tricked by tasks that are designed to be impossible to answer honestly. [paper]
- PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress This agentic framework acts as a digital advisor by providing auditable, evidence-based critiques of scientific drafts. [paper]
- Affect-Prototype Guided Fusion for Open-Vocabulary Incomplete Multi-modal Emotion Recognition This framework helps AI recognize emotions even when some data, like video or audio, is missing. [paper]
- Test-Time Unlearning via Sparse Autoencoder This method allows users to remove specific knowledge from a model during inference without needing to retrain it. [paper]
- Copula Adapted Directed Acyclic Graph for Cluster Representation of Biomedical Data This approach uses causal relationships and statistical models to cluster complex biomedical data more accurately. [paper]
- Splitting the Difference: Interpretable Causal Forests for Treatment Effect Heterogeneity and Bias This algorithm estimates how different individuals respond to treatments while remaining easy for humans to interpret. [paper]
- Bridging the Confidence Gap: Temperature Scaling for Calibrating Test-Time Prompt Tuning This method improves how confident an AI model's predictions are without losing any accuracy. [paper]
- gr-PHYSEC: Real-time Channel-based Key Generation for Physical Layer Secure Wireless Communications This software module enables real-time, secure wireless communication by generating keys from radio signals. [paper]
- Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration This defense protects AI safety features by tricking attackers into targeting useless parts of the model instead of the actual safety guardrails. [paper]
- Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models While different audits can detect bias, they often disagree on which models are more biased than others. [paper]
- When Agents See Differently: Exposing UI Desynchronization Threats in Mobile Agents This study shows that attackers can exploit the difference between what a human sees on a screen and what an AI agent perceives. [paper]
- What Does Layer-Importance Reveal About Transformers and State-Space Models? Analysis shows that transformers and state-space models process information across their layers in fundamentally different ways. [paper]
- The Role of Implicit and Explicit Demographic Signals in Large Language Model-based Student Assessment LLMs can show bias in student grading based on both stated demographics and subtle conversational cues. [paper]
- VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs This framework speeds up video understanding by focusing high-resolution processing only on the most important parts of a video. [paper]
- From Manual Construction to AI-Driven Scenario Emergence: Rethinking Catastrophe Risk Modeling This framework uses AI weather models to generate extreme weather scenarios much faster and cheaper than traditional methods. [paper]
The papers
- Shared Selective Persistent Memory for Agentic LLM Systems — This paper introduces a "shared selective persistent memory" architecture designed to solve the statelessness of agentic LLM systems. [episode]
- Rhythm of the Deep: Two-tier acoustic organization of sperm-whale codas from click waveforms to second-order sequence dependence — This paper investigates the acoustic structure of sperm-whale codas, challenging the traditional view that these signals are merely recurring click-count and timing patterns. [episode]
- The Verifier is the Curriculum: Precision Sets the Return on Search in Code Self-Distillation — This paper investigates how to optimize code generation through self-distillation without falling into "reward hacking," where models exploit learned judges by manipulating surface features rather than improving functionality. [episode]
- FactorEngine: A Program-level Knowledge-Infused Factor Mining Framework for Quantitative Investment — This paper introduces FactorEngine (FE), a program-level factor discovery framework designed to automate the mining of predictive signals from noisy, non-stationary market data. [episode]
- MASCOT: Multi-Agent Socio-Collaborative Companion Systems — This paper introduces M ASCOT, a multi-agent framework designed to develop "multi-perspective socio-collaborative companions" capable of providing emotional and cognitive support. [episode]
- Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation — This paper investigates whether large language models (LLMs) can effectively model "incorrect yet plausible reasoning" through the task of multiple-choice distractor generation. [episode]
- Liberating LLM Capabilities in Full-Duplex Speech Models — This paper presents Listen-Write-Speak (LWS), a "text-first tri-channel paradigm" designed to overcome the inherent limitations of speech-based large language models. [episode]
- Multi-Agent Collaboration for Automated Design Exploration on High Performance Computing Systems — This paper presents MADA (Multi-Agent Design Assistant), an LLM-powered multi-agent framework designed to automate complex scientific design workflows on high-performance computing (HPC) systems. [episode]
- MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning — This paper introduces MiCRo, a two-stage framework designed to enhance personalized preference learning in Large Language Models (LLMs). [episode]
- Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023 — This paper introduces a "reproducible, human-in-the-loop pipeline" designed to measure how undergraduate computer science programs align with international curricular guidelines. [episode]
- LLM-Guided Dynamic Action Spaces for Synthesizable Molecular Optimization — This paper introduces MolReAct, a framework designed to bridge the gap between computational molecular optimization and practical drug development by ensuring that proposed structural modifications correspond to "feasible synthetic routes." It addresses the critical challenge of [episode]
- Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering — This paper investigates position dependence in multimodal Knowledge-Based Visual Question Answering (KB-VQA) systems. [episode]
- Dual Randomized Smoothing: Beyond Global Noise Variance — This paper introduces Dual Randomized Smoothing (Dual RS), a framework designed to overcome the "fundamental limitation" of standard Randomized Smoothing (RS). [episode]
- CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction — CASCADE is an agentic framework designed to predict the "downstream transcriptional consequences of perturbing a focal gene" via directed propagation through tumor-specific ARACNe networks. [episode]
- Hierarchical Modeling of ICD Codes in EHR Foundation Models — This paper investigates the use of ICD-10-CM hierarchy as a "general inductive bias for clinical representation learning" within Electronic Health Record (EHR) foundation models. [episode]
- Autonomous Assessment of Generalizability of AI Agent Capabilities — This paper introduces Monte Carlo Query Synthesis (MCQS), an active query-synthesis method designed to learn symbolic stochastic capability models of black-box AI (BBAI) systems. [episode]
- Robustness as an Emergent Property of Task Performance — This paper investigates whether model robustness—defined as output consistency across prompt variations—is an independent capability or an emergent property of task mastery. [episode]
- Say the Mission, Execute the Swarm: Agent-Enhanced LLM Reasoning in the Web-of-Drones — This paper presents a mission-agnostic, agent-enhanced LLM framework designed for real-time UAV swarm management. [episode]
- Superficial Beliefs in LLM Decision-Making — This paper investigates whether large language models (LLMs) possess a systematic underlying decision structure or merely "imitate the language of belief." By comparing an LLM's actual choices against its explicit justifications, the researchers aim to determine if model reasonin [episode]
- Generating Individual Travel Diaries Using Large Language Models Informed by Census and Land-Use Data — This paper introduces a Large Language Model (LLM) scheme for generating individual travel diaries in agent-based transportation models to address the limitations of traditional approaches. [episode]
- SciNLP: A Domain-Specific Benchmark for Full-Text Scientific Entity and Relation Extraction in NLP —
- MeshODENet: A Graph-Informed Neural Ordinary Differential Equation Neural Network for Simulating Mesh-Based Physical Systems —
- GraphIFE: Rethinking Graph Imbalance Node Classification via Invariant Learning —
- Risk-Calibrated Bayesian Streaming Intrusion Detection with SRE-Aligned Decisions —
- Manifold Dimension Estimation via Local Graph Structure —
- Sockeye: Bug-finding and proofs for platform configurations and hardware based on reference manuals —
- GeoCrossBench: Cross-Band Generalization for Remote Sensing —
- DAPS++: Rethinking Diffusion Inverse Problems with Decoupled Posterior Annealing —
- Training Energy-Based Models with Non-MCMC Samplers and Efficient Temperature Estimation —
- Formalized Hopfield Networks and Boltzmann Machines —
- Nonnegative matrix factorizations and related compositional models: Equivalence, identifiability, and an application on the grain-size analysis of sediments —
- Collaborative Optimization of Multiclass Imbalanced Learning: Density-Aware and Region-Guided Boosting —
- MOZAIK: A Privacy-Preserving Analytics Platform for IoT Data Using MPC and FHE —
- SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages —
- How Clinicians Think and What AI Can Learn From It —
- Window-Diffusion: Accelerating Diffusion Language Model Inference with Windowed Token Pruning and Caching —
- There Is More to Refusal in Large Language Models than a Single Direction —
- Same Answer, Different Representations: Hidden instability in VLMs —
- Neuro-Symbolic Synergy for World Modeling —
- PRISM: Parallel Residual Iterative Sequence Model —
- Partial recovery of meter-scale surface weather —
- Strategic Advice in the Age of Personal AI —
- Routing Absorption in Sparse Attention: Why Random Gates Are Hard to Beat —
- Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization —
- Learning efficient representations of complex constraints for scalable optimization —
- Deep Invertible Autoencoders for Dimensionality Reduction of Dynamical Systems —
- PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark —
- Electrodermal Activity as a Unimodal Signal for Aerobic Exercise Detection in Wearable Sensors —
- Thinking Deeper, Not Longer: Memory-Efficient Test-Time Reasoning with Depth-Recurrent Transformers for Compositional Generalization —
- A Spectral Decomposition Framework for Multiscale Nonlinear Dimensionality Reduction —
- EviDep: Uncertainty-Aware Multimodal Depression Estimation via Disentangled Evidential Learning —
- GPUBreach: Privilege Escalation Attacks on GPUs using Rowhammer —
- GRAFT-ATHENA: Self-Improving Agentic Teams for Autonomous Discovery and Evolutionary Numerical Algorithms —
- YFPO: Yoked Feature Preference Optimization with Neuron-Guided Rewards —
- Structure-Aware Masking for Protein Representation Learning —
- Benchmarking Machine Learning Architectures for Antimicrobial Stewardship in Pediatric ICUs —
- Stream Assembly Is an Uncontrolled Treatment in Streaming Intrusion-Detection Benchmarks —
- Does Continued Pretraining on a Learner Corpus Improve Automated Essay Scoring on English Proficiency Tests? Evidence from EFCAMDAT —
- BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning —
- Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agents —
- Human Factors in Cybersecurity in Icelandic Small and Medium-sized Enterprises —
- ANCHOR: An External LLM-Driven Supervisory Module Facilitating Healthy Evolution in Self-Evolving Systems —
- Do LLMs Make Neural Distinguishers Wise? —
- Shielded Analysis: Certification and Characterization of Defensibility in Systems under Adversarial Interaction —
- Learning aligned EEG representations with subject-specific encoders —
- Tail-Shape Estimation in LLM Evaluation Is Fragile: A Protocol for Diagnosing False Positives —
- Amortized Probabilistic Retrieval of Atmospheric CO2 from OCO-2 Spectra Using Deep Learning with Laplace Approximations and Normalizing Flows —
- The Scissors Effect: When Resize-Based Input Diversity Helps or Hurts Transfer Attacks —
- A Comparative Study of Bayesian Contextual Bandits for Real-Time Warehouse Sorter Optimization —
- FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models —
- Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting —
- Data-Driven Soft Labeling Scales DNA Read Classification to Whole-Body Cell-Type Deconvolution —
- Physics-Informed Machine Learning Under Small-Data Constraints: Lessons from Abrasive Waterjet Milling —
- RegNetAgents: A Multi-Agent Framework for Cross-Network Regulatory Driver Identification in Cancer Genomics —
- The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory —
- surprisal is Not a Theory —
- Neural Operator Learning for Collision-Aware Trajectory Planning of Spacecraft Swarms —
- AutoResearch: Insight In, Hallucination Out —
- Random Hazard Forests —
- Activation-Weighted Seeded Residual Coding for Low-Bit LLM Weight Repair —
- Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 Architectures —
- The Functionalizer: Lossless Functional Decomposition for Subword Tokenization —
- Optimal Model Activation Policies for Inference Networks of Large Language Models —
- Single Document Extractive Summarization using Domination in Hypergraph —
- Latent Undertow: How Ordinary Typos Break Probes —
- Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models —
- Comment on arXiv:2607.01233: Survivorship Bias in Published-Paper Baselines for Research-Idea Distributions —
- Crash Narrative-Guided Countermeasure Recommendation Using Large Language Models: A Retrieval-Augmented Generation Framework for Intersection Safety —
- Self-reported archetypes and behavioral failures in Large Language Models —
- NepKANUN: A RAG-Based Nepali Legal Assistant —
- Nepali Legal Expertise through Generative and Extractive Pre-trained Transformers (NepLEGiT) —
- ViCo: Visual-oriented Coding with Self-Reflection for Chart Replication —
- Are We Grading Properly? Understanding Failure Modes in Medical Benchmarks —
- Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents —
- Causal neural set filtering for online multi-target tracking —
- State of Thought Enables Endogenous Reasoning —
- Managing Action Preconditions in Neuro-Symbolic RL: Three Placement Strategies for Embodied Agents —
- OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning —
- Driver Behavior Estimation at Signalized Intersections Using a Physics-Constrained Decision-Conditioned Autoregressive Transformer —
- Towards Scalable RLVR: Multimodal Instruction Following Data Synthesis and Distillation —
- HintMiner: Automatic Question Hints Mining From Q&A Web Posts with Language Model via Self-Supervised Learning —
- POSPAN: Position-Constrained Span Masking for Language Model Pre-training —
- Signed p-adic Residual Encodings of Finite-Domain All-Different Systems with a Sudoku Case Study —
- You Don't Need To Train: Agentic Heuristic Learning Studio for Executable Human Activity Recognition —
- A panoramic aerodynamic performance prediction method for turbomachinery cascades using transformer-enhanced neural operator —
- A Dynamic Aggregation Strategy Enhanced Efficient Global Optimization Algorithm for Solving High-Dimensional Turbomachinery Design Problems —
- Beyond Distribution Matching: Semantics-Consistent Tabular Diffusion with Weak Semantic Priors —
- Efficient Multimodal Generative Recommendation with Latent Narrative Reasoning —
- World Models for Cross-Machine CNC Transfer under Partial Sensor Overlap —
- The Immutable Past: Formalizing State Mutability and Conflict Resolution in Mutable RAG —
- The Imitation Game: When LLMs Learn to Reason Like Programs via Code-Centric Reasoning Data Synthesis —
- Pseudo-Label Augmentation for Affect Sensing in Small Collaborative Groups —
- Distilling Foundation Models for Agentic What-If Reasoning:Cost, Latency, and Governance in a Hybrid LLM+SLM Architecture —
- Evaluating Open-Weight E-Commerce Agents with Environment-Grounded Verification —
- RAG-CT: Mitigating Privacy Risks on Retrieval-Augmented Generation Systems via Scanning Prompt Distribution —
- Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks —
- SWB-DM: A Calibrated Sliced-Wasserstein-Barycenter Aggregator with Delayed-Momentum Caching for Byzantine-Robust Federated Learning under Partial Participation —
- A Decision-Support Audit Protocol for Supervision Drift in Proxy-Labeled Credit-Risk Prediction —
- Optimal Pruning for Neural Architectures using Fisher Information Distances —
- Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation —
- LLMs as Master Forgers: Generating Synthetic Time Series Data for Manufacturing —
- LLM Inference in a Flash! —
- GPEvac: GNN-Based PPO for Adaptive Evacuation Routing During Shooting Events —
- Skeletal Prototypes on Iterative Nerve Expansions —
- Z-Loss Backward Geometry in Dense Output Heads and Sparse Routers —
- Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It —
- Position: AI Is Not Ready for Strategic Conflicts —
- Permutation-Based Stegomalware in Large Language Models: Threats and Countermeasures —
- Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration —
- Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving —
- Feasibility of Homomorphic Inference for a Genomic Foundation Model —
- Artificial intelligence and biosecurity: capabilities, threat pathways, and defense-in-depth governance —
- Analyzing Multi-Factor Authentication Through Cryptographic Security Properties —
- Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions —
- How I learned to stop worrying and love StopGrads: Stationarity, Convergence, and a case study on Flow Map Learning —
- Test-Time Unlearning via Sparse Autoencoder —
- RuleAutoPilot: Synthesizing Deployable Suricata Rules from Network Traffic —
- Toward Governance-Aware Autonomous GIS: A Narrative Review of Ethical and Privacy Risks in LLM-Enabled GeoAI —
- Copula Adapted Directed Acyclic Graph for Cluster Representation of Biomedical Data —
- Metacognitive Steering: Learning the Structure of Scientific Judgment —
- The Pain Axis: LLMs Represent Self-Directed Harm and Act on It —
- CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design —
- Exploiting and Securing Docker containers and Kubernetes pods from a MitM attack —
- Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning —
- The AI-Enabled Scientific Frontier —
- Compute-Optimal Pretrain--Fine-tune in Ridge Gradient Descent —
- The record is part of the task: matched-record evaluation of text classifiers across maintenance, safety and recall reporting —
- Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act —
- Speaker-Specific and Language-Dependent Temporal Organization in Bilingual Political Speech —
- Speaker or Language? Explaining Variance in Charismatic Prosody Across Luxembourgish and French —
- Scaling Laws for Physics-Aware ACOPF Surrogate Learning —
- Differentially Private Semantic Plans for Aggregate Insight Generation —
- Drift Field Net: Learning Ocean Lagrangian advection fields from in-situ and satellite observations —
- Closing the Loop: Branch-and-Bound for Scalable Verification of Nonlinear Neural Feedback Systems —
- CLEAR: Cross-Source Evidence Adjudication for Large Language Models in Medicine —
- BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents —
- Agentic Search Spaces for Tabular Machine Learning —
- Efficient One-to-Many Translation with Joint Multi-Stream Diffusion —
- Robust Fault Detection in Mechanical Multimodal Time Series via Self-Supervised Cross-Modal Reconstruction —
- Generative models for simulation based filtering: Formulations and Empirical Comparisons —
- Cross-Anatomy Transfer Versus Sparse Interpolation in Digital-Twin-Oriented Aortic Fluid-Structure Interaction Surrogates —
- Understanding the Usability of Cryptographic Verification Tools —
- Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation —
- Breaking the 1.58-bit Barrier for Ternary LLMs —
- StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation —
- Channel-Informed Neural Network for Physical Layer Key Generation —
- Multi-Label Proportion Learning for Sea-Ice Type Prediction —
- Federated stochastic bilevel optimization with fully first-order gradients —
- Mini-batch Sampling Strategies for Long-Tailed Image Classification: An Empirical Study on CIFAR-100-LT —
- How Humans and LLMs Read Gender into "Gender-Neutral" Physical Descriptions —
- Autonomous Droplet Navigation via Model-Based Reinforcement Learning —
- Register Tokens for Bounded-State Reasoning in Diffusion Language Models —
- Certified Uncertainty Propagation in One-Shot Federated Bayesian Models via Posterior Event Transport —
- gr-PHYSEC: Real-time Channel-based Key Generation for Physical Layer Secure Wireless Communications —
- Bounded Adjustment with Reliability-Guided Embedding for Imbalanced Learning with Noisy Labels —
- Attention Mean Fields Predict Average Representation Dynamics and Reveal Context-Specific Computation —
- ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian —
- Negation Beyond the Verbal Channel: Temporal Multimodal Correlates in Dialogue —
- Implementing a White-Box Undetectable Backdoor for Random Fourier Features —
- How Good Are Time-Series Foundation Models for Pedestrian Crowd Count Forecasting? A Cross-Dataset Comparative Study —
- No Bit Left Behind: Using Brute-Force Lifting to Achieve Fully Static Binary Recompilation —
- ReMova: Fine-tuning LLMs for English to Belarusian translation —
- Evaluating the NIST Bugs Framework Against CWE as a Successor for Automated Vulnerability Classification —
- Interpreting and Steering LLM Agents for Social Simulations —
- Learned Look-Ahead Splitting Rule for CART —
- Adaptive Bayesian Partner Selection for Federated Clinical Centers —
- Early-Bird Decoding: Accelerating Diffusion LLMs with Learnable Block Sizes and Parallel Sampling —
- Fine-Tuning Fixes Mode Collapse and Over-Dispersion in LLMs —
- OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation —
- Not All Relations Are Equal: Relation-Balanced and Calibrated Graph Learning for Provenance-Based Intrusion Detection —
- Online Gradient Computation for Warping Gaussian Process Transformations —
- Certified Inference and Training for Deep Equilibrium Networks: A Continuation Framework with Polynomial Complexity Guarantees —
- Skill-based Agentic Evaluation for Real-time Data Science Tasks —
- Decoder Design Matters for ECG Delineation —
- From Manual Construction to AI-Driven Scenario Emergence: Rethinking Catastrophe Risk Modeling —
- High-Performance Tensor Formulation of the Viterbi Algorithm for Hidden Semi-Markov Models —
- Beyond the Name: Demographic Leakage in De-Identified R'esum'es and Evaluation Artifacts in LLM Bias Audits —
- Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening —
- AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research —
- FlowATC: Aircraft Trajectory Prediction via Flow Matching —
- Style-Debiased DPO: Updating LLM Knowledge with Factuality-Aware Synthetic Preference Data —
- What Does Layer-Importance Reveal About Transformers and State-Space Models? —
- On the Importance of Gating: Memorization vs. In-Context Learning in State Space Models —
- A Cyber Range Evaluation of Autonomous Network Incident Response Agents —
- GPUThor: Amplifying Rowhammer Attacks via Non-Uniform Patterns to Exploit ECC-Protected GPUs —
- QueryFormer: Winning Solution for KDD Cup 2026 Tencent UniRec Challenge —
- PunGraph: Retrieval-Enhanced Phonetic-Semantic Graph Reasoning for Pun Understanding —
- The MAL Simulator: Cyber Operations Simulation based on Attack & Defense Graphs —
- Query-Aware Source-Risk Triage for Retrieval-Augmented Generation —
- AsyncCouple-Flow: Asynchronous Cross-Modal Coupling and Flow Matching for Spatio-Temporal Forecasting —
- Recovering Governing Dynamics from Distributed Observations via Exact Spline Merging —
- Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models —
- Challenges of Auditing: Variability in Outputs of Large Language Models for Health —
- A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance —
- A Weighted Kernel Method for Approximation that Adapts to Learned Multivariable Structure —
- RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue —
- Divergence Timing and Cumulative Disagreement under KV-Cache Eviction —
- Stable by Construction: Variational Latent Markov Operators for Long-Horizon PDE Prediction —
- Quantifying Organizational Environmental Action from Web Data and Large Language Models —
- EchoPath: Execution-Level Replayable Memory for GUI Agents —
- ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training —
- GrowMTP: Can RL Grow Its Own Draft Head? —
- Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA —
- DiaWhisper-DPO: Role-Attributed Transcription of Clinical Interviews via Failure-Mined Preference Optimization —
- Right Direction, Wrong Step: Geometric Analysis of Finite-Step Failure in Looped Transformers —
- ANIMASK: What the Model Contributes to Role Play in Simulated Story Worlds —
- From Hypervisor to Container: Cloud Security Vulnerabilities, Defense Mechanisms, and Open Challenges —
- AI for Games in the Foundation Model Era —
- little m: An AI Agent for Industrial Process Optimization —
- MarkSec: Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks —
- Toward Secure AI-Powered Penetration Testing Agents: Security Threats, Guardrails, and Architectural Perspectives —
- Continuous-Time Machine Learning: A Unified Mathematical Perspective —
- VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs —
- LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture —
- When Agents See Differently: Exposing UI Desynchronization Threats in Mobile Agents —
- Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models —
- A Systematic Evaluation of Machine Learning Methods for Fault Detection and Line Identification in Electrical Power Grids —
- TIAO: Token Importance-Aware Policy Optimization for Text Summarization —
- Beyond Episodic AI: Cognitive Field Networks for Biologically Inspired Persistent Cognition —
- TAME: Token Attribution and Masking for Emergent misalignment —
- Turn-level Multiscale Density Ratio Estimation for LLM Agents —
- Coverage-Aware Virtual IMU Augmentation for Low-Resource Human Activity Recognition —
- Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion —
- Integrating the Analytic Hierarchy Process with Large Language Models for Transparent Multi-Criteria Decision-Making —
- Noise2Noise Revisited: Training Pair Distributions Dominate Loss Choice in Self-Supervised Denoising —
- Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models —
- Time-warping estimation via stationarity-based learning of the de-warped signal —
- Smarter by the Moment: Environment-Driven Dynamic Policies for Continual LLM Improvement —
- On the disintegration of the stochastic majority vote: From PAC-Bayesian bounds to a self-bounding algorithm —
- SOTER: A Generative Time-Series Foundation Model for Wearable Human Physiological Signals —
- Geometry of learning dynamics: Gradient descent versus natural gradient on the ridge of optimization —
- Can We Do Interpretable NLI with Graphs Based on Atomic Propositions? —
- ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals —
- InceptionRAG: Stealthy Poisoning Attack Against Retrieval-Augmented Generation —
- Execution Flexibility in Automated Planning: A Comparative Evaluation of Deordering and Reordering Strategies —
- LCAP: Population-Informed Latent Chip Adaptation from Few Output Probes for Photonic Neural Networks —
- Adapting to Decision-Relevant Non-Stationarity in Decentralized Heterogeneous Bandits —
- Information Geometric Self-Organization at the Edge of Stability in High-Capacity Kernel Associative Memories —
- CoAdapt: An LLM-based Framework for Adaptive Collaborative Perception in IIoT Robotic Swarms —
- Can Deep Learning Achieve Cross-Physics Mapping? —
- A Data-free Universal Prior over Syntactic Structures —
- Reduplicative constructions in Mandarin: Socio-emotional profiling through distributional semantics —
- Bridging Learned Visual Perception and Symbolic Belief-Space Planning —
- QART: A Quantum-Classical Hybrid Architecture for Long-Horizon Reasoning -- Exploring a Conditional Path toward Quantum Scaling —
- Cascade: Hierarchical Recoverability Control for Large Language Model Unlearning —
- RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation —
- Deconstructing Stereotypes: Scope-Conditioned Generation for Effective Multilingual Counterspeech —
- Lit3R: Retrieve-Relate-Read for Evidence-Grounded Question Answering over Scientific Literature —
- ROSETTA: Efficient and Accurate Privacy-Preserving LLM Decoding via Hybrid CKKS/TFHE Evaluation —
- HyCoSeq: Contextual Hyperbolic Representation Learning for Genomic Sequences —
- Verbalizing Subliminal Learning Effects Using Text Optimization —
- Cybersecurity in Power Grids: Standards and Research Challenges —
- Repurposing Deep Limit Order Book Forecasting for Scenario-Conditioned Market Impact Modeling —
- When Confidence Signals Disagree: Local and Global Confidence in Autoregressive Language Models —
- Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation —
- AntennaFlow: A Generative Flow Model for Offset Correction in Phaseless Antenna Testing —
- Affect-Prototype Guided Fusion for Open-Vocabulary Incomplete Multi-modal Emotion Recognition —
- HUMAID-NER: A Disaster Tweet Dataset for Joint Named Entity Recognition and Event Classification via Uncertainty-Weighted Multitask Learning —
- Target-Language Generation in Multilingual Models: Activation Steering and Optimal Control —
- Splitting the Difference: Interpretable Causal Forests for Treatment Effect Heterogeneity and Bias —
- Structural Negative Transfer in Federated Graph Neural Networks: Diagnosis, Causal Investigation, and the Limits of Divergence-Aware Mitigation —
- Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs —
- Autoformalizing Argumentative Material Inferences —
- The Role of Implicit and Explicit Demographic Signals in Large Language Model-based Student Assessment —
- PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress —
- Can LLMs Follow the Pulse of a Crisis? Evaluating Crisis Sentiment in Bangladesh's July Uprising —
- FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference —
- ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents —
- ORDER: Task-Conditioned Routing for Retrieval-Augmented Generation —
- SKIP: a Self-knowledge-guided Step-wise Preference Learning Framework for Concise Reasoning —
- CLARE: Scalable Class-Incremental Continual Learning via a Sparsity-Based Framework —
- Distributed JEPA: A Self-Supervised Framework for Energy Forecasting —
- Sparse MLLM Anchors, Dense Adaptation: Breaking the Self-Referential Loop in Wild Test-Time Adaptation —
- Learning Options for Compositional Motor Control with Adapter Banks —
- Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering —
- Repurposing Unified Topological Signatures for Graph Representation Learning —
- Neuro-Symbolic Hierarchical Intention Anticipation in Human Behavior —
- Sample-Conditioned Representation Selection for Audio Few-Shot Learning —
- EviScope: Paired Counterfactual Evidence Diagnostics for Faithful and Efficient Grounded Language Models —
- Interactive Memory Learning for Long-Term Conversations —
- Scaling-Score Conformal Prediction for Multi-Target Regression —
- Semi-Supervised Learning-Based Genetic Biomarkers Dataset for Multiple-Stage Hepatocellular Carcinoma Prediction —
- High-Fidelity Digital Twin Data Models by Randomized Dynamic Mode Decomposition and Deep Learning with Applications in Fluid Dynamics —
- Symbolic Separation: Grounding Deep Agents in Knowledge Graphs for Trustworthy Operational Data Analytics —
- Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs —
- An Empirical Study of Counterfactual Self-Explanations in LLMs —
- FirmCORe: A Benchmark for Structured Reasoning about Inter-Firm Collaboration Opportunities —
- Observational Indistinguishability and Integrity Blind Regions in Hybrid Quantum-Classical Workflows —
- Neural Field Ensembles for Aerodynamic Surface Prediction: Winning Solution to the ONERA CRM Wall Distribution 2025 Challenge —
- Plug 'n' Pray: Agentic LLM-based Detection of Potential Log File Exposures in Third-Party Content Management System Plugins —
- A unified framework for global and local interpretability using adaptive derivative-ordered random explanation —
- IRENE: A Convolutional GRU Ensemble Model for Radar Precipitation Nowcasting over Italy —
- MOCC-R1: Reinforcing Reasoning-Response Consistency for Multimodal Counselor Response Generation —
- LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers —
- End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services —
- MyoFlow: Anchor-Tied Rectified Flow for HD-sEMG Gesture Recognition Across Sessions and Subjects —
- Cross-Domain Inference for Human Localization: Applying Wi-Fi RSSI Data to CSI-Trained Models —
- Memorisation bias in medical AI —
- Psychological Effects of Cultural Upheavals from Millions of Song Lyrics Over 100 Years —
- Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record —
- AraMIP: Extending MIPVU Towards Metaphor Identification in Arabic —
- ECHO: Early-layer Collaborative Hierarchical Orchestration with Bonus Logits in Speculative Decoding —
- Persistent Recurrent Memory Between Transformer Layers - Improves Language Model Generalization —
- SEMA-GUARD: Semantic and Graph-Based Vulnerability Detection in Assembly Code —
- Towards Illusions Awareness in Cyber-Physical System's Design —
- GAUGE: A Formal Framework for Measuring Cryptographic Security under Heterogeneous Adversary Cost Models —
- Personalized Federated Learning through Global Knowledge Distillation and Local Head Adaptation —
- Same Flow, Different Paths: Variance Reduction in Flow Matching —
- Extracting ontology-compliant knowledge from scientific text describing irradiated materials using large language models —
- Zero-shot narrative detection in social messaging —
- Can We Stop The Ads? Taxonomy and Characterization of Smartphone Splash Ads and Existing Countermeasures —
- Towards Detecting AI-Assisted Responses in Online Surveys —
- Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation —
- From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts —
- Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs —
- Self-Emergence Agent Architecture:Behavior-Inertia HMM, Reflexive Metacognition,and Social-Contrastive Self-Modeling —
- Where Should a Document Live: Context, Representations, or Parameters? —
- RobResilience: Implementing and Evaluating a Resilience Framework for Cyber-Physical Embodied Systems —
- Hybrid Variational Quantum Circuits for Multivariate Regression and High-Dimensional Data Reconstruction —
- Same Words, Different Actions: Paired Turn-Taking Evaluation under Rewritten Dialogue Contexts —
- Large Language Models Develop Belief State Geometry In-Context —
- OPEN-1B: A Fully Auditable Training Run —
- Bridging the Confidence Gap: Temperature Scaling for Calibrating Test-Time Prompt Tuning —
- FlashVector: Agent for Hierarchical Model Serving Stack Optimization —
- Closing the Loop: Bidirectional Fully Encrypted Protocols —
- Enhancing Accessibility of Medical Texts through Large Language Model-Driven Plain Language Adaptation —
- SCHERI: Provably Secure Speculation Under the Constant-Time Policy for CHERI (Extended Version) —
- Never Stop Thinking: Continuous-Time Language Agents —
- Knowledge as Orbit: Finite Collections as Phases of an Exactly Periodic Latent Generator —
- World Model Science: Self-Organized Criticality, Weak Chaos, and Metastable Belief Dynamics in Long-Horizon LLM Agents —
- Transformer-Based Token Fusion and Dynamic Graph Planning for Audio-Visual Navigation —
- Talking Head Synthesis with Facial Landmark Guidance via 3D Gaussian Splatting —
- Learning-Guided Planning in Large Dynamic Action Spaces: Budgeted Tree Search for One-to-Many Mobile Charging —
- Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM —
- Reduced-Space Multi-Fidelity Bayesian Optimization of Process Simulation Models —
- Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback —
- JustFit: Just-in-Time State Management for Local LLM Serving —
- LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence —
- FreqSpaNet: Frequency and Spatial Learning of SFPF for Physical Layer Hardware Integrity Detection —
- Verifiable Social Reasoning for LLM Assistants —
- ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation —
- What Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity —
- When Should LLMs Abstain? Chain-of-Self-Questioning for Selective Risk Control —
- ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents —
- SemanticAdv: Generating Adversarial Examples via Attribute-conditional Image Editing —
- You Shall Not Pass into Ring-0! A User Privacy-Friendly Anti-Cheat Architecture for Personal Computers —
- Robust Recurrent Reinforcement Learning under Evolving Hidden Disturbances with Application to Rover Wheel Slip —
- The inherent goodness of well educated intelligence —
- An Initial Introduction to Cooperative Multi-Agent Reinforcement Learning —
- Unleashing Artificial Cognition: Integrating Multiple AI Systems —
- EVINCE: Optimizing Multi-LLM Dialogues Using Conditional Statistics and Information Theory —
- TrafficGamer: Reliable and Flexible Traffic Simulation for Safety-Critical Scenarios with Game-Theoretic Oracles —
- Fairness at Every Intersection: Uncovering and Mitigating Intersectional Biases in Multimodal Clinical Predictions —
- Attention is All You Need Until You Need Retention —
- In-context Learning vs. Instruction Tuning: The Case of Small and Multilingual Language Models —
- CBW: Towards Dataset Ownership Verification for Speaker Verification via Clustering-based Backdoor Watermarking —
- Explainable Graph-theoretical Machine Learning with Application to Alzheimer's Disease Prediction —
- BenSParX: A Robust Explainable Machine Learning Framework for Parkinson's Disease Detection from Bengali Conversational Speech —
- When majority rules, minority loses: bias amplification of gradient descent —
- R3: Robust Rubric-Agnostic Reward Models —
- GPUHammer: Rowhammer Attacks on GPU Memories are Practical —
- Post-Training Large Language Models via Reinforcement Learning from Self-Feedback —
- Observational Multiplicity —
- Script Fragmentation and Format: What Drives the English-Bengali Performance Gap in Open LLMs? —
- Neural Stochastic Differential Equations on Compact State Spaces: Theory, Methods, and Application to Suicide Risk Modeling —
Important terms
- Shared Selective Persistent Memory
- An architecture for AI agents that improves efficiency by stripping away messy reasoning traces and only storing essential information like task specifications and output constraints, rather than trying to remember every single detail from a session.
- Strict-launch Filter
- A simple, deterministic method used to train code-writing models. Instead of using a complex judge model that the AI might learn to cheat, this filter checks if code runs without errors in a headless engine.
- MASCOT
- A framework designed for multi-agent systems to prevent persona collapse. It uses bi-level optimization to ensure that individual agents maintain their unique identities and remain productive during complex social interactions.
- Monte Carlo Query Search
- An approach that treats evaluating an AI agent's capabilities as an active learning problem. It uses tree search to create specific, difficult queries to map out the mathematical boundaries of what an agent can do.
- GRAFT-ATHENA
- A framework that allows autonomous research agents to build cumulative scientific knowledge. By mapping problems to specific methods through a probabilistic structure, these agents can transfer experience from one task to another instead of starting over.