AI papers — 2026-09-02
Researchers are working to make AI systems more efficient, reliable, and secure as they move from experimental labs into the real world. One significant leap in model compression comes through a new framework called QTEA, which allows large language models to be quantized down to just 1.7 bits per weight without the massive accuracy drops usually seen below 2 bits.
By using ternary values and a salient weight system to compensate for errors, this method achieves much lower perplexity on datasets like WikiText and C4 compared to previous methods. It also runs seven times faster than standard FP16 baselines.
This push for efficiency extends into how we optimize models during actual use. Researchers have found that when tuning hyperparameters in production, using local search strategies is much more effective at minimizing performance loss than global strategies that jump around the entire search space.
While global methods like TabPFN are better at finding optimal configurations for complex tasks like Higgs, they are far more expensive when a guess goes wrong. Moving from optimization to reasoning, there is a new approach called Latent Recurrent Thoughts that lets us use frozen language models for much deeper thinking.
Instead of forcing the model to write out every step of its logic in text, which can lead to errors propagating through the conversation, this method uses a small auxiliary network to refine thoughts in continuous vector space through multiple iterative steps. This allows the model to perform complex symbolic and natural language reasoning at a fraction of the usual inference cost.
As these models become more autonomous and capable of interacting with others, we must consider the security of their transactions. A new formal analysis has mapped out 18 shared security principles for agent payment protocols, revealing that many current systems have undocumented vulnerabilities in how they handle delegated authority.
By testing these protocols through rigorous verification, researchers found that ensuring a user's intent remains consistent across every stage of a digital purchase is much harder and more critical than previously assumed. Solving physics within highly irregular and tortuous micro-scale geometries remains a massive hurdle for engineering simulations, but a new approach called GeoLAMP is making significant headway.
By using a dual-encoder architecture on graph representations, the model can translate complex physical spaces into compact latent representations that capture both global topology and fine-scale details. This allows for stable, block-wise autoregressive predictions of temporal dynamics through a flow matching framework, which successfully maintained low error rates across reactive flow and elasticity benchmarks.
This ability to model complex structural dynamics is mirrored in the financial sector by new ways to parse irregular patterns in energy markets. A novel convolutional neural network framework has been developed to detect event-driven price spikes, such as those caused by geopolitical conflicts or weather shifts.
Interestingly, researchers found that this deep learning structure is mathematically equivalent to several traditional statistical measures like maximum drawdown and slope change, though they caution these are classifications of patterns rather than true causal estimates. While these models parse market signals, other researchers are looking at how LLM agents handle the signals of error recovery.
A new protocol called invalidation contracts uses version stamps and cacheability hints to help agents decide when to reuse a previous fix for an API error and when to discard it as stale due to data drift. Using row-level invalidation, they were able to recover up to 33% of baseline token costs across various models, though they found that table-level invalidation actually performed worse by destroying useful co-located data.
The reliability of these agents is further complicated by the inherent biases present in the models themselves. Studies into conversational search engines for academic papers have uncovered a significant authority bias, where LLMs show a systematic preference for prestigious authors and venues even when the content remains identical to less-cited work.
This suggests that surface-level auditing might underestimate how much these models are actually swayed by metadata rather than substance. The most significant breakthrough today comes from a new way to shrink large language models without losing their intelligence.
Most current methods use a simplified mathematical shortcut that freezes the model's error landscape, which leads to information misalignment where quantization errors pile up as you move through the layers. A new approach called REAL-Q fixes this by using dynamic gradient descent to refine the model block by block, essentially correcting its mistakes in real-time.
When tested on LLaMA and Qwen models, this method cut down error levels by up to 49 percent compared to previous state-of-the-art techniques. While researchers work on making models smaller, others are trying to make them more reliable when they perform scientific tasks.
A new library of 163 procedural skills has been released to help AI agents move beyond just writing code and toward performing defensible scientific analysis in fields like genomics and medical imaging. These skills are essentially instruction files that tell an agent how to handle specific workflows, though researchers noted that loading all these references at once could easily overflow a standard context window.
This need for reliability extends into how we protect decentralized learning networks from bad actors. New research shows that where an attacker places their compromised nodes is just as important as how they behave, because certain placements allow malicious influence to spread much faster through the network.
By using a new metric called Byzantine Placement Influence, researchers can predict which node configurations will be most damaging to honest participants. It is becoming increasingly clear that our current ways of measuring how well AI agents actually work are fundamentally broken.
Researchers have identified a skill following paradox where models appear to get better at tasks when given external tools in aggregate tests, yet they actually perform worse on the specific tasks where those tools are actually used. This happens because standard metrics suffer from selection bias, whereas a new metric called the Retrieval-Invoked Actual-Use Effect shows that retrieval can often harm performance rather than help it.
This problem of misleading signals extends into how we train generative models, specifically regarding the loss of variety in diffusion generators. When reward post-training causes a model to collapse onto just a few favored modes, it essentially erases the diversity within a prompt.
A new method called ReNFT attempts to fix this by recalibrating probability mass from within the generator itself rather than using external signals. By using counterfactual proposals to probe for suppressed visual content, ReNFT can recover nearly all of the original model's diversity while retaining almost all of its newly acquired reward.
The difficulty of managing complex, multi-step reasoning is also being addressed through more granular training signals. Instead of just looking at whether a search agent reached the correct final answer, a new approach uses fact utility estimation to provide dense supervision.
By treating a reasoning process as an accumulation of discrete evidence facts and using Bayesian estimation to assign value to those facts, researchers can finally give agents credit for the specific steps that actually lead to success. This need for precision is equally vital when we move from software logic to the physical hardware running these models.
While block quantization is a popular way to make large language model inference more efficient, it often forces a frustrating trade-off between speed and accuracy. A new hierarchical approach called HBQ solves this by using large blocks for hardware efficiency combined with a second-level scaling mechanism.
This allows for much higher energy and area efficiency than previous methods without sacrificing the accuracy required for complex computations. If we want AI to be useful in high-stakes engineering, it has to do more than just guess the next word; it has to understand the actual physics of why things break.
A new framework tackles this by fine-tuning models like Llama and Mistral on thousands of expert-verified corrosion questions, essentially teaching them the why behind material degradation in magnesium alloys. By using a specialized retrieval system, they boosted accuracy significantly, but more importantly, they introduced a tool called Reason Map that can catch when an AI makes a logical leap or flips a causal relationship.
This need for logical rigor extends into the social realm of finance, where cultural cues can actually trick models into being confidently wrong. Researchers found that when presented with Islamic finance questions, large open-weight models almost always recognize the correct framework, yet they frequently fail to provide the correct answer within that framework.
It turns out a model might pick the right rulebook but still stumble on the actual math or logic required by that specific culture. Moving from social logic to physical signals, researchers have found a way to make AI much better at diagnosing respiratory sounds without needing massive amounts of labeled medical data.
By using a medical language model to create semantic anchors from text reports, they can align audio encoders with medical terminology in a shared space. This allows the system to recognize different lung pathologies in a zero-shot manner, performing better than general-purpose audio models by grounding sound in actual clinical meaning.
This drive toward more structured understanding is mirrored at the mathematical level when we look at how Transformers actually move data around during processing. One study investigates whether the complex shifts in data representations are just simple geometric stretches or something much more intricate.
While smaller models show a significant amount of remainder movement that does not fit simple linear maps, larger models appear to follow much more predictable, global transformations as they get deeper. If you are building complex systems where multiple AI agents interact with the same environment, you can no longer rely on fragmented safeguards to keep things safe.
OpenAgentFlow addresses this by establishing a shared enforcement interface at the action-commit boundary, essentially creating a governance layer that sits above the individual agents. In testing across various execution paths—including GUI and API interactions on Android—this architecture achieved a 95.35% attack-block rate and maintained high accuracy in preventing unsafe actions without needing to modify the underlying models or prompts.
This need for robust oversight in complex systems is mirrored by a more fundamental challenge in causal inference: knowing whether your models are actually working. Researchers have found that simply minimizing prediction error for nuisance functions is not a reliable proxy for the quality of a causal estimator, as the method with the lowest error does not always provide the most accurate confidence intervals.
This disconnect suggests that we need more sophisticated metrics than mere prediction accuracy to ensure our causal conclusions are truly valid. Moving from high-level system logic down to the fundamental mechanics of how networks process information, we see that even simple models are shaped by the structure of their inputs.
Researchers have shown that when inputs are positively correlated, linear recurrent neural networks face a specific cost for maintaining memory, which can actually cause them to overshoot and then prune away parts of that memory during training. This effect effectively turns these networks into change detectors, where they only bother to keep the past if the task requires it more than the correlation in the input stream already provides.
The most pressing challenge for long-term AI autonomy lies in how these systems evolve over time, which is why the development of Harness-of-Harness is so significant. By enabling multi-day autonomous software development through a framework of continual improvement, researchers have moved closer to agents that can actually manage complex, evolving codebases rather than just solving isolated tasks.
This drive toward more sophisticated machine learning logic extends into how we train these models, specifically through Iterative GRPO. This method uses single-turn reinforcement learning from human feedback to perform batch-online multi-turn reinforcement learning, essentially allowing a model to learn from its own iterative reasoning processes during training.
The mathematical foundations of these models are also being scrutinized to find better balance, particularly regarding performance-efficiency tradeoffs in Transformers. By applying an approximation theory perspective, researchers are mapping out exactly how much accuracy we sacrifice when we try to make these massive architectures more computationally efficient.
Moving into the realm of data integrity, there is a new way to model information blackouts when dealing with missing not-at-random time series data. This approach addresses the problem of gaps in datasets that aren't just random, but are actually caused by underlying patterns in the data itself.
In more specialized visual tasks, HarmoCore provides a way to perform sparse reconstruction of oscillatory wave fields using functional latent diffusion. This allows for much clearer reconstructions even when very little data is available.
Security and privacy also remain central, as seen in the development of MUGEN to generate unlearnable graph examples for multiple learning tasks. This technique creates data that is intentionally difficult for specific models to learn from, providing a new layer of defense against unauthorized training.
Finally, FedSPDnet introduces geometry-aware federated deep learning by utilizing SPDnet to improve how decentralized models learn from one another. This helps maintain the structural integrity of data as it moves across different nodes in a network.
Today's papers
- Convergence issues in Relational Concept Analysis based on AOC-posets This study investigates why the iterative process used to find relationships between groups of objects may fail to converge when using simplified data structures. [paper] [episode]
- Bandits in Prod: Hyperparameter Optimization at Inference Time The authors analyze how different strategies for exploring hyperparameter settings affect the cost and performance of AI systems during real-world deployment. [paper] [episode]
- QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization This framework compresses large language models into very small ternary values while using a sparse residual to maintain high accuracy and speed. [paper] [episode]
- Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs This method improves reasoning by having a small network iteratively refine continuous mathematical representations before a frozen language model decodes them. [paper] [episode]
- Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts This research shows that an AI's willingness to help with cybersecurity tasks changes significantly depending on the previous conversation history. [paper] [episode]
- Creative Generation via Multi-Agent Debate: Does Debate Suppress Diversity? This study explores whether having multiple AI agents debate improves the quality of creative writing at the expense of original or eccentric ideas. [paper] [episode]
- Beyond Magnitude: Contrastive Routing for Modular Mixture-of-Experts This paper introduces a new routing method for modular AI models that improves how they specialize in different linguistic tasks. [paper] [episode]
- A Formal Analysis of Agent Payment Protocols This work uses formal modeling to identify security gaps and consistency issues in protocols that allow AI agents to make payments. [paper] [episode]
- Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains This model uses graph-based encoders to solve complex physics equations within highly irregular and winding geometric structures. [paper] [episode]
- A convolutional framework for detecting event-driven dynamics in energy price series This paper presents a neural network approach that mathematically connects deep learning to traditional statistical methods for identifying sudden changes in energy prices. [paper] [episode]
- Invalidation Contracts for Cross-Episode Agent Memory This protocol helps AI agents efficiently manage cached solutions by using version stamps to tell them when a saved fix is no longer valid. [paper] [episode]
- Authority Bias in Conversational Search Engines for Academic Paper Recommendation This study finds that AI search engines tend to recommend papers based on author prestige and citation counts rather than just the actual content. [paper] [episode]
- RestoreBench: Can AI Agents Restore Power Flow Convergence? This benchmark evaluates how well AI agents can diagnose and fix technical problems in electrical power grid simulations. [paper] [episode]
- LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts This research demonstrates that prompting an AI to act like a specific demographic group can unintentionally bias its judgments against minority perspectives. [paper] [episode]
- Different representation learning objectives recover distinct latent structures from the same psychometric data This study shows that the way we train AI to represent psychological data determines whether it captures personality traits or how people match with teachers. [paper] [episode]
- Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs This framework uses simulated factual environments to help language models learn new information and know when to admit they are uncertain. [paper] [episode]
- REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent This new quantization method reduces errors in compressed language models by using fine-grained updates that align with the model's overall performance. [paper] [episode]
- Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents This project provides a standardized library of professional research procedures to help AI agents perform scientifically defensible analysis. [paper] [episode]
- Optimizing Byzantine Node Placement in Decentralized Federated Learning This paper studies how malicious actors can strategically place themselves in a decentralized network to cause the most damage during machine learning. [paper] [episode]
- Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing This study uses brain signals to show that while unrestricted chatbots help people pass tests faster, they may actually reduce deep cognitive engagement. [paper] [episode]
- Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning This framework helps remove private information from AI models by first identifying what the model has actually memorized rather than relying on external data. [paper] [episode]
- Local Reference Geometry Residual Augmentation for Imbalanced Time Series Classification This method improves AI performance on rare data patterns by augmenting the feature space around minority examples to make them more reliable. [paper] [episode]
- Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time This approach improves AI safety by only letting a supervisor intervene when it is both confident and the topic is important. [paper] [episode]
- Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition? This paper investigates whether AI rankings are truly accurate or just artifacts of how specific questions are structured in benchmarks. [paper] [episode]
- PCoMoE: Shifting MoE Inference from Monolithic Expert Selection to Fine-Grained Path Composition This framework makes modular AI models more efficient by treating expert selection as a sequence of specialized steps rather than a single choice. [paper] [episode]
- A Version Space Approach for Digital Circuit Analysis This paper uses a mathematical counting method to determine how many possible circuit configurations remain consistent with observed data. [paper] [episode]
- Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations This system provides managers with a realistic voice-based environment to rehearse sensitive employee conversations and receive feedback. [paper] [episode]
- Stochastic complexity of vectors containing cluster structure This study introduces an efficient way to calculate the mathematical complexity of data that contains grouped clusters. [paper] [episode]
- Life Operators: a self-evolving framework for multiscale life modelling This framework proposes a modular way to build biological models that can evolve and be revised as new medical evidence emerges. [paper] [episode]
- Transferable End-to-End Optimization for Indirect Long-Term Memory Poisoning in LLM Agents This research shows how attackers can optimize malicious content to survive the entire process of being written into and retrieved from an AI's memory. [paper] [episode]
- Group Adaptive Clipping Policy Optimization This method improves reinforcement learning by adjusting how much the model ignores certain data based on how difficult a task is. [paper] [episode]
- From Base Rollouts to RL Reasoning: A Budgeted Search Perspective This study suggests that much of the reasoning improvement seen in reinforced language models comes from better sampling efficiency rather than entirely new knowledge. [paper] [episode]
- Breaking the Structural Identity: Personalized Federated LoRA Fine-tuning under Rank Heterogeneity This framework allows multiple users to collaboratively train an AI model while still maintaining their own unique, personalized settings. [paper] [episode]
- ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration This method fixes AI image generators that have become repetitive by redistributing their internal focus to include more diverse visual styles. [paper] [episode]
- HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference This hardware-friendly quantization method uses a hierarchical structure to make large language models much faster and more energy-efficient. [paper] [episode]
- Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents This research reveals that many AI agents actually perform worse on specific tasks when they attempt to use a retrieved tool or skill. [paper] [episode]
- Dense Process Supervision for Search Agents via Fact Utility Estimation This method improves AI search by rewarding the model for finding useful individual facts rather than just looking at the final answer. [paper] [episode]
- Foundation models for electricity price forecasting and battery arbitrage: Can they replace market-specific forecasting models? This study finds that while some general AI models are great at predicting prices, they cannot always replace specialized tools used for making money in energy markets. [paper] [episode]
- Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy This research shows that multimodal AIs often ignore what they see in an image if the accompanying text tells them something different. [paper] [episode]
- JENGA: Exploiting Counter-Based RowHammer Countermeasures to Break Real-Time Predictability This attack exploits security features in memory hardware to cause unpredictable delays in safety-critical real-time computer systems. [paper] [episode]
- Generative artificial intelligence for reliable mechanistic reasoning for corrosion This framework uses specialized AI to help engineers understand the actual chemical processes behind metal corrosion rather than just predicting it. [paper] [episode]
- Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close This study shows that cultural hints can trick AI into choosing a specific financial framework but fail to make it answer correctly within that framework. [paper] [episode]
- Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment This method uses medical language models to help audio encoders recognize lung sounds without needing specific training data for every disease. [paper] [episode]
- Measuring Optimal Transport in Transformer Depth This paper investigates whether the way data moves through a transformer model can be explained by simple geometric transformations. [paper] [episode]
- Automated Event Log Generation from Unstructured Text Using Finetuned LLMs This framework uses fine-tuned language models to turn messy text documents into structured logs for business process analysis. [paper] [episode]
- RW-LoRA: Communication-Efficient Decentralized LoRA Fine-Tuning via Random Walks This method allows multiple devices to collaboratively train an AI model by having a single token move through the network instead of syncing entire models. [paper] [episode]
- Recent Developments in Transformer Inference Deployment on FPGA Platforms: A Survey This paper reviews the latest techniques for running transformer models efficiently on specialized hardware like FPGAs. [paper] [episode]
- GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments This benchmark tests whether AI-generated user interfaces remain consistent and usable when an agent interacts with them over multiple steps. [paper] [episode]
- How Temporal Correlations Shape Memory in Linear Recurrent Neural Networks This study explains how the patterns in input data determine whether a recurrent neural network learns to remember the past or simply reacts to the present. [paper] [episode]
- QILP-0: Constructing Observational Declarative Twins of Quantum Circuits This framework builds logical programs that can perfectly reconstruct and describe the behavior of quantum circuits based on observations. [paper] [episode]
- When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation This research warns that simply minimizing prediction error is not a reliable way to ensure an AI is making accurate causal inferences. [paper] [episode]
- GCA-BULF: A Bottom-Up Framework for Short-Term Load Forecasting Using Grouped Critical Appliances [No summary provided in prompt] [paper]
- Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning [No summary provided in prompt] [paper]
- Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches [No summary provided in prompt] [paper]
- On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous Levels [No summary provided in prompt] [paper]
- OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets This architecture creates a centralized control layer to enforce safety rules across different types of AI agents and tools. [paper]
- The Constitutional Coverage Trilemma in AI Governance [No summary provided in prompt] [paper]
- Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement [No summary provided in prompt] [paper]
- MUGEN: Generating Unlearnable Graph Examples for Multiple Learning Tasks [No summary provided in prompt] [paper]
- HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields [No summary provided in prompt] [paper]
The papers
- Convergence issues in Relational Concept Analysis based on AOC-posets — I am unable to extract the summary because you have only provided a figure caption and associated symbolic data, not the full text of the scientific paper "Convergence issues in Relational Concept Analysis based on AOC-posets." To fulfill your request—which requires synthesizin [episode]
- Bandits in Prod: Hyperparameter Optimization at Inference Time — The paper explores advanced methods for hyperparameter optimization, particularly focusing on how these techniques function when deployed in a production or inference environment. [episode]
- Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts — The paper investigates how varying conversational contexts—specifically, whether a request is presented directly or within a modified history—impact the cybersecurity assistance provided by large language models (LLMs). [episode]
- Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs — The paper introduces a novel framework for enhancing Large Language Model (LLM) reasoning by implementing a latent recurrent refinement process, termed "Latent Recurrent Thoughts" (LRT). [episode]
- QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization — The paper introduces QTEA, a novel framework for developing highly efficient Large Language Models using ternary representation combined with sparse residual salient weights and by-column optimization. [episode]
- Creative Generation via Multi-Agent Debate: Does Debate Suppress Diversity? — The paper "Creative Generation via Multi-Agent Debate: Does Debate Suppress Diversity?" presents a rigorous investigation into utilizing multi-agent debate as a novel mechanism for enhancing creative output in large language models. [episode]
- Beyond Magnitude: Contrastive Routing for Modular Mixture-of-Experts — The paper introduces "Contrastive Routing" (CoRM) as an advancement for Mixture-of-Experts (MoE) architectures, proposing a method that moves "beyond magnitude" by enhancing how computational specialization is analyzed and utilized. [episode]
- A Formal Analysis of Agent Payment Protocols — The paper provides a rigorous, formal analysis of vulnerabilities inherent in modern agent payment protocols and decentralized settlement systems. [episode]
- Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains — The paper outlines a comprehensive suite of advanced deep learning architectures designed for solving partial differential equations (PDEs) within complex and irregular physical domains. [episode]
- Authority Bias in Conversational Search Engines for Academic Paper Recommendation — This paper investigates the susceptibility of large language models (LLMs) used in conversational search engines when recommending academic papers, specifically detailing how inherent "authority bias" influences selection patterns. [episode]
- RestoreBench: Can AI Agents Restore Power Flow Convergence? — The paper, "RestoreBench: Can AI Agents Restore Power Flow Convergence?", presents a comprehensive benchmark evaluation designed to assess the capability of large language model (LLM) agents to autonomously restore power flow convergence in complex electrical grids. [episode]
- Invalidation Contracts for Cross-Episode Agent Memory — This paper introduces "invalidation contracts," a protocol layer designed to manage cross-episode memory for LLM agents interacting with APIs. [episode]
- REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent — REAL-Q presents a novel methodology for end-to-end Large Language Model (LLM) quantization by integrating dynamic gradient descent into the Post-Training Quantization (PTQ) pipeline. [episode]
- LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts — The paper investigates how conditioning on sociodemographic profiles—specifically gender and race—affects Large Language Model (LLM) judgments across various tasks, such as Offensiveness, Politeness, and Intimacy. [episode]
- Different representation learning objectives recover distinct latent structures from the same psychometric data — The paper investigates how different representation learning objectives can recover distinct latent structures from the same psychometric data. [episode]
- A convolutional framework for detecting event-driven dynamics in energy price series — The paper introduces a novel common Convolutional Neural Network (CNN) framework designed for the detection of heterogeneous event-driven dynamics within time series windows, specifically applied to energy price series. [episode]
- Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs — The paper details advanced methodologies for enhancing Large Language Model knowledge updating and performance by integrating structured factual knowledge using a framework called S YNAPSE. [episode]
- Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning — The paper details a sophisticated framework for mitigating "Forget-Set Misalignment" during Large Language Model (LLM) unlearning, addressing the critical issue where models may retain knowledge that should have been erased. [episode]
- Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents — This paper introduces "Scientific Agent Skills," an open library of 163 procedural knowledge modules designed to ensure that language-model agents perform "defensible analysis" rather than merely returning "working code." By providing field-specific conventions, the library addre [episode]
- Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing — As a diligent researcher whose work requires absolute precision, I am prepared to execute this extraction with the highest level of rigor, ensuring every quote and structural element adheres exactly to your specifications. [episode]
- Optimizing Byzantine Node Placement in Decentralized Federated Learning — I apologize, but the text provided is a bibliography and reference list, not the full content of the scientific paper titled "Optimizing Byzantine Node Placement in Decentralized Federated Learning." To generate an accurate summary of 450–600 words, I require the actual body te [episode]
- Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time — The paper "Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time" addresses the critical issue of over-reliance on generative guidance systems in large language models (LLMs). [episode]
- Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition? — The paper details a comprehensive methodological framework designed to rigorously test the robustness of LLM performance rankings when the underlying benchmark structure is intentionally reconfigured based on differential item difficulty (DIF) and source attribution. [episode]
- Local Reference Geometry Residual Augmentation for Imbalanced Time Series Classification — The paper, "Local Reference Geometry Residual Augmentation for Imbalanced Time Series Classification," addresses the critical challenge of accurately classifying time series data when the underlying class distribution is severely skewed. [episode]
- PCoMoE: Shifting MoE Inference from Monolithic Expert Selection to Fine-Grained Path Composition — The paper "PCoMoE: Shifting MoE Inference from Monolithic Expert Selection to Fine-Grained Path Composition" addresses a critical bottleneck in modern Mixture-of-Experts (MoE) model deployment: the inherent inefficiency of treating expert utilization as a single, monolithic selec [episode]
- Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations — The paper "Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations" introduces a novel computational linguistics framework designed to simulate high-stakes professional interactions. [episode]
- Stochastic complexity of vectors containing cluster structure — The paper addresses the critical computational challenge of determining the Normalized Maximum Likelihood (NML) code length for vectors that possess a specific cluster structure, which is vital within the Minimum Description Length (MDL) clustering framework. [episode]
- A Version Space Approach for Digital Circuit Analysis — I apologize, but you have provided a detailed set of instructions and an academic context, but you have not included the actual text of the arXiv paper titled "A Version Space Approach for Digital Circuit Analysis." To perform the summary with the required rigor—adhering to len [episode]
- Transferable End-to-End Optimization for Indirect Long-Term Memory Poisoning in LLM Agents — The paper investigates advanced vulnerabilities in Large Language Model (LLM) agents that utilize long-term memory, detailing methods for indirect memory poisoning. [episode]
- Life Operators: a self-evolving framework for multiscale life modelling — I am unable to extract and summarize the paper because you have provided a bibliography (citations [9] through [27]) but have not provided the actual text or content of the arXiv paper titled "Life Operators: a self-evolving framework for multiscale life modelling." Please provid [episode]
- HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference — This paper introduces Hierarchical Block Quantization (HBQ), a novel quantization scheme designed to address the fundamental "accuracy–efficiency trade-off" in Large Language Model (LLM) inference. [episode]
- Breaking the Structural Identity: Personalized Federated LoRA Fine-tuning under Rank Heterogeneity — The presented research addresses the critical challenge of personalized federated fine-tuning for large language models, specifically focusing on adapting LoRA techniques in environments characterized by structural identity issues and rank heterogeneity. [episode]
- ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration — This paper introduces ReNFT, a novel methodology designed to mitigate mode collapse and recover structural diversity in generative image models that have undergone reward-based post-training. [episode]
- Group Adaptive Clipping Policy Optimization — The paper introduces Group Adaptive Policy Optimization (GAPO), a novel optimization framework designed to enhance the robustness and capability of large language models in solving complex, multi-step reasoning problems. [episode]
- From Base Rollouts to RL Reasoning: A Budgeted Search Perspective — This paper introduces a novel framework for analyzing model reasoning capabilities by framing problem-solving as a "Budgeted Search Perspective." It investigates how models transition from relying on raw base rollouts to incorporating Reinforcement Learning (RL) techniques. [episode]
- Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents — The paper, "Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents," addresses a critical gap in current AI evaluation by moving beyond mere task correctness to assess whether Large Language Model (LLM) agents genuinely utilize retrieved knowledge. [episode]
- Dense Process Supervision for Search Agents via Fact Utility Estimation — This paper introduces a novel framework for enhancing search agents by implementing "Dense Process Supervision for Search Agents via Fact Utility Estimation." The work addresses the critical need to guide complex reasoning agents—which rely on iterative searches and fact assert [episode]
- Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy — The rapid integration of vision and language capabilities into large multimodal models (LMMs) has opened new frontiers in AI understanding; however, this advancement introduces critical reliability concerns regarding how these models reconcile conflicting sensory inputs. [episode]
- Foundation models for electricity price forecasting and battery arbitrage: Can they replace market-specific forecasting models? — The paper investigates whether large, general-purpose Foundation Models can successfully replace or significantly augment specialized, market-specific forecasting models when predicting electricity prices and optimizing battery arbitrage strategies. [episode]
- Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close — The paper "Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close" presents a rigorous investigation into how cultural and domain-specific cues can reveal latent knowledge gaps within large language models (LLMs). [episode]
- JENGA: Exploiting Counter-Based RowHammer Countermeasures to Break Real-Time Predictability — I apologize, but the actual content of the arXiv paper titled "JENGA: Exploiting Counter-Based RowHammer Countermeasures to Break Real-Time Predictability" was not provided. [episode]
- Generative artificial intelligence for reliable mechanistic reasoning for corrosion — I am ready to perform this extraction with extreme diligence. [episode]
- Measuring Optimal Transport in Transformer Depth — The paper examines whether the complex, token-by-token rearrangement of data representations within deep transformer architectures can be simplified or explained by basic geometric transformations, specifically contrasting network behavior against established measures derived fro [episode]
- Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment — As a diligent researcher, I have reviewed the context provided. [episode]
- RW-LoRA: Communication-Efficient Decentralized LoRA Fine-Tuning via Random Walks — The paper presents rigorous mathematical proofs concerning the convergence and stability of expectations within a stochastic process framework. [episode]
- Automated Event Log Generation from Unstructured Text Using Finetuned LLMs — The paper addresses the critical challenge of transforming vast quantities of unstructured textual data—such as customer service transcripts, medical notes, or incident reports—into structured, machine-readable event logs suitable for process mining analysis. [episode]
- Recent Developments in Transformer Inference Deployment on FPGA Platforms: A Survey — The field of deep learning has seen exponential growth in model complexity, particularly with the dominance of Transformer architectures. As these models are deployed from cloud environments to resource-constrained edge devices, efficient inference becomes a critical bottleneck. [episode]
- GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments — This paper introduces GUI-CC, a benchmark designed to evaluate the "contextual consistency of GUI world models as agent environments" rather than treating them as "isolated next-screen predictors." As autonomous GUI agents require interactive execution to complete tasks, world mo [episode]
- QILP-0: Constructing Observational Declarative Twins of Quantum Circuits — As a diligent researcher where precision is paramount, I require the full text of the arXiv paper, "QILP-0: Constructing Observational Declarative Twins of Quantum Circuits," to proceed with this extraction. [episode]
- How Temporal Correlations Shape Memory in Linear Recurrent Neural Networks — I apologize, but you have provided a comprehensive bibliography of scientific literature, but you have not provided the actual text content of the arXiv paper titled "How Temporal Correlations Shape Memory in Linear Recurrent Neural Networks." To fulfill your request—which requ [episode]
- When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation — Simulation studies are necessary to rigorously evaluate whether standard measures of machine learning performance, specifically nuisance-function prediction error, are sufficient for assessing the quality of causal effect estimation. [episode]
- Symbolic Neural Generation with Applications to Lead Discovery in Drug Design —
- M4FC: a Multimodal, Multilingual, Multicultural, Multitask Real-World Fact-Checking Dataset —
- Can machines think efficiently? —
- Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+ —
- Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement —
- Apples on the Table? Evaluating Text-Guided 3D Scene Synthesis via Fine-Grained Constraint Verification —
- KGFR: A Foundation Retriever for Generalized Knowledge Graph Question Answering —
- SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning —
- Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for Incident Response —
- Iterative GRPO: Batch-Online Multi-Turn RL via Single-Turn RLHF —
- Freeze, Diffuse, Decode: Task-Aware Adaptation of Transformer Embeddings for Antimicrobial Peptide Design —
- Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs —
- Multilingual Medical Reasoning for Question Answering with Large Language Models —
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning —
- Control Variate Score Matching for Diffusion Models —
- Modeling Information Blackouts in Missing Not-At-Random Time Series Data —
- Hidden State Poisoning Attacks against Mamba-based Language Models —
- A Hybrid Insider Threat Detection Framework Combining Multi-Agent Simulation, Layered SIEM Correlation, and Theory-of-Mind Reasoning —
- Beyond Static Summarization: Proactive Memory Extraction for LLM Agents —
- DAGGER: Distractor-Aware Graph Generation for Executable Reasoning in Math Problems —
- LifeAgentBench: Benchmarking LLMs for Long-Horizon, Cross-Dimensional Lifestyle Health Reasoning —
- NewsRECON: News Article Retrieval for Image Contextualization —
- Variance-Adaptive Muon: Pre-Orthogonalization Variance Modulation for Efficient Language Model Pretraining —
- FloydNet: A Learning Paradigm for Global Relational Reasoning —
- Breaking the Reasoning Horizon in Entity Alignment Foundation Models —
- Think Like a Doctor: Conversational Diagnosis through the Exploration of Diagnostic Knowledge Graphs —
- MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems —
- Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States —
- Persistent Entropy as a Detector of Phase Transitions —
- Is Knowledge Distillation Actually Greener? A Case Study in Machine Translation —
- Structural Bias Beyond Homophily: A Study of Fairness in Link Prediction —
- PathCRF: Ball-Free Soccer Event Detection via Possession Path Inference from Player Trajectories —
- GRRM: Group Relative Reward Modeling for Machine Translation —
- Ontology-Guided Neuro-Symbolic Inference: Grounding Language Models with Mathematical Domain Knowledge —
- Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning —
- Suffix-Constrained Greedy Search Algorithms for Causal Language Models —
- Inverse Reconstruction of Shock Time Series from Shock Response Spectrum Curves using Machine Learning —
- MMAI Gym for Science: Training Liquid Foundation Models for Drug Discovery —
- HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation —
- CoMMET: A Psychologically Grounded Benchmark for Evaluating Theory of Mind in Multimodal LLMs —
- RetroReasoner: A Reasoning LLM for Strategic Retrosynthesis Prediction —
- Is Human Annotation Necessary? Iterative MBR Distillation for Error Span Detection in Machine Translation —
- SpokenUS: A Spoken User Simulator for Task-Oriented Dialogue —
- SCALE:Scalable Conditional Atlas-Level Endpoint transport for virtual cell perturbation prediction —
- MineDraft: A Framework for Batch Parallel Speculative Decoding —
- CWoMP: Morpheme Representation Learning for Interlinear Glossing —
- Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method —
- Rigorous Error Certification for Neural PDE Solvers: From Empirical Residuals to Solution Guarantees —
- Process-Aware AI for Rainfall-Runoff Modeling: A Mass-Conserving Neural Framework with Hydrological Process Constraints —
- Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs —
- APEX-EM: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay —
- Oblivion: Self-Adaptive Agentic Memory Control through Decay-Driven Activation —
- Scheduling LLM Inference with Uncertainty-Aware Output Length Predictions —
- TRU: Targeted Reverse Update for Efficient Multimodal Recommendation Unlearning —
- KV Cache Offloading for Context-Intensive Tasks —
- What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal —
- DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering —
- Why Fine-Tuning Encourages Hallucinations and How to Fix It —
- Global Attention with Linear Complexity for Exascale Generative Data Assimilation in Earth System Prediction —
- The Topological Trouble With Transformers —
- FedSPDnet: Geometry-Aware Federated Deep Learning with SPDnet —
- Reparameterization through Coverings and Topological Weight Priors —
- GCA-BULF: A Bottom-Up Framework for Short-Term Load Forecasting Using Grouped Critical Appliances —
- Causal Probing for Internal Visual Representations in Multimodal Large Language Models —
- Combating Organized Platform Abuse: Amplifying Weak Risk Signals with Structural Information —
- The Scientific Contribution Graph: Automated Literature-based Technological Roadmapping at Scale —
- Universal Approximation of Nonlinear Operators and Their Derivatives —
- Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road —
- Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs —
- PromptNCE: Conditional Probabilities and PMI Using Only LLMs and Contrastive Estimation Prompts —
- LLM-driven design of physics-constrained constitutive models: two agents are better than one —
- DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking —
- Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior —
- UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities —
- Argument Quality Assessment with Large Language Models: A Pairwise Bradley-Terry Approach —
- Diffusion Large Language Models for Visual Speech Recognition —
- The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic —
- Can Large Language Models Handle Discourse Particles? A Case Study of Colloquial Malay —
- LLMBridge: An LLM Pipeline for End-to-end Referential Bridging Resolution in English —
- MusTBench: Benchmarking and Advancing Temporal Grounding in Music LLMs —
- PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning —
- Skill Reuse as Compression in Agentic RL —
- FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search —
- UME: A Unified Meta-Generalization Framework for Cross-Domain ETA —
- PlanarBench: Evaluating LLM Spatial Reasoning via Planar Graph Drawing —
- Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025 —
- What Do Students Learn? A Feature-Level Analysis of Dark Knowledge —
- RECAP: Regression Evaluation for Continual Adaptation of Prompts —
- Enabling KV Caching of Shared Prefix for Diffusion Language Models —
- DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment —
- DECSELFMASK: Leveraging Unlabeled Text via Self-Relevance-Guided Masking for Decoder-Only Classification —
- AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages —
- HiMPO: Hindsight-Informed Memory Policy Optimization for Less-Entangled Credit in Long-Horizon Agents —
- Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement —
- Learning to Refine Hidden States for Reliable LLM Reasoning —
- ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues —
- Beyond AHI: An Interpretable Causal-Discovery-Guided Framework for Sleep Recovery in Connected Health —
- When Compression Helps and When It Hurts: Condition-Aware Analysis of Chain-of-Thought Distillation —
- Energy-Based Transformers as Predictors of Reading Difficulty —
- Can LLMs Reliably Self-Report Adversarial Prefills, and How? —
- ComputeFHE: A Privacy-Preserving General-Purpose Computation Library —
- Self-Evolving World Models for LLM Agent Planning —
- Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas? —
- DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching —
- ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs —
- Antaeus: Hunting Repository-Level Logic Vulnerabilities via Context-Grounded LLM Reasoning —
- Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models —
- Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories —
- Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation —
- SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction —
- On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous Levels —
- Scope3Trace: Evidence-Based Identification and Extraction of Scope 3 GHG Emissions from Sustainability Reports —
- Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI —
- How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs? —
- A Classifier That Teaches Itself: Self-Improving, Frozen-gate Training (SIFT) for Dynamic Document Classification —
- Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges —
- AREX: Towards a Recursively Self-Improving Agent for Deep Research —
- Shallower ReLU Network Representations via Exact Linear Algebra —
- UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks —
- S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring —
- Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility —
- MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration —
- "I Thought You Were The Uncensored Place": Norms, Rules, and Moderation in AI-Generated Sexual Content Communities —
- Non-Parametric Spatiotemporal Trajectory Prediction via State-Conditioned Transition Sampling —
- A survey of AI-generated voices and their detection —
- Reasoning-supported Robustness Validation of Automotive E/E Components —
- CUSTOS: Toward Forensic-Ready Zero Trust at the Capture-Containment Boundary —
- Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It —
- The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations —
- 0xPass: A Secure Protocol for Universal Cross-Chain Accounts —
- Verifiable abstention makes AI leak diagnosis accountable in urban water distribution networks —
- When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models —
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? —
- FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models —
- Electronic Navigational Chart Change Classification —
- HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning —
- ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents —
- Stress Testing Unlearning Algorithms —
- HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Models —
- I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models —
- Discrete-Time MDP Modeling for Multi-Item Capacitated Lot Sizing with Stochastic Demand Timing —
- Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models —
- Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls —
- Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning —
- OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets —
- SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces —
- UI-Venus-2 Technical Report —
- EULER: Exploring Underused Links with Evidence-Checked Return for Multi-Agent Mathematical Discovery —
- trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories —
- Task-Specific Prompt with Global Context for Multi-Task Graph Pre-Training —
- From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling —
- AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes —
- ValueGraph: Value-Signal Guided Graph Pre-training for Contextualized User Representation —
- CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language —
- DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction —
- RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving — Data contamination poses a significant challenge to the reliable evaluation of large language models (LLMs) when assessing mathematical problem-solving capabilities, as training corpora and evaluation benchmarks often share public sources.
- Medical Causal Hypothesis Verification with Large Language Models —
- Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning —
- OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization —
- Auditing Harness Tampering in Self-Improving Agents —
- MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts —
- AI Morbidity and Mortality: A Framework for Clinical AI Failure Review —
- Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops —
- KItCAT: Knowledge Injection via Input Corruption for Auto-regressive Training —
- Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation —
- Assessing Alignment and Stability of Feature Importance Explanations via Weight of Evidence —
- Safin-1: Safety from Within through Memory-Native State Evolution —
- Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding —
- Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy —
- Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations —
- Flawed in Nature, Perfect through Evolution —
- Recursive Criticality of AI Self-Improvement —
- Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs —
- IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training —
- Explainable Artificial Intelligence for Industrial Cybersecurity: A Review of Methods, Operational Integration, and Research Challenges —
- Do General NLP Embeddings Capture Ontological Reasoning? —
- Asymmetries in Spontaneous and Instructed Deception —
- Elite-Weighted Supervised Fine-tuning for Goal-Directed Molecular Optimization —
- Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models —
- LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark —
- Provably Efficient Federated Reinforcement Learning with Linear Function Approximation and Logarithmic Communication Cost —
- ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide Generation —
- WHALE: A Simple Recipe for Joint Harness-Weight Optimization —
- AI Should Not Only Be Helpful. It Should Be Contingent. Artificial Intimacy, Sycophancy, and the Future of Social Learning —
- Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning —
- ConvDeck: Conversational Paper-to-Slide Generation via Stage-Specific User Feedback —
- Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking —
- Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems —
- LOOMSUM:Weaving Quantitative and Narrative Evidence for Faithful Long Text-Table Summarization —
- Hypotheses-Guided Self Distillation for Continual Personalization —
- NSIDDx: A Design Framework for Neuro-Symbolic, Practitioner-First Differential Diagnosis in Low-Resource Settings —
- DUPIN: Attack Learning Is Still Needed! Demonstrating Few-Shot after Unsupervised Pretraining Is A Nimble Forensics Learner —
- The Answer Is Not the Argument —
- Delegation Without Trust: An Empirical Gap Analysis of Identity, Authorization, and Runtime Governance in Multi-Agent LLM Systems —
- Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching —
- The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systems —
- Exact Global MCMC with Denoising Diffusion —
- Slow to See, Slow to Suppress: Understanding the Effects of Modality in Context-Memory Conflicts —
- Toward Workflow-Aware Benchmarking for Healthcare NLP Agents —
- The Assistant's Ideal Self —
- Workload Identification with Physical Side Channels for AI Governance —
- Emotional Labor Strategy Preferences in LLM Personas —
- Latent Mechanisms of Language Control in Multilingual Language Models —
- The Curse of Multilinguality in Lexical Normalization —
- Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts —
- Human-AI Co-Interpretation for Responsible AI: A Hermeneutic Perspective —
- Two tests of phase-structure features for transition prediction —
- OreProof: Verifiable Provenance with Limited Disclosure for Critical-Minerals Supply Chains Using Zero-Knowledge Proofs —
- SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning —
- From Tool Use to Technological Agency: LoopCAT as a Local-First, Open-Source Tool for Translation Technology Education —
- Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment —
- Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning —
- Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models —
- A Stable Aggregation Method for Quantum Federated Learning —
- Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance —
- Dr. Claw: An AI Scientist Workspace for Vibe Research —
- Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure —
- Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning —
- Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You —
- Removable and Irreducible: A Token-Cost Ledger for the Multilingual Tokenization Tax —
- Neural means and kernel corrections for operator learning —
- NeuroPriv: Adversarial Representation Learning for Privacy in Wearable EEG Systems —
- A Multi-Branch Feature Fusion Approach for Health Misinformation Detection and Propagation —
- Federated Trust for Embodied Robot Capability Marketplaces —
- Dependency-Aware Chain-of-Thought Compression for Financial Reasoning —
- Late Transformer Layers Recode Syntax Canonically: Evidence from Greek Scrambling and Cross-Layer Generalisation —
- SpecMind: Enabling Spectrum Intelligence via Multi-Agent Hybrid Retrieval-Augmented Generation —
- Don't Trust the Code, Check Its Effects: Runtime Refinement for Regenerated Systems Code Under an Adversarial Generator —
- SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents —
- (V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement —
- Capability-Gated Language Models: Security Composes, Utility Does Not —
- CRAD: Class-wise Reliability-Aware Distillation for Decentralized Heterogeneous Federated Learning —
- mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers —
- Location-Aware Language Models via Secondary Embeddings —
- Towards a Belief-Based World Model for LLM Agents —
- Can LLMs Use Relational Transformer Embeddings? —
- Context Window Failures in Relational Foundation Models —
- Toppling the Hierarchy in Byte-level Language Modeling —
- Does Reasoning Mitigate Backdoor Attacks? A Neuro-Symbolic Perspective —
- TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge Poisoning —
- Higher Structures in Deep Learning —
- Exploring Collaboration between a language and a non-language agent —
- EGT-KG: Evidence-Grounded Typed KG Retrieval for Practical Scientific QA with Small Language Models —
- EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities —
- AdaptNTK: Adaptive Uncertainty Quantification and Active Learning for Neural Network Potentials —
- A hybrid quantum-classical neural network for learning to route —
- MemeBridge: A Dataset for Benchmarking and Mitigating the Bidirectional Cultural Gap in Meme Interpretation —
- The Privacy-Hallucination Tradeoff in Differentially Private Language Models —
- Human-Anchored Factuality Evaluation with Strategic Annotation —
- Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models —
- Validity-Aware Jailbreak Evaluation for Large Language Models —
- GlitchLab: A Hardware-in-the-Loop Optimizer for Physical Fault Injection —
- Wave Function Backpropagation with Explicit Temporal-Interval Dynamics —
- VATO: A Vortex-Force-Aware Transformer Operator for Unsteady Separated Aerofoil Flows —
- CoVer: Conflict-Aware Claim Verification —
- When the Algorithm Becomes the Brand Crisis: A Sociotechnical Theory of Distributed Responsibility and Accountable Transparency —
- ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation —
- The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space —
- Learning Task-Specific Antibody Representations via Function-Aware Masking —
- The Safeguard Worked. Is the LLM System Safer? —
- Why Multi-Layer Message Passing Works: Completeness Theory for Graph Neural Network Interatomic Potentials —
- DeSyR: A Decoupled Symbolic Recovery Framework with PINN-Guided Structure Search and Physics-Informed Coefficient Refinement —
- Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented Generation —
- GenONet: A Generative operator Network for High-Resolution Precipitation Nowcasting —
- Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict —
- EM 2Mem: Event-Centric Multimodal Memory for Large Language Models —
- Manifold-Aware General Coded Computing for Straggler-Resilient Distributed Computing —
- EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection —
- VoiceLongMemEval: Do Assistants Remember How You Sounded? —
- Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs —
- Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random —
- GeoPAR: Large-Scale Multi-Agent Combinatorial Optimization with Geometry-Guided Parallel Autoregressive Learning —
- Enoki: Efficient Multi-Level Hallucination Detection —
- Quit While You're Ahead: Quit for Efficient Candidate Generation in Machine Translation Reranking —
- CRAFT: Fine-Tuning Pre-hoc Explainability in AI-native 6G RAN —
- SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems —
- Topological Steering —
- NeuroGraph: An AI Graph-Driven Neuro-Symbolic Framework for Explainable Threat Reasoning in Advanced Manufacturing —
- Investigating Assistant Bias in LLM User Simulators Using a Role Vector —
- Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs —
- ExpArt-KG: Artwork Image Description Generation through Iterative Exploration of Knowledge Graphs —
- REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows —
- DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation —
- DK-GBMKKM: Dynamic Kernel-Space Granular-Ball Multiple Kernel k-Means Clustering —
- Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search —
- EEG-AS: Instance-Level Foundation Model Selection for EEG Foundation Models via Behavior Reconstruction —
- SciTrue: Reliable Scientific Claim Validation with Frontier and Open Language Models at the NTCIR SciClaimEval Task —
- Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets —
- Triple-Bottom-Line Sustainability of Language Models for Edge AI: A Comparison Between SLMs and Quantized LLMs —
- Automating Static Code Analysis Through CI/CD Pipeline Integration —
- HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields —
- Visual Framing for News Stance Detection via Image Generation —
- A Study of Hidden-State Optimization Order in Predictive Coding Networks —
- SCoNE: Selective Context-aware Neuron Editing for Robust Retrieval-Augmented Generation —
- Verdict Instability of OOD Scores under Reference Resampling —
- MUGEN: Generating Unlearnable Graph Examples for Multiple Learning Tasks —
- Patterning in Practice: Debiasing Reward Models with Susceptibilities —
- Value Over Language Model: Detecting Original Contribution in Writing —
- PhantomCall: Evading ML Malware Detectors via Function Call Graph Perturbation —
- A Certificate-Producing Cascade for Equational Implication: The SAIR EQT2 Stage 2 Solver —
- Differentially Private Paired Table-Image Multimodal Synthesis —
- SoK: Motion Data Privacy in Extended Reality —
- ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything —
- A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies —
- SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification —
- Agentic Empirical Asset Pricing: Methodological Foundations —
- Online Self-Weighted Fine-Tuning —
- Escaping Reasoning Basin Collapse with History-Biased Search —
- Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis —
- Can Large Language Models Forecast What Researchers Study Next? —
- ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents —
- How Do Language Models Choose Between Context and Memory? —
- S 3martCirc: Self-supervised Smart Circuit Discovery —
- Joint Training Is Not Enough: Conditioned Cross-Granularity Training for Multimodal Document Understanding —
- Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning —
- A Unified Mechanistic Analysis of Knowledge- and Safety-Based Refusals —
- Frozen Cores Need Task Signal: Fisher-Whitened Cross-Covariance for Low-Resource LLM Adaptation —
- Automated Tree Knowledge Graph Construction using Ontology Expansion and Retrieval from Vietnamese History Textbooks —
- Are You Thinking What I am Thinking?: Examining Conceptual Separation in Neural Architectures —
- DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory —
- Semi-Supervised Classification with Informative Missing Labels in Weibull Mixture Models —
- When Features Become Instances: Inverted Contrastive Learning for Unsupervised Feature Selection —
- MROP: Mask-Region Optimized Purification Against Backdoor Attack in Deep JSCC —
- StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability? —
- Subspace Levenberg Marquardt Algorithms in Training Neural Networks —
- RISA: Response Inspection and Selective Actions for Refusal Calibration in Large Language Models —
- Instella-MoE Technical Report —
- SFAD: Speculative Factuality-Aware Decoding —
- TEIDAN: A Multilingual Multiparty Dialogue Corpus —
- Towards a Reliable and Practical Eval Pipeline —
- Effective Interventions Against AI-Enhanced Scams —
- One Policy, Any Budget: Internalizing Budget-Aware Search via Reinforcement Learning —
- AnalysisBank: An Expert Analysis Pattern Library for Financial Report Generation —
- Polished but Unresolved: Identifying Late-Stage Pressure States in Long-Horizon Tool-Use Agents —
- HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution —
- FLaG: Frequency-Domain Latent-attention Gated Pooling for Token Aggregation —
- TWIX: a Two-Stage Approach for End-To-End Named Entity Recognition and Relation Extraction —
- Staged Linguistic Seeding: Grounded Query Expansion for Verified-Unit QA in AI Contact Centers —
- Towards Generalizable Visually Grounded Exploration of Household Devices —
- Verifiable Disaster Storylines and Causal Knowledge Graphs: A Citation-Grounded Pipeline from Heterogeneous Humanitarian Sources —
- Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems —
- Conditional Flow Matching for ML-Based Inverse Design Problems —
- MemoryWalker: Stop Training Agents on Contexts They Never Saw —
- Membership Inference in Fine-tuned Diffusion Language Models via Token-level Memorization Asymmetry —
- Beyond the Clock: Measuring the Value of Adaptive Revision —
- FractalNet-Based Heterogeneous Federated Learning for Orbital Edge Intelligence in Satellite Mega-Constellations: A Wildfire Case Study —
- Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning —
- iPINN for Broadband CARS Phase Retrieval: A Framework for Function Approximation and Inverse Modeling Problems in Nonlinear Spectroscopy —
- Denoising Diffusion Generative Models Secretly Calculate Attentions —
- Using LLMs to Elicit Security Requirements for Service-Oriented Cyber Ranges —
- CacheBridge: Efficient Cross-Model KV Cache Transfer —
- CARE: Contrastive Anchor-based Rubric Evolution for Large Language Model Post-Training —
- Poisson-Gamma Dynamical Systems with Time-varying Transition Dynamics —
- In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access? —
- When Metropolis and Hastings Meet Bradley and Terry: Exact MCMC From Preference Voting —
- RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation —
- VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences —
- Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO —
- DualStake: Dual-Path Confidence Calibration in Deep Research Agents —
- A Dataset for Modeling Iterative Problem-Solving —
- Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches —
- Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling —
- Influence of Logging Frameworks on Bind9 —
- PersianAnonymizer: Evaluating LLM-Labeled Training for Efficient NER-based Anonymization in Persian —
- Few-Shot Out of Domain Intent Detection with Covariance Corrected Mahalanobis Distance —
- CoBRA: Learning Tool-Use Boundaries via Counterfactual Margins —
- Disclosure-Gated User Simulation for Companion-Agent Evaluation —
- Inspicio: Open-Vocabulary, LLM-Based Sense Retrieval for Historical Languages —
- Figures as Programs: Recursive Generation of Editable Scientific Figures —
- Phrase-Localized Language-Contrastive Guidance: Training-Free Localized Accent Control for Code-Switching Text-to-Speech —
- AKRASIA: Stealthy Backdoor Attack on Reasoning-based Code LLMs —
- The Multiple Timescales of Gradient Descent on the Edge of Stability: A Perturbative Derivation of the Central Flow —
- Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees —
- MultiGait: A Multi-Sensor Multi-Perspective Multi-Session Biometric Inference Benchmark and its Dataset —
- Data-Driven Persona-Conditioned Agents for A/B Test Simulation —
- From Truncation to Commitment: Persistent Context in Uniform Discrete Diffusion —
- AgentFactory: Towards Automated Agentic System Design and Optimization —
- HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation —
- Lagged Coupling: Internal Representations Become Readable Before They Become Causal —
- What Limits Robustness in Deep Image Watermarking: An Analysis of Mechanisms and Their Scaling Across Capacities —
- SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations —
- WorldBench: Culturally Grounded Benchmark for Multilingual Agents —
- User Representation via Cross Multi-source Behavior Pre-training for Mobile Games —
- ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning —
- Space Generative AI with Solar Energy Harvesting —
- OUTLETS: Output-Length Prediction from Speculative Decoding Backbones —
- Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration —
- Post-hoc Alignment of LLM-judges to Human Judgment Distribution —
- StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions —
- Modelpedia: A Catalog of Model Findings for the Meta-Science of AI —
- Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation —
- CRSF: Collusion-Resilient Privacy-Preserving Sensor Fusion with Byzantine-Robust Participation —
- Neural Symbollic Regression Using Deep Learning and Sparse Modelling —
- When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP —
- Replicating TRACE: A Practitioner's Guide to Its Threshold and Particle Budget —
- ClinTraceBench: Source-Verifiable Longitudinal Clinical Reasoning over EHR-Derived Dialogues —
- EDRAC: Benchmarking Arabic Dialect Reading Comprehension —
- When Does Online Adaptation Pay on the Edge? A Leakage-Free Evaluation of Warmup, Learning-Rate Selection, and Resource Trade-offs for Time-Series Forecasting —
- Scaled Idempotence in Transformer Attention: Paired OV Geometry and Shared-Value Algebras —
- Overfitting Mitigation via Singular Value Decomposition in Minimum Bayes Risk Decoding —
- Does task decomposition improve automatic NLG evaluation? —
- Subword Segmental BabyLMs: Learning to Tokenise for Sample-Efficient Pretraining —
- Superposed Latent Autoencoder —
- CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs —
- Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate —
- Pre-carved Niches: The Formation Dynamics of Modular Task Partitions in Early LLM Training —
- Johnny Still Receives Spam SMS: Assessing the Robustness of SMS Spam Detection —
- A SoK for SoCs: Reading the TI Leaves on AI for Cyber Threat Intelligence Generation and Sharing —
- LLMPEDIA: Browsing, Verifying, and Comparing the Parametric Encyclopedic Knowledge of LLMs —
- Reveree: Diagnosing LLM Reverse-Engineering Agents —
- Smart Contracts Claimed Vulnerable by the CVE Database, with Labels and Source Locations —
- PersuaRL: Reinforcement Learning-Driven Multi-Expert Selection for Persuasive Dialogue Generation in Insurance —
- Births are difficult to predict even with rich survey and full-population register data —
- CaRL-EM: Cost-Aware Reinforcement Learning for Entity Matching with LLMs —
- FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking Dialogue —
- Identification of Compositional Risks in Data Protection Impact Assessments and Beyond —
- Towards AI-Assisted Clinical Trial Matching: Practical Considerations, Multicenter Evaluation, and Real-World Deployment —
- Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges —
- REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs —
- H2Table: Hierarchical Hypergraph-Enhanced Large Language Models for Complex Table Reasoning —
- Prompt-Robust Language Models: Which Training Strategies Work? —
- What's in Your Agent's Context? Context Privilege Escalation Attacks against AI Agent Harness —
- Multi-Head Self Attention is a Parameter Identification Mechanism —
- Position Matters: Feature Inversion Attacks in ViT Split Inference with Token Reduction and Shuffling —
- MutMem-V2: Cryptographically Authorized Mutation in Persistent Agent Memory Portable Verification and Reproducible Evidence —
- Post-Training Science for Supervised Fine-Tuning —
- Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents —
- Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation —
- Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations —
- Dual Process Motion Planning —
- Solving In-Table Prediction Problems by Deep Neural Networks with Performance Evaluation Using Synthetic Data —
- Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents —
- Position: Privacy Is a Claim, Not a Property of Synthetic Data —
- The Constitutional Coverage Trilemma in AI Governance —
- Some Emotions Run Deeper: Layer-wise Probing and Causal Intervention in Large Language Models —
- Analog-DB: An Agent-First Analog Integrated Circuit Database, From Blocks to Systems —
- Explore Before Committing: Hypothesis-Guided Search for Deep Research Agents —
- One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context —
- A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation —
- Matched Queries for Curvature and Density at Branching Junctions —
- Exploring Sparse Autoencoders in Text-Based Causal Confounding Adjustment —
- VerTox: Verifiable Reward-Guided Corpus Poisoning Against Neural Ranking Models —
- Hidden Services Protocol for Mixnets —
- LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting —
- Probing Factual Knowledge Transfer with Training Data Interventions —
- SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers —
- Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades —
- CHARM: Character Hallucination for Multicultural Role Play Benchmark —
- SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding —
- Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR —
- Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs —
- EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems —
- Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA —
- How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation —
- Behaviorally Effective LoRA Writes Are Sparse and Structured —
- Polish ModernBERT: The Long and Short of Polish Language Understanding —
- InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations —
- When Tokenization is Secretly Output Supervision —
- Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity —
- Contribution-Aware Bandwidth Allocation for Multimodal Split Learning —
- Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations —
- EdiTikZ: Scientific Figure Editing from Revision Trajectories —
- On the Reliability of Generative Augmentation: A Wasserstein-Based Theoretical and Empirical Study —
- Predicting Subsurface Abnormalities Growth using Physics-Informed Neural Networks —
- Provably Safe Sim-to-Real Transfer —
- From Rollouts to Recipes: Self-Contained Post-Training for LLMs —
- CATeye: Coupled Attribute-Topology Invariance Learning for Voucher Abuse Detection —
- TRIAGE: Three-level Routing and Intelligent Agent Guidance for Efficient Execution —
- Learning Sparse Decision Trees via Transformer Variational Auto-Encoders —
- Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search —
- Edge-Girth as a Structural Edge Feature for Graph Neural Networks —
- Diffusion as a Training Curriculum for Timestep-Free Iterative Reasoning —
- When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning —
- Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers —
- Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement —
- Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents —
- GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions —
- Rethinking Learnability in Offline Data-driven Optimization —
- LatentPress: Context Compression Beyond Text and Vision —
- When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation —
- EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation —
- Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall —
- Quantum Sparse Autoencoders for Q-Matrix Estimation in Cognitive Diagnosis —
- Variable Selection for Feature-Based Newsvendor —
- SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue —
- NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games —
- Linear Reusable Neural Bases Architecture for Network Compression —
- Can LLMs Discover Scientific Laws in Real and Parallel Worlds? —
- Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories —
- Gradient-Update Mismatch: Rethinking Conflict-Free Training of Physics-Informed Neural Networks —
- A systematic Approach to constructing a Chance-and-Risk Matrix for Semiconductor Supply Chains —
- From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification —
- Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers —
- From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix —
- Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs —
- Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics —
- The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally —
- StudentSim: Training LLM-based Student Simulators —
- The Rise of Verbal Reinforcement Learning —
- CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses? —
- PyCSP3: Modeling Combinatorial Constrained Problems in Python —
- Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation —
- Cell Attention Networks —
- Building Expressive and Tractable Probabilistic Generative Models: A Review —
- HOACS: Homomorphic Obfuscation Assisted Concealing of Secrets to Thwart Trojan Attacks in COTS Processor —
- FedReview: Review and Dispose Poisoned Updates without Validation Datasets or Historic Knowledge —
- Predicting unobserved climate time series data at distant areas via spatial correlation using reservoir computing —
- GENIE: Watermarking Graph Neural Networks for Link Prediction —
- Keep Everyone Happy: Online Fair Division of Numerous Items with Few Copies —
- Automatic Item Generation for Personality Situational Judgment Tests with Large Language Models —
- Security Testing Framework for Web Applications: Benchmarking ZAP V2.12.0 and V2.13.0 by OWASP as an example —
- Evaluation of Artificial Intelligence Methods for Lead Time Prediction in Non-Cycled Areas of Automotive Production —
- Generalization Bounds for Markov Algorithms through Entropy Flow Computations —
- X-SG squared S: Safe and Generalizable Gaussian Splatting with X-dimensional Watermarks —
- GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods —
- Multi-View Causal Discovery without Non-Gaussianity: Identifiability and Algorithms —
- Permutation polynomials over finite fields from low-degree rational functions —
- Differential Privacy Meets Invariant Statistics: Some Conundrums in Quantifying Trade-Offs —
- A Few Large Shifts: Layer-Inconsistency Based Minimal Overhead Adversarial Example Detection —
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone —
- ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models —
- Online simultaneous inference for quantiles via smoothed stochastic gradient descent —
- Does Synthetic Data Help Named Entity Recognition for Low-Resource Languages? —
- InComeS: Integrating Compression and Selection Mechanisms into LLMs for Efficient Model Editing —
- UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment —
- Efficient Learning of Balanced Signed Graphs via Sparse Linear Programming —
- On the Existence of Consistent Adversarial Attacks in High-Dimensional Linear Classification —
- BOW: Training Language Models to Reason Over Plausible Next Words —
- Towards Provable and Scalable Training of Quantized Neural Networks with Ising Optimization —
- Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture —
- Test Set Quality in Multilingual LLM Evaluation —
- GeoGR 2:Zero-Shot Geospatial Inference via Geostatistically-Guided Iterative Refinement with LLMs —
- FaST: Feature-aware Sampling and Tuning for Personalized Preference Alignment with Limited Data —
- Unsupervised Partner Design Enables Robust Ad-hoc Teamwork —
- Evaluating Style-Personalized Text Generation: Challenges and Directions —
- BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection —
- SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance —
- MATA: Mindful Assessment of the Telugu Abilities of Large Language Models —
- Bayesian and Multi-Objective Decision Support for Incident Mitigation in Cyber-Physical Systems —
- Recurrent State Encoders for Efficient Neural Combinatorial Optimization —
- A Compositional Kernel Model for Feature Learning —
- ViTSP: A Vision Language Models Guided Framework for Solving Large-Scale Traveling Salesman Problems —
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models —
- Uncovering the Computational Ingredients of Human-Like Representations in LLMs —
- Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective —
- Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages —
- The Five Safes as a Privacy Context —
- One Risk Down, Another Up: Cross-Risk Interactions Induced by LLM Defenses —
- TopoAlign: A Framework for Aligning Code to Math via Topological Decomposition —
- One-shot Style Transfer LLM log-probabilities for Authorship Attribution and Verification —
- Compositional Machine Design as Program Synthesis with LLMs —
- HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning —
Important terms
- QTEA
- A new framework for model compression that allows large language models to be quantized down to just 1.7 bits per weight. It uses ternary values and a salient weight system to maintain high accuracy and speed.
- Latent Recurrent Thoughts
- An approach for deeper reasoning that uses a small auxiliary network to refine thoughts in a continuous vector space. This avoids the errors that occur when a model is forced to write out every logical step in text.
- REAL-Q
- A quantization method that fixes information misalignment by using dynamic gradient descent to refine a model block by block. This prevents errors from piling up across layers, significantly reducing error levels in large models.
- Byzantine Placement Influence
- A metric used to predict how the specific location of compromised nodes in a decentralized learning network affects the spread of malicious influence. It helps identify which node configurations are most damaging to honest participants.
- ReNFT
- A method designed to prevent the loss of variety in generative models. It uses counterfactual proposals to recover suppressed visual content, ensuring the model retains its original diversity while still following new reward signals.