AI papers — 2026-09-10
A new way to manage overlapping memories offers a breakthrough for long-term agent stability. Instead of asking a language model to rewrite its entire history, which is prone to error, a framework called ROAM organizes information into tiny, distinct atomic units.
The system classifies how these units relate as independent, equivalent, or conflicting. By grouping these atoms into primary observations and supporting evidence, the system fuses details into compact views that prevent outdated data from confusing the agent. This approach boosts answer precision by up to 29.8 percentage points while reducing irrelevant token noise.
Moving from how agents remember to how they act, there is a new call to evolve process mining from retrospective dashboards into active tools. The proposed BlueSky agenda suggests building agents capable of deciding whether an action should be taken based on privacy budgets and organizational authority.
This would require moving toward mineable artifacts like governance contracts and evidence packages. These tools would allow an agent to legitimately refuse or defer a task.
In the world of decentralized finance, neural networks are replacing hand-written formulas for calculating wallet reputation. A new network called zScore-N was trained using original formulas as a teacher to ensure perfect reproduction.
This network handles millions of wallets across massive scales and is resilient to gaps in data. It cuts the error caused by missing features nearly in half.
The focus on reliability extends to robotics, where researchers have made deep Q-learning more cautious about risk. By using mini-batch transition risk mappings, they enabled an underwater robot to navigate complex environments while avoiding destruction.
This method allows the robot to learn from a small number of samples without being misled by environmental randomness. Consequently, its policies work even in environments it never saw during training.
The security of massive models is more fragile than previously thought, especially in mixture-of-experts architectures. A study shows that directional ablation can strip a model of its ability to refuse harmful requests without extra training.
This technique works on the 320B parameter GLM-5.3-Flash, but it behaves unexpectedly. Researchers found that seventy-four percent of the effect occurs only when editing attention, dense writers, and routed experts all at once.
This means following old recipes for finding these directions will result in silent failure because researchers are looking in the wrong place. This vulnerability is a reminder of how little we understand about internal representations, a theme that carries over into mobile data privacy.
Researchers developed CrossLink to show that observers can stitch together different temporary identifiers used by phones for LTE, WiFi, and Bluetooth. By linking these identifiers across protocols and locations, they reconstructed full movement traces for eighty-three percent of users in simulations.
The difficulty of managing complex systems is also evident in how we train models for specialized tasks like translation. We need better ways to model the messy reality of human health, and a new generative transformer called NOAH attempts to do just that.
By training on over 559 million clinical events from nearly 300,000 patients, this model simulates the continuous evolution of a patient's journey. It can process medical images, time-series signals, and unstructured clinical notes to perform tasks like zero-shot classification or simulating responses to medical interventions.
This ability to handle multimodal data is also being applied to make large language models more reliable through better grounding. A new method called Evidence-Aligned Entity Verification addresses hallucinations in retrieval-augmented generation by checking if specific entities align with retrieved evidence.
It uses counterfactual stability analysis to ensure these alignments hold up even when evidence is slightly perturbed. This makes it more robust than methods relying solely on a model's internal knowledge.
While NOAH focuses on the depth of patient data, other researchers are looking at making the infrastructure behind these models more efficient. A new benchmark called Φ-Bench evaluates if large language models can engineer the complex software stacks that power them.
Unlike previous tests that look at small code snippets, this benchmark evaluates long-horizon tasks like end-to-end system optimization. This helps determine if we are getting closer to autonomous AI infrastructure.
Efficiency is also being tackled by improving how models talk to each other during inference. A framework called X-CoSD allows a small model on a device to work with a large model on a server, even with different vocabularies.
By using hybrid resampling, it only exchanges data for the shared parts of the vocabulary. This speeds up token generation without losing the quality of the larger model.
The way we manage model interactions is also evolving to include personal security measures. A system called Bio-Memory uses biometrics like face or palmprint recognition to ensure an AI agent gives retrieved memories to the correct person.
In tests, this biometric layer created significant gaps in retrieval accuracy between owners and non-owners. This provides a practical way to keep personalized data safe in shared environments.
Security remains a theme even in the deployment of specialized models for critical infrastructure. A hybrid architecture combining Spiking Neural Networks with XGBoost has been developed to protect electrical distribution networks from cyber attacks.
By using the spiking network as a fixed feature extractor, the system stays lightweight enough for edge deployment. Even when attackers try to poison data, this model's performance drops by only 0.9% in certain tests.
Finally, there is a growing debate about how agents should use their skills to get things done. Some researchers find it is more effective to treat skills as independent subagents rather than loading instructions into a single long context window.
By spawning fresh context windows for each subtask, these subagents can execute complex tasks more reliably than an agent trying to juggle everything at once. However, the security of agentic AI is a major concern because these systems can act on instructions hidden within images or audio.
A new benchmark called MMPIBench shows that while most multimodal prompt injection attacks are caught during planning, the vulnerability remains significant. In 12.8% of attempts, attackers successfully deliver instructions through QR codes or fake interfaces.
While only 1% of these result in a completed tool call, the risk is much higher when audio is involved. In those cases, attacks complete in up to 75% of instances for certain models.
This vulnerability to unexpected inputs mirrors the need for better efficiency in other complex computational structures. For instance, zero-knowledge machine learning circuits often suffer from massive redundancy that can be stripped away without compromising security.
By using whole-circuit abstract interpretation, this new method ensures every removed check is still logically supported. This reduces prover time by up to 72.8% and cuts constraints by nearly half.
The push for efficiency extends into how we handle long contexts in retrieval-augmented generation systems. Instead of choosing between recomputing expensive caches or risking accuracy, a new hybrid approach fine-tunes models to be aware of cache concatenation.
This method selectively recomputes parts of the cache to slash the time to first token by 80%. It also improves accuracy on long-context benchmarks compared to just recomputing everything.
We must solve how models process massive information without slowing down, and a new framework called ConvMem tackles this by treating long-context reasoning like a hierarchical convolution. Instead of reading text in a slow, linear chain, it uses an LLM as a convolutional kernel to summarize segments into a logarithmic tree structure.
This approach is training-free and highly parallelizable. It allows the model to outperform existing methods on complex reasoning tasks without the risk of overfitting.
While ConvMem helps models understand more, we still have to figure out if they truly understand across different languages. A new benchmark called SWORD reveals that many models are fragile regarding factual consistency.
They struggle significantly more with distorted statements in East Asian languages than in others. Models often rely on how familiar a sentence looks rather than true verification, showing a performance gap of up to 28 percentage points.
This gap in deep understanding is also visible in specialized linguistic tasks like Filipino speech synthesis. Researchers found that fine-tuning a ByT5 model using an LLM-assisted pipeline can significantly improve the accuracy of converting text to phonemes.
This helps the model better predict stress markers and disambiguate words that are spelled the same but pronounced differently. We must also build AI assistants that know when to step in proactively.
A new theoretical framework defines symbiotic agency as a system operating under a standing, revocable mandate from a human. The AI must constantly decide whether to act, monitor, or refrain based on its perception of the user's situation.
By treating individual behavioral episodes as the unit of analysis, this approach links an agent's decision to act directly to its authorized perception and authority containment. However, making these agents reliable is difficult because they often fail to understand their own limitations.
New research shows that when we ask a model how it would behave, it provides a generic theory of how AI agents work rather than actual self-knowledge. Even when shown exact data points, their reports are no more accurate than if they were describing a hypothetical third party.
Using first-person language mostly just makes them sound more flattering and less prone to admitting harmful behaviors. This lack of self-awareness makes it harder to use multi-agent debates to improve reasoning, as the benefits might be an illusion.
While having models argue can shift reported agreement by over 50 percentage points, it does not seem to change what they believe or improve final answers. In many cases, a model's dissent is just a temporary reaction to a hostile instruction.
If we cannot rely on self-correction through debate, we might need to look at how models handle specific constraints during generation. For discrete tasks like Sudoku, standard diffusion models often get stuck with early mistakes.
Simply changing the sampling method to draw directly from clean predictions can jump validity from 31% to 95% without retraining. We can also bypass expensive retraining by using training-free task vectors to steer model behavior.
Instead of fine-tuning, this method uses forward-pass statistics to map activation steering into weight-space edits. It allows us to amplify or suppress behaviors while keeping general problem-solving skills intact.
We must also address privacy leaks in federated learning, as new work shows passive attackers can reconstruct almost entire batches of data. By framing gradient inversion as a problem of erasure-correcting codes, researchers developed a peeling attack that recovers every sample and label in a batch from just one round of updates.
On ImageNet, this method recovered between 94 and 100 percent of batches up to size 128. This vulnerability is mirrored by risks in how we use models to generate information via retrieval-augmented generation.
A study found that poisoning just a few retrieved documents can drop accuracy from nearly 78 percent down to 43.5 percent. Interestingly, the model mostly just stops answering altogether when the context gets messy rather than hallucinating new lies.
To trust these systems in high-stakes areas like law, we must audit every individual claim rather than scoring answers as a single block. A new two-agent system called GANDR uses a Drafter and a Critic to verify claims against their sources.
It hits 70.8 percent strict accuracy on legal benchmarks, outperforming existing baselines by over 11 points. This is largely because it forces the model to commit to citations that actually exist in the retrieved text.
The difficulty of managing complex interactions also appears in how we train models on evolving data. A new benchmark called TTGBench addresses the fact that most temporal graph models are great at predicting structural changes but terrible at tracking semantic shifts.
By testing 17 different methods, researchers found that traditional graph neural networks handle structure well but fail at semantic drift, while large language models do the exact opposite. To make these models usable on phones, we must fix how they handle heat and power.
A new approach called PELM manages computational load by realizing that not every generated word needs to go through the entire neural network. By combining frequency scaling with speculative decoding, it can speed up inference by 23.1% while cutting energy use by over 50%.
This efficiency in hardware-level execution is mirrored by a way to speed up the math inside the transformer itself. A method called EFQ-Softmax maps attention scores directly to low-bit operands using a single affine rule.
On specialized hardware like the A5 vector unit, this cuts latency by about 40% without hurting model quality. Moving from hardware efficiency to training stability, researchers found that scale-invariant optimization in normalized networks is governed by a precise mathematical law.
They discovered that learning rates and weight decay interact through the parameter norm to create a hidden feedback loop. This can lead to unstable behavior if not carefully balanced.
The math behind regularization also becomes clearer when looking at the geometry of the loss landscape. By connecting divergence-based regularization to Sharpness-Aware Minimization, it turns out both methods are essentially penalizing local curvature to find flatter, more robust minima.
In the realm of security verification, a framework called AutoTrans is making it easier to protect new hardware. It uses a specialized signal extractor and template-based prompting to translate security assertions for RISC-V processors with a 78% automatic success rate.
This automation of complex tasks extends into business logic through CARRE, which helps companies predict customer churn. By combining retrieval-augmented generation with counterfactual scoring, it predicts risk reductions roughly 80% better than standard methods.
Finally, protecting the integrity of models remains a moving target. New work on preventative steering shows that defending against malicious fine-tuning requires active, progressive intensity scheduling to keep up with how model parameters adapt over time.
Today's papers
- ALIGN-HOLD: Experience Alignment for Real-Time Hold Control in Large-Scale Ride-Hailing Matching at DiDi This framework learns optimal driver-order matching delays by using a reward model trained on implicit marketplace preferences. [paper]
- What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores This study finds that the MMLU benchmark primarily measures factual retrieval rather than actual reasoning ability. [paper]
- Scalable Composition of Byzantine Agreements under Reorder Attacks This research introduces a new security model to protect distributed systems against adversaries that reorder network messages. [paper]
- Teacher Geometry Shapes Learnability in Teacher-Student Networks The geometric structure of a teacher network significantly impacts how easily a student network can learn its function. [paper]
- HOPE: Heterophily-Aware Open-Set Node Classification with Pseudo-Extrapolation This method improves graph classification in networks where connected nodes often have different labels by using pseudo-extrapolation. [paper]
- Subgroup Membership Inference Audits of Differentially Private Synthetic Text Differential privacy protects aggregate data well but can still leave specific subgroups vulnerable to membership inference attacks. [paper]
- Improving Cross-Lingual Token Representations by Adding a Pinch of SALT This technique improves how multilingual models understand individual words by adding span-level supervision during training. [paper]
- Physics-informed neural networks by Gradient-Guided Gaussian Adaptive Sampling (3GAS-PINNs) This method uses gradient information to place more sampling points in complex regions, helping neural networks solve difficult physics equations. [paper]
- ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR This approach uses an offline pass to estimate task difficulty, preventing wasted computational effort during the early stages of reinforcement learning. [paper]
- Privacy-Preserving Split Learning for Federated LLM Fine-Tuning This method allows multiple parties to collaboratively fine-tune large models while keeping their private data secure through an obfuscation and recovery scheme. [paper]
- Who Argues What? Joint Argument-Entity Detection and Classification in Political Debates This framework uses large language models to simultaneously identify arguments and the specific entities they discuss in political debates. [paper]
- Data-Centric Post-Training for Financial Reasoning: Mining, Distillation, and Verifiable Learning This pipeline creates high-quality financial training data through mining, distillation, and knowledge graphs to improve model reasoning. [paper]
- API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces Model performance measured via APIs is often significantly higher than what users experience in actual chatbot interfaces. [paper]
- RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts This new benchmark provides realistic data to help robots recognize the movement patterns of older adults and clinical patients. [paper]
- TrajMark: Ownership Attribution and Segment-Level Tamper Localization for Coding-Agent Trajectories This framework watermarks coding agent trajectories to prove ownership and pinpoint exactly where a user has edited the code. [paper]
- From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora This method embeds personal text data directly into a small language model's weights to create highly individualized AI. [paper]
- LEBGen: An LLM-Enhanced Bayesian Network Framework for Few-Shot Travel Survey Data Generation This framework uses large language models to help generate realistic synthetic travel surveys from very small amounts of initial data. [paper]
- An Exponential Deterministic--Randomized Gap in ERM-Oracle Complexity for Thresholds on an Unknown Order This study proves that randomized learning algorithms can be significantly more efficient than deterministic ones when learning thresholds. [paper]
- Recovery Theory for Projected Power Iterations in Permutation Synchronization This paper provides mathematical proofs for how quickly algorithms can synchronize unknown permutations under noisy conditions. [paper]
- SUN: Reaching for Novelty in Reinforcement Learning This reinforcement learning framework selects goals that are both new to the agent and actually reachable within the environment. [paper]
- Exact-Form Regret for Gradient Descent, Mirror Descent and Follow-the-Regularized-Leader This research uses geometry to explain why certain online learning algorithms can avoid making mistakes over time. [paper]
- Maverick: Private and Verifiable LLM Inference Made Practical via Matrix-Vector Multiplication Delegation This system allows users to run large language models on remote servers privately and verify that the results are correct without high overhead. [paper]
- Safe Harness Self-Evolution: A Theoretical Analysis of Feasibility and Limits This study analyzes the mathematical limits of allowing AI agents to automatically update their own prompts and tools safely.
- Uncertainty-Aware Sea-Ice Type Mapping with Multiple Ice Charts This method accounts for disagreements between different expert ice charts to provide more reliable sea-ice mapping. [paper]
- When Can One Obtain Certificates of Optimality Using Positivstellensaetze? This paper explores how to mathematically prove that a solution is optimal even when the problem involves complex, non-polynomial functions. [paper]
- A statistical approach to bias in zero-shot learning: the lens of handwriting recognition This method uses a two-stage statistical approach to correct the tendency of models to favor known classes over unknown ones. [paper]
- From Version Conflicts to Decision Conflicts: Selective Revalidation for Long-Running AI Agents This framework helps long-running AI agents decide if a change in their environment still justifies the action they were planning to take. [paper]
- FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models This method provides a new way to align flow-based models with human preferences using only fixed, offline data. [paper]
- Sparks of In Silico Cognitive Science: Theories from Simulated Data Can Generalize to Humans Automated discovery systems can find psychological theories from simulated behavior that actually apply to real humans. [paper]
- Decentralized network congestion control for DAG-based distributed ledger system This model uses a proof-of-work mechanism to prevent spamming and manage traffic in decentralized ledger networks. [paper]
- The Surprising Effectiveness of Approximate Value Iteration in Self-Play This study shows that simpler value iteration methods can be highly competitive with complex tree searches in many games. [paper]
- SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design This framework improves how AI agents handle long conversations by rewarding them for balancing goals with social politeness. [paper]
- SCCM: Stream Cruise Control Method for Automated Drift Detection and Adaptation This method provides a real-time way to detect when data patterns change and automatically update models to stay accurate. [paper]
- GoAnt: Quality-Diversity Multi-Agent Search for Alpha Factor Discovery in Market Microstructure Data This framework uses multiple agents to find a diverse set of profitable trading signals while avoiding redundant results. [paper]
- Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning This research shows that careful data mixing allows models to reason effectively in many languages, not just English. [paper]
- Inference-Time Nash Alignment This method enables aligning large language models with human preferences during inference without needing to update the model's weights. [paper]
- When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability This study investigates whether AI digital twins can reliably replace human participants in scientific research without losing statistical validity. [paper]
- Do speech foundation models really learn words? This research shows that speech models learn representations of words that are independent of the specific sounds used. [paper]
- Adversarial Training for Tabular Credit Scoring: A Multi-Attack Robustness Evaluation in P2P Lending This study demonstrates that training against one type of attack does not always protect credit scoring models from other types of manipulation. [paper]
- Through the Looking Glass: Directly Reading and Writing Transformers This analysis reveals that only a tiny fraction of a transformer's parameters are actually responsible for making specific predictions. [paper]
- Chimaera: A Mixture-of-Graph-Experts Architecture for Cross-Task and Cross-Dataset Graph Learning This architecture combines different graph models and large language models to perform well across many different types of graph tasks. [paper]
- Three Types of Negation of Triple and its Elements and an Extension of Triple This paper proposes a new way to represent complex negative information in semantic data models. [paper]
- The Answer Path and the Grounding Instruction in LLM Question Answering over Knowledge Graphs This study finds that providing the correct reasoning path is much more important for accuracy than how the facts are formatted. [paper]
- Unthrottling the Tanh Jacobian in SAC: A Negative Result on Bang-Bang Control and MetaDrive This research shows that trying to fix a specific mathematical issue in reinforcement learning can actually make performance worse. [paper]
- NEXUS-MI: Communication-Aware Federated Personalization for Gateway-Coordinated Motor-Imagery Brain-Computer Interfaces This framework reduces the amount of data sent over networks while personalizing brain-computer interfaces for different users. [paper]
- Understanding the Security Boundary of Obfuscation-based On-Device LLM Protection This study identifies fundamental security flaws in current methods used to protect large language models running on local devices. [paper]
- MUCnoHARM@GermEval Shared Task 2026: Retrieval-based In-Context Learning for Defamatory Offences, and Where It Falls Short Using retrieved examples helps AI detect defamatory speech, but it is still not reliable enough for fully autonomous moderation. [paper]
- SRPO: Setwise Relative Policy Optimization for Multi-Agent LLMs This reinforcement learning method optimizes groups of actions at once rather than treating every single response as a separate event. [paper]
- DiffLUT-Net: Differentiable Training of FPGA LUT Networks with Learnable Connectivity This method allows neural networks to be trained directly on hardware chips by learning the best internal connections and logic. [paper]
- Settling: Equilibrium Inference for Non-Convex Validity Sets This approach uses an iterative process to find stable, valid predictions when standard averaging fails. [paper]
- FrogNano: Training a 4B Coding Agent via Online Task Synthesis This study shows that small coding agents can be trained effectively using only synthetic tasks generated on the fly. [paper]
- A positive resolution of the gap-entropy conjecture This paper provides a mathematical proof for how many samples are needed to identify the best option in a multi-armed bandit problem. [paper]
- What Fixed-Rollout pass@k Evaluations Can Identify This study proves that common ways of measuring AI code accuracy cannot reliably predict how well a model performs on much harder problems. [paper]
- IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications This benchmark tests whether AI can identify missing details in research descriptions that would prevent them from being implemented. [paper]
- StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean This study introduces a specialized math benchmark to test how well AI can prove complex theorems about random processes. [paper]
- ROAM: Robust Organization of Atomic Memories for Agents through Semantic Relations This framework organizes an AI agent's memories by their relationships to prevent redundant or conflicting information. [paper]
- From Event Logs to Governed Action: A BlueSky Agenda for Agentic Process Mining This paper proposes a new direction for using AI agents to turn business logs into automated, accountable actions. [paper]
- zScore-N: A Neural Network for On-Chain Wallet Reputation Scoring This neural network replaces manual formulas to provide more accurate and scalable reputation scores for crypto wallets. [paper]
- Mini-Batch Risk-Averse Deep Q-Learning: A Robot Navigation Case Study This method allows robots to learn how to navigate while specifically avoiding high-risk situations through mini-batch sampling. [paper]
- Aegix Pulse: A Traceable Three-Stage Architecture for Personalized Content Generation and Context-Preserving Revision This architecture helps AI generate brand-consistent content and perform revisions that respect the original context. [paper]
The papers
- A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning —
- Personalized Execution Time Optimization for Billion-Scale Scheduled Jobs —
- On the Importance of Data Size in Probing Fine-tuned Models —
- Reinforcement Learning with Temporal-Logic-Based Causal Diagrams —
- A Smooth Polynomial Lyapunov Certificate for Convergence of Q-Learning and Its Smooth Variants —
- End-to-end Stroke imaging analysis, using reservoir computing-based effective connectivity, and interpretable Artificial intelligence —
- BTBR: A Bayesian-Theory-Driven Probabilistic-Fuzzy Framework for Implicit Bias Removal in Large Language Models —
- Influence-Oriented Personalized Federated Learning —
- Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs —
- Efficient Diversity-based Experience Replay for Deep Reinforcement Learning —
- Reinforcement learning for Quantum Tiq-Taq-Toe —
- Towards On-Device Evidence Gathering for Intimate Partner Infiltration: A Feasibility Study for Joint Identity-Action Detection —
- Safe Learning Under Irreversible Dynamics via Asking for Help —
- Quantifying Logical Consistency in Transformers via Query-Key Alignment —
- Subgroup Performance Analysis in Hidden Stratifications —
- When Do Large Language Models Exhibit Unsolicited Deception? —
- Hallucination Detection in LLMs with Topological Divergence on Attention Graphs —
- Predicting Estimated Times of Restoration for Electrical Outages Using Longitudinal Tabular Transformers —
- Neural Morphological Tagging for Nguni Languages —
- ROTATE: Regret-driven Open-ended Training for Ad Hoc Teamwork —
- Temporal horizons in forecasting: a performance-learnability trade-off —
- ESSA: Evolutionary Strategies for Scalable Alignment —
- Effects of relational graph modularity and depth on the learning performance of neural networks —
- Test-time Prompt Refinement for Text-to-Image Models —
- Instance-Aware Algorithm Selection for Maximum Clique via a Dual-Channel Graph Neural Architecture —
- Clone What You Can't Steal: Black-Box LLM Replication via Logit Leakage and Distillation —
- Constraint-Reduced MILP with Local Outlier Factor Modeling for Plausible Counterfactual Explanations in Credit Approval —
- PAPER: Privacy-Preserving Convolutional Neural Networks using Low-Degree Polynomial Approximations and Structural Optimizations on Leveled FHE —
- MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data Generation —
- Why Do LLM Agents Fail in Exploring New Environments? A World-Modeling Perspective —
- DFNN: A Deep Fr'echet Neural Network Framework for Learning Metric-Space-Valued Responses —
- Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation —
- Beyond One-Size-Fits-All: Neural Networks for Differentially Private Tabular Data Synthesis —
- From Representation to Enactment: The ABC Framework of the Translating Mind —
- Bounty Hunter: Autonomous, Comprehensive Emulation of Multi-Faceted Adversaries —
- MORPHEUS: A Multidimensional Framework for Modeling, Measuring, and Mitigating Human Factors in Cybersecurity —
- Meta-RL with Bayesian Linear Task Models —
- From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges —
- Toward Learning POMDPs Beyond Full-Rank Actions and State Observability —
- Theoretical Analysis of Measure Consistency Regularization for Partially Observed Data —
- Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge —
- Revisiting the Shape Convention of Transformer Language Models —
- "Tab, Tab, Bug": Security Pitfalls of Next Edit Suggestions in AI-Integrated IDEs —
- Central Dogma Transformer II: An AI Microscope for Understanding Cellular Regulatory Mechanisms —
- A Patient Simulation Framework for Risk Assessment of Conversational Healthcare AI: Evaluation of an Antidepressant Decision Aid —
- False positive bias in AI-powered speech-based cognitive screening for multilingual English speakers in the UK —
- LoMime: Query-Efficient Membership Inference using Model Extraction in Label-Only Settings —
- Manifold-Aligned Generative Transport —
- MOSAIC: A Universal Agent-Level Interface for Cross-Paradigm Agent Mixing and Human-AI Collaboration —
- Verify to Amplify: Improving Reasoning via Learned Chain-of-Thought Verification —
- Post-Training Language Models for Crosslingual Consistency —
- FairFinGAN: Fairness-aware Synthetic Financial Data Generation —
- Graph Tokenization for Bridging Graphs and Transformers —
- Unbiased and Biased Variance-Reduced Forward-Reflected-Backward Splitting Methods for Stochastic Composite Inclusions —
- Translation Invariance of Neural Operators for the FitzHugh-Nagumo Model —
- RelayS2S: A Dual-Path Speculative Generation for Real-Time Dialogue —
- How to measure the optimality of word or gesture order with respect to the principle of swap distance minimization —
- DAO to (Anonymous) DAO Transactions —
- Zero-shot World Models Are Developmentally Efficient Learners —
- Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning —
- Where is the Mind? Persona Vectors and LLM Individuation —
- Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning —
- Non-Stationarity Breaks Permutation Surrogates in Multi-Agent Reinforcement Learning: Diagnosis and Remedies —
- Threat-Oriented Digital Twinning for Security Evaluation of Autonomous Platforms —
- Multi-Level Narrative Evaluation Outperforms Lexical Features for Mental Health —
- Enabling Real-Time Training of a Wildfire-to-Smoke Map with Multilinear Operators —
- SurF: A Generative Model for Multivariate Irregular Time Series Forecasting —
- Grounded Continuation: A Linear-Time Runtime Verifier for LLM Conversations —
- Judge Circuits Explain Format-Induced Inconsistency in LLM-as-a-Judge —
- WaveGraphNet: Physics-Consistent Guided-Wave Damage Localization through Coupled Inverse-Forward Graph Learning —
- Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs —
- Tracing Computation Density in LLMs —
- Cultural Binding Heads in Language Models —
- ActTraitBench: Quantifying the Knowledge-Decision Gap in Large Language Models via Human-Grounded Behavioral Validation —
- KairosAgent: Agentic Time Series Forecasting with Fused Semantic Reasoning —
- SEA-LION-Embedding: Open and Reproducible Text Embeddings for Southeast Asia —
- See Better, Foresee Better, Act Wiser: Physically Grounded Proactive Modeling and Decision Making —
- BaltiVoice: A Speech Corpus and Fine-tuned Whisper ASR System for the Balti Language —
- Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning —
- Light or Full Verb? A Minimal-Pair Dataset for Probing Phraseological Competence in Language Models —
- Less is MoE: Trimming Experts in Domain-Specialist Language Models —
- Dead Directions: Geometric Singular Learning —
- Expert-Level Crisis Detection in Mental Health Conversations —
- A Robust Framework for Sybil Attack Detection in Vehicular Ad Hoc Networks —
- A Resource for Enthymeme Detection in Controversial Political Discourse —
- LatentDx: Latent Multi-Agent Communication for Cross-Hospital Rare-Disease Diagnosis —
- Running the Gauntlet: Hard Agentic Tasks —
- AI translation of literary texts is "fine", but readers still prefer human translations —
- MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages —
- Permissions on the Loose: Measuring Overprivilege in Real-World Serverless Applications —
- Builder, Defender, Breaker: Measurable Independence and Bounded Autonomy When Generative Models Build, Defend and Test Software —
- TreeThink: A Modular Tree Search Library for Mathematical Reasoning with LLMs —
- EVOQUANT: Self-Evolving Verifier-Guided Strategy Optimization for Robust Quantitative Trading —
- TIDE: Trustworthy and Interpretable Battery Degradation Estimation with Contextual Learning and Symbolic Distillation —
- Phase Structure in Rotary Attention: A Spectral Framework for Semantic Continuity and Execution-Boundary Governance —
- Sequence prediction under a lying oracle —
- Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models —
- Chameleon: An Adaptive AI-Driven Honeypot Architecture Using Threat-Calibrated Particle Swarm Optimization and Semantic Deception Rapidly-Exploring Random Trees —
- Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents —
- Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning —
- Ask Self, Ask Others: Relation Is All You Need —
- Aegix Pulse: A Traceable Three-Stage Architecture for Personalized Content Generation and Context-Preserving Revision —
- APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents —
- A radiographic world model for clinical reasoning and evidence generation —
- The Profit Alignment Problem: How Profit Mandates Induce Alignment Failures in LLMs —
- When Intelligence Becomes Agency: A Theory of Governed, Proactive Agency for Symbiotic AI Systems —
- xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems —
- What Does an LLM-Agent Leaderboard Rank Actually Compare? —
- Understanding the Impact of Model Pruning on Long-Tail Forgetting and Explanation Reliability in Medical Imaging —
- Do Large Language Models Know What They Don't Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty —
- Quantization Amplifies Determinism, Not Bias: Scale-Dependent Behavioral Effects of Serving-Time Weight Compression —
- PRIMUS: Identity, Governance, and Verification for Multi-Agent Federations —
- FrogNano: Training a 4B Coding Agent via Online Task Synthesis —
- From Event Logs to Governed Action: A BlueSky Agenda for Agentic Process Mining —
- When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability —
- Mini-Batch Risk-Averse Deep Q-Learning: A Robot Navigation Case Study —
- Sparks of In Silico Cognitive Science: Theories from Simulated Data Can Generalize to Humans —
- From Version Conflicts to Decision Conflicts: Selective Revalidation for Long-Running AI Agents —
- What Does Multi-Agent LLM Debate Actually Change? A Layered Analysis of Disagreement and Answer Quality —
- Automated Design of Inventory Policy with Large Language Models: An Exploratory Study —
- Inference-Time Nash Alignment —
- RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts —
- CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information —
- Artificial Intelligence-Assisted Digital Inventory of Cultural Heritage & Traditional Knowledge: Case for Indonesian Open Digital Library of Culture —
- Router Prior Bias: Preserving Base Routing Structure in MoE Post-Training —
- SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents —
- WorldAgen: Unified State-Action Prediction with Test-Time World Model Training —
- Key Path Identification for Resolving Knowledge Conflicts via SAE-based Steering —
- OntologyBench: Can Dense Retrieval Satisfy Structured Biomedical Constraints? —
- A Theory of Reliable Self-Evolution for Agent Harnesses —
- Less Is Personal: Learning Minimal Sufficient User Profiles for Personalized Language Models —
- Bridging the Semantic-Utility Gap in Multimodal RAG via Generator-in-the-Loop Alignment —
- Qiushi Engine on AstaBench E2E-Bench-Hard —
- TTGBench: Benchmarking Topological Evolution and Semantic Drift in Text-attributed Temporal Graphs —
- Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts —
- zScore-N: A Neural Network for On-Chain Wallet Reputation Scoring —
- Agentic ML Exploration (A-MLE) for Ads Ranking —
- CircuTutor: Transforming Static Circuit Problems into Intelligent and Dynamic Tutoring —
- Evidence-Aligned Entity Verification for Hallucination Detection in Retrieval-Augmented Generation —
- Three Types of Negation of Triple and its Elements and an Extension of Triple —
- MemForest: Efficient Agent Memory Management via EventTree Partitioning and Progressive Merging —
- Beyond Coherence: Benchmarking Professional Editing-Technique Execution in Multi-Shot Audio-Video Generation —
- LEBGen: An LLM-Enhanced Bayesian Network Framework for Few-Shot Travel Survey Data Generation —
- FastE: Readout-Triggered Token Compression for LLM Embedding Inference —
- Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models —
- EvolveScaler: Synthesizing Information-Evolution Contexts via Executable State Machines and Natural-Language Rendering —
- SRPO: Setwise Relative Policy Optimization for Multi-Agent LLMs —
- Personalizing LLM Agent Memory Using Biometrics —
- BIO-MEMART: Biometric-Aware KV Cache Memory for Multi-User LLM Agents —
- AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems —
- Graph-Based Personalized Memory for LLM Agents: Representation, Evolution, Retrieval, and Evaluation —
- CLAMP: Constrained Decoding for Vision-Language Embodied Planning —
- Suan: Rectifying Direct Preference Safety Alignment in Large Language Models —
- SUN: Reaching for Novelty in Reinforcement Learning —
- MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models —
- Neither Adversarial Training Nor Purification: Emergent Adversarial Robustness from Oscillatory Predictive Learning —
- HOPE: Heterophily-Aware Open-Set Node Classification with Pseudo-Extrapolation —
- Chimaera: A Mixture-of-Graph-Experts Architecture for Cross-Task and Cross-Dataset Graph Learning —
- GoAnt: Quality-Diversity Multi-Agent Search for Alpha Factor Discovery in Market Microstructure Data —
- BAFF: Bid-Aware Filter Family for Mitigating Training Data Interference in RTB A/B Tests —
- Application of curiosity driven exploration methods for hardware interference identification —
- When Can One Obtain Certificates of Optimality Using Positivstellensaetze? —
- PAC-Bayesian Bounds for Learning Partially Observed Stochastic Linear Time-Invariant State-Space Systems with Inputs and Sub-Gaussian Noise —
- It's All in the Way You Say It: The Role of Information Representation in LLM-Based Glycemic-Event Prediction —
- Adaptive Anisotropic Attention for Axis-Structured Signals —
- Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course —
- Length Generalization for Transformers via Compression —
- Earth System World Model for What-If Simulations: A Case Study for Terrestrial Ecosystems —
- API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces —
- The BatchNorm Illusion: Diagnosing Normalization Artifacts in Machine Unlearning Evaluation —
- SkillAdam: Stable and Efficient Skill Evolution for Agents —
- GraphFAS: A Distributed System for Automated Graph Feature Generation and Selection in Industrial Transaction Networks —
- Physics-Informed Deep Learning for False Ventricular Tachycardia Alarm Reduction in the ICU —
- Deposon: An Auditable, Conservation-Guaranteed, Game-Theoretically Tested Scattering Layer over LLM Reasoning Paths —
- Let It Go or Learn to Self-Correct: Continuous Diffusion for Constrained Discrete Tasks —
- Do Reasoning Representations Help Humans Evaluate LLM Outputs? —
- Training-Free Task Vectors for LLM Behavioral Control —
- Time-Varying Data as Sheaves: an Invitation to Narratives —
- PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games —
- Multi-Task Learning for Sparsely-Labeled Time Series: A Case Study on Cold-Hardiness Modeling —
- ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR —
- Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training —
- The Surprising Effectiveness of Approximate Value Iteration in Self-Play —
- Curriculum Learning as Transport: Understanding Curricula with Wasserstein Geodesics —
- MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents —
- When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay —
- Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails —
- NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting —
- Spectral origin of the topological gap exponent d + eta: mechanism, kernel, decomposition, and scope —
- Physics-informed neural networks by Gradient-Guided Gaussian Adaptive Sampling (3GAS-PINNs) —
- World-Time Compute with Verified Code World Models —
- X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding —
- OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows —
- A Subsampled Davis-Kahan Bound for Large-Scale Eigenspace Estimation —
- AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents —
- Compute-Bounded Security Assurance - Coverage, Verification, and Response under Resource Constraints —
- Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks —
- Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution —
- Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts —
- In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Post-Retrieval Context Tampering —
- Critical initialization destabilizes higher input derivatives in wide scalar-input networks —
- What Fixed-Rollout pass@k Evaluations Can Identify —
- DiffLUT-Net: Differentiable Training of FPGA LUT Networks with Learnable Connectivity —
- CAST: Canonical Approximate Schur Tree for Approximate Cholesky on Graphs —
- Accountable and uncertainty-aware evaluation of sensor-based AI under distribution shift: devices, subjects, and nearly three years underground —
- Tensor Network Moral Graph Recovery of Discrete Probability Distributions —
- StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean —
- Literati: Towards Anytime Optimal Shape Generalized Trees via AO* —
- Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding —
- Cross User/App Network Attacks - Hijacking TCP Connections and DNS Cache Poisoning via a Malicious User/App (Extended Version) —
- SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection —
- Auditable Emergency Triage for Maternal and Newborn Care in India —
- Encrypt What Matters: When Selective Homomorphic Inference Is Efficient —
- Do LLMs Make More Mistakes If They Do Not Believe the Input Data? —
- Explaining f-Divergence-Based Regularization via Local Curvature and Sharpness-Aware Minimization —
- Constraint-Aware Discrete Black-Box Optimization Using Tensor Decomposition —
- What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores —
- LLMSec-AV: A Vulnerability Taxonomy and LLM-Driven Software Weakness Discovery Framework for Autonomous Vehicles —
- XAI-Refine: An Automated Explanation-Knowledge Loop for Brain-Age Prediction —
- X-amine509: Predicting the Practical Risk Level of Enterprise X.509 Certificates —
- An Experimental Evaluation of Multimodal Prompt Injection Attacks on Agentic AI Frameworks —
- Benchmarking Hybrid Deep Research Across Database Querying and Web Search —
- DuplexJail: Spoken Interruption Attacks on Full-Duplex Speech Models —
- FPGA Acceleration of Fully Homomorphic Encryption with Adaptive Key Switching —
- Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements —
- Applying foundation model embeddings towards urban livability evaluation —
- SCCM: Stream Cruise Control Method for Automated Drift Detection and Adaptation —
- Efficient Leakage-Free Neural Architecture Search under Leave-One-Subject-Out Evaluation —
- Tensor-Train Weak SINDy: Identifying High-Dimensional Nonlinear Dynamics —
- Concept drift mitigation through community and spectral graph analysis for the detection of cyberattacks in network traffic —
- Uncertainty-Aware Sea-Ice Type Mapping with Multiple Ice Charts —
- Exact-Form Regret for Gradient Descent, Mirror Descent and Follow-the-Regularized-Leader —
- Code-to-Harness: Distilling Black-Box Optimizers from Self-Play —
- Mode Coverage in Normalizing Flow Boltzmann Generators via Log-Ratio Variation —
- From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function Calls —
- Unthrottling the Tanh Jacobian in SAC: A Negative Result on Bang-Bang Control and MetaDrive —
- Efficient Fairness Auditing Across Guidance Scales in Text-to-Image Diffusion Models via Causal Abstraction —
- Audio Deepfake Detection Using Temporal Coherence Analysis —
- The Mutations of Machine Speech —
- Recovery Theory for Projected Power Iterations in Permutation Synchronization —
- A Statistical Approach to Estimating Sample Size of Machine Learning Models —
- An Efficient and Effective Agentic Group Shilling Attack on Recommender Systems —
- TEFM: Token-Efficient Faithful Modeling for Structured Data —
- Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning —
- BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models —
- High-probability guarantees for linear accessibility in feature superposition —
- Towards Automatic Evolution Tree Generation from Citation Graphs —
- Robust Industrial Cyber Physical Classification Using Neuromorphic Temporal Embeddings and Hybrid SNN XGBoost Under Machine Unlearning Attacks —
- Positional task conditioning for scalable defect detection across product families in large product catalogs —
- Reproducing Omitted Temporal Expressions in Japanese News for Retrieval-Augmented Applications —
- Learning with Synthetic Data via SGD in High-Dimensional Linear Regression —
- When Ad Networks Misbehave: Understanding Risks of Semi-Drive-By Splash Ads —
- Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features —
- CityPlanner: A Sandbox Agent for Executable Urban Planning —
- Distillation of Synthetic Data for Time Series Foundation Models —
- Teacher Geometry Shapes Learnability in Teacher-Student Networks —
- Scalable Composition of Byzantine Agreements under Reorder Attacks —
- Why Learning Rediscovers the Closed-Form Diagonal Regularizer —
- Cascading Gradient Inversion via LT-Code Inspired Peeling in Federated Learning —
- PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling —
- Session Attestation for Unmodified TLS Services in Confidential Virtual Machines —
- SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia —
- Muon-C: Operator-Aligned Muon for Convolutional Kernels —
- X2-NativeCursor: Native-Token Text Progress Tracking for Incremental-Text Streaming Codec TTS —
- Settling: Equilibrium Inference for Non-Convex Validity Sets —
- Which Medical Questions Deserve Rationales? Perturbation-Sensitive Selection for Robust QA —
- ALIGN-HOLD: Experience Alignment for Real-Time Hold Control in Large-Scale Ride-Hailing Matching at DiDi —
- Looped GPT-BERT: Trading Parameters for Computation in Small Language Modeling —
- When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination —
- PrivAudit: A Dual-Lens Auditing Framework for Website Privacy Practices under the CCPA —
- Kernel-Complexity Edge Sanitization for Training-Free Defense against Structural Graph Attacks —
- Scaling E-Commerce Attribute Extraction with Parallel Decoding —
- StreamAlign: Streaming Text-Aligned Speech Tokenization —
- EFQ-Softmax: Exp-Free Quantization for Softmax —
- EEGBind: Detecting Source-Level Interictal Epileptiform Discharges via EEG-Centric Multimodal Binding —
- Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online Safety —
- SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design —
- CARRE: Counterfactual Action Retrieval and Reason Evaluation for Explainable Churn Prescription —
- Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both? —
- SymbolicLight V2: Hybrid Neuromorphic Architecture and Sparse Execution for Low-Energy Language Inference —
- Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward —
- ROAM: Robust Organization of Atomic Memories for Agents through Semantic Relations —
- BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL —
- NEXUS-MI: Communication-Aware Federated Personalization for Gateway-Coordinated Motor-Imagery Brain-Computer Interfaces —
- Evaluating Model Retraining under Drift: Paired Comparisons of Cumulative Subgroup Disparity —
- MUCnoHARM@GermEval Shared Task 2026: Retrieval-based In-Context Learning for Defamatory Offences, and Where It Falls Short —
- How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE —
- Privacy-Preserving Split Learning for Federated LLM Fine-Tuning —
- A practical DIRECT-type algorithm for medium-scale black-box global optimization —
- CS-Guard: Benchmarking LLM Guardrails for Code Generation Security —
- Quantifying IIoT Sensor Node Criticality by Fusing its Data Criticality and Security Vulnerability —
- UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model —
- Lightweight Zero Trust via Automotive SDN —
- HyperTrace: Hypothesis-Based Preference Tracing for Online LLM Personalization —
- Subgroup Membership Inference Audits of Differentially Private Synthetic Text —
- S cubed-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants —
- A Unifying Perspective on Probabilities as Model Predictions —
- When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors —
- Leveraging Fine-grained Error Correction in Korean Speech Recognition for Consultation Services —
- Strangers to Themselves: What Language Models Say About Themselves Is Generic —
- Deep and shallow biases in language models —
- Contrastive Projection: Reading Transformer Internals by Differencing Logit Lenses —
- Beyond Conventional Federated Learning via High-Order Regularization —
- FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models —
- Adversarial Training for Tabular Credit Scoring: A Multi-Attack Robustness Evaluation in P2P Lending —
- Improving Cross-Lingual Token Representations by Adding a Pinch of SALT —
- Decentralized network congestion control for DAG-based distributed ledger system —
- CrossLink: Breaking Location Privacy by Linking Device Identifiers Across Protocols —
- 5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs —
- Towards Stress-Aware Sentence-Level Filipino G2P With Weakly-Supervised ByT5 Fine-Tuning —
- Optimal Value Inference for Reinforcement Learning —
- Dependency-Aware ROM/CBD Correctness Bounds for ML-KEM-768 at the Heuristic Failure Scale —
- Multi-Functional Embedding Models for Funder Name Disambiguation in Scientific Publication Records —
- VLX-VR: An Agentic-Aware Video Reasoning Model —
- Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit Signal —
- SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better Teachers —
- MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes —
- MedDeID enables locally governed clinical-text de-identification from real or synthetic training data —
- Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training —
- OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization —
- AutoTrans: AI-Assisted Automatic Translation of Security Assertions for RISC-V Processors —
- A statistical approach to bias in zero-shot learning: the lens of handwriting recognition —
- RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases —
- Distributed and Private Textual Data Synthesis from Embeddings —
- Data-Centric Post-Training for Financial Reasoning: Mining, Distillation, and Verifiable Learning —
- Understanding the Security Boundary of Obfuscation-based On-Device LLM Protection —
- ProbPlug: A Plugin Uncertainty Network for Reliable Confidence in LLM Binary Classification —
- Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-Tuning —
- Sound Debloating of Redundant Checks in Zero-Knowledge Machine-Learning Circuits —
- YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models —
- From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora —
- Who Argues What? Joint Argument-Entity Detection and Classification in Political Debates —
- An Exponential Deterministic--Randomized Gap in ERM-Oracle Complexity for Thresholds on an Unknown Order —
- Politics of Feelings: Emotional Expression and Legislative Effectiveness in the U.S. Congress —
- Through the Looking Glass: Directly Reading and Writing Transformers —
- -Bench: Can Large Language Models Engineer the Infrastructure That Powers Them? —
- The Answer Path and the Grounding Instruction in LLM Question Answering over Knowledge Graphs —
- Two-Token Features and Small-Large Ensembles for VLM Hallucination Detection —
- Meme Coin Factories: Uncovering Large-Scale Manipulations on pump.fun —
- DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs —
- Are Unreachable Nodes Truly Safe? Fully Eclipsing Monero's P2P Network! —
- Maverick: Private and Verifiable LLM Inference Made Practical via Matrix-Vector Multiplication Delegation —
- KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints —
- GANDR: Claim Auditing for Verifiable Legal Answer Generation —
- An Empirical Analysis of ReDoS Vulnerabilities and ReDoS Detection Tools —
- The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding —
- Learning Intrusion Response Strategies for OT Systems —
- RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding —
- On-Policy Distillation for Vision-Language Model Adaptation, an Effective Paradigm on Low-Quality Multimodal Data —
- From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning —
- Rosetta at AlexandriaX-2026: LoRA-Adapted NileChat for Context-Aware Dialectal Arabic Dialogue Translation —
- Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization —
- TrajMark: Ownership Attribution and Segment-Level Tamper Localization for Coding-Agent Trajectories —
- Do speech foundation models really learn words? —
- ConvMem: Convolutional Memory for Long-Context Reasoning —
- Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning —
- IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier —
- Wicked Problem, Parsimonious Solution: Securing Electric Vehicle Charging Station Software —
- A positive resolution of the gap-entropy conjecture —
- Towards Tackling Application Logic Flaws through Autonomous Formal-Logic Modeling and Automated Reasoning —
- IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications —
Important terms
- ROAM
- A framework for managing long-term AI memory by breaking information into tiny, distinct atomic units. It classifies these units as independent, equivalent, or conflicting to prevent outdated data from confusing the agent.
- zScore-N
- A neural network designed for decentralized finance that replaces manual formulas for calculating wallet reputation. It is trained using original formulas as a teacher to ensure it can handle massive scales and missing data.
- Evidence-Aligned Entity Verification
- A method to reduce hallucinations in retrieval-augmented generation by checking if specific entities match the retrieved evidence. It uses counterfactual stability analysis to ensure these links remain accurate even when data is slightly changed.
- ConvMem
- A framework that handles long-context reasoning by treating it like a hierarchical convolution. It uses an LLM as a kernel to summarize text segments into a logarithmic tree structure, making processing faster and more parallelizable.