mR AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA. The paper proposes a novel generalized framework called multimodal...
QA-Merging: Query-Adaptive Reasoning via Layer Selective Model Merging. The paper introduces Reasoning Pattern Alignment Merging (RPAM), a layer-wise...
A Generative Deep Learning Workflow for Inverse Molecular Design of Fuels. This paper presents a generative deep learning framework for the inverse design...
Decode-Branch Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation. The paper introduces the Dual-Flow Transformer, an architecture that decouples...
DR.GAP: Mitigating Bias in Large Language Models using Gender-Aware Prompting with Decoupled Reasoning. The paper proposes DR.GAP (Demonstration and Reasoning for Gender-Aware...
PhyxMamba: Chaotic System Reconstruction from Short Context Observations with Generative State-Space Models. Problem Statement Long-term forecasting of chaotic systems remains a...
DirMixE: Harnessing Test Agnostic Long-tail Recognition with Hierarchical Label Variations. This paper explores test-agnostic long-tail recognition, a challenging...
BiAxisBias: Evaluating LLM Bias Beyond a Single Prompt and a Single Explanation. The paper introduces BiAxisAudit, a novel framework designed to evaluate bias...
One-shot Robust Federated Learning of Independent Component Analysis. This paper investigates a general robust one-shot aggregation framework for...
Free-Flow Class-Incremental Learning: Towards Robust CIL under Variable Class Arrivals. This paper introduces and formalizes a new, more realistic setting for...
Spectral Certificates and Projection-DPP Rounding for Determinantal MAP Selection. Spectral DPPs via NEPv: A Scalable Continuous Relaxation of Determinantal MAP...
PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows. The paper introduces PolyWorkBench, a benchmark for evaluating LLM agents on...
Multimodal Language Models Benchmarked Against the NRC Reactor Operator Licensing Examination: Fine-Tuning and Retrieval Strategies. This paper presents a systematic benchmark of a 31-billion-parameter...
Evaluating the trustworthiness of the Fréchet Inception Distance with stochastic embedding representations. This paper investigates whether uncertainty quantification (UQ) techniques,...
PRO-Bid: Pareto-Prioritized Regret Optimization for Constraint-Aware Generative Auto-Bidding. This paper introduces PRO-Bid, a constraint-aware generative auto-bidding...
SymbolicLight V1: Spike-Gated Dual-Path Language Modeling at High Activation Sparsity. and Core Contribution The paper presents SymbolicLight V1, "a spike-gated...
The papers
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL — The paper addresses a critical failure mode in multi-agent reinforcement learning for human-AI interaction. Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. [episode]
Decentralized Federated Learning by Partial Message Exchange — The paper introduces PaME (DFL by Partial Message Exchange), a novel decentralized federated learning (DFL) algorithm designed to improve the trade-off among communication efficiency, privacy preservation, and model accuracy. [episode]
MoRFI: Monotonic Sparse Autoencoder Feature Identification — MoRFI: Monotonic Sparse Autoencoder Feature Identification Abstract Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction. [episode]
Red-Teaming the Agentic Red-Team — Summary This paper presents the first in-depth security analysis of agentic systems designed for offensive security operations, termed "agentic-red-teams." The authors analyze 12 open-source tools of this class and demonstrate that they introduce new attack vectors against the or [episode]
In-Context Source and Channel Coding — Summary The paper, titled "In-Context Source and Channel Coding," proposes a receiver-side framework called In-Context Decoding (ICD) to enhance the robustness of Separate Source–Channel Coding (SSCC) systems, particularly in low Signal-to-Noise Ratio (SNR) regimes. [episode]
Annealed Softmax Greedy in Many-Armed Bayesian Bandits — Summary This paper investigates whether annealed softmax (Boltzmann) policies, which are agnostic to epistemic uncertainty, can achieve near-optimal Bayes regret in the many-armed Bayesian Bernoulli bandit setting. [episode]
Differentiable Thermodynamic Phase-Equilibria for Machine Learning — Summary The paper introduces DISCOMAX, a differentiable algorithm for phase-equilibrium calculation that guarantees thermodynamic consistency at both training and inference, only subject to a user-specified discretization. [episode]
The Optimal Sample Complexity of Multiclass and List Learning — Summary This paper resolves a longstanding open conjecture in learning theory, thereby determining the optimal sample complexity of multiclass and list learning. Main Result The central contribution is a positive resolution of a conjecture by Daniely and Shalev-Shwartz (2014). [episode]
A Machine-Learned Comorbidity Index — The paper introduces a Machine-Learned Comorbidity Index (MLCI) that "maps diagnosis codes to a single scalar by maximizing the normalized Hilbert–Schmidt Independence Criterion (nHSIC) between the learned score and multiple clinical outcomes." The authors state that "MLCI capt [episode]
A Theoretical Framework for Statistical Evaluability of Generative Models — Authors: Shashaank Aiyer, Yishay Mansour, Shay Moran, and Han Shao arXiv: 2604.05324v2 [cs.LG], June 2026 --- The paper addresses a fundamental question in generative model evaluation: "Can the performance of generative models be evaluated statistically from finite samples?" The [episode]
Reasoning Pattern Alignment Merging for Adaptive Reasoning — The paper introduces Reasoning Pattern Alignment Merging (RPAM), a layer-wise model merging framework designed to achieve query-adaptive reasoning by combining a Long-CoT reasoning model with a Short-CoT instruction model. [episode]
Towards Understanding Linear Word Analogies — This paper provides a formal explanation for why word analogies can be solved using vector arithmetic in non-linear embedding models such as skip-gram with negative sampling (SGNS) and GloVe, without making the strong assumptions of prior theories. [episode]
DiffImaginE: Imagine to Verify Entity Types with Diffusion — Summary DiffImaginE is a method for multimodal named entity recognition (MNER) that reformulates type verification as conditional latent diffusion inference. The paper states: "DiffImaginE, which formulates MNER type verification as conditional latent diffusion inference. [episode]
Generative Deep Learning Framework for Inverse Design of Fuels — Summary This paper presents a generative deep learning framework for the inverse design of fuels, combining a Co-optimized Variational Autoencoder (Co-VAE) with quantitative structure-property relationship (QSPR) techniques to enable accelerated discovery of fuel molecules with h [episode]
jina-vlm: Small Multilingual Vision Language Model — Summary The paper presents jina-vlm, a token-efficient 2.4B parameter vision-language model (VLM) that achieves state-of-the-art multilingual VQA performance among open 2B-scale VLMs. [episode]
Budget-Aware Tool-Use Enables Effective Agent Scaling — The paper introduces a systematic study of budget-constrained tool-use agents and their test-time scaling behaviors, focusing on web search agents. The authors identify a critical bottleneck: "standard agents lack inherent budget awareness. [episode]
Efficient Code Embeddings from Code Generation Models — The paper introduces jina-code-embeddings, a novel code embedding model suite designed to retrieve code from natural language queries, perform technical question-answering, and identify semantically similar code snippets across programming languages. [episode]
Learning Multi-Timescale Interventions under Safety and Resource Constraints — Summary The paper introduces the Multi-Timescale Intervention Markov decision process (MTI-MDP) and a corresponding reinforcement learning framework called MINT (Multi-timescale Intervention Network Training) to address sequential decision-making problems where actions have effec [episode]
HYGENE: A Diffusion-based Hypergraph Generation Method — Summary This paper introduces HYGENE, a diffusion-based method for hypergraph generation, which the authors claim is "the first attempt to employ diffusion models for hypergraph generation." The core problem addressed is the difficulty of generating realistic and diverse hypergra [episode]
Hybrid quantum recurrent neural network for remaining useful life prediction — Summary This paper introduces a Hybrid Quantum Recurrent Neural Network (HQRNN) framework for predicting the Remaining Useful Life (RUL) of jet engines, evaluated on the NASA Commercial Modular Aero-Propulsion System Simulation (C-MAPSS) dataset, specifically the "FD001" subset. [episode]
Geometric Self-Supervised Pre-training for Neural Combinatorial Optimization — Summary This paper introduces a geometric self-supervised pre-training framework designed to improve the generalization capabilities of Neural Combinatorial Optimization (NCO) models, specifically for routing problems such as the Traveling Salesman Problem (TSP). [episode]