AI papers — 2026-10-10
CMGL aims to improve cancer subtype classification by using confidence guidance within multi-omics graph learning. This approach is important because accurately classifying cancer subtypes is crucial for developing personalized treatment strategies. Researchers explored how this works by looking at predictive multiplicity in cell-fate assignment using label-free rashomon sets, which suggests they are trying to understand the limits of certifying individual cell decisions.
The research also touched upon learning infinite context windows in recurrent architectures through spatial neural computing. This method is a way to give models much longer memories. This connects to characterizing learned dynamical structure in personalized models of brain disorders, where similar predictive fits can arise from different latent dynamics.
Another piece involved developing a certificate-driven agentic harness for scientific program optimization. This suggests a self-improving system for porting software across scientific programs. Finally, they looked at quotient-space exploration for genome-scale metabolic model repair to explore beyond action entropy in these models.
The work on enhancing the reasoning capabilities of large language models through self bootstrapped prolog based chain of thought is what matters most right now. This directly addresses how these systems can perform complex logical steps rather than just pattern matching. This approach, called Thought-Like-Pro, involves using a self bootstrapped prolog based chain of thought to enhance the reasoning of large language models. It means the model is essentially teaching itself better ways to reason by creating its own intermediate reasoning steps before reaching an answer.
This is building on earlier efforts in improving model performance, such as InfiFPO which focuses on implicit model fusion via preference optimization in large language models. That work suggests that by optimizing preferences, you can make different parts of a large language model work together more effectively. Similarly, the investigation into the robustness of llms in mathematical reasoning through mathematically equivalent transformation of advanced mathematical problems is also significant because it tests how well these models handle complex math when they are presented in slightly different but equivalent forms.
Then there is the work on fair gptq which deals with bias aware quantization for large language models. This is important because it tackles fairness issues by adjusting how the model's weights are compressed. This contrasts with the work on policy learning with a language bottleneck, which explores how to learn policies when there is a language bottleneck present in decision-making processes. Finally, enabling quantum natural language processing for hindi language shows an effort to apply cutting-edge computational techniques to specific, complex linguistic challenges.
The work on GUI-KV is particularly important because it tackles the efficiency of agents interacting with graphical user interfaces. This is crucial for making complex systems usable. They introduced a method that leverages a KV cache to improve how these agents process visual information from GUIs while incorporating spatio-temporal awareness to understand changes over time. This means the agent can keep track of evolving visual states more effectively than previous methods allowed.
A related effort focused on the reasoning-planning disconnect in training vision-language driving models. This is significant because it points to a fundamental gap in how these large models actually plan actions based on what they see. The researchers explored this by looking at where the model's perceived understanding diverges from its actual planning capabilities during training.
AyurParam presented a state-of-the-art bilingual language model specifically for Ayurveda. This is valuable because it aims to provide deep linguistic understanding in a specialized domain that often lacks robust digital resources. This model was developed as a significant step toward better handling complex, nuanced medical terminology in both languages.
Towards scalable meta-learning of near-optimal interpretable models involved generating synthetic models to learn how to create better ones. This is important because it seeks a systematic way to build reliable and transparent AI systems without needing extensive manual tuning for every new task.
LMSpell offered spell correction using pre-trained language models, providing a practical application of existing language model capabilities for improving text quality. This work builds on the idea that large language models can be adapted for specific linguistic tasks efficiently.
The most significant work today involved testing how well diffusion language models can scale test-time. This is crucial because it addresses the practical deployment challenges of making these large models useful in real-world scenarios. This was explored through reward-guided stitching, where researchers used a method to stitch together different model outputs based on rewards to improve performance.
This work connects directly to the efforts assessing the limits of language agents with one million benchmarks, which shows how far these agents are from human expertise. Furthermore, there is ongoing research into agentic critical training, which seems designed to refine these agents' decision-making capabilities beyond simple pattern matching.
Another important area involves understanding how social registers shape instruction topology in large language models through imperative interference. This suggests that the way we phrase commands significantly alters how the model responds, a concept related to selective stance accommodation and interaction reorganization in generative agent societies.
Finally, there is work on limited stereotype control within mixture-of-experts language models using routing reweighting. This method attempts to manage biases by selectively directing the model's expertise based on the input context. This contrasts with domain-adapted retrieval for in-context annotation of pedagogical dialogue acts, which focuses more narrowly on adapting models for specific teaching contexts rather than broad social control.
The most significant development today involves the State Stream Transformer V2, which tackles the challenge of latent space reasoning by employing parallel training of nonlinear recurrence. This approach is crucial because it aims to give models a more nuanced way to understand complex information within their internal representations.
Following that, we saw work on SkillGraph, which uses skill-augmented reinforcement learning for agents by evolving skill graphs. This means the agents are not just learning tasks but are actively building and refining a map of how skills connect to each other during the learning process.
Another important piece was CiteVQA, which focuses on benchmarking evidence attribution for trustworthy document intelligence. This work is vital because it helps us figure out exactly where a model gets its answers from in complex documents, which builds trust in automated systems.
We also explored uncertainty-aware budget allocation for adaptive test-time reasoning. This technique is important because it allows the system to dynamically decide how much computational power to use based on how uncertain it is about an answer at that moment.
Then there was MemTrace, which focuses on tracing and attributing errors within large language model memory systems. This helps us understand where mistakes are originating in the model's long-term knowledge storage.
Finally, we looked at Grokking or Glitching, examining how low-precision drives slingshot loss spikes during training. This is a finer detail about the stability of neural network training, showing that even small changes in precision can cause significant instability.
The most significant development was the work on LoRi, which introduces low-rank distillation to improve implicit reasoning in large language models because this is key for making these models more capable of complex inference. This approach involves distilling knowledge from larger models into smaller ones, and the results show that this technique effectively transfers reasoning capabilities.
Another important piece of research explored adaptive red teaming using GRPO to test and defend language models against adversarial attacks. This is crucial for understanding model vulnerabilities in deployment. This work suggests a method for iteratively improving both offensive and defensive strategies by learning from interactions with the target model.
We also saw progress in speech recognition where pretrained self-supervised models were able to recognize consonants that they had not encountered before. This is a step toward more robust audio processing. This capability builds upon the foundation of large pretraining efforts.
The study on answer-choice conformity across forty-four language models provided an interesting look at how different architectures align their responses, offering insights into model behavior under specific prompting conditions. This contrasts with the work on scaling native multimodal pre-training from scratch, which aims to build entirely new foundational models from multimodal data rather than fine-tuning existing ones.
Finally, the development of Wieszcz-XIX involved training a three point one billion word corpus of pre nineteen eighteen Polish language models from scratch. This is a massive undertaking that pushes the boundaries of training large language models on specialized historical text.
The most significant development involves the recurrent self improvement technique applied to looped language models. This matters because it suggests a path toward more robust and continuously refining AI systems. Researchers explored dynamic cross-loop on-policy distillation to enhance these models, aiming for better performance in sequential tasks. This method involves using multiple loops where the output of one loop informs and refines the next, which is a sophisticated way to teach a model to improve itself over time.
A related effort focused on large language model assisted preparation of transportation management plans for WisDOT. This is important because it shows how these models can be practically applied to complex real-world planning scenarios. They used the WisTMP system as a case study to see how LLMs could help generate these plans. This application builds on the general capability of LLMs to handle structured, domain-specific data.
Another area of work delves into cognitive thermometers using machine learning and logical complexity. This is meaningful for understanding how models process intricate information. This research attempts to map the internal state or difficulty a model encounters during reasoning by measuring its logical complexity. This provides a metric for assessing the depth of comprehension achieved by the model.
Then there is the work on lossy compressive text autoencoders, which is valuable because it addresses how to efficiently represent large amounts of text while retaining essential information. These autoencoders are designed to compress text into a smaller format without losing critical meaning, which could be useful for scaling language model applications. This technique complements the focus on efficient representation seen in other areas.
Furthermore, there's research on conversational task disambiguation over tabular data using a leakage-aware formulation and benchmark suite. This is significant for making models better at understanding context in structured conversations. This work tackles the challenge of correctly interpreting user intent when dealing with organized data inputs.
This is connected to the effort on clarify then focus, which deals with statement normalization for conversation analytics at scale. That normalization process helps standardize conversational input so that analytics can be performed reliably across a large volume of interactions. Finally, there is the plan-and-patch approach using diffusion language models for agentic planning. This is important because it moves beyond simple text generation toward creating autonomous agents capable of multi-step planning and execution.
The most significant development today involves a new method that achieves real long-term memory for artificial intelligence using a fifty million token window. This is both faster and more cost-effective than recomputing everything. This breakthrough matters because it fundamentally changes how large language models can maintain context over extended interactions, moving beyond the limitations of fixed context lengths.
This memory mechanism is built upon disentangling linguistic and paralinguistic information through routed sparse autoencoders. This technique tries to separate the actual words from the tone or manner in which they are spoken, which helps build a richer internal representation for the model.
Another important piece of work addresses how sparse attention functions, suggesting that it is actually a matrix approximation rather than simply picking from a collection of values. This clarifies how models focus on different parts of input without needing to process every single token exhaustively.
The paper on stochastic teacher intervention for agentic on-policy distillation explores how to train agents by using a teacher model to guide the learning process in real time. This is crucial for developing more capable autonomous agents that can learn through interaction rather than just static training data.
Finally, the work on storebench provides a live-commerce environment specifically designed for evaluating and training autonomous operator agents, giving these systems practical testing grounds in a simulated retail setting.
Today's papers
- CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification This method uses confidence scores to improve graph learning for classifying cancer subtypes across multiple data types. [paper] [episode]
- Predictive Multiplicity in Cell-Fate Assignment: Label-Free Rashomon Sets and the Limits of Per-Cell Certification This paper explores how to predict cell fate without labels by using sets of cells that share similar outcomes. [paper]
- Learning infinite context windows in recurrent architectures via spatial neural computing This work proposes a way for recurrent models to learn very long contexts by using spatial neural computing. [paper]
- Similar Predictive Fit but Different Latent Dynamics: Characterizing Learned Dynamical Structure in Personalized Models of Brain Disorders This study examines the different underlying dynamics learned by personalized models describing brain disorders that show similar predictive patterns. [paper]
- Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization This paper describes an agentic system that improves scientific programs by using certificates to guide software porting tasks. [paper] [episode]
- Beyond Action Entropy: Quotient-Space Exploration for Genome-Scale Metabolic Model Repair This research uses quotient space exploration to find better ways to repair large metabolic models of genomes. [paper]
- PhysFieldBench: Can Multimodal Models Understand Physical Fields? This paper tests whether multimodal models can correctly interpret physical fields using the PhysFieldBench benchmark. [paper] [episode]
- System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives This paper analyzes how system prompts condition the hidden states of open-weight models and identifies what remains effective. [paper] [episode]
- Introducing Human-Centeredness in AI-Assisted Lexicography This work focuses on making AI tools for language creation more human by incorporating human values into the process. [paper] [episode]
- Enabling Quantum Natural Language Processing for Hindi Language This research explores how to apply quantum computing techniques to improve natural language processing specifically for the Hindi language. [paper] [episode]
- Policy Learning with a Language Bottleneck This paper investigates policy learning when a language bottleneck is introduced into the model's architecture. [paper] [episode]
- Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought This method uses self-bootstrapped Prolog chain of thought to enhance the reasoning capabilities of large language models. [paper] [episode]
- InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models This paper shows how to implicitly fuse different models in a large language model by optimizing based on user preferences. [paper] [episode]
- Optimal Transport Depth Up-Scaling This paper focuses on improving the depth of optimal transport methods for data up-scaling tasks. [paper] [episode]
- An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems This study tests how robust large language models are when faced with mathematically equivalent but differently transformed problems. [paper] [episode]
- Fair-GPTQ: Bias-Aware Quantization for Large Language Models This paper develops a quantization technique for large language models that is aware of and mitigates bias. [paper] [episode]
- GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness This work proposes an efficient method for GUI agents by using a spatio-temporal aware KV cache. [paper] [episode]
- More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models This paper investigates why vision-language driving models often struggle with reasoning and planning despite good training. [paper] [episode]
- AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda This paper presents a state-of-the-art bilingual language model designed to handle the Ayurvedic language. [paper] [episode]
- Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations This research aims to create interpretable models by using meta-learning over synthetically generated models. [paper] [episode]
- LMSpell: Spell Correction with Pre-Trained Language Models This paper demonstrates how pre-trained language models can be effectively used for spell correction tasks. [paper] [episode]
- Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora This study uses translation to show how data contamination can be obscured in Arabic corpora. [paper] [episode]
- Is Peer Review Really in Decline? Analyzing Review Quality across Venues and Time This research analyzes the quality of peer reviews across different venues and over time to assess the state of peer review. [paper] [episode]
- Quantifying Retriever-Generator Alignment in RAG with Local Explanations This paper develops methods to measure how well retrieval-augmented generation systems align using local explanations. [paper] [episode]
- Foundation CAN LM: A Pretrained Language Model For Automotive CAN Data This paper introduces a pretrained language model specifically designed for automotive CAN data. [paper] [episode]
- Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching This work shows how to scale diffusion language models during test time by using reward-guided stitching techniques. [paper] [episode]
- OneMillion-Bench: How Far are Language Agents from Human Experts? This benchmark assesses the capabilities of language agents against human experts at a large scale. [paper] [episode]
- Agentic Critical Training This paper proposes a training method that involves critical evaluation for agentic systems. [paper] [episode]
- Beyond Preset Identities: Selective Stance Accommodation and Interaction Reorganisation in Generative Agent Societies This work explores how generative agents can adapt their stances and reorganize interactions based on context. [paper] [episode]
- Imperative Interference: Social Register Shapes Instruction Topology in Large Language Models This study examines how social register influences the structure of instructions given to large language models. [paper] [episode]
- Limited Stereotype Control Through Routing Reweighting in MoE Language Models This paper shows a way to limit stereotypes in Mixture-of-Experts language models by using routing reweighting. [paper] [episode]
- Domain-Adapted Retrieval for In-Context Annotation of Pedagogical Dialogue Acts This research focuses on retrieving domain-specific information to annotate dialogue acts in pedagogical contexts. [paper] [episode]
- AI Appeals Processor: A Deep Learning Approach to Automated Classification of Citizen Appeals in Government Services This paper uses deep learning to automatically classify citizen appeals in government services. [paper] [episode]
- State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning This paper introduces a parallel training method for nonlinear recurrence using the State Stream Transformer architecture. [paper] [episode]
- Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes This research investigates how low-precision computation causes loss spikes in models during the grokking process. [paper] [episode]
- SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs This paper presents a reinforcement learning approach for agents that uses evolving skill graphs to augment their skills. [paper] [episode]
- CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence This work benchmarks how well document intelligence systems attribute evidence to ensure trustworthiness. [paper] [episode]
- Uncertainty-Aware Budget Allocation for Adaptive Test-Time Reasoning This paper proposes an adaptive budget allocation strategy based on uncertainty to guide test-time reasoning in models. [paper] [episode]
- MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems This study develops methods to trace and attribute errors that occur within large language model memory systems. [paper] [episode]
- EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation This paper introduces a self-evolving rubric driven reinforcement learning approach for open-ended text generation. [paper] [episode]
- LoRi: Low-Rank Distillation for Implicit Reasoning This method uses low-rank distillation to improve the implicit reasoning capabilities of large language models. [paper] [episode]
- Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO This work presents an adaptive red teaming technique using Generative Reinforcement Policy Optimization (GRPO) to test language models. [paper] [episode]
- Pretrained self-supervised speech models can recognize unseen consonants This paper shows that pretrained self-supervised speech models can successfully recognize consonants they have never seen before. [paper] [episode]
- The One-Word Census: Answer-Choice Conformity Across 44 Language Models This study investigates the conformity of answer choices across 44 different language models on a single word question. [paper] [episode]
- Scaling Native Multimodal Pre-Training From Scratch This paper discusses the challenges and methods for scaling native multimodal pre-training from scratch. [paper] [episode]
- An Explainable Header-Centric Framework for Large-Scale Semantic Table Interpretation and Data Quality Assessment This framework provides an explainable way to interpret large semantic tables and assess data quality. [paper]
- Diffu-LoRA: A Novel Low-Rank Adaptation for Personalized Diffusion Models This paper introduces a new low-rank adaptation technique specifically designed for personalized diffusion models. [paper]
- Wieszcz-XIX: A 3.1-Billion-Word Corpus of Pre-1918 Polish and Temporally Bounded Language Models Trained From Scratch This work describes training a large language model from scratch using a massive corpus of pre-1918 Polish text. [paper]
- Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models This paper proposes dynamic cross-loop on-policy distillation to improve looped language models through recurrent self-improvement. [paper]
- Large Language Model-Assisted Preparation of Transportation Management Plans: A Case Study with WisDOT WisTMP System This paper presents a case study on using large language models to assist in preparing transportation management plans. [paper]
- Cognitive Thermometers: Machine Learning and Logical Complexity This work explores the relationship between machine learning and logical complexity using cognitive thermometers. [paper]
- Lossy Compressive Text Autoencoders This paper introduces autoencoders designed for lossy compression of text data. [paper]
- Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark Suite, and Training This study focuses on disambiguating conversational tasks using tabular data by considering leakage awareness. [paper]
- Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale This paper proposes statement normalization as a method for clarifying statements in conversation analytics at scale. [paper]
- Plan-and-Patch: Diffusion Language Models for Agentic Planning This work introduces diffusion language models specifically designed to handle agentic planning tasks. [paper]
- Grammar Concept Annotation at Scale: Deployed Fine-Tuned Small Language Models Outperform Prompted Frontier Models This paper shows that fine-tuned small language models outperform prompted frontier models in grammar concept annotation at scale. [paper]
- Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute This paper introduces a real long-term memory window that is faster and cheaper than recomputing context. [paper]
- Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders This work uses routed sparse autoencoders to separate linguistic and paralinguistic information. [paper]
- Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values This paper argues that sparse attention is an approximation technique rather than just a method for choosing from a bag of values. [paper]
- Stochastic Teacher Intervention for Agentic On-Policy Distillation This paper proposes stochastic teacher intervention as a method for on-policy distillation in agentic systems. [paper]
The papers
- Multimodal Large Language Models as Image Classifiers — The gist The MLLM classification performance depends critically on evaluation protocol and ground truth quality, showing that corrected labels can narrow the performance gap with supervised models and that model outputs must be mapped to predefined classes through protocols like [episode]
- IAD-Unify: Task-Specific Interfaces for Industrial Anomaly Understanding, Segmentation, and Generation — The gist The IAD-Unify framework proposes a dual-encoder unified model that jointly addresses anomaly segmentation, region-grounded understanding, and mask-guided generation across 24 industrial categories. How it works 1. [episode]
- More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models — The gist The planning module from an agent’s output predominantly relies on textual priors (i.e., ego state, history) as shortcuts, largely ignoring the visual context (i.e., surroundings, traffic signals) and the CoT reasoning DriveMind Dataset Creation The researchers built D [episode]
- Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping — The gist: Inverse-LLaVA proposes a multimodal architecture that inverts conventional alignment by projecting text embeddings into continuous visual representation space for fusion within intermediate transformer layers, eliminating the need for explicit alignment pretraining. [episode]
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing — The gist The authors show that by replacing multiple random permutations with a single, deterministic, and optimal permutation, they achieve a method that retains the core principles of permutation-based importance while being non-random, faster, and more stable. [episode]
- LLM-powered Query Expansion for Enhancing Boundary Prediction in Language-driven Action Localization — The gist: This paper proposes an LLM-powered query expansion method to enhance boundary prediction in language-driven action localization by generating detailed textual descriptions of action start and end boundaries and modeling boundary probabilities using semantic similarities [episode]
- Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization — The gist The certificate-driven evolutionary search (CDES) extends evolutionary search with enforceable restrictions derived from failed candidates, recorded as certificates of assumptions, checker evidence, and justified restrictions. [episode]
- Scaling Native Multimodal Pre-Training From Scratch — The gist: Native multimodal pre-training establishes compute-optimal scaling laws for transformer-based vision-language models by showing that language and multimodal objectives follow distinct allocation laws, one composition-invariant and the other composition-variant. [episode]
- DORS: Dynamic Attention Routing for Diffusion-based Object Removal in Dense Scenes — The gist: DORS proposes a training-free, plug-and-play framework that formulates object removal as semantic information flow control in the attention space to effectively suppress misleading information from similar instances while preserving structural consistency in dense scene [episode]
- Agentic Critical Training — The gist The proposed Agentic Critical Training (ACT) is a reinforcement learning paradigm that trains large language models to autonomously develop reasoning about action quality by rewarding correct action selection, leading to genuine self-reflection rather than imitation of p [episode]
- PhysMoDPO: Physically-Plausible Humanoid Motion with Preference Optimization — The gist: PhysMoDPO proposes a Direct Preference Optimization framework that integrates Whole-Body Control into the training pipeline to optimize diffusion motion generators such that their outputs are compliant with both physics and original text instructions. [episode]
- RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models — The gist: Our new model family, RT-DETRv4, achieves state-of-the-art results on COCO, attaining AP scores of 49.7/53.5/55.4/57.0 at corresponding speeds of 273/169/124/78 FPS The core problem addressed is the trade-off between lightweight models and feature representation quality [episode]
- Domain-Adapted Retrieval for In-Context Annotation of Pedagogical Dialogue Acts — The gist The authors present a domain-adapted Retrieval-Augmented Generation (RAG) pipeline for annotating pedagogical dialogue acts, achieving high Cohen’s κ scores by adapting the retrieval component rather than fine-tuning the generative model. [episode]
- Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO — The gist The AdvGRPO framework introduces a co-training method that makes Group Relative Policy Optimization (GRPO) viable for joint attacker-defender optimization using dense multi-channel rewards and decoupled advantage normalization, showing that it produces highly effective a [episode]
- EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation — The gist The EvoRubric framework is a novel single-policy co-evolutionary RL framework that eliminates reliance on static criteria and external rubric generators by unifying response generation and rubric generation under one parameterized policy. [episode]
- LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models — The gist The proposed LearnPruner framework is a twostage token pruning framework that first removes redundant vision tokens via a learnable pruning module after the vision encoder, then retains only task-relevant tokens in the LLM’s middle layer; this approach achieves 95% of [episode]
- ShapeY: A Principled Framework for Measuring Shape Recognition Capacity via Nearest-Neighbor Matching — The gist The ShapeY framework introduces a novel and principled benchmarking system designed to evaluate shape-based recognition capability in object recognition systems using nearest-neighbor matching. [episode]
- State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning — My goal is to synthesize these details into a comprehensive, high-fidelity summary that accurately reflects the core architectural innovations, training methodology, theoretical guarantees, and empirical results. [episode]
- SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs — The gist The proposed framework, SKILLGRAPH, introduces a graph-structured formulation of skill library for LLM agents where skills are connected by explicit prerequisite, enhancement, and co-occurrence relations. [episode]
- Optimal Transport Depth Up-Scaling — The gist Scaling Large Language Models (LLMs) yields performance gains but incurs substantial training costs, and this paper proposes Optimal Transport Depth Up-Scaling (OpT-DeUS) to mitigate neuron permutation mismatch between layers by aligning and fusing Transformer blocks usi [episode]
- Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching — The gist The proposed framework introduces Stitching Noisy Diffusion Thoughts, a self-consistency mechanism that converts cheap diffusion-sampled reasoning into a reusable pool of step-level candidates to improve test-time scaling for large language models. [episode]
- World-Ego Modeling for Embodied Video Generation in Long-Horizon Navigation-Manipulation Tasks — The gist The World-Ego Modeling paradigm decomposes embodied video prediction into persistent world regularities and robot-centric ego dynamics, which addresses long-horizon degradation in hybrid navigation-manipulation tasks. [episode]
- Uncertainty-Aware Budget Allocation for Adaptive Test-Time Reasoning — The gist: Uncertainty-Aware Budget Allocation (UAB) proposes a two-phase inference framework that reallocates a fixed sampling budget based on per-question uncertainty estimated at no additional inference cost to maximize accuracy across multiple questions. [episode]
- MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems — As a fastidious and diligent researcher, I have thoroughly analyzed both provided texts concerning the paper "MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems." The information is rich, detailing a novel framework for debugging LLM memory systems, i [episode]
- 3BASiL: An Algorithmic Framework for Sparse plus Low-Rank Compression of LLMs — The gist The 3BASiL framework introduces an efficient one-shot post-training method for Sparse plus Low-Rank (S + LR) decomposition of Large Language Models (LLMs) that addresses performance degradation seen in existing methods by combining a novel 3-Block Alternating Direction M [episode]
- CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence — The gist The CiteVQA benchmark introduces an evaluation framework that requires models to return element-level bounding-box citations alongside each answer, evaluating both jointly. How it works 1. [episode]
- Stratified Multi-View Aggregation for Score Distillation — The gist The MV-SDI framework reduces gradient variance in score distillation by aggregating distillation gradients from K views per step, which yields significant improvements in asset quality and optimization speed without retraining or adding memory How it works MV-SDI aggrega [episode]
- Learned Image Compression for Vision-Language-Action Models — The gist: SPARC is a learned image compression framework tailored for VLA systems that adaptively allocates bitrate across spatial regions according to their contribution to downstream control, consistently achieving stronger control performance than conventional codecs under the [episode]
- Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs — The gist The method introduces FIRE-MPO, a fine-grained, on-policy alignment framework designed to address the unique challenges of medical Vision-Language Models by utilizing a bidirectional token-wise KL regularizer and a visual-contrastive grounding objective. [episode]
- Trust Region Q Adjoint Matching — The gist The Trust Region Q-Adjoint Matching (TRQAM) is a stable off-policy fine-tuning algorithm that adaptively controls path-space KL with pretrained flow policies through projected dual descent, achieving an overall offline RL success rate of 68% on 50 OGBench tasks Why it ma [episode]
- GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness — The gist The GUI-KV method introduces a plug-and-play KV cache compression technique for GUI agents that exploits spatio-temporal redundancy to achieve near–full-cache accuracy with modest budgets and significant reductions in decoding FLOPs. [episode]
- ASV3D: Adapting Diffusion-Based Single-View 3D Reconstruction with Extra Imagery — The gist: ASV3D is a framework for adapting single-view 3D object reconstruction to test-time data with support from one additional image, proposing two adaptation strategies: zero-shot adaptation and optimised adaptation, which consistently improve reconstruction accuracy and ro [episode]
- PixVL: Self-Supervised Training of Pixel-Level MLLMs via a Unified Mask--Text Consistency Cycle — The gist The authors propose PixVL, a self-supervised post-training framework that introduces a unified Mask–Text Consistency Cycle to enable pixel-level MLLMs to generate and self-verify regional descriptions and learn from unlabeled data. [episode]
- Constrained Dynamic Gaussian Splatting — The gist The Constrained Dynamic Gaussian Splatting framework reformulates dynamic scene reconstruction as a budget-constrained optimization problem to enforce a strict, user-defined Gaussian budget during training. [episode]
- Prognostics for Autonomous Deep-Space Habitat Health Management under Multiple Unknown Failure Modes — The gist: We propose an unsupervised prognostics framework for Remaining Useful Life (RUL) prediction that jointly identifies latent failure modes and selects informative sensors using unlabeled run-to-failure data. [episode]
- OneMillion-Bench: How Far are Language Agents from Human Experts? — The first text (A) is a technical description of a new benchmark, OneMillion-Bench (1M-Bench), designed to evaluate the capabilities of large language models (LLMs) as autonomous agents in complex, real-world professional scenarios. [episode]
- Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts — The gist: Multimodal benchmark designers should proactively try to “game” their own benchmarks first as a key step in the development lifecycle—adopting rigorous diagnostic and debiasing procedures to systematically identify, quantify, and mitigate non-visual biases. [episode]
- WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models — The gist The WorldCraft framework expands interactive video world models from camera navigation to object-level trajectory actions, enabling users to manipulate selected objects while continuing camera navigation. [episode]
- DDL: A Large-Scale Dataset for Deepfake Detection and Localization in Diversified Real-World Scenarios — The DDL dataset introduces a large-scale, diverse, multi-modal deepfake detection and localization benchmark to address the limitations of existing datasets by providing fine-grained spatial and temporal annotations for complex real-world forgery scenarios. [episode]
- Towards Reasonable Concept Bottleneck Models — The gist: We propose Concept REAsoning Models (CREAM), a novel framework for Concept Bottleneck Models (CBMs) that explicitly encodes prior knowledge about concept-concept and concept-task relationships through a reasoning graph, enabling interpretable and effective prediction ev [episode]
- Policy Learning with a Language Bottleneck — The gist: Policy Learning with a Language Bottleneck (PLLB) is a framework that enables AI agents to generate linguistic rules that capture high-level strategies underlying rewarding behaviors, which improves policy learning interpretability and generalization across diverse task [episode]
- CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification — The gist: CMGL proposes a two-stage framework that estimates per-sample modality reliability through evidential deep learning and uses these frozen confidence scores to guide cross-omics fusion and graph construction for cancer subtype classification. [episode]
- V-ECE: Estimating General Expected Calibration Errors — The gist: This paper introduces an extension to variational frameworks for estimating calibration errors that can cover any binary or multiclass Lp calibration error, offering benefits like separating over- and under-confidence and avoiding overestimation through cross-validation [episode]
- Quantum Multi-Armed Bandits and Linear Bandits: Lower Bounds and Algorithms — The gist The authors prove minimax lower bounds for quantum multi-armed bandits and linear bandits, resolving prior questions regarding T-independent regret and improving dimension dependence in finite-action settings. [episode]
- 4D-GSW: Kinematic-Aware Spatio-Temporal Consistent Watermarking for 4D Gaussian Splatting — The gist: 4D-GSW is a kinematic-aware watermarking framework designed to embed robust copyright information while preserving high spatio-temporal consistency by adaptively gating watermark gradients based on SpatioTemporal Curvature (STC) and formulating the embedding as a joint [episode]
- Road Maps as Free Geometric Priors: Weather-Invariant Drone Geo-Localization with GeoFuse — The gist: GeoFuse introduces a cross-modal fusion framework that integrates precisely aligned road map tiles with satellite imagery to yield more discriminative and weather-resilient representations for drone geo-localization. [episode]
- Concomitant DAG Learning: On the Roles of Noise Adaptivity, Sparsity, and Non-negativity — The gist The paper introduces CoLiDE, a novel convex score function for sparsity-aware DAG inference that jointly estimates the DAG adjacency matrix and exogenous noise levels, offering robustness against heteroscedastic noise profiles.<ref:2605.23537#pg11> Problem Formulation an [episode]
- Virtual Smart Metering in District Heating Networks via Heterogeneous Spatial-Temporal Graph Neural Networks — The gist: A heterogeneous spatial-temporal graph neural network (HSTGNN) is proposed to construct virtual smart heat meters by incorporating functional relationships and dedicated branches for flow, temperature, and pressure measurements to achieve joint modeling of cross-variabl [episode]
- Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning — The gist The Similarity-as-Evidence (SaE) framework calibrates text–image similarities by introducing a Similarity Evidence Head (SEH), which reinterprets the similarity vector as evidence and parameterizes a Dirichlet distribution over labels, thereby mitigating overconfidence [episode]
- Reward Observability and the Limits of Offline Checkpoint Selection in RSSM World Models — The gist The Composite Reward Observability Fraction (CROF) is a single-number offline checkpointselection score that combines reward/observation subspace-alignment score with three structural scores, and it selects world models that train a model-based A2C policy that beats a fa [episode]
- Imperative Interference: Social Register Shapes Instruction Topology in Large Language Models — The gist The imperative mood carries different obligatory force across speech communities, leading to an inversion in instruction interaction topology between English and Spanish. [episode]
- Accurate and Efficient Object Pose Estimation via the Aggregation of Diffusion Features — The gist: The authors propose three distinct architectures that can effectively capture and aggregate diffusion features of different granularity to greatly improve the generalizability of object pose estimation. [episode]
- MatLat: Material Latent Space for PBR Texture Generation — The gist The authors propose MATLAT, a generative framework that learns a material latent space to produce high-quality PBR textures by leveraging pretrained image diffusion models and addressing domain gaps through latent-space adaptation and locality regularization. [episode]
- The One-Word Census: Answer-Choice Conformity Across 44 Language Models — The gist The field converges when language models are asked to choose one answer from a large space of options, and this convergence varies structurally across different model types and generations. [episode]
- ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning — The gist The ERASE framework is an adaptive two-stage vision token pruning framework that identifies and retains salient tokens through pruning strategies adaptive to image complexity, demonstrating superior efficiency-performance trade-off over existing methods Motivation Recent [episode]
- InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models — The gist The current few fusion methods on PA phase, like WRPO, simplify the process by utilizing only response outputs from source models while discarding their probability information InfiFPO replaces the reference model in Direct Preference Optimization (DPO) with a fused sour [episode]
- SAM 3D Animal: Promptable Animal 3D Reconstruction from Images in the Wild — The gist The SAM 3D Animal framework introduces a promptable method for joint 3D reconstruction of multiple animals from a single image, addressing challenges in crowded and occluded scenes through flexible prompts like keypoints and masks. [episode]
- Joint Architecture-Token-Bitwidth Multi-Axis Optimization of Vision Transformers for Semiconductor IC Packaging — The gist The proposed multi-axis framework achieves more than 10× improvement in throughput along with over 10× reductions in parameter count, FLOPs, and energy consumption while maintaining the required accuracy on the downstream industrial task How it works The proposed holis [episode]
- Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning — The gist The authors show that prompting MLLMs to reason before prediction does not consistently help, and can even reduce persuasiveness prediction performance, suggesting that naively generated rationales are unreliable signals for this task. [episode]
- Dynamic Free-Rider Detection in Cross-Silo Federated Learning via Simulated Attack Patterns — The gist The proposed method S2-WEF enables dynamic detection of clients that transition into free-riders during training without proxy datasets or pretraining, and it achieves higher robustness than existing approaches across three datasets and five attack types Problem and Moti [episode]
- Memory by Design: Probabilistic Sequence Layers — The gist The design-model framework introduces a way to derive efficient recurrent sequence maps from explicit assumptions about memory, where a design model writes evidence into memory by exact Bayesian filtering and a query-dependent readout produces a predictive distribution w [episode]
- Uncertainty Quantification in Federated Granger Causality Learning — The gist The authors address how uncertainty propagates through Federated Granger Causality (FedGC) to provide a principled basis for identifying reliable cross-client interactions in vertically partitioned data settings. [episode]
- AI Appeals Processor: A Deep Learning Approach to Automated Classification of Citizen Appeals in Government Services — The gist The AI Appeals Processor presents a microservice-based system that integrates natural language processing and deep learning techniques for automated classification and routing of citizen appeals, achieving 78% classification accuracy with Word2Vec+LSTM while reducing pro [episode]
- APEX: Assumption-free Projection-based Embedding eXamination Metric for Image Quality Assessment — The gist The APEX framework introduces an assumption-free projection-based embedding examination metric for image quality assessment by leveraging the Sliced Wasserstein Distance with CLIP and DINOv2 embeddings to overcome limitations in existing metrics. How it works 1. [episode]
- An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems — The gist: The proposed Generalization–and–Perturbation (GAP) framework introduces a systematic methodology to assess LLMs’ mathematical-reasoning robustness by stress-testing them on mathematically equivalent but linguistically and parametrically varied advanced math proble [episode]
- Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency — The gist The proposed Spatial Alignment via Geometric Evolution (SAGE) framework enforces logical consistency in Vision-Language Models through geometric and linguistic duality operations to improve their spatial reasoning capabilities. How it works 1. [episode]
- AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda — The gist The model AyurParam introduces a domain-specialized, bilingual language model for Ayurveda that fine-tuned from Param-1-2.9B using an extensive, expertly curated Ayurveda dataset to achieve state-of-the-art performance on BhashaBench-Ayur. [episode]
- Introducing Human-Centeredness in AI-Assisted Lexicography — The gist This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography, arguing that AI should augment rather than replace lexicographers by focusing on four interrelated dimensions: the augmented lexicographer, the sociotechnical cont [episode]
- BiomedXPro: Prompt Optimization for Explainable Diagnosis with Biomedical Vision Language Models — The gist The clinical adoption of biomedical vision-language models is hindered by prompt optimization techniques that produce either uninterpretable latent vectors or single textual prompts. [episode]
- Networks with Finite VC Dimension: Pro and Contra — The gist: Finite VC dimension is desirable for uniform convergence of empirical errors but may not be desirable for approximation of functions drawn from a probability distribution modeling the likelihood that they occur in a given type of application. [episode]
- Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought — The THOUGHT-LIKE-PRO framework introduces a novel learning framework designed to enhance the logical reasoning capabilities of Large Language Models by imitating verified reasoning processes from a symbolic Prolog logic engine. [episode]
- Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora — The gist The translation into Arabic can suppress conventional contamination indicators while still allowing models to benefit from exposure, particularly those with stronger Arabic capabilities. [episode]
- HyVIC: A Metric-Driven Spatio-Spectral Hyperspectral Image Compression Architecture Based on Variational Autoencoders — The gist The proposed model, HyVIC, introduces a configurable spatio-spectral Variational Autoencoder (VAE) architecture designed to effectively leverage spatio-spectral redundancies in hyperspectral data by allowing independent control over spatial and spectral feature learning. [episode]
- Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes — The gist Deep neural networks exhibit periodic loss spikes during unregularized longterm training, a phenomenon known as the “Slingshot Mechanism” [1]. How it works 1. [episode]
- Learning to Emulate Chaos: Adversarial Optimal Transport Regularization — The gist The authors propose a family of adversarial optimal transport objectives to jointly learn high-quality summary statistics and a physically consistent emulator from a single noisy trajectory, which significantly improves long-term statistical fidelity for emulating chaoti [episode]
- Pretrained self-supervised speech models can recognize unseen consonants — The gist Pretrained self-supervised speech models can recognize click consonants as accurately as other speech sounds, suggesting that self-supervision enables generalization across human speech sounds including rare phonemes. [episode]
- Quantifying Retriever-Generator Alignment in RAG with Local Explanations — The gist The RAG-E framework presents an end-to-end explainability framework that quantifies retriever-generator alignment through mathematically grounded attribution methods, revealing substantial misalignment between these components in RAG systems. [episode]
- PhysFieldBench: Can Multimodal Models Understand Physical Fields? — The first text appears to be an excerpt detailing the methodology, evaluation protocol, and specific task examples of PhysFieldBench, while the second text is a meta-commentary indicating that the provided input *is not* the full paper summary. [episode]
- Discovering Global False Negatives On the Fly for Self-supervised Contrastive Learning — The gist: GLOFND is an optimization-based approach that automatically learns on the fly thresholds for each anchor data to identify its false negatives during training, addressing a critical issue in self-supervised contrastive learning where negative pairs with similar semantics [episode]
- Conformal Data Contamination Tests for In-distribution Data Acquisition — The gist The proposed work introduces a distribution-free, contamination-aware data-sharing framework that uses novel two-sample testing procedures, termed conformal data contamination tests, to identify external data agents whose data is most valuable for model personalization. [episode]
- The Sample Complexity of Membership Inference and Privacy Auditing — The gist: In simple, natural settings for Gaussian mean estimation, any successful membership-inference attack requires a sample complexity of at least omega(n) samples, which is many more than what training algorithms use Main Conceptual Result The main conceptual result shows t [episode]
- Fair-GPTQ: Bias-Aware Quantization for Large Language Models — The gist The first quantization method explicitly designed to reduce unfairness in large language models, Fair-GPTQ, introduces explicit group-fairness constraints into the quantization objective to mitigate bias during compression. [episode]
- HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End-to-end Autonomous Driving — The gist: HENet++ achieves state-of-the-art end-to-end multi-task perception performance on the nuScenes dataset while attaining the lowest collision rate on the nuScenes end-to-end autonomous driving benchmark. [episode]
- SRUG: A Fusion-Driven Generator Network for Medical Image Translation — The gist The proposed SRU-Pix2Pix framework enhances image generation quality and structural fidelity for medical image translation under few-shot learning conditions with fewer than 500 images. How it works 1. [episode]
- LMSpell: Spell Correction with Pre-Trained Language Models — The gist The first empirical study on Large Language Models (LLMs) for spell correction reveals that LLMs outperform their encoder-based and encoder-decoder counterparts when the fine-tuning dataset is large, even in languages where the LLM was not pre-trained. [episode]
- Image Recognition with Vision and Language Embeddings of VLMs — The gist The authors conduct a comprehensive evaluation of both language-guided and vision-only image classification with dual-encoder VLMs, showing that language and vision offer complementary strengths, with some classes favouring textual prompts and others better handled by vi [episode]
- Learning Projection-Aware 360-Degree Image Rectification via Dual-Projection Fusion — The gist: This study presents a dual-stream angle-aware generation network that jointly estimates camera inclination angles and reconstructs upright panoramic images by adaptively fusing local spatial features from equirectangular projections with global contextual cues from cube [episode]
- Bi-temporal Image-driven Acute Stroke Evolution Analysis — The gist The proposed framework explores differences in hypoperfused areas that end up being either infarcted or re-perfused and can be seen as tissue characterization Bi-temporal Analysis Framework The work proposes a bi-temporal analysis framework that characterizes ischemic ti [episode]
- Boosting the Local Invariance for Better Adversarial Transferability — The gist: adversarial perturbation often exhibits poor translation invariance for a given clean image and model, which is attributed to local invariance <ref:2503.06140#pg2>. [episode]
- System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives — The gist The central empirical finding is a qualitative reorganization of the geometric encoding of identity across the instruction-tuning boundary: in the base-weight Gemma-4-E4B, the identity fingerprint is encoded predominantly in the direction of hidden-state vectors, whereas [episode]
- Equilibrium Distribution for t-Distributed Stochastic Neighbor Embedding with Generalized Kernels — The gist The work proves that t-SNE converges to an equilibrium distribution for a wide range of input and output kernels under certain conditions as the number of data points diverges<ref:2505.24311#pg8>. [episode]
- Beyond Preset Identities: Selective Stance Accommodation and Interaction Reorganisation in Generative Agent Societies — The gist: Agents exhibit endogenous stances that override preset identities, and their ability to reconstruct social structures through language practices depends critically on aligning human interventions with these emergent cognitive tendencies. [episode]
- LoRi: Low-Rank Distillation for Implicit Reasoning — The gist: LoRi proposes a low-rank distillation framework that transfers reasoning from explicit chain-of-thought into compact implicit latent processes by aligning teacher and student trajectories in a shared low-rank subspace using first- and second-order statistics. [episode]
- IS-Diff: Improving Diffusion-Based Inpainting with Better Initial Seed — The gist: IS-Diff proposes a training-free approach to improve diffusion-based image inpainting by using initial seeds sampled from unmasked areas and incorporating a dynamic selective refinement mechanism to adjust initialization strength based on distributional cross-entropy. [episode]
- Self-sufficient Independent Component Analysis for Demixing Flows — The gist Learning disentangled features from data using non-linear Independent Component Analysis (ICA) via KL Minimizing Flows. [episode]
- Is Peer Review Really in Decline? Analyzing Review Quality across Venues and Time — The gist The cross-temporal analysis reveals no consistent decline in median review quality across venues and years. [episode]
- Enabling Quantum Natural Language Processing for Hindi Language — The gist The research proposes enabling Quantum Natural Language Processing for Hindi by developing parameterized quantum circuits from Hindi sentences using pregroup grammar and the DisCoCat framework. [episode]
- GPF-Net: Gated Progressive Fusion Learning for Polyp Re-Identification — The gist The Gated Progressive Fusion Network (GPF-Net) is a novel architecture that utilizes a gated progressive fusion strategy to selectively fuse features from multiple levels, achieving layer-wise refinement of semantic information for colonoscopic polyp re-identification. [episode]
- Limited Stereotype Control Through Routing Reweighting in MoE Language Models — The gist: Demographic routing sensitivity is universal across five MoE architectures, but stereotype controllability is not, as routing-level preference modulation does not reliably transfer to decoded generation behavior. [episode]
- OPERA: Object Perception Enhances Single-view 3D Reconstruction — The gist: Learnt object perception can significantly enhance 3D reconstruction by explicitly injecting perceptual signals from pretrained models into existing reconstruction pipelines to drive more accurate and semantically consistent results. [episode]
- A Survey on Industrial Anomaly Synthesis — The gist: This survey comprehensively reviews anomaly synthesis methodologies, introducing the first industrial anomaly synthesis (IAS) taxonomy and exploring cross-modality synthesis and large-scale Vision Language Models (VLM) to boost IAS. [episode]
- Foundation CAN LM: A Pretrained Language Model For Automotive CAN Data — The gist The foundation CAN model demonstrates multi-objective downstream generalization using a single pretrained backbone by treating CAN data as a language and adapting it to various automotive tasks.; <ref:2602.00866#pg2> How it works The core of the approach involves treatin [episode]
- Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations — The gist The work introduces an efficient, scalable method for generating synthetic pre-training data to enable meta-learning of decision trees. Meta-learning Workflow 1. [episode]
- Graphons of Line Graphs — The gist: The method involves mapping original graphs to their line graphs and showing that graphs satisfying a particular property, which we call the square-degree property are sparse, but give rise to dense line graphs, enabling the use of results on graph limits of dense graph [episode]
- Heterogeneous-Modal Unsupervised Domain Adaptation via Latent Space Bridging — The gist: The proposed Latent Space Bridging (LSB) method introduces a novel setting called Heterogeneous-Modal Unsupervised Domain Adaptation (HMUDA) to enable knowledge transfer between completely different modalities by leveraging an unlabeled bridge domain containing samples [episode]
- Compact Multi-level-prior Tensor Representation for Hyperspectral Image Super-resolution — The gist: This paper presents a novel hyperspectral super-resolution model that compactly characterizes multi-level priors of hyperspectral images within the tensor framework, facilitating convergence-guaranteed iterative optimization. [episode]
- Comparing Object Detection Models for Electrical Substation Component Mapping — The gist The research trains and compares three object detection models (YOLOv8, YOLOv11, RF-DETR) to map key electrical substation components in US images, aiming to find the most reliable method for automated mapping. [episode]
- FreshMem: Brain-Inspired Frequency-Space Hybrid Memory for Streaming Video Understanding — The gist The FreshMem framework proposes a Frequency-Space Hybrid Memory network inspired by brain mechanisms to reconcile short-term fidelity with long-term coherence for streaming video understanding. [episode]
- An Axiomatic Assessment of Entropy- and Variance-based Uncertainty Quantification in Regression — The gist: This work provides a formal way of representing uncertainty in continuous space using a general parametric formulation and proposes axioms to rigorously assess total, aleatoric, and epistemic uncertainty measures in regression settings Uncertainty Representation in Regr [episode]
- Gaze Estimation for Human-Robot Interaction: Analysis Using the NICO Platform — The gist: This paper evaluates four state-of-the-art gaze estimation models in a shared workspace scenario using an annotated dataset collected with the NICO robotic platform, finding that while angular errors are comparable to general benchmarks, distance errors are limited to a [episode]
- When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization —
- Harness Compilation: Which Decisions Should a Small Vision-Language Model Keep? —
- MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification —
- Breaking the Group Size Barrier: Parameter-Efficient Group Dance Generation with Chain-of-Dancers —
- Attributing HOW, Not Just WHICH: Counterfactual Response Trajectories for Diffusion Models —
- TAP3D: Thermal-Assisted 3D Human Point Clouds —
- Read What Matters: Query-Adaptive Quantization for KV Caches —
- Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals —
- V-CoLA: Vision Token Compression with Linear Attention —
- LR-V2X: Loss-Resilient Collaborative Perception under Low-Bandwidth Communication —
- Gated Memory: Admission-Controlled Memory Formation for Conversational AI —
- Phonological Interference in Multilingual Speech Models —
- When Scene Text Hijacks the Scene: Uncovering, Exploiting, and Mitigating Rendered-Text Semantic Leakage in Image Generation Models —
- REMORY: Learning Residual Memory for Context Compaction —
- When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better —
- Spatial-Frequency-Aware Implicit Neural Representation of Multidimensional Signals via MLP-KAN Fusion —
- iCATS: Fast Video Generation via Interaction-Aware Sparse Attention and Timestep-Adaptive Sparsity —
- Efficient Multi-Granularity Knowledge Transfer for Radiology Report Generation —
- BeliefScope: Diagnosing Evidence-Driven Revision and Pressure-Induced Shifts in Large Language Models —
- Collaboratively Guided Adversarial Robust Distillation with Teacher-Favorable Examples —
- From Geometry to Generalization: Why Row Normalization Can Beat Adam and Muon —
- FloorSAV: Elucidating Spatial Audio-Visual Context with 2D Floormap for AV-LLMs —
- From Retrieval to Reconstruction: Constructing Evolvable Cognitive Memory for Long-Term Dialogue —
- Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers —
- MetaEncoder: Exploring the Limit of Bi-Encoders for Multimodal System One Decision Making with Natural Language Interface —
- It's Always 10:10: Reference Images Break a Bias That Prompts Only Dent —
- PathLang: A Language-Centered Benchmark for Vision-Language Models in Computational Pathology —
- ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery —
- Point-Focused Attention Meets Context-Scan State Space: Robust Biological Visual Perception for Point Cloud Representation —
- EgoPhys: Estimating Peak Contact Force and Mechanical Work from Egocentric Manipulation Video —
- Deception by Omission: Language Models Knowingly Hide Their Mistakes —
- RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty —
- AdaptEvo: Adaptive Agent Learning with Evolving Supervision —
- Adaptive Adversarial Augmentation for Controllable Face Synthesis —
- Bernoulli Flow Models: Self-Consistent Generative Modeling for Binary Data —
- UniData: Universal Multimodal Instruction Generation Pipeline —
- FlyMark: Training-Free Invisible Watermarking of 3D Gaussian Splatting via a Fruit Fly Connectome —
- SignRAG: Unified Retrieval-Augmented Gloss-Free Sign Language Translation —
- From a Prompt to Repertoires: Evolving Functional REpertoires Enable LLM Continual Learning —
- Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder —
- CRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion Decoding —
- FastJEV: Understanding Redundancy for Compact JEV Inference —
- EvoKnow: Continual Knowledge Evolution for AI-Generated Image Detection —
- Fact over Fiction: Detection of Pathological Hallucinations in Sinhala-to-English Neural Machine Translation —
- CoPoE: Multimodal Fusion via Decomposable Disease-Coordinate Product-of-Experts for Missing-Modality Alzheimer's Diagnosis —
- GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA —
- Beyond Distributional Fidelity: Causal-Penalized Diffusion for Synthetic Tabular Data —
- Fresco++: Frequency-Guided and Canonical-Consistent Optimization for Fine-Grained Head Avatar Modeling —
- Missing Modality-Aware Calibration for Trustworthy Brain Tumor Segmentation —
- BioBigBird: A Sparse Attention Model for Long-Range Dependency Processing in Biomedical Text —
- EchoDiST: Self-distillation-based joint learning for diffusion-conditioned echocardiographic myocardial motion estimation —
- DLC: A Metric-Guided Dynamic Loss Controller for Multi-Objective Training —
- Adversarial Cues in Decision Models Used as Judges: The Role of Request Presentation —
- Learning to Retrieve: Internalizing Memory Retrieval for Video World Models —
- SAIL: Scientific Agentic Intelligence via a Science-Aware Loop —
- Feature Space Adaptation for Effortless Gaussian Process Flows —
- ProtoSemImage: Image-Valued Prototypes with Deformable Row Alignment for Interpretable Document Classification —
- Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents —
- SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders —
- Rare Gate Disagreements Can Limit Plasticity: When Gradient Flow Mispredicts Finite-Batch SGD —
- Conditional Residual Prediction: Improving Autoregressive Video Diffusion without a Bidirectional Teacher —
- PSI-SINDy: Post-Selection Inference for Sparse Identification of Nonlinear Dynamics —
- Beyond Resolution: Object-to-Image Ratio Mismatch in Instance Retrieval —
- Stop My Dancing! Understanding, Detecting and Attributing Motion-Aware Deepfake Videos —
- Parametric Trajectory Distillation for Few-Step Video Generation —
- Beyond Sequences: Distilling Structured Decision Memory for LLM Recommendation —
- Does Modern Standard Arabic (MSA) Dominate Arabic Dialects in LLMs? A Representation-Level Analysis —
- Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards —
- When Can You Prune Your Network? A Study of Intermediate Neurons in Multilingual Speech Parsing —
- MSGAT: Multi-Head Spiking Graph Attention with Similarity-Space Fusion for Image-Text Retrieval —
- HAND: A Biologically-Inspired Activation Function that Improves Generalisation and Sample Efficiency in Image Classification —
- LAIR-Net: Leaky Alignment-Impulse Residual Networks for Tabular Regression —
- Constitutional Gating and Deterministic Recovery for Multi-Agent LLM Negotiation: Ablations Against a Stateful Adversarial Gatekeeper —
- Learning the Loop, Not Just the Page: Execution-Grounded Loop Learning for Web Generation —
- Prosody-to-Text: Predicting text from low-pass filtered speech —
- OX-NeRF: 3D X-ray Tomography Reconstruction from Sparse Views Using Implicit Neural Representation —
- SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction —
- Hankel Subspace Self-Supervised Learning for Parallel MRI Reconstruction —
- Incremental Open-Ended Deep Research with Structured Harness —
- PAM-ToD: Plug-and-Play Appearance Modeling for Cross-Time-of-Day 3D Gaussian Splatting —
- Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts —
- SV-TAD: Native Sparse Convs for Efficient Temporal Action Detection —
- Embedding-Bias in Conditional Independence Testing —
- Adapting English Quality Classifiers for Multilingual LLM Pretraining Data Selection —
- Measuring Cultural Alignment Beyond the Average: A Framework for Evaluating Maternal-Health LLM Interactions in Indian Contexts —
- Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with Prolog —
- Large Language Model Turnover Undermines Screening for Artificial Intelligence-Assisted Scientific Writing —
- S cubed Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-Localization —
- Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence —
- PointVGGT: Zero-Shot Multiview RGB-D Point Cloud Registration with Visual Geometry Foundation Priors —
- TAM: Task-Aware Memory Distillation for Efficient Spatiotemporal Prediction —
- Beyond Action Entropy: Quotient-Space Exploration for Genome-Scale Metabolic Model Repair —
- Minimax Gaussian Mechanisms for Continual Machine Unlearning —
- UXBench Pro: Benchmarking Personalized User Experience in Multi-Turn Dialogue Interactions —
- Revisiting Handcrafted Minutiae Detection: A Simple and Effective Open Source Baseline for Modern Fingerprint Workflows —
- Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study —
- Harness Evolution Hits a Ceiling: When Weight Training Should Begin —
- DisFace3DNet: Explainable Facial Attractiveness Prediction via 3D Component Disentanglement —
- DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation —
- sigma Transfer: Uncertainty Transfer from Small to Large Networks under mu P —
- Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation —
- VESSI - VLM-Enhanced Support for Surveillance and Investigations —
- TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation —
- HI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High Resolution —
- Structured Sentiment Analysis Using Sequence Labeling as Dependency Graph Parsing —
- Internalizer: Portable Context-to-Parameter Mapping for Very Large Language Models —
- 4-Tensor Attention Model for Semantic Physical Reality —
- MultiWorldBench: Do Independently Controlled Views Describe One Shared World? —
- Onboard Marine Anomaly Detection on sat-2: From Simulation-Based Development to In-Orbit Demonstration —
- Towards Unified Evaluation of Prompt Enhancers for Video Generation —
- Streaming-Aware Diffusion for Real-Time Video Super-Resolution via Cross-Step Attention —
- Dino Forcing Flow Models: Do not denoise what you can predict —
- Memory Forcing: Attendable Mid-Horizon History for Streaming Video Generation —
- Thinking Inertia: LLMs Keep Thinking When Told Not To —
- Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation —
- From Video Clips to Creation Trajectory: Sora100K for AI-Native Video Creation —
- RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing —
- DPPM: Dual-Path Parametric Memory for Personalized Language Models —
- Skill-V: Verifiable Self-Evolving Skill Library for Interactive Agents —
- From Sparse Representations to Behavioral Insights for Multimodal Depression Assessment —
- Easy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching misses —
- Phase-aware video generation for physics-grounded dynamics and interactions —
- Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks —
- Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must —
- Seek-and-View Reasoning for Multi-View Spatial Understanding —
- Fast Pose Tracking of Rigid Objects with Compact Pose Graph Optimization —
- From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction —
- From Suppression to Repair: Mitigating Object Hallucination in Large Vision-Language Models via Localized Distribution Alignment —
- Detecting Spin in Clinical Trials with Large Language Models —
- Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion —
- VEDJE: Video-Efficient Discriminative Joint Encoder for Scalable Video-Text Retrieval —
- GRPODropout: Less is More for Online Reinforcement Learning Rollouts —
- Pose-Free Feed-Forward 3D Inpainting via Learnable Mask Attention and Support Token Refinement —
- Conditional Kernel Stein Discrepancy —
- Learning structured linear dynamical systems from missing observations —
- Relative Patch Response Learning for Generalizable AI-Generated Image Detection —
- Forms of LLM-Integrated Applications from LLM-Chats to Autonomous AI Agent System —
- Can Decision Models Understand Stance? Evaluating Jev Against General-Purpose LLMs —
- RobustLDS: Learning linear dynamical systems under adversarial corruptions —
- Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations —
- Not Every Change Is Necessary: Recoverable Drift in Large Language Model Unlearning —
- Event-Centric Memory with Query-Aware Graph Augmentation for Long-Term Conversational Agents —
- Revisiting Identity and Spectra Dispersion in Media-Bridged Time Series Forecasting: Linking Multivariate Signals and Narrative Flows —
- Does Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive Encoders —
- Right Screen, Wrong Transition: World Models as Verifiers for GUI Agents —
- Score-Based Learning of Cluster DAGs from Interventions —
- When History Helps and Hurts: Selective History Use across Multimodal Turns —
- Look Where You Can: Active View Selection for CAD Reconstruction under Occlusion —
- MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement —
- MindFlow: Mind Supernet Powered Thinking Flows for Research Idea Innovation —
- Learning Which Correspondences to Trust: Confidence-Weighted Event-Camera Localization in LiDAR Maps —
- Efficient quadratic entropy with distance sketches —
- Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks —
- FearCaut-Qwen: Affective Steering in a Vision-Language Model Shifts the Decision Criterion for Hazard Assessment —
- DataVista: Diagnosing Multimodal LLMs on Data Video Understanding —
- Agentic-TTT: Training test-time policy for test-time training —
- Examining Social Attribution in LLM Reasoning: A Theory-Guided Probing Methodology —
- InterviewPlayground: A Simulation Environment for Evaluating AI Interviewers —
- Natural Language to First-Order Logic LLM-based Autoformalization —
- Efficient and Generalizable Archetypal Analysis for Discrete Data —
- Do Not Train Away Uncertainty: Early Uncertainty Anchored Calibration —
- Diffusion Removes Langevin's Conditioning Dependence: A Sharp Gaussian Analysis —
- Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal Reasoning —
- When Should Agents Think? Adaptive Reasoning via Cross-Turn Estimation —
- ILM: An AI-Powered Storytelling Educational Tool —
- LIVIN: Benchmarking Spatial and Embodied Intelligence in Digital Twins of Lived-In Homes —
- Exploiting Gradients in Bayesian Inference of Expensive Simulators —
- Perception Test 2026: Challenge Summary and Extension to City-scale Audio-Visual Reasoning —
- All Verdicts are Not Equal: Rethinking LLM Judge Reliability —
- Differentiable Systematic Resampling for Variational Sequential Monte Carlo —
- DVLA-RL++: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning —
- Few-Step Generation via Data-Space Iteration —
- VINCIE-NExT: Unlocking Video Editing from Images via In-Context Modeling —
- ContourVLA: A Closed-Loop Perception-Action Contour Policy for Generalized Referring Expression Segmentation —
- Credal Machine Learning for Risk-Averse Decision Making —
- A persistent accuracy ceiling in automated verbal deception detection —
- LVS: Local View Synthesis from Relative Camera Pose by Reusing Previous Views —
- Rehearse Everything, Remember Nothing: Attic-KV Rehearses What Will Be Read —
- Language-Specific Effects of Tokenizer Choice in Multilingual Language Models —
- Healthy Counterfactual Generation via Diffusion Inpainting for Mammography Classification —
- Large-Scale Benchmarking of Quantum Neural Network Configurations for Financial Time Series Forecasting —
- Connected Self Forcing: Beyond Local Learning in Video Autoregression —
- When KL Regularization Misfires in Group Policy Optimization —
- AI-Based On-Board Maritime Object Detection for Earth Observation Payload Data Reduction on Versal Embedded Hardware —
- Just Weather Scoring: Efficient End-to-end Nowcasting with Distributional Diffusion —
- Quickest Change Detection with Diffusion-Integrated Scores —
- SciTBERT: A family of chronologically consistent language models for scientific and technological language processing —
- Verification with Transfer: Exact Information Frontiers and Their Price in Calls —
- ISBO: Scalable Spatio-Temporal Bayesian Optimization with Log Gaussian Cox Process Models via the INLA-SPDE Approach —
- Stride Independent Patching for Deep Learning —
- VibeEdit: Image Editing with Canvas Instructions —
- From Prompting to Composing: A Spatial Canvas Interface for Poster Generation —
- Language Models as AI Research World Models —
- TokenRouter: Efficient Serving System for Token-Level LLM Routing —
- EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams —
- Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings —
- DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception —
- HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments —
- Slot3R: Set-Associative Spatial Memory for Streaming 3D Reconstruction —
- Testing Algebraic Complete Intersections —
- Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction —
- From What to Which: Decoding Modifier Grounding in Frozen MLLMs —
- BudgetPix: Compute-Adaptive Tokenization for Pixel-Space Image Diffusion —
- Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding? —
- Verdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance Systems —
- Controllable Exaggeration for Generative Motion Models via Training-Time Adaptation and Inference-Time Guidance —
- SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference —
- Composite Online-to-Nonconvex Conversion with Optimal Oracle Complexity —
- Prediction-Powered Data Fusion for Treatment Effect Estimation —
- RiCo: Neural Simulation of Rigid-Body Interactions via Local Contact Reasoning —
- ContiLNN: Mitigating Slice Sampling Discontinuity with Liquid Neural Networks for Medical Image Restoration —
- VFold: Symmetry-Aware Cross-Layer Value Cache Compression —
- Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition —
- Reasoning-Informed Visual Editing —
- Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution —
- Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models —
- Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict —
- Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness —
- Closing the Horizon Gap in Policy Optimization for Adversarial MDPs —
- HANS: A Handwritten Answer Sheet Dataset for Noisy Hybrid Document Parsing —
- Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution —
- Bilevel optimization for data-driven learning of Koopman embeddings using kernel-based autoencoders —
- AgentGarten: Code Worlds for Evolving Agents —
- OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport —
- Latent Core Tokenizer: Compress, but Meaningfully —
- WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation —
- GenIA: Generative Reconstruction with Test-Time Input Alignment —
- Long Text to Predictive Features: LLM-Guided Blockwise Feature Engineering via Executable Program Search —
- SpaceFlow: Locally Controllable 3D Generation —
- SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models —
- ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills —
- Predicting Alignment Generalization with Value Representations —
- WorldCast: Distributed Multiplayer World Models —
- MAMHOI: Factorizing Scene-Aware Human-Object Interaction through Affordances —
- WOVEN: Weaving Visual World Modeling into Multimodal LLMs —
- OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video —
- Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching —
- Pumpire: Unified Benchmark for Metric Distance Estimation —
- FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams? —
- Density Ratio Estimation with Stein Displacement Fields —
- LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation —
- One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts —
- VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation —
- BrickBench: Evaluating Agentic Brick Design —
- OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning —
- WorldGuide: Goal-Directed Video World Model for Procedural Task Execution —
- OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs —
- What 30,000 Hours of Ego-centric Video Does Not Teach —
- Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation —
- Keeping Score: Adaptive, Tuning-Free Loss Weighting for Score-Augmented Neural Ratio Estimation —
- An Explainable Header-Centric Framework for Large-Scale Semantic Table Interpretation and Data Quality Assessment —
- Diffu-LoRA: A Novel Low-Rank Adaptation for Personalized Diffusion Models —
- SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows —
- Wieszcz-XIX: A 3.1-Billion-Word Corpus of Pre-1918 Polish and Temporally Bounded Language Models Trained From Scratch —
- The optimal information complexity of VC learning —
- A Camera-Native Stereo VR180 Dataset —
- Temporal transformer CAN encoder with federated lightweight heads for anomaly detection — The gist The proposed framework introduces a privacy-preserving approach for anomaly detection in in-vehicle networks by combining a Temporal Transformer CAN Encoder with Federated Lightweight Heads to capture subtle temporal and contextual anomalies. [episode]
- JevForest: Path Voting for Budgeted Feature Acquisition —
- When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry — The gist The router-augmented membership inference attack combines conventional outputside signals with aggregated routing features and applies a membership classifier learned from independently fine-tuned shadow models to reveal whether an example was used to fine-tune the deplo [episode]
- Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models —
- Exact SO(3)-Equivariant Isotropic Kernels for Rotation-Robust Neural Dynamics —
- D-SLR: The Disjoint Row-Sparse plus Low-Rank Decomposition —
- From Log-Odds to Shapley Values: An Explanatory Geometry for the Weighted Naive Bayes Classifier —
- Large Language Model-Assisted Preparation of Transportation Management Plans: A Case Study with WisDOT WisTMP System —
- Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning — The gist The Nullify framework proposes a training-free, non-destructive activation steering method for LLM unlearning that redirects privacy-related activations away from memorized answers while satisfying a nullspace constraint to preserve model utility. [episode]
- Explaining the Saliency Map Sparsity of Adversarially-Trained Neural Networks —
- Learning infinite context windows in recurrent architectures via spatial neural computing —
- Seeing Through the Glare: A Multi-Source Benchmark and Ocular-Adaptive Pixel MeanFlow for Eyeglass Reflection Removal —
- LinSlot: Exploiting Linear Representation hypothesis for unsupervised attribute discovery from slot based object representation —
- Cognitive Thermometers: Machine Learning and Logical Complexity —
- Lossy Compressive Text Autoencoders —
- Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark Suite, and Training —
- Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale —
- MESSENGER: Memory-Enhanced Sequential Scene Flow Estimation via Autoregressive Next-Frame Forecasting —
- DOGS: Design-Space Sampling for Prompt-Driven Logo Generation —
- What can linear attention learn from nonlinear teachers in-context? —
- VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning —
- Plan-and-Patch: Diffusion Language Models for Agentic Planning —
- Calibrating Ambiguity Set via Diagnostic Transport for Distributionally Robust Optimization —
- Velocity Scaling in Flow Matching —
- Grammar Concept Annotation at Scale: Deployed Fine-Tuned Small Language Models Outperform Prompted Frontier Models —
- Conformal Prediction under Partial Verification —
- Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute —
- Similar Predictive Fit but Different Latent Dynamics: Characterizing Learned Dynamical Structure in Personalized Models of Brain Disorders —
- Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning —
- Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models —
- Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders —
- Transformed Samplers with Variance Reduction —
- Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values —
- Stochastic Teacher Intervention for Agentic On-Policy Distillation —
- SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages —
- Less from More: Reinforcing Sparse Video Reasoning from Dense References —
- Large Language Models for Machine Translation Quality Annotation: Humans and Models Are Both Challenged —
- Language Models for Page-Level Layout Decisions in E-commerce Search —
- StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents —
- GHARP: Real-time Gaussian Head Animation from Large-scale Reconstruction Prior —
- AI4Fire: Evaluating Large Language Models on Wildfire Tasks —
- SPD-MetaFormer is what you need for small-data brain decoding —
- GPU-Accelerated Computation of Persistent Homology for Topological Analysis of Image Data —
- LVSPM: Long Sequence View Synthesis and Pose Estimation Model —
- When Citations Mislead? A Claim-Level Benchmark for Legal Hallucination Detection —
- Fluid-Gen-Zero: Grounding Pretrained Video Generators in Physics without Training —
- Back in Style: A Sociolinguistic Approach to Authoring and Measuring Persona Fidelity in User Simulation —
- Multi-Bandwidth Distribution Matching Distillation: On the Equivalence of Distribution Matching Distillation and Drifting Models —
- Omni-Diffusion-Distill: Few-Step Distillation of Unified Multimodal Diffusion Large Language Models —
- Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval —
- PCAsplat: Gaussian Splatting with Local PCA Regularization —
- Prompts versus Rules: Auditing and Controlling Speech Naturalness Behaviors in Voice User Simulators —
- Mid-Training Language Models on Raw Video —
- Expression-Diverse References for Identity-Preserving Video Generation —
- FedAlphaEdit: Null-Space-Aligned Merging for Collaborative Knowledge Editing —
- Transforming Image Editors into Video Editors —
- Rendering-Free Lookahead for Question-Guided Active Vision —
- SatFix: Absolute Visual Localization of UAVs in Satellite Maps from a Single Oblique Image —
- Learning What to Trust in Multimodal Learning under Noisy Supervision —
- Refine Connections, Close the Gap: A Reliable Enhancement Framework for Driving Scene Topology —
- AffordDrive3D: Affordance-Aware World-Action Modeling with Spatial Understanding —
- Measuring and Mitigating Solution Mode Collapse in RLVR —
- Diffusion Meta-Prompting and Steering for Generalizable Foundation Model Adaptation —
- Clinician use of language models diverges from how the models are evaluated —
- No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping —
- A General (sqrt T gamma T) Lower Bound for Kernel Bandits —
- Contrast Enhancement or Noise Reduction? On Improving Cervical Cancer Classification —
- Continuous Ground-Truth Construction and a Recovery Policy for Air--Water Robotic Tracking —
- Skeleton-Guided Progressive Test-Time Adaptation for Thin Curvilinear Structures —
- TKCAM: Text and Keyframe to Camera Trajectory Generation —
- Lapras: Latent Reasoning for Time Series Language Models —
- DiscoVL: Unveiling Disentangled C ross-Modal Representation Learning via Orthogonal Adversarial Regularization for V ision-Language Models —
- MCL: Meta Convolution Layer —
- GameCommBench: A Unified Benchmark and Type-Aware Evaluation for AI-Generated Game Commentary —
- SFT-as-Context Mitigates Forgetting in Supervised Fine-Tuning —
- Can a System-One LLM Perform Knowledge Tracing When Few or No Learners Are Logged? —
- The "10th Juror": Open-Set Standpoint Screening for Bureaucratic Bias Detection —
- IntrinSync: Joint Intrinsic Decomposition and Reciprocal Rendering —
- Accelerating Non-Smooth and Heavy-Tailed Sampling —
- ActiveMedAgent: Cost-Aware Trajectory Learning for Multimodal Medical Diagnosis —
- Improving Image-Based Nutrition Estimation Through Multimodal Food-Item Verification and Recovery —
- SP-DocReader: Difference-Aware Self-Play for Precise Document OCR —
- A Unified Score Matching Paradigm for Video Anomaly Detection and Anticipation —
- Do LLMs Learn from Rewards in Context?: Rethinking the role of reward in In-Context Reinforcement Learning —
- CARE: Constrained Attention Refinement for Fine-Grained Visual Classification via Teacher-Student Distillation —
- Local Prototype Reconstruction for Text-Compatible Speech-to-LLM Bridge Pretraining —
- LadderEdit: Edit-Level Residual Compression for Memory-Efficient Lifelong Editing of LLMs —
- VGGTWorld-VLA: Intent-Conditioned 3D World Evolution for Autonomous Driving —
- AutoAdapt: Reliable Few-Shot Adaptation under Clinical Distribution Shifts —
- VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding —
- IntactWorld: Joint World Modeling with Intact Features —
- Multimodal Remote Sensing Image Registration: A Comprehensive Review, Challenges and Prospects —
- MATE4D: Matrix-Guided Editable 4D Generation from a Single Image —
- RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation —
- WorldFact-Bench: Beyond Image-Internal Plausibility to Image-World Consistency —
- Predictive Multiplicity in Cell-Fate Assignment: Label-Free Rashomon Sets and the Limits of Per-Cell Certification —
- RGBD-to-3D Object Mesh Refinement via Depth Matching and Symmetry Propagation —
- 3DTexMOR: 3D Gaussian Multi-Object Removal via Texture-Space Inpainting —
- Dissecting Representation Structure in Vision Transformers: A Rigorous Architectural Study —
- Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration —
- GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images —
- The Lattice of Transition Laws —
Important terms
- Thought-Like-Pro
- This technique uses a self-bootstrapped Prolog chain of thought to enhance large language model reasoning. It allows models to create their own intermediate reasoning steps, making them better at complex logical problems instead of just pattern matching.
- LoRi
- Low-rank distillation is used here to improve implicit reasoning in large language models. This involves transferring knowledge from larger models into smaller ones, effectively boosting the smaller model's ability to perform complex inference.
- GUI-KV
- This method improves how agents interact with graphical user interfaces by using a KV cache and incorporating spatio-temporal awareness. This helps agents track evolving visual states over time more effectively.
- State Stream Transformer V2
- This new method tackles latent space reasoning by training nonlinear recurrence in parallel. It aims to give models a more nuanced way to understand complex information stored within their internal representations.