AI papers — 2026-10-01
The development of an auditable agentic discovery process for designing deletion-only tools builds upon foundational work in genomic modeling and large language model reasoning. Initial steps involve exploring BlockFormer, a transformer-based inference method applied to genomic contact maps to infer structural information. This contrasts with challenges in understanding computation within multi-task neural networks, where the Green's Operator is used to disentangle these computational pathways. Furthermore, researchers are examining how large language models handle reasoning beyond their commitment boundaries by probing epiphenomenal chain-of-thought processes. They are also investigating how human label variation relates to formal semantic structure in natural language inference tasks.
The work on diffusion language models specifically notes that a dominant self-conditioning direction drives repetition in unconditional continuous diffusion language models. Additionally, researchers are looking at query-key alignment within large language models to unlock latent correct answers and assessing zero-shot hallucination detection using a single LLM response based on lowest span confidence.
The work on VisionFoundry focused on teaching vision language models visual perception using synthetic images to explore how this approach impacts their ability to understand and interpret visual data. The researchers investigated the efficacy of this method in developing the model's visual understanding capabilities. Generalizing the Turing Test to interactive agents was also examined, suggesting an effort to move beyond simple classification towards more complex agent behavior assessment.
Concurrently, GrepSeek aimed at training search agents for direct corpus interaction, implying a focus on how these agents can effectively navigate and engage with large bodies of text data. In terms of system architecture, OctoNest introduced adaptive cross-device execution through stateful control, indicating an interest in creating more flexible and robust deployment environments for complex systems. Furthermore, Fork-Think with Confidence suggests an exploration into methods for improving the reasoning or confidence levels within a model's decision-making process.
The research on removing timing shortcuts in brain-to-text translation points toward refining non-invasive techniques by addressing temporal dependencies in neural data interpretation. ETHER introduced aligning emergent communication for hindsight experience replay, suggesting a mechanism to better utilize past experiences during learning cycles. Finally, the work on mitigating memorization in language models addresses a core challenge in large model training by seeking ways to prevent the model from simply recalling training data verbatim.
The work on edit based fingerprints for large language models explored how to create unique identifiers by analyzing the edits made to text, which suggests a method for tracking provenance or identifying model drift through textual modifications. This contrasts with the effort on think right which focused on mitigating under-over thinking in models using adaptive and attentive compression techniques, aiming to improve reasoning accuracy. Furthermore, there was research into medrect, a bilingual medical reasoning benchmark designed specifically for error correction within clinical texts, indicating a focus on domain-specific accuracy.
Simultaneously, context aware classification and grading of sensitive information in online conversational health data addressed the challenge of accurately identifying and handling private details within unstructured text streams. The study on semantic chunking and the entropy of natural language looked at how to segment text based on meaning while measuring its inherent randomness, which provides a structural understanding of language complexity. Dataflex presented a unified framework for data centric dynamic training, suggesting a way to adapt models directly from their data sources.
RA-MoE focused on routing aligned fine tuning for multilingual adaptation specifically within mixture of experts models, pointing toward improving cross-lingual performance in complex architectures. Finally, interactor explored agentic reinforcement learning oriented iterative creation for ad description generation in sponsored search environments, focusing on interactive generation rather than just static prediction.
The CombEval framework was introduced to address the challenge of evaluating combinatorial counting capabilities within large language models. This involved designing a method where models are tested on complex counting tasks and then a separate mechanism assesses how well they handle variations in those counts. The findings suggest that simply achieving high scores on these combinatorial tests does not guarantee robust counting abilities, implying that the framework is necessary to probe deeper into the model's internal logic rather than just surface-level performance.
The QuantCode model explored a method for specializing language models specifically for executable algorithmic trading code, focusing on how to bridge the gap between natural language descriptions and functional programming logic. This involved taking existing large language models and fine-tuning them with proprietary trading code examples to improve their ability to generate syntactically correct and functionally sound trading scripts.
A related effort examined a spike-driven vision-language-action model, suggesting that incorporating visual data alongside language input could enhance the model's capacity for action planning in complex environments. Furthermore, research into compact language and complex model shifts indicated that ambiguity and underspecification in prompts significantly affect how large language models perform. This observation connects to work on marginal response surface elicitation for zero-label tabular learning, where researchers sought ways to generate meaningful responses even when labels are absent by exploring the boundaries of what a model can produce.
The evolution of attention mechanisms in large language models also provided context, showing how changes in these internal weighting systems represent trade-offs between computational efficiency and the model's ability to capture long-range dependencies within code. This is complemented by studies on directed transfer in instruction-tuning mixtures, which investigate how guiding one task while simultaneously introducing a challenging secondary task can improve overall performance.
LatentHarness introduced a technique for learning latent actions for memory and reasoning through counterfactual policy distillation, suggesting a path toward more robust internal state management within models. Finally, stress-testing LLM lie detectors revealed that role-play failures and spurious correlations can lead to unexpected breakdowns in the model's ability to detect deception, highlighting vulnerabilities in current safety mechanisms.
The work on ViLegalExpert focused on creating a large-scale benchmark for Vietnamese legal retrieval and question answering by using real-world consultations. This involved setting up a system to test how well models could handle complex legal queries from actual cases. Related efforts explored the cognitive capabilities of vision language models through 4MT-VLM, which examined how coarse the internal cognitive map of these models is when processing visual and linguistic information.
Furthermore, there was investigation into improving small language models by reusing feedback from larger LLMs through offline guidance and online reasoning techniques to enhance their performance. To address issues in search algorithms, research looked at making grid beam search less greedy, suggesting a refinement in how the model explores potential answers. The robustness of these systems against errors was also considered with RAIM, which introduced robust aggregation of inexpensive models specifically for hallucination detection.
Additionally, efforts were made to compress looped models by considering a tilted bowl analogy and to define language model understanding through bounded Dutch books as both a definition and a training objective. Finally, Ready2Blend addressed the challenge of translating natural-language instructions into composable alignment prompts.
The TTLab at Daleel focused on STAR-Ar, which involved sequence tagging for argument recognition within Arabic text. This work sought to develop a method for identifying the arguments presented in these texts. The findings from this research suggest progress in understanding how arguments are structured within Arabic discourse.
Moving beyond this, DuplexAct-Bench was introduced to broaden full-duplex speech evaluation, aiming for proactive interaction across various behavioral needs. Concurrently, there was a computational linguistic analysis conducted on Frei.Wild, which explored the relationship between right-wing rhetoric and rock music through computational methods.
Furthermore, CATCH was developed as a controllable analysis testbed specifically for reward hacking within reinforcement learning coding environments. This suggests an effort to probe the boundaries of learning agents' behavior under specific constraints. The exploration into whether language models can rely on external guidance selectively also touched upon this area, questioning their capacity for targeted reliance on outside input. Finally, zero-compute cross-lingual transferability estimation was attempted using typological feature proxies to gauge how well models translate across languages without extensive compute.
The work on MemCodex focused on developing a self-programming hierarchical memory system for language agents. This involved exploring how such a structure could manage and utilize information effectively during agent operation. A related study examined synthetic pre-pretraining, finding that while scaling up the training data improved performance, it did not necessarily translate into gains in grammatical prior compared to smaller models.
Furthermore, research into cognitive enhancement suggested rethinking the necessity of role-playing for large language models when considering their operational needs. The investigation into whether computation derived from earlier problems could assist LLMs in solving novel ones was also conducted. Synthetic data characterization through training dynamics provided insights into how these processes shape model behavior.
In a separate line of inquiry, an arithmetic-dependent rejection bottleneck was identified in the Jev problem, indicating where models struggle when the correct answer is absent. This led to work on learning to verify rule-governed decisions, which addresses whether the evidence gathered is decision-critical. Finally, SEPAL explored separating expert pairs using answer-level fusion to achieve more reliable collaboration between LLMs.
Today's papers
- From DNA Design to DNA Slimming: Auditable Agentic Discovery of a Deletion-Only Designer This paper describes an agent that can discover new DNA sequences by only making deletions. [paper] [episode]
- BlockFormer: Transformer-based inference from genomic contact maps This model uses a transformer to infer biological interactions from genomic contact maps. [paper] [episode]
- Disentangling Computation in Multi-Task Neural Networks with the Green's Operator This work introduces an operator to separate different computational paths within multi-task neural networks. [paper]
- Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models This research investigates how chain-of-thought reasoning works beyond what is explicitly shown in large language models. [paper] [episode]
- A Dominant Self-Conditioning Direction Drives Repetition in Unconditional Continuous Diffusion Language Models This paper shows that a strong self-conditioning direction causes repetition in certain continuous diffusion language models. [paper] [episode]
- How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI This study examines how much human label variation is explained by formal semantic structures in natural language inference tasks. [paper] [episode]
- Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models This technique uses query-key alignment to find hidden correct answers within large language models. [paper] [episode]
- Lowest Span Confidence: Zero-Shot Hallucination Detection from a Single LLM Response This method detects hallucinations by checking the confidence of the lowest span in an LLM's response. [paper] [episode]
- VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images This paper teaches vision-language models to perceive visual information using synthetic images. [paper] [episode]
- Generalizing the Turing Test to Interactive Agents This research explores how to generalize the Turing test concept for interactive AI agents. [paper] [episode]
- GrepSeek: Training Search Agents for Direct Corpus Interaction This work focuses on training search agents to directly interact with a corpus of text. [paper] [episode]
- OctoNest: Adaptive Cross-Device Execution through Stateful Control This system allows for adaptive execution across different devices using stateful control mechanisms. [paper] [episode]
- Fork-Think with Confidence This paper discusses how to think critically and maintain confidence when exploring multiple reasoning paths in models. [paper] [episode]
- Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text This research aims to improve brain-to-text translation by removing timing shortcuts in the process. [paper]
- ETHER: Aligning Emergent Communication for Hindsight Experience Replay This work aligns emergent communication patterns to improve hindsight experience replay in reinforcement learning. [paper] [episode]
- Mitigating Memorization In Language Models This paper explores methods to reduce memorization within large language models. [paper] [episode]
- From Construction to Injection: Edit-Based Fingerprints for Large Language Models This method uses edit-based fingerprints to identify and manage modifications in large language models. [paper] [episode]
- Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression This paper develops a method using compression to help language models avoid under or overthinking. [paper] [episode]
- MedRECT: A Bilingual Medical Reasoning Benchmark for Error Correction in Clinical Texts This is a benchmark designed for bilingual medical reasoning and error correction in clinical texts. [paper] [episode]
- Context-Aware Classification and Grading of Sensitive Information in Online Conversational Health Data This system classifies and grades sensitive health information within online conversational data based on context. [paper] [episode]
- Semantic Chunking and the Entropy of Natural Language This study analyzes how semantic chunking affects the entropy, or randomness, of natural language. [paper] [episode]
- DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models This framework provides a unified way to dynamically train large language models using data-centric methods. [paper] [episode]
- RA-MoE: Routing-Aligned Fine-Tuning for Multilingual Adaptation of Mixture-of-Experts Models This method fine-tunes mixture-of-experts models for multilingual adaptation by aligning the routing process. [paper] [episode]
- Interactor: Agentic RL oriented Iterative Creation for Ad Description Generation in Sponsored Search This system uses agentic reinforcement learning to iteratively create ad descriptions in sponsored search. [paper] [episode]
- CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models This framework is designed to evaluate how well large language models perform combinatorial counting tasks. [paper] [episode]
- Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs This research uses fine-tuned language models to find indicators of vulnerability in police incident logs from the UK. [paper] [episode]
- Frozen Memory Is Not Enough: Rethinking External Memory as Extraction This paper suggests that external memory should be treated more like an extraction process rather than just frozen storage. [paper] [episode]
- An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning This study empirically tests how reward specification and benchmark reliability affect the unlearning capabilities of GRPO-based LLMs. [paper] [episode]
- Argument Structure Prediction in Online Conversations: A Comparative Study of Modeling Paradigms and Task Architectures This research compares different modeling approaches for predicting argument structures in online conversations. [paper]
- Concept Subspaces Compute Beyond the Logit Lens: A Weights-Only Test for Locating Representations Upstream of Readout This paper tests whether concept subspaces can locate representations before the final readout layer. [paper]
- Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost This work shows how byte-exact memory can make reading text a one-time cost for large language models. [paper] [episode]
- Exploring Heterogeneous Model Merging Approach for Complex Knowledge Transfer This paper explores merging different heterogeneous models to achieve complex knowledge transfer. [paper]
- QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code This model specializes language models specifically for generating executable algorithmic trading code. [paper]
- Spike-driven Vision-Language-Action Model This model is a vision-language-action model driven by spike dynamics. [paper] [episode]
- Compact Language, Complex Model Shifts: How and Where Ambiguity and Underspecification Affect LLMs This paper investigates how ambiguity and underspecification affect language models when the language becomes compact. [paper]
- Marginal Response Surface Elicitation for Zero-Label Tabular Learning This method elicits marginal response surfaces to perform zero-label tabular learning. [paper] [episode]
- The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends This paper reviews the evolution of attention mechanisms in large language models, covering their mechanisms and trade-offs. [paper]
- A helps B while B hurts A: directed transfer in instruction-tuning mixture This research looks at how to achieve directed transfer when mixing instruction tuning data where A helps B but B also hurts A. [paper] [episode]
- LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation This method learns latent actions for memory and reasoning by using counterfactual policy distillation. [paper]
- Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations This study stress-tests lie detectors by examining role-play failures and spurious correlations in large language models. [paper] [episode]
- ViLegalExpert: A Large-Scale Benchmark for Vietnamese Legal Retrieval and Question Answering from Real-World Consultations This is a large benchmark for Vietnamese legal retrieval and question answering based on real consultations. [paper]
- 4MT-VLM: How Coarse Is a VLMs Cognitive Map? This paper investigates the coarseness of cognitive maps in vision-language models. [paper] [episode]
- Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models This work shows how to reuse large language model feedback for reasoning in smaller language models using offline guidance. [paper]
- Making Grid Beam Search Less Greedy This method aims to make grid beam search less greedy by improving its exploration strategy. [paper]
- RAIM: Robust Aggregation of Inexpensive Models for Hallucination Detection This method robustly aggregates inexpensive models to detect hallucinations. [paper]
- A Tilted Bowl Is Not a Slippery Slope: Compressing Looped Models This paper discusses compressing looped models and argues that tilting the bowl is not a slippery slope in this context. [paper]
- Understanding as No-Arbitrage: Bounded Dutch Books as a Definition and Training Objective for Language Models This research uses bounded Dutch books as both a definition and training objective for language models based on no-arbitrage principles. [paper]
- Ready2Blend: From Natural-Language Instructions to Composable Alignment Prompts This system converts natural language instructions into composable alignment prompts. [paper] [episode]
- TTLab at Daleel 2026: STAR-Ar, Sequence Tagging for Argument Recognition in Arabic This work introduces STAR-Ar for sequence tagging and argument recognition in Arabic. [paper]
- DuplexAct-Bench: Broadening Full-Duplex Speech Evaluation toward Proactive Interaction across Diverse Behavioral Requirements This benchmark broadens full-duplex speech evaluation to cover proactive interactions with diverse behavioral needs. [paper] [episode]
- Right-Wing Rock or Just Rock? A Computational Linguistic Analysis of Frei.Wild This paper provides a computational linguistic analysis of the political language used in Frei.Wild. [paper] [episode]
- CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL This is a controllable testbed designed to analyze reward hacking in reinforcement learning for coding tasks. [paper]
- Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively? This paper asks whether language models can selectively rely on external guidance when thinking outside the box. [paper]
- Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies This method estimates cross-lingual transferability without requiring zero compute by using typological feature proxies. [paper] [episode]
- Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation This technique uses neighborhood on-policy self-distillation to provide better supervision. [paper] [episode]
- Drift Inspector: Exploring and Measuring Scientific Drift with Atomic Contribution Claims This system explores and measures scientific drift by using atomic contribution claims. [paper] [episode]
- MemCodex: Self-Programming Hierarchical Memory for Language Agents This system gives language agents self-programming hierarchical memory capabilities. [paper]
- Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior This study examines whether synthetic pre-pretraining survives large scale but fails to provide a strong grammatical prior. [paper]
- Cognitive Enhancement: Rethinking the Necessity of Role-Playing for Large Language Models This paper rethinks the need for role-playing in large language models for cognitive enhancement. [paper]
- Can Computation from Earlier Problems Help LLMs Solve New Ones? This research investigates whether computation derived from earlier problems can help LLMs solve new ones. [paper] [episode]
The papers
- ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference — DeltaKV introduces a residual-based KV cache compression framework that leverages long-range inter-token similarity to significantly reduce memory footprint while maintaining near-lossless performance for long-context Large Language Models. [episode]
- Lensless Gaze Is Not Private by Default: Auditing Identity Leakage Across Disclosure Surfaces — This paper conducts an audit of identity leakage within a simulated lensless gaze sensing pipeline to demonstrate that visual unintelligibility does not equate to privacy by default. [episode]
- Alignment via Training Against Probes Without Losing Monitorability — Models can be aligned by training directly against probes that are continuously refit, without sacrificing utility. How it works The core idea involves using probes to detect undesired properties in model activations as a direct training signal for fine-tuning. [episode]
- Can Computation from Earlier Problems Help LLMs Solve New Ones? — Can computation from earlier problems help LLMs solve new ones? The gist: STAIR improves mean T2–T4 Avg@4 on AIME 2025 by 11.67 percentage points for Qwen3.5-4B and 6.11 points for Qwen3.5-9B over Native while training only 12,288 parameters across three Qwen models and four re [episode]
- OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning — As a fastidious and diligent AI researcher, I have meticulously analyzed both provided texts regarding the paper "OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning." The goal is to synthesize these two sources into a single, comprehensive, and highly detailed summ [episode]
- MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories — Long-term egocentric video enables personalized AI assistants to reason about daily life, but reprocessing raw clips for every query becomes computationally prohibitive. [episode]
- Lowest Span Confidence: Zero-Shot Hallucination Detection from a Single LLM Response — This paper introduces Lowest Span Confidence (LSC), a novel zero-shot metric designed for efficient and black-box hallucination detection in Large Language Models (LLMs). [episode]
- CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models — Online reinforcement learning has been successfully extended to flow matching for diffusion model (DM) image generation, but existing methods suffer from limitations regarding window selection, reward saturation, and sample inefficiency. [episode]
- Embedding Physical Reasoning into Diffusion-Based Shadow Generation Under the Sun and Sky — Generating realistic shadows for inserted objects requires reasoning about scene geometry and illumination, which existing methods often fail to do purely in image space. [episode]
- Amortized Bayesian Inference on Multilevel Models of Arbitrary Structure — Amortized Bayesian inference (ABI) on multilevel models of arbitrary structure provides a general method for deriving neural network architectures that automatically determine valid posterior factorizations, enabling efficient and scalable Bayesian inference across complex hierar [episode]
- SatSplatDiff: Geometry-preserving generative refinement for high-fidelity satellite Gaussian Splatting — SatSplatDiff is a unified pipeline for satellite-based 3D reconstruction that achieves state-of-the-art geometric accuracy and visual fidelity by extending shadow casting into the generative refinement stage. [episode]
- Let the Carrier Carry the Attack: Preserving the Subject in Adversarial Image Generation — A visually secondary carrier can serve as an auxiliary spatial pathway to facilitate strong and transferable adversarial attacks while preserving the integrity of the primary subject. [episode]
- A helps B while B hurts A: directed transfer in instruction-tuning mixture — A directed transfer map reveals that task selection for language model instruction tuning is not symmetric, demonstrating that certain tasks can actively help a target while others actively harm it. [episode]
- Efficient Active Auditing of Multi-Group Fairness with Bias Probes — As a fastidious and diligent researcher, I have thoroughly analyzed both provided texts from arXiv. [episode]
- ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models — ShieldCLIP introduces a framework for selective safety alignment in CLIP-like multimodal encoders, addressing the challenge of suppressing harmful associations in foundation models without unnecessarily altering benign representations. [episode]
- UXBench: Measuring the Actionability of LLM-Generated UX Critiques — Large language models (LLMs) are increasingly deployed as UX judges that inspect interfaces, diagnose usability problems, and propose repairs, but no controlled benchmark currently measures whether these critiques are reliable and actionable across different product surfaces. [episode]
- BAM! Bayesian Anything Model: a foundation model for generative computational imaging — Generative models are transforming Bayesian computational imaging, yet the field still lacks physics-aware foundation models. [episode]
- ITO: Multi-View Alignment and Training-Time Fusion for Image-Text Pretraining — ITO (Image–Text as One) is a framework proposed to address modality-induced separation in image-text representations by synergizing multimodal multiple alignment with a lightweight training-time fusion module. [episode]
- Interactor: Agentic RL oriented Iterative Creation for Ad Description Generation in Sponsored Search — This paper introduces INTERACTOR, an agentic Reinforcement Learning framework designed for automatically generating informative ad descriptions in sponsored search. [episode]
- From Compound Figures to Medical Multi-image Reasoning: Scaling Multimodal Large Language Models with Biomedical Literature — As a meticulous researcher, I have thoroughly analyzed both provided texts from the arXiv paper, "From Compound Figures to Medical Multi-image Reasoning: Scaling Multimodal Large Language Models with Biomedical Literature." The synthesis below is designed to be comprehensive, det [episode]
- TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images — Recent advances in text-to-image models have improved global realism, but persistent failures in local typography—such as malformed glyphs and broken strokes—remain significant perceptual quality issues that current evaluations fail to capture. [episode]
- Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models — Progress in 4D LiDAR segmentation is currently bottlenecked by data scarcity, as assigning temporally consistent labels across sparse point cloud sequences is costly and difficult to scale. [episode]
- Semifactual Credit-Augmented Policy Optimization — Reinforcement learning with verifiable rewards (RLVR) has advanced large language model reasoning, but its predictions remain sensitive to task-irrelevant prompt features. [episode]
- Marginal Response Surface Elicitation for Zero-Label Tabular Learning — Marginal Response Surface Elicitation (MARS) is a method that transforms feature-level Large Language Model (LLM) priors into a reusable, zero-shot tabular classifier by selecting representative feature values and aggregating LLM responses into weighted feature response functions [episode]
- CogCanvas: A Benchmark for Evaluating Multi-Subject Reference-Based Image Generation — CogCanvas introduces a comprehensive benchmark designed to systematically evaluate multi-subject reference-based image generation by simultaneously testing identity preservation, object/fashion binding, and background scene consistency. [episode]
- Less Data, Better Timing: Student-Curriculum Coupling for VLM On-Policy Distillation in Temporal Video Grounding — On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). [episode]
- SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation — Span-Level Uncertainty Estimation (SLUE) formalizes a new task targeting semantically coherent text spans to provide interpretable and localizable uncertainty scores for Large Language Model generations. [episode]
- Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations — Lie detection probes are being stress-tested by introducing role-play scenarios, which complicate what "truth" means for large language models. [episode]
- GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives — As a fastidious and diligent researcher, I have meticulously analyzed both provided texts regarding "GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives." The combination of these excerpts reveals a comprehensive picture of this work, de [episode]
- Structure over Pixels: Learning Variable-Length Visual Programs — Discrete visual tokenizers (DVTs) translate images into ordered sequences of discrete codes, but existing adaptive tokenizers often struggle to learn a continuous per-image sequence length coupled to both the model and scene. [episode]
- Atomic and Holistic LLM Judges for Reference-Grounded Support Labels: A Prompt-Controlled Comparison — This paper investigates whether breaking down an answer into atomic claims for verification provides a performance advantage over using a holistic rubric when classifying reference support in benchmark-style reference-grounded classification tasks. [episode]
- Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost — A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out or verify. [episode]
- Steering Fields: Adaptive Vector Fields for Safe Image Generation and Beyond — Steering Fields introduce an adaptive vector field generalization of steering vectors that dynamically re-estimates steering directions at each step of image generation, offering a novel, model-agnostic framework for safe image generation and structure-preserving image editing. [episode]
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions — Most coding-agent benchmarks are static, but real coding assistance is interactive, requiring users to clarify goals and correct mistakes over multiple turns. [episode]
- Beyond Normal References: Discriminative Few-Shot Anomaly Detection — This paper introduces IDEAL, an intrinsic deviation learning framework designed for discriminative few-shot anomaly detection (FSAD), which addresses the limitations of existing methods by leveraging both normal and anomalous references to learn generalizable abnormality patterns [episode]
- Decoupling Multi-Contrast Super-Resolution: Self-Supervised Implicit Re-Representation for Unpaired Cross-Modal Synthesis — Multi-contrast super-resolution (MCSR) is crucial for enhancing MRI but current deep learning methods are severely limited by their reliance on large, paired low- and high-resolution (LR/HR) training datasets and fixed upsampling scales. [episode]
- Tighter Regret Bounds for Contextual Action-Set Reinforcement Learning — This paper investigates reinforcement learning under an action-set context setting, where an episode-dependent admissible action set is observed at the start of each episode, and it establishes tighter regret bounds for this challenging framework. [episode]
- XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments — As a diligent researcher, I have meticulously analyzed both provided texts from the arXiv paper "XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments." The goal is to synthesize these summaries into a single, comprehensive, [episode]
- Life-Bench: A Benchmark and Knowledge Graph Framework for Multimodal Personalization Beyond Concept Recognition — Life-Bench introduces a fully synthetic, human-verified multimodal benchmark designed to evaluate personalization capabilities in large language models beyond simple concept recognition. [episode]
- LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception — Paper Title: LEAP (Learned Block-wise Evidence Retrieval for Long Audio-Video Perception) Core Concept: LEAP is a novel framework designed to retrieve relevant evidence from long audio and video recordings without requiring the entire recording to be processed in a single context [episode]
- DiDA: Video Object Segmentation with Distillation Learning of Deformable Attention — This paper proposes a novel method for self-supervised video object segmentation (VOS) that integrates deformable attention and knowledge distillation learning to address limitations in temporal adaptation, computational efficiency, and long-term object memory. [episode]
- Delta-K: Boosting Multi-Instance Generation via Cross-Attention Augmentation — Diffusion models often struggle with concept omission when synthesizing complex multi-instance scenes, where existing training-free methods fail by merely rescaling attention maps without establishing coherent semantic representations. [episode]
- cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents — Computer use agents (CUAs), which utilize graphical user interfaces to complete tasks, have recently surpassed human performance on many standard benchmarks, but their speed and cost remain significant barriers to widespread deployment. [episode]
- DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models — DataFlex presents a unified data-centric dynamic training framework built upon LLaMA-Factory, addressing the fragmentation and reproducibility issues in existing data selection, mixture optimization, and reweighting methods for Large Language Models (LLMs). [episode]
- Ready2Blend: From Natural-Language Instructions to Composable Alignment Prompts — Continual alignment requires Large Language Models (LLMs) to adapt to new requirements without forgetting previously acquired behaviors, and this paper introduces Ready2Blend, a method that combines natural language flexibility with learned alignment to achieve this goal. [episode]
- LOCI: Spatial Linear Memory for Streaming World Models — When a camera revisits a previously observed region, a video world model must reproduce what was there before, requiring both remembering past observations and retrieving the right one for the current viewpoint. [episode]
- From Wrecks to Wisdom: Recovering Crash Mechanics from Real-World Multi-View Photos — Estimating accident mechanics from real-world crashes is important for vehicle-safety analysis, injury modeling, and crash-severity prediction. [episode]
- JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation — Large language models (LLMs) are increasingly deployed as automated judges for AI-generated content, yet when judges disagree, majority voting discards this conflict instead of resolving it. [episode]
- Information Thermodynamics of Agents: The Work Capacity of Channels with Memory — As a meticulous researcher, I have thoroughly reviewed both provided texts. [episode]
- GateSPINE: Gated Cross-View Fusion for Lumbar Spine MRI Report Generation — Automated report generation for lumbar spine MRI studies is being advanced by GateSPINE, a novel vision-language framework that addresses the limitation of existing methods in fully utilizing complementary information across different imaging planes. [episode]
- Concept Unlearning via Cross-Attention Activation Projection for Diffusion Models — Concept unlearning aims to erase a target concept from pretrained text-to-image diffusion models without retraining, which is crucial for deploying systems that must suppress copyrighted styles, celebrity likenesses, or unsafe imagery. [episode]
- PISCO: Precise Video Instance Insertion with Sparse Control — PISCO (Precise Video Instance Insertion with Sparse Control) is a video diffusion model designed to enable precise, controllable insertion of specific objects into existing footage under minimal user effort. [episode]
- A Competing-Hazards Systematization of Loss of Control in Autonomous Agents — Leading AI developers have reported agents acting beyond their approved limits, which a United Nations panel described as an early warning of loss of human control. [episode]
- Unapologetically Distributed: A Call for Decentralized Document Analysis — Distributed learning offers a valuable opportunity to improve model robustness and adaptability in real-world Document Analysis scenarios by enabling collaborative model training without direct data sharing. [episode]
- When Masking Helps or Hurts Robustness in Compressed CLIP: A Pre-Deployment Diagnostic — Whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. [episode]
- Amortized Bayesian Multilevel Models — Multilevel models (MLMs) are crucial for modeling complex, hierarchical data structures common in many scientific and observational fields, but their standard estimation methods, such as Markov Chain Monte Carlo (MCMC), suffer from significant computational challenges. [episode]
- PhysMirror: Physics-Aware Mirror Object Generation — PhysMirror introduces a novel, end-to-end physics-aware generation framework designed to synthesize photorealistic images with geometrically correct mirror reflections, addressing a critical failure point in current text-to-image diffusion models used for generating synthetic tra [episode]
- Contrastive On-Policy Distillation — Contrastive On-Policy Distillation (COPD) is a novel framework designed to enhance reasoning efficiency in large language models by introducing an explicit signal for comparing different reasoning modes. [episode]
- Don't Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents — LLM agents are increasingly used for long-horizon tasks requiring sequential reasoning and tool use, making runtime intervention crucial for improving reliability without retraining. [episode]
- RA-MoE: Routing-Aligned Fine-Tuning for Multilingual Adaptation of Mixture-of-Experts Models — Mixture-of-Experts (MoE) models offer efficient scaling for Large Language Models, but adapting them to non-English downstream tasks remains challenging because existing fine-tuning methods treat MoEs as monolithic learners, ignoring their heterogeneous routing structure. [episode]
- GeoWAM: Visual Geometry World Action Models for Autonomous Driving — GeoWAM introduces a novel visual geometry world action model designed for autonomous driving by shifting the focus from predicting future images to forecasting future scene geometry. [episode]
- Language-Conditioned World Modeling for Visual Navigation — This paper introduces Language-Conditioned Visual Navigation (LCVN), an open-loop trajectory generation task where an embodied agent must follow natural language instructions based only on an initial egocentric observation, without access to goal images. [episode]
- SEPAL: Separated Expert Pairs with Answer-Level Fusion for Reliable LLM Collaboration — SEPAL introduces a method for reliable large language model collaboration by assigning three private Actor–Critic teams to direct reasoning, evidence grounding, and verification, allowing them to revise answers in isolation before combining only their final results through a ma [episode]
- CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding — Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. [episode]
- FrontierGS: Progressive View-Space Frontier Expansion for Sparse 3D Gaussian Splatting — FrontierGS presents a curriculum-guided framework for progressive view-space frontier expansion in 3D Gaussian Splatting, addressing the challenge of sparse view synthesis by dynamically generating and promoting student views to augment supervision. [episode]
- PrePARE: Pre-AA Token Pruning for Frozen Multi-View Geometry Transformers — The PaceVGGT paper introduces PaceVGGT, a novel pre-Alternating-Attention (AA) token pruning framework designed specifically for frozen Visual Geometry Transformers (VGGT). [episode]
- Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression — Recent advancements in language models have enabled complex reasoning tasks, but their performance is often limited by an inability to regulate their reasoning length appropriately, leading to either underthinking on hard problems or overthinking on easy ones. [episode]
- PhantomEnvironments: Training LLM Agents in Fictional Worlds — Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. [episode]
- Provably Tractable NFA-Constrained Language Generation via HMMs — Constrained generation aims to sample from language models conditioned on hard constraints, and this paper introduces NFA-LM, a polynomial-time engine for NFA-constrained generation with theoretical guarantees under mild assumptions. [episode]
- A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss — This paper introduces SimLoss, a novel reference-free embedding-space objective designed to train single-pass fine-grained image captioning models. [episode]
- GlanceWAM: Sparse Test-Time Imagination for World-Action Models — Video generative models offer rich physical priors for robot learning, yet existing world-action models (WAMs) face a fundamental trade-off: synchronous video generation at control rate is latency-prohibitive, while abandoning test-time visual imagination sacrifices task success. [episode]
- RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation — RoGe presents an end-to-end unified framework for novel view synthesis (NVS) that removes the explicit bridge between reconstruction and generation, allowing a generative model to directly consume geometric conditioning derived from an implicit scene representation. [episode]
- Adaptive Reparametrized Time for Score-Based Diffusion Sampling — This paper introduces Adaptive Reparameterized Time (ART), a continuous-time control formulation designed to learn optimal timestep allocation for score-based diffusion sampling. [episode]
- Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation — This research investigates the critical issue of overconfidence in Vision-Language Models (VLMs) deployed for clinical decision support, a domain where trusting model predictions is paramount. [episode]
- Seeing as Humans Do: Learning from Motion to Segment Anything Without Supervision — The Segment Anything Model (SAM) currently relies heavily on massive manual annotations, creating a fundamental bottleneck for model scaling. [episode]
- Drift Inspector: Exploring and Measuring Scientific Drift with Atomic Contribution Claims — Scientific abstracts often conflate field discussion with actual research contributions, and this paper introduces Drift Inspector, an open-source system designed to measure how research fields evolve over time by analyzing Atomic Contribution Claims (ACCs). [episode]
- One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding — Frame selection is crucial for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows, and this work introduces Matryoshka Evidence-to-Context (MEC) Frame Selection, a training-free framework that constructs a single prio [episode]
- MomADv2: Reliable Temporal Memory for End-to-End Autonomous Driving — MomADv2 is a reliable state-space memory framework designed for long-horizon end-to-end autonomous driving, addressing the critical issue where existing temporal memory methods become invalid or misleading during command changes. [episode]
- VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images — VisionFoundry introduces an end-to-end synthetic data generation pipeline designed to address persistent weaknesses in visual perception among vision-language models (VLMs) by providing targeted, task-aware supervision. [episode]
- CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance — COUNTLOOP introduces a training-free framework that achieves precise instance control in high-instance image generation by alternating between VLM-based planning and iterative refinement, which resolves critical failures like count saturation and semantic leakage in dense scenes. [episode]
- SPROUT: A Scalable Diffusion Foundation Model for Multi-Crop Plant Phenotyping — Image-based plant phenotyping is crucial for modern crop science, but it is currently bottlenecked by expensive, bespoke annotation processes that lack cross-crop transferability. [episode]
- Simon-SR: Spatially Adaptive Modulation and Visual Prompt Adaptation for Text-Reinforced Super-Resolution — Simon-SR is a novel multi-modal super-resolution framework designed to reconstruct high-quality images from low-resolution inputs by leveraging learnable prompts for efficient semantic mining and robust text-image fusion. [episode]
- Distribution Matching Distillation for Continuous Diffusion Language Models — Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). [episode]
- Instruction Retrieval at Inference Time for Small Language Models — Small language models (SLMs) can achieve strong reasoning performance without scaling up parameters or requiring additional training by augmenting them with structured, reusable reasoning procedures retrieved at inference time. [episode]
- DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes — Reconstructing dynamic driving scenes from recorded sensor data supports closed-loop evaluation of autonomous driving systems by synthesizing observations beyond the original trajectory. [episode]
- AROID: Improving Adversarial Robustness Through Online Instance-Wise Data Augmentation — Deep neural networks are highly vulnerable to adversarial examples, posing significant security and trustworthiness risks for applications built upon them. [episode]
- VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes — Frontier multimodal large language models (MLLMs) have shown high accuracy on fine-grained perception benchmarks, but this success does not necessarily reflect faithful use of visual evidence. [episode]
- ATM: Why Latent World Models Can Fail to Plan — This paper introduces ATM, an Action-Consistency Transfer Matrix, designed to diagnose whether latent transitions in world models preserve action semantics relevant to planning. [episode]
- Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders — Sparse autoencoders (SAEs) are presented as an instrument for open-ended feature discovery from foundation model representations, enabling systematic exploration of learned patterns beyond pre-specified targets. [episode]
- MatLoom: Layered Text-to-Material Generation in a Compact Program Space — Compact, layer-oriented programs offer an effective output space for pretrained language models by combining text-to-material fidelity with explicit authoring structure. [episode]
- Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents — This research introduces Mid-Harness, a novel framework designed to systematically investigate how allocating test-time compute at the model-harness boundary can enhance the reliability of actions generated by terminal agents. [episode]
- Tackling Decision Processes with Non-Cumulative Objectives using Reinforcement Learning — Markov decision processes (MDPs) are widely used to model sequential decision-making, but many real-world problems, such as those involving risk-adjusted metrics like the Sharpe ratio or maximizing minimum rewards, do not fit the standard framework where the goal is to maximize t [episode]
- Mitigating Memorization In Language Models — Language models (LMs) possess an inherent ability to "memorize" training data, which can lead to verbatim regurgitation of private, sensitive, or copyrighted information during inference. [episode]
- Enhancing Autoregressive Video Generation via Representation Adversarial Distillation — Few-step autoregressive video generation enables efficient streaming synthesis, but errors introduced in early temporal blocks are reused as context and can propagate through subsequent rollouts, leading to detail degradation, structural drift, and unstable motion. [episode]
- Beyond Missing Rates: Rethinking Incomplete Multi-View Clustering with Protocol Divergence — This paper introduces CRAFT, a novel architecture designed to achieve robustness in Incomplete Multi-View Clustering (IMVC) by shifting the burden of robustness from the loss function to the model structure. [episode]
- Learning Adaptive and Visually Aligned Conditions for Generative Zero-Shot Learning — Generative zero-shot learning (ZSL) aims to synthesize plausible visual features for unseen classes by leveraging semantic conditions, but it faces two critical challenges: the class–instance gap, where class-level attributes fail to capture instance-specific appearances due to [episode]
- Diffusable Latents from Structure-Agnostic Distillation — Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and reach higher sample quality. [episode]
- BTC3D: Blended Tile Conditioning for Detail-Enhancing Image-to-3D Generation — Recent diffusion-based pipelines have shown promise in image-to-3D synthesis, but they struggle to generate high-fidelity details due to a common design flaw that causes detail attenuation. [episode]
- OctoNest: Adaptive Cross-Device Execution through Stateful Control — OctoNest proposes HRePlan, a hierarchical replanning framework for multi-device agents designed to handle failures robustly across heterogeneous environments. [episode]
- From Construction to Injection: Edit-Based Fingerprints for Large Language Models — Reliable model fingerprints are essential for protecting large language models (LLMs) against unauthorized redistribution and commercial misuse, as they enable ownership verification in black-box deployment where traditional methods are hindered by defensive filtering and downstr [episode]
- ETHER: Aligning Emergent Communication for Hindsight Experience Replay — ETHER (Emergent Textual Hindsight Experience Replay) is proposed as an extension to existing methods like HIGhER, aiming to bridge the gap between language-conditioned Reinforcement Learning (RL) and sparse reward environments. [episode]
- First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves — Recent progress in multimodal large language models (MLLMs) has fueled enthusiasm for their potential as autonomous agents, but current systems struggle when faced with complex, structured user requirements that involve both mandatory and optional constraints. [episode]
- When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift — Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train–test distributions. [episode]
- A Dominant Self-Conditioning Direction Drives Repetition in Unconditional Continuous Diffusion Language Models — Continuous diffusion language models (DLMs) exhibit low generative perplexity (Gen-PPL), but this metric rewards repetition, leading to samples that repeat far more than human text, which overstates generation quality. [episode]
- One Adapter, Every Resolution: Gated Low-Rank Adaptation for Remote Sensing VLMs — Remote sensing vision-language models (RS-VLMs) face a fundamental mismatch with natural-image counterparts because the same geographic object exhibits radically different visual evidence across ground sampling distances (GSDs) spanning multiple orders of magnitude. [episode]
- Semantic Watermarking for Malicious Image Manipulation Detection — Semantic watermarking for malicious image manipulation detection addresses the challenge of detecting subtle, adversarial edits to images that evade traditional pixel-level classifiers by anchoring an image's original semantic state into an invisible, recoverable reference. [episode]
- Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models — Chain-of-thought (CoT) reasoning is a dominant paradigm for inference-time scaling in large language models, yet the causal influence of individual steps on the final answer poorly understood. [episode]
- Spatial-Temporal Multi-scale Network for Screen Content Video Quality Enhancement — Screen Content Videos (SCVs) are challenging to enhance because they feature abrupt motion, scene switches, and high-frequency details like text and graphics, which conventional video enhancement methods often degrade due to the disruption of temporal correlations. [episode]
- WARP: A Unified Benchmark for Invisible Image Watermarking -- Robustness and Protection Against Attacks — Digital image watermarking is increasingly critical in media contexts, as emerging regulations and industry practices require marking AI-generated content and ensuring traceable sources to prevent manipulation or misuse. [episode]
- An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning — This empirical study investigates how different reward specifications influence unlearning outcomes in Large Language Models using Group Relative Policy Optimization (GRPO). [episode]
- Beyond Static Endpoints: Tool Programs as an Interface for Flexible Agentic Web Services — Tool programs are introduced as an executable representation of tool intent to address the representational bottleneck posed by static API endpoints in agentic web services. [episode]
- Deep Multimodal Fusion Detection through Spatial Mask and Channel Competition — Deep multimodal fusion for object detection has demonstrated good performance through mining modal characteristics, but existing feature-level fusion methods often suffer from overfitting or over-specialization by unifying modalities into a single representation space. [episode]
- Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets — Computer-use agents transform vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejection, calibration, miss-severity ranking, and spatial safety regions. [episode]
- KVEraser: Learning to Steer KV Cache for Efficient Localized Context Erasing — Post-hoc context erasing over the KV cache is challenging because a local edit has a global consequence: once a span has been processed, its influence propagates into the cached states of all subsequent tokens. [episode]
- Cross-Layer Discrete Concept Discovery for Interpreting Language Models — Cross-layer discrete concept discovery for interpreting language models introduces CLVQ-VAE, a novel framework that maps representations from lower to higher layers through a discrete vector-quantization bottleneck to collapse redundant residual features into compact, interpretab [episode]
- Opportunistic Target Selection: Early Directional Commitment for Query-Efficient Black-Box Adversarial Attacks — Black-box adversarial attacks that minimize only ground-truth confidence suffer from class drift, where perturbations wander through feature space without committing to a specific adversarial class, wasting queries on diffuse progress. [episode]
- Long-Tailed 3D Detection via Multi-Modal Fusion — This paper addresses the challenge of Long-Tailed 3D Detection (LT3D) in autonomous vehicle (AV) benchmarks, where existing datasets focus on common classes and neglect crucial rare objects. [episode]
- Dynamics-Inspired Diffusion for Foreground-Preserving Document Background Editing — As a fastidious and diligent AI researcher, I have meticulously analyzed all provided text segments (A, B, and C) to synthesize a comprehensive and highly detailed summary of the scientific paper concerning "Dynamics-Inspired Diffusion for Foreground-Preserving Document Backgroun [episode]
- 3D Software Synthesis Driven by Constraint-Expressive Intermediate Representation — This research introduces Scenethesis, a novel requirement-sensitive 3D software synthesis approach that utilizes ScenethesisLang as a constraint-aware intermediate representation to bridge natural language requirements and executable 3D software. [episode]
- Image AID via continuous-time reinforcement learning — This paper introduces Amortized Inpainting with Diffusion (AID), a novel framework that addresses the inefficiency of existing image inpainting methods by separating inference cost from training complexity. [episode]
- DySurface: Consistent 4D Surface Reconstruction via Bridging Explicit Gaussians and Implicit Functions — DySurface is a novel framework designed to achieve high-fidelity and temporally consistent 3D surface reconstruction in dynamic scenes by bridging the structural gap between explicit 3D Gaussian Splatting and implicit Signed Distance Functions (SDFs). [episode]
- Generalizing the Turing Test to Interactive Agents — The Generalized Turing Test (GTT) introduces a formal framework for comparing arbitrary AI agents based on indistinguishability, offering a dataset- and task-agnostic notion of relative intelligence. [episode]
- Distributionally robust linear regression through the lens of adversarial training — Distributionally robust optimization (DRO) studies parameter estimation under uncertainty in the underlying probability distribution, and this paper introduces Wasserstein DRO linear regression as a general framework that unifies key properties from special cases like square-root [episode]
- Depth Anything in 360: Towards Scale Invariance in the Wild — This paper presents DA360, a panoramic-adapted version of Depth Anything V2, designed to achieve scale-invariant depth estimation from 360° panoramic images in unconstrained environments. [episode]
- Effects of Structural Allocation of Geometric Task Diversity in Linear Meta-Learning Models — This paper investigates how task diversity in linear meta-learning models affects prediction performance, moving beyond simply increasing diversity to analyzing how this variability is structurally allocated relative to an underlying shared low-dimensional structure. [episode]
- Rethinking Cross-Layer Information Routing in Diffusion Transformers — Diffusion Transformers (DiTs) are central to modern visual generation, yet their fundamental residual stream inherited from standard Transformers remains largely unchanged, leading to underexplored issues in cross-layer information flow. [episode]
- Beyond LoRA vs. Full Fine-Tuning: Gradient-Guided Optimizer Routing for LLM Adaptation — The Mixture of LoRA and Full (MoLF) framework addresses the structural limitations of relying solely on either Full Fine-Tuning (FFT) or Low-Rank Adaptation (LoRA) by dynamically routing updates between both regimes at the optimizer level to ensure every expert receives full-batc [episode]
- GrepSeek: Training Search Agents for Direct Corpus Interaction — As a diligent researcher, I have thoroughly analyzed both provided texts concerning "GrepSeek" and its related work on Direct Corpus Interaction (DCI). [episode]
- Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching — Dataset distillation seeks to synthesize a highly compact dataset that achieves performance comparable to the original dataset on downstream tasks, which is crucial for practical applications where large datasets are prohibitive. [episode]
- MedRECT: A Bilingual Medical Reasoning Benchmark for Error Correction in Clinical Texts — MEDRECT introduces a novel, cross-lingual benchmark for medical error detection and correction, focusing on Japanese and English clinical texts. [episode]
- Adaptive Spectral Feature Forecasting for Diffusion Sampling Acceleration — Diffusion models are powerful tools for high-fidelity image and video generation, but their iterative sampling process is computationally expensive, posing a bottleneck for real-time applications. [episode]
- The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space — This research identifies and empirically validates a pervasive vulnerability in current Multimodal Large Language Models (MLLMs) known as the "Cartesian Shortcut." This shortcut occurs because models systematically exploit orthogonal grid-based layouts common in visual reasoning [episode]
- Looped Diffusion Transformer — Looped Diffusion Transformer explores an alternative way to scale computation in text-to-image models by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth without increasing parameter count. [episode]
- SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding — Multimodal Large Language Models (MLLMs) have shown rapid progress in understanding single videos, but their ability to reason across multiple independent video streams remains poorly understood. [episode]
- Frozen Memory Is Not Enough: Rethinking External Memory as Extraction — The paper investigates whether Engram-style hashed external memory remains useful and transferable across different language models, specifically focusing on whether frozen memory alone is sufficient or if it requires a compatible target-side reader for successful knowledge extra [episode]
- Reflectance Multispectral Imaging for Soil Composition Estimation and USDA Texture Classification — This manuscript proposes a robust and field deployable multispectral imaging (MSI) system combined with machine learning to accurately predict soil composition and the United States Department of Agriculture (USDA) texture classes. [episode]
- Prompting Image Generators for Training-free Primitive Shape Abstraction — This paper introduces a training-free pipeline that harnesses generalist visual knowledge from large generative image models to perform semantic shape abstraction, representing complex 3D objects as compact sets of geometric primitives. [episode]
- Context-Aware Classification and Grading of Sensitive Information in Online Conversational Health Data — This study introduces SALP-CG, a large language model-based extraction pipeline designed for classifying and grading privacy risks within large volumes of online conversational health data. [episode]
- ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule — This paper introduces Adaptive Reparameterized Time (ART), a reinforcement learning approach designed to solve the critical problem of timestep scheduling in score-based diffusion models. [episode]
- FaceLinkGen: A Re-evaluation of Identity Leakage in Privacy-Preserving Face Recognition and Face Anonymization Systems Using Simple Distillation — Privacy-preserving face recognition (PPFR) and perception-preserving face de-identification (De-ID) systems both retain identity signals that an adaptive attacker can extract, necessitating a re-evaluation of their security. [episode]
- Fiber-Resolved Microstructure Quantification from Multi-Shell Diffusion MRI using Detection Transformers — Fiber orientation and compartmental microstructure are central to characterizing white matter tissue in diffusion MRI, yet existing methods either resolve fiber orientations without quantifying microstructure or quantify microstructure while assuming fixed numbers of compartments [episode]
- EvoGuard: An Extensible Agentic RL-based Framework for Practical and Evolving AI-Generated Image Detection — The paper introduces EvoGuard, a novel agentic framework designed to address the critical and evolving challenge of AI-Generated Image (AIGI) detection. [episode]
- DuplexAct-Bench: Broadening Full-Duplex Speech Evaluation toward Proactive Interaction across Diverse Behavioral Requirements — Existing full-duplex speech benchmarks cover only subsets of realtime interaction behaviors, often under limited contextual conditions. [episode]
- Front-to-Back: Benchmarking Vision-Language Models for Asymmetric Cross-View Vehicle Re-Identification — Matching vehicles across front and rear cameras is difficult because they do not share a view and the vehicle’s appearance changes substantially. [episode]
- Latent Video Prediction for World Modeling: An Evaluation Uncovering Intriguing Favorable Evidence — This study provides a systematic evaluation of four matched-capacity frontier video foundation models—V-JEPA 2.1, V-JEPA 2, VideoPrism, and VideoMAEv2—across five robustness axes relevant to their deployment as video world models: feature discriminability, corruption robustne [episode]
- Understanding Multimodality in Generative Behavioral Cloning — This paper investigates how multimodality—the existence of multiple valid actions for a single observation—affects behavioral cloning policies, specifically focusing on action-chunking methods. [episode]
- Generalization in Nonlinear Least Squares via Learned Feature Geometry — This paper investigates how generalization in ridge-regularized nonlinear least-squares models is governed by data-dependent geometry rather than worst-case complexity. [episode]
- From Modes to Memories: Characterizing the Scale-Space Dynamics of Diffusion Models — Diffusion models can be interpreted as a family of deterministic dynamical systems indexed by noise scale, allowing researchers to characterize memorization through geometric dynamics. [episode]
- CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models — CombEval introduces a dynamic benchmark framework designed to rigorously evaluate the combinatorial counting capabilities of large language models (LLMs). [episode]
- Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation — On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher, and this research introduces NEIGHBORHOOD OPSD to improve student performance by leveraging complementary corrections from nearby parameter settings. [episode]
- How Much Human Label Variation Does Formal Semantic Structure Explain? Group-Level Effects and Item-Level Ceilings in NLI — Human label variation in Natural Language Inference (NLI) is increasingly viewed as signal rather than noise, prompting researchers to investigate what structures within language—specifically formal semantics—account for this observed disagreement. [episode]
- A Mesoscopic View of Transformer Weights Through Row and Column Scale Fields — Pooled statistics of Transformer weights obscure how magnitude is distributed across functional channels, while individual weights are too numerous to compare directly. [episode]
- ProtScape: A molecular structure and energy-aware representation for protein conformation generation — Understanding protein dynamics on microsecond to millisecond scales remains challenging, necessitating improved methodologies for analyzing and representing protein trajectories. [episode]
- Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves — Sampling several answers and keeping one a verifier scores highest is one of the simplest ways to buy accuracy at test time, and this paper derives a minimax cost law for certifying scaling curves, showing that while simple methods are cheap to draw, they often pay for incorrect [episode]
- FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection — FoCLIP introduces a feature-space misalignment framework designed to fool CLIP-based image quality metrics and simultaneously develop a detection mechanism for tampering. [episode]
- Sequential Bayesian Evaluation of Large Language Model Behavior — Sequential Bayesian Evaluation of Large Language Model Behavior provides a statistical framework for quantifying uncertainty in binary evaluation metrics when assessing black-box LLM behavior. [episode]
- Superficial Reflection or Genuine Thought? A Fine-Grained Cognitive Analysis of Large Reasoning Models — This research introduces a comprehensive taxonomy for analyzing Large Reasoning Model (LRM) reasoning steps, grounded in human cognitive science, to probe their underlying "psyche." By developing this fine-grained classification system and an auxiliary annotation framework called [episode]
- Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning — This work introduces Diffusion-Augmented Markov Decision Processes (DA-MDPs) to extend Maximum Entropy Reinforcement Learning (ME-RL), enabling diffusion models to sample from optimal policy trajectories by minimizing a tractable upper bound on the reverse KL divergence. [episode]
- ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing — Video scene text editing remains underdeveloped, particularly for precise local edits that must preserve original scene dynamics, and this work introduces ViTeX-Bench to systematically study these trade-offs. [episode]
- EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning — EvoIR-Agent introduces a novel, self-evolving image restoration agentic system designed to overcome the limitations of existing methods that struggle with zero-shot planning and lack compatibility with new tools or degradations. [episode]
- Direct Translation between Sign Languages — Direct sign-to-sign translation aims to bridge communication gaps between different sign languages without relying on spoken language intermediaries, which could significantly benefit 1.5 billion deaf and hard-of-hearing people worldwide. [episode]
- A Dynamical Theory of LoRA in Continual Learning — Low-Rank Adaptation (LoRA) dynamics in continual learning are characterized by a closed system of ordinary differential equations that precisely describe how low-rank updates affect feature reorganization and catastrophic forgetting across sequential tasks. [episode]
- EviRover: Reinforcing Agentic Perception Beyond a Glance — Visual perception is conventionally formulated as a one-shot prediction from a single glance at an image, an assumption that often fails in real-world scenarios requiring fine-grained details or external knowledge. [episode]
- Speech-based Psychological Crisis Assessment using LLMs — This paper proposes a novel large language model (LLM)-based framework for automated psychological crisis level classification using authentic speech data from support hotline calls. [episode]
- Fork-Think with Confidence — Parallel thinking has been successful for boosting LLM performance on reasoning tasks without retraining, but existing methods follow a "think-first-then-decide" paradigm, which leads to overgeneration and subsequent pruning. [episode]
- Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring — This study proposes an Edge-Aware and Content-Adaptive Feature Fusion Detector (ECAF-Det) designed to overcome the challenges of detecting faint, small, semi-transparent infrared gas plumes in cluttered industrial thermal scenes. [episode]
- GraphToxin: Reconstructing Full Unlearned Graphs from Graph Unlearning — GraphToxin is a novel full graph reconstruction attack against graph unlearning, designed to exploit residual traces left in graph neural networks (GNNs) after sensitive data has been purportedly removed. [episode]
- 4MT-VLM: How Coarse Is a VLMs Cognitive Map? — An agent that moves must recognize a place from a viewpoint it has never seen, and this paper introduces 4MT-VLM to diagnose how coarse the cognitive maps are in current vision-language models by testing their ability to maintain spatial understanding across viewpoint changes. [episode]
- Conformalized Regression for Continuous Bounded Outcomes — This paper introduces Conformalized Regression for Continuous Bounded Outcomes, developing conformal prediction intervals specifically tailored for regression models where outcomes are bounded (e.g., rates or proportions). [episode]
- ReDiF: Resource-Efficient Few-Step Diffusion Distillation via Reinforcement Learning — ReDiF introduces Reinforcement Learning (RL) as a novel optimization paradigm for accelerating diffusion models through policy-guided distillation, addressing the inherent slow sampling problem by training a few-step student model to approximate a high-step teacher. [episode]
- AD-Relight: Training-Free Banner Relighting via Illumination Translation with Diffusion Priors — The AD-Relight framework is a novel, multi-stage, training-free approach designed to relight custom Photoshop-generated ad banners seamlessly into existing scenes by adapting a diffusion-based relighting model at test time. [episode]
- LSR-Ben: A Logical and Scientific Reasoning Benchmark for Evaluating Process Reward Models — LSR-Ben introduces a comprehensive benchmark designed to evaluate Process Reward Models (PRMs) across both scientific and logical reasoning domains, addressing the critical limitation of existing benchmarks that are narrowly focused on mathematical reasoning. [episode]
- Project and Mix: Task-Semantic Prototypes for Few-Shot Image Classification — This paper introduces a novel, training-free framework for few-shot image classification by leveraging and intelligently mixing cross-modal prototypes derived from Vision-Language Models (VLMs) like CLIP. [episode]
- Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs — This research explores whether fine-tuned Large Language Models (LLMs) can be adapted to estimate the prevalence of four key vulnerability indicators—mental ill health, substance misuse, alcohol dependence, and homelessness—within unstructured UK police incident narratives. [episode]
- Structural Preservation Governs Data Augmentation in Deep Learning-Based Laser Speckle Material Classification — This study investigates how different data augmentation techniques affect deep learning models used for classifying laser speckle material images, arguing that augmentation effectiveness is determined by preserving the underlying structural statistics of coherent scattering rathe [episode]
- Estimation of the Label-Noise Transition Matrix with Performance Guarantees via Selective Classification — The proposed methodology provides a novel way to estimate label-noise transition matrices by framing each column estimation as a one-sided selective classification problem, which offers finite-sample performance guarantees and avoids the curse of dimensionality inherent in existi [episode]
- CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? — This research introduces chi-Bench, a novel benchmark designed to rigorously test the capabilities of AI agents in automating complex, end-to-end, long-horizon healthcare operations. [episode]
- Lens Flare Removal and Reconstruction — Lens flares are artifacts caused by unintended light paths passing through lenses when observing bright light sources, and this work introduces a novel pipeline that decomposes 3D scenes into lens flares and scene Gaussians to enable consistent flare removal and reconstruction. [episode]
- DIVER: Diving Deeper into Distilled Data via Expressive Semantic Recovery — Dataset Distillation (DD) aims to compress massive datasets into compact representations for privacy and efficiency, but classical DD methods often suffer from learning specific patterns that overfit on a prior architecture, leading to suppressed high-level semantics and poor cro [episode]
- Spike-driven Vision-Language-Action Model — Spike-driven VLA introduces a novel framework that enables end-to-end direct training for robotic manipulation using spiking neural networks, offering an energy-efficient alternative to conventional large Transformer models. [episode]
- Rethinking Multi-Image Re-Representation in Multi-Image Understanding — Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them. [episode]
- Minimax Additive Regression under Unknown Dependent Designs — Minimax Additive Regression under Unknown Dependent Designs investigates how to estimate additive models when both the design and the regression function are unknown, even when they depend on a potentially non-product random design. [episode]
- Time-adaptive infinite-dimensional Gaussian process regression on manifolds — This paper proposes a novel formulation of functional Gaussian Process regression tailored for spatiotemporal random fields on manifolds, utilizing an Empirical Bayes approach within an infinite-dimensional framework. [episode]
- Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models — Multimodal reward models have advanced significantly in text and image domains, but progress in video understanding reward modeling remains severely limited due to a lack of robust evaluation benchmarks and high-quality preference data. [episode]
- Functional Subspace, where language models can use vector algebra to solve problems — Large language models (LLMs) are increasingly demonstrating emergent abilities, making it crucial to understand their operating mechanisms for proper diagnostics and repair. [episode]
- Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting — This paper introduces "Ultrafast Spatial Texture-Atlas Splatting," a novel method for novel view synthesis that aims to achieve real-time 4K rendering by decoupling high-frequency texture details from geometry and view-dependent appearance features. [episode]
- Similarity-Distance-Magnitude Activations — We introduce a new activation function and estimator designed to decompose epistemic uncertainty in language models into interpretable signals: SIMILARITY, DISTANCE, and MAGNITUDE. [episode]
- BlockFormer: Transformer-based inference from genomic contact maps — BlockFormer introduces a novel, data-driven transformer architecture designed to infer per-entity parameters from interaction maps, such as those derived from Hi-C techniques, which is crucial for tasks like centromere localization. [episode]
- Prototype-Rule Neurosymbolic Regularization for Rank-Constrained Tensor Neural Networks under Label Scarcity — Rank-constrained tensor neural networks reduce parameterization but lack explicit class geometry constraints, and this study investigates whether a differentiable prototype-rule can provide a complementary inductive bias for Rank-R tensor learning under limited supervision. [episode]
- Cryo-Bench: Benchmarking Foundation Models for Cryosphere Mapping — Geo-Foundation Models (GFMs) have demonstrated strong potential for producing reliable maps even with sparse labels, but benchmarking them for Cryosphere applications has been limited due to a lack of suitable evaluation datasets. [episode]
- Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models — This research introduces novel scoring mechanisms, specifically the Query-Key Score (QK-score) and Attention Score, derived from select-and-copy attention heads within Large Language Models (LLMs), to unlock latent correct answers in Multiple-Choice Question Answering (MCQA) task [episode]
- Structure of Basic Human Values in Russian Social Media — This study presents a multi-stage classification framework for detecting human values in noisy Russian-language social media data, demonstrating that model predictions generally align with human judgments while systematically overestimating the Openness to Change value domain. [episode]
- Towards Formal Verification of Deep Neural Networks for Object Detection — Deep neural networks (DNNs) are increasingly deployed in safety-critical applications, yet their statistical nature makes them vulnerable to errors and adversarial attacks, necessitating robust defense mechanisms. [episode]
- Latent Diffusion Autoencoders: Toward Efficient and Meaningful Unsupervised Representation Learning in Medical Imaging — Latent Diffusion Autoencoders (LDAE) introduce a novel diffusion-based framework for efficient and meaningful unsupervised representation learning in 3D medical imaging, specifically focusing on Alzheimer’s disease (AD) using brain MR data. [episode]
- MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models — World Models are emerging as a frontier in computer vision, but their robustness remains largely unexplored, leading to the identification of latent hallucination—where predicted next latent states decode to scenes that never occur—which compounds autoregressively and silentl [episode]
- Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies — Cross-lingual transfer describes how knowledge in a source language benefits a target language, and this research investigates whether transfer scores can be reliably predicted using only freely available typological features, offering a zero-compute screening tool to prioritize [episode]
- A Generalizable and Explainable Framework for Synthetic Video Detection Using First-Digit Gradient Statistics — Ai video generators have become increasingly sophisticated, necessitating new, interpretable detection mechanisms that move beyond simple artifact recognition. [episode]
- Diagnosing Training Inference Mismatch in LLM Reinforcement Learning via a Zero-Mismatch Reference — Small token-level numerical disagreements between training and inference engines, known as Training–Inference Mismatch (TIM), can independently cause catastrophic failure in LLM Reinforcement Learning (RL) training by fundamentally altering the effective optimization problem. [episode]
- Right-Wing Rock or Just Rock? A Computational Linguistic Analysis of Frei.Wild — Rechtsrock is a subgenre of rock music that spreads right-wing ideology, and this study uses computational linguistic methods to determine whether the band Frei.Wild should be classified as politically right-leaning or as part of the general German rock genre. [episode]
- Mutual Equilibrium: Multimodal Representation Learning through Reciprocal Feedback — This work proposes a mutual feedback architecture, MEQ, that refines two inputs of possibly different modalities into coupled embeddings such that each embedding reflects the information of the other. [episode]
- Rethinking Uncertainty Quantification and Entanglement in Image Segmentation — Uncertainty quantification (UQ) is vital for safety-critical applications like medical image segmentation, but current methods often fail to properly account for how different uncertainty sources interact. [episode]
- Semantic Misalignment in Vision-Language Models under Perceptual Degradation — Vision–Language Models (VLMs) are increasingly deployed in safety-critical domains like autonomous driving, where reliable perception is paramount for decision-making. [episode]
- A Comprehensive Assessment Benchmark for Rigorously Evaluating Deep Learning Image Classifiers — The scientific paper introduces a new comprehensive assessment benchmark designed for rigorously evaluating deep learning image classifiers, addressing the critical need for more reliable and robust models in real-world scenarios. [episode]
- Beyond Uncertainty Sets: Leveraging Optimal Transport to Extend Conformal Predictive Distributions to Multivariate Settings — This paper introduces a novel framework for extending Conformal Prediction (CP) to multivariate settings by leveraging Optimal Transport (OT). [episode]
- NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance — General-purpose sentence embedding models often struggle to capture specialized financial semantics—especially in low-resource languages like Korean—due to domain-specific jargon, temporal meaning shifts, and misaligned bilingual vocabularies. [episode]
- Two-Step Data Augmentation for Masked Face Detection and Recognition: Turning Fake Masks to Real — The COVID-19 pandemic created an urgent need for robust masked face recognition and detection systems, but existing datasets remain insufficient. [episode]
- Differentiable Expectation-Maximisation and Applications to Gaussian Mixture Model Optimal Transport — This work presents methods to differentiate the Expectation-Maximisation (EM) algorithm, which is typically treated as a non-differentiable black box, and applies this differentiability to Gaussian Mixture Models (GMMs) through the Mixture Wasserstein distance (MW2). [episode]
- From DNA Design to DNA Slimming: Auditable Agentic Discovery of a Deletion-Only Designer — Compact regulatory DNA can free up space in vector payloads, reduce synthesis and assay burden, and expose which sequence features drive predicted activity. [episode]
- Semantic Chunking and the Entropy of Natural Language — This paper introduces a first-principles statistical model connecting natural language structure to its inherent uncertainty, quantified by entropy rate. [episode]
- Comparative study of adapting pre-trained models for driving behavior video captioning — This report examines and compares various fine-tuning and prompting methods applied to Large Language Models (LLMs) for driving behavior video captioning, aiming to induce a low-dimensional understanding of driving situations into the SpaceTimeGPT model. [episode]
- SheafStain: Sheaf-Theoretic Schr"odinger Bridge for Spatially and Biologically Coherent Virtual Staining — Current virtual staining approaches for gigapixel whole slide images (WSIs) suffer from spatial discontinuity artifacts when performing patch-wise inference, leading to inconsistent embeddings and catastrophic mismatches with ground truth. [episode]
- Voices of Freelance Professional Writers on AI: Limitations, Expectations, and Fears — The rapid development of AI-driven tools, particularly large language models (LLMs), is reshaping professional writing, and this research investigates how freelance professional writers perceive and utilize these technologies across various linguistic and ethical dimensions. [episode]
- Hyperspectral Image Dataset for Benchmarking on Salient Object Detection — This paper introduces a new hyperspectral image dataset specifically designed for benchmarking salient object detection, addressing a gap where existing models have been tested on limited, non-dedicated public datasets. [episode]
- Determining Vertical Displacement of Agricultural Areas Using UAV-Photogrammetry and a Heteroscedastic Deep Learning Model — The developed algorithm, based on heteroscedastic regression using a U-Net network, demonstrates high effectiveness in determining vertical ground surface displacements in agricultural areas and regions with similar characteristics. [episode]
- Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models —
- COBICount: Separating Object and Background Responses for Remote Sensing Object Counting Without Training on Target Data —
- Making Grid Beam Search Less Greedy —
- Exploring Heterogeneous Model Merging Approach for Complex Knowledge Transfer —
- EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos —
- InfoAgent: Traceable Generation and Repair of Evidence-Grounded Infographics —
- TTLab at Daleel 2026: STAR-Ar, Sequence Tagging for Argument Recognition in Arabic —
- Towards Optimal Inventory Control under Censored Demand: A Biased Sample-Average Approximation Approach —
- Awakening of the Buddha: Subspace Learning During Population-Loss Plateaus —
- QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code —
- PCB-MC: Missing Component Analysis in Printed Circuit Boards —
- Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification —
- Principal Component Regression Dominates all Monotone Spectral Filters for Linear Regression —
- Raw-Routed Mixture of Adapters: A Causal Intervention for Routing Collapse in Time Series Foundation Models —
- Synthetic Data Characterization via Training Dynamics —
- ResARC: Residual-Aware AutoRegressive Coding for Ultra-Low Bitrate Image Compression —
- DensePed-Lite: Quality-Aware Adaptive Detection for Dense Pedestrians under Occlusion —
- CAMOS: Coupled Oscillatory State-Space Model for Multimodal Clinical Time-Series —
- When the Right Answer Is Missing: An Arithmetic-Dependent Rejection Bottleneck in Jev —
- The Geometry of Randomized Smoothing on Feasible Sets —
- PartiCam: Camera Controlled Video Generation with Reward Guidance —
- Mitigating Representation Gaps in Amortized Bayesian Inference with Auxiliary Supervision —
- CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL —
- Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models —
- EffGS: Efficient and High-Fidelity Gaussian Splatting —
- Divide and Collapse: MAPF-Collapse via Exact Decomposition into Independent Sub-Instances —
- RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling —
- From Given to Gathered Evidence: Agentic Learning for Longitudinal Medical Reasoning —
- Invariant Shape Analysis of Surfaces with Spherical Topology —
- Compact Language, Complex Model Shifts: How and Where Ambiguity and Underspecification Affect LLMs —
- Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively? —
- KilometerVision: A New Frontier for Large-Scale Spatial Intelligence in VLMs —
- SPOON: Towards Coherent Compositional 3D Scene Generation from Uncalibrated Multi-view Images —
- Structural Limits of the Information-Theoretic Uncertainty Decomposition —
- GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed —
- FOMO: Forget the Concept, Don't Miss Out on the Scene in Selective Video Unlearning —
- Is This Evidence Decision-Critical? Learning to Verify Rule-Governed Decisions —
- ExpandDiff: Dynamic Range Expanding Diffusion for Single-Image HDR Reconstruction —
- D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders —
- Introduction to Computer Vision —
- SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations —
- FANVIDv2: Evaluating Video Super-Resolution by Face and Licence-Plate Recognition Under Compound Degradation —
- The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends —
- Typographic Attack Against VLM-based AI-generated Image Detection —
- MC-PanDA++: Simpler, Stronger, and More Robust Domain-Adaptive Panoptic Segmentation —
- OverForge: Reasoning Through Strategies and Tactics Helps Cooperative Lifelong Adaptation —
- LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation —
- FAST: Flow Any Scene Transformer —
- Revisiting On-policy Adversarial Black-Box Distillation: Calibrating Groupwise Reward Geometry for Effective Advantage Construction —
- MemCodex: Self-Programming Hierarchical Memory for Language Agents —
- Explore-over-Graph: Hybrid Embedding-LLM Reasoning for Knowledge Graph Question Answering under Incompleteness —
- Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model —
- Probabilistic Adversarial Training —
- Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior —
- P-SRM: Selective Recovery of Rejected Predictions in Visual Tracking —
- Spherical Interpolation for Backward-Compatible Multimodal Representations —
- Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard —
- BayesNDE: Bayesian Generative Modeling for Neural Density Estimation —
- When a Kindergartener Solves Calculus: Measuring Capability Leakage in Role-Prompted Reasoning Models —
- Cognitive Enhancement: Rethinking the Necessity of Role-Playing for Large Language Models —
- FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy —
- GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning —
- A PyTorch Library for Hyperspectral Image Models: Technical Report —
- LLM Persona Unlearning —
- Grounding with Confidence: Controllable Generative Video Temporal Grounding —
- OPSRD: On-Policy Self-Role Distillation —
- Markovian Dynamics Enforcer: Feasibility Preserving Correction on Learned Dynamics Manifolds —
- Learning Where to Look: Anatomical Grounding and Guided Attention for Cardiac MRI Vision-Language Models —
- The Concrete-Arbitrary Gap: Kinship Reasoning in LLMs Is Not Indifferent to Presentation —
- NavHarness: Adaptive Goals for Agentic Vision-Language Navigation —
- MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models —
- CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding —
- Super-Resolving Unseen Hyperspectral Sensors at Any Scale via Spatial Operators —
- AdaGEPA: Adaptive Feedback Allocation for Reflective Prompt Optimization —
- RoPE at the End of Its Rope? Theory, Diagnosis, and Mitigation of Long-Context Failures —
- Reliability-Aware Checkpoint Selection for Domain Generalization —
- Learning to Reason with Compressed Context: Ground-Truth-Free Adaptation of OmniLLMs via Self-Distillation —
- Reconstructing the Dynamic World: A Representation-Centric View of 4D Scene Reconstruction —
- UBTree: Parallel Tree Drafting via Unigram and Bigram Models for Speculative Decoding —
- Overview of BioASQ 2026: The fourteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering —
- Can We Anticipate Violence? Multimodal Learning from Pre-Incident Behavioral Cues —
- OPTS-TTPO: Enhancing Finite-Sample Policy-Gradient Learning with Tree Search —
- MGhana-ST: A Low-Resource Speech Translation Dataset for Ghanaian Languages and an Analysis of Multilingual Training Trade-offs —
- Gromov-Wasserstein Distillation for Inductive Multi-View Embedding —
- Proximal Balancing for Causal Effect Estimation under Unmeasured Confounding —
- LARC: Low-Rank Adaptive Residual Connections for Learning in Frozen Models —
- From Tweets to Trades: Analyzing the Influence of Public Mood over Stock Market Performance in Turkiye —
- LongEmo: Towards Emotion Understanding and Reasoning in Long Videos —
- AutoDataBench: A Data-centric Testbed for Accelerating Auto Research —
- OverdoseMoE: A Multi-Expert Framework for Opioid Overdose Risk Prediction —
- Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training —
- Persistent Context Graphs for Efficient Memory Compaction in LLM Agents —
- On the (In)effectiveness of AMR Augmentation for Large Language Models —
- Debias It Yourself: Teaching LLMs Cognitive Bias Mitigation Interventions —
- Learning Functional Subspaces for Neural Network Compression —
- VR-JEPA: Learning Contrastive-State Latent Guidance for Generation-based Video Reasoning —
- Game-Guided Skill Discovery through Self-Play for Playable Agent Control —
- Index-Translate: A Multilingual Translation Model Family -- Text, Speech, Controlled Dubbing, and Long-Document Translation —
- SCB: SpeechConversationBench for Evaluating Multi-Turn Reasoning in Speech-to-Speech Models —
- Recognition of Urbanized Areas in UAV-Derived Very-High-Resolution Visible-Light Imagery —
- Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models —
- Comparison of techniques for fine-tuning open-weight models for entity extraction from radiology reports —
- StreamRig: Exploiting Intra-Rig Geometry for Streaming Multi-Camera Odometry —
- ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents —
- Belief-Aware Multi-Agent Path Finding under Map Uncertainty —
- Signal Processing over Product DAGs: Causal Shifts and Filters —
- Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning —
- Disentangling Computation in Multi-Task Neural Networks with the Green's Operator —
- How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text —
- Scaling Laws for Looped Mixture of Experts —
- GLARE: Generating Listening Heads with Appropriate Reactions —
- Atomizer-IO: Beyond Pixels, Patches and Grids —
- I Have a Stream: Making Self-Supervised Learning Work on Continuous Video —
- EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery —
- Image Classifiers are Efficient Self-Supervised Video Representation Learners —
- AssemblyWorld: Rethinking 3D Assembly with General-Purpose Agents —
- Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model —
- Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text —
- Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis —
- Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces —
- Biased Heritage: How Datasets Shape Models in Facial Expression Recognition —
- Large Language Models are Approximate Survival Estimators —
- EHR2Trace: Auditable EHR Data Infrastructure for Patient World Models and Clinical Agents —
- TomasuLLM: Out-of-Order Speculative Execution for LLM Agents —
- Automatic estimation of verbal fluency index in people with Motor Neuron Disease using ASR alignment and pause modelling —
- The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models —
- TutlAit v1: a crowdsourced Moroccan Tamazight speech dataset with Arabic transcriptions and regional accent labels —
- Conformal Factuality Control for Multi-Hop Retrieval-Augmented Generation —
- Framing the Narrative: Ideological Mimicry in Large Language Models —
- ContextAdapt: Evaluating Contextual Adaptation and Value Alignment in LLMs —
- NinaXander: Feasibility and Limits of Composing Frozen Language Models Across Architecture Families via a Shared Latent Space —
- Robust LassoNet: Enhancing Feature Selection in Neural Networks via Robust Loss Functions —
- Evaluating Multi-Task Morphological Concept Learning for Pulmonary Nodule Malignancy Assessment in 3D CT —
- Which Models Work Well Together? Measuring Heterogeneity for LLM Team Selection —
- Masked Swingers: Harnessing Data Augmentation to Advance Autoencoders for Self-Supervised Learning —
- GaugeVLM: Structuring Spatial Supervision with Measured Geometric Interventions —
- It Takes Little to Rewrite Perception: Targeted Semantic Substitution in Vision-Language Models at epsilon at most 4/255 —
- Strike a Chord! Modal Kinetic Typography —
- ExploreNet: Learning Where to Explore in Diffusion GRPO —
- Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling —
- EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making —
- Learning Semantic Inpainting for Animatable Gaussian Head Avatars —
- TrackFish3D: Self-Supervised 3D Tracking of Schooling Fish from Multi-view Videos —
- Halluscoring 2026: The first shared task on llms hallucination detection and answer verification —
- Evaluating Language Model Safety Across Long Adversarial Conversations —
- Beyond Mode Collapse: Generating Diverse Synthetic Expert Conversations via Generative Flow Networks —
- On the Off-Policy Teacher in On-Policy Distillation —
- Inductive Visual Logic for Few-Shot Out-Of-Distribution Adaptation in VLMs —
- Acceleration of Diffusion Language Model through Discrete Average Generator —
- Composition, Not Conversation: VLMs Lose the Scene, Not the Thread —
- Doc2LoRA Provides Decodable Representations of Scientific Ideas —
- Lower Bounds for Linear-Oracle Online Learning —
- PhyProbe: Rethinking Physical Consistency Evaluation in Generated Videos —
- Learning to Plan from Random Exploration —
- Fine-Tuning Diffusion Language Models with Context Selection and Target Weighting —
- Team MSU GenText-Forensics Challenge 2026 Technical Report —
- Evaluating Whether LLMs Can Reliably Connect the DOTs? —
- ArgGYM: A Procedural, Engine-Verified Benchmark for Structured Defeasible Reasoning —
- VidHarness: Evolving Agent Harnesses for Cost-Efficient Long Video Understanding —
- LoopVL: Recurrent Visual Intelligence —
- Policy-Conditioned AI-Use Detection: An Evidentiary Framework for Academic Publishing —
- MOBA-VL: Event-Localized Multi-Turn Reinforcement Learning for Real-Time MOBA Commentary —
- Audible World Models: Spatially Aware Sound Generation for 3D Worlds —
- What Pretraining and Midtraining Make Learnable from Rewards? —
- Reach Into The CHOIR: Free-List Elicitation Uncovers Distinct Model Voices in LLM Ensembles —
- Grokking through the Lens of Minimum-Norm Interpolation —
- Does Gradient Conflict Predict the Understanding--Generation Trade-off? A Controlled Audit of Conflict-Metric Validity in Unified Multimodal Models —
- PAMI: Part Anchored Motion for Text to Human-Object Interaction Generation —
- The Backdrop Exposes What the World Around an Agent Costs It —
- Curating Synthetic Data for Task-Specific Visual Perception —
- Caption-Mediated Perceived-Safety Estimation for Pedestrian Routing —
- KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy —
- RetroGEF: Dynamic Graph Edit Flow for Single-Step Retrosynthesis —
- Beyond Layers: Position-Resolved Gradient Conflict and Position-Aware Modulation for Unified Multimodal Models —
- Multidimensional Observer Model and Perceptual Dimensions of Human Image Quality Assessment —
- Personalized State-Transition-Aware Memory for Clinical Agents —
- DEdit: Iterative Draft Editing for Speculative Decoding —
- GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction —
- ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM Weights —
- Generative sequence modeling for infinite memory processes via predictive states —
- Shifting Mechanisms: How Positional Encoding Choice Shapes In-Context Retrieval —
- ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing —
- MedKIT: Evaluating Knowledge Integration and Generalization in Large Language Models —
- Towards Universal Wasserstein Barycenters through Flow Matching —
- Detail in Context: A Dual-Scale Machine Learning Framework for Mycosis Fungoides Detection —
- LongTake: Learning to Sustain Dynamics in Long-Horizon Video Generation —
- Towards Model as a Library: Offline, Community-Sourced AI for Low-Resource African Languages —
- Retargeting Motions to Diverse Skeletons via Learnable Flattening —
- Restoring without Forgetting: Filter-Level Continual Image Restoration via Parameter-Space Integrated Gradients —
- StereoGaussians: Feed-Forward 3D Gaussian Splatting from Stereo Images —
- Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions —
- PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation —
- Aperture: Training-Free Multiscale Concept Bottlenecks for Remote Sensing —
- Beyond Oracle Communication: Benchmarking Interactive Intent Alignment Under Miscommunication and Evolving User Intent —
- After a Decade: Bringing Shadow Removal into the Real World with Agentic Training Data —
- Learning-Enabled Estimation: Tight Characterizations under Sample Selection Biases —
- StreamDecisionBench: Evaluating Decisions in Force on Evolving Language Streams —
- Exo2EgoHOI: Hand-Object-Interaction Aware Exocentric-to-Egocentric Video Generation —
- HIGS: Hierarchical Implicit Grids for Joint Geometric and Semantic Scene Understanding —
- When Scientific Contradictions Are Lost in Translation —
- Eulerian Motion Reconstruction for Water Scenery —
- Marking Contour Tones in Yor` u b' a: A Typographic and Computational Proposal —
- Strong Multilingual Privacy Tagging at Encoder Speed —
- STEPS: Scene Text Editing with Preserved Style Using Diffusion and Contrastive Style Encoding —
- Template-Search Domain Adaptation via Multi-Stage Feature Alignment for Cross-Modal Object Tracking —
- Vision-Language-Action Autonomous Driving Agent with Language-based Memory —
- ChartRevise: A Dataset and Evaluation Protocol for Exact Chart Editing via Code —
- Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivity —
- Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation —
- Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives —
- ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images —
- Unveiling the Value of Motion for Cinematic Camera Trajectories —
- EPIC: Epipolar-Consistent 360 Immersive Stereo Video Generation —
- No Corners Cut: State-Grounded Transitions for Mid-Stream Prompt Switches in Video Generation —
- SCALE: Synthetic Calibration via Agreement Labeling in Embedding Space —
- Hard-Region Supervision: #1 on the Waymo Open Dataset 2D Video Panoptic Segmentation Leaderboard —
- SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models —
- Soft Spatial Reasoning —
- MetaSteer: Context-Conditioned, nonlinear Steering via Attention-Projection Adaptation —
- UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement —
- Matisse: Evidence-Space Reasoning for Active 3D Reconstruction —
- Consensus-Aware Multi-Source Fusion for Reference-Guided Camouflaged Object Detection —
- Hear the World in Stereo: Learning Dynamic Spatial Correspondence for Immersive Joint Video-Audio Generation —
- Agentic Relative Camera Pose Estimation via Learned Ranking and Verification —
- Event-Driven Refresh and Recurrence Memory to Reduce Stale Grounding in Referring Video Object Segmentation —
- Distill the Visual Evidence, Not Just the Answer: Cross-World On-Policy Distillation for Vision-Language Models —
- Training LLM Judges from Language Feedback via Position-Selective Self-Distillation —
- Recovering Off-Policy Supervision for Speculative Decoding —
- Evaluating Persistent Calibration under Evolving Model Knowledge —
- Overlap, Unique and Conflict: Can LLMs Extract What They Can Recognize? —
- Uncovering Uncontrolled Repetition through Residual Stream Dynamics —
- Blackboard Intelligence Can Surpass Autoregressive on Globally Constrained Problems —
- StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning —
- CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models —
- DCM-SAM: Defect-Conditioned Mixture of LoRA Experts for NPU-Deployed AM Defect Segmentation —
- Can Terminal Agents Trust Their Own Verification? Diagnosing and Improving Self-Verification —
- You're Hired: Strategic Model Selection for LLM Collaboration —
- When Reasoning Goes Astray: Attention Dynamics of Uncontrolled Reasoning —
- Whose Voice Survives the Summary? A Voice-Retention Audit of LLM Employee Listening —
- Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video —
- BARRAC: Adaptation of an English Aspect-based Sentiment Analysis Approach for Classification Tasks in Arabic Dialects —
- DecoMoE: Decoupling Visual Propagation and Expert Computation for Efficient Multimodal MoE Inference —
- SparLeak: Privacy Leakage from Sparse Attention in LLM Inference on Shared GPUs —
- Forging LLM Authorship Fingerprints with Targeted Rewriting —
- Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head —
- FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation —
- OpenJev-RLCD: A Working RLCD Implementation —
- Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis —
- Decoupling Spherical Reasoning from Dense Prediction for 360 Depth Estimation —
- TRACE: Target-Aware Retrieval, Attributed Evidence, and Contract-Constrained Extraction for LitTraceQA —
- AdaOcc: Adaptive 3D Occupancy Prediction for Embodied Tasks —
- Sharp Statistical Rates for Asynchronous TD Learning with Markovian Data —
- Learning Continuous Neural Representation of Stochastic Hybrid Systems —
- K2P: Label-Free Knowledge to Prompt Distillation —
- MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding —
- Amortized Data Borrowing with Exchangeability-Aware Neural Posterior Estimation —
- FLOW: Feature-Level Optimal Warping for Generalized Remote Physiological Measurement —
- Warm-starting PDE solvers with any-dimensional machine learning —
- Flow Matching under Noisy Latent Structure: Beyond Exact Low-Dimensional Support —
- GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis —
- From Image Interpretation to Clinical Reasoning: Upstream Physician-Context-Aware Multimodal Learning with Causal Reinforcement Learning —
- World-as-Graph: Relational World Modeling Through Latent Space Graphs —
- On the Relaxation of Conditional Independence Assumption for Image Segmentation —
- Targeted Retrieval, Compact Representations: How CoT Reasoning Improves Long-Context Counting —
- Beyond Spatial-Domain Supervision: A Relation Constrained Space for Multi-Modal Image Fusion —
- Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment —
- Fairness Beyond a Single Run: Training-Seed Variability in Speech LLM Adaptation —
- PARK: Accurate Block Retrieval for Sparse Attention in Video Diffusion Transformers —
- Mitigating Object Hallucination in Large Vision-Language Models via False Discovery Controlled Visual Data Splitting —
- MeshOctave: Vertex Split-and-Rewire Cascades for Native Mesh Generation —
- Anchoring Adversarial Trajectories to Data Manifolds: A Bilevel Transfer Optimization Framework —
- When Clipping Reverses Correction: Failure Dynamics of Pointwise Forward-KL On-Policy Self-Distillation —
- Settle: Learning When to Stop Reasoning —
- Not all solutions are created equal: An analytical dissociation of functional and representational similarity in deep linear neural networks —
- The Invisible Language Tax: Token Premiums of French and Regional Languages in 2026 LLM Tokenizers, and a French-Optimized Prototype —
- When Integral Meets Decomposition: A Signal-Level Self-Supervised Feature Decompose Paradigm for Multi-Modal Image Fusion —
- Evidence First, Arithmetic Second: A System Report and Failure Analysis for DocSem —
- Minimax rates for learning spectral Barron functions by deep ReLU neural networks —
- Frame Differential On-Policy Self-Distillation for Video Reasoning —
- Persistent Watermarking of Text-to-Image Models —
- A Rank Graduation metric for Algorithmic fairness —
- A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review —
- TED:Text-Axis Evidence Decomposition for Prompted Anomaly Localization —
- Switching Linear Attention —
- RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement —
- BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers —
- Structure vs. Chain-of-Thought: Evaluating LLM Criteria Extraction for Depression Severity —
- TSMD: Temporal-Stream Modality Dropout for Robust Video Highlight Detection —
- Agentic Tool-Augmented Reasoning for Explainable Image Forgery Detection —
- CORE: Conflict-Oriented Reasoning Elimination for Verifiable Language-Model Search —
- LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models —
- Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue —
- MRI Super-Resolution with RCDM/WaveMix and Task-Aware Segmentation —
- UGOD: Uncertainty-Guided Opacity and Dropout for Sparse-View 3D Gaussian Splatting —
- Steepest Guidance: A Practical and Principled Approach to Inference-Time Alignment of Flow and Diffusion-based Models —
- DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency —
- False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents —
- Bongard: Training Machine Intuition —
- CamAgent: An LLM-Agent Framework for Multi-Species Camera-Trap Workflows —
- Beyond Local Linearity: Scale-Resolved Geometry of Learned Image Encoders —
- GRC-Pose: Generation-Reconstruction Correspondence for Prior-Free 6D Object Pose Tracking —
- Diagnosing On-Policy Self-Distillation for Reasoning Language Models —
- Is Better Teacher Supervision Enough? Unlocking Student-side Learning in Multimodal On-Policy Distillation —
- How to Reduce Localization Ambiguity? Geometry-Semantic Constrained BEV Representation Learning for Satellite-Ground Localization —
- GeoGAT: Bidirectional Temporal Sampling Meets Hierarchical Graph Attention for Global Video Geo-localization —
- Perceptual Color Difference Modeling Using Machine Learning and Human Similarity Judgments —
- Uncertainty-Aware Consistency Distillation for Few-Step Video Generation —
- Feature-Aware Token Attack for Compression-Triggered Stealthy Failures in Large Vision-Language Models —
- Asking the World: Generalist Physical Reasoning through Agentic World Modeling and Probing —
- When Can Text Replace Vision? Structural Bottlenecks in Diagram Reasoning —
- MindWorldBench: Evaluating Mental-State-to-Behavior Reasoning in Image-to-Video Generation —
- OP-CAD: On-Policy Clean-Audio Distillation for Robust Audio-Visual Reasoning —
- DAGent: Evaluate-then-Grow Planning for Deep Research Agents —
- TripleFlow: Training-Free Video Object Removal by Bridging Residual Editing and Native Generation —
- Asymptotic Properties of Support Vector Machines in High-Dimension, Low-Sample-Size Settings under a Spiked Model —
- Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift —
- ViLegalExpert: A Large-Scale Benchmark for Vietnamese Legal Retrieval and Question Answering from Real-World Consultations —
- Dynamics to decision: A mathematical theory of Lyapunov spectra and decision boundaries in deep classifiers —
- Uruqi: Learning Spatial Cognition from Visual Experience —
- DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence —
- Argument Structure Prediction in Online Conversations: A Comparative Study of Modeling Paradigms and Task Architectures —
- Emergent Multi-View Geometry Through Self-Distillation —
- RAIM: Robust Aggregation of Inexpensive Models for Hallucination Detection —
- Concept Subspaces Compute Beyond the Logit Lens: A Weights-Only Test for Locating Representations Upstream of Readout —
- Universal Cross-Prompt Adversarial Attacks on Promptable Concept Segmentation —
- PLRS-IC: A Dual-Calibration Framework for Chest X-Ray Vision-Language Alignment —
- MegaAvatar: Controllable Talking Avatar Generation —
- A Tilted Bowl Is Not a Slippery Slope: Compressing Looped Models —
- BMASH: Ball-Motion-Aware Soccer Header Spotting —
- Rethinking Generative Image Compression at Extremely Low Bitrates —
- Discrete Score Matching Enables Causal Discovery from Count Data —
- TexTailor: Texture-Preserving Video Virtual Try-On via Adaptive Garment Conditioning —
- Understanding as No-Arbitrage: Bounded Dutch Books as a Definition and Training Objective for Language Models —
Important terms
- BlockFormer
- This is a transformer-based inference method used on genomic contact maps to figure out structural information about DNA.
- Green's Operator
- This tool helps researchers understand the different computational paths inside multi-task neural networks by disentangling them.
- Query-key alignment
- This technique is used in LLMs to find the latent correct answers, helping unlock hidden information within the model.
- CombEval framework
- This framework tests if large language models can actually perform complex counting tasks reliably, going beyond just high scores.
- QuantCode model
- This model specializes LLMs for writing executable trading code by fine-tuning them with real financial examples.