AI papers — 2026-10-07
The CANDLE project aims to improve noninvasive brain source imaging by using cortical null-space decomposition. This method was explored because it has shown more interpretable results than traditional approaches when decomposing the cortical null space into meaningful components.
Another area of work involved observable neural ordinary differential equations for identifying causal forecasting in continuous time. This suggests that researchers can track how brain activity evolves over time with greater precision, which connects to the CANDLE work by potentially helping to better model the underlying neural signals being imaged.
Structural-frontier evaluation was used to uncover hidden failures in ADMET models, meaning current predictive models for drug properties are missing critical failure points when tested against more complex biological structures. This finding is important because it suggests a need for more robust testing environments before moving forward with molecular discovery efforts.
Work on PertMind uses reinforcement learning on cellular perturbation data to elicit emergent reasoning in large language models. This research explores how to push AI beyond simple pattern matching toward genuine biological understanding, which complements the structural insights gained from the ADMET model evaluation.
The most pressing work today concerns building more reliable systems for understanding complex biological data. Promising results were seen from effective biological representation learning by masking gene expression, which shows that strategically hiding parts of a genome helps the model learn better underlying biological structures. This is important because it suggests a path toward creating models that grasp the true function of genes rather than just memorizing sequences.
A key development involves language models and how they evaluate things, specifically breaking the mirror by using activation-based mitigation of self-preference in LLM evaluators. This means researchers are trying to stop these large language models from unfairly favoring their own outputs when judging other things, which is crucial for trustworthy AI applications. This work connects to efforts in integrating high-precision computation and reasoning through PiERN, which uses token-level routing to manage this kind of complex processing within multimodal models.
Research also showed that language model ratings of depression reflect the rater more than the patient, a finding that is significant for understanding bias in mental health assessments. This points to a need for careful calibration when using these tools in clinical settings. Furthermore, there is ongoing research into artificial hivemind concepts exploring the open-ended homogeneity of language models and beyond, suggesting a future where these models might achieve a more unified level of understanding across diverse domains.
The work on stabilizing off-policy training for long horizon agents matters because it directly addresses the reliability of complex AI systems that need to plan many steps ahead. Researchers explored turn-level importance sampling and clipping-triggered normalization to stabilize this process, which essentially means they found a way to make the learning process more stable when an agent is trying to complete very long sequences of actions.
This effort builds upon other alignment work by focusing on practical implementation challenges for agents that operate over extended periods. It is less significant than the foundational work on unbiased reward modeling from implicit feedback, but it provides a necessary technique for making those aligned models actually perform reliably in long-running tasks.
Another piece of research involved dOPT, which differentiates conic optimization through geometric reduction to improve how certain mathematical problems are solved. This is a more theoretical contribution than the practical training stabilization work, though it offers new ways to optimize underlying model behaviors.
SchemaGraphSQL addresses the efficiency of linking text queries to database schemas using pathfinding graph algorithms for text-to-SQL tasks on large databases. This method is distinct from the linguistic analysis done by VietBinoculars, which uses a zero-shot approach to detect Vietnamese LLM-generated text.
Work on cross-lingual activation steering for multilingual language models aims to improve how these models process information across different languages, which is relevant for building globally capable systems. This contrasts with the study on emotion concepts in LLMs and humans, which investigates whether current language models truly grasp human emotional concepts or if they are merely too categorical to do so.
The most pressing work today concerns understanding how different goals within a single system can interfere with each other, which is crucial because it dictates the stability of complex AI applications. Researchers looked at how cooperative profiles predict multi-agent LLM team performance in science workflows, suggesting that knowing who is working together helps predict success in collaborative AI tasks.
This work builds upon research into tone-conditioned curriculum learning for low-resource Bantu speech recognition, showing how adjusting the learning path based on emotional context improves accuracy when data is scarce. Furthermore, researchers explored cross-cultural value attribution in large vision-language models to see if these models can correctly interpret differing cultural perspectives embedded in visual data.
A related effort involved reinforcement learning over predictive distributions for LLM regression, which aims to make the regression predictions more robust by training the model on the distribution of potential errors rather than just single outcomes. This approach is connected to real-time generation of game video commentary with multimodal LLMs, where pause-aware decoding methods are tested to ensure generated text matches visual cues in a dynamic environment.
Finally, researchers investigated token-level off-policy learning for faithful generation under distribution shift, which deals with keeping the generated output accurate even when the data it encounters during testing is different from what it was trained on.
The most significant development centers around FedCoT, which addresses the communication inefficiency when training large language models across distributed systems. This method enhances reasoning capabilities by allowing models to communicate efficiently without needing massive data transfers between nodes, making scaling complex reasoning tasks feasible in real-world, decentralized environments.
TabiBERT introduces a large-scale modern BERT foundation model specifically for the Turkish language, creating a unified benchmark that allows researchers to compare different linguistic approaches systematically. This work builds upon the need for robust models by providing a specific, high-quality starting point for language tasks.
RAM-Net explores linear-time sequence modeling using sparsely addressable state representations. This approach is crucial because it tackles the computational bottlenecks inherent in processing very long sequences efficiently, offering a way to handle longer inputs without incurring quadratic time complexity penalties.
Understanding moral reasoning trajectories in large language models through probing helps explain how these models arrive at certain conclusions, which is vital for building trust in AI systems. This research uses probing techniques to investigate the internal mechanisms of decision-making within these models.
Sentiment analysis on French synthetic social media provides a practical test case for model performance under specific linguistic noise conditions. This work examines how language models handle nuanced and potentially misleading social data streams.
The most important work today involved testing how different adapter placements affect the performance of a dominant adaptation module, which is crucial because it directly impacts how efficiently we can fine-tune large models for specific tasks. Researchers looked at various placement strategies to see which one yielded the best results when adapting a model.
A key finding emerged from this investigation: certain adapter placements significantly improved the model's ability to perform its target function compared to others, suggesting a better way to integrate new knowledge into the existing system. This is important because it points toward a more robust and effective method for tailoring large language models.
Another piece of work focused on emotion recognition in sign language conversation, which matters because it pushes the boundaries of understanding nuanced human communication beyond just spoken words. Researchers explored how to recognize subtle emotional cues conveyed through sign language, aiming to build better systems for interpreting non-verbal social context.
Then there was the study on SubtleMemory, which serves as a benchmark for fine-grained relational memory discrimination in long-horizon AI agents. This work is significant because it establishes a standard for how well an agent can remember and relate distant pieces of information over time, which is vital for complex decision-making.
The research on evaluating large language model raters for German open-response clinical questions provided insights into evaluator bias and agreement among physicians. This matters because it helps us understand the reliability of AI systems when they are tasked with making judgments in sensitive medical contexts.
Finally, researchers saw work on refusal-gated decoding, which attempts to preserve the refusal behavior of models even when they are sampled at high temperatures. This is important for ensuring safety and adherence to guardrails in generative AI applications.
The most critical piece of work today involves understanding how large language models lose coherence across multiple conversational turns, which is vital because sustained interaction requires maintaining context. One study explored component and dimension sparsity within transformer refusal mechanisms to see if these structural simplifications affect the model's ability to refuse appropriately. This relates closely to a different line of inquiry examining zero-shot visualization, where researchers looked into exploring text corpora using user-prompted axes to see if models could correctly interpret spatial relationships without explicit training on those axes.
Another significant finding comes from work characterizing then distilling mechanistic reasoning in large output spaces, which attempts to map out the internal logic of these complex systems. This effort builds upon investigations into seeing isn't knowing, where researchers tested whether vision-language models can correctly identify when they should withhold an answer to a spatial question. This idea of knowing when not to answer is also touched upon by research focused on diagnosing fine-grained inconsistency classification in financial disclosure text, which tries to pinpoint exactly where textual contradictions arise.
Finally, there is the work on verdicts without annotated evidence, which investigates whether rejection sampling or label-only post-training methods can recover missing evidence after a model makes a decision. This moves from understanding internal reasoning to assessing how we can verify the outputs of these complex systems.
The most significant development today involves Mask-Guided KV Cache Eviction in Block Diffusion Language Models because it directly addresses the computational bottleneck in handling long sequences within these models. This technique attempts to manage memory usage by selectively discarding parts of the key-value cache based on masking information, which is crucial for scaling these large diffusion architectures.
This approach builds upon prior work that explored model compression for neural machine translation in the biomedical domain, where researchers investigated methods to reduce model size while maintaining performance in specialized medical contexts. Furthermore, the work on JudgeMoE introduces a new method for distribution aggregation to enable LLM-as-a-judge capabilities, meaning using one large language model to evaluate the outputs of another.
A less immediate but still important piece is TIDE 2.0, an open engine designed for keyed de-identification of clinical notes, which focuses on privacy by removing sensitive information from medical text. This contrasts with the theoretical work exploring a theory of platonic representations in language models, which delves into how these models might represent concepts internally.
Finally, the effort to identify introspection from the inside touches upon understanding model behavior, complementing Kurate's work on scalable scientific quality analysis which assesses the reliability of scientific outputs.
Today's papers
- CANDLE: Cortical Null-Space Decomposition for Noninvasive Brain Source Imaging. [paper]
- Observable Neural ODEs for Identifiable Causal Forecasting in Continuous Time. [paper] [episode]
- Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models. [paper] [episode]
- PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data. [paper] [episode]
- Neural Petri flows for chemical reactions. [paper]
- Confidence-Ordering Reversal under Contextual Priors in Neural Decoding. [paper]
- Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning. [paper] [episode]
- Information-Dense Synthesis for Molecular Discovery. [paper]
- From the Drosophila Visual Connectome to General-Purpose Computer Vision. [paper]
- Language-model ratings of depression reflect the rater more than the patient. [paper]
- Effective Biological Representation Learning by Masking Gene Expression. [paper] [episode]
- Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family. [paper] [episode]
- Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators. [paper] [episode]
- PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning. [paper] [episode]
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond). [paper] [episode]
- Quantifying the Gap between Understanding and Generation within Unified Multimodal Models. [paper] [episode]
- Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks. [paper] [episode]
- Unbiased Reward Modeling from Implicit Feedback for LLM Alignment. [paper] [episode]
- dOPT: Differentiating Conic Optimization via Geometric Reduction. [paper] [episode]
- SchemaGraphSQL: Efficient Schema Linking with Pathfinding Graph Algorithms for Text-to-SQL on Large-Scale Databases. [paper] [episode]
- Too Categorical to be Human: Emotion Concepts in LLMs and Humans. [paper] [episode]
- VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text. [paper] [episode]
- Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization. [paper] [episode]
- Cross-Lingual Activation Steering for Multilingual Language Models. [paper] [episode]
- Uncovering Cross-Objective Interference in Multi-Objective Alignment. [paper] [episode]
- Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches. [paper] [episode]
- Cross-Cultural Value Attribution in Large Vision-Language Models. [paper] [episode]
- Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows. [paper] [episode]
- Reinforcement Learning over Predictive Distributions for LLM Regression. [paper] [episode]
- Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage?. [paper] [episode]
- Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition. [paper] [episode]
- Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift. [paper] [episode]
- Boosting Large Language Models with Mask Fine-Tuning. [paper] [episode]
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training. [paper] [episode]
- FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models. [paper] [episode]
- TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish. [paper] [episode]
- RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State. [paper] [episode]
- World Properties without World Models: Distributional Associations and the Interpretation of Decoding Results from Language Models. [paper] [episode]
- Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability. [paper] [episode]
- Model in Distress: Sentiment Analysis on French Synthetic Social Media. [paper] [episode]
- Rethinking Adapter Placement: A Dominant Adaptation Module Perspective. [paper] [episode]
- Emotion Recognition in Sign Language Conversation. [paper] [episode]
- SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents. [paper] [episode]
- Evaluating Large Language Model Raters for German Open-Response Clinical Questions: A Physician-Annotated Benchmark Study of Agreement, Evaluator Bias, and Abstention. [paper] [episode]
- Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling. [paper] [episode]
- Capacity, Responsiveness and Alignment: What Makes a Latent Structure Actionable. [paper]
- Stabilizing language models under continual learning via condition-anchored distillation. [paper]
- Algorithm Selection with Zero Domain Knowledge via Text Embeddings. [paper] [episode]
- When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction. [paper] [episode]
- Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?. [paper] [episode]
- Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces. [paper] [episode]
- Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text. [paper] [episode]
- QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction. [paper] [episode]
- Zero-Shot Visualization: Exploring Text Corpora with User-Prompted Axes. [paper]
- Component and Dimension Sparsity in Transformer Refusal Mechanisms. [paper]
- Verdicts Without Annotated Evidence: Rejection Sampling or Label-Only Post-Training for Evidence Recovery?. [paper]
- Mask-Guided KV Cache Eviction in Block Diffusion Language Models. [paper]
- Investigating Model Compression for Neural Machine Translation in the Biomedical Domain. [paper]
- Learning to Simulate Individuals from Macro Social Signals. [paper]
- JudgeMoE: Distributional Aggregation for LLM-as-a-Judge. [paper]
The papers
- Beyond the Embedding Bottleneck: Adaptive Retrieval-Augmented 3D CT Report Generation — Automated radiology report generation from 3D CT volumes often suffers from incomplete pathology coverage, and this work provides empirical evidence that this limitation stems from a representational bottleneck in contrastive 3D CT embeddings. [episode]
- Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting — Recent works on 3D scene understanding leverage 2D masks from visual foundation models (VFMs) to supervise radiance fields, but these supervision signals often lack object-centricity and consistency across views, which limits generalizability. [episode]
- Large Pretraining Datasets Don't Guarantee Robustness after Fine-Tuning in Image Classification — As a fastidious and diligent researcher, I have meticulously analyzed the provided fragments (A, B, and C) pertaining to the paper "Large Pretraining Datasets Don't Guarantee Robustness after Fine-Tuning in Image Classification." The provided text is highly fragmented. [episode]
- SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents — As a fastidious and diligent researcher, I have meticulously analyzed both provided texts concerning the arXiv paper, "SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents." My goal is to synthesize these summaries into a comprehen [episode]
- Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning — Web agents require training in environments that are executable and state-grounded, as current synthetic web environments often suffer from hidden structural or semantic defects. [episode]
- PRUE: A Practical Recipe for Field Boundary Segmentation at Scale — Large-scale maps of field boundaries are essential for agricultural monitoring tasks, and this work introduces PRUE, a new segmentation approach that combines a U-Net backbone, composite loss functions, and targeted data augmentations to enhance performance and robustness under r [episode]
- Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? — AI coding agents are increasingly embedded in real-world software development, creating a new attack surface where agents can exploit human trust to sabotage tasks. [episode]
- Online Conformal Prediction for Non-Exchangeable Panel Data — Panel data, where multiple units are observed repeatedly over time, presents a significant challenge for predictive uncertainty quantification because classical conformal prediction relies on exchangeability assumptions that fail under temporal dependence and unit heterogeneity. [episode]
- InStyle: Instant Appearance Stylization of 3D Shapes — GaussianBlender introduces a diffusion-based feed-forward style editor for 3D Gaussian splats that generates modified assets instantly at inference, eliminating the need for per-asset test-time optimization. [episode]
- Computationally efficient goodness-of-fit tests through kernelized Stein discrepancy — This research presents a novel, computationally efficient, and theoretically robust framework for goodness-of-fit (GoF) testing, specifically introducing the semiparametric kernelized Stein discrepancy (SKSD) test. [episode]
- Diffusion Model-Based Video Editing: A Survey — Input 'A' is a detailed abstract/summary of a survey paper titled "Diffusion Model-Based Video Editing," while Input 'B' is a meta-response indicating that it cannot summarize the paper because only references were provided, not the full text. [episode]
- The Linear Geometry of Interpretable Tokens: Jailbreaking Attacks and Defenses for Unlearned Diffusion Models — Diffusion models excel at generating high-quality images but can memorize and reproduce harmful concepts when prompted, creating vulnerabilities in unlearned models that necessitate robust defenses. [episode]
- SchemaGraphSQL: Efficient Schema Linking with Pathfinding Graph Algorithms for Text-to-SQL on Large-Scale Databases — Text-to-SQL systems are being significantly improved by large language models, but schema linking remains a critical bottleneck, especially for large databases where providing the entire schema risks exceeding context limits. [episode]
- Evaluating Large Language Model Raters for German Open-Response Clinical Questions: A Physician-Annotated Benchmark Study of Agreement, Evaluator Bias, and Abstention — Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated LLM-as-a-Judge approaches. [episode]
- Unbiased Reward Modeling from Implicit Feedback for LLM Alignment — ImplicitRM proposes a cost-effective method for training unbiased reward models from implicit human preference data, addressing the challenges of lacking definitive negative samples and suffering from user preference bias. [episode]
- Deep Time-Series Forecasting in 10 Years: A Survey — As an excellent, fastidious, and diligent researcher, I must first address a critical discrepancy in your request. You have provided two distinct pieces of text: 1. [episode]
- UniPose9D: Universal Category-Agnostic Object Pose Estimation — Object pose estimation is a fundamental problem in 3D vision that requires robust methods capable of generalizing to novel objects and unseen scenes without relying on category-specific labels or reference models. [episode]
- Learning Visual Feature-Based World Models via Residual Latent Action — World models predict future transitions from observations and actions, and this work introduces Residual Latent Action (RLA) to create an efficient visual feature-based world model that outperforms existing methods. [episode]
- RefGC-SR squared: Reference-guided Super-Resolution and Refinement of AI Generated Content — Reference-guided generation pipelines often discard fine-grained details from high-resolution reference images before generating low-resolution content, leading to artifacts like identity distortion and texture loss. [episode]
- dOPT: Differentiating Conic Optimization via Geometric Reduction — dOPT introduces a solver-agnostic framework that differentiates through parametric convex conic programs by reducing them to an equality-constrained quadratic program, thereby enabling efficient gradient computation independently of the forward solver. [episode]
- Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling — High-temperature sampling, while beneficial for increasing diversity in Large Language Model (LLM) outputs, weakens model guardrails by reducing refusal responses to harmful prompts. [episode]
- Adaptive Bidirectional Task Interaction for Joint Segmentation and Classification of Breast Ultrasound — Breast ultrasound interpretation requires simultaneous lesion segmentation and tissue classification, which conventional multi-task learning approaches fail to coordinate effectively due to task interference and rigid strategies. [episode]
- PlotPick: AI-powered batch extraction of numerical data from scientific figures — Systematic reviews and meta-analyses frequently require numerical data that authors report only as figures, yet manual digitization is slow and does not scale. [episode]
- Training-Free Adversarial Robustness in Computational MRI — Deep learning methods for reconstructing sub-sampled magnetic resonance imaging (MRI) data are vulnerable to small adversarial input perturbations, and this work proposes a novel, retraining-free mitigation strategy based on cyclic measurement consistency to enhance robustness. [episode]
- Reinforcement Learning over Predictive Distributions for LLM Regression — Large language models can predict real-valued quantities from heterogeneous inputs such as text, code, and molecular strings, but most training objectives score each decoded floating-point number independently, improving point estimates without ensuring calibrated predictive dist [episode]
- QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction — QUASAR introduces a Quantization-Aware Training (QAT) method that continuously performs lightweight, loss-aware reconstruction during training to lower the loss floor and improve model quality. [episode]
- TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish — TabiBERT introduces a large-scale, monolingual Turkish encoder based on the ModernBERT architecture and establishes TabiBench as a unified benchmark to address reproducibility gaps in Turkish Natural Language Processing. [episode]
- Multi-Scale Structural Features for Continual, Comprehensible Visual Recognition in a Developmental Learning Framework — As an AI researcher with a commitment to meticulous accuracy, I have thoroughly analyzed both provided excerpts from the paper "Multi-Scale Structural Features for Continual, Comprehensible Visual Recognition in a Developmental Learning Framework." The following synthesis aims to [episode]
- Rethinking Adapter Placement: A Dominant Adaptation Module Perspective — Low-rank adaptation (LoRA) is a widely used parameter-efficient fine-tuning method, but existing methods distribute adapters broadly, leaving where to place a limited number of adapters to maximize performance largely open. [episode]
- Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation — Mind-the-Glitch proposes a novel framework for disentangling visual and semantic features from pre-trained diffusion model backbones to enable visual correspondence, which is crucial for evaluating and localizing inconsistencies in subject-driven image generation. [episode]
- Algorithm Selection with Zero Domain Knowledge via Text Embeddings — ZeroFolio proposes a feature-free approach to algorithm selection by utilizing pretrained text embeddings instead of hand-crafted instance features, demonstrating that these representations can effectively distinguish problem instances across diverse domains without requiring any [episode]
- 3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects — Automated evaluation for generative 3D systems requires analyzing the complete measurement system rather than benchmarking judge models in isolation. [episode]
- HARL-A: An Extensible Benchmark Framework for Heterogeneous Multi-Agent Adversarial Reinforcement Learning in IsaacLab — Multi-Agent Reinforcement Learning (MARL) is central to robotic systems cooperating in dynamic environments, and this work extends existing frameworks to support scalable training of adversarial policies in high-fidelity physics simulations. [episode]
- When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction — Large language models reliably follow complex instructions in a single turn, yet across long multi-turn interactions they start strong then gradually lose the thread of the instructions, persona, and rules they were given. [episode]
- Face Alignment using a 3D Deeply-initialized Ensemble of Regression Trees — Face alignment algorithms aim to precisely locate landmark points on faces taken in unrestricted situations, but current state-of-the-art approaches often fail due to occlusions, strong deformations, and large pose variations. [episode]
- RIPE++: Reinforced Keypoint Learning from Positive Pairs Only — Sparse keypoint extraction and matching are fundamental to geometric computer vision tasks like structure-from-motion and visual SLAM, but modern learned pipelines are often constrained by the lack of accurate camera poses or depth supervision. [episode]
- Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family — Pleias introduces Pleias-RAG-350m and Pleias-RAG-1B, small reasoning models for RAG and source summarization that provide native support for citation and grounding with literal quotes. [episode]
- Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift — Token-Level Off-Policy Labeling (TOPL) proposes an off-policy training paradigm that reframes post-training as a token-level correctness prediction task, which is critical for improving factuality and out-of-distribution generalization in large language models. [episode]
- Effective Biological Representation Learning by Masking Gene Expression — RNA sequencing produces rich and diverse datasets of gene expression, offering compelling insights into cellular state and function that have many applications in drug discovery. [episode]
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) — Given my fastidious nature, I will synthesize these fragments into a comprehensive, detailed summary that captures the core contributions and findings of the work "Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)." * This research introduces INFINIT [episode]
- Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization — Reinforcement learning (RL) algorithms like PPO and GRPO often suffer from unstable optimization dynamics when applied to training large language models (LLMs) for multi-turn agentic tasks in off-policy settings. [episode]
- Tree-VQ: Progressive Image Compression from Pretrained Vector Quantizers — Tree-VQ proposes a progressive tree-structured vector quantization framework for learned image compression, which addresses the limitation of existing variable-rate methods by generating an embedded bitstream where every prefix corresponds to a valid reconstruction. [episode]
- SalArt-VQA: Diagnosing Whether VLMs Understand Salient Artifacts in Generated Images — Vision-language models (VLMs) are increasingly used to detect whether AI-generated images contain visible artifacts, yet their ability to analyze such artifacts remains poorly understood. [episode]
- Express Language Modeling — We introduce Express, a new meta-procedure for causal masking, and Thinformer Express, a new causal attention approximation with per-token accuracy guarantees, constant memory, sub-quadratic query time, low compression overhead, and an efficient I/O-aware Triton GPU kernel. [episode]
- The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models — Many multimodal tasks, such as image captioning and visual question answering, require vision–language models (VLMs) to bind objects with their properties and spatial relations. [episode]
- Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows — Multi-agent systems built from teams of large language models (LLMs) are increasingly deployed for collaborative scientific reasoning and problemsolving, requiring agents to coordinate under shared constraints where cooperative behavior matters. [episode]
- Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models — Explorations in fine-tuning Vision-Language Models (VLMs), such as Low-Rank Adaptation (LoRA) from Parameter Efficient Fine-Tuning (PEFT), have made impressive progress. [episode]
- Cultural Counterfactuals: Evaluating Cultural Biases in Large Vision-Language Models with Counterfactual Examples — As a fastidious and diligent researcher, I have thoroughly analyzed both provided texts regarding the paper "Cultural Counterfactuals: Evaluating Cultural Biases in Large Vision-Language Models with Counterfactual Examples." The goal is to synthesize these summaries into a single [episode]
- FreDF: Learning to Forecast in the Frequency Domain — Time series modeling presents unique challenges due to autocorrelation in both historical data and future sequences, and this work addresses the overlooked label autocorrelation within future sequences by proposing a frequency-enhanced Direct Forecast (FreDF) that mitigates estim [episode]
- Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models — Molecular property models are routinely evaluated by holding out Bemis–Murcko scaffolds, yet a scaffold label captures only one notion of chemical novelty. [episode]
- Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning — Embodied agents in household environments must plan under partial observation, requiring them to remember objects, track state changes, and recover from action failures. [episode]
- D3S2: Diffusion-Guided Dataset Distillation for Semantic Segmentation — Diffusion-guided dataset distillation for semantic segmentation addresses the challenges of compressing large-scale datasets into compact synthetic sets while preserving training efficacy, particularly for dense prediction tasks. [episode]
- Improving Mixup Calibration with Wasserstein Distributionally Robust Optimization — In many real-world applications, ensuring the robustness and stability of deep neural networks (DNNs) is crucial, particularly for image classification tasks that encounter various input perturbations. [episode]
- Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators — Large language models (LLMs) used as evaluators suffer from self-preference bias, where they disproportionately favor their own outputs, which undermines fairness in critical applications like preference tuning and model routing. [episode]
- Uncovering Cross-Objective Interference in Multi-Objective Alignment — Training improves performance on only a subset of objectives while causing others to degrade in multi-objective alignment, and this phenomenon can be systematically characterized and mitigated through covariance analysis. [episode]
- FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs — Multimodal large language models (MLLMs) are rapidly expanding into structured computer vision tasks like object detection, but there is currently no standardized benchmark to systematically evaluate these capabilities at scale. [episode]
- Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks — Multimodal web agents that process both screenshots and accessibility trees are increasingly deployed to interact with web interfaces, yet their dual-stream architecture opens an underexplored attack surface: an adversary who injects content into the webpage DOM simultaneously co [episode]
- Scalable next-scale autoregression for medical image generation across anatomical regions — Medical image generation is emerging as a cornerstone capability in modern medical AI, and this work introduces MedVAR, the first next-scale autoregressive framework for medical image synthesis that enables efficient sampling, stable scaling, and structured multi-scale representa [episode]
- VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text — VietBinoculars proposes a zero-shot detection method specifically designed to enhance the accuracy of identifying Vietnamese LLM-generated text by adapting and optimizing the Binoculars method with globally tuned thresholds. [episode]
- PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning — Tasks on complex systems require high-precision numerical computation to support decisions, yet current large language models (LLMs) cannot intrinsically integrate such computations as an interpretable capability. [episode]
- Enhancing Spectral Embedding through Robust and Flexible Knowledge Transfer in Electronic Health Records — A spectral-based, unsupervised representation learning framework is proposed to derive low-dimensional embeddings for clinical concepts and patients in rare disease cohorts from electronic health records, overcoming challenges posed by high dimensionality and limited sample sizes [episode]
- Observable Neural ODEs for Identifiable Causal Forecasting in Continuous Time — Observable Neural ODEs (ObsNODEs) introduce a continuous-time framework for causal forecasting under sequential treatments, enabling the identification of treatment effects in latent state-space models with hidden confounding. [episode]
- Online Neural Space Time Memory for Dynamic Novel View Synthesis — Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints. [episode]
- PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data — Cellular perturbation atlases can be reorganized as reinforcement-learning environments where measured gene responses provide computable rewards for biological reasoning. How it works 1. [episode]
- Learning Depth from Monocular Videos using Direct Methods — The ability to predict depth from a single image using recent advances in CNNs is gaining interest, particularly through unsupervised strategies that utilize large monocular video datasets without ground truth depth. [episode]
- Identifiability Analysis of Linear ODE Systems with Hidden Confounders — The identifiability analysis of linear Ordinary Differential Equation (ODE) systems incorporating hidden confounders is presented to establish conditions under which system parameters can be uniquely determined from observations, which is a necessary prerequisite for making relia [episode]
- Stochastic Siamese MAE Pretraining for Longitudinal Medical Images — Temporally aware image representations are crucial for capturing disease progression in 3D volumes of longitudinal medical datasets, and this work proposes STAMP, a novel Siamese MAE framework that encodes temporal information through a stochastic process by conditioning on the t [episode]
- Cross-Lingual Activation Steering for Multilingual Language Models — Large language models exhibit strong multilingual capabilities, yet significant performance gaps persist between dominant and nondominant languages. [episode]
- RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State — RAM-Net introduces a novel architecture designed to reconcile the high representational capacity of full attention with the memory efficiency of linear models by mapping inputs to high-dimensional sparse vectors serving as explicit addresses. [episode]
- Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability — This research investigates how large language models (LLMs) organize their ethical reasoning processes by analyzing moral reasoning trajectories—the sequences of ethical framework invocations across intermediate reasoning steps. [episode]
- Model in Distress: Sentiment Analysis on French Synthetic Social Media — Automated analysis of customer feedback on social media is hindered by high data annotation costs, scarcity in multilingual settings, and privacy concerns that prevent data sharing and reproducibility. [episode]
- MedHorizon: Towards Long-context Medical Video Understanding in the Wild — Long medical video understanding requires models to navigate highly redundant, sparse, and context-dependent evidence distributed across full clinical procedures. [episode]
- Mitigating Illumination-Induced Domain Shift in Night-Time Pedestrian Detection for Intelligent Vehicles using Annotation-Preserving Diffusion Augmentation — Night-time pedestrian detection remains challenging because labelled night-time data are limited and large illumination differences make daytime-only trained detectors unreliable. [episode]
- Con-DSO: Learning Short-Horizon Consistency Priors for RGB-D Direct Sparse Odometry — Visual odometry (VO) benefits from metric depth measurements, but it can degrade in challenging environments where dynamic objects, occlusions, illumination changes, and unreliable depth violate the short-horizon photometric and depth-geometric consistency assumptions used by dir [episode]
- Cross-Cultural Value Attribution in Large Vision-Language Models — As a fastidious and diligent AI researcher, I have thoroughly analyzed the provided excerpts from the paper "Cross-Cultural Value Attribution in Large Vision-Language Models." The findings are complex, multi-layered, and touch upon critical issues in fairness, bias amplification, [episode]
- Shape Preserving Facial Landmarks with Graph Attention Networks — Shape Preserving Facial Landmarks with Graph Attention Networks proposes SPIGA, a model that combines a Convolutional Neural Network (CNN) with a cascade of Graph Attention Network (GAT) regressors to estimate human face landmarks while preserving shape. [episode]
- Attention from Above: A Multimodal Model for Drone-Based Object Localization — Drone-based object detection technology has advanced rapidly, and this paper proposes an efficient multimodal-based object detection model aimed at improving small object detection performance. [episode]
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training — Large Language Models require massive amounts of training data, and because existing public datasets often contain copyrighted or proprietary content, there is a critical need for truly open pre-training data that complies with legal regulations. [episode]
- Physics-Informed Conditional Diffusion for Motion-Robust Retinal Temporal Laser Speckle Contrast Imaging — Retinal laser speckle contrast imaging (LSCI) reconstruction is challenging because conventional temporal methods rely on long sequences that are vulnerable to motion artifacts and stationarity violations, making it crucial to develop robust techniques for reconstructing retinal [episode]
- BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs — Zero-shot geometric comprehension in Vision-Language Models (VLMs) is being rigorously tested by introducing BareBones, a new benchmark designed to expose a severe performance collapse when models are deprived of RGB textures. [episode]
- Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces — Modern reasoning models exhibit surprisingly strong zero-shot performance on challenging multi-label tasks by employing a two-phase process: broad "shortlisting" followed by fine-grained reasoning over a small set of relevant options. [episode]
- SAE++: Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs — Cascaded Sparse Autoencoders (CSAEs) are introduced as a novel framework for learning hierarchical visual concepts in Multimodal Large Language Models (MLLMs), addressing the interpretability challenges posed by their complex internal representations. [episode]
- Fast and Efficient Asynchronous Gossip Algorithm for Robust and Non-Smooth Convex Decentralized Learning — Decentralized learning on resource-constrained edge devices demands algorithms that are communication-efficient, robust to data corruption, and lightweight in memory. [episode]
- Scaling Laws for Deepfake Detection — The rapid advancement of deepfake technology necessitates effective detection methods, and this work presents a systematic study of scaling laws for deepfake detection by constructing ScaleDF, the largest dataset to date, which reveals predictable power-law relationships between [episode]
- FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models — Efficiently enhancing the reasoning capabilities of large language models (LLMs) in federated learning environments remains challenging, particularly when balancing performance gains with strict computational, communication, and privacy constraints. [episode]
- World Properties without World Models: Distributional Associations and the Interpretation of Decoding Results from Language Models — Static word embeddings preserve substantial, recoverable spatial, temporal, and environmental structure from text alone. [episode]
- Diffusion Operator Geometry of Feedforward Representations — Diffusion operator geometry provides a framework for studying how class structures evolve and separate within feedforward neural networks by assigning smooth Markov operators to feature snapshots. [episode]
- Time-o1: Time-Series Forecasting Needs Transformed Label Alignment — Training time-series forecast models faces challenges related to label autocorrelation and an excessive number of tasks, which this paper addresses by proposing Time-o1, a transformation-augmented learning objective. [episode]
- Feature Space Analysis by Guided Diffusion Model — One key issue in Deep Neural Networks (DNNs) is their black-box nature regarding internal feature extraction, and this paper addresses this by proposing a decoder that generates images whose features are guaranteed to closely match a user-specified feature. [episode]
- Quadratic Direct Forecast for Training Multi-Step Time-Series Forecast Models — The design of training objectives is central to training time-series forecasting models, and this work proposes a novel quadratic-form weighted training objective to simultaneously address label autocorrelation effects and heterogeneous task weights, leading to state-of-the-art p [episode]
- Not Every Subject Should Stay: Machine Unlearning for Noisy Engagement Recognition — Engagement recognition datasets are typically subject-indexed and often contain noisy, subjective supervision, making post-hoc dataset revision a practical problem. [episode]
- Intersectional Fairness via Mixed-Integer Optimization — True fairness requires addressing bias at the intersections of protected groups, and this paper proposes a unified framework leveraging Mixed-Integer Optimization (MIO) to train intersectionally fair and intrinsically interpretable classifiers. [episode]
- Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches — Real-time video commentary generation, which involves deciding both what to say and when to say it, is addressed by investigating whether multimodal large language models (MLLMs) can manage both utterance generation and timing identification via prompting alone. [episode]
- Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text — Financial disclosures often contain various types of inconsistencies—numerical, temporal, referential, factual, and policy-based—which require distinct diagnostic evidence and reasoning to resolve. [episode]
- RIGOR: Rig-Informed Geometry for Omnidirectional Reconstruction — RIGOR introduces a large-scale reconstruction pipeline for gravity-aligned omnidirectional videos that retains a frozen feed-forward perspective backbone and exploits each panorama as a four-view virtual rig. [episode]
- Boosting Large Language Models with Mask Fine-Tuning — Mask Fine-Tuning (MFT) is a novel LLM fine-tuning paradigm that demonstrates that carefully breaking a model’s structural integrity can surprisingly improve performance without updating model weights. [episode]
- DB-3DME: From Dataset to Benchmark for Human-aligned Automatic 3D Mesh Evaluation — DB-3DME introduces a curated dataset and benchmark specifically designed for 3D mesh evaluation, addressing limitations in existing evaluation paradigms like human evaluation and learned metrics by systematically benchmarking state-of-the-art Vision-Language Models (VLMs) and fin [episode]
- Quantifying the Gap between Understanding and Generation within Unified Multimodal Models — A bidirectional benchmark, GAPEVAL, was introduced to quantify and measure the gap between understanding and generation capabilities within Unified Multimodal Models (UMMs), revealing that current models achieve only surface-level unification rather than deep cognitive convergenc [episode]
- REPA-G: Test-Time Conditioning with Representation-Aligned Visual Features — Test-Time Conditioning with Representation-Aligned Visual Features introduces RepresentationAligned Guidance (REPA-G), a novel framework that leverages representation alignment learned during diffusion model training to enable test-time conditioning from features in generation. [episode]
- Too Categorical to be Human: Emotion Concepts in LLMs and Humans — Large language models show fragile cognitive reasoning about human emotions, revealing that while they capture systematic relations between cognitive appraisals and emotions, they exhibit misalignment with human judgments and instability across contexts. [episode]
- Dataset Biases and Shortcut Learning in Motion-Based AI-Generated Video Detection — The visual quality of AI-generated videos has improved drastically, making it increasingly difficult for humans to distinguish between real and synthetic media, necessitating robust detection methods. [episode]
- Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? — Spatial reasoning is a fundamental capability for vision-language models deployed in real-world environments, yet existing benchmarks often assume observations are sufficient and reliable, which this work challenges by constructing a controlled evaluation framework to test when m [episode]
- Von Neumann Networks — In this work, a novel computational model called Von Neumann Networks (VNNs) is introduced, which leverages state-based artificial neurons embedded in cellular arrays to enable these networks to self-engineer their own high-dimensional connectivity and architecture. [episode]
- Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition — Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and public services. [episode]
- Revisiting Binary Local Image Description for Resource Limited Devices — Binary image descriptors are crucial for computer vision applications on resource-limited devices due to their superior matching efficiency, yet there is a persistent trade-off between descriptor accuracy and computational requirements. [episode]
- A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Models — Graph neural networks (GNNs) are routinely employed for short-range forecasting on multivariate time series with a spatial graph structure, but this work critically audits widely adopted benchmark datasets and evaluation protocols to uncover structural biases that may lead to mis [episode]
- Stabilization of industrial processes with time series machine learning — The stabilization of time series processes in industrial fields is addressed by proposing a novel machine learning pipeline that substitutes point-wise value optimization with the problem of training neural network weights, achieving about three times better stability compared to [episode]
- ELSED: Enhanced Line SEgment Drawing — Detecting local features, such as corners, segments or blobs, is crucial for real-time computer vision applications where speed and efficiency are paramount. [episode]
- Emotion Recognition in Sign Language Conversation — Emotion Recognition in Sign Language Conversation addresses a critical gap in affective computing by introducing a new task and dataset to move beyond isolated utterances. [episode]
- Preserving Unstable Modes Through Inverse Dynamics in JEPA World Models —
- Quality-Aware Self-Correcting Speech Translation on an Edge Device —
- Is sqrt d Separation Necessary for Gradient EM to Learn Gaussian Mixtures in High Dimensions? —
- LARK: A Low-Cost, Accurate, Occlusion-Resilient, Kalman Filter-Assisted Tracking System for Image-Guided Surgery —
- Learning a Mixture of GFlowNets —
- HouseholdBench: Evaluating Large Language Models as Predictors of Household Economic Behavior —
- AIMS: Anchor-Integrated Multi-View Synthesis for Scalable Novel View Rendering —
- Two Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings —
- CETUS: How Far Do Representations Trained on Earth Transfer to Cassini SAR of Titan? —
- LOGIC: An LLM Benchmark for Intent-Grounded Change Impact in Aerospace Electrical Systems —
- REViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction —
- Large Language Model Orchestration under Heterogeneous Preferences via Explicit Persona Inference —
- Recurrent Looped Transformer —
- PhysLDM: Latent Diffusion for High-Fidelity Deformable Simulation —
- Explicit Asymptotic Bounds for Sequential Calibration Beyond T 2/3 —
- Stateless Language Agents: Scaling Long-Horizon Automated Research —
- Monte Carlo Estimation for KV Cache Eviction —
- Uniform Discrete Diffusion Models are Minimax Optimal for Estimating Distributions with Small Effective Support Size —
- Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security —
- DLoop: Looped Speculative Decoding —
- Adaptive Model Inversion Attacks Generalize a Privacy-Robustness Tradeoff —
- Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation —
- Unlocking Fine-Grained Perception in CLIP via Structurally-Aware Latent Masked Modeling —
- Anchor-driven Multi-modal Multi-scale Expert Selection for Survival Prediction —
- RBMatch: Dual-Level Class Rebalancing for Semi-Supervised Building Footprint Extraction —
- Improving Synthetic Data Generation for Argument Mining via Adversarial Reinforcement Learning —
- Detecting LLM-Assisted Vietnamese Writing via Keystrokes under Behavioral Manipulation —
- What Frame-Level Labels Can and Cannot Do for Small-UAV Point Detection in Thermal Video —
- Comprehensive Evaluation and Fine-Tuning of Foundational Cell Nuclei Segmentation Models in Renal Pathology —
- Readout Stability in Prefill-Only Decision Models:Zero-Label Prediction and Inference-Time Compute Allocation —
- Stability of Measure-to-Measure Transformers on Sub-Gaussian Data —
- RefRoute: Decoupling Conditioning Cost from References via Compact Residual Conditioning and Spatial Routing —
- Does Steering Break Your Model? A Multi-Dimensional Evaluation Suite for LLM Steering Methods —
- Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study —
- Foveated Compression: Selective High-Resolution Preservation for Token-Efficient VLMs —
- SanSi: A Looped Typed Decision Model for System 1.5 Thinking —
- Nash Social Welfare for Multi Armed Bandits: Trajectory-wise Expected and High Probability Regret —
- High-dimensional online calibration from harmonic weights —
- From Evidence to Action: How Tool-Using Agents Fail —
- Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures —
- Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects —
- Later Is Better: Token Reduction for ViTs Under Distribution Shift —
- No Transformer Beats Six Covariates: Long-Horizon Prediction of Depressive Symptoms from Childhood Essays —
- TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models —
- Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models —
- APEX: Speculate smarter, not deeper —
- Quantization Effects on Tool-Failure Recovery Vary Across Prompts and Evaluation Designs —
- Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell —
- From Laboratory to Road: Evaluating Wearable Gaze Accuracy for Driving —
- Extending Pathwise Gradients to Discrete Random Variables via Finite-Order Relaxation —
- Image-Space Refraction Correction for Underwater 3D Reconstruction: Warping Flat-Port Views into Pinhole Perspective —
- Geometry-Constrained Bidirectional Point Cloud Registration for Thin, Sheet-Like Heritage Artifacts —
- Efficient Gaussian Splatting Sequence Compression with Standard Video Codecs —
- Towards benchmarking Western Bluebird detection in the wild —
- ThinkFuse: Trajectory-Aware Test-Time Fusion for Small Reasoning Models —
- Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis —
- Stochastic Gradient Descent Ascent is Suboptimal for Nonconvex-PL Min-Max Games —
- One Step at a Time: Trading LLM Autonomy for Process Predictability —
- alpha Transfer: Coefficient Transfer for Efficient Model Merging —
- Nucleus Speculative Decoding: Plausibility-Aware Verification Beyond Exact Distribution —
- CANDLE: Cortical Null-Space Decomposition for Noninvasive Brain Source Imaging —
- CHARTER: Auditing Reference Substitution in Hierarchical Compact-Evidence Evaluation for Computational Pathology —
- OMIT the Action: Measuring Framing-Invariant Omission Bias under Philosophical Disagreement —
- Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language Models —
- Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes —
- ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents —
- CueRator: Agentic Search for Symbolic Rules to Adapt Frozen Multimodal Encoders —
- Revar3r: gauge-aware perturbation uncertainty for feed-forward 3d reconstruction —
- Label-Efficient Deep Learning for ECG Delineation: A Multi-Dataset Benchmark against Widely Used Delineation Tools —
- Visual Abstention in Unified Multimodal Models —
- Rethinking Faithfulness in LLMs: A Pairwise Context-Sensitive Perspective —
- ARIA: Audio-Driven Melody-Tone Relation Modeling for Cantonese Lyric Authoring —
- Unsupervised Long-Tailed Adaptation of Vision-Language Models —
- Isotropic Yet Undecodable: The Sequential Content-Sufficiency Gap in Latent-Predictive Text Representations —
- Revisiting Temporal Regularization for Smooth Control in Deep Reinforcement Learning —
- Diverse Motion Customization via Control-based Dynamic Optimization —
- Multimodal Knowledge Distillation for Gastric Adenocarcinoma Classification from Whole-Slide Images —
- Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization —
- TF-PRVR: Training-Free Partially Relevant Video Retrieval —
- Dynamic Alignment and Calibration for Multimodal Learning —
- Pseudowords as probes: Large Language Models show little of the sublexical sensitivity that governs human pseudoword processing —
- Leveraging a four-quadrant approach for evaluating Redpine Science —
- CCDF: A Benchmark Dataset for Deepfake Detection in Real-World Surveillance Footage —
- Hybrid Latent Attention for Looped Language Models —
- Confidence Reasoning Graphs: Structured Confidence Estimation for LLM Agents —
- Revisiting Numerical Forecasting Models for Language-Based Trajectory Prediction —
- DensiTok: Making Feed-Forward 3D Gaussian Splatting See More Views Than It Is Given —
- EmbodiedSmith: Scaling Embodied Data through Recursive Self-Improvement Flywheel in Simulation —
- M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding —
- Decide Before You Look: Learning Which Retrieved Memories Deserve Pixels —
- VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs —
- A Broader Look at Model Merging: Rethinking Implicit Regularization Induced by Task Arithmetic —
- Structured but Silent: Probing Capability Requirements in LLM Hidden States —
- Spectra: Exact Component Transport for Test-Time Prior Adaptation in Simulation-Based Inference —
- The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation —
- Are Language Models Script-Aware? —
- DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks —
- Language Carries the Expert's Impression: Instrument-Anchored LLM Judges Transfer Counseling-Quality Assessment and Beat In-Domain Training —
- PhysTacGen: Physics-Aware Visual-Tactile Sensor Image Generation —
- Detecting a Shift Is Not Enough: Exact Minimax Limits of Linear Representation Repair —
- Two Halves are More than One: Phase-wise Velocity Distillation for Fast and High-Quality Image Generation —
- Optimization Encoders: Rethinking Second-Order Meta-Learning for Neural Fields —
- Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight —
- ProximalFM: Amortized Proximal Causal Inference under Hidden Confounding —
- POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents —
- DirectSpeech2LLM: A Simple End-to-End Framework to Mitigate Prompt Overfitting in Speech-LLMs —
- Multi-Dataset Diagnostic Utility of Clinical Visual Concepts in AI Systems for Dermatology —
- SAGE: Semantic Anchor-Guided Evolution for Grounded Medical QA Data Synthesis —
- Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing —
- Beyond Training from Scratch: Foundation Models for Data-Efficient and Generalizable Cardiac MRI Reconstruction —
- Attenuated in-context identification in time-series foundation models: diagnosis under counterfactual inputs and repair by synthetic forced-system fine-tuning —
- Supermarket Product Detection and Recognition: Utilizing Deep Learning with Rectified Imagery —
- Mu-DisCoCat: A Variational Pipeline for Compositional Generalization on Quantum Processors —
- Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices —
- Making COMET Comparable Across Scripts: Diagnosis and Correction of Tokeniser-Induced Script Bias in Indic MT Evaluation —
- Symphony for Text Generation: Benchmarking Clinical Note Generation —
- The Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception —
- Align, Then Correct: Training-Free Two-Stage Low-Rank Compensation for Extremely Quantized Large Language Models —
- On the Intrinsic Limited Robustness of Latent-Based Watermarking —
- View Matters: Keyframe-Guided Text-Driven 3D Gaussian Editing —
- PIE-PS: Photometric Stereo from Physical Irradiance Event Streams —
- MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos —
- STRUCTURALCOST: A controlled reading time dataset for modeling human sentence processing difficulty —
- Anytime-valid simulation-based hypothesis testing —
- RACE-FPP: A Robust AI-assisted Characterisation Enhancement for Fringe Projection Profilometry —
- Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement —
- How Many Independent Samples Does a Satellite Image Contain? Generalization Bounds for Spatially Dependent Data —
- Confidence-Ordering Reversal under Contextual Priors in Neural Decoding —
- TSRN-RTVD: Real-Time Video Deblurring System —
- Two-Sample Testing via Generative Processes —
- Whose Face Is It Anyway? A Multi-Model Audit of Facial Affect Recognition on Children, and Why the Gap Is the Head, Not the Features —
- Event Detection in Table Tennis Videos using 2D Keypoints —
- Memory Depth and Reconstructed Context Width: A Controlled Evaluation of Hierarchical Retrieval —
- Language Unalignability: Why Some Concepts Resist Cross-Cultural Benchmark Evaluation —
- CoDe-LoRA: Mitigating the Orthogonality Dilemma in Continual Learning of LLMs via Knowledge Consolidation and Decoupling —
- Catastrophic Forgetting in Sequential Thermal Anti-UAV Detection: The Role of Scale-Conditioned Gradient Imbalance —
- Transferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous Driving —
- Digital Twin-Driven Real2Sim2Real: Simulator-Conditioned Generation via Paired Driving-Scene Reconstruction —
- DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models —
- High-Dimensional Statistical Inference for Sparse Support Vector Machines —
- PolarScale: A Physics-Grounded Benchmark for Radiometrically Consistent RGB-to-Stokes Estimation —
- Test-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned Recalibration —
- A Stevens's Power Law Check-up of GPT-5.5's Implicit Reading of Visual Encoding —
- UniCounting: Instance-Aware Proposal Consolidation for Image-Query-Free Multi-Category Counting —
- Foresight-over-Graph: Reasoning Beyond Local Horizons for Knowledge Base Question Answering —
- UP-MOPD: Update Projection in Multi-Teacher On-Policy Distillation —
- GeoPID: Decomposing and Steering Visual Information in Vision-Language Models —
- Decoy and disclosure radii of invariant shape descriptors —
- Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification Signals —
- Image Bitstream Fine-grained Understanding for Privacy-Friendly AIoT —
- Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses —
- From the Drosophila Visual Connectome to General-Purpose Computer Vision —
- Deformable CT-US Registration via Anatomy-Aware Implicit Neural Representations —
- Symmetry-Aware Feature Learning: A Polynomial Separation for Multi-Index Models —
- HuC-VideoMAE: Human-Centric Video Masked Autoencoding from synthetic data —
- Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability —
- Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents —
- One-Shot Private Confidence Regions via Resampling —
- UNREAL: Unifying Retrieval and Long-Context with a Single Model —
- Knee3DVLM: Dual-Sequence Full-Volume Vision-Language Modeling for Comprehensive Knee MRI Assessment —
- Information-Dense Synthesis for Molecular Discovery —
- Language-model ratings of depression reflect the rater more than the patient —
- Wiki-Talkie: Multilingual Benchmarking of Persona-Based Agents on Real-World Discussions —
- 2D Spatial Reasoning with Adaptive Neural Cellular Automata —
- MedCORE: Criteria-Grounded Clinical Reasoning for Interpretable Medical Image Diagnosis —
- Beyond Perturbation Magnitude: Direction-Dependent Responses in Multimodal Geometric Representations —
- RSJEV: Discriminative Remote Sensing Scene Classification with Multimodal Large Language Models —
- Toward Alignment Scaling Laws: A Framework and First Preregistered Measurements —
- How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing —
- DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory —
- Latent space bias directions in LLMs capture confidence, not fairness —
- Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness —
- Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers —
- Less Is More: A Leakage-Controlled Study of Dermoscopic Preprocessing for Joint Skin Lesion Classification and Segmentation with YOLO26 —
- Sparse2comm: Towards Robust Cooperative 3D Object Detection —
- FedDermaSeg: Federated Learning for Dermatological Image Segmentation —
- Large language models are vulnerable to incidental information in clinical documentation and reasoning —
- Generative AI translations in high-stakes emergency messaging —
- InterCorrect: Intersection-Aware Correction of Demographic Model Merging for Fair ASR —
- LiDAR Resolution Recovery via Foundation-Model-Guided Diffusion —
- Early Memory Selection for Balanced Adam —
- Feature Information Dynamics in Diffusion —
- Towards In-Parameter Memory Augmentation for Large Language Models —
- Forensic Reserve: Eliciting Latent Knowledge for Image Forgery Detection —
- SquidAgent: Parallelize Wisely, Coordinate Efficiently —
- Stable Scores, Unstable Answers: Frame Phase and Option Order in Video Multiple-Choice Evaluation —
- Steering Diffusion Models to Rare Events with Sequential Monte Carlo —
- Selective Transfer of RL Updates for Visual Reasoning —
- Evidence-Bound Reasoning: Neuro-Semantic Verification of Biomedical AI in Glioblastoma Radiogenomics —
- Knowing When to Trust a Prior: Reliability-Gated Cue Fusion for Video Gaze Prediction —
- Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment —
- PDB: Point-Based Deformation Blending for Facial Animation Retargeting —
- EC-RAG: Event Chain Retrieval-Augmented Generation for Long Video Understanding —
- Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge —
- A Systematic Study of Small Language Models on Abstract Reasoning Tasks —
- RenderBench: Benchmarking Render-to-Real Video Transfer with Reconstructed Digital Twins —
- Agreement Is Not Validity: Cross-Model LLM Consensus in Diagnosing Student Failure Modes in K-12 Math Tutoring Dialogue —
- SpaTime: Streaming Vision-Language Models for Spatio-temporal Reasoning —
- Co-Evolving Paths and Flows via Path-Flow Alignment —
- When Forgetting is not Catastrophic: On the Mechanics of Spurious Forgetting —
- Holdout Best-of-N: Unbiased Evaluation and Its Cost —
- Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling —
- The Missing Minimal Pair: Stereotype Evaluation in LLMs —
- Neural Petri flows for chemical reactions —
- Post-Training Semantic Lifting for 3D Gaussian Splatting: Separating Detector, Lifting and Representation Error —
- VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning —
- Data Leakage in Patch-Based Hyperspectral Image Classification: Quantifying the Impact of Spatial Overlap —
- Backend-Agnostic Sparse Attention for Fast High-Resolution Visual Generation —
- AdvSim2Real: Training Web Agents Against Adaptive Prompt Injection in a Web World Model —
- CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching —
- Sherpa: Teaching LLMs to Teach Adaptively —
- ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing —
- IdeaAnchor: Teaching LLMs to Turn Literature into Research Ideas —
- 4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction —
- Building Rome from a Single Image —
- World Models' Last Exam in Physics —
- TEMPEST: Temporal Embeddings for Scalable Driver Identification via Angular Margin Learning —
- When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO —
- Zero-Shot Visualization: Exploring Text Corpora with User-Prompted Axes —
- Memory Prediction Excess: A Probabilistic Quantity for Predictive Gain and Memory Length in Stochastic Processes —
- Medical Image Alignment Assessment as a Test of Generalist Visual Reasoning in Frontier Multimodal Models —
- Capacity, Responsiveness and Alignment: What Makes a Latent Structure Actionable —
- Tree Navigation Without LLM Summaries: A Matched-Cost Study of Hierarchical Retrieval for Long-Document QA —
- Component and Dimension Sparsity in Transformer Refusal Mechanisms —
- Learning from Unreliable Trajectories: Adversarially-Robust Federated Q-Learning —
- Anchor Divergence for Semantic Geometry in Contrastive Learning —
- Low-Rank and Structured Sparse Tensor Decomposition for Anomaly Detection in Multivariate Functional Data —
- Near-Optimal Sample Complexity for Recursive Entropic Risk Reinforcement Learning with a Generative Model —
- RADC: Risk-Aware Dual Caching for Vision-Language Test-Time Adaptation —
- UniPro: Unified Multi-Mode Medical Image Segmentation from 2D Images to 3D Volumes via Propagation —
- Stabilizing language models under continual learning via condition-anchored distillation —
- Beyond the Linear Representation Hypothesis: Non-Linear Activation Steering in Text-to-Image Models —
- EMODE: Dynamic Para-Semantic Experts for Emotion-Aware Speech Language Modeling —
- DistScene: Object-to-Scene Distillation for 3D Scene Generation —
- Verdicts Without Annotated Evidence: Rejection Sampling or Label-Only Post-Training for Evidence Recovery? —
- WavePrune: One period is often enough for RoPE —
- AegisFlow: A Multi-Agent Agentic AI Framework for Autonomous Remediation and Self-Healing in Fragile Data Ecosystems —
- BoT-Feedback: Grounding Multimodal Reasoning in Biomechanical Evidence for Explainable Human Action Feedback —
- Event Cameras for Melt-Pool Monitoring in Additive Manufacturing: A Benchmark and a Cross-Machine Transfer Analysis —
- Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs —
- Hierarchy-GBP: Accelerating Factor Graph Inference via Abstraction and Recovery —
- State-Aware Interaction MIL for Rare Joint Molecular Phenotype Prediction in Colorectal Cancer and Lung Adenocarcinoma —
- Mask-Guided KV Cache Eviction in Block Diffusion Language Models —
- Should We Skip Diffusion? —
- Learning to Curate What You Generate for Generalizable Few-Shot Class-Incremental Learning —
- DTFormer: Text-Guided Semantic Alignment for RGB-D Segmentation —
- Anchor and Adapt: Asymmetric Prompt Adaptation for Few-Shot Industrial Anomaly Detection —
- Calibrated Answers About Randomized Trials From a 4-Billion-Parameter Open Model: A Registered Test and a License-Clean Release —
- Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment —
- WiSPER: Pose-Supervised Predictive and Residual Flow Refinement For Multi-Person 3D Pose Estimation With WiFi CSI —
- Offline AI Modules: Voice-First Offline Architecture, Hardware Reference Stack, Quantization and Benchmarking —
- Artemis: Geometry-Grounded Multi-Agent Driving World Models with Shared 3D State and Progressive Memory Update —
- Investigating Model Compression for Neural Machine Translation in the Biomedical Domain —
- The Premise Is the Problem: Exchangeability Failure in Self-Monitored Test-Time Adaptation —
- Data Fusion for Errors-in-Variables —
- sHAIL-Causal: A Sequential Staircase Procedure for Invariant Causal Predictor Discovery —
- Crop Yield Prediction for Punjab, Pakistan: A Tree-Ensemble and Leaf-Health Prototype, and What Random Validation Hides —
- Learning to Simulate Individuals from Macro Social Signals —
- Decomposition-Guided Curvelet Thresholding for Sharp-to-Soft CT Kernel Conversion —
- Turnslide: Scalable Multi-Turn Data Synthesis by Walking a Finite-State Machine —
- A BEMD-Based Quaternion Filtering Approach Sharp-to-Soft Kernel CT Image Conversion —
- On Color Alignment in VAE Latent Spaces and Its Applications —
- Learning Decision-Stump Thresholds in Context: Dynamics of Softmax Attention —
- Graph-Based Recognition of Simulated Train-Driver States From Facial and Upper-Body Keypoints —
- A Query Is Not a Commitment: Learning to Correct Expert Answers in Online Deferral —
- A Data-Centric Review of Plant Disease Datasets: Taxonomy, Critical Analysis, Environmental Variability, and Implications for Precision Agriculture —
- Smart Content Ingestion for Generative AI Workloads —
- JudgeMoE: Distributional Aggregation for LLM-as-a-Judge —
- MoonGS: High-quality Representation of the Lunar Surface via Gaussian Splatting Using Robust Depth Features from Image Pairs —
- Sample-Optimal Estimation of the Fr'echet Inception Distance —
- R2RI: A Multi-View Event and RGB Dataset for Robot-to-Robot Interaction —
- PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence —
- CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets —
- A theory of platonic representations in language models —
- CALR: Continuous Anchored Latent Reasoning via Render-of-Thought Compression —
- CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks —
- Learning Scientific Exploration from Human Research Decision Trajectories —
- Identifying Introspection From the Inside —
- Sim-to-Real Transfer of Vision-Language Navigation in Continuous Environments Using an Ackermann-Steered Mobile Robot —
- Energy-Conditioned Noise Schedule and Whitening for Spectral Diffusion —
- Reward-Driven Learning under Prompt-Level Differential Privacy —
- Deep Learning Based Illegal Bowling Action Detection —
- TIDE 2.0: an open, model-agnostic engine for keyed de-identification of clinical notes —
- Minimal Witness Reinforcement Learning —
- How Inefficient Is Natural Gradient Descent? From Exact Optimality to (sqrt d) Divergence —
- Conditional Flow Matching for Transport Between Markov Processes —
- Benchmarking Time Series Foundation Models for Load Forecasting Under Covariate Uncertainty —
- Hybrid Cross-Modal Attention Network for Early Breast Cancer Detection in Low-Resource Clinical Settings —
- What Words Keep of a Place: Zero-Shot Language Reasoning for Cross-View Geo-Localization —
- Assumption-lean logistic regression with missing covariates —
- Kurate: Scalable Scientific Quality Analysis —
- ATLAS-AL: Adaptive Trust-Region for Latent Adversarial Searches via Active Learning —
- Localize Any Object in X-Ray Security Scans without Human Annotation —
- Structuring MoE Expert Selection for Agentic Reinforcement Learning —
- Weight Oracles: Reading Neural Network Weights with Language Models —
- Compositional Concept Erasure in Text-to-Image Diffusion Models via Hierarchically Grounded Semantic Surgery —
- A doctrine-grounded visual question answering dataset for Tactical Combat Casualty Care —
- Stepped MoE: Segment-Level Routing with Configurable Inference Complexity —
- Tracking Is Not Permanence: What Video World Models Keep of a Hidden Object —
- Who Wrote It Is Not Enough: Detecting Who Contributed the Insight —
- Identity-Conditioned Score Fusion for Open-Set Person Re-Identification —
- SimCortex v2: Joint Cortical Surface Reconstruction with Near-Zero Collisions and Self-Intersections —
- GeoWM: Efficient Direct World Modeling in Explicit Geometry —
- HyperNSDE: Personalized Neural SDEs for Joint Static-Longitudinal Clinical Data Generation —
- WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification —
- Semantic Capability Acquisition and Specialization During Vision-Language Model Fine-Tuning —
- Bayesian Optimization on Function Spaces via Sparse RKHS Manifolds —
- Learnable Spectral Activations —
- 2d-fet-bench: from spatial reasoning to fet design on flakes —
- AccentCL: Robust Accent Classification with Incremental Expansion —
- Decoupling What from Where: How Should a Small GUI Grounding Model Receive the Action Type? —
- Active Feature Acquisition for Cost-Efficient Temporal Prediction with Reduced Participant Burden —
- AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generation —
- Auditable Claims about AI Agents —
- ElasticFit: Fit-Aware 3D Object Insertion via VLM Reasoning and Generative Adaptation —
- Protective Perturbations Must Survive the Resize: Scale-Robust Image Immunization against Malicious Editing —
- Structure, Not Belief: Correlated Thompson Sampling from LLM-Derived Covariance in Combinatorial Semi-Bandits —
- In With the Old: Enhancing 'Classical' Document Automation with Generative AI —
- Does Muon Need Fine-Grained Spectral Shaping? —
- Closing Ambient Clinical Documentation Gaps with Automated Provider Queries —
- Two-Sample Testing for Random Graphs without Vertex Correspondence —
- On Open-Ended Information Seeking for Information Elicitation Agents —
- Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates —
- Not What a Child Expressed: Auditing the Sign-to-Text Safety Interface in Child-Facing AI —
Important terms
- CANDLE project
- This project uses cortical null-space decomposition to improve noninvasive brain source imaging, offering more interpretable results than older methods.
- Observable neural ordinary differential equations
- These equations help researchers track how brain activity changes over time with greater precision, aiding in the modeling of neural signals.
- Structural-frontier evaluation
- This method finds hidden failures in drug property models by testing them against complex biological structures, suggesting a need for better testing environments.
- Biological representation learning by masking gene expression
- Strategically hiding parts of a genome helps models learn better underlying biological structures, allowing them to grasp true gene function.
- activation-based mitigation of self-preference in LLM evaluators
- This technique stops large language models from unfairly favoring their own outputs when judging others, which is key for trustworthy AI evaluations.