Home › Digests › DailyResearch papers — 2026-09-14 Today's papers The papers Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to 100+ Languages via Back-Translation — This paper examines the efficacy of multilingual LLM watermarking, revealing that current methods are not "truly multilingual" because they fail to remain robust against translation attacks in medium- and low-resource languages. [episode] UltraQuant: 4-bit KV Caching for Context-Heavy Agents — The paper introduces UltraQuant, an advanced method for efficient Key-Value (KV) caching in large language models, specifically targeting the memory bandwidth bottleneck inherent in context-heavy agent applications. [episode] ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning — ShadowPEFT introduces a novel parameter-efficient fine-tuning paradigm by decoupling the core language model backbone from task-specific adaptation through a centralized "shadow model." This approach is crucial for deploying large models in constrained environments, as it allows [episode] DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models — This paper presents DropVLA, an action-level backdoor attack on Vision–Language–Action (VLA) models. [episode] A Mechanism-Coupled Split Window Network for Medium- to High-Resolution Land Surface Temperature Retrieval — This paper introduces the Parallel Component Decoupled Neural Network (PCD-Net), a framework designed for medium- to high-resolution land surface temperature (LST) retrieval from thermal infrared observations. [episode] A unified self-supervised framework for single-frame Fresnel CDI and overlapped ptychography — This paper presents an extension of the PtychoPINN framework to unify single-exposure Fresnel coherent diffraction imaging (CDI) and overlapped ptychography within a single self-supervised formulation. [episode] AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally — This paper presents AdaRoPE, a method designed to optimize Rotary Position Embedding (RoPE) by allowing individual attention heads to learn unique rotation frequencies and attention scaling factors. [episode] Unified Text-Image Generation with Weakness-Targeted Post-Training — This paper presents a method for "fully unified text-image generation," addressing the limitations of existing multimodal models that rely on "manually controlled modality switching." By enabling models to autonomously transition from textual reasoning to visual synthesis within [episode] MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models — This paper presents Modality-Adaptive Decoding (MAD), a training-free method designed to mitigate "cross-modal hallucinations" in Multimodal Large Language Models (MLLMs). [episode] Measuring Pragmatic Influence in Large Language Model Instructions — This paper introduces a framework for measuring "pragmatic framing"—the use of interpersonal or contextual cues, such as authority or urgency, that shape how instructions are interpreted without altering the task content itself. [episode] FEAT: A Linear-Complexity Foundation Model for Extremely Large Structured Data — FEAT is a linear-complexity foundation model designed to handle extremely large structured datasets by overcoming the scalability and generalization limitations of current structured data foundation models (SFMs). [episode] Robust Trust — This paper characterizes optimal decision-making when an agent relies on an informed but potentially misaligned adviser, such as an AI system. [episode] Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills — The paper introduces "Graph-of-Skills" (GoS), a novel retrieval mechanism designed to improve agent performance in complex, multi-step tasks by understanding the structural dependencies between available skills. [episode] OA-NBV: Occlusion-Aware Next-Best-View Planning for Human-Centered Active Perception on Mobile Robots — This paper introduces OA-NBV, an occlusion-aware Next-Best-View planning pipeline designed for human-centered active perception on mobile robots. [episode] DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction — I apologize, but the source material for the paper titled "DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction" was not provided in your request. [episode] Representation Before Training: A Practical Benchmark for Generative Medical Event Model Tokenization — The paper details a rigorous, practical benchmark for tokenizing medical event data, specifically focusing on the representation of clinical observations before generative model training. [episode] Generative AI Assisted Workflows in Architectural Conceptual Design: Performance, Creative Self-Efficacy, and Cognitive Load — This paper investigates how generative artificial intelligence (GenAI) influences "performance, creative self-efficacy, and cognitive load" during architectural conceptual design tasks. [episode] Project Rachel: Can an AI Become a Scholarly Author? — This paper documents Project Rachel, an action research study that created and tracked a complete "AI academic identity named Rachel So" to investigate how the scholarly ecosystem responds to AI authorship. [episode] AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training — Membership Inference Attacks on Recommender System: A Survey — When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs — Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference — WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics — MedCollab: IBIS-Guided Multi-Agent Collaboration with Hierarchical Disease Relation Chains for Clinical Diagnosis — The Vienna 4G/5G Drive-Test Dataset — Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR — Tunable Latent Generative Priors for Compressed Sensing and Inverse Problems — Much of Geospatial Web Search Is Beyond Traditional GIS — Class-wise Contribution Estimation via Logit Maximization for Federated Learning — UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding — A retrieval conditioned rebinding circuit for dynamic entity tracking in large language models — When Context Misleads: Surprisal, Energy and Attention Entropy as Metrics of Coherence Illusions in LLMs — AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning — Language Models for Portuguese: A Systematic Mapping Study — A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration — AI Safety: Not Optional, Not Later — SoulAuth: An Actor-native Identity Architecture and Rust Reference Implementation for Humans and Long-lived AI Actors — Who Pays for Open Review? Visible Author Reputation and Its Effect on Ratings — Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite — Assessment of Non-Institutional AI Tool Usage Among Clinicians — When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline — Estimating Uncertain Spatial Relationships in Robotics — Continuous Learning of Gravity Field Irregularities Around Small Bodies via Neural Hamiltonian ODEs — COVIDx-US -- An open-access benchmark dataset of ultrasound imaging data for AI-driven COVID-19 analytics — Synthetic Blips: Generalizing Synthetic Controls for Dynamic Treatment Effects — Computing linear sections of varieties: quantum entanglement, tensor decompositions and beyond — Guided Adversarial Robust Transfer Learning with Source Mixing — Protect Your Score: Contact Tracing With Differential Privacy Guarantees — A Training-free Method for LLM Text Attribution — Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks — FLOAT Drone: A Fully-actuated Coaxial Aerial Robot for Close-Proximity Operations — Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision — TestDG: Test-time Domain Generalization for Continual Test-time Adaptation — GLaMoR: Consistency Checking of OWL Ontologies using Graph Language Models — Towards Large Language Models for Lunar Mission Planning and In Situ Resource Utilization — Minimisation of Quasar-Convex Functions Using Random Zeroth-Order Oracles — A Survey on Foundation Models for Personalized Federated Intelligence — LLM-BabyBench: Can Language Models Plan in Worlds They Can Simulate? — RedactOR: An LLM-Powered Framework for Automatic Clinical Data De-Identification — An ab initio foundation model of wavefunctions that accurately describes chemical bond breaking —