AI papers — 2026-09-25
The Qwen-Planner-Agent framework explores a closed-loop system for AI agents to perform real-world mobile planning. This research develops an agent structure that combines planning capabilities with an agent architecture to enable autonomous decision-making in dynamic environments. The study pushes beyond simple task execution by creating a feedback loop where the agent can iteratively refine its plans based on environmental interactions.
This work attempts to build a robust framework for handling complex mobile planning scenarios effectively. While the abstracts do not detail specific quantitative results for Qwen-Planner-Agent, the underlying theme suggests an effort toward more reliable, self-correcting planning agents in practical settings. The broader context touches upon aligning cross-modal attention and using neuro-symbolic AI for industrial configuration.
The investigation into tracking states versus tracking cosets explored the algebraic framework for learned state tracking. Researchers examined how different approaches handle the underlying structure of the system when maintaining a consistent state representation versus focusing on tracking cosets. This theoretical work suggests that the choice between these two paradigms has significant implications for how accurately a model can predict future system behaviors.
The development of Augur presents a synthetic decision lab designed to rehearse reactions to product and policy changes. This provides a controlled environment for testing how learned policies respond to external shifts in the operational landscape. This contrasts with efforts focused on refining model safety through chance-constrained fine-tuning, which aims to manage risk bounds during training.
Simultaneously, work on ADATEX4D addresses texture capacity allocation within 4D gaussian splatting. Also, GHOST-Q investigates grounding hallucinations in quantized vision-language models by focusing on overlooked trade-offs at the same score level. These diverse efforts point toward a broader research agenda involving theoretical modeling of state tracking and practical applications in decision making, safety constraints, and model fidelity across various modalities.
The work on exploiting piecewise smooth tree priors for multi-fidelity bandits involved testing how structural assumptions affect selection when balancing exploration and exploitation across different fidelity levels. Researchers implemented a method to leverage these priors to guide bandit decisions, which indicated a measurable impact on convergence speed compared to standard approaches. This suggests that incorporating prior knowledge about the smoothness of the underlying function space can lead to more efficient resource allocation in scenarios with multiple data collection fidelities.
The synthetic hospital project focused on creating an open and verifiable longitudinal electronic health record benchmark validated by physicians. This provided a crucial real-world context for understanding data integrity and long-term tracking challenges in clinical settings. Furthermore, investigations into how adversarial influence scales in multi-agent systems explored the limits of robustness when agents interact dynamically. These studies revealed complex patterns in influence propagation that are not easily predicted by simple linear models.
The study on style versus self in zero-shot code attribution by large language models demonstrated that superficial surface cues are surprisingly effective predictors for code origin. This is significant for understanding how these models generalize without explicit training examples. This contrasts with efforts to expand neural network verification through NNV3, which aimed to push verification capabilities into novel architectures and domains.
SciWalker addressed the challenge of synthesizing scientific coding problems by using operator graphs and execution feedback. This showed how structured representations can help in generating relevant training data for models. Finally, research into self-play pretraining with zero data explored methods for teaching agents through interaction alone. A self-audit of LLM-inferred prompt structure examined the reproducibility of evaluation conclusions derived from these systems.
The work on reachability-based formal verification of graph neural networks explored methods for rigorously checking properties of these complex models by analyzing their state spaces. One approach involved a framework that leveraged AT-SKM-Net, an accelerated trainable sampling Kaczmarz-Motzkin framework designed for linear hard-constraint feasibility on dynamic graphs. This suggests an effort to find feasible solutions within these constrained systems using this specific iterative sampling method.
Simultaneously, research into auditing large language models focused on PrivDrift, which examined user-secret leakage under topic drift during active conversations. This indicates a concern for privacy in conversational AI. Further advancements were made in accelerating video diffusion through training-free trajectory routing techniques. In parallel, the R-DEIM Net introduced an efficient rationale-augmented dual-expert interaction model specifically for paraphrase detection. These investigations point toward efforts to enhance the robustness, privacy, and efficiency of various machine learning architectures across different domains.
The work on generating and revising strategic plans with agentic AI focused on how models handle rejection and long-horizon reasoning. One line of inquiry explored whether the stated reasons for a model rejecting a candidate actually contribute to the quality of the final output. This suggests that simply providing an explanation might not be sufficient. This connects to efforts aimed at mitigating biases in long-horizon reasoning, where SAGE was developed specifically to counteract these biases by employing topological guidance.
Research into search-aware reinforcement learning for multi-component query understanding within Roblox game search indicates a move toward more sophisticated agentic capabilities in complex environments. Similarly, the Jev-Mobile project introduced Jev as an executor for mobile GUI agents, suggesting practical applications for deploying these reasoning capabilities in real-world interfaces. ExplorationBench measures AI systems' exploration within verifiable alien worlds to gauge their ability to navigate unknown spaces effectively. Finally, work on minimally invasive steering of language models suggests a refinement in how we can guide these agentic systems without completely overriding their intrinsic reasoning processes.
The work on TrackEverything focused on developing a method for long horizon dense tracking by utilizing de-duplicating three dimensional scene representations. This approach suggests that by effectively managing and reducing redundancy within the 3D scene data, the system can maintain tracking capabilities over extended periods. In contrast, PoEM explored predicting reinforcement learning outcomes from existing policies. This implies a focus on leveraging current model behavior to forecast future performance in an RL setting.
Research into retrieval-augmented fact checking in speech addresses the issue of trust by integrating external knowledge sources during verification processes. AD-WM introduced action-discriminative world models specifically designed for counterfactual model predictive control, aiming to understand how different actions affect the predicted world state under uncertainty. A probabilistic approach was also investigated for model alignment with human comparisons. This suggests a method to bridge the gap between learned models and human perception. Finally, there is work on a unified theory of exact inference and learning within exponential family latent variable models.
The work on time-series foundation models that understand data revisions explored methods for modeling sequential data where the underlying process itself is subject to changes. One study focused on how these models can be adapted when new data points arrive, suggesting a path toward more dynamic understanding of evolving systems. This connects conceptually to the research into order-theoretic characterization of consistent inductive inference, which seeks to formalize how inferences remain valid even as the underlying structure might shift.
The exploration of sequential confidence sets for coverage-constrained conformal model selection directly addresses this uncertainty by providing a principled way to select models when data revisions are present and constraints on coverage are important. These approaches suggest that future foundation models should incorporate mechanisms that explicitly track and adapt to temporal shifts in data distributions rather than treating the input as static.
The TAM-Chain focused on using absorbing Markov chains and Shannon entropy uncertainty quantification to tackle false negatives in multi-scale thyroid cytology classification while also addressing domain shift adaptation. Researchers explored how these probabilistic methods could improve the robustness of the classification system when moving between different data domains. The findings suggest that quantifying uncertainty through entropy provides a valuable metric for identifying instances where the model is least confident, which can then guide targeted interventions for suppression of false negatives. What remains open is determining the optimal weighting scheme between the Markov chain transition probabilities and the Shannon entropy measure to achieve maximal suppression without introducing excessive false positives in complex biological samples.
Today's papers
- Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents. [paper]
- Who Holds the Pen? Let Specifications, Not Agents, Sign Off. [paper]
- Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations. [paper]
- An Empirical Study of VLM Pipelines for Long-Document QA. [paper]
- When Temporal Perturbations Act Like Sensor Biases: Label-Free Auditing of Wearable Activity Recognizers. [paper]
- Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration. [paper]
- Neuro-symbolic AI for Industrial Configuration. [paper]
- ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation. [paper]
- Tracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking. [paper]
- Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes. [paper]
- Beyond Average Safety: Chance-Constrained LLM Fine-tuning. [paper]
- ADATEX4D: adaptive texture capacity allocation for 4D gaussian splatting. [paper]
- GHOST-Q: Towards Studying Grounding Hallucinations Overlooked Under Same-score TradeOffs in Quantized VLMS. [paper]
- Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems. [paper]
- Automated Regulatory Compliance Question Answering in Financial Services with Domain-Adapted Retrieval-Augmented Generation. [paper]
- Low-Cost Assays for Measuring Model Behavior Across Vendors and Releases. [paper]
- Canopy: Exploiting Piecewise Smooth Tree Priors for Multi-Fidelity Bandits. [paper]
- Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark. [paper]
- How does Adversarial Influence Scale in Multi-Agent Systems?. [paper]
- Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models. [paper]
- NNV3: Expanding Neural Network Verification to New Architectures and Domains. [paper]
- SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback. [paper]
- Self-Play Pretraining with Zero Data. [paper]
- How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure. [paper]
- Reachability-Based Formal Verification of Graph Neural Networks with Node and Edge Features. [paper]
- AT-SKM-Net: An Accelerated Trainable Sampling Kaczmarz-Motzkin Framework for Linear Hard-Constraint Feasibility on Dynamic Graphs. [paper]
- PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations. [paper]
- Accelerating Video Diffusion via Training-Free Trajectory Routing. [paper]
- R-DEIM Net: An Efficient Rationale-Augmented Dual-Expert Interaction Model for Paraphrase Detection. [paper]
- HEXIS: Compiling Skills into Extended Finite State Machines.
- Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale. [paper]
- EnigmaForge: The Question Is Hidden in the Story. [paper]
- GRASP: Generating, Revising, and Assessing for Strategic Planning with Agentic AI. [paper]
- Does a model's stated reason for rejecting a candidate do any work?. [paper]
- Search-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search. [paper]
- Jev-Mobile: Jev as an Executor for Mobile GUI Agents. [paper]
- SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance. [paper]
- ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds. [paper]
- A Living Benchmark for Information Retrieval from Electronic Health Records. [paper]
- Minimally Invasive Steering of Language Models. [paper]
- TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations. [paper]
- PoEM: Predicting RL Outcomes from Existing Policies. [paper]
- To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech. [paper]
- AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control. [paper]
- A Probabilistic Approach for Model Alignment with Human Comparisons. [paper]
- A Unified Theory of Exact Inference and Learning in Exponential Family Latent Variable Models. [paper]
- Stable and Faithful Explanations for Knowledge Tracing. [paper]
- CFD Correction of Open Tip Clearance Flow in a Compressor Cascade Using VAE Latent Space Adaptation. [paper]
- Time-Series Foundation Models That Understand Data Revisions. [paper]
- Uncovering Residential PV-EV Co-Adoption from Smart-Meter Data: Load Archetypes and Detection for Demand-Side Planning. [paper]
- Sequential Confidence Sets for Coverage-Constrained Conformal Model Selection. [paper]
- Stochastic Inertial Krasnosel'skii-Mann Iteration Achieves Near-Optimal Sample Complexity. [paper]
- An Order-Theoretic Characterization of Consistent Inductive Inference. [paper]
- CARE: Condition-Aware Representation Regularization for Diffusion Models. [paper]
- Matrix Aggregation Operators. [paper]
- Leakage-Safe Machine Learning for Hydrogen Embrittlement Detection in 316L Stainless Steel: A Region-Held-Out Evaluation of Texture and Deep Features in SEM Micrographs. [paper]
- TAM-Chain: Multi-Scale Thyroid Cytology Classification via Absorbing Markov Chains and Shannon Entropy Uncertainty Quantification for False-Negative Suppression and Domain-Shift Adaptation. [paper]
- Physics-Informed Self-Supervised Learning for Joint Wire Calibration and Interaction Position Reconstruction in Multi-Wire Parallel Plate Avalanche Counters. [paper]
- OPDiv: Optimal Selection of Top-K High-Scoring, Diverse Compounds. [paper]
- fable.intermittent: benchmarking probabilistic forecasting methods for intermittent time series. [paper]
The papers
- ESAFusion: LiDAR--4-D Radar Fusion via Local Geometric Complementation and Multiscale Adaptive Interaction for 3-D Object Detection — ESAFusion is an evidence-aware and scale-adaptive framework that combines local geometric complementation with multiscale adaptive interaction for 3-D object detection by fusing LiDAR and 4-D radar data. [episode]
- A Streaming End-to-End Framework For Spoken Language Understanding —
- DCRMTA: Deep Causal Representation Learning for Multi-Touch Attribution —
- A Probabilistic Approach for Model Alignment with Human Comparisons —
- Band-Attention Modulation Network for Robust Face Forgery Detection —
- A Unified Theory of Exact Inference and Learning in Exponential Family Latent Variable Models —
- Generating Interesting Scientific Ideas using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders —
- Deep Positive-Unlabeled Anomaly Detection for Contaminated Unlabeled Data —
- AIR: Analytic Imbalance Rectifier for Continual Learning —
- Evaluation of OpenAI o1: Opportunities and Challenges of AGI —
- Comparing YOLOv11 and YOLOv8 for instance segmentation of occluded and non-occluded immature green fruits in complex orchard environment —
- Foundations of Large Language Models —
- Bias-variance decompositions: the exclusive privilege of Bregman divergences —
- Time-Varying Bayesian Optimization Without a Metronome —
- The Majority Vote Paradigm Shift: When Popular Meets Optimal —
- Language Specific Knowledge: Do Models Know Better in X than in English? —
- The kernel of graph indices for vector search —
- Capturing Unseen Spatial Heat Extremes Through Dependence-Aware Generative Modeling —
- Stacked SVD or SVD stacked? A Random Matrix Theory perspective on data integration —
- Unraveling the cognitive patterns of Large Language Models through module communities —
- Interactive In-Meeting Speaker Correction with Human Feedback —
- Enabling Approximate Joint Sampling in Diffusion LMs —
- Classical AI vs. LLMs for Decision-Maker Alignment in Health Insurance Choices —
- Digital Contrast CT Pulmonary Angiography Synthesis from Non-contrast CT for Pulmonary Vascular Disease —
- MLPerf Automotive —
- Cross-Task Generalization in Handwriting-Based Alzheimer's Screening via Vision Language Adaptation —
- OncoVision: Integrating Mammography and Clinical Data through Attention-Driven Multimodal AI for Enhanced Breast Cancer Diagnosis —
- A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs —
- LeafTrackNet: A Deep Learning Framework for Robust Leaf Tracking in Top-Down Plant Phenotyping —
- Interpretable Similarity of Synthetic Image Utility —
- Data relativistic uncertainty framework for low-illumination anime scenery image enhancement —
- IDRBench: Benchmarking the Interactive Capabilities of Deep Research Agents —
- Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations —
- Do not be greedy, Think Twice: Sampling and Selection for Document-level Information Extraction —
- WaterClear-GS: Optical-Aware Gaussian Splatting for Underwater Reconstruction and Restoration —
- LLM surprisal is necessary but not sufficient to capture English garden-path effects: Evidence from joint latent modeling of reading paradigms —
- TIDE: Temporal Incremental Draft Engine for Self-Improving LLM Inference —
- Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference —
- MDE-VIO: Enhancing Visual-Inertial Odometry Using Learned Depth Priors —
- TabSieve: Explicit In-Table Evidence Selection for Tabular Prediction —
- Context-aware Skin Cancer Epithelial Cell Classification with Scalable Graph Transformers —
- Decoding ML Decision: An Agentic Reasoning Framework for Large-Scale Ranking System —
- MPFlow: Multi-modal Posterior-Guided Flow Matching for Zero-Shot MRI Reconstruction —
- MOOSEnger: A Simulation-Aware AI Agent Framework for the MOOSE Ecosystem —
- Learning Causal Structure of Time Series using Best Order Score Search —
- Match4Annotate: Cross-Video Annotation Transfer in Ultrasound via Implicit Feature Flow-Guided Matching —
- Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety —
- MultiwayPAM: Multiway Partitioning Around Medoids for LLM-as-a-Judge Score Analysis —
- Beyond Pairwise Attention: Higher-Order Modular Attention for Efficient Sequence Learning —
- DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units —
- One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation —
- OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding —
- MemGuard-Alpha: Limits of Membership Inference for Detecting and Filtering Memorization-Contaminated Signals in LLM-Based Financial Forecasting —
- Invertible Query-Key Coupling Composes with Attention Mechanisms —
- LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches —
- COMPASS: Fusion-Matched Supervision for Missing-Modality Human Sensing —
- ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving —
- IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models —
- Continued Pretraining of FinBERT on Finnish Histopathological Reports: Train-Time Signals and Proxy Downstream Correlations —
- Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens —
- GeoBlur: Epipolar Geometry Estimation from a Single Motion-Blurred Image —
- Structured 3D Latents Are Surprisingly Powerful: Unleashing Generalizable Style with 2D Diffusion —
- CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs —
- An Empirical Study of Automating Agent Evaluation —
- BEHAVE: Real-Time Modeling of Human Systems as Observable Complex Dynamical Systems and Operational Objects for Physical AI —
- GOMA: Toward Structure-Driven Multimodal Alignment from a Graph Signal Smoothing Perspective —
- BARRIER: Bounded Activation Regions for Robust Information Erasure —
- Pointwise Generalization in Deep Neural Networks —
- Lossless Anti-Distillation Sampling —
- Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow —
- Every Component Is a Lookup: One Linear Graph for Interaction, Composition and Attribution —
- Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork —
- Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems —
- A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring —
- Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation —
- Planning Takes More Than Token Prediction: Causal Plan for Benchmarking and Building Physically Grounded Embodied Reasoners —
- NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation —
- Deep Learning-based 3D Oral Cavity Reconstruction Using 2D Intraoral Images —
- 3D Oral Modelling with Improved Vertex Distribution Using Matching-Based Learning —
- From Observation to Intervention: A Causal Audit of Expert Importance in Mixture-of-Experts Models —
- JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence —
- Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean —
- Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning —
- SOV-CAD: Stepwise Orthographic Views Guided CAD Modeling Sequence Reconstruction —
- A hybrid analytical-PINN model for subsurface simulation of geothermal heat exchangers in heterogeneous underground —
- An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations —
- Answering Path Queries under Linear and Guarded Existential Rules —
- A Multi-level Information Integration Framework for Physically Verifiable Fault Diagnosis of Rotating Machinery —
- CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models —
- Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization —
- TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation —
- Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection —
- Beyond Forgetting: Diagnosing and Harnessing Shared Reasoning in Continual RLVR —
- Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots —
- Cluster Assignments in Soft Targets Shape Speech Representations: Evidence from S-JEPA —
- CLaST: Context-aware Contrastive VAE for Probabilistic Time Series Forecasting —
- DecoVAE: a Lightweight Interpretable Trend-Seasonal VAE Framework for Efficient Probabilistic Time Series Forecasting —
- How Weight Encoding Affects Language Model Placement and Performance on the Apple Neural Engine —
- Toward a Foundation Plug-and-Play Prior for Computed Tomography Reconstruction via a Multimodal Diffusion Model —
- KISS-GS: 3D Gaussian Splatting Compression Kept Simple —
- SignMimic: Robust High-Quality Sign Language Motion Generation via Human-Shape-Oblivious Pose Transfer Guidance —
- When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability Routing —
- Framing by Wording, Framing by Selection: A Large-Scale Two-Dimensional Audit of French News Headlines, 2022-2025 —
- Stable and Faithful Explanations for Knowledge Tracing —
- TW3Cast: A Frozen Router of Lightly Fine-Tuned Foundation Models for Time-Series Forecasting on GIFT-Eval, Selected Entirely on the Training Split —
- Certified Task-Conditioned Active Observability —
- Sequential Confidence Sets for Coverage-Constrained Conformal Model Selection —
- Heartian: Physiology-Aware Relightable Gaussian Head Avatar —
- Stochastic Inertial Krasnosel'skii-Mann Iteration Achieves Near-Optimal Sample Complexity —
- PAWS: Policy-driven Agentic World Simulation —
- An Order-Theoretic Characterization of Consistent Inductive Inference —
- SMILESGNN: Interpretable Clinical Toxicity Prediction via SMILES-Graph Cross-Attention Fusion —
- Pistis Technical Report —
- BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines —
- CFD Correction of Open Tip Clearance Flow in a Compressor Cascade Using VAE Latent Space Adaptation —
- Speculative Evaluation of Stochastic LLMs —
- CARE: Condition-Aware Representation Regularization for Diffusion Models —
- Matrix Aggregation Operators —
- SpaFactor: Lightweight Spatial Context-Aware Gene Program Modeling for Histology-to-Transcriptomics Inference —
- When Explanations Cannot Be Read: Measuring and Correcting SHAP and LIME Rendering for Right-to-Left Languages —
- Leakage-Safe Machine Learning for Hydrogen Embrittlement Detection in 316L Stainless Steel: A Region-Held-Out Evaluation of Texture and Deep Features in SEM Micrographs —
- DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs —
- TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment —
- Time-Series Foundation Models That Understand Data Revisions —
- Uncovering Residential PV-EV Co-Adoption from Smart-Meter Data: Load Archetypes and Detection for Demand-Side Planning —
- Token Clustering and Semantic Sequence Mamba for Hyperspectral Image Classification —
- Auditability Is Not One Property: Rule Overlap, Behavioural Agreement, and Composition in Reinforcement Learning —
- SGA: Uncertainty Quantification for Multi-Step Forecasting in Time Series Foundation Models —
- NumericJev: Jev-like LLM Numerical Decoding with Multiway Decision Trees —
- TAM-Chain: Multi-Scale Thyroid Cytology Classification via Absorbing Markov Chains and Shannon Entropy Uncertainty Quantification for False-Negative Suppression and Domain-Shift Adaptation —
- Learning to Discover Interesting Mathematics —
- Physics-Informed Self-Supervised Learning for Joint Wire Calibration and Interaction Position Reconstruction in Multi-Wire Parallel Plate Avalanche Counters —
- UO-FIE: Combining Exact-Label Supervision with Graded Utility for Factivity Inference —
- fable.intermittent: benchmarking probabilistic forecasting methods for intermittent time series —
- Adversarial Closed-Loop Curriculum for Evolving Role-Playing Agents —
- UltraBench 2: Towards Robust Evaluation of Vision Foundation Models on Ultrasound —
- Reward Hacking Challenges Oversight of Autonomous Research Agents —
- RLVR landscapes for iterated multiplications can be benign: Insights from spin-glass theory —
- PePESeg3D: Perception Prior Enhances Multi-Scale Segmentation for 3D Gaussian Splatting —
- Training Object Permanence in World Models —
- OPDiv: Optimal Selection of Top-K High-Scoring, Diverse Compounds —
- Beyond Static Graph World Models: Learning Stochastic Latent Dynamics over Evolving Topologies —
- Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks —
- Thinking Leakage: A Causal Audit of NoThink Post-Training in Hybrid Reasoning Models —
- M-plicits: Neural Implicit Surfaces via Nested Multiscale Residuals —
- TRACER: Trajectory-Aligned Learning for Multi-Turn User Simulation —
- Driving Epidemic Models with AI Agents: the Epydemix Agent Framework —
- Progressive Skill Discovery as Access Control for Tool-Using LLM Agents: Structural Governance through Role-Scoped Capability Delivery —
- Federated Learning of AnDE Classifiers —
- LabFactory: Building and Evaluating Executable AI Labs —
- An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection —
- Exact Bayes Regret and Asymptotic Optimality in High-Dimensional Gaussian Bandits —
- Upholding Robustness in Federated Learning: Trends, Emerging Strategies, and Research Opportunities —
- PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs —
- Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning —
- Temporal Taxation Compounds Under Post-Training Compression of Whisper Models —
- GeoNLI - A Natural Language Interpreter for Satellite Imagery —
- Technical Manual for Toolkit for Confidence-Corpus Consistency, Corpus Absorption and Rule Learning via Fine-Tuning on a Fabricated Corpus —
- Evaluating Cross-region Generalization for Wavelet-Diffusion Precipitation Downscaling —
- Selective Inference for Deep Clustering in Latent Spaces —
- Small yet Assistive: Spatially-Aware Post-Training for Low Vision —
- Reinforcement Learning with Verifiable Rewards for Small Search Agents —
- Agent Memory with Episodic Retrieval for Financial Decision-Making —
- Learned Cross-Task Relationships in Multi-Task Models —
- The Mechanics of Delta Learning: Target Design for Generalizable Scientific Machine Learning —
- Script Choice in LLMs: Evidence for Late-Layer Commitment —
- Vector Bellman Theory for Multichain Robust Average-Reward Markov Decision Processes —
- Monitoring Urban Traffic Dynamics at Fine Spatiotemporal Resolution Using Distributed Acoustic Sensing and Deep Learning —
- DrGait: Biomechanically Grounded Visual Reasoning for Interpretable Clinical Gait Analysis —
- Stream Recursion Model (SRM) —
- DeltaWAM: Delta World Action Models for Bimanual Manipulation —
- CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models —
- COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages —
- When Does Unsupervised Learning Succeed or Fail? A PoS Perspective on Reconstruction-Based Anomaly Detection —
- M squared PFN: End-to-End Disentangled Alignment for Generalizable Multimodal In-Context Learning in Alzheimer's Disease —
- LastOPD: Taming Collapse in Latent On-Policy Distillation —
- RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers? —
- Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning —
- Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding —
- MEVL-STP: Multi-Encoder and Vision Language Model for Arbitrarily Shaped Scene Text Spotting —
- Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation —
- Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation —
- Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models —
- Image Fidelity is Not Field Fidelity: Joint Thermodynamic Reconstruction and Error Localization in Neural Tomography —
- Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents —
- Harness Tokenomics: A Router for the Enterprise Agentic Control Plane —
- PFArena: Benchmarking Language Models for Protein Modification —
- ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation —
- PlenoCI: Plenoptic CharacterIstics for View Dependence Aware Change Classification —
- HelloWorld: Towards Practical Applications of Generative Driving World Models —
- Response-state Learning for Transferable Vibrational Spectroscopic Characterization with Electron Prior —
- From Static Personal Values to Contextualized Personalization: Bayesian Personalized Value Alignment for LLMs —
- Exploiting Target Knowledge from MLLMs for Robust Few-Shot Segmentation —
- MoVISA: Multi-Token Reasoning for Video Object Segmentation —
- Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning —
- Passive LWIR Hyperspectral Ranging via Transmittance Extraction and Distance Alignment —
- Spectral Graph Neural Networks with Hermite Polynomials: A Comprehensive Study —
- Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models —
- Only What Was Seen: Observation-Gram Compaction of View-Dependent Appearance in 3D Gaussian Splatting —
- Automatic Rank Allocation for Low-Rank Adaptation in Large Language Models via lp Regularization —
- Learning from Mixed-Quality Deployment Experience for Robot Manipulation —
- Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms —
- FluidRain: Incompressible Rain Flow as an Attention Bias for Loop-in-Loop Video Deraining —
- When Does Action Credit Need Updating? —
- AlphaDiverse: Post-Training Local Quantitative Research Agents for Diverse Exploration in Alpha Factor Mining —
- MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks —
- Growth-Inspired Graph Generation and Inverse Design of Mechanical Lattices via Dot Matrices Database Augmentation and GCNN —
- Generative Atmospheric Super-Resolution from Heterogeneous In Situ Observations through Composable Interfaces —
- RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation —
- Exploiting answer-invariant redundancies in satellite imagery for efficient VLM inference on edge —
- Personalised federated learning for Riemannian and Euclidean EEG decoding —
- Where Hallucinations Live: A Cross-Architecture Circuit in VQ-Tokenized Vision-Language Models —
- SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL —
- From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents —
- Empath: Tracing Multi-Level Emotion Dynamics in Crisis Counseling Dialogues —
- Transformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning —
- EIB-Net: Entropy-Guided Information Bottleneck for Generalizable AI-Generated Image Detection —
- BranchShine-CR: Compact Multilingual IPA Transcription with Self-Conditioned CTC and Consistency Regularization —
- Seeing Is Not Measuring: Tool-Augmented Metric Spatial Reasoning for Vision-Language Models —
- CRISS: A Retrieval-Augmented AI Chatbot for Assisting Cancer Registrars —
- Feature Space Selection and Heterogeneous Effect Estimation for Blood-Brain Barrier Permeability: A Random Forest to the Generalized Random Forest Pipeline —
- A Rapid Pipeline for Training and Deploying ML Models on WeBe Band —
- Physics and Data Driven Transformer-Mamba Framework for Flow Field —
- Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study —
- Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents —
- Downside-Controlled Online Forecast Combination under Delayed and Revised Outcomes —
- Language Specificity vs. Domain Diversity: Benchmarking Transformers for Bangla Medical NER —
- ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks —
- WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model —
- Functional Architecture of European Electricity Trading Markets: Requirements for AI Supported Trading Systems under Regulatory Constraints —
- CounterRoute: Self-Routed Reasoning via Hierarchical Counterfactual Credit Assignment —
- Spectral Amplitude Purification in Distribution Matching for Diffusion Distillation —
- A Concentration Bound for Two-Timescale Actor-Critic Algorithm —
- UpDown-SC: Gravity-Canonicalized Dual-Envelope Scan Context for Indoor LiDAR Place Recognition —
- Less is More: Encoder-only Audio-Visual Segmentation —
- FoCal: Frequency-Oriented Cross-Modal Interaction and Spectral Calibration for Aerial Visible-Infrared Object Detection —
- Tag-Aware Structured Text Translation: Towards a Systematic Understanding —
- Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation —
- Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD —
- Scope Before You Persist: Preventing Cross-Family Interference in Agent Memory —
- Claim-Gated Source-Risk Auditing for Generative Search —
- BanglaKontho: Closing the Long-Form Gap in Bangla Text-to-Speech —
- A Particle-Swarm-Assisted Gradient Meta-Learning Algorithm for Joint Transmit Precoding and STAR-RIS Coefficient Optimization —
- Recoverable Geographic Location Information in Earth-Observation Embeddings —
- A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents —
- Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation —
- Edge AI on Constrained Devices for Binary Sleep-Wake Classification in Dynamic Environments —
- IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking —
- Right Choice of Classification Algorithms Based on Reinforcement Learning for Prediction of Non-Alcoholic Fatty Liver —
- Predicting Emerging Topics from Outliers: A Prospective Study of Weak Signals in Embedding Space —
- An Automated Georeferencing Technique for Multi-Temporal Stope Point Clouds for Downstream Geotechnical Analysis —
- The Entropy Triangle Method (ETM): A novel framework for the prevention of cardiac arrhythmia with a review of more than 10,000 patients —
- When Honesty is Not Enough in AI Debate —
- ASIRF: An Agentic Framework for Context-Dependent Sensitive Information Redaction —
- ImCorr: Sub-pixel Semantic Correspondence via Implicit Feature Decoding —
- FB-GDM: Fully-Bayesian Guided Diffusion Models for High-Dimensional Linear Inverse Problems via Unsupervised Variational Inference —
- FounRef: Robust, Structure-Preserving, and Fast Metric Refinement of Frozen Monocular Foundation Priors with Sparse Anchors —
- ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios —
- Towards An LLM-Driven Unified Conversion Framework for BT and FSM in Autonomous Intelligent Systems —
- EAGER: Enhancing Generative Event Extraction via Reinforcement Learning with Verifiable Rewards —
- Post-Training Leaves Behavioral Shadows on Unrelated Decisions —
- SARFusion: Scene-Aware Routing Fusion for Robust Camera-LiDAR 3D Object Detection —
- TOLA: Text-aware One-Step Latent Adaptation for Diffusion-based Text Image Super-Resolution —
- No More Free Lunch: Corpus Task Complexity Matters as Corpora Grow —
- Policy as Code: A Coroutine-Bridge Harness for Fast-Reasoning Reliability on CAR-bench —
- IronViT: Toward Efficient Generalist Visual Representation Learning —
- Deep learning of longitudinal visual fields predicts glaucoma progression rate and identifies fast progressors —
- Baszta: Data-Centric Fine-Tuning of a Polish Multi-Label Safety Classifier —
- BridgeMem: Causal Dyadic Transition Residuals for Temporal Knowledge Graph Forecasting —
- ALOE: Semantically Addressed Low-Rank Operators for Knowledge Editing —
- Learnable Time-Frequency Masks for Explaining Time-Series Classifiers —
- pylazaro: a Python package for anglicism extraction in Spanish —
- Reasoning Instructions Can Break Answer Decoding in Vision--Language Models —
- Online Task Adaptation via Self-Organisation —
- From Text Decisions to Pixels: An Study of Jev-Style Visual Choice Model —
- PHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark —
- Beyond Feature Reliability: Repeat-Informed Multifractal Curve Regression for Brain-Age Prediction —
- When No One Owns the Judgment: Accountability Under Contribution Dissolution in Human-AI Collaboration —
- Neuralized Multi-Wavelet Decomposition for Time Series Classification and Forecasting —
- TinyCardioUNet: IMU-to-ECG Translation with Graph-Encoded Inter-Axis Dependencies and Tensor Decomposition-Based Parameter Reduction —
- GCUL: Ambiguity Identification in Text Emotion Classification via Cluster-Guided Learning —
- Grammatical "grandmother neurons" are rare in LLMs —
- Hyperbolic Multimodal Continual Learning: A Closest-Admissible Solution —
- FlowAtom: Atom-Based Evidence Aggregation for Multi-Label Website Fingerprinting —
- Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams —
- A Study of the Limits of Collaborative DCT-Based Image Denoising via Interpretable Neural Networks —
- SkinAgent AI: A Safety-Grounded Multimodal Agentic Framework for Non-Diagnostic Skincare Support —
- The Last Human Gate: Forward Deployed Engineering for Governance Automation —
- SEE Challenge 2026: Event-Guided Brightness Adjustment Across a Broad Illumination Range —
- ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts —
- Learning a Flow to Self-Supervised Representations —
- Domain Recentering and Confidence-Weighted Prior Calibration for Vision-Language Models —
- Parts-of-Speech as Emergent Categories in SAE Latent Space —
- Beyond Simple Input-Output Assessment Tasks: Leveraging Automated Programming Assessment for Non-Trivial Courses —
- Epistemic-Probabilistic Model for Guarded Multi-Agent LLM Coordination —
- From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring —
- BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech —
- Shadow Reduction in Ultrasound Imaging Using Differentiable Simulation and Radiance Field Decomposition —
- A Hybrid CNN--State-Space--Attention Backbone with Joint-Embedding Predictive Pretraining for 12-Lead ECG Classification —
- On the second-order optimization for spiking neural networks —
- An auditable conditional-strategy framework for open-ended decision-making in complex lung cancer —
- Lightweight Probabilistic Downscaling from a Deterministic Base Model —
- Segment-Level Risk Discovery in Online Handwriting for Alzheimer's Disease Detection —
- MORE-PLR: multi-output regression employed for partial label ranking —
- When Misalignment Becomes Supervision: Structured Label Noise in Supervised Synthetic CT Generation —
- Likelihood Ranking doesn't Scale Like Prompting in LLMs —
- Wearable ECG Quality Assessment: A Deep Learning and Ambulatory Context-Awareness Approach —
- Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters —
- ICE: Task-Aligned Clifford Latent Fields for Multimodal Graph Foundation Models —
- RD-JEPA: Predictive latent pretraining for few-trajectory transfer across reaction--diffusion equations —
- Machine Unlearning for Gibbs Supervised Learning Algorithms —
- Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code? —
- Neural Transport Nested Sampling —
- Controlling Backchannels in Streamable Full-duplex Models —
- Rufus-Air: An Open LLM Post-Training Recipe —
- agentic-ger: terminology recovery in long-form speech using global context —
- Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures —
- Detecting Glaucoma Across Multi-ethnic Myopic and Non-Myopic Populations Using an Uncertainty-Aware Vision Transformer: A Multicentre Model Development and Validation Study —
- Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition —
- IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis —
- Two Emojis of Difference: What Multilingual Affective Generation Benchmarks Actually Measure —
- SPADE-DFL: Communication-Efficient Decentralized Federated Learning via Derivative-Free Linearized ADMM —
- Frame-to-Panorama Localization and Context-Aware Sampling for Scene-Specific Ship Detection in a Smart Marina Testbed —
- YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech —
- Decoupled Learning and Selection in Slate Recommendation for Privacy and Stability Under Noisy Scores —
- Dense Coverage, Sparse Refinement: Byte-Constrained Cooperative Perception —
- Industrial Anomaly Detection via Defect-Grounded Reasoning in Visual Latent Space —
- Precise Convergence Speed of Clipped SGD —
- AgriCountDINO: Parameter-Efficient Exemplar-Guided Counting and Localization in Agriculture —
- SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories —
- Direct Message Approximation (DMA): A Consistency-Based Framework for Tractable Approximate Inference on Factor Graphs —
- Clinical Intent Extraction: A FHIR-Aligned Representation and the CIRCA Benchmark —
- BLADE: Distilled LLM Regularization for Calibrated Knowledge Graph Completion —
- Who Put the I in AI? Provenance and the Admissibility of Machine Self-Report —
- Evaluating Explanation-Driven Vision-Language Reasoning via Generation Order Interventions —
- Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference —
- PROOF: Profiling Reliability of Object-Level Facts in Large Language Models —
- Spectral-Guided Diffusion: Accelerating Inference via Static Spectral Layer Scheduling —
- What a Cross-Model Fixed-Point Census Can and Cannot Arbitrate About Repetition —
- Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy —
- Delay-of-Gratification as a Multi-Agent Survival Micro-benchmark for Long-Horizon LLMs: Social Exposure, Personas, and Tool Use Budgets —
- EnSiTa - A Trilingual Multi-Domain Parallel Dataset and Benchmark for Domain-Specific Machine Translation —
- AdaPilot: Towards Scene-Adaptive Policy Learning for Cross-Generator Text-to-Image Quality Optimization —
- CataOPD: Catalytic On-Policy Distillation for Large Language Model Reasoning —
- BiGraph-Diffuse: A Bidirectional Diffusion Language Model with Graph-Structured Retrieval For Mental Health Counseling —
- Sample-Weighted End-to-End Trace-Norm Geometry for Multitask Learning —
- Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races —
- Common Covariance Geometry and Certification for Brownian Kernel Ladders —
- CoSWA-YOLOv12: Scale-Invariant Tiny Object Detection and Segmentation of Malaria Parasites —
- The Impossible Trinity of Time-Series Validation: A Conservation Law among Training Sufficiency, Test Coverage, and Temporal Causality —
- Clinical Knowledge Graphs for Chest X-Ray Device Reasoning —
- GeoRefer-Bench: A Benchmark from Referring Pixels to Verifiable Geospatial Reasoning —
- Safe Skill Retirement for Physical Agents —
- ERRAND: Budgeted Maintenance of Agent Memory —
- Generalized Graph Variational Autoencoders: Bounded Divergences Control Posterior Collapse —
- Certified Predictive Value-of-Advice Gating for Cost-Aware Language-Model Guidance in Reinforcement Learning —
- StepCOPS: Closed-Testing Lower-Tail Certificates for Language-Model Policy Selection —
- Spaceborne differential photogrammetry for control-free measurement of large-gradient deformation with structural immunity and a predictable accuracy envelope —
- HiPACE: Hierarchical Phase-Boundary Analysis and Controlled Evaluation of Feature Absorption in Sparse Autoencoders —
- UNWIND: Any-Length Facial Video for Stress Detection without Temporal Windowing —
- Visual Representation and History Modeling for Navigation World Models —
- Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition —
- Benchmarking Arabic--Russian Machine Translation: A Comparison of Fine-tuned NMT and Few-shot LLMs under Rich Morphology and Low Lexical Overlap —
- Is Reasoning Always Useful? Rethinking Reasoning Utility in Universal Multimodal Embeddings —
- Classifier-Dependent Benefits of Pseudo-Labeling for Semi-Supervised Android Malware Attribution —
- ModularSQL: A Runtime Guardrail for the Multiplicity Blind Spot in Text-to-SQL —
- PartHackBench: Certified Equal-Progress Stress Tests for Partial-Credit Tool-Agent Evaluation —
- When Identical Rows Disagree: From Benchmark Identifiability to Replication-Robust Anomaly Detection —
- Long-Tail Adaptive Flow Matching with Explicit Conditional Consistency Guidance for Precise Multimodal Face Synthesis —
- A Computational Framework for Modelling Organisation-Level Semantic Identity from Longitudinal Textual Data —
- Sequential knowledge editing breaks a model's ability to tell good evidence from bad, without costing it accuracy —
- CATCH: Counterfactual Anatomical Tissue Inpainting with Conditional Haar Diffusion —
- QINA: Quantum-Inspired Nonlinear Adapters for Pretrained Vision Models —
- PEEL: Physics-Enabled Evidential Learning for Identifiable Uncertainty in CT Imaging —
- Active Client Selection in Federated Trajectory Prediction with Uncertainty-Awareness and Heterogeneous Complexity —
- PROVE: Proof-guided Regime-aware Operator Verification for Hallucination Detection in Medical Visual Question Answering —
- STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models —
- An Exploratory Ablation of a Small MLA--SSM Hybrid Language Model —
- AgenticCADedit: A Stateful, Tool-Mediated Agentic Approach to Multimodal 3D CAD Editing —
- Optimal Recovery Meets Bayesian Learning: Where Worst-Case Bounds Pay Off —
- Limited Structural Reliability in Public Educational Prediction Benchmarks: A Four-Dimension Audit of Seven Datasets —
- iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model —
- A Manifold-Aware Topic Modeling Approach via Rank-Based Prototypes —
- TTLab at AlexandriaX-2026: A Fine-Tuned Surface Tagger for Arabic Machine-Translation Error-Span Detection and Classification —
- Operator Packages, Proposer Strength, and Construction-Family Plateaus in Office-Scale Verified Search —
- SpectralCTGaussians: Projection-Domain Reconstruction and Basis Material Decomposition for Spectral CT using 3D Gaussian Splatting —
- Albireo: Adaptive, Energy-Efficient Inference Framework for Video Object Detection on the Edge —
- VG-TIE: An interpretable tabular-to-image encoding method based on visibility graphs —
- How To Do Things With Prompts —
- Ingest-Time Fact Compilation for Cost-Efficient and Reliable Question Answering over Revised Corpora —
- Investigating White Blood Cells as a Source of False-Positive Malaria Parasite Detection in African Blood-Smear Images —
- To Think or Not to Think: Allocating Reasoning Where It Helps —
- Graph, Loop, and Harness Engineering for Zero-Trust Agentic Data Engineering and Analytical Processing —
- GBFRVFL: Granular-Ball Computing-Based Fuzzy Random Vector Functional Link Network —
- LLMersion: A Local-First AI Agent Framework for Low-Cost Home Language Learning toward Educational Equity —
- The Sequential Price of Continual Learning —
- ReCalMatch:Reliability-Calibrated Semantic Guidance for Semi-Supervised Fine-Grained Recognition —
- Confident but Wrong: A Constrained Decoding Diagnostic for Low-Resource Automatic Post-Editing —
- Named Entity Recognition using Sliding Window Approach —
- DP-IPI: A Hybrid Differential Privacy Text Rewriting Mechanism for Indirect Personal Identifiers in Clinical Texts —
- Not All Synthetic Data Are Equal: Expert-Committee Audit Screening for Imbalanced Crash-Injury-Severity Prediction in Automated Driving Systems —
- Predicting Symptoms of Amotivation and Anhedonia among University Students with a Novel Oversampling Method —
- Fair Like Us? Auditing LLM Alignment in Resource Allocation —
- Bandit Multiclass PAC Learning: Corrected Lower Bounds, Exact Families, and a Confidence Direct-Sum Phenomenon —
- An Agnostic Sample Compression Scheme for Squared Loss of Near-Linear Size in the Fat-Shattering Dimension —
- Finding Icebergs in Language-Model Workflow: Diagnosing Latent Structural Fragility with Stochastic Semantic Evidence Graphs —
- Safety-oriented pedestrian trajectory prediction at urban intersections using time-to-collision and crossing-zone context —
- Three Ways Classical Test Theory Can Mislead About LLM Judges —
- Decoupling Knowledge and Privacy: Post-Task Self-Distillation Replay for LLM Continual Learning —
- Revalidation Beats Stateful Routing for Scientific Surrogates Under Distribution Shift —
- TopoFuse: Topology-Aware Tri-Planar Fusion for 3D Cryo-Electron Tomography Segmentation —
- PPTBench: Can Coding Agents Reconstruct the Visual World through Structured, Editable Slides —
- SALI: Shot-Aware Late Interaction for Cross-Shot Relation Matching in Text-to-Video Retrieval using Film-Grammar Knowledge —
- A General Framework for Budgeted Threshold Incentives on Request —
- A Multimodal Dataset for Survival Prediction in Resected Pancreatic Ductal Adenocarcinoma —
- The Gold in Bias: Maturing the AI Design Process through Verification —
- TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar) —
- C3M: Cross-Session Multimodal Memory Maintenance for Long-Horizon Tasks —
- TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening —
- AI-based detection of worsening heart failure from low-resolution telemonitoring data —
- JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places —
- WeatherDiagFlow: Evidence-Grounded Radar Nowcasting with Diagnostic Flow Refinement —
- Breaking the Environment Wall: A Unified Framework for Preparing and Evolving Agent-Native Environments —
- An Analytical Theory of Auxiliary Learning —
- Mind the Gap: Mesh-Guided Repair of Broken Vessels —
- Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons —
- Lightweight Vision Transformer-Based U-Net for Brain Tumor Segmentation from MRI —
- OREO: Fidelity Alignment in 3D Generation via On-the-fly Rendering-Editing Optimization —
- TimeBraid: Unifying Time Series and Language for Understanding and Forecasting —
- Benchmarking and Domain Adaptation of Automatic Speech Recognition (ASR) for Adolescent Health Communication in Ghanaian Languages —
- Adaptive Fisher-Whitened Cross-Covariance for Low-Resource Speech Recognition —
- Learning to Ideate for Scientific Impact —
- CORDIAL: Calibrating Ordinal LLM Outputs from Few Labels —
- FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates —
- S2Planner: Multi-Scale Semantic Planner for End-to-End Autonomous Driving —
- SwitchPFN: Shared Switching Dynamics for Frozen In-Context Time Series Classification —
- AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation —
- Decoding Imagined Speech: A Strictly Subject-Independent Approach Using EEG —
- Anatomy-Aligned Surface Field Learning for Myocardial Reconstruction from Sparse Short-Axis Cine MRI —
- ChunkRank: Model-Aware Text Chunking and Abstention-Aware Answer Selection for LLM Pipelines —
- Retrieve-to-Localize: Bridging Large Language Models and LiDAR Geometry for Spatial Grounding —
- SplatLabel: Pseudo-Labelling through 4D Gaussian Splatting —
- PUBG Ally: A Conversational Embodied Agent as an AI Teammate —
- Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs —
- Encoded but Not Decoded: Layer-Localized Evidence for a Three-Level Gap in LLM Syntax —
- Multi-Task Learning by using Contextualized Word Representations for Syntactic Parsing of a Morphologically Rich Language —
- Modelling dynamic systems transfer functions from events in computational neuromorphic imaging —
- Efficient Continuous DEM Reconstruction under Limited Target-Resolution Supervision —
- A Risk-Adaptive and Evidence-Constrained Framework for Generative AI Feedback in Programming Education —
- When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression —
- Ontology-Mediated Neurosymbolic Constraint Acquisition from Multiple Stakeholders —
- From Graphs to Feeders: Constraint-Guided Diffusion for Rule-Compliant Feeder Generation —
- Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents —
- Spatio-temporally complementary feature propagation on graphs for longitudinal AADT estimation —
- MILO: Efficient Many-shot In-Context Learning with Block-wise Low-rank Compression —
- Who Holds the Pen? Let Specifications, Not Agents, Sign Off —
- Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations —
- EndoFSA: Endoscopic Few-Shot Image Generation via Rank-Constrained Parameter Adaptation —
- Improving Calibration of Black-Box Radiology AI Using Test-Time Augmentation —
- An Empirical Study of VLM Pipelines for Long-Document QA —
- Beyond Spatial Benchmarks: From Spatial Reasoning to Navigation —
- Robust Detection of LLM-Generated Text under Contamination —
- When Temporal Perturbations Act Like Sensor Biases: Label-Free Auditing of Wearable Activity Recognizers —
- Path-specific harm decomposition: A partial identification framework —
- Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration —
- MF-SCBO: Multi-fidelity Scalable Constrained Bayesian Optimization —
- Error- and Prediction-Driven Motor Learning in the Cortico-Cerebellar Loop —
- Neuro-symbolic AI for Industrial Configuration —
- ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation —
- Tracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking —
- Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes —
- Not All Confusion Is Equal: A Source-Aware Uncertainty Diagnosis for Fine-Grained Aircraft Detection —
- Beyond Average Safety: Chance-Constrained LLM Fine-tuning —
- A Contraction Framework for Stochastic Operators with Bootstrapping: Application to TD Learning —
- ADATEX4D: adaptive texture capacity allocation for 4D gaussian splatting —
- Diverse Geometries, Frozen Weights: Robust Heterogeneous Treatment-Effect Estimation via Causal Expert Ensembles —
- OceanXL: Large-scale Underwater 3D Gaussian Splatting via Block Partitioning and Adaptive Pruning —
- Let Training Guide Selection: Online Synthetic Data Filtering via Real-Anchored Utility —
- GHOST-Q: Towards Studying Grounding Hallucinations Overlooked Under Same-score TradeOffs in Quantized VLMS —
- Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems —
- VietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching —
- Automated Regulatory Compliance Question Answering in Financial Services with Domain-Adapted Retrieval-Augmented Generation —
- Low-Cost Assays for Measuring Model Behavior Across Vendors and Releases —
- Canopy: Exploiting Piecewise Smooth Tree Priors for Multi-Fidelity Bandits —
- Training-Free Hold-Usage Detection in Sport Climbing with Foundation Pose Models —
- Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark —
- How does Adversarial Influence Scale in Multi-Agent Systems? —
- Artificial Societies Benchmark: A Validation Framework for Synthetic Research —
- Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think —
- AERIAL: Adversarial Evaluation of Robustness in Accuracy-Preserving Low-Precision EEG Decoders —
- ConPro: Contrast Projection Pretraining for Label-Efficient Vessel Segmentation in DSA Sequences —
- Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models —
- NNV3: Expanding Neural Network Verification to New Architectures and Domains —
- SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback —
- Self-Play Pretraining with Zero Data —
- Scoring Both Directions: LLMs realize the MRS they cannot reliably parse —
- How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure —
- A Native-Reference Phone-Class Geometry for Second-Language Pronunciation Analysis —
- Nuclear Norm-Regularized Bayesian Matrix Completion —
- Reachability-Based Formal Verification of Graph Neural Networks with Node and Edge Features —
- Can Frozen Hyperspherical Features Guide the Selection of Pseudo Masks? —
- Residual Correlation as a Diagnostic for Joint-Uncertainty Gains from GP Coregionalisation —
- Return or Revise? Learning When Revision Helps Retrieval-Augmented QA —
- AT-SKM-Net: An Accelerated Trainable Sampling Kaczmarz-Motzkin Framework for Linear Hard-Constraint Feasibility on Dynamic Graphs —
- PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations —
- Accelerating Video Diffusion via Training-Free Trajectory Routing —
- R-DEIM Net: An Efficient Rationale-Augmented Dual-Expert Interaction Model for Paraphrase Detection —
- On the SoS Certifiability of Log-Concave Distributions —
- Smartphone-Based Method for Automated Speed Enforcement —
- What, When, and How: Audio Description as Constrained Global Optimization —
- HEXIS: Compiling Agent Skills into Extended Finite State Machines —
- Multimodal Thinking with Renderable Programs —
- Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale —
- EnigmaForge: The Question Is Hidden in the Story —
- GRASP: Generating, Revising, and Assessing for Strategic Planning with Agentic AI —
- Graph-Based Inference and Topology-Aware Multi-Agent Reinforcement Learning for Large-Scale Railway Network Management —
- Does a model's stated reason for rejecting a candidate do any work? —
- A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition —
- Do Audio Language Models Hear and Read Distinctive Features Alike? —
- Search-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search —
- ARGUS: Role-Aware Event Knowledge Graphs for U.S. Employment-Discrimination Complaints —
- Intrinsic-Extrinsic Coupling in Learning Dynamics —
- Jev-Mobile: Jev as an Executor for Mobile GUI Agents —
- Ego-Exo4D Human Meshes Dataset: 4D Human Motion Reconstruction for Ego-Exo Captures —
- SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance —
- Beyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers —
- ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds —
- A Living Benchmark for Information Retrieval from Electronic Health Records —
- The Alignment Illusion in Multimodal Large Language Models —
- Minimally Invasive Steering of Language Models —
- WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation —
- TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations —
- BiCC: Bidirectional Connected-Component Loss for Instance-Aware Segmentation —
- PoEM: Predicting RL Outcomes from Existing Policies —
- To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech —
- OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction —
- SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data —
- JevOut: Natural Context Can Flip Decision Models —
- Towards Practical Compression of 3D Gaussian Splatting —
- Agentic Detection of Online Conspiracies —
- Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning —
- CatSIM: A Categorical Image Similarity Metric —
- AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control —
Important terms
- Qwen-Planner-Agent
- This framework creates a closed-loop system for AI agents to plan and execute real-world mobile tasks autonomously. It allows agents to refine their plans iteratively based on what they observe in the environment.
- Tracking States vs. Tracking Cosets
- This is an algebraic method used for learned state tracking. It compares keeping a consistent state representation against tracking cosets, which affects how accurately a model predicts future system behaviors.
- Augur
- Augur is a synthetic decision lab designed to simulate reactions to changes in products or policies. It provides a controlled setting to test how learned policies respond to external shifts.
- Piecewise Smooth Tree Priors
- This involves using structural assumptions about the underlying function space, specifically piecewise smooth tree priors, when performing bandit decisions. This helps balance exploration and exploitation efficiently.
- Adversarial Influence Scaling
- This research looks at how the influence of an adversarial agent grows in multi-agent systems. It reveals complex patterns of influence that simple linear models cannot easily predict.