SERUM: State Extraction and Refinement for User Modeling
cs.LG, cs.AI, cs.CV
Submitted: 2026-07-31
Updated: 2026-08-29
Code: https://github.com/minnesotanlp/SERUM
Project page: https://minnesotanlp.github.io/SERUM-web
Terminology
Sources
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen3-VL Technical Report
- ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
- AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators
- Intention-Conditioned Long-Term Human Egocentric Action Forecasting
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
- Creating General User Models from Computer Use
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks