COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation
cs.LG, cs.AI
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/Quark-Medical/COPE
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Training Verifiers to Solve Math Word Problems
- The Llama 3 Herd of Models
- A Survey on Personalized Alignment -- The Missing Piece for Large Language Models in Real-World Applications
- SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation
- A Survey of Personalized Large Language Models: Progress and Future Directions
- Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- Measuring What Makes You Unique: Difference-Aware User Modeling for Enhancing LLM Personalization
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- Proximal Policy Optimization Algorithms
- Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
- Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
- Kimi K2: Open Agentic Intelligence
- Qwen3 Technical Report
- Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized Alignment
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks