Multi-Dimensional Comparative Scale Construction for Efficient Personalized Subjective Judgment in High-Traffic Applications
cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- Advancing Harmful Content Detection in Organizational Research: Integrating Large Language Models with Elo Rating System
- Training Deep Nets with Sublinear Memory Cost
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- LoRA: Low-Rank Adaptation of Large Language Models
- GPT-4o System Card
- Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models
- SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
- PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
- Decoupled Weight Decay Regularization
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models
- AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection