Thesis Proposal: Toward a Human-Centered and Perspective-Aware Framework for Reproducible ML Evaluation and AI Alignment
cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/google-research/vet
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Whose Opinions Matter? Perspective-aware Models to Identify Opinions of Hate Speech Victims in Abusive Language Detection
- Training Verifiers to Solve Math Word Problems
- RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
- D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation
- s1: Simple test-time scaling
- Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models
- Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' Disagreement
- Self-critiquing models for assisting human evaluators
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Reward Is Enough: LLMs Are In-Context Reinforcement Learners
- A Roadmap to Pluralistic Alignment
- Large Language Models as Optimizers
- TextGrad: Automatic "Differentiation" via Text
- Unlocking Recursive Thinking of LLMs: Alignment via Refinement
- Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
- EvoAgentX: An Automated Framework for Evolving Agentic Workflows
- Large Language Models Are Human-Level Prompt Engineers
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering