PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment
cs.AI
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: Accepted by EMNLP 2026
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging applicant volumes have made human-only evaluation costly and inconsistent, while existing AI
Terminology
Abstract
Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging applicant volumes have made human-only evaluation costly and inconsistent, while existing AI approaches yield opaque scores without traceable rationale. We introduce PhoenixNest-Video, an evidence-grounded multimodal agent framework for automated video interview assessment. It builds a semantic video graph as structured working memory, performs rubric-conditioned retrieval with cross-modal verification across visual, audio, and textual streams, and produces per-criterion scores anchored to the candidate's materials. A Scorer trained via Rubrics-based Reinforcement Learning with dual rewards for rubric alignment and score-level differentiation internalizes the discriminative structure of multi-level rubrics. PhoenixNest-Video attains 91.50% grade-level accuracy on VInterview-2025, outperforming substantially larger proprietary models. A compact, rubric-grounded agent therefore scores candidates in closer agreement with an expert panel than direct prompting of much larger models, and exposes the evidence behind each score for human review.
Sources
- Leveraging Multimodal Behavioral Analytics for Automated Job Interview Performance Assessment and Feedback
- Qwen3-VL Technical Report
- Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis
- MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
- Video-R1: Reinforcing Video Reasoning in MLLMs
- A Survey on LLM-as-a-Judge
- Seed1.5-VL Technical Report
- RecruitView: A Multimodal Dataset for Predicting Personality and Interview Performance for Human Resources Applications
- OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis
- How Order-Sensitive Are LLMs? OrderProbe for Deterministic Structural Reconstruction
- VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
- To Trust, or Not to Trust? A Study of Human Bias in Automated Video Interview Assessments
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
- EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model
- Behind the Screens: Uncovering Bias in AI-Driven Video Interview Assessments Using Counterfactuals
- Neptune: The Long Orbit to Benchmarking Long Video Understanding
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection