When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization

arXiv:2609.06646 · cs.CV, cs.AI · Submitted 2026-09-06 · Read on arXiv

cs.CV, cs.AI

Submitted: 2026-09-06

Updated: 2026-09-09

Comments: Accepted to the Workshop on Affective & Behavior Analysis in-the-wild, ECCV 2026

Code: https://github.com/WSCSports/MTLLFM-temporal-laughter-localization

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: Annotators routinely disagree on laughter boundaries and subtle chuckles, yet temporal laughter localization typically evaluates against a single reference annotation.

Terminology

Abstract

Annotators routinely disagree on laughter boundaries and subtle chuckles, yet temporal laughter localization typically evaluates against a single reference annotation. We show that this disagreement is structured rather than random noise. Re-annotating the SMILE-Temporal benchmark (672 videos, 1,683 events) with 3-5 annotators per video (alpha = 0.757), we find systematic patterns: disagreement is 1.73 times larger at offsets than onsets, far more common for chuckles than full laughs (77% vs. 20%), and predictable from event attributes (AUC = 0.831). Evaluating against a single annotator breaks down under this structure: system scores shift by 0.246 F1 depending on the chosen ground truth, correctly ranking systems only 69.7% of the time (vs. 80% against all annotators). We propose a disagreement-calibrated evaluation that scores predictions against the full annotator distribution using conformally calibrated tolerance bands (wider at offsets, 0.727s, than onsets, 0.5s). The per-annotator annotations and analysis code are available at https://github.com/WSCSports/MTLLFM-temporal-laughter-localization.

Related papers