Towards Unified Dynamic Face Landmark Detection
University of Toronto · ModiFace
cs.CV, cs.AI
Submitted: 2026-08-11
Updated: 2026-09-24
Comments: 9 pages, 6 figures in Main Paper. 13 pages, 3 figures in Appendix
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 95/100
The gist: This paper introduces Unified Dynamic Face Landmark Detection, a novel framework that addresses two major functional limitations in existing face landmark detection (FLD) methods: (1) different
Terminology
Summary
This paper introduces Unified Dynamic Face Landmark Detection, a novel framework that addresses two major functional limitations in existing face landmark detection (FLD) methods: (1) different network parameters need to be trained independently for each N-point
benchmark dataset, and (2) a model trained on an N-point
dataset reliably outputs only the N landmarks. The proposed method enables a single model to learn on any number of N-point
datasets and yield any number of specific landmark predictions by loading designated landmark queries at runtime.
The paper makes four main contributions:
-
Face Part-Anchored Landmark Positions (FPALPs): An intuitive representation of face landmarks that are evenly distributed on well-defined face part curves. The FPALP format is universal and allows for compatibility with all existing and future datasets.
-
Unified FLD: The first work to enable, without auxiliary dataset information, the ability of a model to be trained end-to-end on the fusion of multiple
N-point
datasets. -
Dynamic FLD: A novel FPALP-based landmark queried regressor enabling unlimited on-demand landmark prediction without network retraining.
-
Competitive performance: The method matches or outperforms existing state-of-the-art methods on several benchmark datasets.
The paper observes that facial landmark annotations across different datasets (AFLW with 19/21 points, 300W with 68 points, WFLW with 98 points) are semantically anchored to face parts such as eyes, lips, nose, etc., and are often defined to be evenly spaced along a face part boundary.
Based on this, each landmark is associated with one or more containing face parts and represented as a progression value between 0 and 1 denoting its fractional position within the containing face part curve.
Formally, for a landmark l positioned at posl,p within a sequence of Np landmarks composing face part p, the FPALP is defined as: FPALPl,p = posl,p/(Np − 1).
The framework creates a unified face template by taking the union of face templates from all datasets: TU = TD1 ∪ TD2 ∪... ∪ TDD. The alignment of the face templates of AFLW19, 300W, and WFLW resulted in tight proximal clusters having an average intra-cluster distance of 2.22 pixels averaged over all face parts.
To achieve dynamic face landmark detection, target landmarks are represented as landmark queries. The initial image-agnostic representations are constructed by:
-
Encoding FPALPs using a simple MLP with ReLU activation
-
Inputting the face part name into a lightweight pretrained text encoder (SentenceBERT) to get its textual representation
-
Deriving image-agnostic landmark encodings as the summation of the encoded FPALPs and the face part textual representations
The paper states: We hypothesize that pretrained text encoders are more superior since they may already encode the semantics of facial layouts.
The framework conditions the facial image's visual features with the image-agnostic landmark encodings. Using a pretrained image encoder (FaRL pretrained ViT-B or ResNet), it derives an attention map A of the visual features with respect to the required image-agnostic landmark encodings: A = Softmax(EI · EIAT). The initial landmark queries and coordinate predictions are obtained by taking the weighted mean of the image features and grid center coordinates using the attention map.
A cross-modality transformer decoder with ndec decoder layers iteratively hones the landmark queries and predicted coordinates. Each layer executes:
-
Self-attention on landmark queries to exploit inter-landmark dependencies
-
Deformable attention for targeted cross-modality attention between image features and queries
-
Cross-attention between previous-step queries and image-agnostic landmark encodings
-
Feed-forward network to yield the decoder layer's query output
-
MLP to derive coordinate offsets
The model is supervised using Wing Loss on both intermediate and final coordinate predictions.
The framework is trained and evaluated on three benchmark datasets:
-
AFLW: 20,000 training and 4,386 test images with 19 landmarks
-
300W: 3,148 training and 689 test images with 68 landmarks
-
WFLW: 7,500 training and 2,500 test images with 98 landmarks
Cross-dataset evaluation is conducted on COFW (507 test images with 29 landmarks), COFW68, and WFLW68.
Without Dataset Adapters (LoRA modules injected for dataset-specific fine-tuning), the model achieves performance on par with prior state of the art while enabling joint training on multiple datasets and supporting dynamic landmark prediction. With Dataset Adapters, performance improves further and consistently surpasses prior methods across datasets.
Key results on the full datasets (with Dataset Adapters, ViT-B backbone):
-
WFLW: NMEio of 4.02, FR10 of 2.19
-
300W: NMEio of 2.43 (common), 4.19 (challenge)
-
AFLW-19: NMEdiag of 1.01
When trained only on 300W and evaluated on other datasets, the ViT-B model achieves:
-
300W: 3.01 NMEio
-
COFW68: 4.40 NMEio
-
WFLW68: 6.08 NMEio
The paper notes: Using the ViT-B backbone, we demonstrate robustness by significantly improving performance on the challenging WFLW68 dataset, which includes facial images with extreme poses, expressions, occlusions, and makeup.
Training datasets: Training on all datasets combined yields the best performance across most datasets. The results indicate that exposure to varied FPALPs enhances the model's ability to effectively represent face part curves and generalize across varied facial structures and ambient conditions.
Face part representation: Using SentenceBERT (pretrained text encoder) results in faster convergence and better performance compared to learnable embeddings. The paper states: The usage of SentenceBERT to represent face parts results in a faster convergence and a more performant model.
Image encoder: ViT-B proves superior on most evaluation datasets, but both ResNets perform competitively, at a fraction of the size of ViT-B. This suggests that the availability of diverse 'N-point' training datasets is of higher importance than the capacity of the image encoder.
Number of decoder blocks: Performance gains are rapid when increasing from 1 to 3 blocks, followed by degradation starting at 4 blocks, suggesting overfitting. The final configuration uses ndec = 3.
The paper demonstrates qualitative results showing the system can predict landmarks at various granularities within and across face parts. The paper observes: Landmark predictions for face parts with higher FPALP diversity, such as the face contour and the eyes, are more accurate than those with lower FPALP diversity, such as the nose boundary.
The paper includes a quantitative evaluation of dynamically queried landmarks through a controlled held-out experiment. When comparing direct querying at unseen FPALPs versus cubic-spline interpolation:
-
300W (37 retained, 31 held out): 19.0% improvement over interpolation
-
WFLW (73 retained, 23 held out): 16.3% improvement
-
WFLW (50 retained, 46 held out): 15.7% improvement
The paper concludes: These results indicate that the proposed formulation learns more than a geometric interpolation rule between neighboring landmarks.
The paper acknowledges several limitations:
-
Imprecise alignment:
The likelihood of imprecise alignment of the individual datasets' face templates during construction of the unified face template, which may hinder scalability of the FPALP formulation.
-
Interpolation mismatch:
Landmarks generated through interpolation techniques may not align with those predicted via evenly spaced FPALPs.
-
Language bias:
The choice of the training dataset and the text encoder could introduce biases or limitations when face part phrases are described using low-resource languages.
-
COFW exclusion: COFW was excluded from training
due to observed inconsistencies in annotation quality.
The paper concludes: "In this paper, we present our Unified Dynamic Face Landmark Detection method, wherein landmarks are treated as progression points on user-defined face parts, allowing for end-to-end model training on the fusion of diverse 'N-point' datasets and execution of unlimited on-demand landmark predictions. With performance competitive with, if not surpassing, SOTA methods, our simple yet adaptable framework is positioned to meet the requirements of various downstream applications that depend on a wide range of precise face landmarks."
Improvements for AI systems
Improvements to AI systems:
-
Unified multi-dataset training without architectural changes – The FPALP representation allows a single model to train on datasets with different landmark counts (19, 68, 98 points) simultaneously, eliminating the need for separate models per dataset. This improves data efficiency and generalization across annotation schemas.
-
On-demand landmark prediction at arbitrary granularity – By encoding landmarks as continuous progression values (0–1) along face part curves, the system can output any number of landmarks (e.g., 5, 50, 500) at runtime without retraining. This enables adaptive precision for downstream tasks (coarse tracking vs. fine detail analysis).
-
Semantic grounding via text embeddings – Using a pretrained text encoder (SentenceBERT) for face part names injects prior semantic knowledge about facial anatomy into the model, leading to faster convergence and better performance than learnable embeddings, especially for unseen or rare face part combinations.
-
Cross-dataset zero-shot transfer – The unified face template alignment (average intra-cluster distance of 2.22 pixels) enables the model trained on one dataset (e.g., 300W) to generalize to other datasets (COFW68, WFLW68) without fine-tuning, reducing annotation cost for new benchmarks.
-
Improved robustness to challenging conditions – Joint training on diverse datasets (AFLW, 300W, WFLW) with varied poses, expressions, occlusions, and makeup yields significant improvements on hard cases (e.g., WFLW68), making the system more reliable in real-world, uncontrolled environments.
-
Dynamic landmark interpolation superior to geometric methods – The learned querying mechanism outperforms cubic-spline interpolation by 15–19% on held-out landmarks, meaning the model captures actual facial structure rather than just curve geometry, enabling more accurate predictions at unseen positions.
-
Modular dataset adapters for performance boost – Optional LoRA-based adapters allow per-dataset fine-tuning without forgetting, achieving state-of-the-art results (e.g., 4.02 NME on WFLW) while retaining the unified model’s flexibility.
What the improved AI system can do:
-
A single face analysis system that works across all existing landmark benchmarks (19, 21, 29, 68, 98 points) and future datasets without retraining.
-
Real-time adaptive landmark density: switch from 5 landmarks for low-bandwidth mobile tracking to 500 landmarks for high-precision 3D reconstruction or medical analysis on the fly.
-
Zero-shot deployment on new datasets or annotation schemas, eliminating the need for dataset-specific model training.
-
Robust performance on occluded, extreme-pose, and heavily made-up faces, suitable for AR/VR, driver monitoring, and animation.
-
Semantic understanding of facial parts (e.g.,
upper lip
,left eye contour
) enabling natural language queries for landmark extraction (e.g.,give me 10 points along the jawline
). -
Superior interpolation of missing landmarks (e.g., from partial annotations) that respects anatomical structure, improving downstream tasks like face alignment, expression transfer, and 3D morphable model fitting.
Abstract
Although advancements in face landmark detection (FLD) methods continue to push performance boundaries, they overlook two major functional limitations: (1) different network parameters need to be trained independently for each `` N-point'' benchmark dataset, and (2) a model trained on an `` N-point'' dataset reliably outputs only the N landmarks. In our work, we first conceptualize Face Part-Anchored Landmark Positions (FPALPs), wherein each landmark is treated as a progression value between zero (start) and one (end) along a face part's contour. Every landmark can be expressed in the FPALP format, irrespective of its source dataset, hence unlocking the ability to unify all `` N-point'' datasets into a single dataset. Secondly, we represent each landmark with an FPALP-based query, refine it progressively with a cross-modality decoder, and predict its coordinates based on the final representation. Our approach, called Unified Dynamic FLD, embodies these two design choices and streamlines the landmark detection pipeline by enabling (1) a single model to learn on any number of `` N-point'' datasets, and (2) yield any number of specific landmark predictions by loading the designated landmark queries at runtime. Extensive experiments on multiple benchmark datasets show that our method delivers these benefits while remaining competitive with, and in several cases outperforming existing state-of-the-art methods.
Sources
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models