LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection

arXiv:2608.09633 · cs.CV, cs.LG · Submitted 2026-08-13 · Read on arXiv

Peter Lorenz, Anjith George, Marcel Sébastien

Idiap Research Institute · University of Lausanne

cs.CV, cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/timesler/facenet-pytorch

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: This paper presents a comprehensive benchmark evaluating whether Low-Rank Adaptation (LoRA) of foundation models is sufficient for robust face presentation attack detection (PAD).

Terminology

Summary

This paper presents a comprehensive benchmark evaluating whether Low-Rank Adaptation (LoRA) of foundation models is sufficient for robust face presentation attack detection (PAD). The study systematically evaluates 32 pretrained vision encoders and 9 vision-language models (VLMs) under a unified experimental protocol across four MCIO benchmark datasets (MSU-MFSD, CASIA-FASD, Replay-Attack, and OULU-NPU).

Zero-shot VLM prompting performs near chance. The paper states: Among the nine evaluated VLMs, which span four model families and range from 2 billion to 32 billion total parameters, the average ACER remains between 35% and 50%, which is close to chance performance. The authors conclude: language-aligned semantic reasoning does not substitute for task-aligned adaptation of the vision tower.

LoRA achieves strong intra-dataset performance. The paper reports: CLIP ViT-B/32 reaches 0.3% mean ACER, LLaVA-NeXT-7B-ViT 0.4%, InternViT-6B 0.7%, and most adapted backbones remain below 2% across the four datasets while updating fewer than 1% of the foundation model parameters. This demonstrates that parameter-efficient adaptation is sufficient for strong intra-dataset PAD performance.

Cross-dataset transfer remains challenging. The paper states: cross-dataset ACER remains between 20% and 43% for all models, irrespective of pretraining regime, parameter count, or intra-dataset rank. Notably, Transfer performance does not correlate with intra-dataset rank. For example, InternViT-6B achieves 0.7 percent intra-dataset ACER but 39.8 percent cross-dataset ACER.

Pretraining objective matters more than scale. The paper finds: "The pretraining objective matters more than parameter count: CLIP ViT-B/32, with a 90M-parameter backbone, reaches 0.3% mean intra-dataset ACER, whereas the 0.60B-parameter vision tower extracted from the Qwen3-VL-32B VLM averages 5.0%. Furthermore, Contrastive vision–language backbones (CLIP, EVA-CLIP, SigLIP) and self-distillation models (DINOv3) respond most reliably to LoRA."

Pretraining data scale manifests under domain shift. The paper observes: every backbone pretrained solely on ImageNet-1K/22K (≤14M images) sits in the bottom third, every backbone in the top ten saw at least 100M web images, and beyond web scale more data stops helping. The authors conclude: web-scale coverage is necessary but not sufficient, and curation and objective dominate beyond the threshold.

Embedding structure remains domain-organized. The paper states: "Although the attention projections are adapted, the feature geometry is still primarily structured by acquisition domain rather than by presentation-attack class. Samples cluster according to their source dataset rather than the bonafide or attack label. This accounts for the persistent cross-dataset performance degradation."

The paper makes three main contributions: (1) A comprehensive and systematic evaluation of foundation models for face PAD, covering 32 vision encoders, 9 VLMs, specialist PAD baselines, and common datasets under a unified experimental protocol; (2) A bias free generalization score G that jointly captures intra- and cross-dataset skill as the geometric mean below chance; (3) Unveiling the limitation of LoRA adaptation for PAD, analyzing compute, feature-space analysis, and pretraining-scale comparisons besides accuracy.

The authors conclude: "LoRA adaptation substantially improves intra-dataset performance. Most adapted backbones achieve intra-dataset ACER below 2%, whereas zero-shot prompting of the same models remains near chance. Cross-dataset transfer remains challenging after adaptation: ACER ranges from 20–43% and is strongly influenced by the choice of training dataset. Moreover, increasing model scale alone does not consistently improve transferability."

The paper notes: The rather small CLIP ViT-B/32 and the much larger extracted CLIP-vision tower from the MLLM LLaVA-NeXT-7B achieve the strongest detection accuracy, whereas CLIP ViT-B/32 outperforms in overall compute.

The authors raise concerns about previous approaches: "These findings raise concerns regarding the effectiveness of previous approaches that utilize additional data, such as leave-one-out (LOO) or auxiliary datasets, to enhance generalization. Focusing on other post-training methods, including additional fine-tuning with domain-invariant losses or alignment with target-domain data, may mitigate cross-dataset issues."

The paper acknowledges: "Our study is limited to the MCIO benchmark, which primarily covers print and replay attacks and does not include deepfake-based attacks. Cross-dataset generalization is evaluated under a single-source protocol, and lightweight adaptation is restricted to LoRA rather than a broader comparison of parameter-efficient fine-tuning methods. Moreover, zero-shot evaluation uses a single prompt formulation, leaving prompt sensitivity unexplored."

Improvements for AI systems

Improvements to AI Systems Based on This Paper:

  1. Task-Aligned Adaptation Over Semantic Reasoning for Security Tasks
  • Improvement: Replace zero-shot VLM prompting with LoRA fine-tuning of the vision tower for any security-critical classification (e.g., face PAD, deepfake detection, biometric spoofing).

  • Capability: The improved system achieves near-perfect intra-domain accuracy (e.g., 0.3% ACER) instead of chance-level performance (35–50% ACER), while updating <1% of parameters—making it deployable on edge devices with limited compute.

  1. Pretraining Objective-Aware Model Selection
  • Improvement: Prioritize contrastive vision-language (CLIP, SigLIP) or self-distillation (DINOv3) backbones over scale when selecting a foundation model for adaptation.

  • Capability: The system can achieve state-of-the-art results with a 90M-parameter backbone (CLIP ViT-B/32) rather than a 0.6B-parameter tower from a 32B VLM (Qwen3-VL), reducing inference cost by 7x while improving accuracy by 4.7% ACER.

  1. Domain-Invariant Post-Training for Cross-Dataset Generalization
  • Improvement: Augment LoRA with domain-invariant losses (e.g., adversarial domain confusion, gradient reversal) or explicit target-domain alignment during fine-tuning, rather than relying on additional source data.

  • Capability: The system reduces cross-dataset ACER from the current 20–43% range toward single-digit percentages, enabling reliable deployment across unseen acquisition environments, cameras, and lighting conditions without retraining.

  1. Compute-Aware Backbone Selection for Real-Time Deployment
  • Improvement: Use the paper’s bias-free generalization score (G) to jointly optimize for accuracy and compute, selecting the smallest backbone that meets a target G threshold.

  • Capability: The system can automatically choose between CLIP ViT-B/32 (fast, accurate) and LLaVA-NeXT-7B (slightly better but 100x slower) based on hardware constraints, ensuring real-time PAD on mobile devices or embedded cameras.

  1. Feature-Space Re-Organization for Attack-Class Discrimination
  • Improvement: After LoRA, apply a secondary lightweight projection head or contrastive loss that explicitly separates bonafide vs. attack clusters in the embedding space, overriding the observed domain-organized geometry.

  • Capability: The system’s feature representations become attack-class-discriminative rather than dataset-discriminative, directly addressing the root cause of cross-dataset failure and improving interpretability of failure modes.

  1. Web-Scale Pretraining Threshold Detection
  • Improvement: For new PAD tasks, use the paper’s finding (≥100M web images needed for top-tier transfer) as a screening criterion when choosing or designing a pretrained backbone.

  • Capability: The system avoids wasting compute on ImageNet-only backbones (which consistently underperform) and instead selects or fine-tunes backbones with sufficient web-scale coverage, improving worst-case robustness under domain shift.

  1. Hybrid LoRA + Prompt Tuning for Multi-Modal Systems
  • Improvement: For VLMs used in PAD, combine LoRA on the vision tower with a small set of trainable prompt tokens (e.g., 10–20) for the language tower, rather than freezing both or using static prompts.

  • Capability: The system retains some semantic reasoning (e.g., explaining why an attack is detected) while achieving the task-aligned accuracy of pure vision adaptation, enabling human-interpretable security alerts without sacrificing performance.

  1. Benchmark-Driven Architecture Search
  • Improvement: Use the paper’s unified protocol (32 encoders, 9 VLMs, 4 datasets) as a standard evaluation harness for any new PAD model, reporting both intra- and cross-dataset ACER plus the G score.

  • Capability: The system provides a reproducible, bias-free comparison across future models, allowing researchers to quickly identify whether a new architecture genuinely improves generalization or merely overfits to a single dataset.

Sources

Related papers