Cross-Domain Identity Representation for Skull to Face Matching with Benchmark DataSet

arXiv:2507.08329 · cs.CV · Submitted 2025-07-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Cross-Domain Identity Representation for Skull to Face Matching with Benchmark DataSet".

Tom: Craniofacial reconstruction in forensic science is crucial for identifying victims of crimes and disasters by mapping a given skull to its corresponding face using deep learning advancements.

Jane: First, who's behind it and why it matters.

Paper summary: Jane: To wrap up, this paper, "Cross-Domain Identity Representation for Skull to Face Matching with Benchmark DataSet," introduced a framework that uses Siamese networks for cross-domain identity representation and created the IITMandi S2F benchmark dataset.

Tom: The implication is that they've provided researchers with a usable tool to start exploring craniofacial recognition and reconstruction by giving them this dataset.

Lu: What we're really seeing is the development of a novel framework for learning cross-domain identity representation specifically using Siamese networks, which connects these two domains in a way that was previously hard to achieve.

Meng: From an engineering standpoint, the real value is that they've given us something tangible—the benchmark dataset—which other researchers can actually use for other related work.

Jane: So, this research sets up a path for future studies on craniofacial superimposition and reconstruction by giving them the tools to test their theories with real data.

Conclusion: Tom: So we're wrapping up on this one and I gotta say, "Cross-Domain Identity Representation for Skull to Face Matching with Benchmark DataSet." It sounds a lot more technical than just a simple face matching tool, right?

Jane: Yeah, it tackles that skull-to-face problem directly, but the authors are really focused on how they build this representation across different image types.

Lu: What's interesting is the method itself; they use Siamese networks to learn a feature space where similar things are close and dissimilar things are far apart, even when those inputs come from totally different domains like X-rays and photos.

Meng: From an engineering standpoint, that cross-domain part is huge. Most models get stuck when you switch from one kind of image data to another, but they're trying to bridge that gap here.

Lalam: If we think about what this means for culture—for how we handle things like identity in digital systems—it shows AI can learn really deep, structural similarities between completely different visual information.

Tom: Exactly! It means the system isn't just looking at pixels anymore; it's learning a meaningful concept of identity that works regardless of whether the input is a skull scan or a photograph.

Jane: And they created this benchmark dataset, IITMandi S2F, which is super important because it gives everyone else something concrete to test their ideas against.

Lu: That dataset includes X-rays paired with actual face images from volunteers who are from different regions, which adds a lot of necessary diversity to the training data.

Meng: It’s practical because it means other researchers can actually run their models on this specific setup and see how they perform on real, messy data.

Lalam: This opens up avenues for more robust systems that can handle complex forensic tasks without needing perfectly paired data for every single case.

Tom: So, the authors essentially gave us a solid foundation and a testing ground for moving craniofacial recognition past just simple matching toward something more generalizable.

Ravi Shankar Prasada, Dinesh Singha

Indian Institute of Technology Mandi

cs.CV

Submitted: 2025-07-11

Updated: 2025-07-11

Comments: 7 pages, 12 figures, Pattern Recognition Letters

Journal ref: Journal of Visual Communication and Image Representation, 120: 104890 (2026)

DOI: 10.1016/j.jvcir.2026.104890

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 77/100

The gist: Craniofacial reconstruction in forensic science is crucial for identifying victims of crimes and disasters by mapping a given skull to its corresponding face using deep learning advancements.

Key concepts

Siamese Networks
These are twin neural networks with the same structure used to compare inputs. They are trained so that similar images (like a skull and its corresponding face) have close feature representations in the learned space, while dissimilar ones are pushed far apart.
Cross-Domain Identity Representation
This technique aims to create a shared feature space where images from different domains—skull X-rays and facial photos—can be compared effectively. The framework uses a common backbone network to learn features that are meaningful regardless of whether the input is a skull or a face.
Triplet Loss
This loss function guides the training process by comparing three images: an anchor (skull), a positive match (its corresponding face), and one or more negatives. The goal is to ensure the distance between the anchor and its positive pair is smaller than the distance between the anchor and any negative pair.
IITMandi_S2F Dataset
This is a custom benchmark dataset consisting of X-ray images of skulls paired with frontal and side face images from 40 volunteers. It was augmented with various transformations like rotation and color jitter to make the model robust for real-world forensic applications.

Terminology

Summary

Craniofacial reconstruction in forensic science is crucial for identifying victims of crimes and disasters by mapping a given skull to its corresponding face using deep learning advancements. This paper presents a framework that utilizes convolutional Siamese networks for cross-domain identity representation to achieve this skull-to-face matching task, addressing the difficulty of obtaining paired data through the creation of a benchmark dataset.

How it works

The core methodology involves employing Siamese networks, which are twin networks that share the same architecture and can be trained to discover a feature space where nearby observations that are similar are grouped and dissimilar observations are moved apart <ref:2507.08329#pg2> (Page 1). To facilitate cross-domain identity representation, the framework utilizes a backbone neural network where Convolutional Neural Networks (CNNs) serve as the backbone of our proposed Siamese networks, which share a common backbone <ref:2507.08329#pg4> (Page 3). Specifically, for face images, we utilize pretrained deep models with frozen layers, while for skull images, we train the same pretrained model with adjustable weights <ref:2507.08329#pg5> (Page 4).

Dataset Preparation and Data Augmentation

A critical component of this research is the creation of a benchmark dataset, IITMandi S2F, which includes an X-ray image and its corresponding frontal and side face images <ref:2507.08329#pg2> (Page 2). This dataset was prepared from 40 volunteers whose front and side skull X-ray images and optical face images were collected <ref:2507.08329#pg2> (Page 3). The data collection involved X-ray scans with their respective face pair image from volunteers who mostly came from the North and East regions of India, aged 21-30 years females and males <ref:2507.08329#pg5> (Page 3). For each volunteer, 4 (= 2 × 2) anchor-positive pairs are created, while each anchor-positive pair can have 78 (= 39 × 2) negatives <ref:2507.08329#pg5> (Page 3). Furthermore, five types of data augmentation, like rotation, flipping, color jitter, random affine and brightness level changes of the skull and face images are applied to obtain a robust and generalized model <ref:2507.08329#pg6> (Page 5).

Cross-Domain Identity Representation via Triplet Loss

The Siamese networks are trained using a triplet loss function, defined as:

Lt = 1/N Σ Σ i=1 max(0, d(f(x a i)∗, f(x p i))−d(f(x a i)∗, f(x n i))+α) <ref:2507.08329#pg5> (Page 4). In this loss function, the anchor image represents the skull image of an individual, and the positive and negative images are used to establish similarity or dissimilarity <ref:2507.08329#pg5> (Page 4). The objective is to push feature vectors away from input pairs that are labeled as dissimilar and to bring the output feature vectors closer to input pairings that are labeled as similar <ref:2507.08329#pg5> (Page 4).

Experimental Evaluation and Model Performance

The performance of the trained model is evaluated on the benchmark dataset for skull-toface identification <ref:2507.08329#pg2> (Page 1). The system computes a similarity score using a metric based on Euclidean distance, and the confidence score is computed as: p,n = e − δ(f(x a i), f(x p,n i)) where δ(f(x a i), f(x(p,n) i)) = f(x a i) − f(x(p,n) i)2 <ref:2507.08329#pg6> (Page 5). The evaluation compares various backbone networks, with ResNet18, ResNet50, and MobileNet v2 showing better validation accuracy <ref:2507.08329#pg7> (Page 4). For retrieval tasks, the model's capability is assessed by retrieving the top k potential faces based on their confidence score using a mixed gallery of 485 face images <ref:2507.08329#pg7> (Page 6).

Conclusions

The research successfully introduced a framework based on Siamese networks for skull-to-face matching through cross-domain identity representation, and the novel benchmark dataset aims to kickstart the research on craniofacial recognition and reconstruction <ref:2507.08329#pg2> (Page 2). The overall contribution of this work includes preparing the benchmark dataset IITMandi SF, proposing a novel framework for learning cross-domain identity representation using Siamese networks, and providing a performance evaluation of the deep neural networks, such as Siamese networks on this dataset for matching a skull to a corresponding face <ref:2507.08329#pg2> (Page 2). The prepared benchmark dataset can also be used for other related research, such as craniofacial superimposition and craniofacial reconstruction <ref:2507.08329#pg2> (Page 4).

REFERENCES

[1] Bachman, P., Hjelm, R.D., Buchwalter, W., 2019. Learning representations by maximizing mutual information across views. Advances in neural information processing systems 32.

[2] Berar, M., Tilotta, F.M., Glaunès, J.A., Rozenholc, Y., 2011. Craniofacial reconstruction as a prediction problem using a latent root regression model. Forensic science international 210, 228–236.

[3] Claes, P., Vandermeulen, D., De Greef, S., Willems, G., Clement, J.G., Suetens, P., 2010a. Bayesian estimation of optimal craniofacial reconstructions. Forensic science international 201, 146–152.

[4] Claes, P., Vandermeulen, D., De Greef, S., Willems, G., Clement, J.G., Suetens, P., 2010b. Computerized craniofacial reconstruction: conceptual framework and review. Forensic science international 201, 138–145.

[5] Damas, S., Cordón, O., Ibáñez, O., Damas, S., Cordón, O., Ibáñez, O., 2020. Relationships between the skull and the face for forensic craniofacial superimposition. Handbook on craniofacial superimposition: The MEPROCS project, 11–50.

[6] Damas, S., Cordon, O., Ibanez, O., Santamaria, J., Alemán, I., Botella, M., Navarro, F., 2011. Forensic identification by computer-aided craniofacial superimposition: a survey. ACM Computing Surveys (CSUR) 43, 1–27.

[7] Duan, F., Yang, S., Huang, D., Hu, Y., Wu, Z., Zhou, M., 2014. Craniofacial reconstruction based on multi-linear subspace analysis. Multimedia Tools and Applications 73, 809–823.

[8] Hadsell, R., Chopra, S., LeCun, Y., 2006. Dimensionality reduction by learning an invariant mapping, in: 2006 IEEE computer society conference on computer vision and pattern recognition (CVPR’06), IEEE. pp. 1735–1742.

[9] Hu, Y., Duan, F., Yin, B., Zhou, M., Sun, Y., Wu, Z., Geng, G., 2013. A hierarchical dense deformable model for 3d face reconstruction from skull. Multimedia tools and applications 64, 345–364.

[10] Huang, J., Zhou, M., Duan, F., Deng, Q., Wu, Z., Tian, Y., 2011.

Improvements for AI systems

  1. What improved AI system can do: This system will perform robust, cross-domain identification of individuals by mapping a skull X-ray to corresponding optical face images with high accuracy, overcoming limitations of subjective manual reconstruction. The framework utilizes cross-domain identity representation using Siamese networks and triplet loss function to learn a feature space where nearby observations that are similar are grouped and dissimilar observations are moved apart.

  2. What improved AI system can do: The system will offer automated forensic identification by computing a similarity score between an unknown skull and faces in a gallery, as the text states, The confidence score is computed as follows Prasad and Singh: Preprint submitted to Elsevier Page 5 of 7 we computed the Euclidean distance and confidence scores with the feature vector of a given skull.

  3. What improved AI system can do: The model will achieve superior retrieval performance across various deep learning backbones, as evidenced by the comparison showing models like ResNet50 and MobileNet V2 are showing better validation accuracy, enabling faster and more reliable identification in real-world scenarios.

Related papers