REVIEW 3 major objections 4 minor 39 references
Metric for Evaluating Performance of Reference-Free Demorphing Methods
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper proposes a new metric, biometrically cross-weighted IQA, that scores reference-free face demorphing by multiplying face-matcher similarity with image quality and selecting the correct output pairing.
desk verdict Correctly identifies a real evaluation gap in face demorphing, but the proposed metric does not penalize the trivial solution it claims to fix, and the benchmark protocol is murky. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the biometrically cross-weighted IQA score defined in Equation (9) of the paper. Given a morph $x$, outputs $o_1, o_2$, and ground truths $i_1, i_2$, the metric computes biometric similarity $B(o,i)$ with a face matcher and image quality $\text{iqa}(o,i)$ with an IQA function such as SSIM or PSNR, multiplies the two per candidate pair, and takes the maximum over the two possible bijections between outputs and ground truths. The max operation is what handles the fact that demorpher outputs are unordered, and the multiplicative weighting is what forces a high score to require both identity preservation and image fidelity.
What would settle it
Run a demorpher that outputs the same constituent image twice, say $o_1 = o_2 = i_1$, evaluate BW(SSIM) with any face matcher, and check whether its score is comparable to the genuine demorphers reported in Table 2; if it is, the metric does not penalize trivial outputs.
Extended reading notes
Core claim
The central claim is that a demorphing metric should be a product of two factors: the biometric match score between a demorphed output and a ground-truth constituent image, and the image quality of that same output measured by an IQA function. For unordered outputs, the metric computes both possible output-to-ground-truth pairings and keeps the larger weighted sum, then averages over all morphs. This single formula, $\text{BW}(\text{iqa}) = \mathbb{E}_{x \in X} \max\left(\sum_{i\in\{1,2\}} B(o_i,i_i)\cdot\text{iqa}(o_i,i_i), \sum_{i\in\{1,2\}, j=i\%2+1} B(o_i,i_j)\cdot\text{iqa}(o_i,i_j)\right)$, is claimed to overcome the morph-replication loophole that makes TMR and RA uninformative and to correct the identity blindness of SSIM and PSNR. The paper's experiments support this by showing that a trivial demorpher reaches 100% TMR while the proposed metric separates the three methods in a way consistent with visual quality.
Load-bearing premise
The paper assumes that multiplying biometric match scores by image-quality scores makes the trivial demorpher score poorly, but the formula contains no term requiring the two outputs to be distinct, so a demorpher that returns the same image twice could still receive a high score.
Editorial extensions
If this is right
- If the metric is adopted, demorphing papers will need to report a single number that can be compared across methods, datasets, and face matchers.
- TMR and RA should be dropped as primary evidence, because a trivial demorpher that regurgitates the morph achieves a perfect score by construction.
- Methods trained under different assumptions can be ranked under a common protocol using the same score, as the paper does for three open-source demorphing methods.
- The metric extends straightforwardly to any face matcher and any IQA measure, so future demorphing methods can be tuned directly against it.
- The benchmark result, with Identity Preserving Decomposition scoring highest on BW, gives a concrete reference point for subsequent work.
Reading between the lines
- Editorial inference: the metric inherits the blind spots of both components; if the face matcher cannot distinguish two identities, the biometric weight will not penalize identity confusion, and if the IQA is insensitive to artifacts, a distorted output can still score well.
- Editorial inference: the definition of BW(iqa) contains no term enforcing the paper's own Equation (3), the requirement that the two outputs $o_1$ and $o_2$ be dissimilar, so a demorpher that returns the same image for both outputs could still score high if that image matches one ground truth.
- Editorial inference: the Table 2 benchmark numbers are described as supplied by the original authors rather than recomputed under the paper's common protocol, so the reported ranking of methods under the new metric may partly reflect the original papers' own training conditions.
- Editorial inference: a natural testable extension is to add the distinctness penalty $B(o_1,o_2)<\theta$ directly into the metric, or to generalize the formula to morphs created from more than two identities by summing over all possible output-to-identity assignments.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses evaluation of reference-free face demorphing. It argues that existing metrics are flawed: TMR and RA can be gamed by trivial morph replication, while SSIM/PSNR ignore identity information. It proposes a new metric, biometrically cross-weighted IQA (BW), defined in Eq. (9), which multiplies biometric similarity scores with IQA values over possible pairings of outputs and ground truths. The authors benchmark three demorphing methods on six datasets with AdaFace and ArcFace, reporting TMR, RA, PSNR/SSIM, and BW scores, and conclude that BW better reflects visual quality.
Significance. The paper identifies a real problem: reference-free demorphing evaluation lacks a consensus metric, and Table 3 convincingly shows that a trivial solution achieves perfect TMR on all six datasets. If a metric truly penalized morph replication while combining biometric and image-quality information, it would be a useful contribution. The paper also usefully demonstrates that SSIM/PSNR can rank distorted outputs incorrectly. However, the proposed BW metric does not contain any explicit mechanism to penalize the trivial solution, and the experimental protocol is internally inconsistent as reported. As a result, the central claim is not supported by the presented evidence.
major comments (3)
- [§4.2, Eq. (9); §1 contribution 2] The paper claims that BW 'penalizes trivial demorphing outputs,' but Eq. (9) contains no term involving B(o1,o2) and does not enforce the distinctness condition in Eq. (3). If a method outputs o1=o2=x, the two sums inside max(·) become identical, and BW reduces to E_x[ B(x,i1)·iqa(x,i1)+B(x,i2)·iqa(x,i2) ]. Since a morph is constructed so that B(x,ik)>τ, these products can be substantial. The paper never reports the BW score of this trivial solution in Table 2 or elsewhere, so the central claim is neither guaranteed by construction nor verified empirically.
- [§5 vs. Table 2 caption] Section 5 says the methods are trained under a common protocol, but Table 2 states that the scores are 'supplied by the original authors.' If the three methods were evaluated under their originally published training protocols (scenarios 1 and 3 as described in §2), the columns are not comparable; if they were re-trained under the common protocol, the scores cannot be supplied by the original authors. Either way, the cross-method comparison that underpins the efficacy claims for BW is not supported by the reported numbers.
- [§4.1, Eq. (7)] The max over the two possible pairings in Eq. (7) and Eq. (9) means that one well-reconstructed output and one poor output can receive a score similar to two moderately reconstructed outputs. This is inconsistent with the problem statement in Eq. (4), which requires both outputs to align with their corresponding ground-truth images. Thus BW can be inflated by partially successful decompositions, and it does not measure whether the demorpher satisfies the two-output recovery condition.
minor comments (4)
- [Table 2 header] The header column labeled 'Facial Demorphing [33]' refers to [1] in the text; [33] is SDeMorph, so the citation is inconsistent.
- [§5, last sentence] The sentence 'Note that the test morphs include those generated using includes both conventional landmark-based techniques...' contains a typo: 'using includes' should be 'using both' or similar.
- [Table 3 caption] The caption says TMR is 'averaged over two subjects,' but the text in §7 indicates averaging over two face matchers; please clarify which quantity is averaged.
- [§4.1] The impostor score computation described in the Evaluation Criterion paragraph is not used in Eq. (7), Eq. (8), or Eq. (9); either it should be removed or its role in the proposed metric should be explained.
Circularity Check
No significant circularity: the proposed BW metric is an explicit combination of external biometric and IQA components, not fitted to data or derived from a self-citation chain.
full rationale
BW(iqa) in Eq. (9) is defined as the expectation over morphs of the max of two sums, each term being a product of a face-matcher similarity B(o_i, i_k) and an IQA score iqa(o_i, i_k). It is an evaluation formula composed from independently existing components (AdaFace/ArcFace, SSIM/PSNR); no parameter is fitted to the benchmark data and no value is chosen to reproduce a target ranking. The paper's criticism of TMR/RA and SSIM/PSNR is conceptual, and the proposed metric explicitly announces itself as a weighted combination of those two families, so its balancing behavior is built into the definition rather than being a derived prediction equivalent to an input. The authors' prior works ([1], [33], [35], [34]) are cited as benchmarked methods and background, not as the basis for the metric's correctness; the Table 2 note that scores were supplied by the original authors is an independence and reproducibility caveat, not a circular reduction. The known weakness that Eq. (9) has no B(o1,o2) distinctness term and therefore may fail to penalize a trivial solution is a substantive correctness concern, but it is not circularity: the metric does not reduce to its inputs by construction. The derivation chain is self-contained, proceeding from the proposed metric to evaluation on existing methods and comparison with existing metrics, with no fitted parameters or imported uniqueness claims.
Assumptions & free parameters
assumptions (3)
- domain assumption Biometric similarity scores from AdaFace and ArcFace are valid measures of identity match for demorphed images.
- domain assumption The product of biometric similarity and IQA provides a meaningful scalar for ranking demorphing quality.
- domain assumption The max over the two pairings correctly resolves the unordered output problem.
Cite this review
Pith. "Pith review of Metric for Evaluating Performance of Reference-Free Demorphing Methods." pith.science (2026). https://pith.science/paper/XPJAKWXB
@misc{pith2026250112319,
author = {Pith},
title = {Pith review of: Metric for Evaluating Performance of Reference-Free Demorphing Methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/XPJAKWXB}},
note = {Machine review of arXiv:2501.12319}
}
read the original abstract
A facial morph is an image created by combining two (or more) face images pertaining to two (or more) distinct identities. Reference-free face demorphing inverts the process and tries to recover the face images constituting a facial morph without using any other information. However, there is no consensus on the evaluation metrics to be used to evaluate and compare such demorphing techniques. In this paper, we first analyze the shortcomings of the demorphing metrics currently used in the literature. We then propose a new metric called biometrically cross-weighted IQA that overcomes these issues and extensively benchmark current methods on the proposed metric to show its efficacy. Experiments on three existing demorphing methods and six datasets on two commonly used face matchers validate the efficacy of our proposed metric.
Figures
Reference graph
Works this paper leans on
-
[1]
Facial De-morphing: Extracting Component Faces from a Single Morph
Sudipta Banerjee, Prateek Jaiswal, and Arun Ross. Facial De-morphing: Extracting Component Faces from a Single Morph. In Proceedings of IEEE International Joint Confer- ence on Biometrics, 2022. 1, 2, 3, 4, 5
work page 2022
-
[2]
Naser Damer, Fadi Boutros, Alexandra Mosegu ´ı Saladi ´e, Florian Kirchbuchner, and Arjan Kuijper. Realistic Dreams: Cascaded Enhancement of GAN-generated Images with an Example in Face Morphing Attacks. In Proceedings of IEEE 10th International Conference on Biometrics Theory, Appli- cations and Systems (BTAS), pages 1–10, 2019. 1
work page 2019
-
[3]
Naser Damer, Meiling Fang, Patrick Siebke, Jan Niklas Kolf, Marco Huber, and Fadi Boutros. MorDIFF: Recognition Vulnerability and Attack Detectability of Face Morphing At- tacks Created by Diffusion Autoencoders. In Proceedings of 11th International Workshop on Biometrics and Forensics (IWBF), pages 1–6, 2023. 1, 5
work page 2023
-
[4]
Privacy-Friendly Synthetic Data for the Development of Face Morphing Attack Detectors
Naser Damer, C ´esar Augusto Fontanillo L ´opez, Meiling Fang, No ´emie Spiller, Minh Vu Pham, and Fadi Boutros. Privacy-Friendly Synthetic Data for the Development of Face Morphing Attack Detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , pages 1606–1617, June
-
[5]
ReGen- Morph: Visibly Realistic GAN Generated Face Morphing Attacks by Attack Re-generation
Naser Damer, Kiran Raja, Marius S ¨ußmilch, Sushma Venkatesh, Fadi Boutros, Meiling Fang, Florian Kirchbuch- ner, Raghavendra Ramachandra, and Arjan Kuijper. ReGen- Morph: Visibly Realistic GAN Generated Face Morphing Attacks by Attack Re-generation. In Proceedings of Ad- vances in Visual Computing, pages 251–264, 2021. 1
work page 2021
-
[6]
Naser Damer, Alexandra Mosegu ´ı Saladi ´e, Andreas Braun, and Arjan Kuijper. MorGAN: Recognition Vulnerability and Attack Detectability of Face Morphing Attacks Created by Generative Adversarial Network. In Proceedings of 2018 IEEE 9th International Conference on Biometrics Theory, Applications and Systems (BTAS), pages 1–10, 2018. 4
work page 2018
-
[7]
debruine/webmorph morphing software: Beta release 2, January 2018
Lisa DeBruine. debruine/webmorph morphing software: Beta release 2, January 2018. URL: https://doi.org/ 10.5281/zenodo.1162670. 5
-
[8]
Face Research Lab London (FRLL) Image Dataset
Lisa DeBruine and Benedict Jones. Face Research Lab London (FRLL) Image Dataset. May 2017. URL: https : / / figshare . com / articles / dataset / Face_Research_Lab_London_Set/5047666/3. 5
arXiv 2017
Show all 39 references
-
[9]
Arcface: Additive Angular Margin Loss for Deep Face Recognition
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive Angular Margin Loss for Deep Face Recognition. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 4690–4699, 2019. 5
2019
-
[10]
Face Demorphing
Matteo Ferrara, Annalisa Franco, and Davide Maltoni. Face Demorphing. IEEE Transactions on Information Forensics and Security, 13(4):1008–1017, 2018. 1
2018
-
[11]
De- coupling Texture Blending and Shape Warping in Face Mor- phing
Matteo Ferrara, Annalisa Franco, and Davide Maltoni. De- coupling Texture Blending and Shape Warping in Face Mor- phing. In Proceedings of International Conference of the Biometrics Special Interest Group (BIOSIG) , pages 1–5,
-
[12]
Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative Adversarial Nets. In Proceed- ings of the 27th International Conference on Neural Infor- mation Processing Systems, 2014. 1
2014
-
[13]
LADIMO: Face Morph Generation through Biometric Template Inversion with Latent Diffusion
Marcel Grimmer and Christoph Busch. LADIMO: Face Morph Generation through Biometric Template Inversion with Latent Diffusion. In Proceedings of International Joint Conference on Biometrics, 2024. 1
2024
-
[14]
Denoising Dif- fusion Probabilistic Models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising Dif- fusion Probabilistic Models. In Proceedings of Advances in Neural Information Processing Systems , volume 33, pages 6840–6851, 2020. 1
2020
-
[15]
Face Morphing Attack Detection with Denoising Diffusion Probabilistic Models
Marija Ivanovska and Vitomir ˇStruc. Face Morphing Attack Detection with Denoising Diffusion Probabilistic Models. In Proceedings of International Workshop on Biometrics and Forensics (IWBF), pages 1–6. IEEE, 2023. 1
2023
-
[16]
Analyzing and Improv- ing the Image Quality of StyleGAN
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and Improv- ing the Image Quality of StyleGAN. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 5
2020
-
[17]
Adaface: Quality Adaptive Margin for Face Recognition
Minchul Kim, Anil K Jain, and Xiaoming Liu. Adaface: Quality Adaptive Margin for Face Recognition. In Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 5
2022
-
[18]
Davis E. King. Dlib-ml: A Machine Learning Toolkit. vol- ume 10, page 1755–1758. The Journal of Machine Learning Research, December 2009. 5
2009
-
[19]
Face morph using opencv — c++ / python
Satya Mallick. Face morph using opencv — c++ / python. LearnOpenCV, 2016. URL: https://learnopencv. com/face-morph-using-opencv-cpp-python/ . 5
2016
-
[20]
Laws against morphing
Matthias Monroy. Laws against morphing. Security Archi- tectures and the Police Collaboration in the EU 10/01/2020,
2020
-
[21]
Extended StirTrace Benchmarking of Biometric and Forensic Qualities of Mor- phed Face Images
Tom Neubert, Andrey Makrushin, Mario Hildebrandt, Chris- tian Kraetzer, and Jana Dittmann. Extended StirTrace Benchmarking of Biometric and Forensic Qualities of Mor- phed Face Images. IET Biometrics, 7:325–332, 2018. 5
2018
-
[22]
Face Recognition Vendor Test (FRVT) Part 4: MORPH - Performance of Automated Face Morph Detection
Mei Ngan, Patrick Grother, Kayee Hanaoka, and Jason Kuo. Face Recognition Vendor Test (FRVT) Part 4: MORPH - Performance of Automated Face Morph Detection. NIST Interagency/Internal Report (NISTIR), National Institute of Standards and Technology, Gaithersburg, MD, 2020-03-06
2020
-
[23]
Face Morpher
A Quek. Face Morpher. URL: https://github.com/ alyssaq/facemorpher. 5
-
[24]
Raghavendra, Kiran B
R. Raghavendra, Kiran B. Raja, and Christoph Busch. De- tecting Morphed Face Images. In Proceedings of IEEE 8th International Conference on Biometrics Theory, Appli- cations and Systems (BTAS) , page 1–7. IEEE Press, 2016. 1
2016
-
[25]
Raghavendra, Kiran B
R. Raghavendra, Kiran B. Raja, Sushma Venkatesh, and Christoph Busch. Transferable Deep-CNN Features for De- tecting Digital and Print-Scanned Morphed Face Images. In Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 1822– ...
2017
-
[26]
Bommanna Raja, Matteo Ferrara, Annalisa Franco, Luuk J
K. Bommanna Raja, Matteo Ferrara, Annalisa Franco, Luuk J. Spreeuwers, Ilias Batskos, Florens de Wit, Marta Gomez-Barrero, Ulrich Scherhag, Daniel Fischer, Sushma Krupa Venkatesh, Jag Mohan Singh, Guoqiang Li, Lo¨ıc Bergeron, Sergey Isadskiy, Raghavendra Ramachandra, Christian...
2020
-
[27]
Towards making Morphing Attack Detection Robust using Hybrid Scale-Space Colour Texture Features
Raghavendra Ramachandra, Sushma Venkatesh, Kiran Raja, and Christoph Busch. Towards making Morphing Attack Detection Robust using Hybrid Scale-Space Colour Texture Features. In Proceedings of International Conference on Identity, Security, and Behavior Analysis (ISBA), pages 1–8,
-
[28]
Vulnerability Analysis of Face Morphing Attacks from Landmarks and Generative Adversarial Net- works
Eklavya Sarkar, Pavel Korshunov, Laurent Colbois, and S´ebastien Marcel. Vulnerability Analysis of Face Morphing Attacks from Landmarks and Generative Adversarial Net- works. In arXiv, 2020. arXiv:2012.05344. 4
2020 arXiv
-
[29]
Neural Implicit Morphing of Face Images
Guilherme Schardong, Tiago Novello, Hallison Paz, Iurii Medvedev, Vin ´ıcius da Silva, Luiz Velho, and Nuno Gonc ¸alves. Neural Implicit Morphing of Face Images. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR) , pages 7321–7330, Ju...
2024
-
[30]
Detecting Morphed Face Images Us- ing Facial Landmarks
Ulrich Scherhag, Dhanesh Budhrani, Marta Gomez-Barrero, and Christoph Busch. Detecting Morphed Face Images Us- ing Facial Landmarks. In Proceedings of Image and Signal Processing, 2018. 1
2018
-
[31]
Towards Detection of Morphed Face Images in Electronic Travel Documents
Ulrich Scherhag, Christian Rathgeb, and Christoph Busch. Towards Detection of Morphed Face Images in Electronic Travel Documents. In Proceedings of IAPR International Workshop on Document Analysis Systems (DAS), pages 187– 192, 2018. 1
2018
-
[32]
Deep Face Representations for Differential Morphing Attack Detection
Ulrich Scherhag, Christian Rathgeb, Johannes Merkle, and Christoph Busch. Deep Face Representations for Differential Morphing Attack Detection. Proceedings of IEEE Transac- tions on Information Forensics and Security, 15:3625–3639,
-
[33]
SDeMorph: Towards Better Facial De- morphing from Single Morph
Nitish Shukla. SDeMorph: Towards Better Facial De- morphing from Single Morph. In Proceedings of IEEE In- ternational Joint Conference on Biometrics (IJCB) . IEEE,
-
[34]
dc-GAN: Dual-Conditioned GAN for Face Demorphing From a Single Morph
Nitish Shukla and Arun Ross. dc-GAN: Dual-Conditioned GAN for Face Demorphing From a Single Morph. In arXiv,
-
[35]
Facial Demorphing via Iden- tity Preserving Image Decomposition
Nitish Shukla and Arun Ross. Facial Demorphing via Iden- tity Preserving Image Decomposition. Proceedings of IEEE International Joint Conference on Biometrics (IJCB) , 2024. 1, 2, 3, 4, 5
2024
-
[36]
Face Morphing Attack Generation and Detection: A Comprehensive Survey
Sushma Venkatesh, Raghavendra Ramachandra, Kiran Raja, and Christoph Busch. Face Morphing Attack Generation and Detection: A Comprehensive Survey. Proceedings of IEEE Transactions on Technology and Society , 2(3):128– 145, 2021. 1
2021
-
[37]
Zhang, Z
K. Zhang, Z. Zhang, Z. Li, and Y . Qiao. Joint Face Detection and Alignment Using Multitask Cascaded Convolutional Networks. IEEE Signal Processing Letters , 23(10):1499– 1503, 2016. 5
2016
-
[38]
Deep Adversarial Decomposition: A Unified Framework for Separating Superimposed Images
Zhengxia Zou, Sen Lei, Tianyang Shi, Zhenwei Shi, and Jieping Ye. Deep Adversarial Decomposition: A Unified Framework for Separating Superimposed Images . In Pro- ceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12803–12813, 2020. 3
2020
-
[2024]
arXiv:2411.14494. 2, 3
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.