Pith. sign in

REVIEW 4 major objections 6 minor 2 cited by

On the Illusion of Gender Bias in Face Recognition: Explaining the Fairness Issue Through Non-demographic Attributes

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper argues that the gender gap in face recognition is an illusion: it vanishes once male and female test images share the same non-demographic attributes such as hairstyle, facial hair, and accessories.

desk verdict The combination search is a real methodological contribution, but the causal claim that gender bias is an illusion is not supported by the design: the search optimizes iGARBE, so finding near-perfect fairness scores on selected subsets is partly by construction. read the letter →

arxiv 2501.12020 v2 pith:A4C3IBD7 submitted 2025-01-21 cs.CV

classification cs.CV
keywords genderbiasfacerecognitionfairnessnon-demographicattributesattributedecorrelationiGARBECoFairMAAD-Face
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to show that the well-documented gender gap in face recognition accuracy is not caused by gender itself, but by non-demographic appearance attributes socially correlated with gender. It introduces an unsupervised search that decorrelates 40 facial attributes, forms clusters, and looks for combinations that, when used to filter a test set so that both genders share them, push the fairness score toward perfect. Across two recognition models, the top combinations reach fairness scores above 0.99, and the error difference between male and female subjects becomes negligible. If correct, this means fairness evaluations should control for these shared appearance attributes, and mitigations should target robustness to hairstyle, facial hair, and occluding accessories rather than gender per se.

What carries the argument

The method runs on three components. A decorrelation-by-clustering toolchain groups the 40 MAAD-Face non-demographic attributes into 27 clusters so that correlated attributes are treated as one unit. Two new metrics quantify fairness: iGARBE, a fairness score bounded in [0,1] with 1 meaning perfect fairness, and CoFair, which contextualizes a fairness score as a percentile relative to scores of other attribute sets. The search itself is a greedy breadth-first expansion that repeatedly adds the n most fairness-increasing attribute-label assignments to combinations, using a ranking metric that balances sample retention, verification error, and fairness. Applying a combination as a filter predicate to the comparison dataset, requiring both compared images to carry the same attribute labels, is what the authors call equalizing, and this operation is what makes the gender gap collapse.

What would settle it

Take the full comparison database and sample many random attribute combinations of the same size as the reported fair combinations; if a large fraction of them also reach iGARBE values above 0.99 under the same sampling procedure, the vanishing gap is a selection effect rather than evidence about specific attributes.

Watch

Extended reading notes

Core claim

The central claim is that once male and female subjects share specific non-demographic attributes, the gender gap in recognition accuracy vanishes. The paper reports that this holds for both tested face recognition models, ArcFace and FaceNet, across all experimental setups, with the top attribute combinations yielding iGARBE fairness scores of 0.997 to 0.9997. On that basis the authors interpret gender bias in face recognition as likely originating from non-demographic attributes associated with gender, such as hairstyles, facial hair, and occluding accessories, rather than from gender itself. The vanishing gap is presented not as a tweak to the metric but as evidence that the apparent bias is an artifact of how appearance differs between the male and female groups in existing test data.

Load-bearing premise

The argument assumes that filtering the test data to subsets where both genders share the same attribute labels reveals those attributes as the cause of the original gender gap, rather than simply selecting subsets that happen to have smaller bias.

Editorial extensions

If this is right

  • If the claim holds, fairness evaluations of face recognition should report results on test sets balanced not only by gender but also by the relevant non-demographic attributes.
  • The gender gap can be erased by controlling for a small number of appearance categories, so a fair FRS is attainable through targeted test-set design or robustness to these attributes.
  • The finding supports the conclusion that balancing training data by gender alone is insufficient, since the gap is driven by appearance attributes rather than gender representation.
  • Mitigation efforts should shift from removing gender information to making models invariant to hairstyle, facial hair, and occluding accessories.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: because the search explicitly optimizes the fairness metric, finding high-scoring combinations is partially by construction; a stronger test of the causal claim would compare the best combinations against a null distribution of randomly chosen combinations of the same size.
  • Editorial inference: the 'social definition' interpretation predicts that the specific attribute combinations that erase the gap will shift as gender-specific appearance norms change, which is a testable prediction on newer datasets.
  • Editorial inference: the same decorrelation-and-search pipeline could be transferred to other protected attributes such as age or ethnicity, but the attribute set and correlation structure would need to be re-derived for each case.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes an automated search over combinations of non-demographic facial attributes to explain the gender gap in face recognition accuracy. It introduces three components: a clustering-based decorrelation toolchain for MAAD-Face attributes, an inverse GARBE (iGARBE) fairness metric together with a contextualized fairness measure (CoFair), and a greedy search algorithm that forms attribute combinations whose use as filter predicates yields nearly equal error rates for male and female subjects. With ArcFace and FaceNet on VGGFace2/MAAD-Face, the authors report top combinations with iGARBE near 0.99 and conclude that the gender gap vanishes once male and female subjects share these attributes, interpreting this as evidence that gender bias is a matter of social definition of appearance rather than biology.

Significance. If the causal claim were supported, the paper would be a meaningful contribution to the fairness-in-biometrics literature: it extends prior single-attribute analyses to a large set of 40 attributes, makes an explicit attempt to reduce attribute correlation, evaluates on two widely used face recognition models, and reports large genuine-sample counts. The toolchain and search framework are potentially reusable for exploratory fairness analysis. However, the central causal conclusion is not established by the current experimental design; the paper presently reads as an optimization study that finds fair subpopulations, not as a causal explanation of the gender gap.

major comments (4)
  1. [Section III-B3, Eq. (8); Algorithm 1] The near-perfect iGARBE scores reported in Tables IIIa and IIIb are an expected consequence of the optimization procedure, not an empirical discovery. The ranking metric R_f = σ(ω(f_i − f_0)) explicitly rewards increases in iGARBE, and Algorithm 1 discards every candidate with FAIRNESS(m_Al) ≤ f0 (line 10) and branches only into the top-n assignments (line 14). The paper provides no null distribution over random attribute combinations of the same cardinality and no held-out evaluation, so the selected combinations' high scores cannot be used to attribute the original gender gap to those attributes. A random-combination baseline and an out-of-sample replication are necessary before the causal reading in Section VII is warranted.
  2. [Section V-C and Tables IIIa/IIIb] The paper reports no uncertainty or validation for the selected combinations. Although Section III-B2 draws γ=3 sample sets, Tables II and III present only single point estimates, without standard deviations, confidence intervals, or repeated-selection results. Since the same VGGFace2/MAAD-Face data are used both to select and to evaluate the combinations, the reported iGARBE values may reflect selection overfitting; the large genuine-sample counts reduce sampling noise but do not address this issue. The cross-model results (ArcFace combinations evaluated on FaceNet and vice versa) are informative but are still on the same underlying data.
  3. [Section VII] The conclusion that 'Gender bias in FRS is likely no issue of biology but of the social definition of gender-specific appearance' is not supported by the experimental design. The selected combinations include attributes that are at least partly biological or morphological, such as Receding Hairline, Frontal Facial Hair, Big Lips, and Corpulent in Tables IIIa and IIIb. Filtering on these attributes removes biological variation together with social appearance, so the vanishing gap is also consistent with a morphological or biological explanation. To support the social-definition claim, the authors would need to demonstrate that the gap vanishes when biological/morphological factors are controlled for, or that purely social attributes alone are sufficient; neither is shown.
  4. [Section III-A2 and Section V-C1] CoFair is presented as a core metric, but its estimate rests on a kernel density over a very small number of scores. The KDE uses only the 41 single-attribute iGARBE values from Table II, with a bandwidth adjustment factor of 0.5, and the CoFair values in Table III (e.g., 0.9997) are tail probabilities of that smoothed 41-point distribution. The metric is therefore not robust enough to support statements such as 'higher than 99.9% of single attributes'; the paper should either report the sensitivity of CoFair to bandwidth and sample size or substantially weaken the contextual claims.
minor comments (6)
  1. [Section I and Section III-B] Calling Algorithm 1 'unsupervised' is misleading, since it optimizes an explicit fairness objective; 'objective-guided search' or 'automated search' would be more accurate.
  2. [Section V-A] The term 'decorrelated' overstates the outcome; the clustering result still has a mean inter-cluster correlation of about 0.21 and a maximum of about 0.67, so the attribute set is less correlated rather than decorrelated.
  3. [Throughout] There are several typos and stylistic issues: 'seperately' in the Related Work section, 'assigment' in Section III-B3, and 'Expanability' in the Index Terms should be corrected.
  4. [Section III-B2] The notation 'Tγ S(∩∅) g = ∅' is not defined; clarify that it refers to the disjointness of the γ genuine-sample sets.
  5. [Table III] The first data row in Tables IIIa and IIIb appears to be the unfiltered baseline but is not labeled as such; add an explicit 'baseline' label.
  6. [Figure 4] The 90% threshold for 'strongly correlated' is introduced only in the text; state it in the caption or define it before referencing the figure.

Circularity Check

1 steps flagged · score 6.0 of 10

The 'vanishing gender gap' is an artifact of selecting attribute combinations that maximize the iGARBE fairness metric; the causal 'social not biological' interpretation is not independently established.

  1. fitted input called prediction [Section III-B3, Eq. (8), Algorithm 1; conclusion in Section VII]
    "Third, in line with this work’s objective, the assignment combination in question should increase fairness. The bigger the increase, the higher the reward. Decreases are penalized. ... Rf (fi, f0) = σ(ω (fi − f0))"

    Algorithm 1 retains only combinations whose FAIRNESS(mAl) exceeds f0 (line 10) and ranks candidates by Rf, which is a monotone function of the gain in iGARBE (fi − f0). Therefore every reported top combination is, by construction, fairness-increasing; stating that 'the gender gap vanishes' for these combinations (Section VII) restates the selection criterion. The paper provides no random-combination baseline and no held-out validation, so the near-perfect iGARBE values in Table III are expected outputs of the optimization, not independent empirical evidence. The one genuinely non-circular component, cross-model transfer of the selected combinations, shows robustness but cannot distinguish causal attributes from generic subpopulation homogeneity.

full rationale

The derivation chain is: (1) define iGARBE fairness metric (Eq. 3), (2) greedily search attribute combinations with a ranking metric that rewards fairness gain (Eq. 8), (3) report the resulting combinations and their iGARBE scores, and (4) conclude that shared non-demographic attributes eliminate the gender gap, implying social rather than biological causes. Steps (1)-(3) are methodologically coherent, but the empirical claim in step (4) is partially circular because the selected combinations are generated by optimizing the same metric used to verify the claim. The paper acknowledges this in Section V-C1: 'the optimization approach works as expected... these results are also to be expected, since the underlying iGARBE scores already approach values resembling nearly perfect fairness.' No random-combination null distribution is reported, so we cannot tell whether the achieved iGARBE ~0.99 exceeds what any comparably specific subset would achieve. The cross-model evaluation is a valuable independent check and prevents a fully circular verdict, but it only shows transferability of the selected subsets to another FRS, not that the attributes are the cause of the original gap. The final biological conclusion is a further non-sequitur because several of the 'non-demographic' attributes are biological traits, undermining the 'not biology' interpretation. I therefore score the paper 6: the central 'vanishing gap' result is in part forced by the optimization objective, while the cross-model consistency and the specific attribute identities retain some independent content.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The central claim rests mainly on hyperparameters chosen by hand (ranking slopes, sampling ratios, clustering stopping point, unreported search branching), on the accuracy and completeness of MAAD-Face attributes, and on the interpretive assumption that filtering test subsets can identify causes of the original gap. No new physical entities are introduced; iGARBE and CoFair are definitions rather than invented entities.

free parameters (6)
  • Ranking metric slopes mu, lambda, omega = mu=1.3865, lambda=omega=4
    Chosen by hand in Section III-B3 based on assumed baselines; controls how strongly the search rewards fairness, error reduction, and sample retention.
  • Sampling ratio rho_s and sample sets gamma = rho_s=1/5, gamma=3
    Set in Section III-B2 to balance computational complexity; affects variance and which attribute combinations pass the sampling constraints.
  • Minimum genuine samples lambda_g = 1/FMR = 1000
    Set in Section III-B2 as a lower bound per sample set; determines whether a combination is analyzed at all.
  • Clustering stopping iteration imax = 13
    Selected in Section V-A because sampling requirements cannot be met after iteration 13; a data-dependent choice of decorrelation granularity.
  • Greedy search branching n and max depth d_max = Not reported
    Algorithm 1 requires these values, but the experiments do not state them; the top-10 results depend on them.
  • KDE bandwidth adjustment factor = 0.5
    Used in Section V-C to prevent oversmoothing when estimating the CoFair distribution; affects reported CoFair values.
assumptions (4)
  • domain assumption MAAD-Face attribute annotations are accurate enough for the analysis.
    All filtering and equalization rely on these labels; the paper acknowledges possible label noise in Section VI but assumes it does not affect conclusions.
  • domain assumption The 40 non-demographic attributes sufficiently capture the appearance factors relevant to the gender gap.
    If important appearance attributes are missing, attributing the gap to the included ones could be spurious; the Limitations section concedes that a more refined attribute set could change the description.
  • domain assumption Pearson-correlation clustering and label harmonization preserve the information needed to explain the gap.
    Clusters are treated as standalone attributes after inverting labels to make correlations positive, which assumes semantic inverse relationships are valid.
  • ad hoc to paper Filtering test subsets where both genders share attributes allows causal attribution of the original gap to those attributes.
    This is the load-bearing interpretive step in the Conclusion; the experiment only selects subsets and does not control for unmeasured biological or social confounders.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the Illusion of Gender Bias in Face Recognition: Explaining the Fairness Issue Through Non-demographic Attributes." pith.science (2026). https://pith.science/paper/A4C3IBD7

@misc{pith2026250112020,
  author       = {Pith},
  title        = {Pith review of: On the Illusion of Gender Bias in Face Recognition: Explaining the Fairness Issue Through Non-demographic Attributes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/A4C3IBD7}},
  note         = {Machine review of arXiv:2501.12020}
}
read the original abstract

Face recognition systems (FRS) exhibit significant accuracy differences based on the user's gender. Since such a gender gap reduces the trustworthiness of FRS, more recent efforts have tried to find the causes. However, these studies make use of manually selected, correlated, and small-sized sets of facial features to support their claims. In this work, we analyze gender bias in face recognition by successfully extending the search domain to decorrelated combinations of 40 non-demographic facial characteristics. First, we introduce a toolchain to effectively decorrelate and aggregate facial attributes to enable a less-biased gender analysis on large-scale data. Second, we tailor two specialized metrics to quantify the effect of facial attributes on absolute and relative fairness. Based on these grounds, we thirdly present a novel unsupervised joint investigation framework capable of identifying attribute combinations leading to vanishing bias when used as filter predicates for balanced testing datasets. Experiments show the gender gap vanishing when images of male and female subjects share specific attributes, clearly indicating that the disparate performance is not a question of biology but of the social definition of appearance. These findings could reshape our understanding of fairness in face biometrics and provide insights into FRS, helping to address gender bias issues.

Figures

Figures reproduced from arXiv: 2501.12020 by the authors.

Figure 1
Figure 1. Attribute annotation correlations - The correlations are computed using the Pearson coefficient. The depicted attributes are selected such that the 15 highest absolute, i.e., positive or negative, correlations are visible. As can be seen, very strong correlations exist between various attributes, indicating that they are not statistically independent. can be seen, correlations with absolute values of 0.75 and higher… view at source ↗
Figure 2
Figure 2. Evolution of critical clustering metrics over pro￾gressing decorrelation - The chosen measurements reflect the efficacy of the devised algorithm in reducing the correlation of MAAD-Face’s non-demographic attributes. The iteration range reflects the clustering intensity, ranging from 0 (no clustering) to 39 (all attributes in one cluster). The number of clusters, mean and maximum correlation decrease linearly. For an… view at source ↗
Figure 3
Figure 3. Fairness iGARBE distributions for both FRS - The estimations were performed using a KDE with Gaussian kernels with a bandwidth determined using Scott’s rule. These distributions are the basis for CoFair. 0.5 to prevent oversmoothing. For the estimation, we use the samples from Table II. We specifically use only those samples whose FNMR is better than or at max 10 % worse than the corresponding baseline FNMR. In doin… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Relative frequency of occurrence of attribute-label pairs after filtering for attribute combination achieving highest iGARBE score - The filter-dictating combination is chosen based on Tables IIIa and IIIb for ArcFace and FaceNet, respectively. The shown distributions …

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Responsible Face Recognition Approach for Small and Mid-Scale Systems Through Personalized Neural Networks

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MOTE replaces fixed face embeddings with per-identity binary classifiers trained using KDE-generated synthetic samples, improving gender fairness and privacy at the cost of storage and enrollment time.

  2. Review of Demographic Fairness in Face Recognition

    cs.CV 2025-02 conditional novelty 3.0 of 10

    A structured review of demographic fairness in face recognition covering causes, datasets, assessment metrics, and mitigation methods.

Reference graph

Works this paper leans on

56 extracted references · 55 canonical work pages · cited by 2 Pith papers

  1. [1]

    Are face recognition systems accurate? Depends on your race,

    M. Orcutt, “Are face recognition systems accurate? Depends on your race,” MIT Technology Review 2016 , 2016

  2. [2]

    Facial recognition is accurate, if you’re a white guy,

    S. Lohr, “Facial recognition is accurate, if you’re a white guy,” in Ethics of Data and Analytics , Auerbach Publications, 2022, pp. 143–147

  3. [3]

    Racist and sexist’facial recognition cam- eras could lead to false arrests,

    T. Hoggins, “Racist and sexist’facial recognition cam- eras could lead to false arrests,” The Telegraph, 2019

  4. [4]

    Is facial recognition too biased to be let loose?

    D. Castelvecchi, “Is facial recognition too biased to be let loose?” Nature, Nov. 2020

  5. [5]

    Emerging from ai utopia,

    E. Santow, “Emerging from ai utopia,” Science, vol. 368, no. 6486, pp. 9–9, 2020

  6. [6]

    Face Recognition Performance: Role of Demographic Information,

    B. Klare, M. J. Burge, J. C. Klontz, R. W. V . Bruegge, and A. K. Jain, “Face Recognition Performance: Role of Demographic Information,” IEEE Trans. Inf. Forensics Secur., vol. 7, no. 6, pp. 1789–1801, 2012

  7. [7]

    An other-race effect for face recognition algorithms,

    P. J. Phillips, F. Jiang, A. Narvekar, J. H. Ayyad, and A. J. O’Toole, “An other-race effect for face recognition algorithms,” ACM Trans. Appl. Percept. , vol. 8, no. 2, 14:1–14:11, 2011

  8. [8]

    Grother, M

    P. Grother, M. Ngan, and K. Hanaoka, Face Recognition Vendor Test Part 3: Demographic Effects , 2019

Show all 56 references
  1. [9]

    The gender gap in face recognition accuracy is a hairy problem,

    A. Bhatta, V . Albiero, K. W. Bowyer, and M. C. King, “The gender gap in face recognition accuracy is a hairy problem,” in IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACV 2023 - Workshops, Waikoloa, HI, USA, January 3-7, 2023, IEEE, 2023, pp. 1–10

  2. [10]

    A Com- prehensive Study on Face Recognition Biases Beyond Demographics,

    P. Terh ¨orst, J. N. Kolf, M. Huber, et al. , “A Com- prehensive Study on Face Recognition Biases Beyond Demographics,” IEEE Transactions on Technology and Society, vol. 3, no. 1, pp. 16–30, 2022

  3. [11]

    Comparison-Level Mitigation of Ethnic Bias in Face Recognition,

    P. Terh ¨orst, M. L. Tran, N. Damer, F. Kirchbuchner, and A. Kuijper, “Comparison-Level Mitigation of Ethnic Bias in Face Recognition,” in 8th International Work- shop on Biometrics and Forensics, IWBF 2020, Porto, Portugal, April 29-30, 2020 , IEEE, 2020, pp. 1–6

  4. [12]

    Demographic bias in biometric systems: Current research and applicable standards,

    J. W. M. Campbell, “Demographic bias in biometric systems: Current research and applicable standards,” 2017

  5. [13]

    Analysis of Gender Inequality In Face Recognition Accuracy,

    V . Albiero, K. K. S, K. Vangara, K. Zhang, M. C. King, and K. W. Bowyer, “Analysis of Gender Inequality In Face Recognition Accuracy,” in IEEE Winter Applica- tions of Computer Vision Workshops, WACV Workshops 2020, Snowmass Village, CO, USA, March 1-5, 2020 , IEEE, 2020, pp. 81–89

  6. [14]

    How Does Gender Balance In Training Data Affect Face Recog- nition Accuracy?

    V . Albiero, K. Zhang, and K. W. Bowyer, “How Does Gender Balance In Training Data Affect Face Recog- nition Accuracy?” In 2020 IEEE International Joint Conference on Biometrics, IJCB 2020, Houston, TX, USA, September 28 - October 1, 2020 , IEEE, 2020, pp. 1–10

  7. [15]

    Logical consistency and greater descriptive power for facial hair attribute learning,

    H. Wu, G. Bezold, A. Bhatta, and K. W. Bowyer, “Logical consistency and greater descriptive power for facial hair attribute learning,” in CVPR, IEEE, 2023, pp. 8588–8597

  8. [16]

    Facial Hair area in face recognition across demographics: Small size, big Effect,

    H. Wu, S. Tian, A. Bahatta, K. ¨Ozt¨urk, K. Ricanek Jr., and K. W. Bowyer, “Facial Hair area in face recognition across demographics: Small size, big Effect,” presented at the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, 2024

  9. [17]

    Gendered differences in face recognition accuracy ex- plained by hairstyles, makeup, and facial morphology,

    V . Albiero, K. Zhang, M. C. King, and K. W. Bowyer, “Gendered differences in face recognition accuracy ex- plained by hairstyles, makeup, and facial morphology,” IEEE Trans. Inf. Forensics Secur., vol. 17, pp. 127–137, 2022

  10. [18]

    Face recognition accuracy across demographics: Shining a light into the problem,

    H. Wu, V . Albiero, K. S. Krishnapriya, M. C. King, and K. W. Bowyer, “Face recognition accuracy across demographics: Shining a light into the problem,” in CVPRW, IEEE, 2023, pp. 1041–1050

  11. [19]

    Is Face Recognition Sexist? No, Gendered Hairstyles and Biology Are,

    V . Albiero and K. W. Bowyer, “Is Face Recognition Sexist? No, Gendered Hairstyles and Biology Are,” in 31st British Machine Vision Conference 2020, BMVC 2020, Virtual Event, UK, September 7-10, 2020, BMV A Press, 2020

  12. [20]

    On soft-biometric information stored in biometric face embeddings,

    P. Terh ¨orst, D. F ¨ahrmann, N. Damer, F. Kirchbuchner, and A. Kuijper, “On soft-biometric information stored in biometric face embeddings,” IEEE Trans. Biom. Behav. Identity Sci. , vol. 3, no. 4, pp. 519–534, 2021

  13. [21]

    Beyond identity: What information is stored in biometric face templates?

    P. Terh ¨orst, D. F ¨ahrmann, N. Damer, F. Kirchbuchner, and A. Kuijper, “Beyond identity: What information is stored in biometric face templates?” In 2020 IEEE International Joint Conference on Biometrics, IJCB 2020, Houston, TX, USA, September 28 - October 1, 2020, IEEE, 202...

  14. [22]

    Face recognition vendor test 2002,

    P. J. Phillips, P. Grother, R. Micheals, D. M. Blackburn, E. Tabassi, and M. Bone, “Face recognition vendor test 2002,” in 2003 IEEE International SOI Conference. Proceedings (Cat. No. 03CH37443), IEEE, 2003, p. 44

  15. [23]

    Overview of the face recognition grand challenge,

    P. J. Phillips, P. J. Flynn, W. T. Scruggs, et al. , “Overview of the face recognition grand challenge,” in 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2005), 20-26 June 2005, San Diego, CA, USA , IEEE Computer Society, 2005, pp. 947–954

  16. [24]

    Factors that influence algorithm performance in the face recognition grand challenge,

    J. R. Beveridge, G. H. Givens, P. J. Phillips, and B. A. Draper, “Factors that influence algorithm performance in the face recognition grand challenge,” Comput. Vis. Image Underst., vol. 113, no. 6, pp. 750–762, 2009

  17. [25]

    A meta-analysis of face recognition covariates,

    Y . M. Lui, D. Bolme, B. A. Draper, J. R. Beveridge, G. Givens, and P. J. Phillips, “A meta-analysis of face recognition covariates,” in 2009 IEEE 3rd International Conference on Biometrics: Theory, Applications, and Systems, IEEE, 2009, pp. 1–8

  18. [26]

    P. J. Grother, P. J. Grother, P. J. Phillips, and G. W. Quinn, Report on the evaluation of 2D still-image face recognition algorithms. US Department of Commerce, National Institute of Standards and Technology, 2011

  19. [27]

    Face recognition performance: Role of demographic information,

    B. Klare, M. J. Burge, J. C. Klontz, R. W. V . Bruegge, and A. K. Jain, “Face recognition performance: Role of demographic information,” IEEE Trans. Inf. Forensics Secur., vol. 7, no. 6, pp. 1789–1801, 2012

  20. [28]

    Labeled faces in the wild: A database forstudy- ing face recognition in unconstrained environments,

    G. B. Huang, M. Mattar, T. Berg, and E. Learned- Miller, “Labeled faces in the wild: A database forstudy- ing face recognition in unconstrained environments,” in Workshop on faces in’Real-Life’Images: detection, alignment, and recognition , 2008

  21. [29]

    A longitudinal study of automatic face recognition,

    L. Best-Rowden and A. K. Jain, “A longitudinal study of automatic face recognition,” in International Confer- ence on Biometrics, ICB 2015, Phuket, Thailand, 19-22 May, 2015, IEEE, 2015, pp. 214–221

  22. [30]

    Longitudinal study of automatic face recognition,

    L. Best-Rowden and A. K. Jain, “Longitudinal study of automatic face recognition,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 1, pp. 148–162, 2018

  23. [31]

    What should be balanced in a

    H. Wu and K. W. Bowyer, “What should be balanced in a ”balanced” face recognition dataset?” In 34th British Machine Vision Conference 2022, BMVC 2022, Aberdeen, UK, November 20-24, 2023 , BMV A Press, 2023, p. 235

  24. [32]

    Influence of make-up on facial recognition,

    S. Ueda and T. Koyama, “Influence of make-up on facial recognition,” Perception, vol. 39, no. 2, pp. 260–264, 2010

  25. [33]

    Can facial cos- metics affect the matching accuracy of face recognition systems?

    A. Dantcheva, C. Chen, and A. Ross, “Can facial cos- metics affect the matching accuracy of face recognition systems?” In IEEE Fifth International Conference on Biometrics: Theory, Applications and Systems, BTAS 2012, Arlington, VA, USA, September 23-27, 2012 , IEEE, 2012, pp. 391–398

  26. [34]

    Face authentication with makeup changes,

    G. Guo, L. Wen, and S. Yan, “Face authentication with makeup changes,” IEEE Trans. Circuits Syst. Video Technol., vol. 24, no. 5, pp. 814–825, 2014

  27. [35]

    Beard segmentation and recognition bias,

    K. Ozturk, G. Bezold, A. Bhatta, H. Wu, and K. W. Bowyer, “Beard segmentation and recognition bias,” CoRR, vol. abs/2308.15740, 2023. arXiv: 2308.15740

  28. [36]

    Post-comparison mitigation of demo- graphic bias in face recognition using fair score normal- ization,

    P. Terh ¨orst, J. N. Kolf, N. Damer, F. Kirchbuchner, and A. Kuijper, “Post-comparison mitigation of demo- graphic bias in face recognition using fair score normal- ization,” Pattern Recognit. Lett., vol. 140, pp. 332–338, 2020

  29. [37]

    Jointly de-biasing face recognition and demographic attribute estimation,

    S. Gong, X. Liu, and A. K. Jain, “Jointly de-biasing face recognition and demographic attribute estimation,” in Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Pro- ceedings, Part XXIX , A. Vedaldi, H. Bischof, T. Brox, and J. Frahm, Ed...

  30. [38]

    Learning disentangled representation for fair facial attribute clas- sification via fairness-aware information alignment,

    S. Park, S. Hwang, D. Kim, and H. Byun, “Learning disentangled representation for fair facial attribute clas- sification via fairness-aware information alignment,” in Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Ap- ...

  31. [39]

    Distill and de-bias: Mitigating bias in face verification using knowledge distillation,

    P. Dhar, J. Gleason, A. Roy, C. D. Castillo, P. J. Phillips, and R. Chellappa, “Distill and de-bias: Mitigating bias in face verification using knowledge distillation,”CoRR, vol. abs/2112.09786, 2021. arXiv: 2112.09786

  32. [40]

    PASS: protected attribute suppression system for mitigating bias in face recognition,

    P. Dhar, J. Gleason, A. Roy, C. D. Castillo, and R. Chel- lappa, “PASS: protected attribute suppression system for mitigating bias in face recognition,” in 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021 , IEEE, 2...

  33. [41]

    Fairness in Bio- metrics: A Figure of Merit to Assess Biometric Ver- ification Systems,

    T. de Freitas Pereira and S. Marcel, “Fairness in Bio- metrics: A Figure of Merit to Assess Biometric Ver- ification Systems,” IEEE Trans. Biom. Behav. Identity Sci., vol. 4, no. 1, pp. 19–29, 2022

  34. [42]

    Publications Office of the European Union, 2015, ISBN : 978-92- 95205-43-7

    Frontex, Best Practice Technical Guidelines for Au- tomated Border Control (ABC) Systems . Publications Office of the European Union, 2015, ISBN : 978-92- 95205-43-7

  35. [43]

    Demographic Differentials in Face Recognition Algorithms,

    P. J. Grother, “Demographic Differentials in Face Recognition Algorithms,” presented at the EAB Virtual Event Series - Demographic Fairness in Biometric Sys- tems, 2021

  36. [44]

    Evaluating Proposed Fairness Models for Face Recognition Algorithms,

    J. J. Howard, E. J. Laird, R. E. Rubin, Y . B. Sirotin, J. L. Tipton, and A. R. Vemury, “Evaluating Proposed Fairness Models for Face Recognition Algorithms,” in Pattern Recognition, Computer Vision, and Image Processing. ICPR 2022 International Workshops and Challenges - Mont...

  37. [45]

    The Small-Sample Bias of the Gini Coeffi- cient: Results and Implications for Empirical Research,

    G. Deltas, “The Small-Sample Bias of the Gini Coeffi- cient: Results and Implications for Empirical Research,” The Review of Economics and Statistics , vol. 85, no. 1, pp. 226–234, 2003. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 16

  38. [46]

    MAAD-Face: A Mas- sively Annotated Attribute Dataset for Face Images,

    P. Terh ¨orst, D. F ¨ahrmann, J. N. Kolf, N. Damer, F. Kirchbuchner, and A. Kuijper, “MAAD-Face: A Mas- sively Annotated Attribute Dataset for Face Images,” IEEE Trans. Inf. Forensics Secur. , vol. 16, pp. 3942– 3957, 2021

  39. [47]

    VGGFace2: A Dataset for Recognising Faces across Pose and Age,

    Q. Cao, L. Shen, W. Xie, O. M. Parkhi, and A. Zis- serman, “VGGFace2: A Dataset for Recognising Faces across Pose and Age,” in 13th IEEE International Con- ference on Automatic Face & Gesture Recognition, FG 2018, Xi’an, China, May 15-19, 2018 , IEEE Computer Society, 2018, pp. 67–74

  40. [48]

    Using Semi-Joins to Solve Relational Queries,

    P. A. Bernstein and D. W. Chiu, “Using Semi-Joins to Solve Relational Queries,” J. ACM , vol. 28, no. 1, pp. 25–40, 1981

  41. [49]

    ArcFace: Additive Angular Margin Loss for Deep Face Recog- nition,

    J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “ArcFace: Additive Angular Margin Loss for Deep Face Recog- nition,” in IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019 , Computer Vision Foundation / IEEE, 2019, pp. 4690–4699

  42. [50]

    FaceNet: A unified embedding for face recognition and cluster- ing,

    F. Schroff, D. Kalenichenko, and J. Philbin, “FaceNet: A unified embedding for face recognition and cluster- ing,” in IEEE Conference on Computer Vision and Pat- tern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015 , IEEE Computer Society, 2015, pp. 815– 823

  43. [51]

    MS- Celeb-1M: A Dataset and Benchmark for Large-Scale Face Recognition,

    Y . Guo, L. Zhang, Y . Hu, X. He, and J. Gao, “MS- Celeb-1M: A Dataset and Benchmark for Large-Scale Face Recognition,” in Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Nether- lands, October 11-14, 2016, Proceedings, Part III , ser. Lecture Notes in C...

  44. [52]

    Stacked Dense U-Nets with Dual Transformers for Robust Face Alignment,

    J. Guo, J. Deng, N. Xue, and S. Zafeiriou, “Stacked Dense U-Nets with Dual Transformers for Robust Face Alignment,” in British Machine Vision Conference 2018, BMVC 2018, Newcastle, UK, September 3-6, 2018, BMV A Press, 2018, p. 44

  45. [53]

    One Millisecond Face Alignment with an Ensemble of Regression Trees,

    V . Kazemi and J. Sullivan, “One Millisecond Face Alignment with an Ensemble of Regression Trees,” in 2014 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2014, Columbus, OH, USA, June 23-28, 2014, IEEE Computer Society, 2014, pp. 1867– 1874

  46. [54]

    Information technology — Biometric performance testing and reporting — Part 1: Principles and frame- work (ISO/IEC 19795-1),

    “Information technology — Biometric performance testing and reporting — Part 1: Principles and frame- work (ISO/IEC 19795-1),” International Organization for Standardization, Standard, 2021

  47. [55]

    The Effect of Broad and Specific Demographic Homogene- ity on the Imposter Distributions and False Match Rates in Face Recognition Algorithm Performance,

    J. J. Howard, Y . B. Sirotin, and A. R. Vemury, “The Effect of Broad and Specific Demographic Homogene- ity on the Imposter Distributions and False Match Rates in Face Recognition Algorithm Performance,” in 10th IEEE International Conference on Biometrics Theory, Applications ...

  48. [56]

    Responsible AI for Biometrics

    D. W. Scott, Multivariate Density Estimation: The- ory, Practice, and Visualization . Wiley, 1992, ISBN : 9780470316849. Paul Jonas Kurz received the B.Sc. degree in Computer Science from TU Darmstadt, Darmstadt, Germany, in 2023. He currently pursues his M.Sc. degree in Compu...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.