Pith. sign in

REVIEW 4 major objections 6 minor 50 references

Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening

T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read An AI hiring tool can pass a bias audit while being unable to tell qualified from unqualified candidates.

desk verdict The core idea—that an unbiased-looking model may simply be incompetent—is worth taking seriously, but the flagship Experiment 2b is confounded and there is a data inconsistency in the ANOVA reporting that needs fixing before the paper is referable. read the letter →

arxiv 2507.11548 v3 pith:DJAKLGPL submitted 2025-07-11 cs.CY cs.AIcs.CL

classification cs.CYcs.AIcs.CL
keywords IllusionofNeutralityAIrésuméscreeningintersectionalbiasLLMauditevaluativecompetencedual-validationframeworkkeyword-stuffingtestgenerativehiringtools
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that fairness is the wrong first question to ask of an AI résumé screener: an apparent absence of demographic bias can be a symptom of a model that is not actually evaluating candidates. The author introduces the Illusion of Neutrality for this case and argues that bias audits alone are insufficient, proposing instead a dual-validation framework that checks both demographic bias and evaluative competence. Two experiments across eight commercial AI platforms (13 model versions) support the claim: Experiment 1 finds context-dependent, intersectional racial and gender bias, with some models penalizing a résumé simply for carrying a name, while Experiment 2 finds models that rate matched and mismatched résumés almost identically, including one with a discernment score of zero that gave a keyword-stuffed service-industry résumé 92 out of 100 for a Finance Director role.

What carries the argument

The machinery that carries the argument is a two-part audit built on one quantitative instrument: the share of rating variance explained by whether the résumé matches the target role, measured by $\Omega$ Squared ($\omega^2$). Experiment 1 uses matched résumés with name-based demographic signals against a name-redacted control to expose bias; Experiment 2 reuses the same redacted résumés in mismatched role pairings to measure discernment, and adds a keyword-stuffing test in which an irrelevant service-industry résumé is padded with financial jargon. The named concept these tests expose, the Illusion of Neutrality, is the fourth cell of a 2x2 competence-by-bias matrix: the deceptive profile of low competence with low measured bias, which looks fair only because its outputs do not respond meaningfully to candidate quality.

What would settle it

Have an independent panel of experienced recruiters score the same three résumés against the same three roles, and use those scores as ground truth to recompute the $\omega^2$ discernment rankings; in the same study, force the low-discernment models to rank candidates comparatively ('which of these two applicants is more qualified?') rather than assign numeric scores. If the human-expert baselines fail to reproduce ChatGPT-4o's ordering, or if forced-choice ranking shows that models like Grok-fast can order candidates correctly, the Illusion of Neutrality verdict would be an artifact of a self-referential baseline or a compressed scoring scale rather than a genuine failure to evaluate.

Watch

Extended reading notes

Core claim

The paper's central claim is that a low measured-bias score is not evidence of fairness unless the model can demonstrably perform the evaluation task. The argument rests on a two-part audit of eight commercial AI platforms, thirteen model versions in total. Experiment 1 compares matched fictitious résumés that differ only in name and finds bias in contradictory, intersectional, and context-dependent forms: some models penalize every named candidate relative to the redacted control, others reward particular race-gender groups, with effect sizes up to $|d| = 6.9$. Experiment 2 shows that several models cannot distinguish a well-matched résumé from a mismatched one — ChatGPT-fast reaches $\omega^2 = 0.79$ discernment while Grok-fast scores $\omega^2 = 0.00$ — and that keyword stuffing fools the least discerning models while the most discerning ones reject both an irrelevant and a keyword-stuffed résumé. The Illusion of Neutrality names the trap: a model such as Grok-fast appears unbiased precisely because it assigns nearly the same high score to every candidate, so an audit that measures only bias will certify a tool that cannot do its job.

Load-bearing premise

The competence findings assume the three test résumés really are highly qualified, well qualified, and underqualified for their roles, but that ground truth was set by ChatGPT-4o's own baseline scores (92, 83, and 66) rather than by an independent panel of human experts.

Editorial extensions

If this is right

  • A bias audit that returns a clean bill of health can still certify an unusable tool: Grok-fast scored zero discernment ($\omega^2 = 0.00$) between a matched and a mismatched résumé and gave an irrelevant keyword-stuffed résumé 92 out of 100.
  • Fairness metrics and competence tests must be run together; the dual-validation framework treats 'no measurable bias' as meaningful only when the model also demonstrably separates qualified from unqualified candidates.
  • Purchasers and deployers of AI screening tools should run minimum-functionality checks, such as mismatched-résumé and keyword-stuffed-résumé probes, before trusting a vendor's fairness claims.
  • Regulations like New York's Local Law 144 and the EU AI Act should be extended to require competence validation alongside bias auditing, since bias-only mandates cannot catch the Illusion of Neutrality.
  • Provider 'fast' and 'slow' model versions can differ by up to nearly 30 rating points on the same résumé, so deployment decisions must specify and audit the exact model version in use.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The Illusion of Neutrality is a general failure mode: in any domain where a model's scores barely move across very different inputs — loan underwriting, insurance pricing, medical triage, criminal risk assessment — an audit that checks only group-level fairness can certify a tool that is not doing its job; the paper names these domains but does not test them.
  • A flat rating pattern like Grok-fast's (a 92 for nearly every résumé) is observationally the same as a compressed scoring scale; a forced-choice test in which the model must rank candidates rather than score them would separate 'no evaluation' from 'evaluation on a narrow scale,' and if the rankings turn out correct, the incompetence verdict would need softening.
  • Because the baseline qualification levels were assigned by ChatGPT-4o itself, the whole competence ranking inherits that model's judgment; recomputing the $\omega^2$ scores against a human-expert rubric is the direct robustness check, and a different ranking would narrow the conclusion to competence as judged by one particular model.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper reports a two-part audit of eight public AI platforms (13 fast/slow model variants) used for resume screening. Experiment 1 uses name variants of three fictitious resumes to measure racial, gender, and intersectional rating bias against a name-redacted control. Experiment 2 measures evaluative competence by comparing ratings for correctly matched versus mismatched resume-role pairs (2a) and by comparing a well-written irrelevant resume versus a keyword-stuffed irrelevant resume (2b). The paper's central claim is that an apparently unbiased model may be incompetent, producing an 'Illusion of Neutrality' that traditional bias audits miss, and it proposes a dual-validation framework requiring both bias and competence audits for responsible deployment.

Significance. If the findings hold, the paper makes a valuable conceptual contribution: bias audits alone can mislead when the tool under audit lacks the competence to evaluate candidates. The Illusion of Neutrality is a useful, memorable construct with clear practical implications for regulators and HR teams, and the paper's dual-validation 2x2 framework (Section 13) is a sensible policy proposal. The study is also commendable for using publicly available tools, simulating a realistic hiring workflow, and releasing both data and materials via Zenodo (Section 15). The main empirical strengths are the large effect sizes reported (e.g., Grok-fast omega-squared = 0.99 in Experiment 2b) and the internally consistent pattern that models with zero discernment in 2a also show zero bias in Experiment 1. However, the significance is tempered by the internal-validity concerns detailed below, which currently leave the cleanest demonstration of the keyword-stuffing mechanism unestablished.

major comments (4)
  1. [3.2, 8] Section 3.2 sets the qualification levels of the three test resumes using ChatGPT-4o pre-screening scores (92, 83, 66), and the same model family is then audited for competence; this circularity is load-bearing because the 'correct match' versus 'mismatch' classification in Experiment 2a is assumed rather than independently validated. A human expert panel (or an explicit rubric applied by multiple raters) should independently label the resumes, and the discernment analyses should be re-run or re-interpreted under that validation. With only three resumes, the omega-squared estimates in Figure 2 also lack confidence intervals, so the ranking's stability is unknown.
  2. [11, 12] The keyword-stuffing demonstration is confounded. The 'irrelevant, keyword-stuffed' resume (Appendix A, Item 6) adds content beyond keyword substitution: Grok-fast's own justification cites 'mentorship from a former PepsiCo finance director' and 'frequent interactions with finance professionals, including PepsiCo executives' as significant positives. A model, or a human reviewer, could reasonably treat those as credential-like signals, so the 77-point gap between the two resumes cannot be attributed specifically to keywords without context. The authors need matched stimuli in which the only manipulated feature is the presence of financial jargon in nonsensical contexts, with all credential and contact details held constant. In addition, Section 11 does not state the number of submissions per model; without that, the omega-squared values (Section 12) and their precision cannot be assessed.
  3. [3.5, Appendix A] The reported design is internally inconsistent. Section 3.5 says 'six first names ... per group' and '18 unique name variations per resume type,' but Appendix A, Item 4 lists 24 first names (eight per racial group), and the ANOVA degrees of freedom in Appendix B, Item 1 (df = 2, 231) imply 78 ratings per model per job, which is consistent with 24 names times 3 submissions plus 6 control submissions, not with 18 name variations. Please correct the text and provide a per-cell sample-size table.
  4. [5, Appendix B, Item 8] The correlations in Section 5 between a job's perceived femaleness or Whiteness and model bias are based on demographic estimates produced by the same models being audited (Appendix B, Item 8). This is circular if the claim is about properties of the occupations themselves. Use external occupational demographic data (e.g., BLS statistics) or explicitly reframe the analysis as a within-model association between a model's own stereotype and its bias pattern.
minor comments (6)
  1. [3.3] The phrase '100% accuracy' is imprecise; the models agreed with the census-derived labels, but the validation was performed by the same models under test, not by human raters. Please reword and clarify what 'accuracy' means here.
  2. [Figure 1] The heatmap is clamped at -15 to +5, which visually truncates the very large negative effects discussed in the text (e.g., Gemini-fast penalties around -20). Consider re-plotting without clamping or adding a note about the truncation.
  3. [8] The formula for omega-squared is not given; please provide an explicit expression or a precise citation so the reported values can be reproduced.
  4. [4.2] Define the coefficient of variation (CV) formula and state why the 25% threshold from psychometric standards [12] is the relevant benchmark.
  5. [15] The data and stimuli DOI is a strength; consider adding model version identifiers (e.g., specific checkpoint or 'as of June 2025') to the main text and the archived materials, since free-tier model versions change frequently.
  6. [2.3] Section 2.3 cites [23] for hallucination, but that reference addresses self-reflection; a more standard hallucination reference would be easier for readers to verify.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the Illusion of Neutrality claim rests on independent, face-valid audit observations; the ChatGPT-4o pre-screen is confirmatory, not definitional, and the contested femaleness correlation is auxiliary.

full rationale

The paper's central claim is an empirical existence claim derived from direct model outputs rather than from equations that encode the conclusion. The qualification labels in Section 3.2 are introduced as design intent ('designed to represent high, moderate, and low fits'), with the ChatGPT-4o pre-screening offered only as confirmation; the key competence demonstrations, such as Grok-fast rating a call-center/graphic-design resume 92 for an SVP Fraud role and its zero discernment score in Experiment 2a, do not depend on that pre-screen and would survive its removal. The Section 5 correlation between perceived job 'femaleness' and pro-female bias is an auxiliary internal-consistency observation; even if the demographic estimates come from the audited models themselves, the correlation is not asserted as an external ground truth and is not load-bearing for the Illusion of Neutrality concept. The paper uses no fitted parameters renamed as predictions, no load-bearing self-citations, and no uniqueness or ansatz imported from the authors' prior work. The real weaknesses are methodological rather than circular: Experiment 2b's two stimuli differ in more than keywords, since the keyword-stuffed resume adds PepsiCo-mentorship and executive-interaction content that the well-written control lacks, and the limitations section (5.1) does not flag this stimulus mismatch or the ChatGPT-4o ground-truth dependence. These concerns reduce evidential strength, but they do not make the derivation equivalent to its inputs by construction.

Assumptions & free parameters 0 free parameters · 5 assumptions · 1 invented entities

The ledger is small because this is an empirical audit rather than a derivation. The key non-empirical load-bearing items are the ChatGPT-4o competence labels, the redacted-resume baseline assumption, and the census-based name mapping. No numerical free parameters are fitted; the discernment category labels are arbitrary cutoffs.

assumptions (5)
  • domain assumption Name-to-demographic mapping: census-derived name statistics (Appendix A Item 4) are assumed to be the demographic signals the AI models perceive.
    The whole bias design relies on names signalling race and gender to the models, validated by asking the models themselves.
  • domain assumption The redacted resume is a neutral control.
    Differences between named and redacted resumes are interpreted as bias, but a redacted resume may simply be a lower-information input.
  • ad hoc to paper ChatGPT-4o's baseline scores define the true qualification levels of the three resumes.
    Competence judgments in Experiment 2 are evaluated against this LLM-generated ordering, not against independent human expert ratings.
  • standard math Numeric 1-100 ratings are treated as interval-scale measurements across models.
    ANOVA and Cohen's d are applied to Likert-like ratings; this is a conventional but not guaranteed assumption.
  • domain assumption The three resume templates are representative of the white-collar qualification spectrum.
    Generalizations about 'competence' of platforms rest on three hand-written resumes.
invented entities (1)
  • Illusion of Neutrality independent evidence
    purpose: Describes and names the pattern where an AI appears demographically unbiased because it cannot meaningfully evaluate candidates.
    It yields a falsifiable handle: models with near-zero measured bias should show low discernment scores and be susceptible to keyword manipulation, testable on any platform. The current paper provides one strong instance (Grok-fast) but no external validation yet.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening." pith.science (2026). https://pith.science/paper/DJAKLGPL

@misc{pith2026250711548,
  author       = {Pith},
  title        = {Pith review of: Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DJAKLGPL}},
  note         = {Machine review of arXiv:2507.11548}
}
read the original abstract

The use of publicly available generative AI systems for resume evaluation is often justified by the assumption that these tools reduce bias relative to human judgment. However, this framing leaves a prior question unresolved: whether these systems are capable of performing the evaluative task at all. This study presents a two-part audit of eight widely used AI platforms used for resume screening. Drawing on the concept of the Illusion of Neutrality, the study examines cases in which systems appear demographically unbiased because they lack the ability to meaningfully differentiate among candidates. Experiment 1 evaluates racial and gender bias using matched fictitious resumes and finds that bias persists in context-dependent and intersectional forms. Some models penalize candidates for the presence of demographic signals, while others exhibit inconsistent patterns across roles and identities under controlled conditions. Experiment 2 evaluates task competence using resume-role mismatch and keyword-manipulation tests and finds that several models that appear relatively unbiased fail to distinguish relevant from irrelevant candidate experience in relation to the target role. In these cases, outputs appear stable or neutral not because the systems are fair, but because they fail to perform meaningful evaluation. Fairness assessments alone are therefore insufficient. The paper proposes a dual-validation framework that requires auditing for both demographic bias and evaluative competence as a minimum condition for responsible deployment.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

50 extracted references · 29 canonical work pages

  1. [1]

    The Question of Bias (Experiment 1): To what extent do current generative AI models show racial and gender bias when evaluating identical résumés?

  2. [2]

    fast” and more advanced “thinking/slow

    The Question of Competence (Experiment 2): Are these AI models sufficiently competent to differentiate between qualified, mismatched, and unqualified résumés? This experiment investigates the Illusion of Neutrality, where an apparent lack of bias may simply be a symptom of a model’ s inability to perform the task. 3 EXPERIMENT 1 METHOD 3.1 AI Models Evalu...

  3. [3]

    Name Redacted

    Underqualified (“Fraud”): This résumé described a candidate with a bachelor’s degree in Graphic Design who had advanced through call center roles to a Team Manager position in banking. The candidate was applying for the Senior Vice President role in American Express’s fraud divisi on (see Supplemental Materials, Appendix A, Item 3). 3.3 Demographic Variab...

  4. [4]

    The candidate was applying for the Director of Finance position in PepsiCo’s Snacks Divi sion (see Supplemental Materials, Appendix A, Item 1)

    Highly Qualified (“Finance”): This résumé described a candidate with both a bachelor's and master’s degree in Business and Finance, along with a clear progression to a Senior Finance Manager role. The candidate was applying for the Director of Finance position in PepsiCo’s Snacks Divi sion (see Supplemental Materials, Appendix A, Item 1)

  5. [5]

    The candidate was applying for an HR Manager role at a Fortune 500 company (see Supplemental Materials, Appendix A, Item 2)

    Well Qualified (“HR”): This résumé described a candidate with an online bachelor’s degree in Human Resources and a clear trajectory toward an HR Manager position at a retail location. The candidate was applying for an HR Manager role at a Fortune 500 company (see Supplemental Materials, Appendix A, Item 2)

  6. [6]

    Joy Buolamwini and Timnit Gebru. 2018. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Proceedings of the 1st conference on fairness, accountability and transparency (Proceedings of machine learning research), February 23,

  7. [7]

    Chan and Walter J

    Michael W. Chan and Walter J. Eppich. 2020. The Keyword Effect: A Grounded Theory Study Exploring the Role of Keywords in Clinical Communication. AEM Education and Training 4, 4 (October 2020), 403–410. https://doi.org/10.1002/aet2.10424

  8. [8]

    Chinyere Linda Agbasiere and Goodness Rex Nze-Igwe. 2025. Algorithmic Fairness in Recruitment: Designing AI-Powered Hiring Tools to Identify and Reduce Biases in Candidate Selection. PoS 11, 4 (April 2025), 5001. https://doi.org/10.22178/pos.116-10

Show all 50 references
  1. [9]

    Jiafu An, Difang Huang, Chen Lin, and Mingzhu Tai. 2025. Measuring gender and racial biases in large language models: Intersectional evidence from automated resume evaluation. PNAS Nexus 4, 3 (February 2025). https://doi.org/10.1093/pnasnexus/pgaf089

  2. [10]

    Eitan Anzenberg, Arunava Samajpati, Sivasankaran Chandrasekar, and Varun Kacholia. 2025. Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions. https://doi.org/10.48550/arXiv.2507.02087

  3. [11]

    Marianne Bertrand and Sendhil Mullainathan. 2004. Are Emily and Greg More Employable Than Lakisha and Jamal? A Field Experiment on Labor Market Discrimination. American Economic Review 94, 4 (September 2004), 991–1013. https://doi.org/10.1257/0002828042002561

  4. [12]

    Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016. Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings. https://doi.org/10.48550/ARXIV.1607.06520

  5. [13]

    2025 Job Seeker Nation Report

    Employ. 2025 Job Seeker Nation Report. Employ. Retrieved July 7, 2025 from https://pages.employinc.com/Employ-2025-Job-Seeker- Nation-Report.html

  6. [14]

    Kawin Ethayarajh and Dan Jurafsky. 2020. Utility is in the Eye of the User: A Critique of NLP Leaderboards. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020. Association for Computational Linguistics, Online. https://doi.o...

  7. [15]

    Carmine Ferrara, Giulia Sellitto, Filomena Ferrucci, Fabio Palomba, and Andrea De Lucia. 2024. Fairness-aware machine learning engineering: how far are we? Empir Software Eng 29, 1 (January 2024). https://doi.org/10.1007/s10664-023-10402-y

  8. [16]

    Zhisheng Chen. 2023. Ethics and discrimination in artificial intelligence-enabled recruitment practices. Humanit Soc Sci Commun 10, 1 (September 2023). https://doi.org/10.1057/s41599-023-02079-x

  9. [17]

    Correll, Stephen Benard, and In Paik

    Shelley J. Correll, Stephen Benard, and In Paik. 2007. Getting a Job: Is There a Motherhood Penalty? American Journal of Sociology 112, 5 (March 2007), 1297–1339. https://doi.org/10.1086/511799

  10. [18]

    Insight - Amazon scraps secret AI recruiting tool that showed bias against women | Reuters

    Jeffrey Dastin. Insight - Amazon scraps secret AI recruiting tool that showed bias against women | Reuters. Retrieved July 7, 2025 from https://www.reuters.com/article/us-amazon-com-jobs-automation-insight-idUSKCN1MK08G/

  11. [19]

    Robert A. Divine. 1962. The Illusion of Neutrality. The Mississippi Valley Historical Review 49, 3 (December 1962), 534. https://doi.org/10.2307/1902595

  12. [20]

    Downing and Thomas M

    Steven M. Downing and Thomas M. Haladyna (Eds.). 2006. Handbook of test development. L. Erlbaum, Mahwah, N.J

  13. [21]

    These methods aim to correct for bias by re-weighting data, adding fairness constraints to algorithms, or adjusting model outputs

    confirms are broadly categorized as pre-processing, in-processing, and post-processing techniques [1]. These methods aim to correct for bias by re-weighting data, adding fairness constraints to algorithms, or adjusting model outputs. Recent comparative studies highlight the po...

  14. [22]

    Dan Hendrycks and Kevin Gimpel. 2018. A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks. https://doi.org/10.48550/arXiv.1610.02136

  15. [23]

    Ziwei Ji, Tiezheng Yu, Yan Xu, Nayeon Lee, Etsuko Ishii, and Pascale Fung. 2023. Towards mitigating LLM hallucination via self reflection. In Findings of the association for computational linguistics: EMNLP 2023, December 2023. Association for Computational Linguistics, Singap...

  16. [24]

    Fuller, Manjari Raman, Eva Sage-Gavin, and Kristen Hines

    Joseph B. Fuller, Manjari Raman, Eva Sage-Gavin, and Kristen Hines. Hidden Workers: Untapped Talent

  17. [25]

    Thinking

    Shaz Furniturewala, Surgan Jandial, Abhinav Java, Pragyan Banerjee, Simra Shahid, Sumit Bhatia, and Kokil Jaidka. 2024. “Thinking” Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models. In Proceedings of the 2024 Conference on Empirical Methods in ...

  18. [26]

    Gaebler, Sharad Goel, Aziz Huq, and Prasanna Tambe

    Johann D. Gaebler, Sharad Goel, Aziz Huq, and Prasanna Tambe. 2024. Auditing the Use of Language Models to Guide Hiring Decisions. https://doi.org/10.48550/arXiv.2404.03086

  19. [27]

    Nikhil Garg, Londa Schiebinger, Dan Jurafsky, and James Zou. 2018. Word embeddings quantify 100 years of gender and ethnic stereotypes. Proc. Natl. Acad. Sci. U.S.A. 115, 16 (April 2018). https://doi.org/10.1073/pnas.1720347115

  20. [28]

    The Myth in the Methodology: Towards a Recontextualization of Fairness in Machine Learning

    Ben Green and Lily Hu. The Myth in the Methodology: Towards a Recontextualization of Fairness in Machine Learning

  21. [29]

    Yufei Guo, Muzhe Guo, Juntao Su, Zhou Yang, Mengqiu Zhu, Hongfei Li, Mengyang Qiu, and Shuo Shuo Liu. 2024. Bias in Large Language Models: Origin, Evaluation, and Mitigation. https://doi.org/10.48550/arXiv.2411.10915

  22. [30]

    Midtbøen

    Lincoln Quillian, Devah Pager, Ole Hexel, and Arnfinn H. Midtbøen. 2017. Meta-analysis of field experiments shows no change in racial discrimination in hiring over time. Proc. Natl. Acad. Sci. U.S.A. 114, 41 (October 2017), 10870–10875. https://doi.org/10.1073/pnas.1706255114 14

  23. [31]

    Manish Raghavan, Solon Barocas, Jon Kleinberg, and Karen Levy. 2020. Mitigating bias in algorithmic hiring: evaluating claims and practices. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, January 27, 2020. ACM, Barcelona Spain, 469–481. ht...

  24. [32]

    Kusner, Joshua R

    Matt J. Kusner, Joshua R. Loftus, Chris Russell, and Ricardo Silva. 2017. Counterfactual Fairness. https://doi.org/10.48550/ARXIV.1703.06856

  25. [33]

    Something Fast and Cheap

    Tina B. Lassiter and Kenneth R. Fleischmann. 2024. “Something Fast and Cheap” or “A Core Element of Building Trust”? - AI Auditing Professionals’ Perspectives on Trust in AI. Proc. ACM Hum.-Comput. Interact. 8, CSCW2 (November 2024), 1–22. https://doi.org/10.1145/3686963

  26. [34]

    The Future of Recruiting 2025 | LinkedIn

    LinkedIn. The Future of Recruiting 2025 | LinkedIn. Retrieved July 7, 2025 from https://business.linkedin.com/talent- solutions/resources/future-of-recruiting

  27. [35]

    Timothy Niven and Hung-Yu Kao. 2019. Probing Neural Network Comprehension of Natural Language Arguments. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019. Association for Computational Linguistics, Florence, Italy. https://doi.or...

  28. [36]

    Lin, and Lee Ross

    Emily Pronin, Daniel Y. Lin, and Lee Ross. 2002. The Bias Blind Spot: Perceptions of Bias in Self Versus Others. Pers Soc Psychol Bull 28, 3 (March 2002), 369–381. https://doi.org/10.1177/0146167202286008

  29. [37]

    Natasha Quadlin. 2018. The Mark of a Woman’s Record: Gender and Academic Performance in Hiring. Am Sociol Rev 83, 2 (April 2018), 331–360. https://doi.org/10.1177/0003122418762291

  30. [38]

    Pauleen, and Nazim Taskin

    Melika Soleimani, Ali Intezari, James Arrowsmith, David J. Pauleen, and Nazim Taskin. 2025. Reducing AI bias in recruitment and selection: an integrative grounded approach. The International Journal of Human Resource Management (March 2025), 1–36. https://doi.org/10.1080/09585...

  31. [39]

    Bryan Chen Zhengyu Tan and Roy Ka-Wei Lee. 2025. Unmasking Implicit Bias: Evaluating Persona-Prompted LLM Responses in Power- Disparate Social Scenarios. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguisti...

  32. [40]

    Half of Managers Use AI To Determine Who Gets Promoted and Fired

    ResumeBuilder. Half of Managers Use AI To Determine Who Gets Promoted and Fired. ResumeBuilder.com. Retrieved July 7, 2025 from https://www.resumebuilder.com/half-of-managers-use-ai-to-determine-who-gets-promoted-and-fired/

  33. [41]

    Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond Accuracy: Behavioral Testing of NLP Models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020. Association for Computational Lingui...

  34. [42]

    Help Wanted

    Aaron Rieke and Miranda Bogen. Help Wanted. Upturn. Retrieved July 7, 2025 from https://upturn.org/work/help-wanted/

  35. [43]

    Alejandro Salinas, Amit Haim, and Julian Nyarko. 2024. What’s in a Name? Auditing Large Language Models for Race and Gender Bias. https://doi.org/10.48550/ARXIV.2402.14875

  36. [44]

    Selbst, Danah Boyd, Sorelle A

    Andrew D. Selbst, Danah Boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi. 2019. Fairness and Abstraction in Sociotechnical Systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency, January 29, 2019. ACM, Atlanta GA USA, 59–...

  37. [45]

    Kamal Shah, Manish Rana, and Trusha Pimple. 2025. Fair and transparent AI-driven resume screening: Enhancing recruitment with bias- aware machine learning. South Eastern European Journal of Public Health (February 2025), 2346–2361. https://doi.org/10.70135/seejph.vi.4674

  38. [48]

    András Tilcsik. 2011. Pride and Prejudice: Employment Discrimination against Openly Gay Men in the United States. American Journal of Sociology 117, 2 (September 2011), 586–626. https://doi.org/10.1086/661653

  39. [49]

    Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C. Schmidt. 2023. A prompt pattern catalog to enhance prompt engineering with ChatGPT. In Proceedings of the 30th conference on pattern languages o...

  40. [50]

    Finance” résumé (Highly Qualified) 16 Item 2: “HR

    Kyra Wilson and Aylin Caliskan. 2024. Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval. AIES 7, (October 2024), 1578–1590. https://doi.org/10.1609/aies.v7i1.31748 17 REFERENCES IN APPENDICES [b1] US Census Bureau. Frequently Occurring Surn...

  41. [100]

    Illusion of Neutrality,

    Because it cannot grasp the meaning behind the data, the process is fundamentally broken. This reve als a critical blind spot for the industry. For example, while the audit by Gaebler [18] provided valuable insights into LLM bias, the authors explicitly noted that assessing th...

  42. [2018]

    Retrieved from https://proceedings.mlr.press/v81/buolamwini18a.html

    PMLR, 77–91. Retrieved from https://proceedings.mlr.press/v81/buolamwini18a.html

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.