Pith. sign in

REVIEW 3 major objections 7 minor 74 references

From Expectation to Habit: Why Do Software Practitioners Adopt Fairness Toolkits?

T0 review · 3 major / 7 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper claims that software practitioners adopt fairness toolkits chiefly because they expect the tools to improve bias mitigation and because using them has become habitual, with performance expectancy the strongest driver of…

desk verdict A competent UTAUT2 application to fairness toolkit adoption whose headline habit finding may be partly tautological, and whose manuscript shows careless copy-paste slips. read the letter →

arxiv 2412.13846 v2 pith:GH7B6O54 submitted 2024-12-18 cs.SE cs.AI

classification cs.SEcs.AI
keywords fairnesstoolkitstechnologyadoptionUTAUT2PLS-SEMhabitperformanceexpectancymachinelearningsoftwarepractitioners
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper set out to explain why software practitioners actually use fairness toolkits, which are libraries and metrics designed to detect and mitigate bias in machine-learning models. It surveyed 181 industry practitioners and analyzed their responses with a structural equation model based on the Unified Theory of Acceptance and Use of Technology (UTAUT2). The central finding is that two individual factors dominate: performance expectancy, or the belief that the toolkit will help produce fairer software, is the strongest predictor of intention to adopt, while habit is the strongest predictor of actual use behavior. Social influence, effort expectancy, facilitating conditions, and enjoyment were not significant in this sample. If the finding holds, organizations and tool vendors should promote adoption less through mandates or social pressure and more by demonstrating concrete bias-mitigation results and by embedding toolkit use into routine workflows.

What carries the argument

The central machinery is the UTAUT2 model, a validated survey-based framework for technology adoption that explains intention and use through constructs such as performance expectancy, effort expectancy, social influence, facilitating conditions, hedonic motivation, and habit. The study drops price value, excludes the standard demographic moderators, and links the remaining latent constructs to self-reported use via partial least squares structural equation modeling, a regression-based path analysis suited to small and non-normal samples. The habit construct, defined as the extent to which behavior becomes automatic through learning, carries the paper's strongest result because its direct path to use behavior dominates all other paths in the model.

What would settle it

A follow-up measurement study in which habit is assessed with an automaticity scale rather than a frequency scale, checked against the same UTAUT2 paths, would settle the question: if the habit-to-use coefficient drops far below 0.543 or becomes non-significant, the central claim that habit formation drives adoption is not supported.

Watch

Extended reading notes

Core claim

The study reports that performance expectancy and habit are the primary drivers of fairness toolkit adoption. In the estimated model, performance expectancy has the largest path to behavioral intention (coefficient 0.465), habit has the second-largest path to intention (0.323), and habit has by far the largest path to actual use behavior (0.543), with behavioral intention itself contributing a smaller path to use (0.161). The model explains 63 percent of the variance in intention and 40.7 percent of the variance in use. A mediation analysis indicated that habit's effect on use is direct rather than mediated by intention. The authors interpret this as evidence that practitioners decide to adopt fairness toolkits based on perceived effectiveness at mitigating bias, and that sustained usage comes from the tools becoming part of their regular working routines.

Load-bearing premise

The load-bearing premise is that the survey's habit items measure an automatic tendency built through learning, distinct from how often practitioners say they use the toolkit; if the items mostly capture past frequency of use, the strong habit-to-use path largely restates that past use predicts current use.

Editorial extensions

If this is right

  • If the paper is right, awareness campaigns for fairness toolkits should emphasize demonstrated bias-mitigation effectiveness and real-world success cases, since performance expectancy is the strongest driver of intention.
  • Tool vendors and organizations should focus on making toolkit use easy to embed in daily work, for example through well-designed APIs and integration into pipelines, because habit is the strongest predictor of actual use.
  • The non-significant paths for social influence, effort expectancy, facilitating conditions, and hedonic motivation imply that organizational support, peer pressure, ease of use, and enjoyment are not reliable levers for increasing adoption in this population.
  • The variance explained suggests that individual perceptions and routines account for a large share of adoption behavior, but that roughly 37 percent of intention and 59 percent of use remain unexplained by the factors studied.
  • The direct, unmediated habit-to-use path implies that forming a routine is at least as important for real adoption as strengthening a practitioner's conscious intention to use the toolkit.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural testable extension is whether habit-forming interventions, such as scheduled fairness checks in regression pipelines, actually increase adoption over time; the paper's cross-sectional design cannot establish that habitual use causes adoption, so longitudinal or experimental studies would sharpen the claim.
  • The null results for effort expectancy and social influence may not generalize to less experienced practitioners or to organizations where fairness toolkits are mandatory, because the sample skews toward experienced developers and self-reported professional contexts.
  • The strong habit-to-use path may partly reflect measurement overlap: if the habit items ask mainly about past frequency of use, then part of that coefficient restates that past use predicts current use rather than showing an automatic tendency that managers can deliberately cultivate.
  • A useful replication would measure habit with an automaticity-focused scale rather than a frequency-based scale; if habit's path to use drops substantially under that measurement, the practical advice to 'build habits' would need to be reframed around workflow redesign.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. This paper reports a survey-based study of 181 software practitioners recruited through Prolific, using the UTAUT2 framework and PLS-SEM to explain behavioral intention (BI) and actual use behavior (UB) regarding fairness toolkits. The model includes performance expectancy, effort expectancy, social influence, hedonic motivation, facilitating conditions, and habit as predictors. The main findings are that performance expectancy has the strongest effect on BI (path 0.465) and habit has the strongest effect on UB (path 0.543), with BI also significantly predicting UB. The authors conclude that practitioners adopt fairness toolkits primarily for performance reasons and that habitual use drives sustained adoption, leading to recommendations for organizations and tool vendors to integrate fairness toolkits into routine workflows.

Significance. If the findings hold, the paper would provide one of the first quantitative, theory-grounded accounts of fairness toolkit adoption, complementing existing qualitative work and offering concrete design and organizational implications. The study has methodological strengths: an a priori power analysis, iterative pilot testing, attention checks, reliability and validity reporting (Cronbach's alpha, rho_c, rho_A, AVE), bootstrapping with 10,000 subsamples, and PLSpredict benchmarking. However, the paper's central empirical claim—that habit is the primary driver of actual use—depends on the discriminant validity between the habit construct and the single-item frequency measure of use behavior, and the evidence for that discriminant validity is not currently verifiable because the item wordings are omitted and the online appendix reference is a placeholder. This gap must be resolved before the practical recommendations about cultivating habit can be accepted.

major comments (3)
  1. [Section III.B, Section IV.C.1, Table II] The strongest path in the model, HB→UB (0.543, Table II), is the central empirical result, but the paper does not report the wording of the four HB items or the single UB frequency item, and the cited online appendix [53] is listed as 'A. Authors' with no accessible content. The hypothesis text defines habit partly as 'use them regularly' (Section III.B) and H6b states that 'practitioners who consistently integrate new tools are more inclined to use fairness toolkits consistently.' If the HB items ask about regularity, frequency, or amount of use, then the HB→UB path is partly tautological in the sense that past use predicts current use. Please report the full item wordings, provide the UB item, and show that the HB items capture automaticity rather than mere frequency, for example by reporting a factor analysis that separates HB from UB and by re-estimating the model without any frequency-based HB indicator.
  2. [Section V.A, 'Discriminant Validity'] The claim that habit and use behavior are empirically distinct rests entirely on HTMT values that are only mentioned as being in the online appendix; the appendix is not verifiable because reference [53] is a placeholder with no author and no accessible link. Given that UB is a single-item frequency scale and HB items are not shown, the reported HTMT results are insufficient to rule out the tautology concern. Provide the full HTMT matrix (or at least the HB–UB value), and ideally a robustness analysis that treats UB and HB as indicators of a single factor to show they are not empirically indistinguishable.
  3. [Section VI.B and Section VIII] The practical recommendations state that organizations should 'help employees develop a habit' and that habitual use can lead to higher adoption. These causal prescriptions go beyond what a cross-sectional, self-report survey can support; the data establish associations only. The paper should either add an explicit limitation acknowledging that the direction of causality is not tested, or soften the causal language throughout the discussion and conclusion. This is directly relevant to the central claim because the practical value of the habit construct depends on whether habit can be deliberately cultivated, which the present design cannot establish.
minor comments (7)
  1. [Section V.B, 'Explanatory Power'] The text reports R² values for Behavioral Intention and Use Behavior but then says 'the model successfully explains 63% of the variance in the intention to use large language models (LLMs).' This appears to be a copy-paste error; it should refer to fairness toolkits.
  2. [Reference [53]] The reference for the online appendix is incomplete: it lists 'A. Authors' and a figshare URL but no author names, title, or year-specific identifier. Please provide a complete citation and ensure the replication materials are accessible.
  3. [Table II] For HM→BI the reported p-value is 0.94 with a negative path coefficient; this is likely a one-tailed probability for the wrong direction. Clarify whether one-tailed or two-tailed tests are used for structural paths and report the corresponding p-values consistently.
  4. [Section IV.B] The survey title is written as '["Fairness in Software Development] Pre-Screening — Fairness Toolkit Adoption' with mismatched quotation marks; please correct the formatting.
  5. [Section VII, Internal Validity] The internal-validity discussion does not address common method bias, which is a standard concern for single-source, self-report surveys. Consider adding a Harman's single-factor test or a marker-variable analysis, or at least discuss why this is not a concern here.
  6. [Section III.A] The claim that age, gender, and experience moderators were not significant is supported only by a reference to the online appendix. Report the relevant test statistics or model comparison in the main text to make this exclusion auditable.
  7. [Section V.A, 'Indicator Reliability'] The paper states that EE4 and HB2 had outer loadings below 0.70 but provides no loading values. Please include the complete loading table or add the values in the appendix rather than referring only to an inaccessible source.

Circularity Check

1 steps flagged · score 3.0 of 10

The strongest path (HB→UB = 0.543) risks being self-definitional: the paper's own H6b rationale defines habit as regular/consistent use while UB is a single-item frequency scale, and the item wordings that would disprove the tautology are absent.

  1. self definitional [Section III.B (H6b); Section IV.C.1; Table II]
    "Practitioners are more likely to adopt fairness toolkits if they are familiar with certain tools and use them regularly. Furthermore, habit influences actual usage behavior, as practitioners who consistently integrate new tools are more inclined to use fairness toolkits consistently in their development activities. ... The dependent variable, Use Behavior (UB), was measured using a single-item frequency scale."

    The paper's central claim is that Habit is the primary driver of actual use (HB→UB = 0.543, f² = 0.150, Table II). Its own H6b rationale defines habit as regular/consistent use ("use them regularly," "consistently integrate new tools"), and UB is measured with a single-item frequency scale (Section IV.C.1). If the four HB items operationalize that same regularity, predictor and outcome are the same behavior by construction. The item wordings are not reported, and the replication appendix [53] is a placeholder ("A. Authors"), so the operational separation of HB from UB cannot be verified; the practical recommendation to "help employees develop a habit" would then reduce to encouraging regular use to obtain regular use.

full rationale

Most of the derivation chain is not circular. The study applies the externally validated UTAUT2 framework [23]; constructs are measured with instruments adapted from Venkatesh et al. (2012), and the empirical paths PE→BI (0.465) and BI→UB (0.161) are standard theoretical relationships whose estimation is data-driven, not forced by construction. PLS predict uses a proper train/holdout split against an LM benchmark, so it is not a re-fit of the same data. Self-citations exist ([20], [43], [44]) but are not load-bearing: moderators were excluded because the authors ran their own preliminary analysis (results cited to [53]), and [43]/[44] are background cataloging work. The one genuine circularity risk is internal to H6b: the hypothesis rationale verbalizes habit as "use them regularly"/"consistently integrate new tools," which is the same behavior captured by the single-item frequency scale for UB; if the unshown HB items followed that frequency framing, HB→UB = 0.543 would be partly a self-correlation and the "cultivate habit" recommendation becomes definitional. The items are not in the main text, and the online appendix reference [53] ("A. Authors, Online appendix, 2024") cited for data, item lists, HTMT values, and PLS predict results is a placeholder, so the reduction can be neither confirmed nor excluded. Score 3 reflects this flagged, load-bearing operationalization risk in the strongest path, without treating it as demonstrated circularity; the remaining findings retain independent empirical content.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The study does not introduce new theoretical entities or fitting parameters. The path coefficients and R2 values are estimated from the survey data and are the outputs of the analysis, not free inputs. The central claim rests on the validity of the UTAUT2 instruments and on self-report and sample representativeness.

assumptions (4)
  • domain assumption UTAUT2 and its measurement instruments (Venkatesh et al., 2012) are a valid and appropriate lens for studying fairness toolkit adoption.
    The entire hypothesis set is drawn from UTAUT2; the paper assumes the constructs capture the relevant drivers. Cited in Section III.
  • domain assumption Self-reported questionnaire responses accurately reflect participants' actual intentions and behaviors.
    All constructs are measured via self-report; no behavioral observations are used. The paper relies on self-report in construct validity and external validity in Section VII.
  • domain assumption The Prolific-recruited sample (181 respondents) is representative enough of software practitioners to generalize the findings.
    External validity relies on Prolific filters and the demographic distribution; the paper acknowledges that most participants are from Europe.
  • standard math The constructs are measured reflectively, justifying the use of PLS-SEM reliability and validity criteria.
    Section IV.C states all indicators are reflective, following Hair et al.'s guidelines.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Expectation to Habit: Why Do Software Practitioners Adopt Fairness Toolkits?." pith.science (2026). https://pith.science/paper/GH7B6O54

@misc{pith2026241213846,
  author       = {Pith},
  title        = {Pith review of: From Expectation to Habit: Why Do Software Practitioners Adopt Fairness Toolkits?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GH7B6O54}},
  note         = {Machine review of arXiv:2412.13846}
}
read the original abstract

As the adoption of machine learning (ML) systems continues to grow across industries, concerns about fairness and bias in these systems have taken center stage. Fairness toolkits, designed to mitigate bias in ML models, serve as critical tools for addressing these ethical concerns. However, their adoption in the context of software development remains underexplored, especially regarding the cognitive and behavioral factors driving their usage. As a deeper understanding of these factors could be pivotal in refining tool designs and promoting broader adoption, this study investigates the factors influencing the adoption of fairness toolkits from an individual perspective. Guided by the Unified Theory of Acceptance and Use of Technology (UTAUT2), we examined the factors shaping the intention to adopt and actual use of fairness toolkits. Specifically, we employed Partial Least Squares Structural Equation Modeling (PLS-SEM) to analyze data from a survey study involving practitioners in the software industry. Our findings reveal that performance expectancy and habit are the primary drivers of fairness toolkit adoption. These insights suggest that by emphasizing the effectiveness of these tools in mitigating bias and fostering habitual use, organizations can encourage wider adoption. Practical recommendations include improving toolkit usability, integrating bias mitigation processes into routine development workflows, and providing ongoing support to ensure professionals see clear benefits from regular use.

Figures

Figures reproduced from arXiv: 2412.13846 by the authors.

Figure 1
Figure 1. Overview of the UTAUT2 theoretical model with our hypothesis. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. presents an overview of our research methodology, which we elaborate upon in this section. We initiated our pro￾cess by carefully defining participant selection criteria for our survey and calculating the sample size using G*Power [58], taking into account the complexities of our theoretical model. Data (181) Data Collection UTAUT2 Theoretical Framework Main Questionnaire Data Collection Pilot Selection Criteria Pre… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

74 extracted references · 39 canonical work pages

  1. [53]

    Technology acceptance model,

    F. D. Davis, R. Bagozzi, and P. Warshaw, “Technology acceptance model,” J Manag Sci , vol. 35, no. 8, pp. 982–1003, 1989

  2. [1]

    Zhou and F

    J. Zhou and F. Chen, Human and Machine Learning . Springer, 2018

  3. [2]

    Software engineering for ai-based systems: A survey,

    S. Mart ´ınez-Fern´andez, J. Bogner, X. Franch, M. Oriol, J. Siebert, A. Trendowicz, A. M. V ollmer, and S. Wagner, “Software engineering for ai-based systems: A survey,” ACM Transactions on Software Engineering and Methodology , vol. 31, no. 2, 2022. [Online]. Available: http://dx.doi.org/10.1145/3487043

  4. [3]

    Comparative analysis of image clas- sification algorithms based on traditional machine learning and deep learning,

    P. Wang, E. Fan, and P. Wang, “Comparative analysis of image clas- sification algorithms based on traditional machine learning and deep learning,” Pattern recognition letters, vol. 141, pp. 61–67, 2021

  5. [4]

    A survey on theories and applications for self-driving cars based on deep learning methods,

    J. Ni, Y . Chen, Y . Chen, J. Zhu, D. Ali, and W. Cao, “A survey on theories and applications for self-driving cars based on deep learning methods,” Applied Sciences, vol. 10, no. 8, p. 2749, 2020

  6. [5]

    Can an algorithm hire better than a human,

    C. C. Miller, “Can an algorithm hire better than a human,” The New York Times, vol. 25, 2015

  7. [6]

    The algorithm that beats your bank manager,

    P. Olson, “The algorithm that beats your bank manager,” CNN Money March, vol. 15, 2011

  8. [7]

    A survey on bias and fairness in machine learning,

    N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A survey on bias and fairness in machine learning,” ACM Computing Surveys (CSUR), vol. 54, no. 6, pp. 1–35, 2021

Show all 74 references
  1. [8]

    Bias and unfairness in machine learning models: a systematic review on datasets, tools, fairness metrics, and identification and mitigation methods,

    T. P. Pagano, R. B. Loureiro, F. V . Lisboa, R. M. Peixoto, G. A. Guimar˜aes, G. O. Cruz, M. M. Araujo, L. L. Santos, M. A. Cruz, E. L. Oliveira et al., “Bias and unfairness in machine learning models: a systematic review on datasets, tools, fairness metrics, and identificatio...

  2. [9]

    A review on fairness in machine learning,

    D. Pessach and E. Shmueli, “A review on fairness in machine learning,” ACM Computing Surveys (CSUR) , vol. 55, no. 3, pp. 1–44, 2022

  3. [10]

    Machine learning, ethics and law,

    S. Miller, “Machine learning, ethics and law,” Australasian Journal of Information Systems, vol. 23, pp. 1–13, 2019

  4. [11]

    Software fairness,

    Y . Brun and A. Meliou, “Software fairness,” in Proceedings of the 2018 26th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineering , 2018, pp. 754–759

  5. [12]

    Facebook apologizes after a.i. puts ’primates’ label on video of black men,

    R. Mac, “Facebook apologizes after a.i. puts ’primates’ label on video of black men,” The New York Times . [Online]. Available: https://www. nytimes.com/2021/09/03/technology/facebook-ai-race-primates.html

  6. [13]

    ’gay writing’ falls foul of amazon sales ranking system,

    B. Johnson and H. Pidd, “’gay writing’ falls foul of amazon sales ranking system,” The Guardian . [Online]. Available: https: //www.theguardian.com/culture/2009/apr/13/amazon-gay-writers

  7. [14]

    Ai ethics issues in real world: Evidence from ai incident database,

    M. Wei and Z. Zhou, “Ai ethics issues in real world: Evidence from ai incident database,” 2022

  8. [15]

    Bias mitigation for machine learning classifiers: A comprehensive survey,

    M. Hort, Z. Chen, J. M. Zhang, M. Harman, and F. Sarro, “Bias mitigation for machine learning classifiers: A comprehensive survey,” ACM Journal on Responsible Computing , vol. 1, no. 2, pp. 1–52, 2024

  9. [16]

    Ai fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias,

    R. K. Bellamy, K. Dey, M. Hind, S. C. Hoffman, S. Houde, K. Kannan, P. Lohia, J. Martino, S. Mehta, A. Mojsilovi ´c et al. , “Ai fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias,” IBM Journal of Research and Development , vol. 63, no. 4/5, pp. ...

  10. [17]

    Fairlearn,

    “Fairlearn,” 2019. [Online]. Available: https://fairlearn.github.io/

  11. [18]

    Exploring how machine learning practitioners (try to) use fairness toolkits,

    W. H. Deng, M. Nagireddy, M. S. A. Lee, J. Singh, Z. S. Wu, K. Holstein, and H. Zhu, “Exploring how machine learning practitioners (try to) use fairness toolkits,” in 2022 ACM Conference on Fairness, Accountability, and Transparency , ser. FAccT ’22. ACM, Jun. 2022. [Online]. ...

  12. [19]

    The landscape and gaps in open source fairness toolkits,

    M. S. A. Lee and J. Singh, “The landscape and gaps in open source fairness toolkits,” in Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , ser. CHI ’21. New York, NY , USA: Association for Computing Machinery, 2021. [Online]. Available: https://doi...

  13. [20]

    Investigating the role of cultural values in adopting large language models for software engineering,

    S. Lambiase, G. Catolino, F. Palomba, F. Ferrucci, and D. Russo, “Investigating the role of cultural values in adopting large language models for software engineering,” arXiv preprint arXiv:2409.05055 , 2024

  14. [21]

    The extended unified theory of acceptance and use of technology (utaut2): A systematic literature review and theory evaluation,

    K. Tamilmani, N. Rana, S. Fosso Wamba, and R. Dwivedi, “The extended unified theory of acceptance and use of technology (utaut2): A systematic literature review and theory evaluation,” International Journal of Information Management , vol. 57, p. 102269, 11 2020

  15. [22]

    Where responsible ai meets reality: Practitioner perspectives on enablers for shifting organizational practices,

    B. Rakova, J. Yang, H. Cramer, and R. Chowdhury, “Where responsible ai meets reality: Practitioner perspectives on enablers for shifting organizational practices,” Proc. ACM Hum.-Comput. Interact., vol. 5, no. CSCW1, apr 2021. [Online]. Available: https://doi.org/10.1145/3449081

  16. [23]

    Consumer acceptance and use of information technology: Extending the unified theory of acceptance and use of technology,

    V . Venkatesh, J. Y . L. Thong, and X. Xu, “Consumer acceptance and use of information technology: Extending the unified theory of acceptance and use of technology,” MIS Quarterly , vol. 36, no. 1, pp. 157–178, 2012

  17. [24]

    A primer on partial least squares structural equation modeling (pls-sem),

    J. F. Hair Junior, G. T. M. Hult, C. M. Ringle, and M. Sarstedt, “A primer on partial least squares structural equation modeling (pls-sem),” 2014

  18. [25]

    Fairness per- ceptions of algorithmic decision-making: A systematic review of the empirical literature,

    C. Starke, J. Baleis, B. Keller, and F. Marcinkowski, “Fairness per- ceptions of algorithmic decision-making: A systematic review of the empirical literature,” 2021

  19. [26]

    Fairness definitions explained,

    S. Verma and J. Rubin, “Fairness definitions explained,” in 2018 ieee/acm international workshop on software fairness (fairware). IEEE, 2018, pp. 1–7

  20. [27]

    Fair enough: Searching for sufficient measures of fairness,

    S. Majumder, J. Chakraborty, G. R. Bai, K. T. Stolee, and T. Menzies, “Fair enough: Searching for sufficient measures of fairness,” ACM Trans. Softw. Eng. Methodol. , vol. 32, no. 6, sep 2023. [Online]. Available: https://doi.org/10.1145/3585006

  21. [28]

    Data augmentation for discrimination prevention and bias disambiguation,

    S. Sharma, Y . Zhang, J. M. R ´ıos Aliaga, D. Bouneffouf, V . Muthusamy, and K. R. Varshney, “Data augmentation for discrimination prevention and bias disambiguation,” in Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , 2020, pp. 358–364

  22. [29]

    Optimized pre-processing for discrimination prevention,

    F. Calmon, D. Wei, B. Vinzamuri, K. Natesan Ramamurthy, and K. R. Varshney, “Optimized pre-processing for discrimination prevention,” Advances in neural information processing systems , vol. 30, 2017

  23. [30]

    Bias in machine learning software: why? how? what to do?

    J. Chakraborty, S. Majumder, and T. Menzies, “Bias in machine learning software: why? how? what to do?” in Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2021, pp. 429–440

  24. [31]

    Mitigating unwanted biases with adversarial learning,

    B. H. Zhang, B. Lemoine, and M. Mitchell, “Mitigating unwanted biases with adversarial learning,” in Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society , 2018, pp. 335–340

  25. [32]

    Fairness-aware classifier with prejudice remover regularizer,

    T. Kamishima, S. Akaho, H. Asoh, and J. Sakuma, “Fairness-aware classifier with prejudice remover regularizer,” in Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2012, Bristol, UK, September 24-28, 2012. Proceedings, Part II

  26. [33]

    Springer, 2012, pp. 35–50

  27. [34]

    Data preprocessing techniques for classi- fication without discrimination,

    F. Kamiran and T. Calders, “Data preprocessing techniques for classi- fication without discrimination,” Knowledge and information systems , vol. 33, no. 1, pp. 1–33, 2012

  28. [35]

    Fairway: a way to build fair ml software,

    J. Chakraborty, S. Majumder, Z. Yu, and T. Menzies, “Fairway: a way to build fair ml software,” in Proceedings of the 28th ACM joint meeting on European software engineering conference and symposium on the foundations of software engineering , 2020, pp. 654–665

  29. [36]

    Fairness testing: testing software for discrimination,

    S. Galhotra, Y . Brun, and A. Meliou, “Fairness testing: testing software for discrimination,” in Proceedings of the 2017 11th Joint meeting on foundations of software engineering , 2017, pp. 498–510

  30. [37]

    Automated directed fairness testing,

    S. Udeshi, P. Arora, and S. Chattopadhyay, “Automated directed fairness testing,” in Proceedings of the 33rd ACM/IEEE international conference on automated software engineering , 2018, pp. 98–108

  31. [38]

    Fairness improvement with multiple protected attributes: How far are we?

    Z. Chen, J. M. Zhang, F. Sarro, and M. Harman, “Fairness improvement with multiple protected attributes: How far are we?” in Proceedings of the IEEE/ACM 46th International Conference on Software Engineering , 2024, pp. 1–13

  32. [39]

    Refair: Toward a context-aware recommender for fairness requirements engineering,

    C. Ferrara, F. Casillo, C. Gravino, A. De Lucia, and F. Palomba, “Refair: Toward a context-aware recommender for fairness requirements engineering,” in Proceedings of the IEEE/ACM 46th International Con- ference on Software Engineering , 2024, pp. 1–12

  33. [40]

    Fairness-aware machine learning engineering: how far are we?

    C. Ferrara, G. Sellitto, F. Ferrucci, F. Palomba, and A. De Lucia, “Fairness-aware machine learning engineering: how far are we?” Em- pirical Software Engineering , vol. 29, no. 1, p. 9, 2024

  34. [41]

    Lift: A scalable framework for measuring fairness in ml applications,

    S. Vasudevan and K. Kenthapadi, “Lift: A scalable framework for measuring fairness in ml applications,” in Proceedings of the 29th ACM International Conference on Information & Knowledge Management , 2020, pp. 2773–2780

  35. [42]

    “ignorance and prejudice

    J. M. Zhang and M. Harman, ““ignorance and prejudice” in software fairness,” in 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE). IEEE, 2021, pp. 1436–1447

  36. [43]

    Understanding fairness in software engineering: Insights from stack exchange,

    E. Sesari, F. Sarro, and A. Rastogi, “Understanding fairness in software engineering: Insights from stack exchange,” arXiv preprint arXiv:2402.19038, 2024

  37. [44]

    A catalog of fairness-aware practices in machine learning engineering,

    G. V oria, G. Sellitto, C. Ferrara, F. Abate, A. De Lucia, F. Ferrucci, G. Catolino, and F. Palomba, “A catalog of fairness-aware practices in machine learning engineering,” arXiv preprint arXiv:2408.16683, 2024

  38. [45]

    Fairness-aware practices from developers’ perspective: A survey,

    ——, “Fairness-aware practices from developers’ perspective: A survey,” Available at SSRN 4949224 , 2024

  39. [46]

    “fairness toolkits, a checkbox culture?

    A. Balayn, M. Yurrita, J. Yang, and U. Gadiraju, ““fairness toolkits, a checkbox culture?” on the factors that fragment developer practices in handling algorithmic harms,” in Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society , ser. AIES ’23. New York, NY ,...

  40. [47]

    The what-if tool: Interactive probing of machine learning models,

    J. Wexler, M. Pushkarna, T. Bolukbasi, M. Wattenberg, F. Vi ´egas, and J. Wilson, “The what-if tool: Interactive probing of machine learning models,” IEEE transactions on visualization and computer graphics , vol. 26, no. 1, pp. 56–65, 2019

  41. [48]

    scikit-fairness,

    “scikit-fairness,” 2019. [Online]. Available: https://github.com/koaning/ scikit-fairness

  42. [49]

    Improving fairness in machine learning systems: What do industry practitioners need?

    K. Holstein, J. Vaughan, I. Daum ´e, H., M. Dud ´ık, and H. Wallach, “Improving fairness in machine learning systems: What do industry practitioners need?” Association for Computing Machinery, 2019, cited By 214

  43. [50]

    Why don’t men ever stop to ask for directions? gender, social influence, and their role in technology acceptance and usage behavior,

    V . Venkatesh and M. Morris, “Why don’t men ever stop to ask for directions? gender, social influence, and their role in technology acceptance and usage behavior,” MIS Quarterly , vol. 24, pp. 115–139, 03 2000

  44. [51]

    User acceptance of information technology: Toward a unified view,

    V . Venkatesh, M. G. Morris, G. B. Davis, and F. D. Davis, “User acceptance of information technology: Toward a unified view,” MIS quarterly, pp. 425–478, 2003

  45. [52]

    The unified theory of acceptance and use of technology: A new approach in technology acceptance,

    A. Momani, “The unified theory of acceptance and use of technology: A new approach in technology acceptance,” International Journal of Sociotechnology and Knowledge Development , vol. 12, pp. 79–98, 07 2020

  46. [54]

    Online appendix,

    A. Authors, “Online appendix,” 2024. [Online]. Available: https: //figshare.com/s/583036589ad8b8abf666

  47. [55]

    Ethical perceptions of ai in hiring and organizational trust: The role of performance expectancy and social influence,

    M. Figueroa-Armijos, B. Clark, and S. Veiga, “Ethical perceptions of ai in hiring and organizational trust: The role of performance expectancy and social influence,” Journal of Business Ethics , vol. 186, 06 2022

  48. [56]

    Oneto and S

    L. Oneto and S. Chiappa, Fairness in Machine Learning . Cham: Springer International Publishing, 2020, pp. 155–196

  49. [57]

    Perceived usefulness, perceived ease of use, and user acceptance of information technology,

    F. D. Davis, “Perceived usefulness, perceived ease of use, and user acceptance of information technology,” MIS quarterly , pp. 319–340, 1989

  50. [58]

    User acceptance of hedonic information systems,

    H. Van der Heijden, “User acceptance of hedonic information systems,” MIS quarterly, pp. 695–704, 2004

  51. [59]

    Statistical power analyses using g* power 3.1: Tests for correlation and regression analyses,

    F. Faul, E. Erdfelder, A. Buchner, and A.-G. Lang, “Statistical power analyses using g* power 3.1: Tests for correlation and regression analyses,” Behavior research methods , vol. 41, no. 4, pp. 1149–1160, 2009

  52. [60]

    Personal opinion surveys,

    B. A. Kitchenham and S. L. Pfleeger, “Personal opinion surveys,” in Guide to advanced empirical software engineering . Springer, 2008, pp. 63–92

  53. [61]

    Conducting research on the internet:: Online survey design, development and implementation guidelines,

    D. Andrews, B. Nonnecke, and J. Preece, “Conducting research on the internet:: Online survey design, development and implementation guidelines,” 2007

  54. [62]

    Pls-sem for software engineering research: An introduction and survey,

    D. Russo and K.-J. Stol, “Pls-sem for software engineering research: An introduction and survey,” ACM Computing Surveys (CSUR), vol. 54, no. 4, pp. 1–38, 2021

  55. [63]

    Do you really code? designing and evaluating screening questions for online surveys with programmers,

    A. Danilova, A. Naiakshina, S. Horstmann, and M. Smith, “Do you really code? designing and evaluating screening questions for online surveys with programmers,” in 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 2021, pp. 537– 548

  56. [64]

    Empirical standards for software engineering research,

    P. Ralph, N. b. Ali, S. Baltes, D. Bianculli, J. Diaz, Y . Dittrich, N. Ernst, M. Felderer, R. Feldt, A. Filieri et al., “Empirical standards for software engineering research,” arXiv preprint arXiv:2010.03525 , 2020

  57. [65]

    Ethical issues in software engineering research: a survey of current practice,

    T. Hall and V . Flynn, “Ethical issues in software engineering research: a survey of current practice,” Empirical Software Engineering , vol. 6, no. 4, pp. 305–317, 2001

  58. [66]

    Smartpls 4,

    C. M. Ringle, S. Wende, and J.-M. Becker, “Smartpls 4,” B ¨onningstedt,

  59. [67]

    The partial least squares approach to structural equation modeling,

    W. W. Chin et al. , “The partial least squares approach to structural equation modeling,” Modern methods for business research , vol. 295, no. 2, pp. 295–336, 1998

  60. [68]

    A new criterion for assessing discriminant validity in variance-based structural equation modeling,

    J. Henseler, C. M. Ringle, and M. Sarstedt, “A new criterion for assessing discriminant validity in variance-based structural equation modeling,” Journal of the academy of marketing science , vol. 43, pp. 115–135, 2015

  61. [69]

    The elephant in the room: Predictive performance of pls models,

    G. Shmueli, S. Ray, J. M. V . Estrada, and S. B. Chatla, “The elephant in the room: Predictive performance of pls models,” Journal of business Research, vol. 69, no. 10, pp. 4552–4564, 2016

  62. [70]

    On the value rel- evance of customer satisfaction. multiple drivers and multiple markets,

    S. Raithel, M. Sarstedt, S. Scharf, and M. Schwaiger, “On the value rel- evance of customer satisfaction. multiple drivers and multiple markets,” Journal of the academy of marketing science , vol. 40, pp. 509–525, 2012

  63. [71]

    Personal computing: Toward a conceptual model of utilization,

    R. L. Thompson, C. A. Higgins, and J. M. Howell, “Personal computing: Toward a conceptual model of utilization,” MIS quarterly, pp. 125–143, 1991

  64. [72]

    Predictive model assessment in pls- sem: guidelines for using plspredict,

    G. Shmueli, M. Sarstedt, J. F. Hair, J.-H. Cheah, H. Ting, S. Vaithilingam, and C. M. Ringle, “Predictive model assessment in pls- sem: guidelines for using plspredict,” European journal of marketing , vol. 53, no. 11, pp. 2322–2347, 2019

  65. [74]

    Wohlin, P

    C. Wohlin, P. Runeson, M. H ¨ost, M. C. Ohlsson, B. Regnell, A. Wessl´en et al. , Experimentation in software engineering . Springer, 2012, vol. 236

  66. [2024]

    Available: https://www.smartpls.com/

    [Online]. Available: https://www.smartpls.com/

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.