Pith. sign in

REVIEW 4 major objections 4 minor 2 references

Artificial Intelligence Generates Stereotypical Images of Scientists but Can Also Detect Them: A Pilot Study Using the Draw-A-Scientist Test

T0 review · 4 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This pilot study reports that Midjourney v6.1, prompted to "draw a scientist," produces images that score 97% for lab coats and eyeglasses, 81% for male gender, and 85% for Caucasian appearance in a 100-image human-scored sample, and that…

desk verdict Human-scored stereotype percentages are plausible pilot data, but the 79% machine-human agreement is in-sample raw agreement, so the detection claim needs held-out validation before it can be accepted. read the letter →

arxiv 2504.19005 v1 pith:HI7Y7EYQ submitted 2025-04-26 physics.ed-ph

classification physics.ed-ph
keywords draw-a-scientisttestsciencestereotypesgenerativeAIbiasMidjourneyvision-languagemodelsautomaticscoringeducationDAST-C
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Using the classic Draw-A-Scientist prompt, this pilot study asked Midjourney v6.1 to produce 1,100 images of a scientist, had a science education researcher score 100 of them with a 15-item stereotype checklist, and prompted gpt-4.1-mini to score the same images. The human scores found the familiar stereotype pattern: 97% lab coats, 97% eyeglasses, 81% male, and 85% Caucasian. The vision-language model matched the human on 79% of checklist items on average and produced similar stereotype rates on the remaining 1,000 images. If the result holds, popular image generators may amplify the very stereotypes science educators have tried to break, and off-the-shelf AI could audit that bias at scale.

What carries the argument

The central instrument is the Draw-A-Scientist Checklist (DAST-C), a 15-item binary rubric that converts an image into stereotype flags such as lab coat, eyeglasses, male gender, Caucasian, lightning bolts, and secrecy. The second load-bearing object is gpt-4.1-mini, a vision-language model prompted with a role-playing instruction to act as an unbiased science education researcher and score each image against those same 15 items. The checklist supplies the common scoring language: human scores on 100 images are the reference, machine scores are compared item by item to produce machine-human agreement, and the machine's 1,000-image scores extend the stereotype counts to the full dataset. The prompt engineering step—assigning a role, supplying the checklist, and requesting a rationale—is what makes the machine scores interpretable as DAST-C responses.

What would settle it

Have a second trained rater independently score the same 100 Midjourney images: if the two raters disagree widely on features like male gender or Caucasian appearance, or if a fresh random sample of 100 from the 1,100 images yields lab-coat and eyeglasses rates far from the reported 97%, then the stereotype percentages and the 79% machine-human agreement are not reproducible.

Watch

Extended reading notes

Core claim

The paper's central claim is that a current image-generation model reproduces the long-documented Draw-A-Scientist stereotypes, and that a current vision-language model can detect them. On 100 images generated by Midjourney v6.1 with the prompt "draw a scientist," the researcher's DAST-C scoring found lab coats in 97%, eyeglasses in 97%, male gender in 81%, and Caucasian appearance in 85% of images. When the same 100 images were given to gpt-4.1-mini through a role-prompted checklist, its scores agreed with the human on 79% of items on average, with highest agreement on concrete features and lower agreement on ambiguous ones. On 1,000 additional Midjourney images scored only by the model, lab coats appeared in 97%, eyeglasses in 95%, male gender in 82%, and Caucasian appearance in 67%. The author concludes that generative AI both perpetuates stereotypical scientist imagery and can be used to identify that imagery automatically.

Load-bearing premise

Everything rests on the 100 human-scored images being a fair stand-in for all 1,100 generated images, and on one researcher's checklist judgments being the ground truth, with no random sampling and no second rater described.

Editorial extensions

If this is right

  • If the pilot result generalizes, a user typing "draw a scientist" into Midjourney v6.1 will almost always receive a lab-coated, bespectacled scientist, and about four-fifths of the time that scientist will be male.
  • AI-generated scientist imagery can be audited automatically: the same kind of model that creates the images can be prompted to flag stereotype features with roughly 79% average agreement with one human scorer.
  • VLM scoring returns rationales along with each checklist decision, which could help teachers see why a particular drawing is or is not marked stereotypical.
  • Machine scoring of the larger set suggests the stereotype pattern is not confined to the 100 scored images: the model found lab coats in 97%, eyeglasses in 95%, male gender in 82%, and Caucasian appearance in 67% of the 1,000 additional images.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: The paper does not randomize which generated images receive human scoring, so its percentages are not yet population estimates; a random sample and a second rater would tell whether the 79% figure is stable.
  • Inference: Because the prompt "draw a scientist" was designed to elicit stereotypes, the high stereotype rates partly reflect the prompt itself; comparing other prompts or neutral descriptions would separate model bias from prompt-induced bias.
  • Inference: The same audit chain—generate, score with a VLM, compare to human—could be run on other image generators and on other demographic categories, turning this one-off test into an automated stereotype monitor.
  • Inference: The item-level disagreement examples show that several checklist items, such as technology, working indoors, and middle-aged appearance, are ambiguous; clarifying those definitions would likely raise machine-human agreement beyond 79%.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This pilot study uses the Draw-A-Scientist Test (DAST) framework to examine whether AI-generated images of scientists are stereotypical and whether a vision-language model can detect such stereotypes. The author generated 1,100 images with Midjourney v6.1 using the prompt "draw a scientist", had one researcher (the author) score 100 of these images on the 15-item DAST-C rubric, and then used gpt-4.1-mini to score the same 100 images and an additional 1,000 images. Human scoring shows high prevalence of lab coats (97%), eyeglasses (97%), male gender (81%), and Caucasian ethnicity (85%). Machine scoring on the same 100 images reaches an average of 79% raw agreement with the human rater, and the machine's scores on the 1,000 additional images show similar stereotype rates (lab coat 97%, eyeglasses 95%, male 82%, Caucasian 67%). The author concludes that AI generates stereotypical scientist images and that a VLM can detect these stereotypes, with implications for science education and AI bias. The paper explicitly frames itself as a pilot study and lists plans for a larger, multi-rater, multi-model follow-up.

Significance. If the results are valid, the paper addresses a timely and educationally relevant question: whether widely used image-generation models amplify scientist stereotypes, and whether modern VLMs can be used to audit such stereotypes at scale. The use of an established instrument (DAST-C), the provision of machine rationales, and the explicit acknowledgment of the pilot nature are strengths. The paper's empirical claims, however, rest on two unsupported pillars: (1) the 100-image human-scored sample is assumed representative of the 1,100-image corpus without any sampling description or inter-rater reliability, and (2) the 79% machine-human agreement is computed in-sample on the very images used for prompt engineering, with no held-out validation and no chance-corrected metric. Consequently, the headline stereotype percentages for the 1,000-image set are not established. The contribution is a useful proof-of-concept with a clear experimental template, but the present evidence is too weak to support the paper's general conclusions.

major comments (4)
  1. [Methods (Data collection; Human scoring)] The paper does not describe how the 100 images were selected from the 1,100 generated images, and the human scoring was performed by a single rater with no inter-rater reliability check. Because the central descriptive claim (lab coat 97%, eyeglasses 97%, male 81%, Caucasian 85%) is computed on this sample, the representativeness of the sample and the reliability of the scoring are load-bearing. Without random sampling or a second independent rater, these percentages could reflect selection bias or idiosyncratic rubric application, and the study cannot support its broad claim that "AI-generated images of scientists represent stereotypical perceptions of them." This is a fixable but necessary limitation for the pilot to support even provisional conclusions.
  2. [Methods (Prompt engineering); Results (Machine-human agreement)] The 79% average machine-human agreement is computed on the same 100 images that were used to iteratively engineer the gpt-4.1-mini prompt. The paper says "the researcher went through prompt engineering" and then reports agreement on "the same 100 images." This is an in-sample evaluation: the prompt was tuned to match the human scores on those exact images, so the agreement reflects prompt optimization rather than the model's ability to score new images. No held-out validation set is used. The claim that "gpt-4.1-mini could also detect those stereotypes in the accuracy of 79%" is therefore not supported by the reported evidence.
  3. [Results (Machine scoring); Table 1] The 1,000-image machine scores (lab coat 97%, eyeglasses 95%, male 82%, Caucasian 67%) are presented without any independent validation on those images. The only evidence offered for the machine's scoring validity is the in-sample 79% agreement discussed above. Since that agreement is not an out-of-sample estimate and is not chance-corrected, the validity of the machine scores on the 1,000 unseen images is unestablished. The paper should either validate the machine scorer on a held-out set of human-scored images or explicitly reframe the 1,000-image numbers as unvalidated machine output rather than as findings.
  4. [Results (Table 1); Future works (Data imbalance)] The average machine-human agreement of 79% is a raw percentage over 15 features with highly imbalanced base rates. For example, features 10 (Indications of Danger) and 13 (Indications of Secrecy) are scored 0% by both human and machine, so perfect agreement on those features is trivially achieved by always answering "absent." The paper acknowledges in Future works that data imbalance "makes judging whether the MHA performance index acceptable complicated," but it does not implement a chance-corrected metric such as Cohen's kappa, nor does it report per-feature kappa or exclude degenerate features. As reported, the 79% average overstates the machine's detection ability and does not support the conclusion that the VLM "can detect" stereotypes with the claimed accuracy.
minor comments (4)
  1. [Introduction and Backgrounds] There is a typo in the first paragraph: "Sudent-drawn images" should be "Student-drawn images."
  2. [Results (Table 1)] The text cites "5. Symbols of Research" (72%) but Table 1 lists row 4 as "4. Symbols of Research" with 72% and row 5 as "5. Symbols of Knowledge"; the citation should be to feature 4. Similarly, the Machine scoring paragraph cites "3. Male Gender" but the table lists "8. Male Gender."
  3. [Results (Machine-human agreement)] The paper uses the term "accuracy" to describe machine-human agreement; "agreement" is more appropriate because no external ground truth is established, and "accuracy" conflates agreement with correctness.
  4. [Methods (Machine scoring)] The prompt engineering process is described only as "the researcher went through prompt engineering"; the number of iterations, the criteria for stopping, and the prompts attempted are not reported, which would be useful for reproducibility.

Circularity Check

1 steps flagged · score 6.0 of 10

The 79% machine-human agreement is computed on the same 100 images used for prompt engineering and is then used to validate the machine's scores for the 1,000 unseen images, so the detection claim is in-sample rather than independently predictive.

  1. fitted input called prediction [Methods – Prompt engineering; Results – Machine-human agreement and Table 1]
    "Using the data, the researcher went through prompt engineering to instruct gpt-4.1-mini to automatically analyze the remaining 1,000 images. ... However, gpt-4.1-mini could also detect those stereotypes in the accuracy of 79% in the same 100 images."

    The 100 human-scored images are explicitly the data used for prompt engineering, and the same 100 images are then used to compute the 79% machine-human agreement that is presented as 'validity of the scoring machine.' No held-out set is described before the machine is used to score the additional 1,000 images. The detection claim (RQ2) therefore rests on an in-sample evaluation of a prompt tuned to those exact images, and the Table 1 machine scores for the 1,000 images inherit that unvalidated accuracy. The 79% is not an estimate of how well gpt-4.1-mini would detect stereotypes in new images; it is a measure of agreement on the tuning set.

full rationale

The paper's RQ1 conclusion (Midjourney outputs are stereotyped) is supported by the researcher's direct DAST-C scoring of 100 images and does not depend on the machine; that part is not circular. The machine-scored 1,000-image extension, however, is only as trustworthy as the MHA validation, and that validation is contaminated because the same 100 images used to engineer the prompt are used to compute the 79% agreement. The prompt is a fitted input; reporting its agreement on the fitting set as 'validity' is the 'fitted input called prediction' pattern. No independent, held-out human scores are reported, so the RQ2 detection claim and the Table 1 rates for the additional 1,000 images are not established beyond in-sample tuning. The paper's own Future Works note on data imbalance acknowledges one related limitation but does not mention that the prompt-development images and the evaluation images are the same set. Self-citations (Lee & Zhai, 2025; Lee et al., 2023) are background and not load-bearing. No definitional circularity or self-citation chain is present, so the circularity is partial rather than total.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no fitted parameters or invented entities. It rests on the DAST-C rubric being a valid measure of stereotypes, on the unstated representativeness of the 100-image sample, and on treating the single researcher's scoring as ground truth.

assumptions (3)
  • domain assumption DAST-C checklist is a valid and accepted measure of scientist stereotypes.
    The stereotype rates and the machine scoring are defined by this checklist from Finson et al. (1995); the paper does not question its validity.
  • ad hoc to paper The 100 images selected for human scoring are representative of the 1,100 generated images.
    No randomization or sampling procedure is described; all stereotype prevalence estimates depend on this unstated assumption.
  • ad hoc to paper The researcher's single-rater DAST-C scoring is an acceptable ground truth.
    No second rater or inter-rater reliability is reported; machine-human agreement is interpreted as validity against this single scoring.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Artificial Intelligence Generates Stereotypical Images of Scientists but Can Also Detect Them: A Pilot Study Using the Draw-A-Scientist Test." pith.science (2026). https://pith.science/paper/HI7Y7EYQ

@misc{pith2026250419005,
  author       = {Pith},
  title        = {Pith review of: Artificial Intelligence Generates Stereotypical Images of Scientists but Can Also Detect Them: A Pilot Study Using the Draw-A-Scientist Test},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HI7Y7EYQ}},
  note         = {Machine review of arXiv:2504.19005}
}
read the original abstract

How the general public perceives scientists has been of interest to science educators for decades. While there can be many factors of it, the impact of recent generative artificial intelligence (AI) models is noteworthy, as these are rapidly changing how people acquire information. This report presents the pilot study examining how modern generative AI represents images of scientist using the Draw-A-Scientist Test (DAST). As a data, 1,100 images of scientist were generated using Midjourney v 6.1. One hundred of these images were analyzed by a science education scholar using a DAST scoring rubric. Using the data, the researcher went through prompt engineering to instruct gpt-4.1-mini to automatically analyze the remaining 1,000 images. The results show that generative AI represents stereotypical images of scientists, such as lab coat (97%), eyeglasses (97%), male gender (81%), and Caucasian (85%) in the 100 images analyzed by the researcher. However, gpt-4.1-mini could also detect those stereotypes in the accuracy of 79% in the same 100 images. gpt-4.1-mini also analyzed the remaining 1,000 images and found stereotypical features in the images (lab coat: 97%, eyeglasses: 95%, male gender: 82%, Caucasian: 67%). Discussions on the biases residing in today's generative AI and their implications on science education were made. The researcher plans to conduct a more comprehensive future study with an expanded methodology.

Figures

Figures reproduced from arXiv: 2504.19005 by the authors.

Figure 3
Figure 3. The researcher simply coded each image for the 15 features defined in DAST￾C (Finson et al., 1995). When each image was fed to gpt-4.1-mini API with the prompt ( [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

2 extracted references · 1 canonical work pages

  1. [1]

    Arantes, J. (2024). Understanding intersections between genAI and pre-service teacher education: What do we need to understand about the changing face of truth in science education? Journal of Science Education and Technology. https://doi.org/10.1007/s10956-024-10189-7 Chambers, D. W. (1983). Stereotypic images of the scientist: The draw-a-scientist test....

  2. [798]

    Shanahan, M., McDonell, K., & Reynolds, L. (2023). Role play with large language models. Nature, 623(7987), 493-498. Zhai, X., He, P., & Krajcik, J. (2022). Applying machine learning to automatically assess scientific models. Journal of Research in Science Teaching, 59(10), 1765-

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.