{"id":"e42ecd97-8eb7-4c07-a4d3-53430327a5f3","arxiv_id":"2504.19005","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Midjourney-generated scientist images are heavily stereotyped, and gpt-4.1-mini can detect those stereotypes with 79% raw agreement with a human scorer.","lead":"This pilot study generated 1,100 images of scientists with Midjourney and found that most show lab coats, eyeglasses, male gender, and white skin. It also tested whether gpt-4.1-mini can spot those stereotypes, and the model matched a human scorer on 79% of checklist features.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 79% machine-human agreement is computed on the same 100 images used to engineer the prompt, with no held-out validation set and no chance-corrected metric, so the machine scores on the 1,000 unseen images lack demonstrated validity.","rationale":"The reader identified sampling representativeness and single-rater scoring as the weakest assumptions. I agree those are important, but the more directly load-bearing issue is that the 100 human-scored images serve simultaneously as the prompt-engineering set and the evaluation set for the 79% MHA. Even a perfectly representative sample and a flawless human rater would not make an in-sample agreement figure a valid estimate of the model's accuracy on the 1,000 unseen images. The paper's central generation claim is directionally credible because the human-scored subset already shows high stereotype rates, but the quantitative extension to the full 1,100-image dataset depends on machine labels that are not independently validated. The detection claim is further weakened by the use of raw percentage agreement on heavily imbalanced features; trivial all-negative responses can produce perfect agreement on rare features. I do not think this requires rejecting the paper: it is explicitly a pilot study, the direction of the findings is plausible, and the author lists several relevant limitations in Table 2. The appropriate disposition remains CONDITIONAL, with the conditions being an out-of-sample validation of the machine scorer, chance-corrected agreement metrics, and a transparent description of how the 100 images were selected. Therefore the reader's verdict should be UNCHANGED, and my agreement with the reader's specific weakest-assumption framing is partial rather than full.","tokens_in":5367,"tokens_out":4417,"duration_ms":49022,"concrete_test":"Hold out 30 of the 100 human-scored images before any prompt engineering, use the remaining 70 to develop and freeze the gpt-4.1-mini prompt, then run the frozen prompt on the 30 held-out images and compute MHA, Cohen's kappa, and per-feature chance-adjusted agreement; if held-out agreement falls materially below 79% or kappa is low on features such as male gender and Caucasian, the detection claim and the 1,000-image stereotype rates are not yet supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's detection claim and the 1,000-image stereotype percentages both rest on a validation metric that is in-sample and not chance-corrected. The Methods state that 'the researcher went through prompt engineering' using the 100 human-scored images, and the Results then report machine-human agreement on 'the same 100 images.' There is no held-out set, so the 79% MHA is a measure of how well the prompt was tuned to those particular images, not an estimate of how accurately gpt-4.1-mini would score new images. This is load-bearing because Table 1's machine scores for the additional 1,000 images are accepted without any independent check; the full-dataset stereotype rates (lab coat 97%, eyeglasses 95%, male 82%, Caucasian 67%) are therefore not established. The average is also a raw percentage over 15 features with extreme base rates: Features 10 and 13 occur 0% in both human and machine scores, so perfect agreement there is achievable by always answering 'absent,' and Feature 6 differs sharply (3% human vs 18% machine) yet still contributes to the average. Without Cohen's kappa or per-feature chance-adjusted agreement, 79% overstates detection ability. The paper's own Future Works section concedes that data imbalance 'makes judging whether the MHA performance index acceptable complicated,' but it does not acknowledge that the prompt-development images and the evaluation images are the same set.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This pilot study uses the Draw-A-Scientist Test (DAST) framework to examine whether AI-generated images of scientists are stereotypical and whether a vision-language model can detect such stereotypes. The author generated 1,100 images with Midjourney v6.1 using the prompt \"draw a scientist\", had one researcher (the author) score 100 of these images on the 15-item DAST-C rubric, and then used gpt-4.1-mini to score the same 100 images and an additional 1,000 images. Human scoring shows high prevalence of lab coats (97%), eyeglasses (97%), male gender (81%), and Caucasian ethnicity (85%). Machine scoring on the same 100 images reaches an average of 79% raw agreement with the human rater, and the machine's scores on the 1,000 additional images show similar stereotype rates (lab coat 97%, eyeglasses 95%, male 82%, Caucasian 67%). The author concludes that AI generates stereotypical scientist images and that a VLM can detect these stereotypes, with implications for science education and AI bias. The paper explicitly frames itself as a pilot study and lists plans for a larger, multi-rater, multi-model follow-up.","tokens_in":5681,"tokens_out":3133,"duration_ms":31981,"significance":"If the results are valid, the paper addresses a timely and educationally relevant question: whether widely used image-generation models amplify scientist stereotypes, and whether modern VLMs can be used to audit such stereotypes at scale. The use of an established instrument (DAST-C), the provision of machine rationales, and the explicit acknowledgment of the pilot nature are strengths. The paper's empirical claims, however, rest on two unsupported pillars: (1) the 100-image human-scored sample is assumed representative of the 1,100-image corpus without any sampling description or inter-rater reliability, and (2) the 79% machine-human agreement is computed in-sample on the very images used for prompt engineering, with no held-out validation and no chance-corrected metric. Consequently, the headline stereotype percentages for the 1,000-image set are not established. The contribution is a useful proof-of-concept with a clear experimental template, but the present evidence is too weak to support the paper's general conclusions.","major_comments":[{"comment":"The paper does not describe how the 100 images were selected from the 1,100 generated images, and the human scoring was performed by a single rater with no inter-rater reliability check. Because the central descriptive claim (lab coat 97%, eyeglasses 97%, male 81%, Caucasian 85%) is computed on this sample, the representativeness of the sample and the reliability of the scoring are load-bearing. Without random sampling or a second independent rater, these percentages could reflect selection bias or idiosyncratic rubric application, and the study cannot support its broad claim that \"AI-generated images of scientists represent stereotypical perceptions of them.\" This is a fixable but necessary limitation for the pilot to support even provisional conclusions.","section":"Methods (Data collection; Human scoring)"},{"comment":"The 79% average machine-human agreement is computed on the same 100 images that were used to iteratively engineer the gpt-4.1-mini prompt. The paper says \"the researcher went through prompt engineering\" and then reports agreement on \"the same 100 images.\" This is an in-sample evaluation: the prompt was tuned to match the human scores on those exact images, so the agreement reflects prompt optimization rather than the model's ability to score new images. No held-out validation set is used. The claim that \"gpt-4.1-mini could also detect those stereotypes in the accuracy of 79%\" is therefore not supported by the reported evidence.","section":"Methods (Prompt engineering); Results (Machine-human agreement)"},{"comment":"The 1,000-image machine scores (lab coat 97%, eyeglasses 95%, male 82%, Caucasian 67%) are presented without any independent validation on those images. The only evidence offered for the machine's scoring validity is the in-sample 79% agreement discussed above. Since that agreement is not an out-of-sample estimate and is not chance-corrected, the validity of the machine scores on the 1,000 unseen images is unestablished. The paper should either validate the machine scorer on a held-out set of human-scored images or explicitly reframe the 1,000-image numbers as unvalidated machine output rather than as findings.","section":"Results (Machine scoring); Table 1"},{"comment":"The average machine-human agreement of 79% is a raw percentage over 15 features with highly imbalanced base rates. For example, features 10 (Indications of Danger) and 13 (Indications of Secrecy) are scored 0% by both human and machine, so perfect agreement on those features is trivially achieved by always answering \"absent.\" The paper acknowledges in Future works that data imbalance \"makes judging whether the MHA performance index acceptable complicated,\" but it does not implement a chance-corrected metric such as Cohen's kappa, nor does it report per-feature kappa or exclude degenerate features. As reported, the 79% average overstates the machine's detection ability and does not support the conclusion that the VLM \"can detect\" stereotypes with the claimed accuracy.","section":"Results (Table 1); Future works (Data imbalance)"}],"minor_comments":[{"comment":"There is a typo in the first paragraph: \"Sudent-drawn images\" should be \"Student-drawn images.\"","section":"Introduction and Backgrounds"},{"comment":"The text cites \"5. Symbols of Research\" (72%) but Table 1 lists row 4 as \"4. Symbols of Research\" with 72% and row 5 as \"5. Symbols of Knowledge\"; the citation should be to feature 4. Similarly, the Machine scoring paragraph cites \"3. Male Gender\" but the table lists \"8. Male Gender.\"","section":"Results (Table 1)"},{"comment":"The paper uses the term \"accuracy\" to describe machine-human agreement; \"agreement\" is more appropriate because no external ground truth is established, and \"accuracy\" conflates agreement with correctness.","section":"Results (Machine-human agreement)"},{"comment":"The prompt engineering process is described only as \"the researcher went through prompt engineering\"; the number of iterations, the criteria for stopping, and the prompts attempted are not reported, which would be useful for reproducibility.","section":"Methods (Machine scoring)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is more of a research note or pilot-report than a full paper, and the author explicitly frames it as such. The editor may wish to consider whether the journal's standards allow a pilot with a single human rater and an in-sample validation metric as the basis for general conclusions. The topic is timely and the planned follow-up is well scoped, but the present version's central claims outrun its evidence. A revision that adds sampling details, a second rater, held-out validation, and chance-corrected agreement measures would substantially strengthen the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a small pilot with a genuinely useful idea and an honest limitations section. The direction of the result is expectable, but the machine-scoring validity claim is weaker than the 79% headline suggests. Treat the human-scored 100-image percentages as promising pilot data, not established rates.\n\nWhat is actually new: applying the Draw-A-Scientist Checklist to Midjourney-generated images, and trying a vision-language model as an automated DAST scorer that returns rationales. That is a legitimate extension of two established literatures, and the study design is straightforward: one prompt, one image model, one human rater, one VLM, and an external benchmark for the machine. Credit where due: the paper uses the published DAST-C instrument, calls itself a pilot, and lists a sensible set of future expansions. It also reports the raw numbers rather than hiding them, which lets you see the weak spots. The citation pattern is fine; the relevant DAST and VLM work is there.\n\nThe weak spots are real, though. The human scoring is one person scoring 100 images, with no second rater, no reported sampling procedure for those 100, no confidence intervals, and no data release. That alone makes the 97% lab-coat figure suggestive, not conclusive. The larger problem is the machine-scoring validity claim. The researcher used the same 100 human-scored images for prompt engineering and for reporting the 79% machine-human agreement, so that 79% is in-sample. It tells you how well the prompt was tuned to those particular images, not how accurately the VLM would score new images. The machine scores on the other 1,000 images therefore rest on an unvalidated tool. On top of that, agreement is reported as raw percent agreement with no kappa or per-feature chance correction, and the average is inflated by features that never occur in either human or machine scoring.\n\nThe paper's own Future Works section concedes that the data imbalance makes the MHA hard to judge, but it does not acknowledge the same-set problem. That omission matters.\n\nBottom line: the human-scored pattern (male, Caucasian, lab coat, eyeglasses) is plausible and consistent with decades of DAST literature, but the machine-detection result is not established as stated. This deserves a serious referee as a pilot; I would ask for a held-out validation set, chance-corrected agreement, a second human rater on a subset, and a description of how the 100 were sampled. With those additions it could be a citable pilot. I'd take it in a working group to talk about validation norms, but I wouldn't cite it yet.","headline":"Human-scored stereotype percentages are plausible pilot data, but the 79% machine-human agreement is in-sample raw agreement, so the detection claim needs held-out validation before it can be accepted.","tokens_in":6160,"tokens_out":2771,"would_cite":false,"duration_ms":29075,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This pilot study reports that Midjourney v6.1, prompted to \"draw a scientist,\" produces images that score 97% for lab coats and eyeglasses, 81% for male gender, and 85% for Caucasian appearance in a 100-image human-scored sample, and that…","keywords":["draw-a-scientist test","science stereotypes","generative AI bias","Midjourney","vision-language models","automatic scoring","science education","DAST-C"],"falsifier":"Have a second trained rater independently score the same 100 Midjourney images: if the two raters disagree widely on features like male gender or Caucasian appearance, or if a fresh random sample of 100 from the 1,100 images yields lab-coat and eyeglasses rates far from the reported 97%, then the stereotype percentages and the 79% machine-human agreement are not reproducible.","tokens_in":5168,"feed_emoji":"🥼","tokens_out":6964,"duration_ms":63480,"temperature":0.7,"pith_summary":"Using the classic Draw-A-Scientist prompt, this pilot study asked Midjourney v6.1 to produce 1,100 images of a scientist, had a science education researcher score 100 of them with a 15-item stereotype checklist, and prompted gpt-4.1-mini to score the same images. The human scores found the familiar stereotype pattern: 97% lab coats, 97% eyeglasses, 81% male, and 85% Caucasian. The vision-language model matched the human on 79% of checklist items on average and produced similar stereotype rates on the remaining 1,000 images. If the result holds, popular image generators may amplify the very stereotypes science educators have tried to break, and off-the-shelf AI could audit that bias at scale.","feed_headline":"AI draws scientists as stereotypes — and can spot them in itself","feed_subtitle":"Midjourney portraits of 'a scientist' were 97% lab coats in a 100-image pilot; a vision model flagged the same pattern.","key_machinery":"The central instrument is the Draw-A-Scientist Checklist (DAST-C), a 15-item binary rubric that converts an image into stereotype flags such as lab coat, eyeglasses, male gender, Caucasian, lightning bolts, and secrecy. The second load-bearing object is gpt-4.1-mini, a vision-language model prompted with a role-playing instruction to act as an unbiased science education researcher and score each image against those same 15 items. The checklist supplies the common scoring language: human scores on 100 images are the reference, machine scores are compared item by item to produce machine-human agreement, and the machine's 1,000-image scores extend the stereotype counts to the full dataset. The prompt engineering step—assigning a role, supplying the checklist, and requesting a rationale—is what makes the machine scores interpretable as DAST-C responses.","core_discovery":"The paper's central claim is that a current image-generation model reproduces the long-documented Draw-A-Scientist stereotypes, and that a current vision-language model can detect them. On 100 images generated by Midjourney v6.1 with the prompt \"draw a scientist,\" the researcher's DAST-C scoring found lab coats in 97%, eyeglasses in 97%, male gender in 81%, and Caucasian appearance in 85% of images. When the same 100 images were given to gpt-4.1-mini through a role-prompted checklist, its scores agreed with the human on 79% of items on average, with highest agreement on concrete features and lower agreement on ambiguous ones. On 1,000 additional Midjourney images scored only by the model, lab coats appeared in 97%, eyeglasses in 95%, male gender in 82%, and Caucasian appearance in 67%. The author concludes that generative AI both perpetuates stereotypical scientist imagery and can be used to identify that imagery automatically.","pith_inferences":["Inference: The paper does not randomize which generated images receive human scoring, so its percentages are not yet population estimates; a random sample and a second rater would tell whether the 79% figure is stable.","Inference: Because the prompt \"draw a scientist\" was designed to elicit stereotypes, the high stereotype rates partly reflect the prompt itself; comparing other prompts or neutral descriptions would separate model bias from prompt-induced bias.","Inference: The same audit chain—generate, score with a VLM, compare to human—could be run on other image generators and on other demographic categories, turning this one-off test into an automated stereotype monitor.","Inference: The item-level disagreement examples show that several checklist items, such as technology, working indoors, and middle-aged appearance, are ambiguous; clarifying those definitions would likely raise machine-human agreement beyond 79%."],"forward_implications":["If the pilot result generalizes, a user typing \"draw a scientist\" into Midjourney v6.1 will almost always receive a lab-coated, bespectacled scientist, and about four-fifths of the time that scientist will be male.","AI-generated scientist imagery can be audited automatically: the same kind of model that creates the images can be prompted to flag stereotype features with roughly 79% average agreement with one human scorer.","VLM scoring returns rationales along with each checklist decision, which could help teachers see why a particular drawing is or is not marked stereotypical.","Machine scoring of the larger set suggests the stereotype pattern is not confined to the 100 scored images: the model found lab coats in 97%, eyeglasses in 95%, male gender in 82%, and Caucasian appearance in 67% of the 1,000 additional images."],"supporting_citations":[{"why":"Supplies the Draw-A-Scientist Test, including the prompt \"draw a scientist,\" and the construct of stereotypic scientist images.","marker":"(Chambers, 1983)"},{"why":"Supplies the Draw-A-Scientist Checklist (DAST-C), the 15-item rubric used for both human and machine scoring.","marker":"(Finson et al., 1995)"},{"why":"Documents persistent DAST stereotypes across grade levels, genders, racial groups, and national borders, motivating the search for them in AI images.","marker":"(Finson, 2002)"},{"why":"Systematic review and meta-analysis establishing global consistency of scientist stereotypes, which frames the expected pattern.","marker":"(Ferguson & Lezotte, 2020)"},{"why":"Shows that vision-language models such as GPT-4V can score drawn images and provide rationales, supporting the machine-scoring approach.","marker":"(Lee & Zhai, 2025)"},{"why":"Provides the role-prompting method used to instruct gpt-4.1-mini to act as an unbiased science education researcher.","marker":"(Shanahan et al., 2023)"},{"why":"Prior work on automatically assessing student-drawn scientific models, setting up the extension to DAST images.","marker":"(Zhai et al., 2022)"},{"why":"Prior automated scoring of student hand drawings, which supports the researcher's claimed experience and the feasibility of machine scoring.","marker":"(Lee et al., 2023)"}],"fun_headline_variants":["Midjourney draws scientists as stereotypes; GPT-4 mini spots them","AI generates stereotypical scientist images then flags them itself","Midjourney's scientist portraits are 97% lab coats; AI can see it","AI both draws and catches its own scientist stereotypes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the 100 human-scored images being a fair stand-in for all 1,100 generated images, and on one researcher's checklist judgments being the ground truth, with no random sampling and no second rater described.","fun_headline_variants_meta":{"raw":{"variants":["Midjourney draws scientists as stereotypes; GPT-4 mini spots them","AI generates stereotypical scientist images then flags them itself","Midjourney's scientist portraits are 97% lab coats; AI can see it","AI both draws and catches its own scientist stereotypes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00048,"raw_usage":{"total_tokens":2428,"prompt_tokens":1054,"completion_tokens":1374,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":670,"completion_tokens_details":{"reasoning_tokens":1301}},"tokens_in":670,"tokens_out":1374,"duration_ms":9524,"temperature":1.0,"reasoning_tokens":1301,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:02:56.699755+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a second trained rater independently score the same 100 Midjourney images: if the two raters disagree widely on features like male gender or Caucasian appearance, or if a fresh random sample of 100 from the 1,100 images yields lab-coat and eyeglasses rates far from the reported 97%, then the stereotype percentages and the 79% machine-human agreement are not reproducible.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the role-prompting method used to instruct gpt-4.1-mini to act as an unbiased science education researcher."}],"review_version":1}