{"id":"e1611e62-b739-4428-9667-de3262637385","arxiv_id":"2507.18633","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new 1.95M-image benchmark measures how well vision models identify artist names explicitly prompted into text-to-image systems, across artists, prompts, generators, and artist counts.","lead":"This paper releases a 1.95 million image benchmark for guessing which artist's name was written in the text prompt that generated an image. It tests vision models on held-out artists, complex prompts, multiple artists, and different image generators, and shows the task is far from solved.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Midjourney evaluation uses a different artist split (35 seen / 95 held-out) than the stated 100/10 split, undermining cross-model comparisons and the held-out generalization claim.","rationale":"The reader's verdict is CONDITIONAL, and the reader's rationale explicitly flags the Midjourney artist-split inconsistency and the '91%' statement as needing correction before held-out generalization conclusions can be taken at face value. My stress-test identifies the Midjourney split inconsistency as the single most load-bearing concern because it directly undermines the comparability of one of the four stated generalization axes (held-out artists) across the four text-to-image models, and it is an internal data inconsistency rather than a debatable modeling choice. The reader's stated weakest_assumption instead emphasized the benchmark's assumption that artist names reliably alter the generated image; that concern is real and is partially addressed by the paper's own Section 3.5 analysis and by the exclusion of SD3.5 and FLUX, so I do not elevate it above the split inconsistency. The concrete check I propose is straightforward: verify the released Midjourney metadata against the 100/10 split and recompute the affected tables if needed. If the split is as reported in Table 5(c), the paper must either correct the split to match the stated benchmark design or clearly document and justify the different Midjourney label space; otherwise the cross-model generalization comparisons and the 'substantial headroom' conclusion are not supported on a consistent evaluation. This does not change the reader's CONDITIONAL verdict; it strengthens the reason for it.","tokens_in":25942,"tokens_out":3781,"duration_ms":39488,"concrete_test":"Inspect the released dataset metadata (e.g., the Midjourney image-prompt pairs and artist lists on HuggingFace) and count the number of unique artist labels in the Midjourney \"seen train\", \"seen test\", \"held-out reference\", and \"held-out test\" sets. Verify whether they match the 100/10 split used for SDXL, SD1.5, and PixArt. If they do not, recompute Table 13 and the Midjourney panels of Figure 7 using only the canonical 100 seen and 10 held-out artists, and re-check whether the qualitative conclusions in Section 5 (CSD better on held-out, prototypical networks worse on held-out Midjourney, no method exceeds 91%) survive on a consistent label space.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the benchmark measures seen vs. held-out artists consistently across generators and that the reported patterns (e.g., CSD transfers better on held-out artists, prototypical networks on seen artists) describe current vision methods. Section 3.1 states the benchmark uses 100 seen and 10 held-out artists. Yet the Midjourney subset is reported in Table 5(c) and Table 13 as \"Seen artists (34-way)\" and \"Held-out (96-way)\", i.e., 35 seen and 95 held-out artists, with chance accuracy 2.9% and 1.0% respectively. This directly contradicts the stated split and the 10-way held-out setting used for SDXL, SD1.5, and PixArt. As a result, the \"seen vs. held-out\" axis is not aligned across models: Figure 6 and Figure 7 cross-model comparisons mix a 100-way/10-way problem with a 34-way/96-way problem, so the held-out accuracy numbers for Midjourney (e.g., CSD 42.8%, prototypical network 16.8%) are not on the same label space as the other models and the claim that style descriptors generalize better on held-out artists may be an artifact of the different label-space difficulty. This is an internal inconsistency in the reported data, not a disagreement with outside consensus. The released dataset metadata must clarify which split was actually used; the supplement's Section 7.1 describes selecting 10 held-out artists, which conflicts with the tables. Separately, Section 5.1's statement that \"none exceed 91% accuracy\" is contradicted by Table 8's CSD held-out simple accuracy of 92.0%, though that table is a prototype-retrieval baseline; the Midjourney split issue is more load-bearing because it changes the interpretation of the central generalization axis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a large-scale benchmark for identifying which artist name was invoked in a text-to-image prompt, given only the generated image. The dataset comprises roughly 1.95M images across 110 artist names and four generalization axes: seen versus held-out artists, simple versus complex prompts, multiple text-to-image generators (SDXL, SD1.5, PixArt-Σ, and a collected Midjourney subset), and prompts with two or three artist names. The authors evaluate retrieval-based baselines (CLIP, DINOv2, CSD, AbC) and trained classifiers (prototypical networks and a vanilla classifier), reporting accuracy and mAP@10 with bootstrapped confidence intervals. The central empirical claims are that supervised and few-shot methods generalize better on seen artists and complex prompts, style descriptors transfer better on simple prompts and held-out artists, multi-artist prompts are the most difficult, and no method approaches saturation.","tokens_in":26294,"tokens_out":3627,"duration_ms":37504,"significance":"If the benchmark is valid, it is a valuable public testbed for the responsible moderation of text-to-image content and for studying the relationship between prompted artist names and generated image style. The paper's strengths are its controlled prompt construction, the breadth of evaluated method families, transparent dataset release, bootstrapped uncertainty estimates, and several informative ablations (training data composition, prototype sources, and k-NN behavior). The problem is timely and the reported headroom is plausibly of practical interest. The main weaknesses are internal inconsistencies in the Midjourney evaluation split and the lack of a direct test of whether the remaining benchmark items are visually answerable, both of which affect the interpretation of the cross-model and 'headroom' conclusions.","major_comments":[{"comment":"The benchmark's validity rests on the assumption that an artist name explicitly placed in a prompt reliably and detectably alters the generated image. The authors themselves exclude SD3.5 and FLUX because this assumption fails, and Section 3.5 quantifies that PixArt images and complex prompts substantially dilute the artist's influence. Yet the benchmark retains PixArt and complex-prompt test items without establishing that these items are answerable in principle. As a result, the 'substantial headroom' conclusion in Section 5.1 may conflate genuinely difficult vision problems with test items where the image simply does not contain enough artist-specific signal to identify the prompted name. Please add an answerability analysis, such as human accuracy on a representative sample or an oracle-style upper bound per condition, and discuss whether the observed accuracy gaps reflect method limitations or unanswerable items.","section":"Section 3.3, Section 3.5, Section 5.1"}],"minor_comments":[{"comment":"The statement that 'none exceed 91% accuracy' is contradicted by Table 8 in the supplement, where CSD with artist-average retrieval reaches 92.0% on held-out artists with simple prompts; please qualify the claim to the main evaluation setting or adjust the text.","section":"Section 5.1, Table 8"},{"comment":"The main text says simple prompts use '500 different contents sampled from ChatGPT', while Section 7.2 states that 100 subjects were curated and lists a ChatGPT request for 100 subjects; these numbers should be reconciled.","section":"Section 3.2, Section 7.2"},{"comment":"The text refers to 'our observation in Table 3.5'; this should be Section 3.5, since 3.5 is not a table.","section":"Section 5.3"},{"comment":"The caption of Figure 6 describes the axes as 100-way and 10-way for all models, but Table 13 reports Midjourney as 34-way and 96-way; the figure caption and/or the figure itself should be corrected to reflect the actual label spaces.","section":"Figure 6"},{"comment":"There is a typo in 'by Leottaet al.'; it should read 'Leotta et al.'.","section":"Supplement, related work paragraph"}],"recommendation":"major_revision","confidential_remarks":"The dataset and benchmark are the main contribution, and the paper's qualitative findings are generally consistent across the open-weight generators. The Midjourney split inconsistency and the answerability question are both fixable, but they are load-bearing for the cross-model generalization claims. I do not see grounds for rejection, provided the authors address these issues in a revised submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The authors have built something the field needs: a large, transparent benchmark for deciding whether a generated image was prompted with an artist's name. The dataset is 1.95M images, 110 artists, four generators, four generalization axes, and the evaluation covers six method families, including retrieval baselines, style descriptors, data attribution, classifiers, and few-shot prototypes. Construction is careful: held-out artists were filtered against CSD's training captions, so the main benchmark is not circular; prompt splits are documented; bootstrapping is standard. The headline findings—trained classifiers win on seen artists and complex prompts, CSD wins on simple and held-out prompts, multi-artist prompts are hardest—are consistent with the tables for SDXL, SD1.5, and PixArt.\n\nThe problem is Midjourney. Section 3.1 states the benchmark uses 100 seen and 10 held-out artists. Supplement Table 5(c) reports the Midjourney subset as 35 seen and 95 held-out artists, and Table 13 evaluates it as 34-way and 96-way. Figure 6 and Figure 7 then plot Midjourney in the same seen/held-out frame as the other models. That is not a cosmetic mismatch: a 96-way held-out problem with 1% chance is not comparable to a 10-way problem with 10% chance, and the claim that CSD transfers better on held-out Midjourney images may simply reflect the different label space. This needs to be corrected in the paper or fully disclosed per model, with the cross-model generalization conclusions re-stated.\n\nSmaller items: the \"none exceed 91%\" statement in Section 5.1 is safe only if you exclude the supplement's prototype-retrieval experiment, which reaches 92.0 on held-out simple prompts; that claim should be scoped. The held-out set is only 10 artists, so open-set conclusions are necessarily provisional. And Midjourney is complex prompts only, which the limitations section admits. None of these sink the core contribution—the dataset and the SDXL/SD1.5/PixArt comparisons are solid—but the Midjourney split inconsistency is real and load-bearing for the generalization story.\n\nMy bottom line: this deserves a serious referee. A competent reviewer can verify the split and ask for a revision; the resource is valuable enough that desk-rejecting it would be a mistake. I would send it out.","headline":"Large, transparent benchmark for prompted-artist identification; method comparison is solid for three generators, but the Midjourney subset uses a different artist split than the stated one, so the held-out generalization story needs a fix.","tokens_in":26883,"tokens_out":3968,"would_cite":true,"duration_ms":38105,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces a 1.95M-image benchmark and shows that no current vision model can reliably identify which artist name was used in the prompt of a generated image.","keywords":["prompted artist identification","text-to-image generation","style attribution","generated image detection","benchmark","generalization","prototypical networks","artist names"],"falsifier":"Regenerate a held-out set of the benchmark's complex SDXL prompts twice—once with the artist name and once with the identical prompt and seed but no artist name—and have a strong binary detector or human raters pick which image was artist-prompted. If accuracy is at chance on complex prompts, many test items are visually unanswerable, and the measured headroom would overstate the limits of vision methods.","tokens_in":25761,"feed_emoji":"🎨","tokens_out":9632,"duration_ms":86978,"temperature":0.7,"pith_summary":"The paper's central claim is that prompted-artist identification—predicting, from the image alone, which artist's name appeared in the text prompt that generated it—is a well-defined task that current vision methods cannot yet solve reliably. To back this claim, the authors build the first large-scale benchmark for the task: 1.95 million generated images, 110 frequently prompted artist names, spanning simple and real-user complex prompts, four text-to-image generators, and prompts with one, two, or three artists. Their experiments show consistent generalization patterns: classifiers and few-shot prototypical networks trained on artist-prompted images do best on seen artists and complex prompts, while style descriptors trained on real artwork transfer better to simple and held-out prompts; multi-artist prompts are the hardest, and no method approaches saturation. The benchmark matters because online platforms ban artist-named generations, yet without access to the original prompt there is currently no reliable way to detect them.","feed_headline":"No vision model can reliably name the artist behind a generated image","feed_subtitle":"A 1.95M-image benchmark with 110 artists reveals large gaps in detecting artist-name prompts from pixels alone.","key_machinery":"The load-bearing object is the structured benchmark dataset itself. For each of 110 frequently prompted artist names, images are generated by inserting the name into the same content prompts—simple prompts written for the task and complex prompts scraped from real users—across SDXL, SD1.5, PixArt-Σ, and Midjourney, with 100 seen artists, 10 held-out artists, and separate multi-artist subsets with two and three names. The 'same content prompt, different artist name' design isolates the effect of the artist name from the effect of the content, and the held-out artists and prompts force methods to generalize rather than memorize. The evaluation then compares retrieval-based baselines (CLIP, DINOv2, contrastive style descriptors, data attribution features) with fine-tuned classifiers and prototypical networks, using CLIP embedding similarity to quantify how strongly an artist name influences the output.","core_discovery":"The paper establishes that recognizing the artist an image generator was told to imitate is a distinct problem from recognizing artistic style in real paintings, and that the gap between the two is measurable. On the released benchmark, methods trained on real artwork (contrastive style descriptors) generalize well to simple prompts and held-out artists, where the artist's style is visually apparent, but trained classifiers and prototypical networks—which learn from generated, artist-prompted images—surpass them on complex prompts and seen artists. Across all settings, the best methods stay below 91% accuracy, performance drops consistently as prompts become more complex and on images from PixArt compared with SDXL and SD1.5, and prompts with multiple artists remain the hardest case. The authors also show that adding training images from one generator does not improve performance on an unseen generator, and that each added artist name in a prompt has a smaller visual effect than the previous one.","pith_inferences":["If the measured patterns hold for future generators that do respond to artist names, moderation could be framed as open-set retrieval against a maintained artist database, with held-out artist accuracy as the primary deployment metric rather than closed-set accuracy.","The persistent gap between real-art style descriptors and generated-image classifiers suggests a generator-aware style encoder—one that models how each text-to-image model attenuates artist influence—could transfer better across generators than either current family.","A testable extension of the paper's CLIP similarity analysis is a per-artist 'promptability' score (how much an artist name changes the generated image); if such scores predict classification difficulty, they could be used to choose which artists need extra reference data or targeted training.","Because the benchmark excludes SD3.5 and FLUX on the grounds that they ignore artist names, the paper implies a moderation system for those models would need to detect style imitation from prompt description rather than from artist names, a different and possibly harder signal."],"forward_implications":["Prompted-artist identification cannot be equated with style recognition of real artwork; the two tasks have different generalization curves, so moderation tools need training on generated, artist-prompted images.","Deployed detection will be most reliable on simple prompts and familiar artists, and least reliable on multi-artist prompts and generators like PixArt where the artist's influence is weak.","No current method generalizes across text-to-image generators: adding training data from one generator does not help on another, so cross-generator robustness must be handled explicitly.","The benchmark's headroom is large—best methods are far below saturation on every setting—so artist-name detection is an open problem rather than a solved one.","Multi-artist prompts are the clearest failure mode; because each additional name dilutes the visual trace, a practical system would need to predict a set of names rather than a single label."],"supporting_citations":[{"why":"Supplies the contrastive style descriptor baseline and the real-artist reference set used for style comparison and prototypes.","marker":"[76]"},{"why":"Supplies real-user complex prompts and the Midjourney image-prompt pairs used in the benchmark.","marker":"[78]"},{"why":"Supplies the few-shot prototypical network method used as a strong trained baseline.","marker":"[73]"},{"why":"Provides the image and text features used for prototypes, retrieval, and similarity analyses.","marker":"[60]"},{"why":"Provides the SDXL generator, which produces the largest and most thoroughly evaluated image subset.","marker":"[59]"},{"why":"Supplies the data attribution baselines compared against style descriptors and trained classifiers.","marker":"[81]"},{"why":"Provides the closest prior artist-inference dataset, which the benchmark extends to a much larger scale.","marker":"[43]"}],"fun_headline_variants":["Multi-artist prompts trip up every artist-attribution model","1.95M-image benchmark finds artist-prompt recognition still hard","Pixel-based artist attribution: the multi-artist case remains unsolved","New benchmark: no model cracks multi-artist generated image prompts","Even top classifiers fail to detect multiple artists in prompts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark assumes that an artist name placed in a prompt leaves a detectable visual trace in the generated image for the four generators studied; when prompts are complex or the generator is PixArt, the paper's own measurements show that trace weakens, and it excluded SD3.5 and FLUX because they showed almost no trace.","fun_headline_variants_meta":{"raw":{"variants":["Multi-artist prompts trip up every artist-attribution model","1.95M-image benchmark finds artist-prompt recognition still hard","Pixel-based artist attribution: the multi-artist case remains unsolved","New benchmark: no model cracks multi-artist generated image prompts","Even top classifiers fail to detect multiple artists in prompts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001081,"raw_usage":{"total_tokens":4503,"prompt_tokens":907,"completion_tokens":3596,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":3511}},"tokens_in":523,"tokens_out":3596,"duration_ms":26582,"temperature":1.0,"reasoning_tokens":3511,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:09:56.064258+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Regenerate a held-out set of the benchmark's complex SDXL prompts twice—once with the artist name and once with the identical prompt and seed but no artist name—and have a strong binary detector or human raters pick which image was artist-prompted. If accuracy is at chance on complex prompts, many test items are visually unanswerable, and the measured headroom would overstate the limits of vision methods.","supporting_citations":[{"cited_title":"Investigating style similarity in diffusion models","cited_arxiv_id":null,"evidence_quote":"Supplies the contrastive style descriptor baseline and the real-artist reference set used for style comparison and prototypes."},{"cited_title":"Journeydb: A benchmark for generative im- age understanding","cited_arxiv_id":null,"evidence_quote":"Supplies real-user complex prompts and the Midjourney image-prompt pairs used in the benchmark."},{"cited_title":"Prototypical networks for few-shot learning","cited_arxiv_id":null,"evidence_quote":"Supplies the few-shot prototypical network method used as a strong trained baseline."},{"cited_title":"Learn- ing transferable visual models from natural language super- vision","cited_arxiv_id":null,"evidence_quote":"Provides the image and text features used for prototypes, retrieval, and similarity analyses."},{"cited_title":"Sdxl: Improving latent diffusion models for high-resolution image synthesis","cited_arxiv_id":null,"evidence_quote":"Provides the SDXL generator, which produces the largest and most thoroughly evaluated image subset."},{"cited_title":"Evaluating data attribution for text-to-image mod- els","cited_arxiv_id":null,"evidence_quote":"Supplies the data attribution baselines compared against style descriptors and trained classifiers."},{"cited_title":"Not with my name! inferring artists’ names of input strings employed by diffusion models","cited_arxiv_id":null,"evidence_quote":"Provides the closest prior artist-inference dataset, which the benchmark extends to a much larger scale."}],"review_version":2}