Pith. sign in

REVIEW 3 major objections 6 minor 75 references

Understanding Generative AI Capabilities in Everyday Image Editing Tasks

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Current AI image editors can handle about one-third of real-world editing requests, according to 4,359 human votes on 328 real requests.

desk verdict Valuable dataset and a solid VLM-judge result, but the headline 'one-third of requests' is a vote-level statistic, not a request-level estimate. read the letter →

arxiv 2505.16181 v2 pith:Q6KU3PXI submitted 2025-05-22 cs.CV cs.AI

classification cs.CVcs.AI
keywords imageeditinggenerativeAIhumanevaluationVLMasjudgeRedditdatasetrequesttaxonomyidentitypreservationperceptualquality
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper analyzes 83,000 real image-editing requests and 305,000 human-made edits collected from a large Reddit community over 12 years, together with edits produced by 49 AI editors. By running a controlled human preference study on 328 of these requests, it claims that human judges prefer the human-made edits over the AI edits 66% of the time, and that after weighting by how often each editing action appears in the corpus, only 33.35% of real-world requests can be satisfactorily handled by the best current AI editors. The paper further claims that AI editors are weakest on low-creativity, precise requests, that they frequently change the identity of people and animals, and that they make unrequested aesthetic touch-ups. Finally, it claims that vision-language model judges disagree sharply with human judges, so automated VLM ratings are not a reliable proxy for human preference.

What carries the argument

The load-bearing object is the PSR dataset itself: 82,976 real-world requests with 305,806 human edits, annotated with 15 user-intent editing actions, WordNet-based subject labels, and three creativity levels. The quantitative estimate is produced by a stratified sample (PSR-328) balanced across creativity levels, pairwise human and VLM judgments between human and AI edits, and a weighting formula: overall handleable percentage equals the sum over actions of the action's frequency in the full corpus times the AI win-plus-tie rate for that action.

What would settle it

Conduct a new human preference study on a larger random sample of requests drawn from the same PSR corpus, stratified by editing action to match the corpus distribution, and compute the action-weighted AI win-plus-tie rate; if the result falls clearly outside 30–36%, the paper's headline estimate does not generalize.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is an estimate of the current ceiling of automated image editing on real user requests: human raters prefer human edits over AI edits 66.0% of the time, and the weighted AI win-plus-tie rate across all editing actions is 33.35%. The per-action breakdown shows AI editors succeed most often on open-ended actions such as 'add', 'apply', and 'merge' and least often on spatially precise actions such as 'zoom', 'crop', and 'move'. A separate controlled experiment shows that both GPT-4o and Gemini-2.0-Flash gradually change a person's facial identity and body shape over a sequence of simple shirt-color edits, measured by growing DINOv2 feature distance, confirming identity preservation as a systematic weakness.

Load-bearing premise

The 33.35% figure assumes that the 328 requests selected for human evaluation represent every type of editing request in the full 83k corpus in proportions that allow the per-action success rates to be reweighted into a global estimate.

Editorial extensions

If this is right

  • If 33.35% is accurate, current text-to-image editors are not close to replacing human editors on general real-world requests, and benchmarks built from synthetic requests likely overstate real capability.
  • Identity preservation and avoidance of unrequested changes are the two highest-leverage targets for improving AI editors.
  • AI editors are comparatively more useful for open-ended, high-creativity requests than for precise, low-creativity ones.
  • VLM judges should not be used as the primary evaluator for image editing without human calibration, since they can exhibit strong model-specific biases (e.g., preferring one editor's output up to 85% of the time).

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the 33.35% figure is a snapshot of models available in early 2025; rerunning the same protocol on newer editors would give a direct measure of progress in closing the largest gaps.
  • Beyond the paper: the per-action weighting rests on the PSR-328 subset's action mix, and a study sampled to match the corpus's action frequencies could produce a different overall estimate, so the number should be treated as a point estimate rather than a tight bound.
  • Beyond the paper: the observed 'polish bias'—AI edits raising aesthetic scores even when not requested—might be part of why VLM judges over-prefer AI edits, since aesthetic quality is easier for a VLM to verify than faithfulness to the request.
  • Beyond the paper: the taxonomy of 15 intent-level actions could be reused as a template for building future training sets that match real user request distributions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces PSR, a large-scale dataset of 83k real-world image-editing requests collected from r/PhotoshopRequest over 2013-2025, together with 305k human-made edits. The authors annotate each request with WordNet-derived subjects, 15 editing-action labels, and three creativity levels. On a stratified subset of 328 requests, they collect 4,359 human preference votes comparing 1,644 human edits against 2,296 AI edits from 49 models, and report that human raters prefer human edits 66% of the time. The paper further estimates that AI editors can satisfactorily handle 33.35% of requests, finds that VLM judges (GPT-4o, o1, Gemini-2.0-Flash-Thinking) agree poorly with human raters and show strong biases (e.g., o1 prefers GPT-4o edits 83.9% of the time), and documents that AI editors tend to make unrequested aesthetic enhancements and struggle to preserve identity.

Significance. If the headline results hold, the dataset and findings would be a valuable contribution: PSR is the largest real-world image-editing request dataset with human edits, the action/creativity taxonomy is more user-intent-oriented than prior tool-based taxonomies, and the human study with 4.3k votes plus VLM-judge comparisons is unusually large for this task. The paper also ships code, qualitative examples, and a controlled identity-drift experiment with DINOv2 distances, all of which strengthen reproducibility. The observation that VLM judges can be severely biased in edit evaluation is practically important for automated benchmarking. However, the central quantitative claim of '33.35% of requests can be handled by AI' is not supported by the computation as reported, because the estimate is built from vote-level win/tie rates rather than request-level outcomes; this must be corrected before the headline result can be accepted.

major comments (3)
  1. [Sec. 5.4 and Table A8] The 33.35% estimate conflates vote-level win/tie rates with request-level handleability. The formula multiplies D_v (the request-level proportion of each action in the full dataset) by AI_v, which Table A8 defines as the percentage of pairwise votes won or tied by the AI across all requests and all 49 models. Because each request has roughly five human edits and seven AI edits, a single request contributes many votes and can have both AI wins and AI losses across those pairs. The stated definition of 'satisfactorily handled' is per request ('if it is rated Tie or AI wins by human raters'), so AI_v must be computed after aggregating votes to the request level (e.g., a request is handled if at least one AI edit ties or beats all human edits for that request). As written, 33.35% is the expected probability that a randomly selected AI-edit versus human-edit pair is won or tied by the AI, not the fraction of requests for which an AI editor produces a satisfactory result.
  2. [Abstract and Sec. 5.4] The abstract attributes the 33% figure to 'the best AI editors (including GPT-4o, Gemini-2.0-Flash, SeedEdit)', but the computation in Sec. 5.4 pools all 49 models, including many low-performance Hugging Face tools, as shown in Table A8. This is not equivalent to best-model performance. For example, Table A9 reports SeedEdit's per-action AI win+Tie rates that are substantially higher than the pooled rates for most actions (merge 63.0% vs 39.3%, add 52.0% vs 38.0%, delete 46.5% vs 34.9%). The paper should either report request-level handleability separately for each of the three SOTA models, or explicitly characterize 33.35% as an aggregate over all 49 models, not as the capability of the best AI editors.
  3. [Sec. 5.4 and Table A8] The weighted estimate is fragile because per-action vote counts are small for several actions and the sample was not stratified by action. PSR-328 was stratified by creativity level only, and Table A8 shows very low counts for clone (n=25), zoom (n=75), crop (n=129), and specialized operation (n=116). The formula weights these noisy per-action rates by the full-dataset action frequencies, but the paper reports only the point estimate 33.35% with no confidence intervals, bootstrap, or sensitivity analysis. Given the central role of this number, the authors should provide uncertainty quantification and show that the conclusions are robust to excluding low-count actions.
minor comments (6)
  1. [Sec. 5.1] The headline comparison (66.0% human preference vs. 25.8% AI preference) is reported without confidence intervals or significance tests. Because votes are clustered by request and by rater, cluster-robust intervals would be appropriate for assessing the strength of the claim.
  2. [Appendix E] The human rater pool is a convenience sample of 122 volunteers, one-third of whom are professional image editors. The paper should discuss how this composition may affect the measured human-preference rates and the generalizability of the comparison.
  3. [Fig. 1] The figure labels '2k AI edits' and '1.6k human edits', while the text and Table 3 report 2,296 AI edits and 1,644 human edits; the labels should be made consistent with the exact counts.
  4. [Sec. 6 Limitations] The limitations paragraph acknowledges LLM-based annotation biases and the exclusion of some unavailable models, which is appropriate. It should also mention the vote-level versus request-level aggregation issue in Sec. 5.4 and the small per-action sample sizes as limitations of the 33.35% estimate.
  5. [Table A8] The column header 'AI Win+Tie' is described as 'indicating the percentage of % requests that can already be handled', which is misleading given that the values are vote-level rates. The header should say 'percentage of votes' unless the quantity is redefined at the request level.
  6. [Sec. 5.4] The mathematical notation 'Pv v=1 Dv ×AIv' is garbled; the summation symbol and indices need to be typeset correctly, and the equation should explicitly state that D_v and AI_v are defined at different levels (request-level and vote-level, respectively).

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the 33% figure is an independently measured, reweighted vote statistic; its interpretation caveat is a validity concern, not a definitional loop.

full rationale

The paper's central claims are empirical and self-contained. Human ratings were collected independently via pairwise votes between human and AI edits; the AI win/tie rates are directly measured, not fitted to the target claim. The 33.35% figure in Sec. 5.4 is a weighted average of measured per-action AI win+tie rates using dataset action proportions; whether this vote-level statistic supports a request-level interpretation is a statistical validity/external-validity concern, not a circular reduction. VLM judges are explicitly compared against human votes, and using GPT-4o-mini for annotation and instruction rewriting does not make the evaluation circular because the objects of evaluation are AI editors, not the annotator. The only self-citation is reference [40], used to support the qualitative observation that VLMs are blind to image details; that observation is demonstrated by the paper's own examples and reasoning, so the citation is not load-bearing. No equation is defined in terms of the claimed result, and no fitted parameter is renamed as a prediction. Therefore, no significant circularity is present.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper does not fit any numeric parameters, so the free_parameters list is empty. The central claims rest on domain assumptions about representativeness (the subreddit, the wizard edits, the 328-request subset, the LLM labels) and about the validity of pairwise human votes as a measure of satisfaction. No invented entities are introduced.

assumptions (5)
  • domain assumption The /r/PhotoshopRequest community is representative of everyday image editing requests.
    Used in Sec. 1 and 3 to generalize from PSR to 'everyday' editing needs, and the abstract claims the request distribution reflects real-world needs.
  • domain assumption PSR-wizard human edits are high-quality reference edits.
    Used as the human baseline in the pairwise comparison in Sec. 5.1; the paper assumes these edits are a strong or typical human performance.
  • domain assumption The PSR-328 subset, stratified by creativity, yields unbiased per-action AI win/tie rates for the full dataset.
    Sec. 4 describes stratified sampling by creativity only, yet Sec. 5.4 combines per-action rates from PSR-328 with full-dataset action proportions to estimate the 33% figure.
  • domain assumption LLM-generated taxonomy labels (action, subject, creativity) from GPT-4o-mini are accurate enough for the analysis.
    Sec. 3.2 uses zero-shot GPT-4o-mini labeling; the paper notes in limitations that this may introduce biases or inaccuracies, but the main statistics rely on these labels.
  • domain assumption Human raters' pairwise preferences are a valid measure of request satisfaction.
    Sec. 5.1 and 5.4 define 'satisfactorily handled' based on human votes; no validation that pairwise win/tie corresponds to actual satisfaction.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Understanding Generative AI Capabilities in Everyday Image Editing Tasks." pith.science (2026). https://pith.science/paper/Q6KU3PXI

@misc{pith2026250516181,
  author       = {Pith},
  title        = {Pith review of: Understanding Generative AI Capabilities in Everyday Image Editing Tasks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q6KU3PXI}},
  note         = {Machine review of arXiv:2505.16181}
}
read the original abstract

Generative AI (GenAI) holds significant promise for automating everyday image editing tasks, especially following the recent release of GPT-4o on March 25, 2025. However, what subjects do people most often want edited? What kinds of editing actions do they want to perform (e.g., removing or stylizing the subject)? Do people prefer precise edits with predictable outcomes or highly creative ones? By understanding the characteristics of real-world requests and the corresponding edits made by freelance photo-editing wizards, can we draw lessons for improving AI-based editors and determine which types of requests can currently be handled successfully by AI editors? In this paper, we present a unique study addressing these questions by analyzing 83k requests from the past 12 years (2013-2025) on the Reddit community, which collected 305k PSR-wizard edits. According to human ratings, approximately only 33% of requests can be fulfilled by the best AI editors (including GPT-4o, Gemini-2.0-Flash, SeedEdit). Interestingly, AI editors perform worse on low-creativity requests that require precise editing than on more open-ended tasks. They often struggle to preserve the identity of people and animals, and frequently make non-requested touch-ups. On the other side of the table, VLM judges (e.g., o1) perform differently from human judges and may prefer AI edits more than human edits. Code and qualitative examples are available at: https://psrdataset.github.io

Figures

Figures reproduced from arXiv: 2505.16181 by the authors.

Figure 1
Figure 1. We propose PSR, the largest dataset of real [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Example cases from PSR dataset where PSR wizard edits were preferred by human raters over the AI edits (a-c) and samples where AI edits was preferred (d-f). (a): The human edit completes the request, but the AI edit removes the people, which was not requested. (b): The human edit completes the request, but the AI edit generates a similar image with people resembling those in the source image, although with different… view at source ↗
Figure 3
Figure 3. Over 12 years of Reddit data, delete, adjust, and add are the top-3 most wanted actions (a). Specifically, humans, body parts, text, and pets are the most frequent WordNet subjects for such common actions (b). While delete and adjust are the top-2 most common actions in the low - and medium -creativity requests, add takes up the largest share (34.9%) in high -creativity, e.g., inserting some “interesting” background… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: AIs can significantly alter both a person’s identity and the overall image quality through iterative image [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: o1 judge occasionally fails to notice details in edited images, here, overlooking the position of the hand and the configuration of the fingers. ten touch up human faces, making skin appear smoother and more polished (data not shown due to concerns of revealing identit…
Figure 6
Figure 6. Figure 6: AI models tend to increase overall image aesthetics. While both requests (a) and (b) ask for changes in the [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: AI edits have higher LAION aesthetics scores than the source image and human edits (a). AI edits are [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

75 extracted references · 67 canonical work pages

  1. [1]

    Samyadeep Basu, Mehrdad Saberi, Shweta Bhard- waj, Atoosa Malemir Chegini, Daniela Massiceti, Maziar Sanjabi, Shell Xu Hu, and Soheil Feizi

  2. [2]

    Jason Baumgartner, Savvas Zannettou, Brian Kee- gan, Megan Squire, and Jeremy Blackburn. 2020. The pushshift reddit dataset. InProceedings of the international AAAI conference on web and social media, volume 14, pages 830–839. 3

  3. [3]

    Jacqueline Brixey, Ramesh Manuvinakurike, Nham Le, Tuan Lai, Walter Chang, and Trung Bui. 2018. A system for automated image editing from natural lan- guage commands.arXiv preprint arXiv:1812.01083. 2

  4. [4]

    Tim Brooks, Aleksander Holynski, and Alexei A Efros. 2023. Instructpix2pix: Learning to follow image editing instructions. InProceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 18392–18402. 1, 2, 3, 9

  5. [5]

    Vladimir Bychkovsky, Sylvain Paris, Eric Chan, and Frédo Durand. 2011. Learning photographic global tonal adjustment with a database of input / output image pairs. InThe Twenty-Fourth IEEE Conference on Computer Vision and Pattern Recognition. 3

  6. [6]

    Dongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang, Yinuo Liu, Huichi Zhou, Qihui Zhang, Yao Wan, Pan Zhou, and Lichao Sun. 2024. Mllm-as- a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark. InForty-first Interna- tional Conference on Machine Learning. 7

  7. [7]

    Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, and 1 others. 2024. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling.arXiv preprint arXiv:2412.05271. 5

  8. [8]

    Zhihao Chen, Bin Hu, Chuang Niu, Tao Chen, Yuxin Li, Hongming Shan, and Ge Wang. 2024. Iqagpt: computed tomography image quality assessment with vision-language and chatgpt models.Frontiers in Radiology, 4:131. 2, 7

Show all 75 references
  1. [9]

    Marx D. 2025. If i was good at photoshop or graphic design, i’d do this side hustle. Reports r/Photo- shopRequest receives an average of 226 posts per day, based on Subredditstats and Social Rise analyt- ics as of April 2025. 2

  2. [10]

    Xiaoliang Dai, Ji Hou, Chih-Yao Ma, Sam Tsai, Jialiang Wang, Rui Wang, Peizhao Zhang, Simon Vandenhende, Xiaofang Wang, Abhimanyu Dubey, and 1 others. 2023. Emu: Enhancing image genera- tion models using photogenic needles in a haystack. arXiv preprint arXiv:2309.15807. 3, 9

  3. [11]

    Feedspot. 2025. Top 15 Photoshop Forums in 2025 — forums.feedspot.com. https://forums. feedspot.com/photoshop_forums/. [Ac- cessed 14-05-2025]. 2

  4. [12]

    Christiane Fellbaum. 2010. Wordnet. InTheory and applications of ontology: computer applications, pages 231–243. Springer. 3

  5. [13]

    Yuying Ge, Sijie Zhao, Chen Li, Yixiao Ge, and Ying Shan. 2024. Seed-data-edit technical report: A hybrid dataset for instructional image editing.arXiv preprint arXiv:2405.04007. 2

  6. [14]

    Google DeepMind. 2024. Gemini 2.0 flash thinking. https://deepmind.google/ technologies/gemini/. Experimental AI model. 2, 7

  7. [15]

    Simon Hentschel, Konstantin Kobs, and Andreas Hotho. 2022. Clip knows image aesthetics.Frontiers in Artificial Intelligence, 5:976235. 7

  8. [16]

    Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-or. 2023. Prompt-to-prompt image editing with cross-attention control. InThe Eleventh International Conference on Learning Representations. 3

  9. [17]

    Yi Huang, Jiancheng Huang, Yifan Liu, Mingfu Yan, Jiaxi Lv, Jianzhuang Liu, Wei Xiong, He Zhang, Shifeng Chen, and Liangliang Cao. 2024. Diffusion model-based image editing: A survey.arXiv preprint arXiv:2402.17525. 1

  10. [18]

    Hugging Face. 2023. Hugging Face Spaces. https://huggingface.co/spaces. Ac- cessed on March 02, 2025. 6

  11. [19]

    Fortune Business Insights. 2025. Ai image gen- erator market size, share & industry growth 2030. [Online; accessed 2025-03-06]. 1

  12. [20]

    Dongfu Jiang, Max Ku, Tianle Li, Yuansheng Ni, Shizhuo Sun, Rongqi Fan, and Wenhu Chen. 2025. Genai arena: An open evaluation platform for gen- erative models.Advances in Neural Information Processing Systems, 37:79889–79908. 3, 7

  13. [21]

    Kat Kampf and Nicole Brichtova

  14. [22]

    Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. 2023. Imagic: Text-based real image editing with diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6007–6017. 3 11

  15. [23]

    Yoonjeon Kim, Soohyun Ryu, Yeonsung Jung, Hyunkoo Lee, Joowon Kim, June Yong Yang, Jaery- ong Hwang, and Eunho Yang. 2024. Augmentation- driven metric for balancing preservation and modifi- cation in text-guided image editing.arXiv preprint arXiv:2410.11374. 3

  16. [24]

    Jari Korhonen and Junyong You. 2012. Peak signal- to-noise ratio revisited: Is simple beautiful? In 2012 Fourth international workshop on quality of multimedia experience, pages 37–38. IEEE. 3

  17. [25]

    Benno Krojer, Dheeraj Vattikonda, Luis Lara, Varun Jampani, Eva Portelance, Chris Pal, and Siva Reddy. 2025. Learning action and reasoning-centric image editing from videos and simulation.Advances in Neural Information Processing Systems, 37:38035– 38078. 3

  18. [26]

    Black Forest Labs. 2024. Flux. https:// github.com/black-forest-labs/flux. 55

  19. [27]

    Gierad P Laput, Mira Dontcheva, Gregg Wilen- sky, Walter Chang, Aseem Agarwala, Jason Linder, and Eytan Adar. 2013. Pixeltone: A multimodal interface for image editing. InProceedings of the SIGCHI Conference on Human Factors in Comput- ing Systems, pages 2185–2194. 2

  20. [28]

    Seongyun Lee, Seungone Kim, Sue Park, Gee- wook Kim, and Minjoon Seo. 2024. Prometheus- vision: Vision-language model as a judge for fine- grained evaluation. InFindings of the Association for Computational Linguistics ACL 2024, pages 11286– 11315. 7

  21. [29]

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation. InInternational conference on ma- chine learning, pages 12888–12900. PMLR. 3

  22. [30]

    Yiwei Ma, Jiayi Ji, Ke Ye, Weihuang Lin, Zhibin Wang, Yonghan Zheng, Qiang Zhou, Xiaoshuai Sun, and Rongrong Ji. 2024. I2ebench: A comprehensive benchmark for instruction-based image editing. In Advances in Neural Information Processing Systems (NeurIPS). 3

  23. [31]

    Ramesh Manuvinakurike, Jacqueline Brixey, Trung Bui, Walter Chang, Doo Soon Kim, Ron Artstein, and Kallirroi Georgila. 2018. Edit me: A corpus and a framework for understanding natural language image editing. InProceedings of the Eleventh In- ternational Conference on Language...

  24. [32]

    OpenAI. 2024. Hello GPT-4o. https:// openai.com/index/hello-gpt-4o/. Ac- cessed: March 2, 2025. 7, 9

  25. [33]

    OpenAI. 2024. Openai o1 model series. Accessed 2025-03-02. 2, 7

  26. [34]

    OpenAI. 2025. Gpt-4o mini: advancing cost- efficient intelligence. Accessed on February 13,

  27. [35]

    OpenAI. 2025. Introducing 4o image generation. Accessed: 2025-05-10. 2, 6

  28. [36]

    OpenAI. 2025. Introducing 4o Image Generation — openai.com. https://openai.com/index/ introducing-4o-image-generation/ . [Accessed 14-05-2025]. 1

  29. [37]

    OpenAI. 2025. Introducing ChatGPT pro. https://openai.com/index/ introducing-chatgpt-pro/. Accessed: February 28, 2025. 5

  30. [38]

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Fran- cisco Massa, Alaaeldin El-Nouby, and 1 others. 2023. Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193. 55

  31. [39]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, and 1 others. 2021. Learning transferable vi- sual models from natural language supervision. In International conference on machi...

  32. [40]

    Pooyan Rahmanzadehgervi, Logan Bolton, Mo- hammad Reza Taesiri, and Anh Totti Nguyen. 2024. Vision language models are blind. InProceedings of the Asian Conference on Computer Vision (ACCV), pages 18–34. 8

  33. [41]

    Christoph Schuhmann. 2023. Im- proved aesthetic predictor. https: //github.com/christophschuhmann/ improved-aesthetic-predictor. 7

  34. [42]

    Shelly Sheynin, Adam Polyak, Uriel Singer, Yu- val Kirstain, Amit Zohar, Oron Ashual, Devi Parikh, and Yaniv Taigman. 2024. Emu edit: Precise image editing via recognition and generation tasks. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti...

  35. [43]

    Jing Shi, Ning Xu, Trung Bui, Franck Dernoncourt, Zheng Wen, and Chenliang Xu. 2020. A benchmark and baseline for language-driven image editing. In Proceedings of the Asian Conference on Computer Vision. 2, 3

  36. [44]

    Jing Shi, Ning Xu, Yihang Xu, Trung Bui, Franck Dernoncourt, and Chenliang Xu. 2021. Learning by planning: Language-guided global image edit- ing. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13590–13599. 1, 3

  37. [45]

    Yichun Shi, Peng Wang, and Weilin Huang. 2024. Seededit: Align image re-generation to image editing. arXiv preprint arXiv:2411.06686. 1, 2, 6 12

  38. [46]

    SubredditStats.com. 2025. r/PhotoshopRequest Subreddit Stats (Photoshop Request) — subreddit- stats.com. https://subredditstats.com/ r/PhotoshopRequest. [Accessed 14-05-2025]. 2

  39. [47]

    Peter Sushko, Ayana Bharadwaj, Zhi Yang Lim, Vasily Ilin, Ben Caffee, Dongping Chen, Moham- madreza Salehi, Cheng-Yu Hsieh, and Ranjay Kr- ishna. 2025. Realedit: Reddit edits as a large-scale empirical dataset for image transformations. InPro- ceedings of the IEEE/CVF Conferen...

  40. [48]

    Hao Tan, Franck Dernoncourt, Zhe Lin, Trung Bui, and Mohit Bansal. 2019. Expressing visual relation- ships via language. InProceedings of the 57th An- nual Meeting of the Association for Computational Linguistics, pages 1873–1883, Florence, Italy. Asso- ciation for Computation...

  41. [49]

    Jianyi Wang, Kelvin CK Chan, and Chen Change Loy. 2023. Exploring clip for assessing the look and feel of images. InProceedings of the AAAI con- ference on artificial intelligence, volume 37, pages 2555–2563. 7

  42. [50]

    Su Wang, Chitwan Saharia, Ceslee Montgomery, Jordi Pont-Tuset, Shai Noy, Stefano Pellegrini, Ya- sumasa Onoe, Sarah Laszlo, David J Fleet, Radu Soricut, and 1 others. 2023. Imagen editor and edit- bench: Advancing and evaluating text-guided image inpainting. InProceedings of t...

  43. [51]

    Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612. 3

  44. [52]

    Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, Qiong Yan, Xiongkuo Min, Guangtao Zhai, and Weisi Lin

  45. [53]

    Tianyi Xiong, Xiyao Wang, Dong Guo, Qing- hao Ye, Haoqi Fan, Quanquan Gu, Heng Huang, and Chunyuan Li. 2024. Llava-critic: Learning to evaluate multimodal models.arXiv preprint arXiv:2410.02712. 7

  46. [54]

    Kai Zhang, Lingbo Mo, Wenhu Chen, Huan Sun, and Yu Su. 2023. Magicbrush: A manually anno- tated dataset for instruction-guided image editing. Advances in Neural Information Processing Systems, 36:31428–31449. 2, 3

  47. [55]

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018. The unreason- able effectiveness of deep features as a perceptual metric. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 586–

  48. [56]

    Shu Zhang, Xinyi Yang, Yihao Feng, Can Qin, Chia-Chih Chen, Ning Yu, Zeyuan Chen, Huan Wang, Silvio Savarese, Stefano Ermon, and 1 oth- ers. 2024. Hive: Harnessing human feedback for instructional visual editing. InProceedings of the IEEE/CVF Conference on Computer Vision and ...

  49. [57]

    Haozhe Zhao, Xiaojian Shawn Ma, Liang Chen, Shuzheng Si, Rujie Wu, Kaikai An, Peiyu Yu, Minjia Zhang, Qing Li, and Baobao Chang. 2024. Ultraedit: Instruction-based fine-grained image editing at scale. Advances in Neural Information Processing Systems, 37:3058–3093. 10

  50. [58]

    description

    Haozhe Zhao, Xiaojian Shawn Ma, Liang Chen, Shuzheng Si, Rujie Wu, Kaikai An, Peiyu Yu, Minjia Zhang, Qing Li, and Baobao Chang. 2025. Ultraedit: Instruction-based fine-grained image editing at scale. Advances in Neural Information Processing Systems, 37:3058–3093. 1, 2, 3 13 ...

  51. [62]

    Offensive content (discriminatory themes, extreme political content, severe profanity)

  52. [63]

    }, "image_editing_relevance

    Inappropriate but non-harmful content (crude humor, mild toilet humor, silly/whimsical inappropriate gestures, playful trolling). Note: Mild humorous or whimsical content that might be considered ’silly inappropriate’ (like tongue-in-cheek jokes, mild pranks, or playful memes)...

  53. [64]

    Evaluate both the original and clarified requests

  54. [65]

    Only include actions that are: - Explicitly stated OR - Logically necessary to achieve the described result

  55. [66]

    Consider the final image’s appearance to identify implicit actions

  56. [67]

    Return your response as a JSON object with an ’actions’ array containing only valid categories from the list provided

    Exclude actions that are: - Only potentially useful but not required - Vaguely related but not essential - Could be used as alternatives Focus only on actual image manipulation actions that match our predefined categories. Return your response as a JSON object with an ’actions...

  57. [68]

    Generate a list of candidate keywords that are **highly likely to already exist as a WordNet synset **

    **Candidate Keyword Selection ** - Extract the most relevant subject from the instruction and image. Generate a list of candidate keywords that are **highly likely to already exist as a WordNet synset **. Prioritize concrete nouns and commonly recognized entities

  58. [69]

    "" Figure A8: Prompt for assigning WordNet synsets to subjects of editing requests 23 Full Details for Subcategories (Part 1) categories_desc = {

    **WordNet Synset Matching ** - Use the refined candidate keywords to query WordNet, -confirming and selecting the best-fitting synset(s) for the subject. ### **Inputs Provided: ** - A textual instruction describing the image edit. - An image associated with the request. - The ...

  59. [70]

    Remove the red-eye effect from this photo

    **Request:** "Remove the red-eye effect from this photo." - The edits will be similar, focusing on correcting the eyes. - **Creativity Score: ** Low

  60. [71]

    Transform this portrait into a work of abstract art

    **Request:** "Transform this portrait into a work of abstract art." - There are countless ways to interpret and edit the image. - **Creativity Score: ** High

  61. [72]

    Adjust the brightness and contrast to enhance the image

    **Request:** "Adjust the brightness and contrast to enhance the image." - There are some variations in how this can be done. - **Creativity Score: ** Medium

  62. [73]

    Crop the image to focus on the main subject

    **Request:** "Crop the image to focus on the main subject." - Limited variations in the final image. - **Creativity Score: ** Low

  63. [74]

    Add a dramatic sky to this landscape photo

    **Request:** "Add a dramatic sky to this landscape photo." - Several ways to interpret a ’dramatic sky.’ - **Creativity Score: ** Medium

  64. [75]

    Reimagine this landscape in a fantasy setting

    **Request:** "Reimagine this landscape in a fantasy setting." - Numerous possibilities for how the image could be edited. - **Creativity Score: ** High. Provide the creativity score ( **Low**, **Medium**, or **High**) along with a brief explanation for your assessment. """ Fig...

  65. [2023]

    Editval: Benchmarking diffusion based text- guided image editing methods.arXiv preprint arXiv:2310.02426. 3, 7

  66. [2024]

    InProceedings of the 41st International Conference on Machine Learning, ICML’24

    Q-align: teaching lmms for visual scoring via discrete text-defined levels. InProceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org. 7

  67. [2025]

    https: //developers.googleblog.com/en/ experiment-with-gemini-20-flash-native-image-generation/

    Experiment with gemini 2.0 flash native image generation. https: //developers.googleblog.com/en/ experiment-with-gemini-20-flash-native-image-generation/ . Accessed: 2025-03-21. 2, 6, 7, 9

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.