REVIEW 4 major objections 4 minor 130 references
Clutter Detection and Removal by Multi-Objective Analysis for Photographic Guidance
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A camera guidance system identifies photo clutter by learning each object's negative contribution to a photo's aesthetic and content scores, and removes it with fast iterative GAN inpainting.
desk verdict A well-designed HCI camera-guidance paper whose central clutter measure is a plausible but unvalidated construction; worth reviewing seriously, but the evaluation does not yet support the 'accurate algorithms' claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the linear decomposition model of Eqs. (1)-(2) together with the contribution formula of Eq. (5). The model learns the overall scores $s_{\mathrm{aes}}$ and $s_{\mathrm{content}}$ as linear combinations of sub-image scores $s_{p_i}^{\mathrm{aes}}$ and $s_{p_i}^{\mathrm{content}}$, where the weights $\beta_i$ and $\gamma_i$ come from a SoftMax mixing network; a negative $q_i$ means the object lowers overall quality and is therefore clutter. A second mechanism is the iterative inpainting generator with an artifact-probability branch that re-fills regions whose generated pixels are uncertain, capped at three iterations for capture-time response and ten for high-fidelity processing after saving.
What would settle it
On a held-out set of photos, compare the system's clutter labels and $q_i$ rankings with independent human importance rankings; chance-level agreement would refute the claim. A stricter check is to edit objects out of the photos rather than blurring them and see whether the linear score prediction still holds.
Extended reading notes
Core claim
The central claim is that the overall aesthetic and content scores of an image can be learned as a weighted sum of the scores of sub-images, each formed by Gaussian-blurring one detected object, with a mixing network conditioned on the original image providing the weights through a softmax. With this decomposition, the contribution of the $i$-th object is $q_i = \beta_i(s_{\mathrm{aes}} - s_{p_i}^{\mathrm{aes}}) + \gamma_i(s_{\mathrm{content}} - s_{p_i}^{\mathrm{content}})$, and objects with negative $q_i$ are classified as clutter. The paper further claims that an iterative GAN inpainting algorithm, which fills a missing region by repeatedly accepting only the generated pixels that look most trustworthy, can remove such clutter in real time at capture time and with higher fidelity for saved photos. The user studies are presented as evidence that these clutter labels and inpainted previews are accurate enough to help amateurs identify distractions and produce better compositions with less effort.
Load-bearing premise
The paper assumes that a photo's aesthetic and content quality combine additively from its parts: blurring one object leaves all other contributions intact, so a weighted sum of sub-image scores equals the overall score, and the learned weights truthfully measure what each object contributes to the photo's quality.
Editorial extensions
If this is right
- Objects flagged with negative $q_i$ are classified as clutter, so the system can prompt the photographer to zoom in, move, or change orientation before shooting.
- For clutter that cannot be physically removed, the inpainting tool generates a preview with the clutter gone, letting users keep photos they would otherwise discard.
- In the user study, photos taken under the system's guidance were rated higher by users (median 6 vs 4 on a 7-point scale, $p = 0.007$) and preferred by photography experts in 15, 14, and 17 out of the paired comparisons.
- Users reported spending less time on composition decisions and feeling encouraged to attempt bolder scenes, because the system acted as a trusted second opinion.
- A double-click interaction lets users override the system's clutter classification, making the definition of clutter negotiable rather than fixed.
Reading between the lines
- If the additive decomposition holds, the same $q_i$ machinery could be run in reverse to recommend which objects to keep or emphasize, since large positive contributions identify elements that support the photo's intended story.
- The decomposition is essentially a credit-assignment scheme, so the same idea could extend to video frames or to other aesthetic attributes beyond a single aggregate score, such as color harmony or depth.
- The paper approximates object absence by Gaussian blurring; an inpainting-based leave-one-out scoring would be a more faithful counterfactual and a natural stress test of the linearity assumption, at higher computational cost.
- Because the model is trained on two aggregate human ratings, its notion of 'content' is tied to that dataset's definition of interestingness; a different content definition would likely change which objects count as clutter.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an in-camera guidance system that detects and removes clutter from photographs. The system uses Mask R-CNN for object detection, then estimates each object's contribution to the image's aesthetic and content scores via a linear decomposition of sub-images in which the object is blurred (Eqs. 1-2). Objects with a negative contribution q_i (Eq. 5) are labeled as clutter. The system also provides an iterative GAN-based inpainting algorithm to remove selected objects. The evaluation consists of user studies with 18 participants, including module-level tests on clutter detection and inpainting, and a system-level comparison to a default camera app, with photography experts judging the resulting photos.
Significance. If the core contribution-estimation assumption holds, the system could offer a useful, practical tool for amateur photographers by providing real-time, interpretable guidance on which objects to remove and how to reframe a scene. The system is well-motivated by a design survey and integrates several existing techniques (object detection, aesthetic scoring, GAN inpainting) into a cohesive interface. The paper is also transparent about the algorithm's limitations and includes user feedback that is largely positive. However, the central claim that q_i measures an object's causal contribution to aesthetics and content is not validated outside the model's own definition, and the evaluation is primarily qualitative with no baseline comparison. The significance is therefore conditional on addressing these validation gaps.
major comments (4)
- [Sec. 4.1, Eq. (5)] The contribution q_i is computed from sub-images p_i in which object i is only Gaussian-blurred (13x13 kernel, variance 1), not removed. The text claims that s_p_i is 'the quality of the image without the ith object,' but a blurred object remains visually present and may create an artificial smudge artifact. Thus q_i measures the aesthetic/content effect of blurring the object, not the causal contribution of the object itself. The author should either replace the blur with a more faithful object-removal operation (e.g., using the inpainting model) or provide evidence that blurred-object scores approximate removed-object scores. At minimum, this limitation must be clearly acknowledged, as it directly affects the validity of every clutter label produced by Eq. (5).
- [Sec. 4.1, Eqs. (1)-(2)] The linear decomposition is trained only to minimize prediction error of the overall aesthetic and content scores (Eqs. 3-4). No supervision signals that the learned weights beta_i and gamma_i correspond to an object's true importance or causal contribution. Because sub-images of the same photo are highly correlated, many different weight configurations can fit the overall scores equally well, so the learned weights may be arbitrary decompositions. The paper provides no validation of the decomposition against human judgments of object importance, which is essential given that the clutter classification and all downstream guidance rely on q_i. An appropriate validation could involve asking human annotators to rank object importance and comparing those rankings with the model's contributions, or an ablation that tests whether the learned weights generalize across different object configurations.
- [Sec. 5.2.1] The evaluation of clutter detection is circular. The model's own q_i defines clutter (negative q_i), and the user study measures user agreement with these model-generated labels, after the author manually flipped 6 of 67 clutter labels. This does not provide independent evidence that the model identifies clutter in any objective sense. The reported 92.05% agreement is computed after those flips and therefore overstates the model's performance. A proper evaluation would compare the model's labels against an external ground truth (e.g., clutter labels provided by independent human annotators who have not seen the model output) or against a baseline clutter-detection method such as Fried et al. [28], which the paper discusses but does not compare with experimentally.
- [Sec. 5.3 and Sec. 5.2] The quantitative evaluation of the complete system is limited to a comparison with the default camera app using a small sample (18 participants, 3 experts). The significant Wilcoxon result (V=3.5, p=0.007) suggests an overall preference for the system, but it does not isolate the contribution of the proposed clutter-detection or inpainting algorithms. The modular tests rely on Likert scales and subjective feedback, with no statistical tests or inter-rater reliability metrics. To support the paper's central claim that the system 'allows users to better identify distractions and take higher quality images within less time,' the author should provide ablations that isolate the effect of the contribution-estimation method (e.g., comparing against a random or saliency-based baseline) and report effect sizes and confidence intervals for the user-study results.
minor comments (4)
- [Sec. 4.1, Eq. (2)] The notation for the content-score weights is inconsistent: Eq. (2) uses gamma_i, but the text later writes 'vector {@code γ} = ⟨γ,γ 2,...,γ k⟩', which should be '⟨γ_1,γ_2,...,γ_k⟩'.
- [Sec. 2.2] Reference [49] is cited as 'L. et al.' in the text, which is incomplete; the full citation in the bibliography is 'Jane L. et al.' and the text should use the first author's full name or a consistent abbreviation.
- [Sec. 2.3] The related-work paragraph lists reference [107] twice: 'composition features [8, 9, 20, 86, 107, 107, 120]'. This is a duplicate and should be cleaned up.
- [Sec. 3.2] The sentence 'In the next section, we describe in detail how to develop algorithms...' appears at the end of Sec. 3.2, but the algorithms are actually described in Sec. 4, which is more than one section later. The text should say 'In the following section' or similar.
Circularity Check
Clutter is defined as the sign of the model's own q_i, and the reported agreement is computed after the author manually flipped labels; the core clutter-detection claim is therefore partially circular.
-
self definitional
[Section 4.1.1, immediately after Eq. (5)]
"A negative value of 𝑞𝑖 indicates that the 𝑖th object lowers the overall quality of the photo. Therefore, we classify these objects as clutter and the others as normal objects."
Clutter is stipulated to be exactly the sign of q_i, and q_i is the system's own output. Thus the algorithm's clutter classification is true by definition for any input image: objects with negative q_i are called clutter because the paper defines clutter that way. No external clutter ground truth is needed to produce the label, and the label cannot be wrong under the definition. The only independent check would be a human-judgment comparison, but that comparison is weakened by the author flips reported in the user study.
-
fitted input called prediction
[Section 4.1, Eqs. (1)-(5), and Section 4.1.1]
"The semantic meaning of these scores is the quality of the image without the 𝑖th object. With the estimated overall scores for the original image, 𝑠aes and 𝑠content, we can calculate the contribution of the 𝑖th object to the whole image as: 𝑞𝑖 = 𝛽𝑖(𝑠aes − 𝑠𝑝𝑖 aes) + 𝛾𝑖(𝑠content − 𝑠𝑝𝑖 content). (5)"
Both s_aes and s_p_i_aes are outputs of the same trained network; s_aes is the Eq. (1) linear combination of the sub-image scores. Eq. (5) is therefore an algebraic identity among fitted variables and never references the human ground-truth y_aes/y_content used in the training losses. Calling this quantity 'the contribution of the i-th object' and then using its sign as the clutter label imports a fitted decomposition as if it were an independently measured effect. The 'prediction' of object contribution is a rearrangement of the model's own predictions rather than an estimate validated against external importance judgments.
1 more flagged steps
-
fitted input called prediction
[Section 5.2.1, Clutter Detection user study]
"For the cluttered objects, the category for 6 (8.96%) of them was flipped by the author. These flips lead to an overall agreement degree of 92.05% between our clutter detection algorithm and the judgment of experimenters."
The reported 92.05% agreement is computed after the author, who built the model, manually changed the model's labels on six objects. The statistic is therefore not the unmodified algorithm's blind agreement with the experimenters; it is a post-hoc fitted agreement in which the author corrected the very outputs being evaluated. The headline accuracy is partly forced by the author's own flips, so it cannot serve as independent validation of the clutter-detection claim.
full rationale
The main circularity lies in the clutter-distinguishment module. The paper defines clutter as q_i < 0, so the algorithm's clutter output is a stipulated definition rather than an empirical prediction; the only external check is a user study in which the author manually flipped 6 labels before computing the 92.05% agreement, making part of the reported accuracy a fitted evaluation. In addition, q_i in Eq. (5) is computed entirely from the model's own predicted scores (Eqs. (1)-(2)) rather than from the ground-truth aesthetic/content labels used in the loss, so the claimed contribution is an algebraic rearrangement of fitted quantities presented as a measured effect. These issues make the central clutter-detection claim partially circular. The inpainting component is independently trained on ImageNet/Places2 with original images as ground truth and is not circular; the system-level photo comparison and Likert results are independent user judgments, although they inherit the questionable clutter labels. No load-bearing self-citation or uniqueness-theorem pattern is present; the author's self-citations [105, 106] are contextual only. Overall score: 6.
Assumptions & free parameters
free parameters (4)
- Decomposition weights beta_i, gamma_i =
Learned by mixing network f_mix with SoftMax activation; trained on AADB
- Gaussian blur kernel size and variance =
13x13, sigma=1
- Loss weight lambda_aes =
1
- Inpainting maximum iterations =
3 (capture-time), 10 (off-line)
assumptions (4)
- domain assumption The overall aesthetic/content score of an image is a linear combination of scores of sub-images with each object blurred (Eqs. 1-2).
- domain assumption A sub-image with a Gaussian-blurred object region represents the image without that object.
- domain assumption A negative q_i indicates the object is clutter that lowers photo quality.
- domain assumption The iterative inpainting with at most 3 iterations provides visually plausible results in real time.
Cite this review
Pith. "Pith review of Clutter Detection and Removal by Multi-Objective Analysis for Photographic Guidance." pith.science (2026). https://pith.science/paper/LLZ2DSNG
@misc{pith2026250714553,
author = {Pith},
title = {Pith review of: Clutter Detection and Removal by Multi-Objective Analysis for Photographic Guidance},
year = {2026},
howpublished = {\url{https://pith.science/paper/LLZ2DSNG}},
note = {Machine review of arXiv:2507.14553}
}
read the original abstract
Clutter in photos is a distraction preventing photographers from conveying the intended emotions or stories to the audience. Photography amateurs frequently include clutter in their photos due to unconscious negligence or the lack of experience in creating a decluttered, aesthetically appealing scene for shooting. We are thus motivated to develop a camera guidance system that provides solutions and guidance for clutter identification and removal. We estimate and visualize the contribution of objects to the overall aesthetics and content of a photo, based on which users can interactively identify clutter. Suggestions on getting rid of clutter, as well as a tool that removes cluttered objects computationally, are provided to guide users to deal with different kinds of clutter and improve their photographic work. Two technical novelties underpin interactions in our system: a clutter distinguishment algorithm with aesthetics evaluations for objects and an iterative image inpainting algorithm based on generative adversarial nets that reconstructs missing regions of removed objects for high-resolution images. User studies demonstrate that our system provides flexible interfaces and accurate algorithms that allow users to better identify distractions and take higher quality images within less time.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[28]
Ohad Fried, Eli Shechtman, Dan B Goldman, and Adam Finkelstein. 2015. Finding distractors in images. In Proceedings of the IEEE Conference on Computer Vision and pattern Recognition. 1703–1712
2015
-
[1]
Remove clutter from your photography
2013. Remove clutter from your photography. https://digital-photography- school.com/remove-clutter/
2013
-
[2]
Macarena Aspillaga. 1991. Screen design: Location of information and its effects on learning. Journal of Computer-Based Instruction 18, 3 (1991), 89–92
1991
-
[3]
Tunç Ozan Aydın, Aljoscha Smolic, and Markus Gross. 2014. Automated aes- thetic analysis of photographic images. IEEE transactions on visualization and computer graphics 21, 1 (2014), 31–42
2014
-
[4]
Soonmin Bae, Aseem Agarwala, and Frédo Durand. 2010. Computational repho- tography. ACM Trans. Graph. 29, 3 (2010), 24–1
2010
-
[5]
Coloma Ballester, Marcelo Bertalmio, Vicent Caselles, Guillermo Sapiro, and Joan Verdera. 2001. Filling-in by joint interpolation of vector fields and gray levels. IEEE transactions on image processing 10, 8 (2001), 1200–1211
2001
-
[6]
Connelly Barnes, Eli Shechtman, Adam Finkelstein, and Dan B Goldman. 2009. PatchMatch: A randomized correspondence algorithm for structural image editing. ACM Trans. Graph. 28, 3 (2009), 24
2009
-
[7]
Marcelo Bertalmio, Guillermo Sapiro, Vincent Caselles, and Coloma Ballester
Show all 130 references
-
[8]
Subhabrata Bhattacharya, Rahul Sukthankar, and Mubarak Shah. 2010. A frame- work for photo-quality assessment and enhancement based on visual aesthetics. In Proceedings of the 18th ACM international conference on Multimedia . 271–280
2010
-
[9]
Subhabrata Bhattacharya, Rahul Sukthankar, and Mubarak Shah. 2011. A holistic approach to aesthetic enhancement of photographs. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) 7, 1 (2011), 1–21
2011
-
[10]
Jennifer Blessing. 2002. Gina Pane’s Witnesses: The Audience and Photography. Performance Research 7, 4 (2002), 14–26
2002
-
[11]
Bruce Block. 2020. The visual story: Creating the visual structure of film, TV, and digital media. Routledge
2020
-
[12]
Scott Carter, John Adcock, John Doherty, and Stacy Branham. 2010. NudgeCam: Toward targeted, higher quality media capture. In Proceedings of the 18th ACM international conference on Multimedia . 615–618
2010
-
[13]
Filippos Christianos, Lukas Schäfer, and Stefano Albrecht. 2020. Shared expe- rience actor-critic for multi-agent reinforcement learning. Advances in neural information processing systems 33 (2020), 10707–10717
2020
-
[14]
Neil Cohn. 2013. Visual narrative structure. Cognitive science 37, 3 (2013), 413–452
2013
-
[15]
Soheil Darabi, Eli Shechtman, Connelly Barnes, Dan B Goldman, and Pradeep Sen. 2012. Image melding: Combining inconsistent images using patch-based synthesis. ACM Transactions on graphics (TOG) 31, 4 (2012), 1–10
2012
-
[16]
Ritendra Datta, Dhiraj Joshi, Jia Li, and James Z Wang. 2006. Studying aesthetics in photographic images using a computational approach. InEuropean conference on computer vision. Springer, 288–301
2006
-
[17]
Ritendra Datta, Jia Li, and James Z Wang. 2007. Learning the consensus on visual quality for next-generation image management. In Proceedings of the 15th ACM international conference on Multimedia . 533–536
2007
-
[18]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition . Ieee, 248–255
2009
-
[19]
Yubin Deng, Chen Change Loy, and Xiaoou Tang. 2017. Image aesthetic assess- ment: An experimental survey. IEEE Signal Processing Magazine 34, 4 (2017), 80–106
2017
-
[20]
Sagnik Dhar, Vicente Ordonez, and Tamara L Berg. 2011. High level describable attributes for predicting aesthetics and interestingness. In CVPR 2011. IEEE, 1657–1664
2011
-
[21]
Heng Dong, Tonghan Wang, Jiayuan Liu, Chi Han, and Chongjie Zhang. 2021. Birds of a feather flock together: A close look at cooperation emergence via multi-agent rl. arXiv preprint arXiv:2104.11455 (2021)
2021 arXiv
-
[22]
Heng Dong, Tonghan Wang, Jiayuan Liu, and Chongjie Zhang. 2022. Low- rank modular reinforcement learning via muscle synergy. Advances in Neural Information Processing Systems (NeurIPS) 35 (2022), 19861–19873
2022
-
[23]
Heng Dong, Junyu Zhang, Tonghan Wang, and Chongjie Zhang. 2023. Symmetry-aware robot design with structured subgroups. In International Con- ference on Machine Learning (ICML) . PMLR, 8334–8355
2023
-
[24]
Will Eisner. 2008. Graphic storytelling and visual narrative . WW Norton & Company
2008
-
[25]
Jakob Foerster, Ioannis Alexandros Assael, Nando de Freitas, and Shimon White- son. 2016. Learning to communicate with deep multi-agent reinforcement learning. In Advances in Neural Information Processing Systems . 2137–2145
2016
-
[26]
Jakob N Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. 2018. Counterfactual multi-agent policy gradients. In Thirty-Second AAAI Conference on Artificial Intelligence
2018
-
[27]
Ohad Fried, Jingwan Lu, Jianming Zhang, Radomír Mech, Jose Echevarria, Pat Hanrahan, James A Landay, et al. 2020. Adaptive photographic composition guidance. In CHI
2020
-
[29]
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. Advances in neural information processing systems 27 (2014)
2014
-
[30]
R Scott Grabinger. 1993. Computer screen designs: Viewer judgments. Educa- tional Technology Research and Development 41, 2 (1993), 35–73
1993
-
[31]
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. 2017. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision . 2961–2969
2017
-
[32]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. InProceedings of the IEEE conference on computer vision and pattern recognition . 770–778
2016
-
[33]
Jesse M Heines. 1984. Screen design strategies for computer-assisted instruction . Digital Press
1984
-
[34]
Safwan Hossain, Tonghan Wang, Tao Lin, Yiling Chen, David C Parkes, and Haifeng Xu. 2024. Multi-sender persuasion: a computational perspective. In International Conference on Machine Learning (ICML) . 18944–18971
2024
-
[35]
Dima Ivanov, Paul Dütting, Inbal Talgam-Cohen, Tonghan Wang, and David C Parkes. 2024. Principal-Agent Reinforcement Learning: Orchestrating AI Agents with Contracts. arXiv preprint arXiv:2407.18074 (2024)
2024 arXiv
-
[36]
LE Jane, Ohad Fried, and Maneesh Agrawala. 2019. Optimizing Portrait Lighting at Capture-Time Using a 360 Camera as a Light Probe.. In UIST. 221–232
2019
-
[37]
Hyeongnam Jang and Jong-Seok Lee. 2021. Analysis of Deep Features for Image Aesthetic Assessment. IEEE Access 9 (2021), 29850–29861
2021
-
[38]
Jiechuan Jiang, Chen Dun, Tiejun Huang, and Zongqing Lu. 2019. Graph Convolutional Reinforcement Learning. In International Conference on Learning Representations
2019
-
[39]
Yipeng Kang, Tonghan Wang, Qianlan Yang, Xiaoran Wu, and Chongjie Zhang
-
[40]
Yueying Kao, Kaiqi Huang, and Steve Maybank. 2016. Hierarchical aesthetic quality assessment using deep convolutional neural networks. Signal Processing: Image Communication 47 (2016), 500–510
2016
-
[41]
Yan Ke, Xiaoou Tang, and Feng Jing. 2006. The design of high-level features for photo quality assessment. In2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06) , Vol. 1. IEEE, 419–426
2006
-
[42]
Richard S Kelster and Glen R Gallaway. 1983. Making software user friendly: An assessment of data entry performance. In Proceedings of the Human Factors Society Annual Meeting, Vol. 27. SAGE Publications Sage CA: Los Angeles, CA, 1031–1034
1983
-
[43]
Jongyoo Kim, Jiaolong Yang, and Xin Tong. 2021. Learning High-Fidelity Face Texture Completion Without Complete Face Texture. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 13990–13999. 12
2021
-
[44]
Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimiza- tion. In Proceedings of the International Conference on Learning Representations (ICLR)
2015
-
[45]
Shu Kong, Xiaohui Shen, Zhe Lin, Radomir Mech, and Charless Fowlkes. 2016. Photo aesthetics ranking network with attributes and content adaptation. In European Conference on Computer Vision . Springer, 662–679
2016
-
[46]
Jakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen, Fanglei Sun, Jun Wang, and Yaodong Yang. 2021. Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning. In International Conference on Learning Representations
2021
-
[47]
Masaaki Kurosu and Kaori Kashimura. 1995. Apparent usability vs. inherent usability: experimental analysis on the determinants of the apparent usability. In Conference companion on Human factors in computing systems . 292–293
1995
-
[48]
Jaime L Kurtz and Sonja Lyubomirsky. 2013. Happiness promotion: Using mindful photography to increase positive emotion and appreciation. (2013)
2013
-
[49]
Zhai, Jose Echevarria, Ohad Fried, Pat Hanrahan, and James A
Jane L., Kevin Y. Zhai, Jose Echevarria, Ohad Fried, Pat Hanrahan, and James A. Landay. 2021. Dynamic Guidance for Decluttering Photographic Compositions. In The 34th Annual ACM Symposium on User Interface Software and Technology (Virtual Event, USA) (UIST ’21). Association fo...
2021
-
[50]
Congcong Li, Andrew Gallagher, Alexander C Loui, and Tsuhan Chen. 2010. Aesthetic quality assessment of consumer photos with faces. In 2010 IEEE Inter- national Conference on Image Processing . IEEE, 3221–3224
2010
-
[51]
Jia Li, Zhaoyang Li, Jie Cao, Xingguang Song, and Ran He. 2021. FaceInpainter: High Fidelity Face Adaptation to Heterogeneous Domains. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5089–5098
2021
-
[52]
Qifan Li and Daniel Vogel. 2017. Guided selfies using models of portrait aes- thetics. In Proceedings of the 2017 Conference on Designing Interactive Systems . 179–190
2017
-
[53]
Liang Liao, Jing Xiao, Zheng Wang, Chia-Wen Lin, and Shin’ichi Satoh. 2021. Image inpainting guided by coherence priors of semantics and textures. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6539–6548
2021
-
[54]
Arnaud Lienhard, Patricia Ladret, and Alice Caplier. 2015. Low level features for quality assessment of facial images. In10th International Conference on Computer Vision Theory and Applications (VISAPP 2015) . 545–552
2015
-
[55]
Guilin Liu, Fitsum A Reda, Kevin J Shih, Ting-Chun Wang, Andrew Tao, and Bryan Catanzaro. 2018. Image inpainting for irregular holes using partial con- volutions. In Proceedings of the European conference on computer vision (ECCV) . 85–100
2018
-
[56]
Hongyu Liu, Bin Jiang, Yibing Song, Wei Huang, and Chao Yang. 2020. Rethink- ing image inpainting via a mutual encoder-decoder with feature equalizations. In European Conference on Computer Vision . Springer, 725–741
2020
-
[57]
Hongyu Liu, Ziyu Wan, Wei Huang, Yibing Song, Xintong Han, and Jing Liao
-
[58]
Li-Yun Lo and Ju-Chin Chen. 2012. A statistic approach for photo quality assess- ment. In 2012 International Conference on Information Security and Intelligent Control. IEEE, 107–110
2012
-
[59]
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017. Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems. 6379–6390
2017
-
[60]
Xin Lu, Zhe Lin, Hailin Jin, Jianchao Yang, and James Z Wang. 2015. Rating image aesthetics using deep learning. IEEE Transactions on Multimedia 17, 11 (2015), 2021–2034
2015
-
[61]
Xin Lu, Zhe Lin, Xiaohui Shen, Radomir Mech, and James Z Wang. 2015. Deep multi-patch aggregation network for image style, aesthetics, and quality esti- mation. In Proceedings of the IEEE international conference on computer vision . 990–998
2015
-
[62]
Yiwen Luo and Xiaoou Tang. 2008. Photo and video quality evaluation: Focusing on the subject. In European Conference on Computer Vision . Springer, 386–399
2008
-
[63]
Shuai Ma, Zijun Wei, Feng Tian, Xiangmin Fan, Jianming Zhang, Xiaohui Shen, Zhe Lin, Jin Huang, Radomír Měch, Dimitris Samaras, et al . 2019. SmartEye: assisting instant photo taking via integrating user preference with deep view proposal network. In Proceedings of the 2019 CH...
2019
-
[64]
Long Mai, Hailin Jin, and Feng Liu. 2016. Composition-preserving deep photo aesthetics assessment. In Proceedings of the IEEE conference on computer vision and pattern recognition. 497–506
2016
-
[65]
Luca Marchesotti, Florent Perronnin, Diane Larlus, and Gabriela Csurka. 2011. Assessing the aesthetic quality of photographs using generic image descriptors. In 2011 international conference on computer vision . IEEE, 1784–1791
2011
-
[66]
Luca Marchesotti, Florent Perronnin, and France Meylan. 2013. Learning beau- tiful (and ugly) attributes.. In BMVC, Vol. 7. 1–11
2013
-
[67]
Jon McCormack and Andy Lomas. 2020. Understanding aesthetic evaluation using deep learning. arXiv preprint arXiv:2004.06874 (2020)
2020 arXiv
-
[68]
Jon McCormack and Andy Lomas. 2021. Deep learning of individual aesthetics. Neural Computing and Applications 33, 1 (2021), 3–17
2021
-
[69]
Hiroko Mitarai, Yoshihiro Itamiya, and Atsuo Yoshitaka. 2013. Interactive photo- graphic shooting assistance based on composition and saliency. In International Conference on Computational Science and Its Applications . Springer, 348–363
2013
-
[70]
Naila Murray, Luca Marchesotti, and Florent Perronnin. 2012. AVA: A large- scale database for aesthetic visual analysis. In2012 IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 2408–2415
2012
-
[71]
Kamyar Nazeri, Eric Ng, Tony Joseph, Faisal Z Qureshi, and Mehran Ebrahimi
-
[72]
Masashi Nishiyama, Takahiro Okabe, Imari Sato, and Yoichi Sato. 2011. Aesthetic quality classification of photographs based on color harmony. In CVPR 2011. IEEE, 33–40
2011
-
[73]
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros. 2016. Context encoders: Feature learning by inpainting. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2536–2544
2016
-
[74]
Bei Peng, Tabish Rashid, Christian Schroeder de Witt, Pierre-Alexandre Kami- enny, Philip Torr, Wendelin Böhmer, and Shimon Whiteson. 2021. Facmac: Factored multi-agent centralised policy gradients. Advances in Neural Informa- tion Processing Systems 34 (2021), 12208–12221
2021
-
[75]
Jialun Peng, Dong Liu, Songcen Xu, and Houqiang Li. 2021. Generating diverse structure for image inpainting with hierarchical VQ-VAE. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10775–10784
2021
-
[76]
Rong-Jun Qin, Feng Chen, Tonghan Wang, Lei Yuan, Xiaoran Wu, Yipeng Kang, Zongzhang Zhang, Chongjie Zhang, and Yang Yu. 2022. Multi-Agent Policy Transfer via Task Relationship Modeling. InDeep Reinforcement Learning Workshop NeurIPS 2022
2022
-
[77]
Tabish Rashid, Gregory Farquhar, Bei Peng, and Shimon Whiteson. 2020. Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning. Advances in Neural Information Processing Systems 33 (2020)
2020
-
[78]
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. 2018. QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning. In International Conference on Machine Learning . 4292–4301
2018
-
[79]
Yogesh Singh Rawat and Mohan S Kankanhalli. 2015. Context-aware photog- raphy learning for smart mobile devices. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) 12, 1s (2015), 1–24
2015
-
[80]
Amanpreet Singh, Tushar Jain, and Sainbayar Sukhbaatar. 2019. Learning when to Communicate at Scale in Multiagent Cooperative and Competitive Tasks. In Proceedings of the International Conference on Learning Representations (ICLR)
2019
-
[81]
Hsiao-Hang Su, Tse-Wei Chen, Chieh-Chi Kao, Winston H Hsu, and Shao-Yi Chien. 2011. Scenic photo quality assessment with bag of aesthetics-preserving features. In Proceedings of the 19th ACM international conference on Multimedia . 1213–1216
2011
-
[82]
Rongju Sun, Zhouhui Lian, Yingmin Tang, and Jianguo Xiao. 2015. Aesthetic vi- sual quality evaluation of Chinese handwritings. In Twenty-Fourth International Joint Conference on Artificial Intelligence
2015
-
[83]
Xiaoshuai Sun, Hongxun Yao, Rongrong Ji, and Shaohui Liu. 2009. Photo assessment based on computational visual attention model. In Proceedings of the 17th ACM international conference on Multimedia . 541–544
2009
-
[84]
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al. 2018. Value-decomposition networks for cooperative multi- agent learning based on team reward. In Proceedin...
2018
-
[85]
Michael Szabo and Heather Kanuka. 1999. Effects of violating screen design principles of balance, unity, and focus on recall learning, study time, and com- pletion rates. Journal of educational multimedia and hypermedia 8, 1 (1999), 23–42
1999
-
[86]
Xiaoou Tang, Wei Luo, and Xiaogang Wang. 2013. Content-based photo quality assessment. IEEE Transactions on Multimedia 15, 8 (2013), 1930–1943
2013
-
[87]
Xinmei Tian, Zhe Dong, Kuiyuan Yang, and Tao Mei. 2015. Query-dependent aesthetic model with deep learning for photo quality assessment. IEEE Transac- tions on Multimedia 17, 11 (2015), 2035–2048
2015
-
[88]
SC Toh. 1998. Cognitive and motivational effects of two multimedia simulation presentation modes on science learning. Unpublished Ph. D. thesis, University of Science Malaysia (1998)
1998
-
[89]
Hanghang Tong, Mingjing Li, Hong-Jiang Zhang, Jingrui He, and Changshui Zhang. 2004. Classification of digital photos taken by photographers or home users. In Pacific-Rim Conference on Multimedia . Springer, 198–205
2004
-
[90]
Noam Tractinsky. 1997. Aesthetics and apparent usability: empirically assessing cultural and methodological issues. InProceedings of the ACM SIGCHI Conference on Human factors in computing systems . 115–122
1997
-
[91]
Thomas S Tullis. 1981. An evaluation of alphanumeric, graphic, and color information displays. Human Factors 23, 5 (1981), 541–550
1981
-
[92]
Thomas Stuart Tullis. 1984. Predicting the usability of alphanumeric displays . Ph. D. Dissertation. Rice University. 13
1984
-
[93]
Tengfei Wang, Hao Ouyang, and Qifeng Chen. 2021. Image inpainting with external-internal learning and monochromic bottleneck. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5120–5129
2021
-
[94]
Wenguan Wang and Jianbing Shen. 2017. Deep cropping via attention box prediction and aesthetics assessment. In Proceedings of the IEEE International Conference on Computer Vision . 2186–2194
2017
-
[95]
Weining Wang, Mingquan Zhao, Li Wang, Jiexiong Huang, Chengjia Cai, and Xiangmin Xu. 2016. A multi-scene deep learning model for image aesthetic evaluation. Signal Processing: Image Communication 47 (2016), 511–518
2016
-
[96]
Yi Wang, Xin Tao, Xiaojuan Qi, Xiaoyong Shen, and Jiaya Jia. 2018. Image in- painting via generative multi-column convolutional neural networks. Advances in neural information processing systems 31 (2018)
2018
-
[97]
Zhangyang Wang, Shiyu Chang, Florin Dolcos, Diane Beck, Ding Liu, and Thomas S Huang. 2016. Brain-inspired deep networks for image aesthetics assessment. arXiv preprint arXiv:1601.04155 (2016)
2016 arXiv
-
[98]
Muning Wen, Jakub Kuba, Runji Lin, Weinan Zhang, Ying Wen, Jun Wang, and Yaodong Yang. 2022. Multi-agent reinforcement learning is a sequence modeling problem. Advances in Neural Information Processing Systems 35 (2022), 16509–16521
2022
-
[99]
Steve J Westerman, S Kaur, C Dukes, and J Blomfield. 2007. Creative industrial design and computer-based image retrieval: The role of aesthetics and affect. In International Conference on Affective Computing and Intelligent Interaction . Springer, 618–629
2007
-
[100]
Wendy R Williams. 2019. Attending to the visual aspects of visual storytelling: using art and design concepts to interpret and compose narratives with images. Journal of Visual Literacy 38, 1-2 (2019), 66–82
2019
-
[101]
Lai-Kuan Wong and Kok-Lim Low. 2009. Saliency-enhanced image aesthetics class prediction. In 2009 16th IEEE International Conference on Image Processing (ICIP). IEEE, 997–1000
2009
-
[102]
Rongliang Wu and Shijian Lu. 2020. Leed: Label-free expression editing via disentanglement. In European Conference on Computer Vision. Springer, 781–798
2020
-
[103]
Rongliang Wu, Gongjie Zhang, Shijian Lu, and Tao Chen. 2020. Cascade ef-gan: Progressive facial expression editing with local focuses. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5021–5030
2020
-
[104]
Siyang Wu, Tonghan Wang, Chenghao Li, Yang Hu, and Chongjie Zhang
-
[105]
Xiaoran Wu. 2022. Interpretable Aesthetic Analysis Model for Intelligent Pho- tography Guidance Systems. In 27th International Conference on Intelligent User Interfaces. 661–671
2022
-
[106]
Xiaoran Wu and Jia Jia. 2021. Tumera: Tutor of Photography Beginners. arXiv preprint arXiv:2109.11365 (2021)
2021 arXiv
-
[107]
Yaowen Wu, Christian Bauckhage, and Christian Thurau. 2010. The good, the bad, and the ugly: Predicting aesthetic image labels. In 2010 20th International Conference on Pattern Recognition . IEEE, 1586–1589
2010
-
[108]
Mei-Chen Yeh and Yu-Chen Cheng. 2012. Relative features for photo quality assessment. In 2012 19th IEEE International Conference on Image Processing. IEEE, 2861–2864
2012
-
[109]
arXiv preprint arXiv:2110.08169 (2021)
Containerized distributed value-based multi-agent reinforcement learning. arXiv preprint arXiv:2110.08169 (2021)
2021 arXiv
-
[110]
Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang
-
[111]
Yingchen Yu, Fangneng Zhan, Shijian Lu, Jianxiong Pan, Feiying Ma, Xuansong Xie, and Chunyan Miao. 2021. WaveFill: A Wavelet-based Generation Network for Image Inpainting. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 14114–14123
2021
-
[112]
Yingchen Yu, Fangneng Zhan, Rongliang Wu, Jianxiong Pan, Kaiwen Cui, Shi- jian Lu, Feiying Ma, Xuansong Xie, and Chunyan Miao. 2021. Diverse image inpainting with bidirectional and autoregressive transformers. In Proceedings of the 29th ACM International Conference on Multimed...
2021
-
[113]
Yanhong Zeng, Jianlong Fu, Hongyang Chao, and Baining Guo. 2019. Learning pyramid-context encoder network for high-quality image inpainting. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1486–1494
2019
-
[114]
Wenyuan Yin, Tao Mei, and Chang Wen Chen. 2012. Assessing photo quality with geo-context and crowdsourced photos. In 2012 Visual Communications and Image Processing. IEEE, 1–6
2012
-
[115]
Fangneng Zhan, Yingchen Yu, Kaiwen Cui, Gongjie Zhang, Shijian Lu, Jianxiong Pan, Changgong Zhang, Feiying Ma, Xuansong Xie, and Chunyan Miao. 2021. Unbalanced feature transport for exemplar-based image translation. In Proceed- ings of the IEEE/CVF Conference on Computer Visio...
2021
-
[116]
In Proceedings of the IEEE/CVF International Conference on Computer Vision
Free-form image inpainting with gated convolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 4471–4480
-
[117]
Fangneng Zhan, Hongyuan Zhu, and Shijian Lu. 2019. Spatial fusion gan for image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3653–3662
2019
-
[118]
Edwin Zhang, Sadie Zhao, Tonghan Wang, Safwan Hossain, Henry Gasztowtt, Stephan Zheng, David C Parkes, Milind Tambe, and Yiling Chen. 2024. Position: social environment design should be further developed for AI-based policy- making. In International Conference on Machine Learn...
2024
-
[119]
Luming Zhang. 2016. Describing human aesthetic perception by deeply-learned attributes from flickr. arXiv preprint arXiv 1605 (2016)
2016
-
[120]
Yu Zeng, Zhe Lin, Jimei Yang, Jianming Zhang, Eli Shechtman, and Huchuan Lu
-
[121]
Yunfan Zhao, Tonghan Wang, Dheeraj Mysore Nagaraj, Aparna Taneja, and Milind Tambe. 2025. The bandit whisperer: Communication learning for restless bandits. In AAAI Conference on Artificial Intelligence (AAAI) , Vol. 39. 23404– 23413
2025
-
[122]
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba
-
[123]
Fangneng Zhan, Yingchen Yu, Rongliang Wu, Kaiwen Cui, Aoran Xiao, Shi- jian Lu, and Ling Shao. 2021. Bi-level feature alignment for versatile image translation and manipulation. arXiv preprint arXiv:2107.03021 (2021)
2021 arXiv
-
[127]
Luming Zhang, Yue Gao, Roger Zimmermann, Qi Tian, and Xuelong Li. 2014. Fusion of multichannel local and global structural cues for photo aesthetics evaluation. IEEE Transactions on Image Processing 23, 3 (2014), 1419–1429
2014
-
[2000]
In Proceedings of the 27th annual conference on Computer graphics and interactive techniques
Image inpainting. In Proceedings of the 27th annual conference on Computer graphics and interactive techniques . 417–424
-
[2017]
Places: A 10 million image database for scene recognition.IEEE transactions on pattern analysis and machine intelligence 40, 6 (2017), 1452–1464. 14
2017
-
[2019]
arXiv preprint arXiv:1901.00212 (2019)
Edgeconnect: Generative image inpainting with adversarial edge learning. arXiv preprint arXiv:1901.00212 (2019)
2019 arXiv
-
[2020]
In European conference on computer vision
High-resolution image inpainting with iterative confidence feedback and guided upsampling. In European conference on computer vision . Springer, 1–17
-
[2021]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Pd-gan: Probabilistic diverse gan for image inpainting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 9371–9381
-
[2022]
Advances in Neural Information Processing Systems (NeurIPS) 35 (2022), 25655–25666
Non-linear coordination graphs. Advances in Neural Information Processing Systems (NeurIPS) 35 (2022), 25655–25666
2022
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.