REVIEW 3 major objections 6 minor 102 references
HAIFAI: Human-AI Interaction for Mental Face Reconstruction
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims that a face held only in memory can be reconstructed by having users rank small sets of auxiliary faces, with the system fusing those rankings into a StyleGAN2 latent vector and beating prior interactive methods on…
desk verdict Genuine new synthetic-user-model training for interactive face reconstruction, but the abstract overstates usability gains and one evaluation metric is circular. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a pair of stacked transformer encoders: a Siamese transformer reads each iteration's six ranked auxiliary latents plus a class token, and a second transformer combines the per-iteration feature vectors into a single predicted latent $w_{\mathrm{rec}}$. The other essential component is the computational user model (Algorithm 1), which orders auxiliary faces by cosine similarity to the target in a fine-tuned ArcFace embedding space and adds uniform noise $U(-\sigma,\sigma)$ with $\sigma=0.22$ so simulated rankings have human-like variability; the ablation shows this noise raises test embedding similarity from 0.397 to 0.421 and prevents overfitting. Stage two uses UP-FacE's 24 landmark-derived sliders for manual refinement.
What would settle it
Recruit new users to rank the same target and auxiliary sets used in training, and compare each user's rank orders with the model's noisy rankings: if per-user Kendall tau is near zero, or if a model trained on deterministic rankings matches the noisy model on the lineup task, the claim that user-model fidelity drives reconstruction quality is contradicted. A second check is to run lineups that do not include the true target, a condition the paper notes its own study does not use.
Extended reading notes
Core claim
The central claim is that a user's mental face image can be recovered from rank orders over small sets of auxiliary faces, because those ordinal signals carry enough information to locate the target in a generative latent space. The reconstruction network is trained end-to-end to map ranked auxiliary latents to a reconstructed latent $w_{\mathrm{rec}}$, using a loss that combines squared error in latent space with cosine similarity in a fine-tuned ArcFace embedding space, and the computational user model supplies synthetic ranking data by adding uniform noise to embedding similarities. Against CG-GAN and MFRS, HAIFAI's first stage is faster (8.3 vs. 10.2 vs. 17.8 minutes), rates higher on usability (SUS 87 vs. 85 vs. 59), and scores better on visual similarity and embedding similarity; the full system achieves a 60.6% identification rate in a forensic-style lineup, up from 56.1% for CG-GAN, with the target ranked in the top three in 98.5% of cases.
Load-bearing premise
The whole training scheme rests on the assumption that ranking faces by noisy cosine similarity in a fine-tuned face-embedding space reproduces how real people rank faces held in memory; the paper's support for this is an average Kendall tau of 0.284 between model and human rankings, close to the 0.267 human-human agreement, which does not guarantee that individual users or forgotten memories behave like the model.
Editorial extensions
If this is right
- If the central claim holds, forensic witnesses can produce identification-ready composites by ranking six-face sets, with a 60.6% chance that the true target is placed first in a four-person lineup and a 98.5% chance it is in the top three.
- Ranking-based interaction shifts high-dimensional face search from user to system, which is why first-stage reconstruction time drops to 8.3 minutes compared with 17.8 for CG-GAN and 10.2 for MFRS.
- Training on a synthetic user model removes the need for large-scale human ranking data, making the approach transferable to other generative domains at low collection cost.
- The two-stage design implies that the AI should produce a global approximation first and let humans make only targeted local refinements, reducing the 'mental shift' that heavy manual editing causes.
Reading between the lines
- A natural extension the paper leaves implicit is a personalized user model: calibrating the noise level or embedding per user could reduce the required ranking iterations below the observed 10-19 while extracting a stronger signal from each ranking.
- The same ranking-plus-fusion pipeline could apply to other mental-image domains, such as objects, scenes, or voices, whenever a pretrained generator and a similarity embedding exist, since neither component is face-specific except for the fine-tuned embedding.
- Because the lineup study always includes the true target, the 60.6% figure is an upper bound for realistic open-set identification; an open-set lineup with foils replacing the target would measure performance when memory is imperfect and the target may not be present.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents HAIFAI, a two-stage interactive system for reconstructing a face from a user's mental image. In the first stage, users iteratively rank sets of six StyleGAN2-generated auxiliary faces; a transformer-based reconstruction network consumes the ranked latent vectors and predicts a latent/embedding for the target face, trained with a latent-space MSE plus an ArcFace cosine-similarity term. To avoid collecting large amounts of human ranking data, the authors introduce a computational user model that ranks auxiliary faces by cosine similarity with added uniform noise, and they fine-tune this model on a 275-participant AMT dataset of human rankings. The second stage uses UP-FacE, a slider-based face-editing tool, for manual refinement. The system is evaluated in a 12-participant user study against CG-GAN and MFRS, and in an 18-participant lineup study reporting a 60.6% identification rate. The abstract claims superiority in reconstruction quality, usability, perceived workload, and reconstruction speed.
Significance. The core idea of replacing expensive human ranking data with a computational user model is timely and potentially impactful, and the paper includes real user studies, attention checks, a lineup protocol, and ablations that isolate the contributions of fine-tuned embeddings, variable iteration counts, and the noisy user model. If the reconstruction-quality and identification-rate results survive the concerns below, HAIFAI would be a meaningful advance over prior interactive composite-generation methods. However, two issues currently limit the strength of the paper: the headline usability/workload/speed claim is contradicted by Table 1 for the full two-stage system, and the embedding-similarity metric is the same objective used to train the model, making that column an optimized-by-construction measure rather than an independent quality metric.
major comments (3)
- [Abstract; Section 5.2; Table 1] The abstract and conclusion state that HAIFAI outperforms the previous state of the art on usability, perceived workload, and reconstruction speed, but Table 1 shows the opposite for the complete system: full HAIFAI scores SUS 77 versus MFRS 85, NASA-TLX 31 versus 27, and time 11.3 versus 10.2 minutes. Only the ablated first stage (HAIFAI w/o UP-FacE) beats MFRS on these three metrics (87, 25, 8.3). Section 5.2 itself states that adding UP-FacE causes a significant performance degradation in these metrics. Since HAIFAI is presented as a two-stage system and contribution (3) claims improvements in usability, speed, and quality, the headline claim must be revised: usability, workload, and speed gains should be attributed to the ranking-only stage, while reconstruction quality and identification rate are the claims for the full system.
- [Table 1; Eq. (1)] The 'Embedding Sim.' column is not an independent evaluation metric for HAIFAI. Equation (1) trains the reconstruction network to maximize the cosine similarity between ArcFace embeddings of reconstructed and target faces, and the same embedding model is used to compute the reported evaluation values. It is therefore expected that HAIFAI scores higher than CG-GAN and MFRS, which were not optimized for this objective. Please either evaluate with a held-out face embedding model not used anywhere in training, or relabel this column as an optimization-objective match and base reconstruction-quality claims on the human ratings and lineup results.
- [Section 6.3; Algorithm 1] The validation of the computational user model is currently at the level of aggregate distributions: the model-human Kendall tau is 0.284, comparable to the human-human value of 0.267, and the noise level sigma=0.22 is chosen by minimizing the Wasserstein distance between the human-human and model-model ranking matrices. Because all training data for the reconstruction network are generated by this model, the claim that it faithfully simulates human ranking behavior should be supported by a stronger check, for example per-user ranking accuracy, agreement at each rank position, or consistency across the 20 iterations. Without such validation, the synthetic training source remains a generalizability risk even though the final human studies are encouraging.
minor comments (6)
- [Title on page 1] The title contains a stray space in 'Mental Fa ce Reconstruction'; please correct the typo.
- [Section 3.2, Eq. (3)] Equation (3) writes squared differences as (f_a - f_p)^2, but the operands should be the embedding vectors E(f_a) and E(f_p), not the images; please make the notation consistent.
- [Section 5.2, Figure 6] The text and caption refer to 'the last two columns' for HAIFAI reconstructions, but the figure layout is described in rows; please correct the direction.
- [Section 3.3; Reference [77]] The system is called UP-FacE in the text, but the bibliography entry for [77] is titled 'SeFFeC'; please reconcile the naming.
- [Section 6.3] The Kendall tau coefficients are reported without confidence intervals or the number of ranking pairs used; please provide these so the aggregate comparison to human-human agreement can be assessed.
- [Section 5.3, Eq. (4)] The definition IR = #Rank 1 / #Votes x 100 should state whether #Votes counts raters, trials, or lineup judgments; as written it is ambiguous.
Circularity Check
One headline metric is self-definitional: the reported 'embedding similarity' is exactly the cosine-similarity term of the training loss in Eq. 1, so HAIFAI is optimized to maximize it; other quality and usability metrics are independent.
-
self definitional
[Eq. (1), Sec. 3.1; 'Embedding similarity' metric, Sec. 5.2 and Table 1]
"L =∥wrec− wM∥2−λe E(G(wrec))· E(G(wM)) / (|E(G(wrec))||E(G(wM))|) (1) ... We used the loss function defined in Equation 1 for training with λe = 1. ... Embedding similarity: We calculated the cosine similarity between the embeddings of the target and reconstructed images to assess the reconstruction quality."
Eq. 1's second term is precisely the metric later reported as 'Embedding similarity' in Table 1 and Fig. 8: the cosine similarity between the target and reconstructed face embeddings. With λe = 1, training HAIFAI's reconstruction network directly maximizes this number. Reporting that same number as evidence that HAIFAI 'outperforms the previous state of the art regarding reconstruction quality' is therefore partly self-definitional: the evaluation metric is the training objective, not an independent measurement. The comparison is also biased in HAIFAI's favor because the baselines (MFRS, CG-GAN) were not trained with this loss. The other headline metrics (visual rating, SUS, NASA-TLX, time, identification rate) are not the training objective and are independent.
full rationale
The only identified circularity is the embedding-similarity evaluation metric being the training objective of Eq. 1. This makes one of the reported quality metrics favorable to HAIFAI by construction; however, the headline claims also rest on independent user-study measures (visual rating, mental rating, SUS, NASA-TLX, time) and on a separate human lineup identification-rate study (IR 60.6%), which are not optimized by the loss. The computational user model is a training-data simulator, and its agreement with human rankings (Kendall tau 0.284 vs. 0.267) is validated against human ranking data; the noise level sigma is tuned to a distributional distance, not to the reported model-human agreement, so it does not constitute a fitted-input-called-prediction step. Citations of the authors' prior work [76] and [77] are used as a comparison baseline and as an off-the-shelf editing tool, respectively; neither is invoked to forbid alternatives or to supply the paper's core derivation. Lineup construction using ArcFace nearest neighbors makes identification harder, not easier, so it is not circular. Overall, the paper has partial circularity in one headline metric, but the central system claim retains substantial independent empirical content.
Assumptions & free parameters
free parameters (4)
- sigma =
0.22
- alpha =
0.1
- lambda_e =
1
- triplet margin m =
0.1
assumptions (3)
- domain assumption StyleGAN2's latent space W supports meaningful face generation and editing, and generated auxiliary faces cover the target population.
- domain assumption Fine-tuned ArcFace embeddings approximate human perceptual face similarity after triplet fine-tuning.
- ad hoc to paper Human ranking behaviour can be simulated by adding uniform noise to cosine similarities in embedding space.
Cite this review
Pith. "Pith review of HAIFAI: Human-AI Interaction for Mental Face Reconstruction." pith.science (2026). https://pith.science/paper/WLMV5WXR
@misc{pith2026241206323,
author = {Pith},
title = {Pith review of: HAIFAI: Human-AI Interaction for Mental Face Reconstruction},
year = {2026},
howpublished = {\url{https://pith.science/paper/WLMV5WXR}},
note = {Machine review of arXiv:2412.06323}
}
read the original abstract
We present HAIFAI - a novel two-stage system where humans and AI interact to tackle the challenging task of reconstructing a visual representation of a face that exists only in a person's mind. In the first stage, users iteratively rank images our reconstruction system presents based on their resemblance to a mental image. These rankings, in turn, allow the system to extract relevant image features, fuse them into a unified feature vector, and use a generative model to produce an initial reconstruction of the mental image. The second stage leverages an existing face editing method, allowing users to manually refine and further improve this reconstruction using an easy-to-use slider interface for face shape manipulation. To avoid the need for tedious human data collection for training the reconstruction system, we introduce a computational user model of human ranking behaviour. For this, we collected a small face ranking dataset through an online crowd-sourcing study containing data from 275 participants. We evaluate HAIFAI and an ablated version in a 12-participant user study and demonstrate that our approach outperforms the previous state of the art regarding reconstruction quality, usability, perceived workload, and reconstruction speed. We further validate the reconstructions in a subsequent face ranking study with 18 participants and show that HAIFAI achieves a new state-of-the-art identification rate of 60.6%. These findings represent a significant advancement towards developing new interactive intelligent systems capable of reliably and effortlessly reconstructing a user's mental image.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Rameen Abdal, Peihao Zhu, Niloy J Mitra, and Peter Wonka. 2021. Styleflow: Attribute-conditioned exploration of stylegan-generated images using conditional continuous normalizing flows. ACM Transactions on Graphics (ToG) 40, 3 (2021), 1–21
2021
-
[2]
Rudolf Arnheim. 1969. Visual thinking. Univ of California Press
1969
-
[3]
Roman Beliy, Guy Gaziv, Assaf Hoogi, Francesca Strappini, Tal Golan, and Michal Irani. 2019. From voxels to pixels and back: Self-supervision in natural-image reconstruction from fMRI. In Advances in Neural Information Processing Systems . 6517–6527
2019
-
[4]
Philip Bontrager, Wending Lin, Julian Togelius, and Sebastian Risi. 2018. Deep interactive evolution. In International Conference on Computational Intelligence in Music, Sound, Art and Design . 267–282
2018
-
[5]
Andrea Bruera and Massimo Poesio. 2022. Exploring the representations of individual entities in the brain combining EEG and distributional semantics. Frontiers in artificial intelligence (2022), 25
2022
-
[6]
Andreas Bulling and Daniel Roggen. 2011. Recognition of Visual Memory Recall Processes Using Eye Movement Analysis. InProc. ACM International Joint Conference on Pervasive and Ubiquitous Computing (UbiComp) . 455–464. https://doi.org/10.1145/2030112.2030172
arXiv 2011
-
[7]
Shu-Yu Chen, Feng-Lin Liu, Yu-Kun Lai, Paul L Rosin, Chunpeng Li, Hongbo Fu, and Lin Gao. 2021. DeepFaceEditing: deep face generation and editing with disentangled geometry and appearance control. ACM Transactions on Graphics (TOG) 40, 4 (2021), 1–15. HAIFAI: Human-AI Interaction for Mental Face Reconstruction 23
2021
-
[8]
Shu-Yu Chen, Wanchao Su, Lin Gao, Shihong Xia, and Hongbo Fu. 2020. DeepFaceDrawing: Deep generation of face images from sketches. ACM Transactions on Graphics (TOG) 39, 4 (2020), 72–1
2020
Show all 102 references
-
[9]
Chia-Hsing Chiu, Yuki Koyama, Yu-Chi Lai, Takeo Igarashi, and Yonghao Yue. 2020. Human-in-the-loop differential subspace search in high- dimensional latent space. ACM Transactions on Graphics (TOG) 39, 4 (2020), 85–1
2020
-
[10]
Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo. 2018. Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In Proceedings of the IEEE conference on computer vision and pattern recognition . 8789–8797
2018
-
[11]
Donald F Christie and Hadyn D Ellis. 1981. Photofit constructions versus verbal descriptions of faces. Journal of Applied Psychology 66, 3 (1981), 358
1981
-
[12]
Alan S Cowen, Marvin M Chun, and Brice A Kuhl. 2014. Neural portraits of perception: reconstructing face images from evoked brain activity. Neuroimage 94 (2014), 12–22
2014
-
[13]
Thirza Dado, Yağmur Güçlütürk, Luca Ambrogioni, Gabriëlle Ras, Sander Bosch, Marcel van Gerven, and Umut Güçlü. 2022. Hyperrealistic neural decoding for reconstructing faces from fMRI activations via the GAN latent space. Scientific reports 12, 1 (2022), 1–9
2022
-
[14]
Hiroto Date, Keisuke Kawasaki, Isao Hasegawa, and Takayuki Okatani. 2019. Deep learning for natural image reconstruction from electrocorticog- raphy signals. In 2019 IEEE International Conference on Bioinformatics and Biomedicine . 2331–2336
2019
-
[15]
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. 2019. Arcface: Additive angular margin loss for deep face recognition. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4690–4699
2019
-
[16]
Yu Deng, Jiaolong Yang, Dong Chen, Fang Wen, and Xin Tong. 2020. Disentangled and controllable face image generation via 3d imitative-contrastive learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5154–5163
2020
-
[17]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)
2018 arXiv
-
[18]
Hadyn D Ellis, Graham M Davies, and John W Shepherd. 1978. A critical examination of the Photofit system for recalling faces. Ergonomics 21, 4 (1978), 297–307
1978
-
[19]
special
Martha J Farah, Kevin D Wilson, Maxwell Drain, and James N Tanaka. 1998. What is" special" about face perception? Psychological Review 105, 3 (1998), 482
1998
-
[20]
Charlie D Frowd, Peter JB Hancock, and Derek Carson. 2004. EvoFIT: A holistic, evolutionary facial imaging technique for creating composites. ACM Transactions on applied perception 1, 1 (2004), 19–39
2004
-
[21]
Yue Gao, Fangyun Wei, Jianmin Bao, Shuyang Gu, Dong Chen, Fang Wen, and Zhouhui Lian. 2021. High-fidelity and arbitrary face editing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 16115–16124
2021
-
[22]
Stuart J Gibson, Chris J Solomon, Matthew IS Maylin, and Clifford Clark. 2009. New methodology in facial composite construction: From theory to practice. International Journal of Electronic Security and Digital Forensics 2, 2 (2009), 156–168
2009
-
[23]
Shuyang Gu, Jianmin Bao, Hao Yang, Dong Chen, Fang Wen, and Lu Yuan. 2019. Mask-guided portrait editing with conditional gans. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition . 3436–3445
2019
-
[24]
Yağmur Güçlütürk, Umut Güçlü, Katja Seeliger, Sander Bosch, Rob van Lier, and Marcel A van Gerven. 2017. Reconstructing perceived faces from brain activations with deep adversarial neural decoding. Advances in Neural Information Processing Systems 30 (2017)
2017
-
[25]
Jingtao Guo, Zhenzhen Qian, Zuowei Zhou, and Yi Liu. 2019. Mulgan: Facial attribute editing by exemplar. arXiv preprint arXiv:1912.12396 (2019)
2019 arXiv
-
[26]
Yuxuan Han, Jiaolong Yang, and Ying Fu. 2021. Disentangled face attribute editing via instance-aware latent space search. IJCAI (2021)
2021
-
[27]
Erik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, and Sylvain Paris. 2020. Ganspace: Discovering interpretable gan controls. Advances in neural information processing systems 33 (2020), 9841–9850
2020
-
[28]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 770–778
2016
-
[29]
Zhenliang He, Meina Kan, Jichao Zhang, and Shiguang Shan. 2020. Pa-gan: Progressive attention generative adversarial network for facial attribute editing. arXiv preprint arXiv:2007.05892 (2020)
2020 arXiv
-
[30]
Zhenliang He, Wangmeng Zuo, Meina Kan, Shiguang Shan, and Xilin Chen. 2019. Attgan: Facial attribute editing by only changing what you want. IEEE transactions on image processing 28, 11 (2019), 5464–5478
2019
-
[31]
Xianxu Hou, Linlin Shen, Or Patashnik, Daniel Cohen-Or, and Hui Huang. 2022. Feat: Face editing with attention. arXiv preprint arXiv:2202.02713 (2022)
2022 arXiv
-
[32]
Xianxu Hou, Xiaokang Zhang, Hanbang Liang, Linlin Shen, Zhihui Lai, and Jun Wan. 2022. Guidedstyle: Attribute knowledge guided style manipulation for semantic face editing. Neural Networks 145 (2022), 209–220
2022
-
[33]
Ziqi Huang, Kelvin CK Chan, Yuming Jiang, and Ziwei Liu. 2023. Collaborative diffusion for multi-modal face generation and editing. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6080–6090
2023
-
[34]
Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning . 448–456
2015
-
[35]
Marc Jeannerod. 1995. Mental imagery in the motor context. Neuropsychologia 33, 11 (1995), 1419–1432
1995
-
[36]
Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. 2017. Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196 (2017). 24 Florian Strohm, Mihai Bâce, & Andreas Bulling
2017 arXiv
-
[37]
Tero Karras, Samuli Laine, and Timo Aila. 2019. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4401–4410
2019
-
[38]
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8110–8119
2020
-
[39]
Siavash Khodadadeh, Shabnam Ghadar, Saeid Motiian, Wei-An Lin, Ladislau Bölöni, and Ratheesh Kalarot. 2022. Latent to latent: A learned mapper for identity preserving editing of multiple face attributes in stylegan-generated images. In Proceedings of the IEEE/CVF Winter Confer...
2022
-
[40]
Minjeong Kim, Jung-Hwan Kim, Minjung Park, and Jungmin Yoo. 2021. The roles of sensory perceptions and mental imagery in consumer decision-making. Journal of Retailing and Consumer Services 61 (2021), 102517
2021
-
[41]
Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)
2014 arXiv
-
[42]
Christine E Koehn and Ronald P Fisher. 1997. Constructing facial composites with the Mac-a-Mug Pro system. Psychology, Crime and Law 3, 3 (1997), 209–218
1997
-
[43]
Marek Kowalski, Stephan J Garbin, Virginia Estellers, Tadas Baltrušaitis, Matthew Johnson, and Jamie Shotton. 2020. Config: Controllable neural face image generation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI ...
2020
-
[44]
Jeong-gi Kwak, David K Han, and Hanseok Ko. 2020. CAFE-GAN: Arbitrary face attribute editing with complementary attention feature. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIV 16 . Springer, 524–540
2020
-
[45]
Laughery and Richard H
Kenneth R. Laughery and Richard H. Fowler. 1980. Sketch artist and Identi-kit procedures for recalling faces. Journal of Applied Psychology 65, 3 (June 1980), 307–316
1980
-
[46]
Cheng-Han Lee, Ziwei Liu, Lingyun Wu, and Ping Luo. 2020. Maskgan: Towards diverse and interactive facial image manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5549–5558
2020
-
[47]
Yunfeng Lin, Jiangbei Li, and Hanjing Wang. 2019. DCNN-GAN: Reconstructing Realistic Image from fMRI. In 2019 16th International Conference on Machine Vision Applications. 1–6
2019
-
[48]
Huan Ling, Karsten Kreis, Daiqing Li, Seung Wook Kim, Antonio Torralba, and Sanja Fidler. 2021. Editgan: High-precision semantic image editing. Advances in Neural Information Processing Systems 34 (2021), 16331–16345
2021
-
[49]
Yongyi Lu, Yu-Wing Tai, and Chi-Keung Tang. 2018. Attribute-guided face generation using conditional cyclegan. InProceedings of the European conference on computer vision (ECCV) . 282–297
2018
-
[50]
Safa C Medin, Bernhard Egger, Anoop Cherian, Ye Wang, Joshua B Tenenbaum, Xiaoming Liu, and Tim K Marks. 2022. MOST-GAN: 3D morphable StyleGAN for disentangled face image manipulation. In Proceedings of the AAAI conference on artificial intelligence , Vol. 36. 1962–1971
2022
-
[51]
Samuel T Moulton and Stephen M Kosslyn. 2009. Imagining predictions: mental imagery as mental emulation. Philosophical Transactions of the Royal Society B: Biological Sciences 364, 1521 (2009), 1273–1280
2009
-
[52]
Thomas Naselaris, Cheryl A Olman, Dustin E Stansbury, Kamil Ugurbil, and Jack L Gallant. 2015. A voxel-wise encoding model for early visual areas decodes mental images of remembered scenes. Neuroimage 105 (2015), 215–228
2015
-
[53]
Dan Nemrodov, Matthias Niemeier, Ashutosh Patel, and Adrian Nestor. 2018. The neural dynamics of facial identity processing: insights from EEG-based pattern analysis and image reconstruction. Eneuro 5, 1 (2018)
2018
-
[54]
Adrian Nestor, David C Plaut, and Marlene Behrmann. 2016. Feature-based face representations and image reconstruction from behavioral and neural data. Proceedings of the National Academy of Sciences 113, 2 (2016), 416–421
2016
-
[55]
Yongjie Niu, Mingquan Zhou, and Zhan Li. 2023. Disentangling the latent space of GANs for semantic face editing.Plos one 18, 10 (2023), e0293496
2023
-
[56]
Furkan Ozcelik and Rufin VanRullen. 2023. Natural scene reconstruction from fMRI signals using generative latent diffusion. Scientific Reports 13, 1 (2023), 15666
2023
-
[57]
Xingang Pan, Ayush Tewari, Thomas Leimkühler, Lingjie Liu, Abhimitra Meka, and Christian Theobalt. 2023. Drag your gan: Interactive point-based manipulation on the generative image manifold. In ACM SIGGRAPH 2023 Conference Proceedings . 1–11
2023
-
[58]
Omkar M Parkhi, Andrea Vedaldi, and Andrew Zisserman. 2015. Deep face recognition. In Proceedings of the British Machine Vision Conference
2015
-
[59]
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. 2021. Styleclip: Text-driven manipulation of stylegan imagery. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 2085–2094
2021
-
[60]
Joel Pearson, Thomas Naselaris, Emily A Holmes, and Stephen M Kosslyn. 2015. Mental imagery: functional mechanisms and clinical applications. Trends in cognitive sciences 19, 10 (2015), 590–602
2015
-
[61]
Tiziano Portenier, Qiyang Hu, Attila Szabó, Siavash Arjomand Bigdeli, Paolo Favaro, and Matthias Zwicker. 2018. Faceshop: deep sketch-based face image editing. ACM Transactions on Graphics (TOG) 37, 4 (2018), 1–13
2018
-
[62]
Amir Sadovnik, Wassim Gharbi, Thanh Vu, and Andrew Gallagher. 2018. Finding your lookalike: Measuring face similarity rather than face identity. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops . 2345–2353
2018
-
[63]
Antonios Saravanos, Stavros Zervoudakis, Dongnanzi Zheng, Neil Stott, Bohdan Hawryluk, and Donatella Delfino. 2021. The hidden cost of using Amazon Mechanical Turk for research. In HCI International 2021-Late Breaking Papers: Design and User Experience: 23rd HCI International ...
2021
-
[64]
Hosnieh Sattar, Andreas Bulling, and Mario Fritz. 2017. Predicting the category and attributes of visual search targets using deep gaze pooling. In Proceedings of the IEEE International Conference on Computer Vision Workshops . 2740–2748. HAIFAI: Human-AI Interaction for Menta...
2017
-
[65]
Hosnieh Sattar, Mario Fritz, and Andreas Bulling. 2020. Deep gaze pooling: Inferring and visually decoding search intents from human gaze fixations. Neurocomputing 387 (2020), 369–382
2020
-
[66]
Paul Scotti, Atmadeep Banerjee, Jimmie Goode, Stepan Shabalin, Alex Nguyen, Aidan Dempster, Nathalie Verlinde, Elad Yundler, David Weisberg, Kenneth Norman, et al. 2024. Reconstructing the mind’s eye: fMRI-to-image with contrastive learning and diffusion priors. Advances in Ne...
2024
-
[67]
Katja Seeliger, Umut Güçlü, Luca Ambrogioni, Yagmur Güçlütürk, and Marcel AJ van Gerven. 2018. Generative adversarial networks for reconstructing natural images from brain activity. NeuroImage 181 (2018), 775–785
2018
-
[68]
Sophia M Shatek, Tijl Grootswagers, Amanda K Robinson, and Thomas A Carlson. 2019. Decoding images in the mind’s eye: The temporal dynamics of visual imagery. Vision 3, 4 (2019), 53
2019
-
[69]
Guohua Shen, Tomoyasu Horikawa, Kei Majima, and Yukiyasu Kamitani. 2019. Deep image reconstruction from human brain activity. PLoS computational biology 15, 1 (2019), e1006633
2019
-
[70]
Yujun Shen, Ceyuan Yang, Xiaoou Tang, and Bolei Zhou. 2020. Interfacegan: Interpreting the disentangled face representation learned by gans. IEEE transactions on pattern analysis and machine intelligence 44, 4 (2020), 2004–2018
2020
-
[71]
Yujun Shen and Bolei Zhou. 2021. Closed-form factorization of latent semantics in gans. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 1532–1540
2021
-
[72]
Andrew Sims and Marcus Missal. 2019. Perceptual Decision-Making and Beyond: Intention as Mental Imagery. In Free Will, Causality, and Neuroscience. Brill, 13–34
2019
-
[73]
Pawan Sinha, Benjamin Balas, Yuri Ostrovsky, and Richard Russell. 2006. Face recognition by humans: Nineteen results all computer vision researchers should know about. Proc. IEEE 94, 11 (2006), 1948–1962
2006
-
[74]
Linsen Song, Jie Cao, Lingxiao Song, Yibo Hu, and Ran He. 2019. Geometry-aware face completion and editing. In Proceedings of the AAAI conference on artificial intelligence , Vol. 33. 2506–2513
2019
-
[75]
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. Dropout: a simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research 15, 1 (2014), 1929–1958
2014
-
[76]
Florian Strohm, Mihai Bâce, and Andreas Bulling. 2023. Usable and fast interactive mental face reconstruction. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–15
2023
-
[77]
Florian Strohm, Mihai Bâce, Markus Kaltenecker, and Andreas Bulling. 2024. SeFFeC: Semantic Facial Feature Control for Fine-grained Face Editing. arXiv preprint arXiv:2403.13972 (2024)
2024 arXiv
-
[78]
Florian Strohm, Ekta Sood, Sven Mayer, Philipp Müller, Mihai Bâce, and Andreas Bulling. 2021. Neural Photofit: Gaze-based Mental Image Reconstruction. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 245–254
2021
-
[79]
Florian Strohm, Ekta Sood, Dominike Thomas, Mihai Bâce, and Andreas Bulling. 2022. Facial Composite Generation with Iterative Human Feedback. In Proceedings of Machine Learning Research
2022
-
[80]
Jianxin Sun, Qiyao Deng, Qi Li, Muyi Sun, Min Ren, and Zhenan Sun. 2022. Anyface: Free-style text-to-face synthesis and manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18687–18696
2022
-
[81]
Jingxiang Sun, Xuan Wang, Yichun Shi, Lizhen Wang, Jue Wang, and Yebin Liu. 2022. Ide-3d: Interactive disentangled editing for high-resolution 3d-aware portrait synthesis. ACM Transactions on Graphics (ToG) 41, 6 (2022), 1–10
2022
-
[82]
Jingxiang Sun, Xuan Wang, Yong Zhang, Xiaoyu Li, Qi Zhang, Yebin Liu, and Jue Wang. 2022. Fenerf: Face editing in neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 7672–7682
2022
-
[83]
Qiushi Sun, Jingtao Guo, and Yi Liu. 2022. PattGAN: Pluralistic Facial Attribute Editing. IEEE Access 10 (2022), 68534–68544
2022
-
[84]
Yu Takagi and Shinji Nishimoto. 2023. High-resolution image reconstruction with latent diffusion models from human brain activity. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 14453–14463
2023
-
[85]
Ayush Tewari, Mohamed Elgharib, Florian Bernard, Hans-Peter Seidel, Patrick Pérez, Michael Zollhöfer, and Christian Theobalt. 2020. Pie: Portrait image embedding for semantic control. ACM Transactions on Graphics (TOG) 39, 6 (2020), 1–14
2020
-
[86]
Ayush Tewari, Mohamed Elgharib, Gaurav Bharaj, Florian Bernard, Hans-Peter Seidel, Patrick Pérez, Michael Zollhofer, and Christian Theobalt
-
[87]
Rufin VanRullen and Leila Reddy. 2019. Reconstructing faces from fMRI patterns using deep generative neural networks. Communications biology 2, 1 (2019), 1–10
2019
-
[88]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)
2017
-
[89]
Feng Wang, Jian Cheng, Weiyang Liu, and Haijun Liu. 2018. Additive margin softmax for face verification. IEEE Signal Processing Letters 25, 7 (2018), 926–930
2018
-
[90]
Hao Wang, Yitong Wang, Zheng Zhou, Xing Ji, Dihong Gong, Jingchao Zhou, Zhifeng Li, and Wei Liu. 2018. Cosface: Large margin cosine loss for deep face recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 5265–5274
2018
-
[91]
Yi Wei, Zhe Gan, Wenbo Li, Siwei Lyu, Ming-Ching Chang, Lei Zhang, Jianfeng Gao, and Pengchuan Zhang. 2020. Maggan: High-resolution face attribute editing with mask-guided generative adversarial network. In Proceedings of the Asian Conference on Computer Vision . 26 Florian St...
2020
-
[92]
Zongze Wu, Dani Lischinski, and Eli Shechtman. 2021. Stylespace analysis: Disentangled controls for stylegan image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12863–12872
2021
-
[93]
Weihao Xia, Yujiu Yang, Jing-Hao Xue, and Baoyuan Wu. 2021. Tedigan: Text-guided diverse face image generation and manipulation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2256–2265
2021
-
[94]
Taihong Xiao, Jiapeng Hong, and Jinwen Ma. 2018. Elegant: Exchanging latent encodings with gan for transferring multiple face attributes. In Proceedings of the European conference on computer vision (ECCV) . 168–184
2018
-
[95]
Caie Xu, Ying Tang, Masahiro Toyoura, Jiayi Xu, and Xiaoyang Mao. 2019. Generating Users’ Desired Face Image Using the Conditional Generative Adversarial Network and Relevance Feedback. IEEE Access 7 (2019), 181458–181468
2019
-
[96]
Guoxing Yang, Nanyi Fei, Mingyu Ding, Guangzhen Liu, Zhiwu Lu, and Tao Xiang. 2021. L2m-gan: Learning to manipulate latent space semantics for facial attribute editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2951–2960
2021
-
[97]
Huiting Yang, Liangyu Chai, Qiang Wen, Shuang Zhao, Zixun Sun, and Shengfeng He. 2021. Discovering interpretable latent space directions of gans beyond binary attributes. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 12177–12185
2021
-
[98]
Xu Yao, Alasdair Newson, Yann Gousseau, and Pierre Hellier. 2021. A latent transformer for disentangled face editing in images and videos. In Proceedings of the IEEE/CVF international conference on computer vision . 13789–13798
2021
-
[99]
Nicola Zaltron, Luisa Zurlo, and Sebastian Risi. 2020. Cg-gan: An interactive evolutionary gan-based approach for facial composite generation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 2544–2551
2020
-
[100]
Gang Zhang, Meina Kan, Shiguang Shan, and Xilin Chen. 2018. Generative adversarial network with spatial attention for face attribute editing. In Proceedings of the European conference on computer vision (ECCV) . 417–432
2018
-
[101]
Xiao Zheng, Wanzhong Chen, Mingyang Li, Tao Zhang, Yang You, and Yun Jiang. 2020. Decoding human brain activity with deep learning. Biomedical Signal Processing and Control 56 (2020), 101730
2020
-
[2020]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Stylerig: Rigging stylegan for 3d control over portrait images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6142–6151
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.