{"id":"38e301b8-407c-4c4f-abf7-f06f44c87ef0","arxiv_id":"2412.06323","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"HAIFAI reconstructs a user's mental face image from iterative ranking feedback and optional manual refinement, achieving a 60.6% lineup identification rate.","lead":"HAIFAI is a two-stage system that reconstructs the face a person has in mind by having them rank randomly shown faces and then fine-tune the result with sliders. It reports a 60.6% identification rate in lineups, a new best for deep-learning-based face composite systems, which matters for forensic witness applications.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract overstates HAIFAI's gains: full HAIFAI is worse than MFRS on SUS, NASA-TLX, and time; only the ranking-only stage delivers the advertised usability, workload, and speed advantages.","rationale":"The reader's stated weakest_assumption is the fidelity of the computational user model (Section 6.3, Kendall tau 0.284), but their rationale also flags the abstract overstatement relative to Table 1. My stress test identifies the abstract overstatement as the more load-bearing concern for the central claim, because the claim as written is directly contradicted by the paper's own data: full HAIFAI is worse than MFRS on SUS, NASA-TLX, and time. The user-model concern is real but weaker: the model is validated against human-human agreement (tau 0.267 vs 0.284), and the system is ultimately tested on real users, so a mismatch in synthetic ranking distributions would affect explainability and potential scalability more than the reported empirical results. The abstract issue, by contrast, makes a false comparative statement about the deployed system, and the paper's own Section 5.2 acknowledges the degradation. The fix is a rewrite of the abstract and possibly a reordering of the contribution claims; this does not invalidate the reconstruction-quality or identification-rate findings. Therefore I keep the reader's CONDITIONAL verdict unchanged. I mark agreement as partial because the reader's rationale already contained this observation even though it was not listed as the weakest assumption.","tokens_in":22876,"tokens_out":5426,"duration_ms":44300,"concrete_test":"Take the per-participant SUS, NASA-TLX, and completion-time data from the 12-participant user study. For each metric, compute the paired difference between full HAIFAI and MFRS and test it against zero with a paired t-test or Wilcoxon signed-rank test. If the differences are not significantly in HAIFAI's favor (or are significantly worse), the abstract's 'outperforms ... usability, perceived workload, and reconstruction speed' statement is false and must be revised. Run the same test for HAIFAI w/o UP-FacE vs MFRS to confirm at which stage the advantage actually holds. This is a re-analysis of existing data and requires no new user study.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that HAIFAI 'outperforms the previous state of the art regarding reconstruction quality, usability, perceived workload, and reconstruction speed.' Table 1 does not support this for the full HAIFAI system. On SUS, HAIFAI scores 77 vs MFRS's 85; on NASA-TLX, 31 vs 27; on time, 11.3 vs 10.2 minutes. Only the ablated first stage (HAIFAI w/o UP-FacE) beats both baselines on those usability metrics (87, 25, 8.3). The paper itself states in Section 5.2 that adding UP-FacE causes 'a significant performance degradation in these three usability metrics.' The quality and identification-rate claims are better supported: visual rating 4.4 and embedding similarity 0.43 are the best in Table 1, and IR 60.6% beats CG-GAN's 56.1% in the lineup study (Section 5.3). But the usability, workload, and speed half of the headline claim is false for the complete system as presented. This matters because the contribution is framed as a two-stage system; a reader who adopts the full pipeline for its advertised usability would be misled. The fix is straightforward: restrict the superiority claim to reconstruction quality and identification rate, or separately attribute the usability gains to the ranking-only stage.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents HAIFAI, a two-stage interactive system for reconstructing a face from a user's mental image. In the first stage, users iteratively rank sets of six StyleGAN2-generated auxiliary faces; a transformer-based reconstruction network consumes the ranked latent vectors and predicts a latent/embedding for the target face, trained with a latent-space MSE plus an ArcFace cosine-similarity term. To avoid collecting large amounts of human ranking data, the authors introduce a computational user model that ranks auxiliary faces by cosine similarity with added uniform noise, and they fine-tune this model on a 275-participant AMT dataset of human rankings. The second stage uses UP-FacE, a slider-based face-editing tool, for manual refinement. The system is evaluated in a 12-participant user study against CG-GAN and MFRS, and in an 18-participant lineup study reporting a 60.6% identification rate. The abstract claims superiority in reconstruction quality, usability, perceived workload, and reconstruction speed.","tokens_in":23160,"tokens_out":6585,"duration_ms":66416,"significance":"The core idea of replacing expensive human ranking data with a computational user model is timely and potentially impactful, and the paper includes real user studies, attention checks, a lineup protocol, and ablations that isolate the contributions of fine-tuned embeddings, variable iteration counts, and the noisy user model. If the reconstruction-quality and identification-rate results survive the concerns below, HAIFAI would be a meaningful advance over prior interactive composite-generation methods. However, two issues currently limit the strength of the paper: the headline usability/workload/speed claim is contradicted by Table 1 for the full two-stage system, and the embedding-similarity metric is the same objective used to train the model, making that column an optimized-by-construction measure rather than an independent quality metric.","major_comments":[{"comment":"The abstract and conclusion state that HAIFAI outperforms the previous state of the art on usability, perceived workload, and reconstruction speed, but Table 1 shows the opposite for the complete system: full HAIFAI scores SUS 77 versus MFRS 85, NASA-TLX 31 versus 27, and time 11.3 versus 10.2 minutes. Only the ablated first stage (HAIFAI w/o UP-FacE) beats MFRS on these three metrics (87, 25, 8.3). Section 5.2 itself states that adding UP-FacE causes a significant performance degradation in these metrics. Since HAIFAI is presented as a two-stage system and contribution (3) claims improvements in usability, speed, and quality, the headline claim must be revised: usability, workload, and speed gains should be attributed to the ranking-only stage, while reconstruction quality and identification rate are the claims for the full system.","section":"Abstract; Section 5.2; Table 1"},{"comment":"The 'Embedding Sim.' column is not an independent evaluation metric for HAIFAI. Equation (1) trains the reconstruction network to maximize the cosine similarity between ArcFace embeddings of reconstructed and target faces, and the same embedding model is used to compute the reported evaluation values. It is therefore expected that HAIFAI scores higher than CG-GAN and MFRS, which were not optimized for this objective. Please either evaluate with a held-out face embedding model not used anywhere in training, or relabel this column as an optimization-objective match and base reconstruction-quality claims on the human ratings and lineup results.","section":"Table 1; Eq. (1)"},{"comment":"The validation of the computational user model is currently at the level of aggregate distributions: the model-human Kendall tau is 0.284, comparable to the human-human value of 0.267, and the noise level sigma=0.22 is chosen by minimizing the Wasserstein distance between the human-human and model-model ranking matrices. Because all training data for the reconstruction network are generated by this model, the claim that it faithfully simulates human ranking behavior should be supported by a stronger check, for example per-user ranking accuracy, agreement at each rank position, or consistency across the 20 iterations. Without such validation, the synthetic training source remains a generalizability risk even though the final human studies are encouraging.","section":"Section 6.3; Algorithm 1"}],"minor_comments":[{"comment":"The title contains a stray space in 'Mental Fa ce Reconstruction'; please correct the typo.","section":"Title on page 1"},{"comment":"Equation (3) writes squared differences as (f_a - f_p)^2, but the operands should be the embedding vectors E(f_a) and E(f_p), not the images; please make the notation consistent.","section":"Section 3.2, Eq. (3)"},{"comment":"The text and caption refer to 'the last two columns' for HAIFAI reconstructions, but the figure layout is described in rows; please correct the direction.","section":"Section 5.2, Figure 6"},{"comment":"The system is called UP-FacE in the text, but the bibliography entry for [77] is titled 'SeFFeC'; please reconcile the naming.","section":"Section 3.3; Reference [77]"},{"comment":"The Kendall tau coefficients are reported without confidence intervals or the number of ranking pairs used; please provide these so the aggregate comparison to human-human agreement can be assessed.","section":"Section 6.3"},{"comment":"The definition IR = #Rank 1 / #Votes x 100 should state whether #Votes counts raters, trials, or lineup judgments; as written it is ambiguous.","section":"Section 5.3, Eq. (4)"}],"recommendation":"major_revision","confidential_remarks":"The paper's own Section 5.2 concedes the usability/workload/speed degradation when the second stage is added, so the abstract overclaim should be straightforward to fix. The embedding-similarity metric is more serious because it is the training objective itself, but it is fixable by using an evaluation embedding model that was not part of training. The work fits the scope of the journal and the underlying system is interesting; I see no indication of data fabrication, only framing and metric-validity issues."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about HAIFAI. The genuine contribution is the computational user model: a fine-tuned ArcFace embedding with calibrated noise that generates synthetic ranking data to train the reconstruction network, plus the variable-length transformer that takes any number of ranked iterations. That is new relative to MFRS and CG-GAN, both of which need real user feedback at run time. Second, the abstract overstates the usability and speed gains. In Table 1, the full HAIFAI system is worse than MFRS on SUS (77 vs 85), NASA-TLX (31 vs 27), and time (11.3 vs 10.2 minutes). Only the ranking-only ablation beats both baselines on those metrics. The authors are transparent about this in Section 5.2, which makes it a reporting problem rather than a hidden flaw, but the abstract as written is not accurate.\n\nWhat the paper does well: the user model is validated against human-human agreement (Kendall tau 0.284 vs 0.267), and the ablation study shows that adding noise and variable iterations improves generalization. The lineup study is a sensible practical evaluation, and the 60.6% identification rate beats CG-GAN's 56.1%, even though lineups always include the target, which the authors acknowledge.\n\nSoft spots. The embedding similarity metric in Table 1 is the same cosine similarity used as the training objective in Eq. 1, so that column is circular. The visual rating (4.4 vs 4.1) gives independent evidence for quality, but the circular metric should be replaced or supplemented. The user model is only validated at the aggregate level; a Kendall tau of 0.284 does not tell you whether it captures individual ranking patterns or memory drift. That is a limitation, not a fatal flaw.\n\nThis paper is for researchers in interactive face reconstruction and forensic composite systems. It is a solid incremental advance, not a breakthrough. I would send it to peer review, because the core idea is novel and the limitations are fixable with a rewritten abstract and an independent quality metric. I would cite it for the synthetic user model. Reading group: maybe.","headline":"Genuine new synthetic-user-model training for interactive face reconstruction, but the abstract overstates usability gains and one evaluation metric is circular.","tokens_in":23644,"tokens_out":3697,"would_cite":true,"duration_ms":33527,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a face held only in memory can be reconstructed by having users rank small sets of auxiliary faces, with the system fusing those rankings into a StyleGAN2 latent vector and beating prior interactive methods on…","keywords":["mental image reconstruction","faces","user modelling","deep learning","human-AI interaction","facial composite generation","interactive systems","computational user model"],"falsifier":"Recruit new users to rank the same target and auxiliary sets used in training, and compare each user's rank orders with the model's noisy rankings: if per-user Kendall tau is near zero, or if a model trained on deterministic rankings matches the noisy model on the lineup task, the claim that user-model fidelity drives reconstruction quality is contradicted. A second check is to run lineups that do not include the true target, a condition the paper notes its own study does not use.","tokens_in":22678,"feed_emoji":"🧠","tokens_out":8154,"duration_ms":72285,"temperature":0.7,"pith_summary":"HAIFAI is a two-stage interactive system for reconstructing a face that exists only in a person's memory. In the first stage, the user repeatedly ranks six candidate faces by resemblance to their mental image, and a transformer-based reconstruction network fuses these rankings and predicts a latent vector that a StyleGAN2 generator decodes into an initial portrait. To avoid collecting enough human ranking data for training, the paper introduces a computational user model that adds uniform noise to cosine similarities computed by a fine-tuned ArcFace embedding network, simulating the variability of human ranking behaviour. In a 12-participant comparison with CG-GAN and the earlier MFRS system, HAIFAI reports better visual similarity, embedding similarity, usability, perceived workload, and reconstruction speed for its first stage, and a lineup study with 18 raters reports a 60.6% identification rate. The paper's central claim is that ranking-based interaction plus a learned fusion model gives faster, less burdensome, and more accurate facial composites than evolutionary or manual-edit alternatives.","feed_headline":"Ranking faces from memory reconstructs them at a 60.6% ID rate","feed_subtitle":"Ranking-based composites beat prior systems on speed, quality, workload, and forensic identification.","key_machinery":"The load-bearing mechanism is a pair of stacked transformer encoders: a Siamese transformer reads each iteration's six ranked auxiliary latents plus a class token, and a second transformer combines the per-iteration feature vectors into a single predicted latent $w_{\\mathrm{rec}}$. The other essential component is the computational user model (Algorithm 1), which orders auxiliary faces by cosine similarity to the target in a fine-tuned ArcFace embedding space and adds uniform noise $U(-\\sigma,\\sigma)$ with $\\sigma=0.22$ so simulated rankings have human-like variability; the ablation shows this noise raises test embedding similarity from 0.397 to 0.421 and prevents overfitting. Stage two uses UP-FacE's 24 landmark-derived sliders for manual refinement.","core_discovery":"The central claim is that a user's mental face image can be recovered from rank orders over small sets of auxiliary faces, because those ordinal signals carry enough information to locate the target in a generative latent space. The reconstruction network is trained end-to-end to map ranked auxiliary latents to a reconstructed latent $w_{\\mathrm{rec}}$, using a loss that combines squared error in latent space with cosine similarity in a fine-tuned ArcFace embedding space, and the computational user model supplies synthetic ranking data by adding uniform noise to embedding similarities. Against CG-GAN and MFRS, HAIFAI's first stage is faster (8.3 vs. 10.2 vs. 17.8 minutes), rates higher on usability (SUS 87 vs. 85 vs. 59), and scores better on visual similarity and embedding similarity; the full system achieves a 60.6% identification rate in a forensic-style lineup, up from 56.1% for CG-GAN, with the target ranked in the top three in 98.5% of cases.","pith_inferences":["A natural extension the paper leaves implicit is a personalized user model: calibrating the noise level or embedding per user could reduce the required ranking iterations below the observed 10-19 while extracting a stronger signal from each ranking.","The same ranking-plus-fusion pipeline could apply to other mental-image domains, such as objects, scenes, or voices, whenever a pretrained generator and a similarity embedding exist, since neither component is face-specific except for the fine-tuned embedding.","Because the lineup study always includes the true target, the 60.6% figure is an upper bound for realistic open-set identification; an open-set lineup with foils replacing the target would measure performance when memory is imperfect and the target may not be present."],"forward_implications":["If the central claim holds, forensic witnesses can produce identification-ready composites by ranking six-face sets, with a 60.6% chance that the true target is placed first in a four-person lineup and a 98.5% chance it is in the top three.","Ranking-based interaction shifts high-dimensional face search from user to system, which is why first-stage reconstruction time drops to 8.3 minutes compared with 17.8 for CG-GAN and 10.2 for MFRS.","Training on a synthetic user model removes the need for large-scale human ranking data, making the approach transferable to other generative domains at low collection cost.","The two-stage design implies that the AI should produce a global approximation first and let humans make only targeted local refinements, reducing the 'mental shift' that heavy manual editing causes."],"supporting_citations":[{"why":"Defines the MFRS baseline that HAIFAI must beat and supplies the target face pool and lineup protocol used in the user studies.","marker":"[76]"},{"why":"Provides the CG-GAN baseline, the previous identification rate of 56.1%, and the hybrid evolutionary-manual interaction paradigm HAIFAI contrasts with.","marker":"[99]"},{"why":"Supplies the StyleGAN2 generator and its W latent space, which the reconstruction network predicts into and UP-FacE edits.","marker":"[38]"},{"why":"Supplies the ArcFace embedding network used for computing face similarity in the computational user model and in the training loss.","marker":"[15]"},{"why":"Supplies the UP-FacE slider interface for the second-stage manual refinement of facial shape.","marker":"[77]"},{"why":"Provides the FFHQ face dataset that StyleGAN2 is trained on, the source of the face distribution HAIFAI operates in.","marker":"[37]"}],"fun_headline_variants":["Human-AI face reconstruction hits 60.6% ID rate","Ranking faces reveals mental images at 60.6% ID","Interactive AI reconstructs remembered faces with 60.6% accuracy","Human-in-the-loop face recovery achieves 60.6% identifications","New interactive method tops prior face-reconstruction systems"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole training scheme rests on the assumption that ranking faces by noisy cosine similarity in a fine-tuned face-embedding space reproduces how real people rank faces held in memory; the paper's support for this is an average Kendall tau of 0.284 between model and human rankings, close to the 0.267 human-human agreement, which does not guarantee that individual users or forgotten memories behave like the model.","fun_headline_variants_meta":{"raw":{"variants":["Human-AI face reconstruction hits 60.6% ID rate","Ranking faces reveals mental images at 60.6% ID","Interactive AI reconstructs remembered faces with 60.6% accuracy","Human-in-the-loop face recovery achieves 60.6% identifications","New interactive method tops prior face-reconstruction systems"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000509,"raw_usage":{"total_tokens":2514,"prompt_tokens":1017,"completion_tokens":1497,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":633,"completion_tokens_details":{"reasoning_tokens":1409}},"tokens_in":633,"tokens_out":1497,"duration_ms":11176,"temperature":1.0,"reasoning_tokens":1409,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:46:02.345419+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recruit new users to rank the same target and auxiliary sets used in training, and compare each user's rank orders with the model's noisy rankings: if per-user Kendall tau is near zero, or if a model trained on deterministic rankings matches the noisy model on the lineup task, the claim that user-model fidelity drives reconstruction quality is contradicted. A second check is to run lineups that do not include the true target, a condition the paper notes its own study does not use.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the MFRS baseline that HAIFAI must beat and supplies the target face pool and lineup protocol used in the user studies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the CG-GAN baseline, the previous identification rate of 56.1%, and the hybrid evolutionary-manual interaction paradigm HAIFAI contrasts with."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the StyleGAN2 generator and its W latent space, which the reconstruction network predicts into and UP-FacE edits."},{"cited_title":"UP-FacE: User-predictable Fine-grained Face Shape Editing","cited_arxiv_id":"2403.13972","evidence_quote":"Supplies the UP-FacE slider interface for the second-stage manual refinement of facial shape."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the FFHQ face dataset that StyleGAN2 is trained on, the source of the face distribution HAIFAI operates in."}],"review_version":1}