{"id":"1a3c8d54-4120-4f27-8dd7-608d749ebbb4","arxiv_id":"2411.10489","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature review on biometric authentication, attacks, photorealistic avatars, and datasets in extended reality, built around a proposed vulnerability taxonomy.","lead":"This paper surveys how biometrics are used in extended reality (XR), covering authentication methods, attack surfaces, avatar generation, and datasets. It proposes a taxonomy of XR vulnerability points and claims to be the first systematic treatment of this topic.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed first taxonomy of XR biometric vulnerability gateways is internally inconsistent as presented: §2.3 declares four points, Figure 6 says three, §2.3.3 promises a point 4 that never appears, and RQ1 answers three.","rationale":"I read the vulnerability section in full because the paper's headline novelty rests on it. The four/three mismatch, the three numbered subsections, the unreachable 'vulnerability point 4' promised in §2.3.3, and the RQ1 answer's three-point summary are all present in the manuscript text; no external corpus is needed to see them. This is an internal failure of the central deliverable: a taxonomy that cannot be enumerated from its own text cannot serve as a benchmark for a 'first in the literature' claim. I regard this as more load-bearing than the Section 1 article-count inconsistency (318 vs 391), because even a perfectly transparent search procedure would not repair a taxonomy that is inconsistent in its own definitions. I do not reject the paper outright: the modality, dataset, and avatar sections assemble a usable set of references, and a revision that reconciles the vulnerability-point definitions, removes the dangling point-4 promise, and adds the missing list of the 62 shortlisted papers could make the survey valuable. The reader's CONDITIONAL verdict already requires fixing these inconsistencies, so my read leaves the verdict unchanged. I mark agreement as partial because the reader's weakest_assumption was sampling adequacy, whereas I locate the load-bearing weakness in the internal consistency of the vulnerability taxonomy that the paper itself claims as its first contribution.","tokens_in":34169,"tokens_out":5563,"duration_ms":54053,"concrete_test":"Produce a completeness audit of the vulnerability taxonomy: create one row per nominal vulnerability point (1 through 4) with columns for section heading, defining sentence, attacks assigned, and where the point is counted in the 'three' or 'four' statements. Then enumerate the distinct definitions that actually appear. If vulnerability point 4 has no defining sentence and no section, the taxonomy as published is incomplete, and the 'first taxonomy' claim cannot be verified until the section is written and the three/four statements are reconciled.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's advertised contribution is a first systematic taxonomy of biometric vulnerability gateways in XR (Abstract; Section 1). For that claim to hold, the taxonomy must be determinately enumerable. As published it is not. Section 2.3 says 'we identified four different vulnerability points, as illustrated in Figure 6,' but Figure 6's caption states 'In total, three vulnerability points need to be addressed,' and the RQ1 answer in Section 7 says 'we identified three primary points of vulnerability.' The body provides subsections 2.3.1, 2.3.2, and 2.3.3; the opening of 2.3.3 says 'we discuss the possible attacks at vulnerability point 4,' but no Vulnerability Point 4 definition or subsection follows. If the taxonomy cannot be enumerated consistently from the paper itself, the claimed first taxonomy cannot be checked against prior work, and the central contribution fails as written. This is an internal correctness problem, not a dispute about external literature coverage. The missing list of the 62 shortlisted papers is a separate coverage issue; the taxonomy inconsistency is more directly load-bearing.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript surveys biometrics in extended reality (XR), with sections on authentication modes, biometric vulnerability gateways and attacks, physiological/behavioral verification methods, photorealistic avatar generation and verification, datasets, and performance metrics. The paper is organized around four research questions and claims two firsts: a systematic discussion of biometric vulnerability gateways in XR with a taxonomy, and an exclusive survey of XR biometrics. The survey aggregates a broad set of references into tables of biometric modalities, datasets, avatar-generation methods, and prior surveys.","tokens_in":34416,"tokens_out":3443,"duration_ms":35336,"significance":"If corrected, the survey would address a real gap in the literature: prior reviews cover VR security, gaze privacy, or metaverse threats, but none ties the XR biometric workflow to a vulnerability taxonomy while also covering avatar generation and verification. The main value would be the organizing framework—the vulnerability gateways—and the consolidated tables of datasets and methods. The paper also makes a checkable priority claim ('first in the literature'), which becomes testable once the reviewed corpus is fully listed. The contribution is taxonomic and expository rather than computational, so the central quality bar is internal consistency of the taxonomy and reproducibility of the selection procedure; both currently need work.","major_comments":[{"comment":"The central claim of a first systematic taxonomy of biometric vulnerability gateways cannot currently be enumerated from the paper as written. Section 2.3 states that 'we identified four different vulnerability points, as illustrated in Figure 6,' while the Figure 6 caption says 'In total, three vulnerability points need to be addressed.' The RQ1 answer in Section 7 likewise says 'we identified three primary points of vulnerability.' Section 2.3.3 is headed 'Vulnerability Point 3' but its final sentence promises 'we discuss the possible attacks at vulnerability point 4,' and no Vulnerability Point 4 definition or subsection appears anywhere. Please reconcile the count, align the figure caption with the text, and make the one-to-one mapping from each numbered vulnerability point to a subsection and figure element explicit.","section":"§2.3, Figure 6, §2.3.3, §7 RQ1"},{"comment":"The reported article pool is inconsistent and the shortlist is not auditable. Section 1 says 'we collected 318 articles from Google scholar,' but the second screening step says 'we conducted a comprehensive review of the 62 shortlisted papers out of the initial 391 articles.' The reader cannot determine whether 318 or 391 articles formed the initial corpus, and the 62 selected papers are never listed, nor are the exclusion criteria defined. This undermines the 'systematic' and 'comprehensive' claims and the priority claim, because a reader cannot verify the coverage without reconstructing the full search and screening. Please add a PRISMA-style flow diagram with exact counts, a complete list of the 62 shortlisted papers, and a description of inclusion/exclusion decisions.","section":"Section 1, Methodology"}],"minor_comments":[{"comment":"The Sun et al. [2022] 2023 Hand Movement row appears twice with identical sample counts; one duplicate should be removed.","section":"Table 7"},{"comment":"Table 1 lists Giaretta [2022] twice with different descriptions and labels De Guzman et al. [2019] under year 2023; please unify these entries and correct the years to match the reference list.","section":"Table 1"},{"comment":"Some attribution details are inconsistent: Table 4 lists Szymanowicz et al. [2023] for PointAvatar and PiCA, whereas the text attributes PointAvatar to Zheng et al. [2022] and PiCA to Ma et al. [2021]; Table 5 is also very terse for a survey that promises comprehensive coverage. Please align the table entries with the cited works.","section":"Tables 4 and 5 and surrounding text"},{"comment":"The paragraph beginning 'The The Brain signals...' contains a duplicated definite article; please correct this and similar typographical errors elsewhere, including 'formualating' (Section 1), 'Vulnerbality' (Section 2.3.3), 'Realilty' (Appendix), 'characterics' (Abstract and Section 3), and the spurious spaces in 'V oice' and 'V AE'.","section":"Section 5.1"},{"comment":"The figure summarizing research-question answers should be harmonized with Section 7: the RQ1 box lists attack types rather than the 'three primary points of vulnerability' stated in the text, and the RQ4 box lists modalities rather than dataset categories, which makes the figure potentially misleading.","section":"Figure 11"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the vulnerability-point inconsistency lands directly and is load-bearing for the paper's main novelty claim. The methodology inconsistency (318 vs 391 articles) is also reproducible and should be fixed with a transparent screening flow. The paper is a reasonable survey effort, and the issues appear fixable within the manuscript's scope, so I do not recommend rejection. The heavy self-citation pattern is worth editorial awareness, but it does not by itself affect my recommendation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a survey, not a new result. What it does well is collect a lot of material on biometrics in XR—authentication modalities, avatar generation and verification, datasets, metrics—into one place. The avatar generation timeline and the dataset tables are genuinely handy if you're starting in this area. It also has a reasonable open-challenges discussion, and the individual citations for specific biometric methods look about right.\n\nThe soft spots are real and not minor. The claimed contribution is a 'first' systematic taxonomy of biometric vulnerability gateways. As published, that taxonomy cannot be consistently enumerated. Section 2.3 says there are four vulnerability points; Figure 6's caption says three; the RQ1 answer says three; and subsection 2.3.3 explicitly promises a vulnerability point 4 that never appears. That is an internal inconsistency in the paper's central organizing framework, not a dispute about external coverage. The stress-test note is accurate.\n\nThe priority claim also needs work. The paper says it is the first to discuss biometric vulnerability gateways in XR, yet Table 1 lists Heruatmadja et al. 2023 (a systematic literature review of biometric authentication in VR) and Qamar et al. 2023 (a systematic threat analysis for XR/metaverse). Maybe the authors can argue those papers don't do exactly what they do, but they need to make that argument instead of asserting priority.\n\nThe methodology has inconsistencies too: 318 articles collected in one place, 391 in another, and the 62 shortlisted papers are never listed. That makes the 'comprehensive' claim unverifiable. There are also many small errors—duplicate rows in Table 7, wrong years in Table 4, repeated typos. These are fixable.\n\nOn the other hand, I did not find evidence that the survey misrepresents the individual papers it cites. The descriptions of iris, periocular, EEG, gaze-based authentication, and the avatar generation works match what I know of those works. So the substance is largely sound; the presentation is not.\n\nVerdict: conditional. This deserves a serious referee, but not acceptance as is. The authors need to fix the taxonomy numbering, list the selected papers, reconcile the article counts, and soften or substantiate the novelty claim. If they do, this becomes a useful reference for XR security researchers and device makers. If they don't, the central contribution remains uncheckable and I would not rely on it.","headline":"A useful but sloppy survey of biometrics in XR; the central vulnerability taxonomy is internally inconsistent as written, so it needs a real revision before it can be trusted.","tokens_in":34963,"tokens_out":3230,"would_cite":false,"duration_ms":31246,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review claims to be the first systematic map of biometrics in extended reality, organizing attacks around three vulnerability points and classifying modalities and avatars into taxonomies.","keywords":["extended reality","biometrics","authentication","virtual reality security","presentation attacks","avatar generation","behavioral biometrics","XR datasets"],"falsifier":"To test the priority claim, a reader could search the literature up to the paper's November 2024 cutoff for any earlier survey that already taxonomizes biometric vulnerability points in XR systems; if a pre-2024 paper presents the same three-gate model (device authentication, interaction channel, avatar space), the 'first' claim fails. To test the coverage claim, an independent team could reproduce the 62-paper shortlist from the stated keywords and screening steps; the reported pool sizes of 318 and 391 are inconsistent, so the shortlist is not currently auditable from the manuscript.","tokens_in":33979,"feed_emoji":"🕶️","tokens_out":9300,"duration_ms":78900,"temperature":0.7,"pith_summary":"The paper tries to establish that biometrics in Extended Reality form a research area with a recognizable structure, and that the area has been missing a security map. It claims to supply that map for the first time, in the form of three taxonomies: vulnerability points in a biometric XR system, physiological and behavioral biometrics for authentication, and photorealistic avatar generation and verification. A sympathetic reader should care because commercial XR devices already capture iris and gaze data, and a photorealistic avatar is effectively a biometric credential that can be stolen or forged; without an agreed map, security measures, datasets, and benchmarks develop in isolation. If the survey is right, the field gains a common vocabulary and a checklist of attack surfaces to defend.","feed_headline":"Three gates where XR biometrics can be attacked","feed_subtitle":"A new review maps iris, gaze, and avatar biometrics onto device, channel, and virtual-space attack points.","key_machinery":"The organizing device is the four-block XR workflow—input processor, simulation processor, rendering processor, XR environment—with three named vulnerability points. Vulnerability point 1 sits at initial authentication and is the home of presentation and shoulder-surfing attacks; vulnerability point 2 sits on the interaction and sensor channel and hosts side-channel and injection attacks; vulnerability point 3 sits in the rendered virtual space and hosts avatar impersonation, morphing, and deepfake attacks. The same device is paired with a two-axis biometric classification (physiological versus behavioral) and an avatar classification (2D versus 3D, static versus continuous, generation versus verification), so that each reviewed paper is slotted by which block of the workflow it touches and which gate it defends or exposes.","core_discovery":"On its own terms, the paper's central discovery is that a biometric-enabled XR system has a small, enumerable set of attack gates, and that every reviewed technology can be placed at one of them. The first gate is initial device authentication in the physical world, where presentation attacks (printed artifacts, masks, replay) and shoulder surfing occur. The second is the channel that carries the user's captured motion and behavior into the virtual environment, where side-channel extraction and biometric sample injection occur. The third is the rendered virtual space, where an attacker can impersonate an avatar through man-in-the-room techniques, morphing, deepfakes, or stream exfiltration. The paper further claims that physiological biometrics (iris, periocular region, fingerveins, brain and heart signals) and behavioral biometrics (gaze, hand motion, gait, air handwriting, voice) are both in use in XR, and that deep-learning-generated photorealistic avatars now carry enough biometric fidelity to act as identity tokens. It concludes that secure XR needs layer-specific defenses at all three gates plus continuous avatar verification.","pith_inferences":["The three-gate model implies, though the paper does not say so directly, that continuous behavioral biometrics are the only layer able to verify identity after an avatar is rendered, because static enrollment credentials cannot follow the user into the scene.","A direct test of the survey's reproducibility would be to reconstruct the 62-paper shortlist from the reported search keywords and screening steps; the manuscript reports inconsistent starting pool sizes (318 versus 391 articles) and never lists the selected papers, so an independent check could confirm or overturn the coverage claim.","The vulnerability model transfers naturally to multi-user XR collaboration, where one participant's continuous authentication stream becomes another participant's attack surface—a scenario the survey leaves implicit in its man-in-the-room discussion.","The separation of avatar generation from avatar verification suggests an unstudied attack class: adversarially generating an avatar that preserves a victim's biometric identity well enough to pass face or gait verification while looking visually different to human observers."],"forward_implications":["Defenses can be assigned to distinct layers: anti-spoofing and liveness checks at device entry, tamper-resistant capture and encryption on the interaction channel, and continuous biometric verification of avatars inside the rendered scene.","Dataset construction can become systematic, with new XR biometric collections classified by modality and by the vulnerability point they stress, as the survey does for the pupil, iris, periocular, EEG, ECG, gaze, and gait datasets it lists.","Benchmarking should move beyond EER and accuracy toward XR-specific measures such as immersion score and latency, which the paper argues are needed to judge real-time authentication in a headset.","Commercial iris-authenticated headsets can be stress-tested against the enumerated presentation-attack instruments, since the survey identifies iris as an effective modality that remains vulnerable to presentation attacks."],"supporting_citations":[{"why":"Closest prior discussion of biometric authentication in VR; the survey positions its detailed biometric treatment against this work.","marker":"Kürtünlüoğlu et al. [2022]"},{"why":"Prior review of VR authentication systems that helps establish the gap the survey aims to fill.","marker":"Jones et al. [2021]"},{"why":"Systematic literature review of biometric authentication in VR, used as a comparison baseline in the survey's Table 1.","marker":"Heruatmadja et al. [2023]"},{"why":"Threat analysis for metaverse and XR systems that the survey compares against its own vulnerability taxonomy.","marker":"Qamar et al. [2023]"},{"why":"Defines presentation attacks and presentation attack instruments, the conceptual basis for vulnerability point 1.","marker":"Ramachandra and Busch [2017a]"},{"why":"OcuLock gaze-and-iris HMD authentication, the anchor example for multimodal physiological-behavioral verification in XR.","marker":"Luo et al. [2020]"},{"why":"FitMe photorealistic avatar generation pipeline, anchor for the avatar generation and biometric fidelity discussion.","marker":"Lattas et al. [2023]"},{"why":"VRBiom periocular dataset captured with a consumer headset, anchor for the XR biometric dataset inventory.","marker":"Kotwal et al. [2024]"}],"fun_headline_variants":["Three gates that guard—or betray—XR biometrics","Every XR biometric attack fits one of three gates","Review pins XR biometric hacks to three chokepoints","XR identity risk: three gates, one map","The three entry points for XR biometric attacks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's usefulness and its 'first in the literature' claim rest on the 62 shortlisted papers being a fair and complete sample of XR-biometrics work, but the paper never lists those 62 papers or the reasons for exclusion, and its method section reports two different starting numbers, so a reader cannot currently check whether one missed prior survey or a skewed selection would have changed the taxonomy.","fun_headline_variants_meta":{"raw":{"variants":["Three gates that guard—or betray—XR biometrics","Every XR biometric attack fits one of three gates","Review pins XR biometric hacks to three chokepoints","XR identity risk: three gates, one map","The three entry points for XR biometric attacks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000195,"raw_usage":{"total_tokens":1353,"prompt_tokens":940,"completion_tokens":413,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":556,"completion_tokens_details":{"reasoning_tokens":335}},"tokens_in":556,"tokens_out":413,"duration_ms":4347,"temperature":1.0,"reasoning_tokens":335,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:24:39.232152+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"To test the priority claim, a reader could search the literature up to the paper's November 2024 cutoff for any earlier survey that already taxonomizes biometric vulnerability points in XR systems; if a pre-2024 paper presents the same three-gate model (device authentication, interaction channel, avatar space), the 'first' claim fails. To test the coverage claim, an independent team could reproduce the 62-paper shortlist from the stated keywords and screening steps; the reported pool sizes of 318 and 391 are inconsistent, so the shortlist is not currently auditable from the manuscript.","supporting_citations":[{"cited_title":"Biometric as secure authentication for virtual reality environment: A systematic literature review","cited_arxiv_id":null,"evidence_quote":"Systematic literature review of biometric authentication in VR, used as a comparison baseline in the survey's Table 1."}],"review_version":1}