{"id":"797d65e9-4128-4118-9d58-5d33f2b08134","arxiv_id":"2509.20106","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Texture resolution drives perceived realism and quality of VR digital twins, geometric detail affects only quality, and prior exposure to the physical object had no significant effect.","lead":"In a virtual-reality study with 24 people, the authors compared how texture resolution and geometric detail change people's ratings of quality and realism for digital copies of everyday objects, and whether seeing the real object beforehand changes those ratings. The results point to texture resolution as the main driver of perceived realism, while a prior glimpse at the physical object made no measurable difference.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Null exposure effect rests on an unverified manipulation the authors themselves doubt; explicit-exposure replication is needed before interpreting the null.","rationale":"The reader's weakest_assumption correctly identifies the exposure manipulation as the most load-bearing weak point. The paper's own limitation statement in Section 5.4 validates this: the authors question whether placing objects on the table was sufficient. The statistical power analysis in Section 3.2 further confirms that the between-subjects comparison was underpowered for all but large effects, making the null result fragile. The fidelity findings, by contrast, are robust: the within-subjects texture effect has enormous effect sizes (partial eta^2=0.820 and 0.709) and is confirmed by robust resampling tests; geometry effects on quality are consistent. These conclusions are not threatened. I considered whether the pseudoreplication in the correlation analysis (pooling repeated measures) is more serious, but that affects only the secondary correlational claims (H7, H8), not the central exposure/fidelity claims. Therefore I agree with the reader's conditional verdict: the paper contributes solid fidelity evidence, but the exposure null should be treated as inconclusive pending a stronger manipulation and larger sample. No verdict change is needed.","tokens_in":13716,"tokens_out":2349,"duration_ms":18905,"concrete_test":"Run a pre-registered replication with the same stimuli and procedure but two changes: (1) exposure group receives an explicit 60-second structured inspection of the physical objects with instructions to memorize surface details, followed by a brief recall/recognition check to confirm encoding; (2) recruit at least n=30 per group (total n=60) to achieve power for medium effects (f=0.25) on the between-subject factor. If the exposure effect remains non-significant, the original null is corroborated; if it becomes significant, the original null was an artifact of the weak manipulation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central conclusion that prior exposure has no effect on perceived quality/realism (H1, H2 rejected) depends entirely on the exposure manipulation actually inducing usable knowledge of the physical objects. Section 3.7 describes only passive placement of objects on a table during the instructor's explanation, with no dedicated inspection time, no instruction to memorize, and no manipulation check. Section 5.4 explicitly concedes: 'we may have overestimated the impact of placing objects on the table.' If participants in the exposure group did not encode the objects' texture and geometry details, the non-significant between-group effect (F(1,22)=1.059, p=.315 for quality; F(1,22)=0.374, p=.547 for realism) is uninformative rather than evidential. This is compounded by severe underpowering for the between-subject factor: with n=12 per group, the design could only detect large effects (f=0.58), whereas the observed partial eta-squared of 0.046 corresponds to a medium effect (f≈0.22). Thus the headline null could reflect manipulation failure, type II error, or both.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a mixed-design VR experiment (N=24, 12 per group) that varies texture resolution (256 vs 2084 px) and geometric detail (~100 vs ~40,000 quads) across three photogrammetric digital twins, and manipulates prior exposure to the physical objects by placing them on the table for the exposure group. Perceived quality and realism are measured via in-VR 7-point items. The authors report a very large effect of texture on both quality (F=100.549, p<.001, η²=0.820) and realism (F=53.705, p<.001, η²=0.709), a significant but smaller effect of geometry on quality only, no significant exposure effect, and positive correlations between quality/realism and age/realism. They conclude that texture resolution should be prioritized over geometry in DT development.","tokens_in":13995,"tokens_out":5416,"duration_ms":37748,"significance":"If the fidelity effects are valid, the study provides actionable evidence that texture resolution is the dominant fidelity dimension for perceived quality and realism of object-level DTs in VR. The within-subjects results are statistically robust, including resampling-based WTS/ATS checks, and the effect sizes are very large. The exposure null, however, is not yet a substantive contribution: the manipulation is implicit and unvalidated, and the between-subjects design is underpowered. The paper's explicit acknowledgment of these limitations is commendable, but as presented the exposure finding is inconclusive rather than evidential.","major_comments":[{"comment":"The null result for H1/H2 is uninterpretable because the exposure manipulation was not validated. The procedure (Section 3.7) consisted only of placing the physical objects on the table next to consent forms while the instructor explained the procedure; no dedicated inspection time, no instruction to memorize, and no manipulation check are reported. Section 5.4 acknowledges this ('we may have overestimated the impact of placing objects on the table'). The non-significant between-group F-tests (F(1,22)=1.059, p=.315 for quality; F(1,22)=0.374, p=.547 for realism) therefore cannot support the conclusion that prior exposure has no effect. A manipulation check (e.g., comparing Q1 familiarity ratings between groups) should be reported, or the manuscript should be reframed as an inconclusive pilot.","section":"§3.7, §5.4, Tables 2–3"},{"comment":"The a priori power analysis shows the between-subjects factor was only powered to detect f=0.58 (partial eta²≈0.252). The observed quality effect (partial eta²=0.046, f≈0.22) is far below this, so the null is equally consistent with a Type II error as with a true absence of effect. The conclusion in §5.1 that H1 and H2 'must both be rejected' is too strong; at most, the data fail to support them. This should be stated explicitly in the abstract and conclusion.","section":"§3.2, Tables 2–3"},{"comment":"Perceived quality is measured with a single questionnaire item (Q2), while realism uses four items. The additional fidelity-specific items (Q3, Q4) were excluded post hoc due to 'unsuitable wording' (§3.8). The reliability of a single-item quality measure is not assessed. Because perceived quality is one of the two central dependent variables, the manuscript should either use a multi-item quality scale or explicitly justify the single-item measure and discuss its limitations.","section":"§3.6, §3.8"}],"minor_comments":[{"comment":"Q1 familiarity ratings are collected but never analyzed; report them at least as a check on the exposure manipulation.","section":"§3.7 / §4.3"},{"comment":"The robust analysis reports WTS and ATS values identical to the F-statistic for some effects; clarify which test statistic is being reported in each case.","section":"§4.1, §4.2"},{"comment":"The photograph of the physical objects would benefit from a scale reference to convey object size.","section":"Figure 3"},{"comment":"The VRCS scale is described as 'Visual Realism Classification Scale' but the cited source uses 'Complexity Scale'; verify the name.","section":"§2 Measurements"}],"recommendation":"major_revision","confidential_remarks":"The within-subjects fidelity results are solid and likely publishable. The main obstacle is the exposure claim: the manipulation is unvalidated and the test underpowered, making the headline null inconclusive. Because the authors already identify the appropriate fix (explicit exposure with dedicated inspection time), major revision rather than rejection seems appropriate. If the paper is reframed around the fidelity findings, the exposure result may be reported as a pilot or secondary outcome."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The within-subject fidelity results are the real content here: texture resolution had very large effects on both perceived quality (partial eta² = 0.82) and realism (0.71), while geometry mattered for quality only (0.17) and not realism. That is a clean, useful dissociation, and it supports the practical takeaway that asset pipelines should prioritize textures over mesh detail. The robust resampling checks and the power analysis add credibility; the authors also write plainly about what they did and what they didn't find.\n\nThe soft spot is the exposure claim. The manipulation was passive — objects placed on a table during the briefing, no dedicated inspection, no manipulation check. With n=12 per group, the between-subjects comparison could only detect large effects (f≥0.58), and the observed effect (partial eta² ≈ 0.046) is consistent with either a small effect or noise. The authors themselves admit they \"may have overestimated the impact of placing objects on the table,\" which is honest — but it means the abstract's flat statement that exposure had \"no significant effect\" overstates the evidence. That null should be read as a failed or weak manipulation, not evidence of absence. If I were reviewing, I'd ask them to either drop the exposure hypotheses or reframe as a pilot.\n\nTwo smaller issues. First, the correlation analysis pools 288 repeated measures (24 participants × 12 conditions) as if they were independent; the reported age-realism correlation (rho=0.231, p<0.001) is probably inflated by this. Second, Q3 and Q4 were excluded after the fact because they \"did not provide additional value\" — plausible, but a robustness check or pre-registration would have been cleaner.\n\nOverall, this is a competent, honest study with one robust finding and one under-supported claim. It deserves peer review, not because the exposure story is convincing, but because the fidelity result is useful and the methodological caution about manipulation checks is worth airing. I'd send it, with the expectation that the exposure section gets rewritten or scaled back.","headline":"Texture dominates perceived quality and realism in VR digital twins; the exposure null is not to be trusted because the manipulation is unverified and underpowered.","tokens_in":14421,"tokens_out":2128,"would_cite":true,"duration_ms":39432,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Texture resolution dominates perceived quality and realism of VR digital twins.","keywords":["perceived realism","perceived quality","digital twin","virtual reality","texture resolution","geometric detail","prior exposure","fidelity"],"falsifier":"A follow-up study that gives participants explicit, time-controlled physical inspection (e.g., 60 seconds of handling and a familiarity check) and finds a significant effect of exposure on realism or quality ratings would directly contradict the paper's null exposure result; likewise, showing that geometry changes realism ratings for larger or room-scale digital twins would limit the generality of the texture-dominance claim.","tokens_in":13657,"feed_emoji":"🥽","tokens_out":5139,"duration_ms":40152,"temperature":0.7,"pith_summary":"This paper asks whether seeing a real object before inspecting its virtual replica in VR changes how people judge the replica's quality and realism, and which component of visual fidelity matters most. The authors created photogrammetric digital twins of three everyday objects, varied texture resolution (256 vs 2084 pixels) and geometric detail (~100 vs 40,000 quads), and had 24 participants rate each variant in a headset. They find that texture resolution has a very large effect on both perceived quality and realism, while geometric detail significantly affects quality but not realism. They find no significant effect of prior exposure to the physical objects. The practical upshot: when building VR digital twins, put effort into high-resolution textures and simplify geometry freely.","feed_headline":"Texture, not geometry, drives VR digital twin perception","feed_subtitle":"Users judge VR copies of real objects mainly by texture; geometry affects quality but not realism.","key_machinery":"The central mechanism is systematic independent variation of the two fidelity factors in photogrammetry-based replicas of office objects, combined with a mixed-design ANOVA and robust resampling tests. Texture resolution is set to 256x256 versus 2084x2084 pixels, and geometry to roughly 100 versus 40,000 quads, with identical lighting and material settings so that each factor's effect can be isolated. The in-VR questionnaire, which adapts the realism subscale of the Igroup Presence Questionnaire plus a Mean Opinion Score quality item, provides the dependent measures.","core_discovery":"Using a mixed design with 24 participants, the study shows that texture resolution is the dominant visual cue for judgments of both perceived quality (F(1,22)=100.55, p<.001, partial eta-squared=0.820) and perceived realism (F(1,22)=53.71, p<.001, partial eta-squared=0.709) of digital twins in VR. Geometric detail has a statistically significant but much smaller effect on perceived quality and no significant effect on perceived realism. Prior exposure to the physical counterpart produced no significant effect on either rating. The authors interpret the texture dominance as consistent with its immediate visibility and a known texture-substitutes-for-geometry effect, and they caution that the","pith_inferences":["Because the paper's exposure was implicit (objects on a table, no forced inspection), the null result should not be read as evidence that object knowledge never shapes VR judgment; an explicit-exposure replication, as the authors themselves call for, is the needed test.","Ratings were made from memory after the object was hidden, so the fidelity effects reflect recalled impressions; simultaneous side-by-side rating might change the geometry-texture balance.","The texture-over-geometry dominance may be specific to small, isolated objects; when scaled to rooms or buildings, where geometry defines spatial structure, geometry could matter more for realism.","A practical implication left implicit: photogrammetry capture should invest in high-resolution texture acquisition and compression, since aggressive mesh decimation appears relatively safe for perceived realism."],"forward_implications":["VR digital twin developers can prioritize texture resolution over geometric detail when aiming for perceived quality and realism.","Mesh simplification down to roughly 100 quads may preserve perceived realism as long as textures stay high-resolution, offering a low-cost optimization path.","The null exposure effect, if taken at face value, challenges the assumption that physical familiarity raises expectations and lowers realism ratings.","Perceived quality and realism move together strongly (rho=0.802), so single-metric evaluations may be sufficient for many design decisions.","Age correlates weakly with realism perception (rho=0.231), consistent with younger users having higher expectations."],"fun_headline_variants":["Texture beats geometry in VR twin perception","Geometry only shapes quality in VR digital twins","Prior exposure has no effect on VR twin realism","VR twin ratings hinge on texture, not geometry","Texture dominates how we judge VR digital twins"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The prior-exposure manipulation may not have actually made participants familiar with the physical objects, because exposure consisted only of the objects sitting on a table during instruction, so the null result could reflect weak manipulation rather than evidence that prior knowledge does not matter.","fun_headline_variants_meta":{"raw":{"variants":["Texture beats geometry in VR twin perception","Geometry only shapes quality in VR digital twins","Prior exposure has no effect on VR twin realism","VR twin ratings hinge on texture, not geometry","Texture dominates how we judge VR digital twins"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000301,"raw_usage":{"total_tokens":1542,"prompt_tokens":681,"completion_tokens":861,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":425,"completion_tokens_details":{"reasoning_tokens":794}},"tokens_in":425,"tokens_out":861,"duration_ms":7465,"temperature":1.0,"reasoning_tokens":794,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T15:12:57.114324+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A follow-up study that gives participants explicit, time-controlled physical inspection (e.g., 60 seconds of handling and a familiarity check) and finds a significant effect of exposure on realism or quality ratings would directly contradict the paper's null exposure result; likewise, showing that geometry changes realism ratings for larger or room-scale digital twins would limit the generality of the texture-dominance claim.","supporting_citations":[],"review_version":1}