{"id":"7a470b25-e650-4cd9-805a-1a78083a98e0","arxiv_id":"2501.10338","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In VR, curved wall-sized displays improve visual variable perception accuracy over flat displays, but tasks take longer; interaction techniques further reduce errors.","lead":"This paper ran two virtual reality experiments asking people to match the size of lines, angles, and circles on large virtual display walls. It found that curved virtual walls led to fewer estimation errors than flat ones, and that letting users move or select displays improved accuracy.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-study comparison with the physical flat wall is not matched on viewing geometry and apparatus; the headline curved-wall error advantage over RFlat depends on an uncontrolled between-study contrast.","rationale":"I agree with the reader's assessment. The within-VR comparisons are internally valid: the study uses a between-subjects design for arrangements, bootstrapped confidence intervals with Bonferroni correction, and the Flat versus Cylinder/Cockpit difference is a controlled contrast. The load-bearing weakness is the external comparison with RFlat, which is exactly where the abstract makes its strongest claim. The reader's weakest assumption captures this concern; I would add the specific start-position mismatch and emphasize that the paper's own limitation statements in Sections 4.5 and 4.6 concede the point. The remedy is either a matched follow-up experiment or a softened claim, which is what the reader's conditional verdict already reflects. Therefore no change to the verdict is needed: the conditional status appropriately signals that the headline cross-study result requires verification before it can be accepted as a general finding.","tokens_in":25771,"tokens_out":8288,"duration_ms":84741,"concrete_test":"Run a within-subjects follow-up with roughly 20-25 participants covering four conditions: a physical flat wall at 3.2 m from the display center, virtual Flat, virtual Cylinder, and virtual Cockpit, using identical stimulus and modulus values, identical start position relative to the wall, identical adjustment controls, and identical trial counts. Compute the AbsErr differences between conditions. If the curved-wall advantage over the physical flat wall does not reproduce when geometry and controls are matched, the abstract's cross-study claim must be downgraded to an exploratory finding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that virtual curved walls yield smaller absolute errors than a physical flat wall rests on the Section 4.5 comparison with Bezerianos and Isenberg's RFlat. That comparison is not a controlled experiment: RFlat is an aggregate from a separate study with different participants, a different physical display, and no randomized assignment to the conditions being contrasted. Two specific mismatches matter. First, viewing geometry: in the VR study the participant starts 3.2 m from ColA, the leftmost column, whereas the prior study's 3.2 m condition is described as 3.2 m away from the display; unless the original study also anchored the origin at the left edge, the modulus distances and viewing angles are not matched. Second, rendering and input: the Vive Pro Eye's 110-degree FOV, 1440x1600 per-eye resolution, vergence-accommodation conflict, and touchpad adjustment mechanism differ from the physical wall and its adjustment controls. The paper itself acknowledges this in Section 4.5 ('we can only compare a portion of their experimental results') and in Section 4.6 notes the absence of a direct comparison between curved conditions in real and virtual environments. The within-VR Flat versus curved-wall contrast is controlled and credible; what is not established is the quantitative superiority over RFlat. The similar AbsErr between virtual Flat and RFlat (1.68 ppt difference, CI [-1.11, 3.46]) mitigates some hardware concerns, but it does not validate the curved-wall comparison, because curvature interacts with HMD resolution, depth cues, and viewing geometry in ways that the unmatched comparison cannot separate.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports two user studies on magnitude reproduction of length, angle, and area on virtual wall-sized tiled displays in VR. Study 1 (between-subjects, n=49) compares Flat, Cylinder, and Cockpit display arrangements and compares the results to a previously published physical-wall study (RFlat) [10]. Study 2 (within-subjects, n=25) evaluates four interaction techniques (Selection, Walking, Steering, Teleportation) on the Flat layout. The main claims are that curved virtual arrangements reduce absolute error relative to virtual Flat and to the physical RFlat, at the cost of longer task completion time, and that interaction techniques improve perception accuracy. The paper includes detailed bootstrap confidence-interval analyses, pre-specified hypotheses, and supplementary materials.","tokens_in":26021,"tokens_out":7323,"duration_ms":66963,"significance":"If the within-VR comparisons are taken as the paper's core contribution, the study provides useful and fairly rigorous evidence for immersive analytics: curved virtual wall layouts reduce absolute error in a magnitude-reproduction task, and several interaction techniques improve accuracy over a no-interaction baseline. The statistical analysis is careful by current HCI standards—pre-specified hypotheses, Latin-square counterbalancing, bootstrap CIs with Bonferroni correction, and explicit reporting of effect magnitudes. The movement-strategy observations and subjective measures enrich the contribution. However, the headline comparison to the physical wall is an uncontrolled between-study contrast, and the second study's Personal-condition claims lack a within-condition baseline. These issues are central to the abstract's strongest statements, so the paper's overall scientific contribution is currently overstated, though the underlying empirical work is valuable and largely sound.","major_comments":[{"comment":"The claim in the Abstract and Section 4.5.1 that Cylinder and Cockpit yield smaller absolute errors than the physical flat wall (RFlat) rests on an uncontrolled between-study comparison. The viewing geometry is not matched: participants in the present study start 3.2 m from the leftmost column ColA (Section 3.3), whereas Bezerianos and Isenberg [10] describe a 3.2 m condition as 3.2 m away from the display; unless the original study also anchored the origin at the left edge, the modulus distances and viewing angles differ. The VR apparatus (Vive Pro Eye, 110° FOV, 1440×1600 per eye, touchpad adjustment) also differs from the physical wall and its input. The authors acknowledge in Section 4.5 that they 'can only compare a portion of their experimental results.' The observed 6.93 ppt and 8.47 ppt AbsErr advantages for Cylinder and Cockpit over RFlat may therefore be artifacts of these mismatches rather than genuine perceptual benefits of curved virtual layouts. This is load-bearing because the abstract's primary comparison to the physical wall depends on it.","section":"Section 4.5, Figure 4"},{"comment":"The statistical comparison to RFlat is under-specified. Section 3.4 describes a bootstrap procedure with 10,000 BCa iterations, but the paper does not state whether the raw trial data from [10] were available or whether only published summary statistics were used. If only published means and CIs were used, the RFlat CIs (e.g., 15.9% [15.1, 16.8]) and the pairwise difference CIs (e.g., RFlat–Cylinder 6.93 ppt [3.90, 8.68]) cannot be produced by the described bootstrap; the procedure would be treating a fixed published aggregate as if it were a random sample from a comparable population. The authors should clarify the data source and, if raw data are unavailable, re-frame the RFlat comparison as descriptive and label the difference CIs as informal or adjust them to account for the aggregate nature of the external benchmark.","section":"Section 4.5, statistical method"},{"comment":"The conclusion that interaction techniques 'further improved task performance' (Abstract, Section 5.3.1) is not fully supported for the Personal stimulus location because no No Interaction baseline was collected for Personal. The reported improvements for Personal conditions (Selection, Walking, Steering, Teleportation) are all relative to 'No Interaction and Frontal' (Section 5.3.1, 'pairwise comparisons to the baseline condition (No Interaction and Frontal)'). Since Study 1 showed that Personal without interaction has substantially higher AbsErr than Frontal (7.19 ppt; Section 4.4.1), the observed improvements for Personal could reflect the absence of the Personal-display depth mismatch rather than the effect of the interaction techniques per se. I recommend either collecting a Personal No Interaction baseline or explicitly limiting the claim to Frontal conditions and to relative comparisons among interaction techniques.","section":"Section 5.1.1, Section 5.3.1"}],"minor_comments":[{"comment":"The sentence 'Frontal had a lower EstErr than Personal by 3.59ppt [0.60, 6.13]' contradicts the reported means (Frontal 6.69% vs Personal 3.10%); the direction should be reversed to 'Personal had a lower EstErr than Frontal'.","section":"Section 4.4.2"},{"comment":"Please state in the text or figure caption how the RFlat CI values were obtained (e.g., from [10]'s reported CIs, or from raw data), as this is essential for interpreting the pairwise comparisons.","section":"Section 4.5, Figure 3/4 captions"},{"comment":"The listing of initial magnitudes ('65 cm for length..., 178 degrees, and 41 cm in diameter') is grammatically ambiguous for Angle and Area. Also, the modulus multipliers are applied to 180 degrees for Angle, while the initial Angle stimulus is 178 degrees; clarify whether the Angle stimulus can initially exceed the largest modulus (0.7×180°=126°).","section":"Section 4.1.3"},{"comment":"The figure caption uses CoWall, CyWall, and FWall, but the text uses Cockpit, Cylinder, and Flat; unify the abbreviations to avoid confusion.","section":"Figure 3 caption / Section 4.4"},{"comment":"The term 'Z-fighting' is used to describe participants' strategy of overlapping the stimulus and modulus; consider explaining or glossing the term for readers unfamiliar with rendering artifacts.","section":"Section 5.4"},{"comment":"The phrase 'rejection of H5' (also in Section 4.4.1) is stronger than the CI-based analysis warrants; 'no evidence supporting H5' would be more consistent with the paper's own statistical framework.","section":"Section 4.4.1, Section 4.6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid empirical study with careful within-experiment statistical analysis, and it fits the journal's scope. The main concern is that the abstract and conclusion overstate the comparison to the physical wall: the RFlat benchmark is not a controlled condition, and the paper does not clarify whether raw data from [10] were used for the CIs. If revision appropriately softens the cross-study claim and addresses the Personal-baseline issue, I would be willing to see a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core contribution is the first careful comparison I know of magnitude reproduction for length, angle, and area on virtual wall-sized tiled displays in three layouts: flat, cylindrical, and cockpit. The within-VR result—curved walls produce lower absolute error than flat—is credible and useful for immersive analytics design. Study 2's comparison of selection, walking, steering, and teleportation is also new and mostly well executed. The statistics are careful: bootstrapped CIs, Bonferroni corrections, pre-specified hypotheses. The authors also deserve credit for stating their limitations explicitly.\n\nThe soft spot is the headline comparison with the physical wall from Bezerianos and Isenberg. That is a between-study contrast with different participants, hardware, and viewing geometry. The VR starting position is 3.2 m from the leftmost column; the prior study's 3.2 m condition is 3.2 m from the display. With a 110-degree FOV HMD and a touchpad adjustment mechanism, the conditions are not matched. The fact that virtual Flat and RFlat came out close (1.68 ppt difference) is reassuring but doesn't validate the curved-wall comparison, because curvature interacts with resolution and depth cues in ways only a controlled study can separate. The authors acknowledge they \"can only compare a portion of their experimental results,\" yet the abstract and conclusion still state the curved walls beat the physical wall. I'd want that claim softened to \"in this uncontrolled comparison\" or backed by a follow-up with matched geometry.\n\nSecond soft spot: in Study 2, the Personal condition has no No-Interaction baseline; its interaction comparisons are against Frontal No Interaction. So the claim that selection \"improved\" Personal is partly about stimulus location, not just interaction. The paper notes this but still frames it as an interaction effect.\n\nMinor: no data or code appears to be shipped, and the detailed analysis relies on supplementary figures in the appendix. Not a blocker for a venue, but reproducibility would be improved.\n\nWho's this for? Researchers and practitioners in immersive analytics who want guidance on display layouts and interaction techniques. It deserves a serious referee. The within-VR results should hold up; the cross-study claim needs revision.","headline":"Well-run VR perception studies with a credible within-VR layout result; the cross-study claim that curved virtual walls beat a physical wall is not controlled and should be softened.","tokens_in":26585,"tokens_out":3048,"would_cite":true,"duration_ms":28918,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Virtual curved walls beat flat walls for reading sizes in VR.","keywords":["virtual reality","wall-sized tiled displays","visual variables","magnitude reproduction","perception","immersive analytics","3D user interaction","display curvature"],"falsifier":"A direct replication that compares a physical curved wall display with the virtual curved displays at the same viewing distance, using identical stimulus sizes and adjustment rates, would settle the claim; if the physical curved wall does not show the same error advantage, or if the virtual advantage disappears when headset resolution is increased, the central claim is challenged.","tokens_in":25542,"feed_emoji":"🖥️","tokens_out":3776,"duration_ms":37993,"temperature":0.7,"pith_summary":"This paper investigates how accurately people can judge the magnitude of visual variables (length, angle, and area) on virtual wall-sized tiled displays in immersive VR. Two formal user studies were conducted: the first compares three virtual display arrangements (Flat, Cylinder, and Cockpit) and finds that the curved arrangements produce smaller absolute errors than the flat one, and also smaller errors than a physical flat wall display from prior work, though tasks take longer. The second study shows that adding 3D interaction techniques (Selection, Walking, Steering, and Teleportation) further reduces errors on a virtual flat display, with the Personal display plus Selection performing especially well. The work suggests that virtual curved displays could serve as a practical workspace for visual analytics, offering flexibility and better perception accuracy than real-world flat wall displays in certain conditions.","feed_headline":"Curved virtual walls beat flat for reading sizes in VR","feed_subtitle":"Two VR studies find curved displays cut estimation errors, at the cost of longer task times.","key_machinery":"The central object is the magnitude reproduction task, in which participants adjust a blue stimulus (a line segment, an angle, or a circle area) to match a red modulus target shown elsewhere on a virtual wall-sized tiled display of 32 individual tiles. The three display arrangements — Flat, Cylinder, and Cockpit — differ in curvature and orientation while sharing identical dimensions and aspect ratios; Cylinder wraps the tiles in a quarter-circle facing the participant, and Cockpit orients each tile toward the participant along both axes. The mechanism is that curved arrangements bring the modulus physically closer to the viewer and reduce acute viewing angles, which should make size comparisons easier, while the interaction techniques (Selection, Walking, Steering, Teleportation) let users reposition themselves or bring the modulus or stimulus closer to reduce depth disparity.","core_discovery":"In a magnitude reproduction task on a 32-tile virtual wall-sized display, participants made smaller estimation errors when the display was arranged in a curved configuration (Cylinder or Cockpit) than when it was flat (Flat), both in absolute error and in directional overestimation. When compared with a prior real-world flat wall study at the same 3.2 m viewing distance, the virtual curved displays also yielded smaller errors than the physical flat wall, at the cost of longer task completion times. In a second study, all four interaction techniques improved accuracy relative to a no-interaction baseline, and the Selection technique, which copies a distant display tile onto a controller-held personal display, gave the lowest absolute errors of any condition tested. These findings support the claim that virtual curved wall displays, and interactive techniques unique to VR, can make immersive environments a viable workspace for visual analytics.","pith_inferences":["The error advantage of curved virtual walls may partly compensate for the limited resolution and field of view of current VR headsets, since curvature reduces the effective angular distance to displayed content; a direct test with higher-resolution headsets could separate these factors.","The lack of significant differences among length, angle, and area in VR might stem from the different adjustment rates used in the task (e.g., 0.25 cm per frame for length and area versus 1 degree per frame for angle) rather than a genuine change in perceptual ranking; a replication with equalized adjustment rates would test this explanation.","The Selection technique, which copies a distant display tile to the controller-held personal display, could be valuable in collaborative VR scenarios, but the paper notes it risks losing the spatial context of the original display; a follow-up that highlights the source tile could mitigate this.","The finding that curved arrangements reduce error but increase time suggests a practical design guideline: use curvature for accuracy-critical data reading and flat layouts for time-efficient browsing, which could be validated in a task that measures both metrics together."],"forward_implications":["Curved virtual wall displays can be adopted in immersive analytics as a substitute for physical flat wall displays, offering lower estimation errors for elementary magnitude-reading tasks.","The longer task completion times observed for curved displays mean designers face a speed-accuracy trade-off when choosing display curvature.","Interaction techniques, especially Selection with a Personal display, can make distant comparisons more accurate and may be particularly useful when physical navigation is impossible or impractical.","Teleportation and Steering can serve as viable alternatives to physical Walking in VR, providing similar accuracy improvements without requiring real-world space, though with longer completion times.","The absence of clear accuracy differences among length, angle, and area in VR suggests that the established real-world perceptual ranking may not transfer directly, which would affect how visual encodings are chosen in immersive analytics."],"supporting_citations":[{"why":"Supplies the physical wall-sized display study whose 3.2 m results are the real-world baseline for comparison.","marker":"[10]"},{"why":"Provides the magnitude reproduction task methodology used in both VR studies.","marker":"[16]"},{"why":"Defines the three visual variables (length, angle, area) and the elementary perceptual tasks framework.","marker":"[21]"},{"why":"Defines the steering and teleportation travel techniques used in the second study.","marker":"[14]"},{"why":"Provides the display layout design space from which the Flat, Cylinder, and Cockpit arrangements are selected.","marker":"[54]"}],"fun_headline_variants":["Curved VR walls beat flat for size reading","Interactive VR cuts visual estimation errors","Curved tiled displays improve VR precision","Virtual curved walls outperform flat in VR study","VR interaction techniques boost reading accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison with the real-world flat wall display assumes that the VR setup (headset resolution, field of view, adjustment rates, and task implementation) is perceptually comparable to the physical display, so that any error difference is due to display curvature rather than to equipment or procedural differences.","fun_headline_variants_meta":{"raw":{"variants":["Curved VR walls beat flat for size reading","Interactive VR cuts visual estimation errors","Curved tiled displays improve VR precision","Virtual curved walls outperform flat in VR study","VR interaction techniques boost reading accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000161,"raw_usage":{"total_tokens":1198,"prompt_tokens":872,"completion_tokens":326,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":488,"completion_tokens_details":{"reasoning_tokens":263}},"tokens_in":488,"tokens_out":326,"duration_ms":4147,"temperature":1.0,"reasoning_tokens":263,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:10:51.172762+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct replication that compares a physical curved wall display with the virtual curved displays at the same viewing distance, using identical stimulus sizes and adjustment rates, would settle the claim; if the physical curved wall does not show the same error advantage, or if the virtual advantage disappears when headset resolution is increased, the central claim is challenged.","supporting_citations":[{"cited_title":"Buchsbaum","cited_arxiv_id":null,"evidence_quote":"Provides the magnitude reproduction task methodology used in both VR studies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the three visual variables (length, angle, area) and the elementary perceptual tasks framework."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the display layout design space from which the Flat, Cylinder, and Cockpit arrangements are selected."}],"review_version":1}