Pith. sign in

REVIEW 2 major objections 5 minor 75 references

Towards Deeper Understanding of Natural User Interactions in Virtual Reality Based Assembly Tasks

T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read When an assembly workspace offers no nameable landmarks, untrained users fall back on 'put that there' phrasing, longer pointing, and richer multimodal combinations, and the paper argues natural user interfaces should be built to expect…

desk verdict A useful new multimodal VR elicitation dataset with large behavioral differences between two tasks, but the central causal claim about spatial anchors is underdetermined by a confounded two-condition design. read the letter →

arxiv 2508.17124 v1 pith:HL4CFZSL submitted 2025-08-23 cs.HC cs.SYeess.SY

classification cs.HCcs.SYeess.SY
keywords virtualrealitynaturaluserinterfacesmultimodalinteractionWizard-of-Ozassemblytasksput-that-therespatiallanguagehuman-robot
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks how people who have received no training on the interface naturally tell a virtual robot arm what to assemble, using voice, hand gestures, gaze, and head position. In a Wizard-of-Oz study with 34 participants, the authors compared a collaborative brick assembly, where recently placed pieces give concrete spatial anchors, with an instructive printed-circuit-board assembly, where placement locations are hard to name. Their analyses indicate that users adapt their instructions to the spatial clarity of the scene: descriptive, explicit commands dominate when anchors exist, while spatially vague 'put-that-there' phrasing paired with prolonged pointing and gaze dominates when locations are ambiguous. The same data show task type shifting the modality mix, with more unimodal commands in the anchored task and more trimodal commands in the ambiguous one, and the authors release an annotated multimodal dataset. If the pattern holds, interface designers could let users choose their own instruction style instead of forcing fixed command patterns.

What carries the argument

The load-bearing apparatus is a single-factor within-subjects elicitation design whose two conditions are meant to differ in one property: whether the workspace supplies concrete, nameable spatial anchors. The analytical core is an annotation scheme built on the put-that-there command paradigm, defined as a voice-plus-gesture pattern in which an utterance names an object and a gesture or gaze supplies the location. Each utterance is classified as explicit (descriptive) or implicit (spatially vague, relying on gestures or gaze), and a multimodal sequence miner assembles each voice command with any stationary gaze or pointing gesture occurring within five seconds. This machinery lets the authors compare not only utterance content but also the duration of points, the gaze direction, the number of modalities per instruction, and how these change over task completion phases.

What would settle it

Run the same circuit-board assembly with visually anchored locations, such as outlined slots or labelled positions, while keeping the board layout and part count identical; if users still produce mostly implicit utterances and long pointing times, the claim that spatial ambiguity drives the behavior is wrong. Alternatively, within the released dataset, compare average pointing duration on implicit versus explicit utterances: if pointing is not longer for implicit commands, the implied linkage between vague language and prolonged deictic compensation fails.

Watch

Extended reading notes

Core claim

The paper's central discovery is that untrained users spontaneously follow a put-that-there instruction pattern, an utterance that names an object plus a deictic indication of where it goes, but the linguistic and gestural balance of that pattern shifts with the availability of spatial anchors. In the brick task, where pieces provide nameable reference points, participants used significantly more explicit descriptive commands, pointed for shorter periods, looked more toward the reference model, and leaned on unimodal speech. In the circuit-board task, where the blank board offers no obvious landmarks, participants used significantly more implicit commands such as 'put a resistor here' and compensated with significantly longer pointing, more gaze use, and more trimodal combinations; instruction frequency also predicted completion time far more strongly in this task. These differences were not static: command style evolved across task phases, with brick-task users gradually adding implicit commands while circuit-board users stayed implicit throughout.

Load-bearing premise

The study assumes the two assembly tasks differ only in the factor being tested, whether the scene provides concrete spatial anchors, so that all observed differences in user behavior can be attributed to that factor, even though the tasks also differ in piece count, part types, board geometry, and interaction demands.

Editorial extensions

If this is right

  • Natural multimodal interfaces for assembly should accept put-that-there style instructions that pair vague voice with gesture and gaze instead of requiring users to follow preset command grammars.
  • In workspaces with concrete anchors, designers can expect explicit descriptive speech to suffice and can keep gesture interpretation lightweight, while in anchor-free workspaces the interface must tolerate long pointing episodes and combine eye and head gaze with speech.
  • Task phase matters: instruction style changes as an anchored assembly progresses, so session-level assumptions about modality use should give way to phase-adaptive interpretation.
  • The released annotated dataset, containing voice transcripts, hand tracking, eye gaze, head pose, and command labels, gives other researchers a basis for training models of natural multimodal assembly instruction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test that varies only anchor availability, keeping the same board layout and part count while adding or removing visual placeholders, would separate the spatial-anchor explanation from other task differences such as piece count and board geometry; the current study conflates those factors.
  • Because instruction frequency predicted completion time sharply in the ambiguous task, an interface that auto-completes repeated component types could cut instruction overhead more in layout-heavy tasks than in anchored ones.
  • Modality mix could serve as a real-time uncertainty signal: rising implicit utterances combined with longer pointing would flag regions the user finds hard to describe, letting an adaptive system highlight candidate placements.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper presents a Wizard-of-Oz elicitation study (N=34) in which participants instruct a virtual robot arm through a collaborative LEGO assembly task and an instructive PCB assembly task in VR. Voice, hand tracking, eye gaze, and head pose are captured and analyzed. The main empirical claims are that explicit/descriptive utterances dominate the LEGO task while implicit "put-that-there" language dominates the PCB task, that pointing durations are longer in the PCB task, and that users produce more trimodal commands in the PCB task and more unimodal commands in the LEGO task. The paper also reports correlations between utterance frequency and task completion time and contributes a publicly available annotated dataset.

Significance. If the descriptive findings are reliable, the dataset is a useful resource for future natural multimodal interface design, and the large effects on utterance ratio (t=8.14), pointing time, and trimodal usage are noteworthy. The study is unusual in capturing untrained users' raw multimodal input with a Wizard-of-Oz robot, and the public dataset is a concrete contribution. However, the paper's headline causal interpretation—that spatial-anchor availability causes the observed differences—is not supported by the operationalization, and the central utterance-ratio measure lacks demonstrated coding reliability. The work would be publishable as a descriptive, dataset-focused study once these issues are addressed.

major comments (2)
  1. [§3.1, §6.2, abstract] The study is framed as a single-factor comparison of "task nature," but the two conditions differ on many dimensions at once: the participant manually manipulates objects only in Task 1; the object sets, geometries, and board layouts differ; the piece counts are 25 versus 20 (§3.6); and the observed completion times differ as well (M=382.84 s versus 451.40 s, Table 3). The abstract and §6.2/§6.3 nevertheless attribute the differences to the availability of concrete spatial anchors ("due to concrete spatial anchors... due to a lack of spatial anchors"). Because no condition manipulates anchor availability while holding the other task properties fixed, the causal attribution in the abstract is underdetermined. The authors should either soften the claims to descriptive differences between two assembly scenarios or add a follow-up condition that isolates anchor availability; the current Limitations section does not acknowledge this confound.
  2. [§4.1, §5.1.1] The headline effect—explicit versus implicit utterance ratios differing across tasks with t33 = 8.14—rests entirely on manual annotation of transcripts into nine labels, but the paper provides no information about the number of annotators, the annotation protocol, or inter-rater agreement. Without coding reliability evidence, the large ratio difference could reflect a single annotator's interpretation of the Bolt-based scheme rather than a stable behavioral difference. The authors should report inter-rater reliability (e.g., Cohen's kappa or equivalent) or provide a robustness check such as re-annotation of a subset.
minor comments (5)
  1. [Tables 1–4 and §5.4.3] Statistical reporting needs alignment: §5.4.3 reports F3,30=5.296, R=0.530, R²=0.281 for the Task 1 multimodal correlation, whereas Table 2 reports F3,30=5.438, R=0.535, R²=0.287; please use consistent values and correct notation, since Wilcoxon results are reported with t statistics in places (e.g., §5.2.1 reports "t33=37.0").
  2. [§3.3] The paper states that participants "underwent training" for task objectives and manipulation mechanics, which appears to conflict with the repeated claim that interactions come from untrained users; please clarify what was trained and why this does not undermine the "untrained" framing.
  3. [§4.2] The pointing-detection thresholds (z-rotation differences of 10° and 30°, minimum duration 0.5 s) are stated without justification; please cite prior work or report a sensitivity analysis showing the findings are robust to these choices.
  4. [Figure 10] The task-sequence visualization is dense and the legend ("G – Gaze, P – Pointing Gesture, U – Utterance") does not explain how sequences are encoded; the figure would benefit from a concrete example sequence annotation.
  5. [§6.4] Minor typographical issues (e.g., "V oice ismost effective" in §6.4.1 and "cuessuch" in §6.4.2) should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an observational elicitation study whose central results are descriptive statistical comparisons grounded in collected data, not derived from fitted parameters or self-citation chains.

full rationale

This paper reports a Wizard-of-Oz elicitation study in which participants completed two VR assembly tasks while voice, hand tracking, and gaze were recorded. The main findings are empirical: explicit utterances were more frequent in the LEGO task, implicit utterances and trimodal commands were more frequent in the PCB task, and average pointing time was longer in the PCB task. These results come from direct statistical comparisons (paired t-tests, Wilcoxon signed-rank tests, correlations) of the collected data, not from a derivation that assumes its own conclusion. The explicit/implicit annotation scheme follows Bolt's external 'put-that-there' work and labels utterances by their linguistic content, not by the task condition, so the measured difference is an empirical observation rather than a definitional artifact. The authors cite their own prior work (e.g., [24], [25]) only for background in related work; those citations are not load-bearing for the study's conclusions. The paper does not fit parameters to a subset and then 'predict' a related quantity, nor does it invoke a self-citation chain to force its interpretation. The confound between task nature and spatial anchor availability is a validity and generalizability concern, appropriately categorized as correctness risk rather than circularity, because the paper's causal language is an interpretive claim about a two-condition design, not a mathematical reduction. Under the hard rules requiring a quotable reduction of Eq. X to Eq. Y or a fitted parameter renamed as prediction, no circular step can be exhibited. The appropriate finding is therefore no significant circularity.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The empirical claims do not rest on mathematical axioms, but they assume the measurement pipeline (transcription, manual annotation, feature thresholds) and the equivalence between Wizard-of-Oz behavior and an autonomous interface. The hand-chosen thresholds for pointing and sequence windows are free parameters because no sensitivity analysis is provided.

free parameters (3)
  • Pointing pose detection thresholds
    Hand-chosen thresholds in Section 4.2: index finger z-rotation difference below 10 degrees, middle finger difference above 30 degrees, and a minimum pointing duration of 0.5 seconds. No sensitivity analysis is provided, and changing these thresholds would change point counts and average point times.
  • Sequence context window = 5 seconds
    Section 5.4.1 defines a multimodal sequence by including stationary gaze and/or pointing gestures within five seconds of an utterance start. This window is hand-chosen and no alternative windows are tested.
  • Task phase boundaries = 0-20%, 20-80%, 80-100%
    Section 5.1.3 groups utterances into three phases using these percentage thresholds of task completion. The boundaries are arbitrary and affect the reported phase-wise differences.
assumptions (4)
  • domain assumption Whisper-Large V2 transcription is accurate enough for command annotation
    Section 4.1 uses Whisper-Large V2 to transcribe voice into text without reporting transcription accuracy or manual correction, yet the later utterance labeling depends on these transcripts.
  • domain assumption Manual annotation labels are reliable despite a single annotator and no inter-rater agreement
    Section 4.1 describes nine annotation labels but provides no inter-rater reliability metric, no annotation guide, and no report of multiple annotators. The central explicit/implicit distinction rests on this unverified reliability.
  • domain assumption Wizard-of-Oz behavior approximates an autonomous natural interface
    Section 3.2 and the procedure in Section 3.6 rely on an experimenter controlling the robot in the background, interpreting user instructions. The study assumes this setup elicits natural behavior equivalent to what an autonomous system would receive.
  • domain assumption Higher-level interaction behaviors transfer from VR to real-world assembly
    Sections 3.3 and 6.5 assert that language use for spatial references and modality coordination transfer to physical scenarios, while acknowledging that fine motor control and error recovery may not transfer. This transferability is assumed rather than demonstrated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Deeper Understanding of Natural User Interactions in Virtual Reality Based Assembly Tasks." pith.science (2026). https://pith.science/paper/HL4CFZSL

@misc{pith2026250817124,
  author       = {Pith},
  title        = {Pith review of: Towards Deeper Understanding of Natural User Interactions in Virtual Reality Based Assembly Tasks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HL4CFZSL}},
  note         = {Machine review of arXiv:2508.17124}
}
read the original abstract

We explore natural user interactions using a virtual reality simulation of a robot arm for assembly tasks. Using a Wizard-of-Oz study, participants completed collaborative LEGO and instructive PCB assembly tasks, with the robot responding under experimenter control. We collected voice, hand tracking, and gaze data from users. Statistical analyses revealed that instructive and collaborative scenarios elicit distinct behaviors and adopted strategies, particularly as tasks progress. Users tended to use put-that-there language in spatially ambiguous contexts and more descriptive instructions in spatially clear ones. Our contributions include the identification of natural interaction strategies through analyses of collected data, as well as the supporting dataset, to guide the understanding and design of natural multimodal user interfaces for instructive interaction with systems in virtual reality.

Figures

Figures reproduced from arXiv: 2508.17124 by the authors.

Figure 1
Figure 1. Task setups for LEGO and PCB tasks during completion. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Flow diagram of data collection, processing, and storage [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Task space for Task 1 (top) and Task 2 (bottom). [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Ratios of Explicit and Implicit Utterances. More Explicit [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Stacked bar graphs show the distribution of utterance labels across the two tasks for each participant. [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Distributions of Explicit and Implicit Utterances across Task [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 9
Figure 9. Figure 9: Average Look Durations for Task 1 and Task 2. Participants [PITH_FULL_IMAGE:figures/full_fig_p007_9.png]
Figure 10
Figure 10. Figure 10: Stacked Bar Graph showing the Task Sequence Distribution across Task 1 (left) and Task 2 (right) for all participants (G – Gaze, P – [PITH_FULL_IMAGE:figures/full_fig_p008_10.png]
Figure 11
Figure 11. Figure 11: Unimodal, Bimodal and Trimodal Command usages per [PITH_FULL_IMAGE:figures/full_fig_p008_11.png]
Figure 12
Figure 12. Figure 12: Task space for Task 1 (top) and Task 2 (bottom). [PITH_FULL_IMAGE:figures/full_fig_p009_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

75 extracted references · 67 canonical work pages

  1. [1]

    Ahmad, S

    A. Ahmad, S. Darmoul, W. Ameen, M. Abidi, and A. Al-Ahmari. Rapid prototyping for assembly training and validation. vol. 48, 05

  2. [2]

    A. J. Aubrey, D. Marshall, P. L. Rosin, J. Vendeventer, D. W. Cun- ningham, and C. Wallraven. Cardiff conversation database (ccdb): A database of natural dyadic conversations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp. 277–282, 2013. 2

  3. [3]

    Azagra, Y

    P. Azagra, Y . Mollard, F. Golemo, A. C. Murillo, M. Lopes, and J. Civera. A multimodal human-robot interaction dataset. In NIPS 2016, workshop future of interactive learning machines, 2016. 2

  4. [4]

    Multimodal Dataset of Human-Robot Hugging Interaction

    K. Bagewadi, J. Campbell, and H. B. Amor. Multimodal dataset of human-robot hugging interaction. arXiv preprint arXiv:1909.07471,

  5. [5]

    Ben-Youssef, C

    A. Ben-Youssef, C. Clavel, S. Essid, M. Bilac, M. Chamoux, and A. Lim. Ue-hri: a new dataset for the study of user engagement in spontaneous human-robot interactions. In Proceedings of the 19th ACM international conference on multimodal interaction , pp. 464– 472, 2017. 2

  6. [6]

    Bilakhia, S

    S. Bilakhia, S. Petridis, A. Nijholt, and M. Pantic. The mahnob mimicry database: A database of naturalistic human interactions. Pat- tern recognition letters, 66:52–61, 2015. 2

  7. [7]

    put-that-there

    R. A. Bolt. “put-that-there” voice and gesture at the graphics interface. In Proceedings of the 7th annual conference on Computer graphics and interactive techniques, pp. 262–270, 1980. 1, 3, 4, 8

  8. [8]

    Borghi, F

    S. Borghi, F. Zucchi, E. Prati, A. Ruo, V . Villani, L. Sabattini, and M. Peruzzini. Unlocking human-robot dynamics: Introducing sensec- obot, a novel multimodal dataset on industry 4.0. In Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot In- teraction, HRI ’24, p. 880–884. Association for Computing Machin- ery, New York, NY , USA,...

Show all 75 references
  1. [9]

    Camburn, V

    B. Camburn, V . Viswanathan, J. Linsey, D. Anderson, D. Jensen, R. Crawford, K. Otto, and K. Wood. Design prototyping methods: state of the art in strategies, techniques, and guidelines. Design Sci- ence, 3:e13, 2017. 9

  2. [10]

    Can ´evet, W

    O. Can ´evet, W. He, P. Motlicek, and J.-M. Odobez. The mummer data set for robot perception in multi-party hri scenarios. In 2020 29th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), pp. 1294–1300, 2020. doi: 10.1109/RO -MAN47096.2020.9223340 2

  3. [11]

    Carlson, A

    P. Carlson, A. Peters, S. B. Gilbert, J. M. Vance, and A. Luse. Virtual training: Learning transfer of assembly tasks. IEEE transactions on visualization and computer graphics, 21(6):770–782, 2015. 3

  4. [12]

    Celiktutan, E

    O. Celiktutan, E. Skordos, and H. Gunes. Multimodal human-human- robot interactions (mhhri) dataset for studying personality and engage- ment. IEEE Transactions on Affective Computing , 10(4):484–497,

  5. [13]

    Chang, X

    Y . Chang, X. Wang, J. Wang, Y . Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y . Wang, et al. A survey on evaluation of large language models. ACM transactions on intelligent systems and technology , 15(3):1–45, 2024. 10

  6. [14]

    Chu and Y .-L

    C.-H. Chu and Y .-L. Liu. Augmented reality user interface design and experimental evaluation for human-robot collaborative assembly. Journal of Manufacturing Systems, 68:313–324, 2023. 1

  7. [15]

    Cohen, C

    P. Cohen, C. Swindells, S. Oviatt, and A. Arthur. A high-performance dual-wizard infrastructure for designing speech, pen, and multimodal interfaces. In Proceedings of the 10th international conference on Multimodal interfaces, pp. 137–140, 2008. 2

  8. [16]

    P. R. Cohen, M. Johnston, D. McGee, S. Oviatt, J. Pittman, I. Smith, L. Chen, and J. Clow. Quickset: Multimodal interaction for distributed applications. In Proceedings of the fifth ACM international conference on Multimedia, pp. 31–40, 1997. 2

  9. [17]

    L. M. Daling and S. J. Schlittmeier. Effects of augmented reality-, virtual reality-, and mixed reality–based training on objective perfor- mance measures and subjective evaluations in manual assembly tasks: a scoping review. Human factors, 66(2):589–626, 2024. 3

  10. [18]

    Devillers, S

    L. Devillers, S. Rosset, G. D. Duplessis, M. A. Sehili, L. B ´echade, A. Delaborde, C. Gossart, V . Letard, F. Yang, Y . Yemez, et al. Mul- timodal data collection of human-robot humorous interactions in the joker project. In 2015 international conference on affective computin...

  11. [19]

    Dong and H

    G. Dong and H. Liu. Feature engineering for machine learning and data analytics. CRC press, 2018. 3

  12. [20]

    J. Epps, S. Oviatt, and F. Chen. Integration of speech and gesture inputs during multimodal interaction. In Proc Aust. Int. Conf. on CHI,

  13. [21]

    Fechter, B

    M. Fechter, B. Schleich, and S. Wartzack. Comparative evaluation of wimp and immersive natural finger interaction: A user study on cad assembly modeling. Virtual Reality, 26(1):143–158, 2022. 1

  14. [22]

    Fiorentino, R

    M. Fiorentino, R. Radkowski, C. Stritzke, A. E. Uva, and G. Monno. Design review of cad assemblies using bimanual natural interface. International Journal on Interactive Design and Manufacturing (IJI- DeM), 7:249–260, 2013. 1

  15. [23]

    Geiger, E

    A. Geiger, E. Brandenburg, and R. Stark. Natural virtual reality user interface to define assembly sequences for digital human models. Ap- plied System Innovation, 3(1):15, 2020. 1

  16. [24]

    R. K. Ghamandi, Y . Hmaiti, T. T. Nguyen, A. Ghasemaghaei, R. K. Kattoju, E. M. Taranta, and J. J. LaViola. What and how together: a taxonomy on 30 years of collaborative human-centered xr tasks. In 2023 IEEE International Symposium on Mixed and Augmented Real- ity (ISMAR), pp...

  17. [25]

    R. K. Ghamandi, R. K. Kattoju, Y . Hmaiti, M. Maslych, E. M. Taranta, R. P. McMahan, and J. LaViola. Unlocking understanding: An inves- tigation of multimodal communication in virtual reality collaboration. In Proceedings of the 2024 CHI Conference on Human Factors in Computin...

  18. [26]

    Gottsacker, Y

    M. Gottsacker, Y . Hmaiti, M. Maslych, G. Bruder, J. J. LaViola Jr, and G. F. Welch. Xr-first design for productivity: A conceptual 10 © 2025 IEEE. This is the author’s version of the article that has been published in the proceedings of IEEE Visualization conference. The fina...

  19. [27]

    Hmaiti, M

    Y . Hmaiti, M. Maslych, A. Ghasemaghaei, R. K. Ghamandi, and J. J. LaViola Jr. Visual perceptual confidence: Exploring discrepancies between self-reported and actual distance perception in virtual reality. IEEE Transactions on Visualization and Computer Graphics, 2024. 2

  20. [28]

    Hmaiti, M

    Y . Hmaiti, M. Maslych, E. M. Taranta, and J. J. LaViola. An explo- ration of the effects of head-centric rest frames on egocentric distance judgments in vr. In 2023 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pp. 263–272. IEEE, 2023. 2

  21. [29]

    S. Howard. User interface design and hci: identifying the training needs of practitioners. SIGCHI Bull., 27(3):17–22, jul 1995. doi: 10. 1145/221296.221302 1

  22. [30]

    Husainy, A

    A. Husainy, A. Joshi, V . Chougule, R. Thomake, H. Kamat, and H. Jadhav. The impact of virtual reality integration in robotics: En- hancing efficiency, safety, and human-robot interaction.Asian Review of Mechanical Engineering, 12:28–34, 12 2023. doi: 10.70112/arme -2023.12.2.4228 3

  23. [31]

    Iftikhar, M

    M. Iftikhar, M. Saqib, M. Zareen, and H. Mumtaz. Artificial in- telligence: revolutionizing robotic surgery. Annals of Medicine and Surgery, 86(9):5401–5409, 2024. 1, 9

  24. [32]

    P. G. Ikonomov and E. D. Milkova. Virtual assembly/disassembly system using natural human interaction and control. Virtual and aug- mented reality applications in manufacturing, pp. 111–125, 2004. 1

  25. [33]

    Inamura and Y

    T. Inamura and Y . Mizuchi. Sigverse: A cloud-based vr platform for research on multimodal human-robot interaction. Frontiers in Robotics and AI, 8:549360, 2021. 2

  26. [34]

    D. B. Jayagopi, S. Sheiki, D. Klotz, J. Wienke, J.-M. Odobez, S. Wrede, V . Khalidov, L. Nyugen, B. Wrede, and D. Gatica-Perez. The vernissage corpus: A conversational human-robot-interaction dataset. In 2013 8th ACM/IEEE International Conference on Human- Robot Interaction (H...

  27. [35]

    Jiang, A

    Y . Jiang, A. Gupta, Z. Zhang, G. Wang, Y . Dou, Y . Chen, L. Fei-Fei, A. Anandkumar, Y . Zhu, and L. Fan. Vima: General robot manip- ulation with multimodal prompts. arXiv preprint arXiv:2210.03094, 2(3):6, 2022. 2

  28. [36]

    Johnston, P

    M. Johnston, P. R. Cohen, D. McGee, S. Oviatt, J. A. Pittman, and I. Smith. Unification-based multimodal integration. In 35th Annual Meeting of the Association for Computational Linguistics and 8th Conference of the European Chapter of the Association for Compu- tational Lingu...

  29. [37]

    J. S. Joyner, M. Vaughn-Cooke, and H. L. Benz. Comparison of dex- terous task performance in virtual reality and real-world environments. Frontiers in Virtual Reality, 2:599274, 2021. 1

  30. [38]

    Kesim, T

    E. Kesim, T. Numanoglu, O. Bayramoglu, B. B. Turker, N. Hussain, M. Sezgin, Y . Yemez, and E. Erzin. The ehri database: a multimodal database of engagement in human–robot interactions. Language Re- sources and Evaluation, 57(3):985–1009, 2023. 2

  31. [39]

    J. J. LaViola Jr, S. Buchanan, and C. Pittman. Multimodal input for perceptual user interfaces. Interactive Displays: Natural Human- Interface Technologies, pp. 285–312, 2014. 1

  32. [40]

    J. J. LaViola Jr, E. Kruijff, R. P. McMahan, D. Bowman, and I. P. Poupyrev. 3D user interfaces: theory and practice . Addison-Wesley Professional, 2017. 3

  33. [41]

    grip-that- there

    K. Mahadevan, M. Sousa, A. Tang, and T. Grossman. “grip-that- there”: An investigation of explicit and implicit task allocation tech- niques for human-robot collaboration. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems, pp. 1–14, 2021. 2

  34. [42]

    A. D. Marshall, P. L. Rosin, J. Vandeventer, and A. Aubrey. 4d cardiff conversation database (4d ccdb): A 4d database of natural, dyadic conversations. Auditory-Visual Speech Processing,{AVSP} 2015, pp. 157–162, 2015. 2

  35. [43]

    Martin, S

    D. Martin, S. Malpica, D. Gutierrez, B. Masia, and A. Serrano. Multimodality in vr: A survey. ACM Computing Surveys (CSUR) , 54(10s):1–36, 2022. 2

  36. [44]

    Medell ´ın-Castillo, J

    H. Medell ´ın-Castillo, J. Corney, J. Ritchie, R. Sung, and T. Lim. Virtual assembly rapid prototyping of near net shapes. Ingenier´ıa Mec´anica. Tecnolog´ıa y Desarrollo, 3:66–76, 03 2009. doi: 10.1115/ WINVR2009-723 1

  37. [45]

    B. A. Newman, R. M. Aronson, S. S. Srinivasa, K. Kitani, and H. Ad- moni. Harmonic: A multimodal dataset of assistive human–robot col- laboration. The International Journal of Robotics Research, 41(1):3– 11, 2022. 2

  38. [46]

    S. Oviatt. Mulitmodal interactive maps: Designing for human perfor- mance. Human–Computer Interaction, 12(1-2):93–129, 1997. 3

  39. [47]

    S. Oviatt. Ten myths of multimodal interaction. Commun. ACM , 42(11):74–81, nov 1999. doi: 10.1145/319382.319398 2, 7, 8

  40. [48]

    S. Oviatt. Advances in robust multimodal interface design. IEEE computer graphics and applications, 23(05):62–68, 2003. 2

  41. [49]

    Oviatt and P

    S. Oviatt and P. Cohen. Perceptual user interfaces: multimodal inter- faces that process what comes naturally.Communications of the ACM, 43(3):45–53, 2000. 2

  42. [50]

    Oviatt, P

    S. Oviatt, P. Cohen, L. Wu, L. Duncan, B. Suhm, J. Bers, T. Holzman, T. Winograd, J. Landay, J. Larson, et al. Designing the user inter- face for multimodal speech and pen-based gesture applications: State- of-the-art systems and future research directions. Human-computer inte...

  43. [51]

    Oviatt, R

    S. Oviatt, R. Coulston, and R. Lunsford. When do we interact multi- modally? cognitive load and multimodal communication patterns. In Proceedings of the 6th International Conference on Multimodal Inter- faces, ICMI ’04, p. 129–136. Association for Computing Machinery, New York...

  44. [52]

    Oviatt, A

    S. Oviatt, A. DeAngeli, and K. Kuhn. Integration and synchronization of input modes during multimodal human-computer interaction. In Proceedings of the ACM SIGCHI Conference on Human factors in computing systems, pp. 415–422, 1997. 7, 8

  45. [53]

    Petersen and D

    N. Petersen and D. Stricker. Continuous natural user interface: Re- ducing the gap between real and digital world. In ISMAR, vol. 1, pp. 23–26, 2009. 1

  46. [54]

    Radford, J

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever. Robust speech recognition via large-scale weak supervi- sion. In International Conference on Machine Learning , pp. 28492– 28518. PMLR, 2023. 4

  47. [55]

    Reddy, P

    K. Reddy, P. Gharde, H. Tayade, M. Patil, L. S. Reddy, D. Surya, and L. srivani Reddy. Advancements in robotic surgery: a comprehen- sive overview of current utilizations and upcoming frontiers. Cureus, 15(12), 2023. 1, 9

  48. [56]

    Ringeval, A

    F. Ringeval, A. Sonderegger, J. Sauer, and D. Lalanne. Introducing the recola multimodal corpus of remote collaborative and affective inter- actions. In 2013 10th IEEE international conference and workshops on automatic face and gesture recognition (FG), pp. 1–8. IEEE, 2013. 2

  49. [57]

    F. D. Rose, E. A. Attree, B. M. Brooks, D. M. Parslow, and P. R. Penn. Training in virtual environments: transfer to real world tasks and equivalence to real task training. Ergonomics, 43(4):494–511,

  50. [58]

    Rozgic, B

    V . Rozgic, B. Xiao, A. Katsamanis, B. R. Baucom, P. G. Georgiou, and S. S. Narayanan. A new multichannel multi modal dyadic interaction database. In INTERSPEECH, pp. 1982–1985, 2010. 2

  51. [59]

    Sahoo and C.-Y

    S. Sahoo and C.-Y . Lo. Smart manufacturing powered by recent tech- nological advancements: A review. Journal of Manufacturing Sys- tems, 64:236–250, 2022. 9

  52. [60]

    Sanchez-Cortes, O

    D. Sanchez-Cortes, O. Aran, and D. Gatica-Perez. An audio visual corpus for emergent leader analysis. In Workshop on multimodal cor- pora for machine learning: taking stock and road mapping the future, ICMI-MLMI. Citeseer, 2011. 2

  53. [61]

    Selfridge and W

    M. Selfridge and W. Vannoy. A natural language interface to a robot assembly system. IEEE Journal on Robotics and Automation , 2(3):167–171, 1986. 1

  54. [62]

    B. S. Shafique, A. Vayani, M. Maaz, H. A. Rasheed, D. Dissanayake, M. I. Kurpath, Y . Hmaiti, G. Inoue, J. Lahoud, M. S. Rashid, et al. A culturally-diverse multilingual multimodal video benchmark & model. arXiv preprint arXiv:2506.07032, 2025. 10

  55. [63]

    Shrestha, Y

    S. Shrestha, Y . Zha, S. Banagiri, G. Gao, Y . Aloimonos, and C. Fer- muller. Natsgd: A dataset with speech, gestures, and demonstrations for robot learning in natural human-robot interaction. arXiv preprint arXiv:2403.02274, 2024. 2, 3

  56. [64]

    Siltanen, M

    S. Siltanen, M. Hakkarainen, O. Korkalo, T. Salonen, J. Saaski, C. Woodward, T. Kannetis, M. Perakakis, and A. Potamianos. Mul- 11 © 2025 IEEE. This is the author’s version of the article that has been published in the proceedings of IEEE Visualization conference. The final ve...

  57. [65]

    H. Su, W. Qi, J. Chen, C. Yang, J. Sandoval, and M. A. Laribi. Recent advancements in multimodal human–robot interaction. Frontiers in Neurorobotics, 17:1084000, 2023. 2

  58. [66]

    N. T. V . Tuyen, A. L. Georgescu, I. Di Giulio, and O. Celiktutan. A multimodal dataset for robot learning to imitate social human-human interaction. In Companion of the 2023 ACM/IEEE International Con- ference on Human-Robot Interaction, HRI ’23, p. 238–242. Associa- tion for...

  59. [67]

    G. C. van der Veer, T. , Michael J., W. , Yvonne, , and B. v. Muylwijk. On the interaction between system and user characteristics.Behaviour & Information Technology, 4(4):289–308, Oct. 1985. Publisher: Tay- lor & Francis eprint: https://doi.org/10.1080/01449298508901809. doi:...

  60. [68]

    Vayani, D

    A. Vayani, D. Dissanayake, H. Watawana, N. Ahsan, N. Sasikumar, O. Thawakar, H. B. Ademtew, Y . Hmaiti, A. Kumar, K. Kukreja, et al. All languages matter: Evaluating lmms on culturally diverse 100 lan- guages. In Proceedings of the Computer Vision and Pattern Recogni- tion Con...

  61. [69]

    X. Wang, S. Ong, and A. Y .-C. Nee. Multi-modal augmented-reality assembly guidance based on bare-hand interface.Advanced Engineer- ing Informatics, 30(3):406–421, 2016. 1

  62. [70]

    J. O. Wobbrock, M. R. Morris, and A. D. Wilson. User-defined gestures for surface computing. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems , CHI ’09, p. 1083–1092. Association for Computing Machinery, New York, NY , USA, 2009. doi: 10.1145/15187...

  63. [71]

    C. Yu. Robot Behavior Generation and Human Behavior Understand- ing in Natural Human-Robot Interaction . Theses, Institut Polytech- nique de Paris, June 2021. 8

  64. [72]

    Zachmann and A

    G. Zachmann and A. Rettig. Natural and robust interaction in virtual assembly simulation. In Eighth ISPE International Conference on Concurrent Engineering: Research and Applications (ISPE/CE2001), vol. 1, pp. 425–434. Citeseer, 2001. 1

  65. [73]

    Zhang, Q

    Q. Zhang, Q. Liu, J. Duan, and J. Qin. Research on teleoperated virtual reality human–robot five-dimensional collaboration system. Biomimetics, 8(8):605, 2023. 3 12

  66. [2015]

    doi: 10.1016/j.ifacol.2015.06.116 1

  67. [2019]

    doi: 10.1109/TAFFC.2017.2737019 2

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.