Pith. sign in

Paper Citation Record · LEDGER

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

As of 15 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 6 inbound Pith citation observations for arXiv:2503.16492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.16492 v3

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T01:10:55.850274Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:46:48.174002Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T04:57:38.604855Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy43
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a90f630d-ffd0-4942-a234-0f17297e9db1 · outbound

This paper cites Social robots in therapy and care.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Social robots in therapy and care

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.498346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:e18e7b4852f5f68afee66edbdf8117c58376cb62910404b521519fef0fb1a5b2

Observation 2c5c96ee-84a7-45b9-b6a9-82fa5872d50b · outbound

This paper cites Jubileo: An open- source robot and framework for research in human-robot social interac- tion.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Jubileo: An open- source robot and framework for research in human-robot social interac- tion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.128820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:ed3e0bee488497c873d5bcc20f1dacf502cdb699a6cfa2a625231f4fc575216b

Observation f34bb0eb-1fa4-4202-b70c-34b47f84741b · outbound

This paper cites Human-aware physical human–robot collaborative transportation and manipulation with multiple aerial robots.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Human-aware physical human–robot collaborative transportation and manipulation with multiple aerial robots

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.119264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:85fd417d3b4f8aec8492a7cc1440c91e5b8db3847ae0de303e8ab11d564f8bca

Observation d98a3c25-71e5-4f90-b9c5-728f256fc094 · outbound

This paper cites Communicating human intent to a robotic companion by multi-type gesture sentences.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Communicating human intent to a robotic companion by multi-type gesture sentences

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.134707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:7b12b711ee44b28faf51433ff50a9edec1f012e817352dfc67190e74b6d20842

Observation 5b6af122-e4a8-40b3-a9be-a587238033ad · outbound

This paper cites Nvp-hri: Zero shot natural voice and posture-based human–robot interaction via large language model.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Nvp-hri: Zero shot natural voice and posture-based human–robot interaction via large language model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.122045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:55b636c038190d246e1473d3440b76e592211fc1bcc84befe7c90a82b7f460a9

Observation e82bc16b-485c-4eec-99fc-960afc86de6d · outbound

This paper cites Robot reading human gaze: Why eye tracking is better than head tracking for human-robot col- laboration.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Robot reading human gaze: Why eye tracking is better than head tracking for human-robot col- laboration

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.125007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:36c9e53024b295f180eaa156ae10dde0dc61c4795a6d966624017cfefa4ca47e

Observation 8b016c46-0c14-406b-9e6a-28c479563a7d · outbound

This paper cites A gaze-speech system in mixed reality for human-robot interaction.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech A gaze-speech system in mixed reality for human-robot interaction

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.115991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:db9f46fccf9244d17a29ee18a68531f6b9f15b7a0984f016393beb7ca42717e0

Observation de8dca7e-1ef4-48da-8c8b-9c8aebb24d00 · outbound

This paper cites Human–robot interaction through eye tracking for artistic drawing.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Human–robot interaction through eye tracking for artistic drawing

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.110198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:264e7552deb621e42ad3cec2ef3fce272620567d26e29eda751db3d6d9ddfcd4

Observation 5c486e0b-366d-4f44-8c52-95c382ebf73b · outbound

This paper cites Is it possible to recognize a speaker without listening? unraveling conversation dynamics in multi-party interactions using continuous eye gaze.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Is it possible to recognize a speaker without listening? unraveling conversation dynamics in multi-party interactions using continuous eye gaze

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.501692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:69ce581be6db627ee1a4fd87b9f094724cc310b843146b0a4623ff2751bfd486

Observation 8466c8fb-5ca1-4abc-a755-aedd77d4e226 · outbound

This paper cites Eye-gaze control of a wheelchair mounted 6dof assistive robot for activities of daily living.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Eye-gaze control of a wheelchair mounted 6dof assistive robot for activities of daily living

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.137422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:ccc3971faaae0e98f6e46f2ffba42fc041e8d47924de199423c1489838275178

Observation 935e96a0-07e1-43b4-abba-59deb811e27b · outbound

This paper cites Free-view, 3d gaze-guided, assistive robotic system for activities of daily living.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Free-view, 3d gaze-guided, assistive robotic system for activities of daily living

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.132103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:d9f7b6488601550db3fb9235aad47150fcbf92baf595d9ac187b92012c22f01d

Observation 2ec08df2-47da-4ebf-b7ba-6311edd74354 · outbound

This paper cites Microsaccade-inspired event camera for robotics.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Microsaccade-inspired event camera for robotics

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.113054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:4b6ceecc7276a0d0336288dbf7341135c462bdb76b98ebd6ad31a8c2ed0f320b

Observation 84a328ec-92e3-407c-bad0-6116ac31cee4 · outbound

This paper cites Project Aria: A New Tool for Egocentric Multi-Modal AI Research.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:12:20.528544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:fe1c80fd285049aecca25363a9c24be1d2f1850520a0f9cc2c031329b51c979d

Observation 9599ee3f-bdf7-433d-b998-3f8e424e18fd · outbound

This paper cites Robust gaze- based intention prediction for real-world scenarios.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Robust gaze- based intention prediction for real-world scenarios

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.620189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:16cbc46cece91f449cf4a56ca1b17668ed70099f2bb812d3bde406ac5f60b6c7

Observation ea247167-bc67-49e9-afed-9df33b676395 · outbound

This paper cites Getting to know your robot customers: Automated analysis of user identity and demographics for robots in the wild.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Getting to know your robot customers: Automated analysis of user identity and demographics for robots in the wild

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.617076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:c1928441f18949633334a744ce1c090c4a6d0757b4ea54fc832402ccbb3a1b79

Observation 910f1d15-3185-4d1a-b07e-28c592d36a60 · outbound

This paper cites A personalized comfort space with variable shape based on environmental information for robot navigation in homes.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech A personalized comfort space with variable shape based on environmental information for robot navigation in homes

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.613531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:80faed61da239213a166e9e34e5880f7ef48c26933c4aeb6405ed32c74f007d5

Observation afe87ad1-6531-44e1-a26d-41153c03a66c · outbound

This paper cites Improving the collision tolerance of high-speed industrial robots via impact-aware path planning and series clutched actuation.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Improving the collision tolerance of high-speed industrial robots via impact-aware path planning and series clutched actuation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.610101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:f52f82abdcb160ae28c2b1d994d5c5285c73ce522d6b87a69c27dff845ecba06

Observation a57f45cf-4fa9-494c-8990-9392a22f9248 · outbound

This paper cites In situ calibration of six- axis force–torque sensors for industrial robots with tilting base.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech In situ calibration of six- axis force–torque sensors for industrial robots with tilting base

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.606522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:61e99eac2bd0e3e419f12f6a6c7a5df1f7d1a4d091dce7aea51c3db96c209a17

Observation e243694a-600a-4025-af69-5088701a14c1 · outbound

This paper cites Gesture-informed robot assistance via foundation models.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Gesture-informed robot assistance via foundation models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.603049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:c3643b86370a12f35b1ecf878fd33c5576f4d871695bbd7ad960bc3c4bc620e8

Observation 940510b7-d6c0-4d94-92e9-14a5af3548a0 · outbound

This paper cites Interactive multimodal robot dialog using pointing gesture recognition.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Interactive multimodal robot dialog using pointing gesture recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.599715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:409839da4f5885fcc371ea88f9c28b70a8ed08b9752e01db7e44009ecfbfb61f

Observation 496eb716-9ef3-4144-b255-ac71be72844e · outbound

This paper cites Code as policies: Language model programs for embodied control.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Code as policies: Language model programs for embodied control

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.596254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:7c14712e2031f6bbc6476be5488492c68937213225e1e0ad81ca442439ff39f2

Observation 0b0e1534-af26-4086-bd52-341690d9cfd4 · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Progprompt: Generating situated robot task plans using large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.592956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:7f10a52d8cca232f631092d5b6fcba57d49682bc58e7ded1ed35fa4e72c59731

Observation de495b96-5a16-4344-bf7c-f97d58b016be · outbound

This paper cites Semi-autonomous robotic arm reaching with hybrid gaze–brain machine interface.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Semi-autonomous robotic arm reaching with hybrid gaze–brain machine interface

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.589241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:8e4dd8be5e3730d044872ca2b2aeca061abc7e57bb8473e6e6ed5b306c27aef0

Observation 0c579db8-c1f5-44ad-abe7-36cdd32dcd11 · outbound

This paper cites Investigating the usability of collabo- rative robot control through hands-free operation using eye gaze and augmented reality.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Investigating the usability of collabo- rative robot control through hands-free operation using eye gaze and augmented reality

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.585779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:3d9df48f29747674f382174add8a8ccf37dafdceada1b8aea30c1c7605d7b25f

Observation c424de04-57ed-4f11-bde6-5bdfb2588bf8 · outbound

This paper cites Human gaze following for human-robot interaction.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Human gaze following for human-robot interaction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.582187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:41b35d5aa59418d1c3ea12812b30a5425f0b3f7cdc3161e2314914609044919d

Observation f52a08cc-f2d2-4f1a-9642-e0fb93061f79 · outbound

This paper cites Gaze-based attention recognition for human-robot collaboration.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Gaze-based attention recognition for human-robot collaboration

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.578972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:0f2a6ab6e82d12b65786af0e04780434bad59f93b54665df95dda1dffa0bcff6

Observation 5751a08f-385d-46d0-b29c-f6e50498ead3 · outbound

This paper cites A novel human-in-the-loop multimodal intention fusion method for human-robot interaction.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech A novel human-in-the-loop multimodal intention fusion method for human-robot interaction

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.575616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:ceb2e6bb67fcb64806b202000f1d081f53da8dd173c1d3e3a2335aafec86e963

Observation 8578005a-437a-419a-96ab-d558537fcaee · outbound

This paper cites Alchemist: Llm-aided end-user development of robot applications.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Alchemist: Llm-aided end-user development of robot applications

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.572180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:9013c4225ffa6d561cc8d302510aba45b6356ee6e1b0f2de835cb08985839e46

Observation ecbd186f-2683-4c15-8ba5-4f0afafe9617 · outbound

This paper cites Lami: Large language models for multi-modal human-robot interaction.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Lami: Large language models for multi-modal human-robot interaction

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.568621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:f122b45b86fef7fdabbbd9b34fd7d9e4b4a4f5cba2c00f8bc24b6ae10b8586a0

Observation bb177b39-f85f-4513-92e2-f0658bc631a8 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Robust speech recognition via large-scale weak supervi- sion

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.565092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:73da487f2f69bc2f3a0cec5d00b5f1ca7953dbc8149130b0112e76769ac91714

Observation a72528e2-30a4-4431-aa57-a0b3192fc074 · outbound

This paper cites Grounding DINO: marrying DINO with grounded pre-training for open-set object detection.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Grounding DINO: marrying DINO with grounded pre-training for open-set object detection

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.561791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:9ceed6ce726584a4b838d7e860fb51e74c063dfa7b84916c6e309a9bb31578fb

Observation ce1f1fe4-5005-4905-ad46-a85942036ea9 · outbound

This paper cites Sam 2: Segment anything in images and videos.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Sam 2: Segment anything in images and videos

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.558174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:5c47c608b8a3b5cee052f1d97b983cc7f89891edb1a9360f03075569c9bf7023

Observation 9fda874c-469f-4bb3-a6f7-507ac3e52dcc · outbound

This paper cites ORB-SLAM3: An accurate open-source library for visual, visual-inertial and multi-map SLAM.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech ORB-SLAM3: An accurate open-source library for visual, visual-inertial and multi-map SLAM

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.554943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:7e34693d228e8f0a9252a90840a58a8155dca4fa969975e51427fab30be4fa09

Observation 2c5ad54f-e294-4151-b0f0-2c4d1dcff7a8 · outbound

This paper cites Super- glue: Learning feature matching with graph neural networks.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Super- glue: Learning feature matching with graph neural networks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.551583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:843871d4aee1f5bd21abc5a182f4be85f144c848c0e4bf2ae5cdd1e4816edddd

Observation 1f99c207-0fb2-4a07-bf49-3715be11343f · outbound

This paper cites Design and implementation of a haptic measurement glove to create realistic human-telerobot interactions.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Design and implementation of a haptic measurement glove to create realistic human-telerobot interactions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.548125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:4f379b49569824328dba5329e498bae69e56488f96430d3d1c285a5a4a12b243

Observation 7b540175-6910-4a4e-a4f2-d200020897b2 · outbound

This paper cites Don’t yell at your robot: Physical correction as the collaborative interface for language model powered robots.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Don’t yell at your robot: Physical correction as the collaborative interface for language model powered robots

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.544376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:d986ba18e981240135390a645c48f935a0a364d01484fca7914ba57801f433b1

Observation 178de768-70af-43cc-90b0-4aaf91676fdc · outbound

This paper cites Lora: Low-rank adaptation of large language models.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Lora: Low-rank adaptation of large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.540619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:20a082ef331774fd9b5499435f9162d60bf5e701ed98db1d50dd1cfccc80830c

Observation acb66482-e5a4-4b3c-9f17-d229533c0146 · outbound

This paper cites Model compression.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Model compression

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.537242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:77dc1adb2e8cf01c46b6a28e7d317a723e6a68f77c11315a0732e5b00dca3756

Observation 814818ae-8916-4fba-aca1-5f51523fa296 · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Retrieval- augmented generation for knowledge-intensive nlp tasks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.534150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:ae539730e49755bad52a229e9c8e9db6f1b34f14cf9176f8860879748dc1076e

Observation b1c3cf0a-1399-45fd-85ed-783303144749 · outbound

This paper cites Human gaze improves vision transformers by token masking.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Human gaze improves vision transformers by token masking

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.530595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:165a3893dc53bf4cc98fa7ba5588d4a8ac19ca58adb66bcaf832aa9a89dfc3d7

Observation c56d6b22-0910-4d2f-940e-506fbbffe23a · outbound

This paper cites Egolife: Towards egocentric life assistant.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Egolife: Towards egocentric life assistant

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.527137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:a29c4ada3acd725089857da11e396415c8bba40edff8db5538a721d338e69a3a

Observation a4db9305-d286-4410-9fa0-ff7150e5c81d · outbound

This paper cites Interactive multimodal robot dialog using pointing gesture recognition.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Interactive multimodal robot dialog using pointing gesture recognition

Reference 42

Resolution
verified exact
doi, observed 2026-05-23T01:12:19.807631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:40859cfbc8823c8b02fc8289eef04210e0ae0a6b46b95f9229086af4634aca32

Observation 539233a1-3390-4a59-a9d1-24b70220ad98 · outbound

This paper cites The speech recognition error in our system primarily due to misinterpretation of similar-sounding words.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech The speech recognition error in our system primarily due to misinterpretation of similar-sounding words

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.523573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:ab9ebd3d0fadeaa389ebc6e73cd20a81432c67935f38ddceec12e65fcf337df9

Observation 1bf8564e-aeb3-4089-85b9-e3865399e044 · outbound

This paper cites an unresolved cited work.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:15:18.519970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:fe2f493599a2c73d90e735180cfd27bcf77c13ddc26a22204d41370e026b7c48

Observation eff27f64-7fbf-4972-a976-62c64a22edbc · outbound

This paper cites Fea- ture matching using superglue becomes unreliable in the presence of weak object textures, repetitive patterns, or partial occlusions, leading to incorrect correspondences.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Fea- ture matching using superglue becomes unreliable in the presence of weak object textures, repetitive patterns, or partial occlusions, leading to incorrect correspondences

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.516832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:24f6b106edde7db3062b9109178953b10b237088bcea5acba5e4617da9750a22

Observation afcba079-64b8-4e3e-ae7d-5b6cbcf93483 · outbound

This paper cites Since FAM-HRI requires the LLM’s response to strictly follow a prede- fined prompt format, any deviation renders the output unusable by subsequent system modules.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Since FAM-HRI requires the LLM’s response to strictly follow a prede- fined prompt format, any deviation renders the output unusable by subsequent system modules

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.513854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:de1a1378332b9c3237c0953e8130a1fbaf59ebe3d946f60452030d7f59d07095

Observation ee2e4734-419e-463b-9803-3a8b807a24d2 · outbound

This paper cites an unresolved cited work.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:15:18.510143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:a5ee221eb100fad9b874f1ceeafa99fad4a1eaaeb6a5d51059dab1b1efdf79be

Observation cc3d4649-3672-4fd2-80a2-53745af9c675 · outbound

This paper cites an unresolved cited work.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:15:18.507120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:12cef095def9c4a4c8b0ffcc44ae48759261e212829ac96793d26d2516baf6b7

Observation 11f0fb02-9e1d-4821-9e3e-14599f2dd333 · outbound

This paper cites an unresolved cited work.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:15:18.504367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:06ec3efce3f9db35e776f2b770bbd06abf76a08efff75f9789f46163a7e76d19

Pith citing papers

Observation 829ca9da-7935-4272-bb01-b912871b234b · inbound

Natural Multimodal Fusion-Based Human-Robot Interaction: Application With Voice and Deictic Posture via Large Language Model cites this paper.

Natural Multimodal Fusion-Based Human-Robot Interaction: Application With Voice and Deictic Posture via Large Language Model FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:46:48.174002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:46:48.174002Z digest=sha256:d951c41394946c0765096fce27aeb5d215df7e5f73e58775c20db2c4207d4787

Observation fb998fb0-166b-48cd-a1b3-444875bef31d · inbound

Gaze-supported Large Language Model Framework for Bi-directional Human-Robot Interaction cites this paper.

Gaze-supported Large Language Model Framework for Bi-directional Human-Robot Interaction FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:10.299425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:10.299425Z digest=sha256:87a4da4c16dcf436b8d6ef7ea3a65cec1f9d7a7c68c5969caaf267e658597c4f

Observation f9ad8420-686a-4da6-9528-a95aa196d81e · inbound

Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective cites this paper.

Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-04T19:39:07.075356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:39:07.075356Z digest=sha256:507a049d558ab7bae5ada478c1c9ab14d8d4631acaa6309d19c5a9e71e2e0df9

Observation b9f2c4a1-f946-4ca1-afb4-b51790544ce6 · inbound

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving cites this paper.

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-22T12:21:31.264373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T12:17:54.325055Z digest=sha256:870c1ed2f4016db7e3f289bd9a8493370d8e6b339ee0ba4434ff14652b4b5dae

Observation 0601019d-8154-4643-a96d-f50bc647a580 · inbound

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving cites this paper.

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T17:07:42.809307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:07:42.809307Z digest=sha256:20e2d8f9819dad4867727ab8c1f3d4d3cdf3576c18d2ce86882b3a48f64b1b55

Observation 8c624a78-df58-47cc-82e1-a232ae2d3f6a · inbound

Hierarchical Policies from Verbal and Egocentric Human Signals for Natural Human-Robot Interaction cites this paper.

Hierarchical Policies from Verbal and Egocentric Human Signals for Natural Human-Robot Interaction FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:57:38.606095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T13:31:02.640975Z digest=sha256:7e6d19cfee0147813f95fd6febd012cf18551f5554f103c828eb5255272b2180