Pith. sign in

Paper Citation Record · LEDGER

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2605.00963.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.00963 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T18:37:32.865795Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact5
  • verified fuzzy22
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 183148c1-d87f-4995-936b-ad28412c9836 · outbound

This paper cites An advanced medical robotic system augment- ing healthcare capabilities-robotic nursing assistant.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task An advanced medical robotic system augment- ing healthcare capabilities-robotic nursing assistant

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.203037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:a432d0a248357b39dd17575ddb0b3d6a4440d01b3433e4a017d331c5774641d2

Observation 4c39b7fb-c675-4a5e-b346-a92b0075a63a · outbound

This paper cites A human-robot interac- tion applicution based on augmented reality (ar) for industrial robot grasping process.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task A human-robot interac- tion applicution based on augmented reality (ar) for industrial robot grasping process

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.192863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:7287e5b54a0e6c52d325bb68fb15dc0b880134c59b85ba08b3694f8b73608c36

Observation 5f767e74-d661-4836-97ee-93e2fa331a0d · outbound

This paper cites An educational robot system of visual question answering for preschoolers.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task An educational robot system of visual question answering for preschoolers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.196266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:568a78a8d3abfb828264f6edb7583ee77712db9e59d108d6e3d277578af77587

Observation 1b97e164-789b-43b4-bb42-d7ad824ecdbb · outbound

This paper cites Home robot service by ceiling ultrasonic locator and microphone ar- ray.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Home robot service by ceiling ultrasonic locator and microphone ar- ray

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.199644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:596a80165888f85823b75497e5287b1c46b5e9731cc064a880fa5a857d2df328

Observation 4cb39b36-ef97-4a59-97be-d4ef7086790e · outbound

This paper cites The human intention: a taxonomy attempt and its applications to robotics.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task The human intention: a taxonomy attempt and its applications to robotics

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.209649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:9f6794814bf7246e241d549da40ea27b0fdc4724316d1838efc4bc2b940cac7a

Observation dad20ba3-845b-467a-8972-462403d80010 · outbound

This paper cites Anticipatory robot control for efficient human-robot collaboration.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Anticipatory robot control for efficient human-robot collaboration

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.212910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:bf43cec06691f119c6e3ffdd91049cda951c792f944a3b14a55e5d26a851f150

Observation 2497d62a-9f5e-4f98-8efc-6ffa89c4794e · outbound

This paper cites Pointing gestures for human-robot interaction with the humanoid robot digit.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Pointing gestures for human-robot interaction with the humanoid robot digit

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.175250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:c68bf476f4bc106df55a86b32585ec1fcda5d77684f3eaa04ff85a7827bbb124

Observation b4cf76bc-5a15-4b13-bf98-bd7e7f99f3ed · outbound

This paper cites Autonomous laparoscopic robotic suturing with a novel actuated suturing tool and 3d endoscope.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Autonomous laparoscopic robotic suturing with a novel actuated suturing tool and 3d endoscope

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.178563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:de188c74a5b66ecd4ee68fe65658c161de072ee9b8126015ef3d18295b716cca

Observation 57376c0c-cd56-4d10-b263-a654bf42da18 · outbound

This paper cites Perception– intention–action cycle in human–robot collaborative tasks: the col- laborative lightweight object transportation use-case.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Perception– intention–action cycle in human–robot collaborative tasks: the col- laborative lightweight object transportation use-case

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.182051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:89a496e8ebe1a77a4795adb9c96d663fdebd22185f7a8f86125f8a50f9f21280

Observation 72930273-a607-48c6-8b7a-fa1bf6619a00 · outbound

This paper cites Exploring transformers and visual transformers for force prediction in human-robot collaborative transportation tasks.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Exploring transformers and visual transformers for force prediction in human-robot collaborative transportation tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.185616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:c2e4ba8bba6d48a628cfcada4543470df8045dcd82e71a3fc74d88d69342aa0f

Observation 0bae2e8b-63e6-4330-afbc-a4e404d56cb2 · outbound

This paper cites Force and velocity predic- tion in human-robot collaborative transportation tasks through video retentive networks.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Force and velocity predic- tion in human-robot collaborative transportation tasks through video retentive networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.167802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:815f9b97e297a337f004222768282c7a572d1c9f9725e98cf5c7a22671a9f050

Observation f2372dc3-f71f-4b7e-afbe-ae0f8c7f8912 · outbound

This paper cites Language and sketching: An llm-driven interactive multimodal multitask robot navigation framework.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Language and sketching: An llm-driven interactive multimodal multitask robot navigation framework

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.160638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:d698e168ce3c85a91089c4a2e7de70e3aa878fb579c6eb68f44b3a4d181d4e9b

Observation 22cf6f80-e16a-4957-ba16-d774eb3240a8 · outbound

This paper cites Interactive navigation in environments with traversable obstacles using large language and vision-language models.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Interactive navigation in environments with traversable obstacles using large language and vision-language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.164292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:5be5ca422215af019ce7d683a493958b9a50f6ee887fc219d93375ebbb5dd88d

Observation 8ed54e66-2d45-4925-a56b-e6218c2ba471 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Physically grounded vision-language models for robotic manipulation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.171128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:afea8e03a60dded7cb0c10d025725115252d7b35257e08a5d57327471edb0111

Observation 030fc9cd-05b7-426a-b924-e44f6f4c3c36 · outbound

This paper cites When the inference meets the explicitness or why multimodality can make us forget about the perfect predictor.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task When the inference meets the explicitness or why multimodality can make us forget about the perfect predictor

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.189369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:a68d89372eddd3c26d10d94133dca20569d088ba2150d712cc7f802b0b4ffcc9

Observation f8fd0236-cd83-4c0c-ae9e-ac6ae973bd37 · outbound

This paper cites Anticipation and proactivity. unraveling both concepts in human-robot interaction through a han- dover example.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Anticipation and proactivity. unraveling both concepts in human-robot interaction through a han- dover example

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.206291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:3333eab4b31071a691f883a1edab9d153be24bde32d2d2608e41f02ba34b445f

Observation f2799689-edbc-4285-8b07-4686b201b390 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task LLaMA: Open and Efficient Foundation Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:11:08.303610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:f6c6bdc57bc4d8e1ad845216530be5c87ebbff89278aa511518f694c198a7824

Observation 3b09c363-9f43-462d-82f8-163f0228952d · outbound

This paper cites GPT-4 Technical Report.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task GPT-4 Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:11:08.332019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:6adfe8512d0a19581d93cc9c76f6101f4241a286874bd757b10ba6e5df840a2d

Observation de45a080-e331-48cc-8793-67ce3e8e4c70 · outbound

This paper cites Leveraging Large Language Models in Human-Robot Interaction: A Critical Analysis of Potential and Pitfalls.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Leveraging Large Language Models in Human-Robot Interaction: A Critical Analysis of Potential and Pitfalls

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.312260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:2e372ae387e609feaa02a3ca5074f6c2941e5fc1c58405894878064e31213d01

Observation 2d6f11dc-a3a9-4e3e-9ad5-b08864f22db5 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.216378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:6a2250109a55e3d6b7bfeacd8641014b1f5283b35443aaf506031f60399d1075

Observation 648ebf6d-0598-46ac-96bf-602e3686f29a · outbound

This paper cites Robust speech recognition via large-scale weak super- vision.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Robust speech recognition via large-scale weak super- vision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.219864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:2618d836a818bfb604ff7f1143c87d29e16b2621c6ccc284b45f8f6900d78772

Observation 1e947acc-fc63-438d-920b-2950bf62b7eb · outbound

This paper cites AST: Audio Spectrogram Transformer.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task AST: Audio Spectrogram Transformer

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.290066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:686907ef1e94c0442a87a9e5209dbc12f70a6c5938e2121c0f078afb18a188f8

Observation b3a39b44-fbaa-42af-a3df-96d3d1081215 · outbound

This paper cites Fuzzy logic systems for engineering: a tutorial.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Fuzzy logic systems for engineering: a tutorial

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.149913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:f008bed44528f24230634dead7dd26adf6d711281574bcca8ef5d459cf289ccc

Observation 3c8abc05-8cd9-455e-82ed-4a69708aef0f · outbound

This paper cites Interval type-2 fuzzy logic systems: theory and design.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Interval type-2 fuzzy logic systems: theory and design

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.153373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:4e6f7b6c140f8f77a4a1494ff3ccafc9a6b384204783fc16f140960f24075125

Observation 8491ac5d-1db2-4bc8-8f56-ca3ac55856c1 · outbound

This paper cites Fuzzy logic introduction.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Fuzzy logic introduction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.156900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:ba19461195630dc71744cb7caa32a62cc6fadaa1fcbb0794b6d70e5d38fcd6eb

Observation b35d1674-42f5-4558-8681-a1df0b965274 · outbound

This paper cites A type-2 fuzzy logic controller for autonomous mobile robots.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task A type-2 fuzzy logic controller for autonomous mobile robots

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.146744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:032a9b4722abb870256868ad0abc07037cd8532012ea6e493936857291b5b457

Observation d5daa156-1f5f-444d-8b37-f23969b1d5f3 · outbound

This paper cites An approach to combining video and speech with large language models in human-robot interaction.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task An approach to combining video and speech with large language models in human-robot interaction

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.325436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:68937361fcaaedd4c055d76fcf3add309a73d39539860f6c142fc16de274b180

Pith citing papers

No inbound Pith citation observations are available.