Pith. sign in

Paper Citation Record · LEDGER

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction

As of 14 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 2 inbound Pith citation observations for arXiv:2506.20566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20566 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:49:10.881232Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:43:33.391528Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a90adff-d994-4835-914a-c1ef00a2389d · outbound

This paper cites Evaluating Vision-Language Models as Evaluators in Path Planning.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Evaluating Vision-Language Models as Evaluators in Path Planning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:49:11.200446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:07.509682Z digest=sha256:49823635d41c6c2ca8d258e0419bf522ace0c4b4b166291d08ec2918d254582d

Observation 962eeeb6-509b-491b-a217-70702f5a12f5 · outbound

This paper cites In: Proceedings of the 19th ACM International Conference on Mul- timodal Interaction, pp.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: Proceedings of the 19th ACM International Conference on Mul- timodal Interaction, pp

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:17.057335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:07.590463Z digest=sha256:552c159674ea4c4e8a792086152d4f730762dfb6f36de21f3849341f805ac989

Observation 71e0cde9-4a6e-48b7-bf08-c46222e1b4e5 · outbound

This paper cites Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:07.721472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:07.721472Z digest=sha256:d725399c32519db1421f4b0cd8da03248d7827cebbcae197c5bcdc1a5cd99daf

Observation 6138fe75-e541-450c-a7b0-2affee204481 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:16.795102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:07.813078Z digest=sha256:2465adfcbe26deb1da0877acf627e052d6c042ba15e4261c4adcc7dde53fd235

Observation a50c31f4-58d4-4301-bb48-80366c6c1157 · outbound

This paper cites In: International Conference on Learning Representations (ICLR) (2025).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: International Conference on Learning Representations (ICLR) (2025)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:16.485703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:07.915292Z digest=sha256:92bdb59e1dda818e6029576f3f41252df5860c91a16f4949d77c85f9e52df586

Observation 500ca7e8-8220-4aff-b648-776631922f9f · outbound

This paper cites In: 2008 8th Ieee International Conference on Automatic Face & Gesture Recognition, pp.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: 2008 8th Ieee International Conference on Automatic Face & Gesture Recognition, pp

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:16.193859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:08.077537Z digest=sha256:ae3861d338ede29621fb4095ba08e5b4a5bd15593eee0c9006ba97e5ecc2653d

Observation 43fee879-e3ea-4fe5-9578-b03d930b3efa · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Gemini: A Family of Highly Capable Multimodal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:08.200579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:08.200579Z digest=sha256:855d89c804c8dd8c5a66aa803a7c2006c9a1da446618b7abf2bac9a5efe960d3

Observation 04f16770-e207-4dac-b03e-0b76b6c81956 · outbound

This paper cites The Llama 3 Herd of Models.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:08.371492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:08.371492Z digest=sha256:7bc72cfc21fdb0f36e34c8e0aa5dbd43e305bbdac55d32cddba6e81f364fd605

Observation 7cfc5b71-6772-4b10-abd6-4f8f340cb62e · outbound

This paper cites In: Neural Information Processing Systems (NeurIPS) (2018).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: Neural Information Processing Systems (NeurIPS) (2018)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:15.866164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:08.475881Z digest=sha256:5fd52edf07effd3bac840c6602c1f63c73ae5261b4709a51c07f30b9a78ab67b

Observation 84759529-e6ae-4686-a506-72c9db046e09 · outbound

This paper cites Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:08.555776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:08.555776Z digest=sha256:24f429830b71436a9799f586eed3a81df7a17cf18515181fa05542a0d881dfd7

Observation eaf7af65-57fd-49c9-b145-d09867dd3e2f · outbound

This paper cites Human factors53(5), 517–527 (2011).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Human factors53(5), 517–527 (2011)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:15.496895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:08.658326Z digest=sha256:eb8d1b30c84613e4b331bdfcd32a4989df41cfb1fbe92d3ef034581f1dea61ab

Observation 59c4f372-2edd-4317-aaea-6ba8c91afd8c · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:08.758255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:08.758255Z digest=sha256:7045ecdd54437de1a243f9c4800feb8ba0a0de6f392e0d741bba81d98e9ea1a4

Observation 3440f35e-8fb1-4ba9-9e23-b74c19f90f4d · outbound

This paper cites In: Conference on Robot Learning (2023).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: Conference on Robot Learning (2023)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:15.105483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:08.876306Z digest=sha256:ca1d58a4f4dbe1cb1d5b18dcc4054aaf849833fbfe58e0dba8fe042c24f6d125

Observation 85b25a58-aa63-465c-9dac-fc4084039e6b · outbound

This paper cites Discourse Processes52(4), 255–289 (2015).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Discourse Processes52(4), 255–289 (2015)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:14.752505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:08.993107Z digest=sha256:1ce51f4bb41fe51065f4bf855314650c8cc9dbd27bc6f7774a135f245ed5e83c

Observation 2a620a9f-dda5-4e84-9be3-0afa2b7e443a · outbound

This paper cites A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:09.088375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:09.088375Z digest=sha256:b0400ce31619b66f27d2e106959c8767fc393213492dadc7d9b1c8e242a759cd

Observation f7deb7e5-620d-4d07-afaa-baa9253c1041 · outbound

This paper cites In: Robotics: Science and Systems (RSS) (2024).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: Robotics: Science and Systems (RSS) (2024)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:14.371159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:09.199127Z digest=sha256:8f9ce77641372ca1a14a82e89b6ea8e5cc4cb47c5cca58d1ebabeb96428b5378

Observation 771afcae-c807-47e1-bc8d-dfd15e5b7eca · outbound

This paper cites ACM Transactions on Human-Robot Interaction12(3), 1–39 (2023).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction ACM Transactions on Human-Robot Interaction12(3), 1–39 (2023)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:14.066289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:09.350710Z digest=sha256:c4765a62c02dae44bd9be5c82bb0518a14f482f0326fd566fc6e646488156c71

Observation f17af41b-8e60-4a27-bc26-6d2157040c84 · outbound

This paper cites https://robotsguide.com/robots/kuri (n.d.).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction https://robotsguide.com/robots/kuri (n.d.)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:13.769324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:09.505490Z digest=sha256:fb7c962707e23b85647274f7c1cd67b1a9028afec492d20e9a693dd77ac7d1d0

Observation c8b0254b-7d9e-4bb3-8206-9beb73ad23f3 · outbound

This paper cites In: 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:13.346604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:09.666785Z digest=sha256:7c457c674cda8040d6af6fc2bcddbac9530007df2638a519f415f27cb7b245fe

Observation 7fc6ee28-5da2-48ca-b3e4-943eb4eaed6d · outbound

This paper cites Image and Vision Computing25(12), 1875–1884 (2007).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Image and Vision Computing25(12), 1875–1884 (2007)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:12.972703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:09.869220Z digest=sha256:38109708d26c4d56327f913b8de7ba3009219f51a74627e3429356580bc81797

Observation da941580-0f76-4fb9-a786-311d074a0bac · outbound

This paper cites GPT-4o System Card.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction GPT-4o System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:10.058113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:10.058113Z digest=sha256:595747eee1d92303309e4f7e0965326388421b10603d17ba89d463da03553166

Observation e9872746-cccb-424c-8a30-bae22d03fa1b · outbound

This paper cites ACM Transactions on Human-Robot Interaction12(1), 1–66 (2023).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction ACM Transactions on Human-Robot Interaction12(1), 1–66 (2023)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:12.673737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:10.170501Z digest=sha256:a2b02bbc9989cbb6f76a1e822bdb2393b534c112a9b54962b705fba2d9369b90

Observation ac043cf2-4523-4df1-a8a3-7a424b53de3f · outbound

This paper cites IEEE Transactions on Robotics38(3), 1755–1772 (2021).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction IEEE Transactions on Robotics38(3), 1755–1772 (2021)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:12.374698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:10.326554Z digest=sha256:f0cc42f960e4000b8384ede282d325f6e89eb39dabf82d4c3d80711c361793d7

Observation d24db9c4-7666-4fdf-9c03-eba2412b1b6d · outbound

This paper cites ACM Transactions on Human-Robot Interaction 12(2), 1–21 (2023).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction ACM Transactions on Human-Robot Interaction 12(2), 1–21 (2023)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:12.095926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:10.477606Z digest=sha256:6f941ea7b0e112ed30639b0963e7b16da21dc911f240b374953f06ce80c6221d

Observation e218a7f9-66ac-44ac-9c42-c00fa81747a2 · outbound

This paper cites Advances in Neural Information Processing Systems35, 12,014–12,026 (2022) 10 Zhonghao Shi et al.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Advances in Neural Information Processing Systems35, 12,014–12,026 (2022) 10 Zhonghao Shi et al

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:11.776077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:10.621525Z digest=sha256:e188f14cf23903dbcb6db1be215e342a3b8e194dc6f678d6f1a144eea602903e

Observation f2ad57d4-4681-4293-8b70-f9451405badc · outbound

This paper cites ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:10.769599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:10.769599Z digest=sha256:f96c8315363c7069a7b29d44ca5f94d912b68ab03ea683176e50c3b89b51f3b6

Observation d6210fe0-1cc8-4a6a-9776-53739c97e593 · outbound

This paper cites In: Artificial In- telligence and Machine Learning in Defense Applications II (2020).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: Artificial In- telligence and Machine Learning in Defense Applications II (2020)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:11.460335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:49:10.881232Z digest=sha256:7871c32bbe2eaa83c01214c9492e3cf92a67b0eb0738c206bf9f88ffa330d382

Pith citing papers

Observation e94ac13b-5b75-40fb-92fe-c1c5ee6012dc · inbound

ERR@HRI 3.0 Challenge: Multimodal Detection of Errors and Anticipation in Human-Robot Interactions cites this paper.

ERR@HRI 3.0 Challenge: Multimodal Detection of Errors and Anticipation in Human-Robot Interactions HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T04:42:38.419346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:42:38.419346Z digest=sha256:37c86f60974deee07a7002ac043a0d94460d8ebc20e6db7854faf6fb6d0c4a7a

Observation f3cf56be-7b34-4131-812e-7362bbde9172 · inbound

HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration cites this paper.

HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:43:33.391528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:43:33.391528Z digest=sha256:9a43318dd018c8a064c9e25737627126d1c05240ef212a4a801d49c4843da045