Pith. sign in

Paper Citation Record · LEDGER

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition

As of 21 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2506.20373.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20373 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:53:19.917077Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact3
  • verified fuzzy9
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 522050a9-df17-427e-b02e-1a90539db50f · outbound

This paper cites Large language models for human–robot interaction: A review,.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Large language models for human–robot interaction: A review,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.728747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.728747Z digest=sha256:8ef65478ef3e411f4aba3f57e51e75e1b8c6521adb2bd69d09c9be8f4e3b64db

Observation fb58932b-dda3-409d-b3c1-e5a7b9db7f3a · outbound

This paper cites A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.736138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.736138Z digest=sha256:5f2ad269525eb5e79804b93384a5ba770525dc6769c2e8ed2cb8f77539f41c46

Observation a67e63f2-50c3-4244-b6a0-ffe57fa8d91f · outbound

This paper cites MUTEX: Learning Unified Policies from Multimodal Task Specifications.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition MUTEX: Learning Unified Policies from Multimodal Task Specifications

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.742757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.742757Z digest=sha256:4ff45703b224c58597d6e88757bb3752c6847099496c5b33c58ba1c467d96cb4

Observation e7e3ad0e-3de1-4113-b5b5-6c1193f37026 · outbound

This paper cites Vision- language model-driven scene understanding and robotic object manip- ulation,.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Vision- language model-driven scene understanding and robotic object manip- ulation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.750227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.750227Z digest=sha256:563673b5a6a7e3bc095b621cd4eaede4d5001919904341805a077a33c0575abe

Observation f2dd8171-240b-4c6d-965e-dc82a58a169d · outbound

This paper cites CoPAL: Corrective Planning of Robot Actions with Large Language Models.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition CoPAL: Corrective Planning of Robot Actions with Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:53:20.119684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.770138Z digest=sha256:3a892e94f22db584ab6ab6dada6cb165845abed0a71ee2343a3a667cd0932064

Observation 7d07ab79-1a41-4486-b734-31d71a195f24 · outbound

This paper cites VLM-Social-Nav: Socially Aware Robot Navigation through Scoring using Vision-Language Models.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition VLM-Social-Nav: Socially Aware Robot Navigation through Scoring using Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.778811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.778811Z digest=sha256:e3a3b1227506a229c8ab7a5d06596643b538141fcd26d24d69e6aeb5c1831e0e

Observation baffae8f-7b45-4d4f-8380-a6bb2ed302bc · outbound

This paper cites VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.786689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.786689Z digest=sha256:5c9eee7c755515e02cdf6d25897cfc94769aa47f30d1f365738afca845bda942

Observation 5fcfe77a-5771-4099-aae6-c6b2b638058f · outbound

This paper cites ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.792331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.792331Z digest=sha256:a2d0571013a755195d518ffdefed970fb298cb40599881b895672ffa515b367a

Observation 932c51cd-ff69-4386-ae63-34c696512b8d · outbound

This paper cites LaMI: Large Language Models for Multi-Modal Human-Robot Interaction,.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition LaMI: Large Language Models for Multi-Modal Human-Robot Interaction,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.797931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.797931Z digest=sha256:6994e4614bb40f94d3e713d8dc3a3fd09945a4c05aa5faf0b6b1fde1ffcc12f9

Observation 406d83e2-0e16-451a-aed4-c583e86af768 · outbound

This paper cites To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions,.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.802938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.802938Z digest=sha256:88a93586fd476f2f4c0571c7cb98c4fe7318a123a539314d7ba93278265fe0b9

Observation ce2a863c-e62b-4d0f-a10e-5e61b1b19f45 · outbound

This paper cites VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model,.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model,

Reference 12

Resolution
verified exact
doi, observed 2026-08-06T22:53:20.230056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.808029Z digest=sha256:e1c0cca836aa137b3c713704375e0622ec55bac09f249d8e6b5e0f56fdb16aeb

Observation 8c86f276-7e64-4c68-be46-d2b17b4e4b8f · outbound

This paper cites Robots Can Multitask Too: Integrating a Memory Architecture and LLMs for Enhanced Cross-Task Robot Action Generation.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Robots Can Multitask Too: Integrating a Memory Architecture and LLMs for Enhanced Cross-Task Robot Action Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:53:20.036658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.813106Z digest=sha256:6b5757d6b4c8daf66fe09f1cabc6f92ade5f12150610762da8cb58735f055db2

Observation 0cf753cb-4b92-40f3-8b1a-40e2c1f0132d · outbound

This paper cites “Exploring large language models as a source of common-sense knowledge for robots“.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition “Exploring large language models as a source of common-sense knowledge for robots“

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.929041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.818252Z digest=sha256:1fb745a02301368367d1b9138b54644e17705360968113b916ef1d4cf7cb7e53

Observation 85859591-6b94-439b-b285-fd18f478b147 · outbound

This paper cites an unresolved cited work.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:53:20.911871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.823043Z digest=sha256:7245ec021be459648ac6c16f8bdb01632fc588acf2b277ce521f1f11c3d5b1d9

Observation 0d7d1a81-4741-41e6-8852-3370bab662d9 · outbound

This paper cites an unresolved cited work.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:53:20.893612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.827857Z digest=sha256:e867d0eff0092db9c8899a76c21b65ade2c19ddbe359693d0dae253538b0bc08

Observation e61757eb-223c-4dd2-8d87-f9d87ac45006 · outbound

This paper cites “Quo vadis, action recogni- tion? A new model and the kinetics dataset.“ Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition “Quo vadis, action recogni- tion? A new model and the kinetics dataset.“ Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.876697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.833424Z digest=sha256:fbcb6cea833c58a8d4fa938fa15d96e60e38fb3c2acc578dbca7fbfd0966cf9a

Observation 29d4c07f-d1b3-4335-b8e8-f7e59a6d1aba · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.844287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.844287Z digest=sha256:c2a929eb69f726b9949ce26a337cdf819439eee2967eba331f200cc1cd197a0d

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.849314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.849314Z digest=sha256:4c8a3da5e71bf0c27caa7e2930554093080c721c33dab3871b9887ebe79f5b41

Observation baf6a17e-7231-4b82-9dfd-887a732ae259 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.854918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.854918Z digest=sha256:f0cc667de29d3034af0efc68c8112a9b7b627d7a7cfef91c8e03f198af7af169

Observation 6299b6e2-5fb4-475f-a796-8ce59e463f18 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.860241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.860241Z digest=sha256:93b05bfec418ac723d0ca7561a4443a2da3d4d7dd8083976dcb07f6b8d9dbb8f

Observation 7d5f813b-6236-424a-a817-6f9a4bdba0ad · outbound

This paper cites and Kembhavi, A., 2020.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition and Kembhavi, A., 2020

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.860482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.865807Z digest=sha256:71767de949e0f535e21e0c376158e9c00a6e737f9b0b66c94027c90bc00c3b2f

Observation 05318ce4-1021-43ad-8a6d-e17ca64811f7 · outbound

This paper cites and Chen, L.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition and Chen, L

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.843586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.870783Z digest=sha256:e94fefda4b91f73e4e8b55b291a0113ad0c5e53854950fab20ce724ab49aa02f

Observation 6db1a9ce-b900-41af-8861-827df54ad46e · outbound

This paper cites and Saffiotti, A.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition and Saffiotti, A

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.823612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.876678Z digest=sha256:28a0e9f4002c2eb2eb0b4eec320adcc1663dfe4e062f8b9427dd0db28eb549d0

Observation 23e42ec6-a853-4166-87e9-2a1f9b80d6a0 · outbound

This paper cites and Ros, R.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition and Ros, R

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.805943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.883161Z digest=sha256:1ae61f34b4fa84038e439fe1fd244dea286bcfdc111074fee37e2e35087d5d80

Observation 40fe6b39-86f5-46bb-9b4d-9ed17c63b2aa · outbound

This paper cites and Deigmoeller, J.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition and Deigmoeller, J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.789418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.888879Z digest=sha256:ea5d527629273f61f88282dcf5e4d72693b538e4f8df2ee34b19b117fd6786c6

Observation e14c0a4f-93e7-4c41-bc1f-3dc66ed2ff6f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.893891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.893891Z digest=sha256:be8304c1572d0b7fd2b061faa1f458b673f0d28a0164ca6eb1cd7ba8c6220402

Observation 516b5a2c-9614-4e39-8142-89ae274d8c21 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.903289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.903289Z digest=sha256:c2e7fbdd5b658546dc161b369990e51c7d1b090c9eae3b11f4382b54184f9fcb

Observation a38c89f6-c7b7-4bf7-8390-90fe92ccbfe2 · outbound

This paper cites Vila: On pre-training for visual language models,.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Vila: On pre-training for visual language models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.772519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.909474Z digest=sha256:1610901e61fa779b4ed46d0016d01db0baffd73f9c9994005cb1d17e858e9f6f

Observation 64ac326d-5dcd-4fd6-aa71-ce4ec4b72c73 · outbound

This paper cites Kr `‘uger, C.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Kr `‘uger, C

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.754860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:53:19.917077Z digest=sha256:c670fd824ac9046d39f802ac0bb67e207c6521341936187d458a104130b3c40c

Pith citing papers

No inbound Pith citation observations are available.