Pith. sign in

Paper Citation Record · LEDGER

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2506.20373.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20373 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:53:19.917077Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact3
  • verified fuzzy9
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 522050a9-df17-427e-b02e-1a90539db50f · outbound

This paper cites Large language models for human–robot interaction: A review,.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Large language models for human–robot interaction: A review,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.728747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.728747Z digest=sha256:f8fab3539368abbcacde3b342af7a2c4292f8f34a774ceb69bcf519bf3e43247

Observation fb58932b-dda3-409d-b3c1-e5a7b9db7f3a · outbound

This paper cites A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.736138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.736138Z digest=sha256:63f1ca082cbe5932fe02485855a020add8e3fbe28ad45d0542571febe8deedfe

Observation a67e63f2-50c3-4244-b6a0-ffe57fa8d91f · outbound

This paper cites MUTEX: Learning Unified Policies from Multimodal Task Specifications.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition MUTEX: Learning Unified Policies from Multimodal Task Specifications

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.742757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.742757Z digest=sha256:af652210bd860b59a402f2951a0b3eb10160ce5e537ca99dd0078ee8b97a1d70

Observation e7e3ad0e-3de1-4113-b5b5-6c1193f37026 · outbound

This paper cites Vision- language model-driven scene understanding and robotic object manip- ulation,.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Vision- language model-driven scene understanding and robotic object manip- ulation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.750227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.750227Z digest=sha256:960565d769c31b0fdfc2d6f94e222017aca4c4498d64db3842bb74f3bf4c2da3

Observation f2dd8171-240b-4c6d-965e-dc82a58a169d · outbound

This paper cites CoPAL: Corrective Planning of Robot Actions with Large Language Models.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition CoPAL: Corrective Planning of Robot Actions with Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:53:20.119684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.770138Z digest=sha256:bf5291e9f76efc023d7a5f0e612a3170cb99ad53fcefbf68c6e862b4ca827aa2

Observation 7d07ab79-1a41-4486-b734-31d71a195f24 · outbound

This paper cites VLM-Social-Nav: Socially Aware Robot Navigation through Scoring using Vision-Language Models.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition VLM-Social-Nav: Socially Aware Robot Navigation through Scoring using Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.778811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.778811Z digest=sha256:be466d8b2155805189bd4e74b73f7b1473ffb589c34834f4086002ab80e4665a

Observation baffae8f-7b45-4d4f-8380-a6bb2ed302bc · outbound

This paper cites VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.786689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.786689Z digest=sha256:accf6866b5e4aa225e3aad31df55042714db5975a314625893c7b37817e36cac

Observation 5fcfe77a-5771-4099-aae6-c6b2b638058f · outbound

This paper cites ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.792331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.792331Z digest=sha256:1720a216e6b914df5023b7f0381df2a39dfb76524c1ed9371ae94dd5133e1044

Observation 932c51cd-ff69-4386-ae63-34c696512b8d · outbound

This paper cites LaMI: Large Language Models for Multi-Modal Human-Robot Interaction,.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition LaMI: Large Language Models for Multi-Modal Human-Robot Interaction,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.797931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.797931Z digest=sha256:16dc250c69d9effdac2bb4a0be04d82a82fac75c4d6ed56ecbc7740a6bb2bee0

Observation 406d83e2-0e16-451a-aed4-c583e86af768 · outbound

This paper cites To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions,.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.802938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.802938Z digest=sha256:0ad094cc32d009407760ff7d666cafb2798f553fd57458b18097dc901aad0b3e

Observation ce2a863c-e62b-4d0f-a10e-5e61b1b19f45 · outbound

This paper cites VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model,.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model,

Reference 12

Resolution
verified exact
doi, observed 2026-08-06T22:53:20.230056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.808029Z digest=sha256:e74bff6e85e892cd8f736509be28ccdb012ddc27efee7cacb574845d1fab9805

Observation 8c86f276-7e64-4c68-be46-d2b17b4e4b8f · outbound

This paper cites Robots Can Multitask Too: Integrating a Memory Architecture and LLMs for Enhanced Cross-Task Robot Action Generation.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Robots Can Multitask Too: Integrating a Memory Architecture and LLMs for Enhanced Cross-Task Robot Action Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:53:20.036658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.813106Z digest=sha256:e1b3b2917dc71e2b06da1dcae1474ecef4d47c7ed27fab07050630d4eddf8abb

Observation 0cf753cb-4b92-40f3-8b1a-40e2c1f0132d · outbound

This paper cites “Exploring large language models as a source of common-sense knowledge for robots“.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition “Exploring large language models as a source of common-sense knowledge for robots“

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.929041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.818252Z digest=sha256:7f9628040fb05f22cf8342c259857e672755f972d7f1784c330ed534d25d8464

Observation 85859591-6b94-439b-b285-fd18f478b147 · outbound

This paper cites an unresolved cited work.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:53:20.911871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.823043Z digest=sha256:e3583297e70f78fb19d77e6b66d62521a9a98e7dc9962fb7201701ab96808350

Observation 0d7d1a81-4741-41e6-8852-3370bab662d9 · outbound

This paper cites an unresolved cited work.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:53:20.893612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.827857Z digest=sha256:463fb32d65707ce0dad26da8f48b23759e09f071659fd5ab571f793eb6a8f4e0

Observation e61757eb-223c-4dd2-8d87-f9d87ac45006 · outbound

This paper cites “Quo vadis, action recogni- tion? A new model and the kinetics dataset.“ Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition “Quo vadis, action recogni- tion? A new model and the kinetics dataset.“ Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.876697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.833424Z digest=sha256:d687a2d58c3f4f826f408944e5b4edf015a9b9d63d2f89d67ba4ef5ea0882cd7

Observation 29d4c07f-d1b3-4335-b8e8-f7e59a6d1aba · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.844287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.844287Z digest=sha256:5b1788ba643b825cd65aa150c2950ab02256e8a5fe4f0f4bdca4779c4d05a051

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.849314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.849314Z digest=sha256:36f7646aa29dd640b09502775e2e6661970d6266b92c8a032145db34f98b0493

Observation baf6a17e-7231-4b82-9dfd-887a732ae259 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.854918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.854918Z digest=sha256:4ed5fb702055c159952fedc8ef2653f3e7182a51e4e5e3f302c7aa6a7c6e37d4

Observation 6299b6e2-5fb4-475f-a796-8ce59e463f18 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.860241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.860241Z digest=sha256:2754b182e7be53db200a91c947aef47e6836a8d7d3a73c05fe1e3639c275be64

Observation 7d5f813b-6236-424a-a817-6f9a4bdba0ad · outbound

This paper cites and Kembhavi, A., 2020.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition and Kembhavi, A., 2020

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.860482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.865807Z digest=sha256:0f925f358a804d7545356688e33621e641855605c8e46ae0113371b95577d6f3

Observation 05318ce4-1021-43ad-8a6d-e17ca64811f7 · outbound

This paper cites and Chen, L.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition and Chen, L

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.843586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.870783Z digest=sha256:70764db109fc8cf685320cad20dd320a5c5e06f05cd5def0d701dfe411d87593

Observation 6db1a9ce-b900-41af-8861-827df54ad46e · outbound

This paper cites and Saffiotti, A.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition and Saffiotti, A

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.823612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.876678Z digest=sha256:f43d0ab5b12bff018408795e539504dbb35550b00f7462e44ecbde68c3bc6439

Observation 23e42ec6-a853-4166-87e9-2a1f9b80d6a0 · outbound

This paper cites and Ros, R.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition and Ros, R

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.805943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.883161Z digest=sha256:4b2b8145780afee78b46e37850d108af72acb0ba1e556af304bbb7c523b7fd04

Observation 40fe6b39-86f5-46bb-9b4d-9ed17c63b2aa · outbound

This paper cites and Deigmoeller, J.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition and Deigmoeller, J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.789418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.888879Z digest=sha256:feb14e7dd5ddfbe73b07d6053470ede45f3c555b525cfd214a929ddfe7f1094a

Observation e14c0a4f-93e7-4c41-bc1f-3dc66ed2ff6f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.893891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.893891Z digest=sha256:2ba4d56e5d63a9a241b379ce2f4c2ca6bb2be1e75af353a27d0d6c6a4d8fd6a2

Observation 516b5a2c-9614-4e39-8142-89ae274d8c21 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:19.903289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:19.903289Z digest=sha256:6747446c722224aaedc5ac81a298d85489fa530c610cfa31340d760b8aef164e

Observation a38c89f6-c7b7-4bf7-8390-90fe92ccbfe2 · outbound

This paper cites Vila: On pre-training for visual language models,.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Vila: On pre-training for visual language models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.772519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.909474Z digest=sha256:3cb900e0d4abaeac42f87c6ce211e1285b789d700d8f5437ee10227599b26807

Observation 64ac326d-5dcd-4fd6-aa71-ce4ec4b72c73 · outbound

This paper cites Kr `‘uger, C.

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Kr `‘uger, C

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:20.754860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:19.917077Z digest=sha256:9178247513ce3e9c23d6f19fe62144e16c5eabc4f0ccc3f95db79bf97d8fd43b

Pith citing papers

No inbound Pith citation observations are available.