Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:53:19.917077Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2506.20373.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:53:19.917077Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 522050a9-df17-427e-b02e-1a90539db50f · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Large language models for human–robot interaction: A review,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb58932b-dda3-409d-b3c1-e5a7b9db7f3a · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a67e63f2-50c3-4244-b6a0-ffe57fa8d91f · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition MUTEX: Learning Unified Policies from Multimodal Task Specifications
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7e3ad0e-3de1-4113-b5b5-6c1193f37026 · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Vision- language model-driven scene understanding and robotic object manip- ulation,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2dd8171-240b-4c6d-965e-dc82a58a169d · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition CoPAL: Corrective Planning of Robot Actions with Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7d07ab79-1a41-4486-b734-31d71a195f24 · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition VLM-Social-Nav: Socially Aware Robot Navigation through Scoring using Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baffae8f-7b45-4d4f-8380-a6bb2ed302bc · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fcfe77a-5771-4099-aae6-c6b2b638058f · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 932c51cd-ff69-4386-ae63-34c696512b8d · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition LaMI: Large Language Models for Multi-Modal Human-Robot Interaction,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 406d83e2-0e16-451a-aed4-c583e86af768 · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce2a863c-e62b-4d0f-a10e-5e61b1b19f45 · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8c86f276-7e64-4c68-be46-d2b17b4e4b8f · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Robots Can Multitask Too: Integrating a Memory Architecture and LLMs for Enhanced Cross-Task Robot Action Generation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0cf753cb-4b92-40f3-8b1a-40e2c1f0132d · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition “Exploring large language models as a source of common-sense knowledge for robots“
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 85859591-6b94-439b-b285-fd18f478b147 · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0d7d1a81-4741-41e6-8852-3370bab662d9 · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e61757eb-223c-4dd2-8d87-f9d87ac45006 · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition “Quo vadis, action recogni- tion? A new model and the kinetics dataset.“ Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 29d4c07f-d1b3-4335-b8e8-f7e59a6d1aba · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba1ec72c-99c8-49a7-96ef-4e90884b8175 · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Pixtral 12B
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baf6a17e-7231-4b82-9dfd-887a732ae259 · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6299b6e2-5fb4-475f-a796-8ce59e463f18 · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition LLaMA: Open and Efficient Foundation Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5f813b-6236-424a-a817-6f9a4bdba0ad · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition and Kembhavi, A., 2020
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 05318ce4-1021-43ad-8a6d-e17ca64811f7 · outbound
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6db1a9ce-b900-41af-8861-827df54ad46e · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition and Saffiotti, A
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 23e42ec6-a853-4166-87e9-2a1f9b80d6a0 · outbound
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 40fe6b39-86f5-46bb-9b4d-9ed17c63b2aa · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition and Deigmoeller, J
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e14c0a4f-93e7-4c41-bc1f-3dc66ed2ff6f · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 516b5a2c-9614-4e39-8142-89ae274d8c21 · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition LLaVA-OneVision: Easy Visual Task Transfer
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a38c89f6-c7b7-4bf7-8390-90fe92ccbfe2 · outbound
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition Vila: On pre-training for visual language models,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 64ac326d-5dcd-4fd6-aa71-ce4ec4b72c73 · outbound
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.