Pith. sign in

Paper Citation Record · LEDGER

AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2406.13807.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.13807 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:38:43.950532Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:35:04.661111Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation da1fcaa3-c76b-4cb6-a7b7-28e65f9900a1 · inbound

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent cites this paper.

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:43.950532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:38:43.950532Z digest=sha256:325afe4b0f15321f0ea099304b8daa7882cf0d8026b90d69d35d6c813122cb8f

Observation 295e2f7d-6221-4401-874a-24cfa57b4430 · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:43.060746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:43.060746Z digest=sha256:e348e429b600deb1ac837beb2d88f3878e128c2623fcf310b0611f707710283b

Observation ea651a9d-1b27-4bff-8db3-46c22c15aa32 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.341363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.341363Z digest=sha256:5ba7fd44767093cebcbd69de1270f4ec54c2de2987e17c64a00d2825fe451c45

Observation e4e6362a-cf70-4623-b7e9-d766c9cd45bd · inbound

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection cites this paper.

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:13.433486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:31:28.091617Z digest=sha256:b974d4877da662e7b9da14e18c5b3d8e9d8a0ffed3672d92e931049a0f52b346

Observation 8a1deaec-c68e-47f1-987f-0db9b217f58b · inbound

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding cites this paper.

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.662610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:29:27.063028Z digest=sha256:c326f8b1c1a5f71c86f453f5717e92dc59a9252ed578b48877495fb5b3668f16

Observation 8f5c2a54-2c1c-46d1-a50e-8ff3b43a3dcf · inbound

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation cites this paper.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.186399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.186399Z digest=sha256:8dcad06d2404dfc18f999d781181f34b6e53e74fee1fca59f6383f78926f96d6