Pith. sign in

Paper Citation Record · LEDGER

VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2410.11623.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.11623 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:30:00.262834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T06:41:12.615349Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c3229e19-a5d5-4dad-a488-c6a4aacb021a · inbound

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios cites this paper.

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:00.262834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:00.262834Z digest=sha256:58ecde95a01c544a74bc0d6cf3631fb45c00ba18f12e19c3454ffe91327dd2e5

Observation d4e4993a-07d0-402b-ad54-f9c698aa5629 · inbound

EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World cites this paper.

EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T21:30:03.924813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:30:03.924813Z digest=sha256:2ab15b7907fa49b55cd26a53da6391d4aebd9af78fe44d19ab5bba26acf1b2b1

Observation 4b68f2c0-dcc7-4251-8541-cfa9be868bbd · inbound

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering cites this paper.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.151522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.151522Z digest=sha256:a290cc063e0f1f23fb4d1fd6d9819420fd783323d898e836a96e7a32e39cbd4b

Observation 466cd78f-6508-472d-b095-c69c619c2e8e · inbound

EgoVLM: Policy Optimization for Egocentric Video Understanding cites this paper.

EgoVLM: Policy Optimization for Egocentric Video Understanding VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.902583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.902583Z digest=sha256:8f813fd3e6d7ba9d29a8c06330e805d46708cbc2d2bcdf2424116929df0a71eb

Observation 7cbc93df-fee0-4c19-be81-59e670680f03 · inbound

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models cites this paper.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.756138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.756138Z digest=sha256:691d480c8b300ae4499bbe850701a18fd5517fb02885c3af17143f66510c5498

Observation c22ae530-88db-429e-a09c-eb3d51c257fe · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:53.068893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:53.068893Z digest=sha256:e480ae07354d6c41066a6e9b234f4764a9cf1cb716fa58a1720534ad7934bb11

Observation 55295991-c419-4bdc-8535-c022ed56cf83 · inbound

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next cites this paper.

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T05:50:27.155757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:50:27.155757Z digest=sha256:a346392fab0527efafc6370da550a259aeabeb8b56cb940dc5684589cc61a1f0

Observation 6651e0c6-564f-41b4-8796-26cfb2abd227 · inbound

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection cites this paper.

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:12.685790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T17:31:28.091617Z digest=sha256:cca5e2d0f43e1e98e18745db976eed67e868554faec3a3640b256557f0092991

Observation d7319970-128d-43ee-9474-4cdc11218371 · inbound

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions cites this paper.

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:53:04.284252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T08:49:33.107658Z digest=sha256:1e1112f0432e45a0b5de0833c7af04b46b01ed20a2b46f4a4d470082d465afe3

Observation edc05bd7-0390-45f0-82c1-7042fad95dfe · inbound

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation cites this paper.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.211777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.211777Z digest=sha256:631be7254248dcad666c0ce8f2e1a767b2a8f6a3e8617e7b52c4cd2b9d8d8183