Pith. sign in

Paper Citation Record · LEDGER

EgoVLM: Policy Optimization for Egocentric Video Understanding

As of 9 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 5 inbound Pith citation observations for arXiv:2506.03097.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03097 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:14:29.581142Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:50:27.250841Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T10:41:29.637290Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cf0adb5-e990-427e-b9ba-fcc2aabeae8c · outbound

This paper cites an unresolved cited work.

EgoVLM: Policy Optimization for Egocentric Video Understanding Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:31.912954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:26.484755Z digest=sha256:45fb75e48ca54bf83f39b277f20fbdafa15b67ee06acd395376ed2a6e771e2af

Observation 7c62262a-f97b-4a42-a49e-e7c8c4d7f46d · outbound

This paper cites GPT-4 Technical Report.

EgoVLM: Policy Optimization for Egocentric Video Understanding GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.585399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.585399Z digest=sha256:5c352789a5aa077f5860f6dc8feb0d77c648e8ff127a4756d3104c00c0058f61

Observation 268c3634-4b6b-4f1d-8dc7-fdbe3453239c · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

EgoVLM: Policy Optimization for Egocentric Video Understanding Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.665985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.665985Z digest=sha256:061e61fd9466115acb1dd5e50943b4b6d2c4388edd0f0b284cb15c8588ab0f60

Observation 6e54bb2d-fa9d-4c34-aa24-d369516496b3 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

EgoVLM: Policy Optimization for Egocentric Video Understanding Qwen2.5-vl technical report, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.747337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.747337Z digest=sha256:bf1742d0136cc8188040bfffedaf974614e6e38c5ef27e78b2fe5f5dee9fbf9d

Observation 829fa18c-862f-4e6b-8a1a-caaeada02d0f · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

EgoVLM: Policy Optimization for Egocentric Video Understanding EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.803299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.803299Z digest=sha256:8a34325ccf21b4a65c828d50217cd7a8c7f3935dc8cf491e16457289ac6356cf

Observation 466cd78f-6508-472d-b095-c69c619c2e8e · outbound

This paper cites VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI.

EgoVLM: Policy Optimization for Egocentric Video Understanding VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.902583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.902583Z digest=sha256:2e852ecb3c116bbdceec159ecc292856cf7cf54b28e678bc2e578b5c6776428d

Observation 6531419f-4afd-4eed-8a25-94de004f008c · outbound

This paper cites an unresolved cited work.

EgoVLM: Policy Optimization for Egocentric Video Understanding Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:31.779038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:26.966208Z digest=sha256:2f7b4669e89fd4bfad2e43278f03ce388210791cf5e38047ea63d1dd49e2eaa7

Observation 7281fecc-e61b-40e9-aba9-dc314e630fbf · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

EgoVLM: Policy Optimization for Egocentric Video Understanding ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.043868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.043868Z digest=sha256:a152dc222f720304943a9e7dd7b7c46b5b84767e43c1024a310b18688409d0b0

Observation fd5018df-37a8-459b-bd86-54df58edf779 · outbound

This paper cites GPT-4o System Card.

EgoVLM: Policy Optimization for Egocentric Video Understanding GPT-4o System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.136926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.136926Z digest=sha256:ae4ca7096101899a48e58fd5a6da8dafc3b23b67717a5b30504cfa992e790af0

Observation 8e1f2ebc-5652-4ef5-9a09-f499df2d6a3d · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark, 2024.

EgoVLM: Policy Optimization for Egocentric Video Understanding Mvbench: A comprehensive multi- modal video understanding benchmark, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.650068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:27.251656Z digest=sha256:ed790b9275f882cba66660a9444328a8eb720697f1466894465cbcf443e1f77b

Observation 9d089f33-2fa1-42a4-b067-705125e34387 · outbound

This paper cites Dual-Difficulty Curriculum Learning for Direct Preference Optimization.

EgoVLM: Policy Optimization for Egocentric Video Understanding Dual-Difficulty Curriculum Learning for Direct Preference Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.351094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.351094Z digest=sha256:56cbe536706b8c1193eba40f365fbd3dba940c0c66a78c6cfbfbffddf754ccb3

Observation 94afe2fa-f34c-4ffc-9b8f-dce191566858 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries.

EgoVLM: Policy Optimization for Egocentric Video Understanding ROUGE: A package for automatic evaluation of summaries

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.558170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:27.482076Z digest=sha256:d6a2eecad5e36a8ed05ee155b9e4bb1dd913e350be1fb1e44ad00e4625a8b24d

Observation 3094ddd2-5790-45dc-8488-ba0939c6f1c1 · outbound

This paper cites Microsoft coco: Common objects in context.

EgoVLM: Policy Optimization for Egocentric Video Understanding Microsoft coco: Common objects in context

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.604637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.604637Z digest=sha256:eb52eb1a391928f8a711ee995551b63f471b4cb9fa5910fba02c4897ac6b45a5

Observation 1f733e8e-e71c-4a73-b8e3-a4d59c859591 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

EgoVLM: Policy Optimization for Egocentric Video Understanding Understanding R1-Zero-Like Training: A Critical Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.701631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.701631Z digest=sha256:3bed043b7bca3fa950e4e00dbec0fa017d66b6863f526221925e2cd9a5689622

Observation 52c4be18-2d30-49e9-ba1d-c1880b371685 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

EgoVLM: Policy Optimization for Egocentric Video Understanding Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.794624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.794624Z digest=sha256:bc7e1bb4ee5189094c74ad19b74fcffd7f764fb3e776ab5213ed584d5be53bfa

Observation 32e78173-37f0-4384-9aa4-b0f083b520eb · outbound

This paper cites Reasoning models can be effective without thinking, 2025.

EgoVLM: Policy Optimization for Egocentric Video Understanding Reasoning models can be effective without thinking, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.418970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:27.897137Z digest=sha256:37e7cfac59affbd23fb44329a0064f79b480a0dd3dbbd1a4e8a31f519a445eef

Observation a91b1c91-5fe9-4fc6-8fb2-4495f31efc8e · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

EgoVLM: Policy Optimization for Egocentric Video Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.299087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:27.989782Z digest=sha256:5112bd4fa6ff43c81cda5570368863cb6b1daaa7739fea4a138301952a711dbe

Observation dd2855f9-5bca-45ea-969e-eb567c9996a8 · outbound

This paper cites Training language models to follow instructions with human feedback.

EgoVLM: Policy Optimization for Egocentric Video Understanding Training language models to follow instructions with human feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.107807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.107807Z digest=sha256:c7c2bb45e2f4a65a19ce582f6426fd408831b3fa8119b04f8d122ba1cdb050b1

Observation 08c17e3c-3a5d-403c-b440-72646bd5320c · outbound

This paper cites Medvlm-r1: Incentivizing medical reasoning ca- pability of vision-language models (vlms) via reinforcement learning, 2025.

EgoVLM: Policy Optimization for Egocentric Video Understanding Medvlm-r1: Incentivizing medical reasoning ca- pability of vision-language models (vlms) via reinforcement learning, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.178831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:28.186843Z digest=sha256:a7c9450b2d530f543041186e3177131c63525cf31ed27ccf13314960a988e2f9

Observation 298fb0c7-5478-40fc-ad63-b781ba56dcf8 · outbound

This paper cites Learning transferable visual models from natural language supervision.

EgoVLM: Policy Optimization for Egocentric Video Understanding Learning transferable visual models from natural language supervision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.041234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:28.261259Z digest=sha256:0e21f601a4f77db12e5515fae1fd8b8dfc85d53e11e45d5ae40b61720cd464ba

Observation 7742fe99-1aa1-48bf-9fde-6792242a6b3b · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

EgoVLM: Policy Optimization for Egocentric Video Understanding Direct preference optimization: Your language model is secretly a reward model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.361137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.361137Z digest=sha256:47687019e482f5a74bf532bb96277302cc79b05b81b17a2d25472270169ba819

Observation f9c58009-197e-4dfe-8e46-e8a0dc3e55c6 · outbound

This paper cites Proximal Policy Optimization Algorithms.

EgoVLM: Policy Optimization for Egocentric Video Understanding Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.448038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.448038Z digest=sha256:89d4d9b37d5d8d49858065d29afedc64de901f1c8da852f66eb01a0e1c65ea05

Observation e43a7e88-60ed-4972-a6b7-ff2fe5ebbd17 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

EgoVLM: Policy Optimization for Egocentric Video Understanding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.572354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.572354Z digest=sha256:5baa594838af711c7782d55f13f367bcb4ec64741d1b46a6567d986112847980

Observation 42623fe7-8624-4ba8-aad3-ed55fcf989cf · outbound

This paper cites an unresolved cited work.

EgoVLM: Policy Optimization for Egocentric Video Understanding Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:30.870663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:28.653266Z digest=sha256:6f43b0562930345275a91c26fc583cc6a8ff1fa237493273088a532d7b531f16

Observation 071969d7-a674-4a0c-ae25-d3cf448472a5 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

EgoVLM: Policy Optimization for Egocentric Video Understanding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.744187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.744187Z digest=sha256:452652e13e82cc7c11b2c9d6d2490654c8097cb578dcdcfcb15f2de3b2d77ad2

Observation 3badcadc-13ad-476b-9356-b973be507cc3 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EgoVLM: Policy Optimization for Egocentric Video Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.861844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.861844Z digest=sha256:fa6b7ef000f2df6149d34267ab28063a605854dcf07b5b7cb7dd2bc55fff01b9

Observation e7ae7185-273b-42c0-a295-50a794bcd589 · outbound

This paper cites Internvideo2.5: Empowering video mllms with long and rich context modeling, 2025.

EgoVLM: Policy Optimization for Egocentric Video Understanding Internvideo2.5: Empowering video mllms with long and rich context modeling, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.924627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.924627Z digest=sha256:7fabe52050761b3e904024fb2061b30a65cf095cdcde8c6a92f303e1b448f555

Observation 350801aa-6cc5-4d4c-a2e3-37346c54cf80 · outbound

This paper cites Wizardlm 2, 2024.

EgoVLM: Policy Optimization for Egocentric Video Understanding Wizardlm 2, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.711275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:29.032858Z digest=sha256:a52ec014bdf89ce7e50cd909c08b38aba13ffffdbf6be15f5581afec36d3d240

Observation 7825ea91-ba7e-42fb-a506-0c2806d66248 · outbound

This paper cites ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos.

EgoVLM: Policy Optimization for Egocentric Video Understanding ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:29.112728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:29.112728Z digest=sha256:bb06edd1c905bc307f8b60cdf96547f4fd94f1205a599a3afab69b44e5abd998

Observation 819dae82-f2af-4964-b142-c92a2946656e · outbound

This paper cites Egolife: Towards egocentric life assistant.

EgoVLM: Policy Optimization for Egocentric Video Understanding Egolife: Towards egocentric life assistant

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.571100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:29.229419Z digest=sha256:ff3968a0478bc85ebcebde16a2f0f6361bb10a8a35444d70e7b4581e66d2d6da

Observation f48211bc-4251-41c3-bd38-c7686456d4e3 · outbound

This paper cites Mm-ego: Towards build- ing egocentric multimodal llms.

EgoVLM: Policy Optimization for Egocentric Video Understanding Mm-ego: Towards build- ing egocentric multimodal llms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.450637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:29.289694Z digest=sha256:4a4ba10f3063a6825cf3224b15022160ade7a3ba9245ab102a402ca06b8b9589

Observation 2a23480f-3d85-4f22-964e-c9c8e56192d7 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

EgoVLM: Policy Optimization for Egocentric Video Understanding Video instruction tuning with synthetic data, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:29.397148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:29.397148Z digest=sha256:64a4d8421624d5947214bbd823e5bcfc57aefead10e99639b2a76ebf2182c6e5

Observation 6781ea2e-d1f7-4430-a462-1821dc6db063 · outbound

This paper cites Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els.

EgoVLM: Policy Optimization for Egocentric Video Understanding Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.310972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:29.444276Z digest=sha256:73bfbaab7fe401be746184afbe12ca5bfc69758586dbe3f7eddf9f7b964b6956

Observation c3d95f10-d7c8-4aaa-b7c9-c88f519f0f44 · outbound

This paper cites R1-zero’s ”aha moment” in visual reasoning on a 2b non-sft model, 2025.

EgoVLM: Policy Optimization for Egocentric Video Understanding R1-zero’s ”aha moment” in visual reasoning on a 2b non-sft model, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.161229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:29.544152Z digest=sha256:825e58d65df40de8ec23285c9ed2767d0547083d2d06117701be3c9d3a14fc3e

Observation da53d852-81dc-4541-ae7c-4a469c6a3fa0 · outbound

This paper cites Egotextvqa: Towards egocentric scene-text aware video question answering, 2025.

EgoVLM: Policy Optimization for Egocentric Video Understanding Egotextvqa: Towards egocentric scene-text aware video question answering, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.027639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:29.581142Z digest=sha256:49e008c4a050b2fa2062b48711be8b7cd3e3b39501627ef7a540cb58b885c3b0

Pith citing papers

Observation 358e3010-f86b-4f49-a0c8-0dcd3bb9fe7c · inbound

EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning cites this paper.

EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning EgoVLM: Policy Optimization for Egocentric Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:52:51.049418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:52:51.049418Z digest=sha256:1339194fcf40f7a7f936aebf9da16ed5f60c54752da4abd77952a2706d1b2688

Observation db8a76be-cd62-4481-9ca2-4f1014065a3b · inbound

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next cites this paper.

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next EgoVLM: Policy Optimization for Egocentric Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T05:50:27.250841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:50:27.250841Z digest=sha256:213fc07518a32f7fadefffa98d06b12865e1d384e7f31f24083b323b28d53e3f

Observation 1fdb557d-46fb-422e-b3b0-ef98a0901133 · inbound

Robot Learning from Human Videos: A Survey cites this paper.

Robot Learning from Human Videos: A Survey EgoVLM: Policy Optimization for Egocentric Video Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:41:29.640044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T04:55:44.273643Z digest=sha256:18a9d98660a62e2a974c0583697b651d586449a8fd5c5ffa469535eef2f3f2c0

Observation 6dbe324b-7bf7-4845-adb5-069be76a7df9 · inbound

Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks cites this paper.

Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks EgoVLM: Policy Optimization for Egocentric Video Understanding

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:06.581439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:16:31.820718Z digest=sha256:63eb018af6daa5cd9d75a049342ea4da87074a3e19c8575c13604fd3cb4ad2e9

Observation 2ed5ebf4-0079-4353-a6bd-f0edb17c0757 · inbound

Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks cites this paper.

Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks EgoVLM: Policy Optimization for Egocentric Video Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T05:19:56.255880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:19:56.255880Z digest=sha256:a5d5ff8f63b76120a64cc95f3a2826a84a779c94a67b8985e2b5916a4f289444