Pith. sign in

Paper Citation Record · LEDGER

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models

As of 21 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2504.13351.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.13351 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:15:06.119010Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T22:31:13.391859Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T22:32:10.973323Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b5e9176-f274-4bf5-8f78-19c8935e8aaf · outbound

This paper cites GPT-4 Technical Report.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.823996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.823996Z digest=sha256:d257ae69d29b1e9707b4985eea5b5d2d9b967cecf3ca9d415cd23afe34bb5976

Observation 953d85a8-f64c-4f59-9276-8907087ead28 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.829473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.829473Z digest=sha256:2b01a288fcee070c0ae180bfe82c951af9b6191bdedff2f7287f7a563161fec0

Observation 1b7c2986-8ea6-40be-ac2a-4df0580fb00d · outbound

This paper cites Human-to-Robot Imitation in the Wild.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Human-to-Robot Imitation in the Wild

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.834129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.834129Z digest=sha256:63d684baf15ea25033f0ca421b2f5abf34cc6abe5425e36587410b1e83dcaf4a

Observation c24c9671-cfc9-4b7a-a0f7-8eda2e820442 · outbound

This paper cites Affordances from human videos as a versatile representation for robotics,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Affordances from human videos as a versatile representation for robotics,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.838737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.838737Z digest=sha256:b3042d80cd7cf69879f02bb96e1609c5bf9c13c980f3ad3c24d67c61de457c97

Observation 6566c015-5722-4ca1-9442-d503edf5b012 · outbound

This paper cites Towards generalizable zero-shot manipulation via translating human interaction plans,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Towards generalizable zero-shot manipulation via translating human interaction plans,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:07.051450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.843956Z digest=sha256:4a4c324144b3d7f53a29e5d9967ae9a3a4b3d13362c35ec48895e92d3eb323dd

Observation f66ebc5d-31b3-4d1c-ba56-abeb9bd5b526 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Activitynet: A large-scale video benchmark for human activity understanding,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:07.037710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.849102Z digest=sha256:f481b9119b0f26f9159330a063c1569589dbad7fdb304f21975dd8a25fb422d8

Observation a6de1d3a-4ac8-4cc1-a888-30e335918dbc · outbound

This paper cites Procedure planning in instructional videos,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Procedure planning in instructional videos,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:07.023693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.854359Z digest=sha256:cc6d0bc7482fedcce7a20f61e497aa77109a0946dfa8a5fb1f2f7b4220ea5e33

Observation d1a7a79a-8866-4e3c-bfaf-b82b29ebe58d · outbound

This paper cites Learning generalizable robotic reward functions from.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Learning generalizable robotic reward functions from

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:07.010164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.858352Z digest=sha256:d6844637d62f61ee5d771935039f2a2dcec0859c170df51f147bdf085145b0cd

Observation 1ec69fac-60e0-488d-806f-08500ddda4d0 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.996361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.862463Z digest=sha256:83732c1a073e06b79b272cd80c5438f798b66f0ae2619038bc90f733d98482ed

Observation 8e045913-0882-4d8d-a633-ff8ff2c3ea22 · outbound

This paper cites Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.866636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.866636Z digest=sha256:04e2357a071252f2055eeff982154bb3eca9ee7df45c7a58e22f4c58fd74993d

Observation 562bd8c0-423d-4c76-af4c-4b8d28ab45b8 · outbound

This paper cites Can foundation models perform zero-shot task specification for robot manipulation?.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Can foundation models perform zero-shot task specification for robot manipulation?

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.981665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.871228Z digest=sha256:9d3bccdc0644040ababf93ebc0e5ac425d5549d19569f9515b5410ce8696f361

Observation 51339741-d041-4973-b24c-15288bb4f70f · outbound

This paper cites Scaling egocentric vision: The epic- kitchens dataset,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Scaling egocentric vision: The epic- kitchens dataset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.967585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.875734Z digest=sha256:6a2d56093219400e227340efaf2900574b38a3d1e2f1770b9c591e5d56188e25

Observation 1b415199-27db-4206-a35b-1a90ea1d0228 · outbound

This paper cites Model-based inverse reinforcement learning from visual demonstrations,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Model-based inverse reinforcement learning from visual demonstrations,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.954253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.880003Z digest=sha256:7aa8823edf88b34116aba9c27a7b8f7a7a00941f60f64e3cb6b9adce719c433e

Observation 07a1c60b-2f91-4c1c-b40c-78ee1eb6c582 · outbound

This paper cites Perceptual Values from Observation.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Perceptual Values from Observation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.884007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.884007Z digest=sha256:eef2f4aead82a06c6470c84ebc276614e225a2927a28333ca8c007ad3778c18f

Observation 2363f667-350e-4d73-98b8-e7f87053b07d · outbound

This paper cites Foundation Models in Robotics: Applications, Challenges, and the Future.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Foundation Models in Robotics: Applications, Challenges, and the Future

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.888688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.888688Z digest=sha256:f5b56f8dfb9447c2a6d69380397945b02c7dffb6af3638a862ca473aa1b8d010

Observation 73711770-4d0f-494a-ad7c-18cacc6d290a · outbound

This paper cites The” something something.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models The” something something

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.940935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.893962Z digest=sha256:1a65acd49475a40e9813a5b8ff4f86ad4ceaf2078f0608b12645a30f11d9d658

Observation aba618a8-bbf5-45a9-a975-2baf93792f60 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.928077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.898804Z digest=sha256:cf955593d9007e2c1de166c3e5abf5a4fc533356f008a7c80130277647d32b50

Observation 332cdff3-e8e4-46a1-a16e-80336f017610 · outbound

This paper cites Gu et al., Rt-trajectory: Robotic task generalization via hindsight trajectory sketches , 2023.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Gu et al., Rt-trajectory: Robotic task generalization via hindsight trajectory sketches , 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.903048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.903048Z digest=sha256:783fb1b039c2e0ea7736b412739d327f0f9c2ac93e930cee30d2caefe9656779

Observation 9df6f5cb-68e4-4f18-ae1b-b25a61b06d3e · outbound

This paper cites Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.907167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.907167Z digest=sha256:379ca13f62979691d2f4944719b1e4e61209119bda8e4fbaaf48652d01981e3e

Observation 153673ef-0b34-408b-aea2-b7b00219c727 · outbound

This paper cites Neural task graphs: Generalizing to unseen tasks from a single video demonstration,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Neural task graphs: Generalizing to unseen tasks from a single video demonstration,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.906112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.911401Z digest=sha256:7e004dd8d28a2c96cd4352e346e280dc07d4b51914e4fe85b85efc38acb91295

Observation 4e09e530-64a9-4abf-bffc-429cbe81a950 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.893030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.915400Z digest=sha256:cd66fbdd123fde81166a6d867a42376903bed9b6e163c422880cf86884a460ad

Observation ad73f741-e752-4740-8e4b-86eedd14fd18 · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.920527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.920527Z digest=sha256:01854ee13b07db8bc936d4a1d7836c115a7809b5e192770ff629af3b26d655eb

Observation 626817ea-66c6-4e8f-bc38-27a963ee36c8 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.924813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.924813Z digest=sha256:e3d50e09fdee1c761ab87a2d8542a45cf27a9c2b94b91ca5fce50909957e6844

Observation dea3d730-a29c-4b51-b11c-129b69886541 · outbound

This paper cites Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.929096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.929096Z digest=sha256:7568ba7255e79c5aeba31d91f81fb8afa027f34911aa9ebb296e44c9d27faf08

Observation ec8b220b-43bd-46c8-8859-31c04e85b903 · outbound

This paper cites Prompting visual-language models for efficient video understanding,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Prompting visual-language models for efficient video understanding,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.879575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.933595Z digest=sha256:94086f05e4582b729a3974329aba872764c4d15df8b23710ebc27b2c87cfc1dc

Observation 0f4db8b3-28f7-406b-9960-8089c67be071 · outbound

This paper cites Human action recognition and prediction: A survey,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Human action recognition and prediction: A survey,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.865923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.937756Z digest=sha256:6bc1c2d39438c6191410e9e18efd294ff3f8cab4f8a8e51f5880ca53070258f7

Observation 7ffb5d53-ac82-452c-b8ac-cc0832839844 · outbound

This paper cites Anticipating human activities using object affordances for reactive robotic response,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Anticipating human activities using object affordances for reactive robotic response,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.853103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.941701Z digest=sha256:e6281c28318ce809a20e5231990d614780106a13f1fe508d4279e73ee2014bca

Observation edf41b8e-010d-427b-9192-4445b94ad7da · outbound

This paper cites Graph inverse reinforcement learning from diverse videos,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Graph inverse reinforcement learning from diverse videos,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.840008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.945640Z digest=sha256:ffe5c2336e785e4c19fb5e91d383b2bbbb3c59630cd1ba01cf5678afde993d82

Observation 524e4606-1f1d-494e-a216-e69ba2bf4290 · outbound

This paper cites HAKE: Human Activity Knowledge Engine.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models HAKE: Human Activity Knowledge Engine

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.949799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.949799Z digest=sha256:7fd9483107baf97b51dac7c19dd5a695f71601a628225faa3348c26060ec4b9d

Observation 43e694a3-e04c-4a4c-9f6f-834dba6310ba · outbound

This paper cites Code as policies: Language model programs for embodied control,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Code as policies: Language model programs for embodied control,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.826654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.954444Z digest=sha256:d997193fde4c46652ce8ece6d81a972e474ab199bb6f19758c9e5551f5bea47a

Observation ddfeea2d-bcaf-4cdb-8549-2177c55fee32 · outbound

This paper cites Learning to Learn Faster from Human Feedback with Language Model Predictive Control.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Learning to Learn Faster from Human Feedback with Language Model Predictive Control

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.958643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.958643Z digest=sha256:1dc8dc5bdc335fd4c4e161d288029eb8c4cd234850eec0c5019aa834c18cb720

Observation 029c77d3-91c1-44d6-a7fb-7be115bd0f9d · outbound

This paper cites Text2motion: From natural language instructions to feasible plans,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Text2motion: From natural language instructions to feasible plans,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.813390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.963232Z digest=sha256:4792a1614d62321d79757b9dd0ac81456afb8220943c4d3327cc9906ffb13c65

Observation 8334d377-0654-4f39-8059-f769e0eaf7d6 · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.967780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.967780Z digest=sha256:e9a66d940edb9fbe6e590a068c8124660241395eba30d79d7c72cb2603e9911a

Observation af4a4e74-bada-4530-8546-c78ec375f196 · outbound

This paper cites Imitation from observation: Learning to imitate behaviors from raw video via context translation,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Imitation from observation: Learning to imitate behaviors from raw video via context translation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.799915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.972747Z digest=sha256:58ad3b74c3914d0f3c2d597d17019e4d67159f14ed718a773afb3927a69bd995

Observation 62ae18b9-ac15-436c-b5d5-dc9e7809140d · outbound

This paper cites VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.977805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.977805Z digest=sha256:cb8593259e44d2cee1ea8916e49c9cc349ce85d9e8209636f92b67955ef6e8aa

Observation ca246feb-1907-4498-abb9-91b4e55c323f · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.982335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.982335Z digest=sha256:8cefd49d997c0ecff5e67baeb9b72c308b9552e18d425f3de5cc720db8318632

Observation 8012cf00-5557-49e5-a3d7-f5eac04cd3be · outbound

This paper cites R+X: Retrieval and Execution from Everyday Human Videos.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models R+X: Retrieval and Execution from Everyday Human Videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.986843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.986843Z digest=sha256:ec34fd91f23c6d91bf67550dd20b716c96cb1985372099ee46fe94c37dbb8c30

Observation ecf863bd-b2cf-4e7f-92a8-d099dd57f61f · outbound

This paper cites Reconstructing hands in 3D with transformers,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Reconstructing hands in 3D with transformers,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.785927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.991085Z digest=sha256:65b94c07387498eda2d8d5f2cbfe981b9676470c6ae815dc5f528b43b28afa3e

Observation 3d0306c3-cfbf-44a0-ad5e-08931ec73954 · outbound

This paper cites Planning with large language models via corrective re-prompting,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Planning with large language models via corrective re-prompting,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.772228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:05.995264Z digest=sha256:8c90649918bfb7f80877ce6c40a84116099969a39b432d5b6b56e446d8511021

Observation ce308fe6-2918-458c-a391-e35be8176faa · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.999412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.999412Z digest=sha256:1575aa27f46af7a9ba4828b97f0c0a05f30828630800de58a6b9e42cec1aa96b

Observation 494da36c-0ae1-49eb-b631-6315d0951129 · outbound

This paper cites First-person activity forecasting with online inverse reinforcement learning,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models First-person activity forecasting with online inverse reinforcement learning,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.756810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.003775Z digest=sha256:6ec0001c2a7feb4f25a9ad45c8f7638d899c97db1b081636458cce4597e51ebc

Observation 35af371c-7f12-42ac-9df6-63f1efbfe7c0 · outbound

This paper cites Reinforcement Learning with Videos: Combining Offline Observations with Interaction.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Reinforcement Learning with Videos: Combining Offline Observations with Interaction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.008300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.008300Z digest=sha256:a22c83e6a01a6a6748316b390c46d82f412c8910b35a603d13107666d4b507a3

Observation c017f492-57b7-4e75-bb03-e249e5bd876e · outbound

This paper cites Learning predictive models from observation and interaction,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Learning predictive models from observation and interaction,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.743315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.012623Z digest=sha256:15c46ac57f9ca0514b564a28db80f04197bd75183bc6a842063b0c6cde1dea68

Observation fbfba823-2a36-4ea3-a700-7213cf749988 · outbound

This paper cites Time-contrastive networks: Self- supervised learning from video,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Time-contrastive networks: Self- supervised learning from video,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.729483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.016846Z digest=sha256:b8b9f8bb6c776935b9bd87e940a8dd345379e10e462afbd90f99de30fd889d82

Observation 5936cc87-13c4-47de-8a41-75dc2ba438d8 · outbound

This paper cites Unsupervised Perceptual Rewards for Imitation Learning.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Unsupervised Perceptual Rewards for Imitation Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.020797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.020797Z digest=sha256:85cf690ee714a853c35b08ca0497f93cb0f7dd6d471a414f9504192df8c31de4

Observation 5d5cca8d-92fe-425d-bbe4-94d46b2f6314 · outbound

This paper cites Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.716265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.025092Z digest=sha256:6cdd2fbc5d1554bed1d24e9e135fcd74ec9bb9562cfeb67b44931508da0995d3

Observation 53ea4167-f4bc-4391-9a69-c5eb279eaa1c · outbound

This paper cites Concept2robot: Learning manipulation concepts from instructions and human demonstrations,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Concept2robot: Learning manipulation concepts from instructions and human demonstrations,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.702427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.029196Z digest=sha256:e300ef767c5855e5215699965b41594da24bc84e129282b8c6d30094dc2dd216

Observation 64ac1d08-59bd-40d4-ae6d-1577697185fd · outbound

This paper cites Third-person visual imitation learning via decoupled hierarchical controller,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Third-person visual imitation learning via decoupled hierarchical controller,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.033284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.033284Z digest=sha256:a055549c8e52a7e0812e5958731aa1908be663988aa3e3ea5a1443d1523c523a

Observation 7fdefc1e-e9de-46d6-b321-ccb76c1d509b · outbound

This paper cites Videodex: Learning dexterity from internet videos,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Videodex: Learning dexterity from internet videos,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.679337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.037238Z digest=sha256:ac44f35d58f99824037f524bf27398fe95155e96f70344f174fd49d2e79e0c28

Observation 9cbb0e9d-2e59-4dce-b269-e94561a12242 · outbound

This paper cites Cliport: What and where pathways for robotic manipulation,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Cliport: What and where pathways for robotic manipulation,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.666143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.041470Z digest=sha256:5c2a3ec596217ac76c2fad6859cae7aa3777434b7338b016f8bda35cdfabea5b

Observation b100e8b7-b69d-47f4-8f2b-9c497e515b8d · outbound

This paper cites Generalized planning in pddl domains with pretrained large language models,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Generalized planning in pddl domains with pretrained large language models,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.652030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.045413Z digest=sha256:412102ba2073b0f25f451fc63d62bc8a8d566323319331b14415ae696bfc56db

Observation c8ce1d2a-f309-4ba7-9fbe-079e8c5f9771 · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Progprompt: Generating situated robot task plans using large language models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.638380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.049345Z digest=sha256:ba79fbfcd2adf719c19d036d42287f91648b9df0f3323451e46f62a044f83ef6

Observation 5e08acf3-1ab6-4dbe-b8fb-d79b87d0f243 · outbound

This paper cites AVID: Learning Multi-Stage Tasks via Pixel-Level Translation of Human Videos.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models AVID: Learning Multi-Stage Tasks via Pixel-Level Translation of Human Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.053396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.053396Z digest=sha256:b7119b216e13762bd63a269c3c4fba125d95a7d12e679591dc3b9b6d9125d31e

Observation bae1268d-aecf-4e21-afd8-befc75498813 · outbound

This paper cites Rt-sketch: Goal-conditioned imitation learning from hand-drawn sketches,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Rt-sketch: Goal-conditioned imitation learning from hand-drawn sketches,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.624307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.058134Z digest=sha256:adaa47aea2dcccfd40808e97be74054494caa4a1e286554e5af84b014b1c453e

Observation 70c2604c-d6d6-4a9b-820f-01e4eaac77c2 · outbound

This paper cites MimicPlay: Long-Horizon Imitation Learning by Watching Human Play.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models MimicPlay: Long-Horizon Imitation Learning by Watching Human Play

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.062178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.062178Z digest=sha256:4d7372ae87b5d769210b7e53c5f208d384a7c6b1a9c9bb555a95806059431fc8

Observation e029ff48-0892-4c15-ae94-e09bb9b4a6a5 · outbound

This paper cites Temporal segment networks for action recognition in videos,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Temporal segment networks for action recognition in videos,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.609454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.066608Z digest=sha256:156d4e1eed8be1aed13a6118df5c0810918a89b1138aebb092d863268967b257

Observation ccfc0c3a-e781-4fc8-b1e8-42124a900da9 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.070735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.070735Z digest=sha256:89854ac206c1c1256e67c214afa404c48be71dbe33759351470ce83beb9e1f17

Observation 9bad7664-3736-4466-96e9-284f630f4666 · outbound

This paper cites You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.075484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.075484Z digest=sha256:023f74d3e69da9a3e66f53deed3da13e2c8cb7def2022f6ee34e4e0508c5322c

Observation 5a8437c9-a347-4a60-bd11-f753709bfdbd · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Any-point Trajectory Modeling for Policy Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.080003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.080003Z digest=sha256:8fe1beb171bb0ecd8078321dcc125946af7dbf81bfc58261036c4211aa2d1659

Observation 0fc53998-4711-45aa-89a6-21b936be4c15 · outbound

This paper cites Learning by watching: Physical imitation of manipulation skills from human videos,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Learning by watching: Physical imitation of manipulation skills from human videos,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.594823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.084224Z digest=sha256:d5d5f9d6334b42960b74dd75f0138556added68ddc99def33602bc59c0f7535a

Observation c7e02b1c-fd93-45ed-83c9-62b7ff4b66f5 · outbound

This paper cites R-c3d: Region convolutional 3d network for temporal activity detection,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models R-c3d: Region convolutional 3d network for temporal activity detection,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.580137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.088928Z digest=sha256:e189bf84bcefd235d544abe4db37ed05fb1411f3fe508ce945dc676cd9f75e57

Observation 543b7814-6507-4f88-a2f9-a6cb528eb278 · outbound

This paper cites Xskill: Cross embodiment skill discovery,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Xskill: Cross embodiment skill discovery,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.566205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.092975Z digest=sha256:9ef540dd5b6648c2931bead90ca864633d3e82ecfbdbd0cae808d5bdadc7adcd

Observation b81a5940-fb40-436a-aafb-63f13630f2c4 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.097274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.097274Z digest=sha256:bb8812516b8c7908ee7cb04d3ea6a884cc7345cef9fa32462bea8813a23229ed

Observation f5a88b61-30fd-43b7-85df-92a144c2dde7 · outbound

This paper cites Language to Rewards for Robotic Skill Synthesis.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Language to Rewards for Robotic Skill Synthesis

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.101518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.101518Z digest=sha256:fa44a33c22135d17f55f40051a2d26870248a60d024693b91c52ef0b35cf51ee

Observation b4d1615e-b21a-4e42-8b67-7f36d58670b9 · outbound

This paper cites Xirl: Cross-embodiment inverse rein- forcement learning,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Xirl: Cross-embodiment inverse rein- forcement learning,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.551893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.105941Z digest=sha256:166be06f238e87147ab67d9d5ff69754475e6c1cb17e0d9fa5af00273409df5b

Observation 8f7de078-182a-49d8-938e-d4883ebcce9f · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.110419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.110419Z digest=sha256:3b8d9a99508de2b0897d131747aad90e5842bbb3656a07a08bda5e1bb202baef

Observation dda92637-4a19-45f2-bcf8-3250d59a31c1 · outbound

This paper cites Actionformer: Localizing moments of actions with transformers,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Actionformer: Localizing moments of actions with transformers,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.538148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T12:15:06.114777Z digest=sha256:182deaa09e84a2eb0f2a3b5b85e80246098b809efb69804e8700272405261bfa

Observation 530ccaf9-7f76-4434-9786-94972fb0ade5 · outbound

This paper cites Vision-based Manipulation from Single Human Video with Open-World Object Graphs.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Vision-based Manipulation from Single Human Video with Open-World Object Graphs

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.119010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.119010Z digest=sha256:6b8ead37087f062344d805c3c0803f6be5e2bc758fd65cae135e7fd8da98ebf3

Pith citing papers

Observation 4aaef4b5-f425-411a-b59c-31eadb2ff0d2 · inbound

Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views cites this paper.

Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:32:10.976303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T22:31:13.391859Z digest=sha256:dfe99e83c15c9a46c06b22e6e62a974e260b3071e63fc14fcbbc7efc04595d9a