Pith. sign in

Paper Citation Record · LEDGER

Video as the New Language for Real-World Decision Making

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2402.17139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.17139 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:00.956992Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:57:27.914181Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cb68f435-9ef9-496f-9a99-9d4e3f324fae · inbound

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation cites this paper.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Video as the New Language for Real-World Decision Making

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.397994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:6e2c137702988287d61c7e5bb260cf533b841d415815201ff3d06dfa2b5ff8a6

Observation 01be40f5-6ffc-4b25-a2d8-0c8a4b18caf3 · inbound

Generative World Explorer cites this paper.

Generative World Explorer Video as the New Language for Real-World Decision Making

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T18:09:24.602792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:09:24.602792Z digest=sha256:f2abb25669ff02234959ffa8c4dfb761d4baec5657cc194039e1f0c37a7fa098

Observation a447bdb1-2a06-4ba4-b034-53ae484edf1f · inbound

Motion Prompting: Controlling Video Generation with Motion Trajectories cites this paper.

Motion Prompting: Controlling Video Generation with Motion Trajectories Video as the New Language for Real-World Decision Making

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:09.918380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:09.918380Z digest=sha256:6ed03d64971338ffdf83c6ab5f16ad88a5cb7f323b14592b86b8862b2e29067f

Observation 33864465-2512-42e9-8726-982687fcfe7f · inbound

GenEx: Generating an Explorable World cites this paper.

GenEx: Generating an Explorable World Video as the New Language for Real-World Decision Making

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:55:49.296224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:55:49.296224Z digest=sha256:0fe91d167bd38f0a25a7b399e986030de5cab60a8a9a4fede54c998a0de2aad6

Observation 8e15a14c-8dc4-4df3-b87f-5034072c8e3d · inbound

A theory of appropriateness with applications to generative artificial intelligence cites this paper.

A theory of appropriateness with applications to generative artificial intelligence Video as the New Language for Real-World Decision Making

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:06.681901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:06.681901Z digest=sha256:817100e57fe6b071f133a4fe9c5a748b2a9f4694fe15401d84f303c798e3885f

Observation cc8839f7-62bd-437a-92f1-1ae556bb0999 · inbound

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos cites this paper.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Video as the New Language for Real-World Decision Making

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.110723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.110723Z digest=sha256:7820accf0b6c294369e3858ffb26552bcb51ab9e1ba9774ea29dda67e4cfed98

Observation 0f5792e0-2ac2-4ada-b082-a272046578aa · inbound

Generative Physical AI in Vision: A Survey cites this paper.

Generative Physical AI in Vision: A Survey Video as the New Language for Real-World Decision Making

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:00.136358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:53:00.136358Z digest=sha256:42e41b23a4b5c7a00f5d0b809403d043f13345e2a2939f644e8eca8236178666

Observation d9e99fea-0c5a-4df0-bae7-857ae8ae0f22 · inbound

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models cites this paper.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models Video as the New Language for Real-World Decision Making

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:21:45.130034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T05:21:44.903048Z digest=sha256:3a16f177b49678246a6be7eccc5379e2aa78fea07976977ede839ec6b564385f

Observation cb130f2f-0356-44c7-9d0f-3ae20fa53072 · inbound

Solving New Tasks by Adapting Internet Video Knowledge cites this paper.

Solving New Tasks by Adapting Internet Video Knowledge Video as the New Language for Real-World Decision Making

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:00.956992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:00.956992Z digest=sha256:c80314158c9ce4e19a374f525d9f0766e0fd32735358de018c4b840611fbfa03

Observation 0d826409-6c2b-487e-b4b8-b08728619982 · inbound

Latent Diffusion Planning for Imitation Learning cites this paper.

Latent Diffusion Planning for Imitation Learning Video as the New Language for Real-World Decision Making

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T10:56:16.059353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:56:16.059353Z digest=sha256:e214e51eb400910de899373874a00ae371701bef0e49a1e034961073cde0038a

Observation 55d0f200-1768-4633-a8f8-cfeef38356e1 · inbound

Video-GPT via Next Clip Diffusion cites this paper.

Video-GPT via Next Clip Diffusion Video as the New Language for Real-World Decision Making

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:30.562182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:30.562182Z digest=sha256:c39b280c97a36ca40106f5f62c9931f1ea0d919533d33e975d6cb58d1dfb90b8

Observation b9323015-c13c-4360-b99d-9cbe805562ee · inbound

Humanoid World Models: Open World Foundation Models for Humanoid Robotics cites this paper.

Humanoid World Models: Open World Foundation Models for Humanoid Robotics Video as the New Language for Real-World Decision Making

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:57.292927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:57.292927Z digest=sha256:467a0119b88318ae4cdf0b54c06d0f8403a519b36ab3a865bae77b983532b865

Observation 2ada284e-c6b7-4b1d-9580-983545486f75 · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents Video as the New Language for Real-World Decision Making

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:57.884964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:57.884964Z digest=sha256:54124f872b5476e9e12c6893f54b119e8ff94021d147444ab7f0ef27fffc3664

Observation bd6a97b2-880f-4f4a-81a5-1f4dba080336 · inbound

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics cites this paper.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video as the New Language for Real-World Decision Making

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.250329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.250329Z digest=sha256:26e6e5b80243d8d4e44206ba3c75e8550dd8942e9a3dd4b2cb29963733ae71c4

Observation 37625214-519f-46b9-9d50-251d091c359f · inbound

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models cites this paper.

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models Video as the New Language for Real-World Decision Making

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:31.134881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:31.134881Z digest=sha256:b2ec1cb26a5069a0ff99181093d2ae9e47600d57ba040d99dba7cb6c750f27aa

Observation ebdd81f7-2e05-4cc2-aa21-62ef55f20a77 · inbound

Whole-Body Conditioned Egocentric Video Prediction cites this paper.

Whole-Body Conditioned Egocentric Video Prediction Video as the New Language for Real-World Decision Making

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:27:13.090738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:27:13.090738Z digest=sha256:66ab44dd787f179b9dfb7be1c48d7273f53b45003145397bd8f2b673f0f3e40f

Observation 9d3b7601-69a2-4004-9bce-c21bbd9453f9 · inbound

T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation cites this paper.

T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation Video as the New Language for Real-World Decision Making

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:28.927615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:28.927615Z digest=sha256:e94ddcad69763166ae42b9b3c1596c150d105598646df7187f8cd5461d1d2b2a

Observation 4a4b10f8-090a-45da-8050-70aa98552be8 · inbound

Video models are zero-shot learners and reasoners cites this paper.

Video models are zero-shot learners and reasoners Video as the New Language for Real-World Decision Making

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.722401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:6e057c3dca56c2b74602f864c86e5d945c87f70a080ff53fbf30ef332801bf6b

Observation 87e01a35-60f0-4852-9adf-c989b5237486 · inbound

Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning cites this paper.

Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning Video as the New Language for Real-World Decision Making

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:53:54.888033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T02:52:22.919967Z digest=sha256:22cf94578cce694ccc7aa959f8676f4446eee9f82301a00f08ec50023dc97b54

Observation 52267ed1-33aa-4896-8491-9ed18aef500c · inbound

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery cites this paper.

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery Video as the New Language for Real-World Decision Making

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T05:23:36.964436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:23:36.964436Z digest=sha256:75900dd250765e48fb35f010dcb38c0915ab8ccf6255646a486be47d7781522d

Observation a645df80-184e-4802-96b5-9920ab6b6856 · inbound

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models cites this paper.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Video as the New Language for Real-World Decision Making

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:46:02.852584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:51:12.604102Z digest=sha256:2dcb1b96ffd509c13ab6094f1c0fab71772629fcb6b75ade501fcaf9c6fa84c9

Observation 16b3b3ec-13db-46c2-b9a2-84c7af644737 · inbound

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models cites this paper.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Video as the New Language for Real-World Decision Making

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:27.915621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:99b62ba37e957801311a234d323e17e9099a2b80302056fe0d62a483eba1659a

Observation a8759541-caa2-4281-a86f-58fc7310dc89 · inbound

Visual prompt engineering for video models cites this paper.

Visual prompt engineering for video models Video as the New Language for Real-World Decision Making

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T02:13:11.587979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:13:11.587979Z digest=sha256:31b72d21e63cbb2686178b89e7b317183cfc056f0210804ecd5e65518c3ca37b