Pith. sign in

Paper Citation Record · LEDGER

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting

As of 15 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2412.11621.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11621 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:49:08.954450Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3704557-7b7c-4f31-8cd6-39bd177236fd · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.560441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.771367Z digest=sha256:e6caa393c0514c4479c16cee4feb1f3212110a14f461738a55de40893e1ce222

Observation 335ed295-1e65-46e1-88cc-596fc51371fc · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.775952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.775952Z digest=sha256:3e909e36aa6f2717405ffedf160ee08c99bfe13e8b94951ba574520672d9d514

Observation d53646f4-ed9d-48e2-9d1c-bd94b34e96ca · outbound

This paper cites W.; Fidler, S.; and Kreis, K.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting W.; Fidler, S.; and Kreis, K

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.536762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.779774Z digest=sha256:f239a126893d9579c3795fbc783ddc5dd366d59c1a990c025eb5670077954ef5

Observation fc4ce99e-b9b6-439a-bbf5-690cbf334087 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.524899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.783638Z digest=sha256:78eae8cc9193dc9bba237fd98285ddb22da221aee8695100da82b681a43843e8

Observation 69e13dbf-1d91-4e02-8842-ebf95ba1bfcf · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.512793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.787366Z digest=sha256:36e39d904f3c85f19ad6f5d69e1fd209cd9c1392cfe38a250b2221a69d807edb

Observation f0f099a0-f06a-4b87-accf-3b84683ab146 · outbound

This paper cites Prompting Large Language Models With the Socratic Method.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Prompting Large Language Models With the Socratic Method

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.791812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.791812Z digest=sha256:cb5b6f87724b06ab18d637cdbc2ad2dffed91aae83cdcecb49b33c3ffc445cae

Observation a090152c-e966-4ce8-8258-ab4665fec9f4 · outbound

This paper cites M.; and Cardie, C.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting M.; and Cardie, C

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.501437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.796874Z digest=sha256:e998336465b663684aac332d6af64371e6420bd62ea0e6b8a412c956b96e7412

Observation f6d5fc53-e486-42f4-8cf6-11c77dc60bc4 · outbound

This paper cites G.; Wildes, R.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting G.; Wildes, R

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.490949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.800722Z digest=sha256:fc3f2cf6f8ccebcbe710744347e3ea47c88b95bb7b65c16ec44235c3f3132261

Observation 42092e8e-5e8b-40ae-a81d-9e27fd1ec5f3 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.479735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.804970Z digest=sha256:504d277f9635bdd79fdbf74808510f5bedb819d6be2a4ea1f7bdf62f4f0f9ec9

Observation 95726092-4085-4a18-8bf5-f65d9fe42131 · outbound

This paper cites Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.808773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.808773Z digest=sha256:7ecce9983a1bd2b540798c9d3e69ca936d847c98ff7289e9b234e6a2c60d3e93

Observation d32aa3d6-7551-4f38-b0ec-450175a8d828 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.467848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.813282Z digest=sha256:a14ebeedcd70d0e503812c50b8b9a9bfc1400fa10537124d47c1331eb8a2670f

Observation bbb01c25-d681-4ccf-b080-b6caf2b07682 · outbound

This paper cites F.; Song, C.; Chen, J.; Gao, D.; Lei, W.; Xu, Q.; Lim, J.; and Shou, M.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting F.; Song, C.; Chen, J.; Gao, D.; Lei, W.; Xu, Q.; Lim, J.; and Shou, M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.454996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.817561Z digest=sha256:083f0b2f5df78e65569fa3425880eba6ffd26cc4eb249f9b3f3b44b0c2a2b133

Observation 149f6b9a-9f9f-42e0-b150-3b4cc7aab241 · outbound

This paper cites Mistral 7B.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Mistral 7B

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.821527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.821527Z digest=sha256:53179a33092b8df791794ae976324f0341674546c6bed4aa44b7b54c83d596a9

Observation 04983b0e-a886-42a0-8d35-e1aeba6aec35 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.439353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.825759Z digest=sha256:abb3de8e6f307a17b6b28a6ba105376a9f5e707b5a99157fb7b184adc1521b4c

Observation 74f526f3-895a-4219-895c-7e73c1c056af · outbound

This paper cites S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.424290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.829227Z digest=sha256:8eb3ead7c4ad6a3365c5952b60c52df0895d72879a97cf0e5e1b675dc424ff4e

Observation 033db58e-43ba-42b5-9e20-57cec483aad4 · outbound

This paper cites B.; and Serre, T.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting B.; and Serre, T

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.412834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.832764Z digest=sha256:1cfce8ee50939c8fff45dfb1bb90773cc7c58b26d54423508532d643f2a9a839

Observation c790b25b-f0cb-4854-9dfb-80459f26cf61 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.400947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.836447Z digest=sha256:04c0db580ca489e0a8b6f4324f6858a48bc7955a352f863b3bac05c7e58c8d45

Observation 545c1ac0-e76b-4f31-9eb4-25ccb4cc5f66 · outbound

This paper cites LLM-grounded Video Diffusion Models.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting LLM-grounded Video Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.840008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.840008Z digest=sha256:a9acd5087f78af0af1c1c068b2dc16ac7c5b67b8d3be16eace001a867f3ce0fc

Observation 00aa1965-8e9b-4a68-bbe8-4f771952e744 · outbound

This paper cites VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.843675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.843675Z digest=sha256:eea884397716987db6b9085af4001d624ec34045c3f4fd1f6750845adb5bf977

Observation 24affb3b-56d8-43e6-a6a7-830cf003d20b · outbound

This paper cites Q.; and Lei, S.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Q.; and Lei, S

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.388594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.847944Z digest=sha256:ece6e784ee68f433d4772218115481503a4c50665d55d945708f56d533c72afe

Observation 7b76ba1a-ecfa-401e-89e3-b9798d9a14b0 · outbound

This paper cites E.; Eckstein, M.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting E.; Eckstein, M

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.374221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.851942Z digest=sha256:c146b06551cbf380763afabc49b9a5ba6e5e78b5935924bc9f1c4109d02ef60b

Observation 62fe9c8b-9799-4a95-8852-8ef2bc278224 · outbound

This paper cites Multimodal Procedural Planning via Dual Text-Image Prompting.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Multimodal Procedural Planning via Dual Text-Image Prompting

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.856049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.856049Z digest=sha256:133f0dd5e46cb20c752e8013933bcc36033693158d78b80d595b422e1173ebd6

Observation 376e6aa1-4f9b-4510-aa98-6c91b535c724 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.359713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.860635Z digest=sha256:b078e4317a88756c2374f9ca2316f381e83b259cd59a0fa12ac66f55d2f9715c

Observation ea943d56-b61a-41b4-951f-c5ab3e59e352 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.346541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.864619Z digest=sha256:dd77d46059b23266001f4b27315f0370e4325eafb4b4cbabcead705dfc2bb7a3

Observation 4ec25088-2703-44a6-bcba-f3cfc7e71df8 · outbound

This paper cites GPT-4 Technical Report.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting GPT-4 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.868882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.868882Z digest=sha256:33d144570f965791173b8c4ff8d603ee199352b218474689a4b0d7489110f06a

Observation 94ae40cd-cc7f-4ce5-b8a7-c7dbc941ceab · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.872878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.872878Z digest=sha256:d5d75febba29d214828d62adab700dbfab47971b683b60ac58cdc57f30269ecd

Observation 7c8f02b2-111a-42ec-bdf7-73e276cba1c2 · outbound

This paper cites W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.328255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.877444Z digest=sha256:e98d2de7d8dc442ca74ba8dd5edfea6bc4a362976e3f2715823060b275e61f0c

Observation 80bbef9c-ece4-43e5-9016-23f1f6c5d06e · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.317989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.881421Z digest=sha256:198e86d3845214300a0e9b8abe1adc994dc15a380b390ad91f0a4b016bfba948

Observation 14ecd0dd-a6f4-472e-961b-460215eba19a · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.306590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.884976Z digest=sha256:842aa0eaf45815b917cb299160d80399d4a0324fd9e5966b4377b2bbededfbae

Observation 23b1f492-326d-4dd8-b1b0-c5a864c906f6 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.295544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.888324Z digest=sha256:4d2798ba545290dae2de4cd65422a3a4a91f3bd4923c48f27993045e5918a724

Observation 598ba009-f5cd-4530-8272-9e2c58842eb2 · outbound

This paper cites H.; Sadler, B.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting H.; Sadler, B

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.283311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.891605Z digest=sha256:2a74e167d6799ea0e1287fe0ed8a13a045b368e92b9ec6c243482a800b0af4e1

Observation 5c6ef65e-f41b-4872-84ed-b9c6a83fd178 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.270981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.895322Z digest=sha256:8b83e1ba6bbe40519bc6e8ccfe0c8d35674f907d448931bbb8804fc65078b099

Observation 9124f636-ed3c-4f6e-ba19-9b26190776b6 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.258180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.899103Z digest=sha256:127ac01ba7959f42f4386f11a062003ae635952c3f668c0e8714be65626f70c4

Observation dd4fa7a6-4250-42e6-ac1d-289d73622c29 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.902451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.902451Z digest=sha256:e90ba0b0b3af158f35d120cd3341965f151b9bd3bb12da34f852d218a1133a8f

Observation fc121910-51f0-486a-a7f3-67ca2a509bc0 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.244619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.905627Z digest=sha256:731602ca24479f1f12a58df8da3f3eb4475429d0a908779ac59de5f17ec46595

Observation 263817a5-8948-426c-bbda-94e44b92593e · outbound

This paper cites ModelScope Text-to-Video Technical Report.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting ModelScope Text-to-Video Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.909066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.909066Z digest=sha256:a7e41dad1fa82fb509e2188e15624fd4ecfc1f915b2d7d454f84557b1f5a1fa4

Observation 2deac228-7cb7-4aa9-9c5f-88a36854967d · outbound

This paper cites Z.; Ge, Y.; Wang, X.; Lei, S.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Z.; Ge, Y.; Wang, X.; Lei, S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.233322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.912657Z digest=sha256:58e2ab5cec7ee0c58f5eaaf706f72ff0a374759e5726bb5ae59e36a576c93a16

Observation 3a5dc640-8b66-4a9b-a0ea-dc86b1d8a70f · outbound

This paper cites H.; Miech, A.; Pont - Tuset, J.; Laptev, I.; Sivic, J.; and Schmid, C.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting H.; Miech, A.; Pont - Tuset, J.; Laptev, I.; Sivic, J.; and Schmid, C

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.220251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.915869Z digest=sha256:7363b75ef3373670df18463cb5016034fc4edeb27fd1b313062f3e71d319be12

Observation 39972daa-76e7-4027-bbe8-784006803196 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.206886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.919591Z digest=sha256:b47012665c1b8749620592bcd31033c31b3ded7e5534478aec3cb5dc2a0b342c

Observation 784b6a82-36b1-4ffe-9c22-1619ecd18db8 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.923237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.923237Z digest=sha256:ff77d79a053adf60c767fded8d5a52ee1b6aa97e8acc31f32dc63361c03b729f

Observation c17ebd86-dcb9-49a2-8308-8f94be3ee6bc · outbound

This paper cites G.; Wildes, R.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting G.; Wildes, R

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.193535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.927426Z digest=sha256:680975b0c26615e0dd648e86840f46d2ee4463156e26247fe842bd784999d537

Observation 9820123e-4482-43f6-bbaf-d247a33215a4 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.931033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.931033Z digest=sha256:a1248b437708762e19f872c2d11de4ff1240f5ff9945a79981ead7401c1e07ad

Observation 10658eb0-4a2c-472f-9925-e6ed54ea9942 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.180515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.935214Z digest=sha256:05a5f46f2a0f4fdc3d7c09f599177a082931c1c50bf8e170fdc9ba4b9f526918

Observation 6dcd6dfa-435f-49f7-aabc-98b375f1bcb8 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.167551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.938944Z digest=sha256:db8a2d6d05e85d1ddc816d48303301422bdd5b0177e08914b1e8d8a74f049431

Observation 1e4c5177-b20d-49d4-92e9-a7b8df3ffaf0 · outbound

This paper cites an unresolved cited work.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:49:09.154446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.942721Z digest=sha256:b6fd3e2d63863df1e525e2fff3eddb2dea8a2d82e1c51d7efcd58c5cea7331be

Observation ad2dbf87-0579-46ce-bd2d-1804976dbdf5 · outbound

This paper cites G.; Fouhey, D.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting G.; Fouhey, D

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:49:09.140752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:49:08.946423Z digest=sha256:791a4682af14e1f5f9a7af8df4eabbfa937fa6c0115cb7548f53517693140148

Observation 63549908-fe39-45f7-ae68-14993bd8d06a · outbound

This paper cites , " * write output.state after.block = add.period write newline.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting , " * write output.state after.block = add.period write newline

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.950286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.950286Z digest=sha256:1d2a93c90cb16d03721cb2bc7ead550d1238299c1f8d8fe6fd3465949645093f

Observation b0daf325-8131-4ca7-8b22-a331bb7f1e31 · outbound

This paper cites write newline.

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T14:49:08.954450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:49:08.954450Z digest=sha256:d8412ac3040a32612ac66cf85eb8fb46713a70a4ecce41056fd5025cce1e7423

Pith citing papers

No inbound Pith citation observations are available.