Pith. sign in

Paper Citation Record · LEDGER

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

As of 17 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 2 inbound Pith citation observations for arXiv:2412.01987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01987 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:04:09.060651Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-23T00:30:55.729900Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T00:32:18.226311Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact2
  • verified fuzzy35
  • unresolved33
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbfc89c9-4c74-41e5-acbe-0ae8d7cc9fc5 · outbound

This paper cites GPT-4 Technical Report.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.631013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.631013Z digest=sha256:19c43c87f1a38e30f10cf2bc3a137ebabb941b4a5da353617f3c7ac82236e6fc

Observation 0474e1f2-d6b2-4e26-a975-3bf24cdf8c0b · outbound

This paper cites Ht-step: Aligning instructional articles with how-to videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Ht-step: Aligning instructional articles with how-to videos

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.169513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.639434Z digest=sha256:bd337d5492066a1e5295ac7676d87226d9dc10a1e018da04f62865742a493b3a

Observation 8f61ac1d-c42a-4270-a247-83f6c736d298 · outbound

This paper cites Introducing computer use, a new claude 3.5 son- net, and claude 3.5 haiku.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Introducing computer use, a new claude 3.5 son- net, and claude 3.5 haiku

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.154558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.644101Z digest=sha256:aef9b57456b7a2273572a927d5112a32a05858f58fda5a7e484b4edd82321d22

Observation 875724f0-374e-43f8-b6f3-ef6ab842f263 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.140176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.648560Z digest=sha256:6789492b5987ca0ac8026b178e10e9d26834a63596a42002175ba7e5736b45cd

Observation 8ab9f193-3f26-4f9d-a311-62485c7168ad · outbound

This paper cites Whisperx: Time-accurate speech transcription of long- form audio.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Whisperx: Time-accurate speech transcription of long- form audio

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.126539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.652991Z digest=sha256:1a822732e300cef3b8ef0d0ce5bca32696542b582a69cff7a2be0ac31668bfc1

Observation cb136022-b1d8-437c-a9af-85786097d1b9 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.657233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.657233Z digest=sha256:b37004ea13a8a6512bef5468e2f848dde6f1a280ba4e805c44b4556593275e38

Observation 2228cb1a-90c3-4c2f-af15-114b78151792 · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.661872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.661872Z digest=sha256:297a004d2744c2e142fd80ebf1c30e24595c974fb399a763aefbc7ef5ff0a275

Observation 628aaaac-5f4e-46f2-ac6e-73cfb7c3be57 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.670093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.670093Z digest=sha256:512b688fb3d297c896eb8848b9f96623ff94c28c7d66d06eddb2ac804a966552

Observation 746ffd17-70f9-4af0-a183-6a1a1a6b9442 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.675581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.675581Z digest=sha256:2a343c50c38b903b2c5b7c722b252bfe306ab81fa1255995e4569c1fff241ea6

Observation f136592c-ea7d-406d-809b-5d5afe95275f · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.685044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.685044Z digest=sha256:636d9d92f6b007481b44c89f8a945e19e6caacbb699af5f89db4979f7a4f7edf

Observation 06fae0db-32da-468c-a6bd-38a1976d1f43 · outbound

This paper cites Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:04:09.430891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.697307Z digest=sha256:90538b4653fd2a590a43fec9a092c5fa5a9eb69188c3c6846863b749a4a81b5f

Observation 72bfeceb-e09b-43f6-b98c-ac53ebc28a20 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions In- structpix2pix: Learning to follow image editing instructions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.101887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.704361Z digest=sha256:681d0906026301aa9a26305f3d6ffdd734d59ebdbcd7829dbaa379be88edb4ed

Observation 38ec8158-98fa-4360-8dc1-1a764f8a2779 · outbound

This paper cites Video generation models as world simulators.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Video generation models as world simulators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.709987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.709987Z digest=sha256:768d9d7e50512dc739765f248368f9d823e66b2713c65c2376feda51e6975412

Observation fb6aa962-a943-45e2-9180-a2aec0997299 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.714910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.714910Z digest=sha256:fdf90a2338fc4c2109dc4cb2fbe166c1713a872f032554a52ff3a8613432f383

Observation 9f601e5c-8a61-48e9-9bff-bd37a1e99ac1 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.724998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.724998Z digest=sha256:0bb3d4d853ea4455fa664077a691897cf7a02c3b5fec6ca9987e74687b7a8102

Observation 3be16d03-b94e-4d79-aace-587c08d7e986 · outbound

This paper cites Seine: Short-to-long video diffusion model for generative transition and prediction.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Seine: Short-to-long video diffusion model for generative transition and prediction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.730901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.730901Z digest=sha256:84b4385695434fc32cfb7e5098f6d5252d82d2af048b76668cb5397f1d61691a

Observation de56ccef-0ea9-4e96-a12d-695c9042cc4d · outbound

This paper cites Palm: Scaling language modeling with pathways.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Palm: Scaling language modeling with pathways

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.058051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.735941Z digest=sha256:5f9e1c5aa11845ea3b7fc952242bc74280eac36c7d5b041551ff90b510da4dcb

Observation d8a4ef4d-ce20-4221-a470-e5b41932c615 · outbound

This paper cites Learning universal policies via text-guided video generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning universal policies via text-guided video generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.043067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.742913Z digest=sha256:116b80d014b21c7df7e35100512a96fdc50981d630e42d115bac57c7dc7ebc41

Observation 63ad7054-90a2-4ab5-abfb-c809dea7d116 · outbound

This paper cites The Llama 3 Herd of Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.747773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.747773Z digest=sha256:d021f2e3bd2441d06e34de36d8bb06553261f81f1ab4d48280eeb39cbaa117dc

Observation efc81969-1507-4b68-bb28-67ee39e1c0c9 · outbound

This paper cites Step- former: Self-supervised step discovery and localization in instructional videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Step- former: Self-supervised step discovery and localization in instructional videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.028020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.753963Z digest=sha256:4741543a847b4b41a2e0113f86b00e62508d36d24c75cd7dfbd82f304de8e4cc

Observation f12d4f78-d0a0-4581-a0f7-9b5029f6f0f1 · outbound

This paper cites Data Filtering Networks.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Data Filtering Networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.758855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.758855Z digest=sha256:462feca52ff3890297adacd8b444f2c38dc7cab037d2e43b354d8a6b57e5b69a

Observation 2df30a7a-d2c4-4f11-bc5e-911df71131bd · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Preserve your own correlation: A noise prior for video diffusion models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.763399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.763399Z digest=sha256:528512fba7a92f0e4adbbce1a9f05039e275054727d22935782f8fce4c088764

Observation 71c850b4-401d-4bf5-9ee1-348ffdf0f099 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.767955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.767955Z digest=sha256:b49d6b909ed698b5788b043d0f4bd88c2173d0c722e1698d6b791effc5f79966

Observation 087efbe0-b9a4-4ab9-9432-f7c3f38a3c80 · outbound

This paper cites Photorealistic Video Generation with Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Photorealistic Video Generation with Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.772500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.772500Z digest=sha256:5172d069cdb8a53e6fe029f5576b6d33015a383c178b1c82aa2318eaa341cf21

Observation 01f158c6-5bd0-4186-8ecd-12540da6f884 · outbound

This paper cites Temporal alignment networks for long-term video.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Temporal alignment networks for long-term video

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.001952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.778529Z digest=sha256:ec30c8f2fbeac427d5b725b8dcd44cb2778687d47983790ba648dfe02f8aeea5

Observation d77c6347-6bce-4269-bd9d-e48011e07a49 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.783762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.783762Z digest=sha256:91779e915d354c21f800409453c9f0ff0cf5ce6f5f325c055b1df78deb1eca37

Observation 2b027e12-2a04-4346-9ad4-79bb40dc6edc · outbound

This paper cites Denoising diffu- sion probabilistic models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Denoising diffu- sion probabilistic models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.789267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.789267Z digest=sha256:b15144023dd1db4fc3bfca67e5f5cb4ad04b921bdd45af4da85b62773bb78418

Observation 3abf3e7a-6051-4e66-9f39-dc79b95c8b6b · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Imagen Video: High Definition Video Generation with Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.793860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.793860Z digest=sha256:048bc3f3682d222e71284742d01c19ecc0b493f68aae5fb8526f0327c8a1e62f

Observation 26c9e021-b692-4deb-88b7-002fa62d32bd · outbound

This paper cites Video dif- fusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Video dif- fusion models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.800354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.800354Z digest=sha256:17391ccbb23eee74f6ddecaae1ae0355253cea2d0269adcbb711c3249630a9c6

Observation b396e86c-3622-448c-8055-ac46ab252cf8 · outbound

This paper cites Make it move: controllable image-to-video generation with text de- scriptions.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Make it move: controllable image-to-video generation with text de- scriptions

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.969266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.806210Z digest=sha256:cd0390aea44195a1825181648f267eb41247da57ee69dc4737ef11025d2acd7a

Observation 9b7a549b-ad21-4a77-8b66-2143d34c7663 · outbound

This paper cites An edit friendly ddpm noise space: Inversion and manipulations.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions An edit friendly ddpm noise space: Inversion and manipulations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.953814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.811584Z digest=sha256:7a5e1b917852fdb6c7cf36ebfe482b92edc03d434225cb16d97c256ef2dbf509

Observation 7324c0f5-d6b0-43d8-ade9-fd63baecc2ec · outbound

This paper cites Incorporating Task Progress Knowledge for Subgoal Generation in Robotic Manipulation through Image Edits.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Incorporating Task Progress Knowledge for Subgoal Generation in Robotic Manipulation through Image Edits

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.817423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.817423Z digest=sha256:631b8b1b1d83b08816b44a4353c7b2a0c4043bc340cab4c9d0f4dcb4080e94b4

Observation 55037516-e7ce-47bd-bed1-a644e7532620 · outbound

This paper cites Large language models are zero-shot reasoners.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Large language models are zero-shot reasoners

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.940558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.828636Z digest=sha256:b1dc3e8046c761e45455693f792ca3fc93c5ea9cde00b4999e018768723ae7d3

Observation 04de4a20-014e-45dd-9ebe-c5fb5bb9f704 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.833207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.833207Z digest=sha256:dc1e291f84ba6561da534e83a0980029c70ced54a1ab3ebcf17ea3a166e6d86e

Observation 784a9e7e-f9ad-483e-b297-44665d9fe130 · outbound

This paper cites Learning Action and Reasoning-Centric Image Editing from Videos and Simulations.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.838693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.838693Z digest=sha256:0ba2245ab7274204a9b1956b56c83228c1d2e8cf955afdfe2adf10ab432207a0

Observation 97fb824f-cc38-46a4-880d-f76455855370 · outbound

This paper cites LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.843887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.843887Z digest=sha256:075e081ddbad29c233a170efda48ab2d8e159725a07d9703115c6d09d84e29c7

Observation 6cf761fe-5126-447f-9cd1-21b14f35bc65 · outbound

This paper cites Multi-sentence grounding for long- term instructional video.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Multi-sentence grounding for long- term instructional video

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.926952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.848623Z digest=sha256:7a9f2df3845e34aff7e654dbfb25fa2085588ce4d980743be5d8cdc00ebeb229

Observation e15e0d37-d20b-425f-9fbd-a0e04fa81b58 · outbound

This paper cites Dreamitate: Real-World Visuomotor Policy Learning via Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.852974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.852974Z digest=sha256:fd37b08df626b26a9a5bdd4e168136b5b271255b844ff7d377e3df5eb3fc5327

Observation 2e6c8c9a-86e6-4774-bc4c-faed00a04766 · outbound

This paper cites Learning to ground instructional articles in videos through narrations.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning to ground instructional articles in videos through narrations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.913479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.857693Z digest=sha256:b52a5d7304ea8d6c35d4612f731a5f7814b9c74d9525f36ff8d755feae8fbedf

Observation d71a7704-a7bb-4419-9c33-a5e13bd09666 · outbound

This paper cites Vidm: Video implicit diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Vidm: Video implicit diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.899081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.861703Z digest=sha256:c32ef7f0d1c2112ae22305fc342b7afae3ac1ed8d343febac8646a1625382a94

Observation b36f24c2-8541-4af3-847b-8808db6542bb · outbound

This paper cites Generating illustrated instructions.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Generating illustrated instructions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.886483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.865718Z digest=sha256:d1afd5ed5007907eec297ac66c52989bbac129b86f3288ec5b20b518d163d46f

Observation 3997a520-f7c9-4b35-8c17-146d1d48f8a8 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.873064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.884360Z digest=sha256:f744bf75522d1a9426b8514d486dbd5fe93f4f44ac06026f430e110b1d8d6cde

Observation 75ab12e5-d43e-4af6-bc3b-3f0068300426 · outbound

This paper cites Visual reinforcement learn- ing with imagined goals.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Visual reinforcement learn- ing with imagined goals

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.858339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.889614Z digest=sha256:1558c7a17b484ad226ab598109beee85ca117e81bb5e44dd9ed7a141df9a2bd4

Observation 3159ebe7-26d7-4e29-a8c7-7056b9841d72 · outbound

This paper cites Dinov2: Learning robust visual features without su- pervision.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dinov2: Learning robust visual features without su- pervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.844241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.894028Z digest=sha256:323bdf935addb910589b44579445a58d3ac4be2d6b0b5b8d889a49d20362942e

Observation 91487550-b785-4f9b-98c3-6241dcfd339d · outbound

This paper cites Coherent Zero-Shot Visual Instruction Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Coherent Zero-Shot Visual Instruction Generation

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:04:09.220057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.900527Z digest=sha256:dc6a294b0eda867c04f2e58b1df7d3ae3ff16b30f78661a11a6f46ea2e39e0b4

Observation 60140c13-e07b-4231-85d3-a63fc6abda67 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Movie Gen: A Cast of Media Foundation Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.907577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.907577Z digest=sha256:0d4620d81619ed1db6c02587a1ad841280d7b937d059ca0a3b6d18f0f048be9e

Observation 4619a875-76d8-4c97-b9cc-687d634fb7a6 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learn- ing transferable visual models from natural language super- vision

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.829895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.913694Z digest=sha256:85d2390a1cba72c0b7dd40c5cdbed2992b0d6db1c1ade849e0d03c8a81c647bd

Observation e2eb7c0b-4895-4727-8a31-4d58a0bef8ac · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions High-resolution image syn- thesis with latent diffusion models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.815433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.922999Z digest=sha256:276b46e2c5081b1d879d338840176dde11e5aed4eaee5639cdc34f46febe047d

Observation ff5beb7d-edc1-49b3-b812-878915810340 · outbound

This paper cites Gen-3 alpha.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Gen-3 alpha

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.800306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.929037Z digest=sha256:1e08358bfae95d0360002a8951c44aab8b0ccfeb1ba818d65847d58ae93331d4

Observation abd3eec9-2225-48e8-a13b-a20ccc259fa1 · outbound

This paper cites Howtocap- tion: Prompting llms to transform video annotations at scale.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Howtocap- tion: Prompting llms to transform video annotations at scale

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.786673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.935087Z digest=sha256:879d30604314e98660ccb5faf0148794d77f2baa5d504edeb353fffd71125ff4

Observation d89dc2ff-12b3-47c0-9b6e-d4c0834ef508 · outbound

This paper cites Ego4d goal-step: Toward hierarchical understanding of procedural activities.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Ego4d goal-step: Toward hierarchical understanding of procedural activities

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.772416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.941594Z digest=sha256:2c540f006423979d86798e5fa91858948e8ae3788559c0bad5b153f0f47d46be

Observation f4c24efa-4fec-44fd-8d0f-eb90fd4941d9 · outbound

This paper cites Multi-task learning of object states and state-modifying actions from web videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Multi-task learning of object states and state-modifying actions from web videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.755919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.947814Z digest=sha256:1def798bc62aa4795cbcdc45f273839d48205f149ac3e838aa119bbfc912f02d

Observation 91449512-da44-444c-9e3c-6da5db8539b9 · outbound

This paper cites Genhowto: Learning to generate actions and state transformations from instructional videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Genhowto: Learning to generate actions and state transformations from instructional videos

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.739336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.954781Z digest=sha256:f96d132afc0a951275eb99f434bbb85751db1b1ed7590fec09558e465f73f0c8

Observation a16c4a41-8d5a-4f5c-8f2b-4e07dd8542a1 · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.707756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.967511Z digest=sha256:64f18524c75230f8de7a184b83f4bcb600cfaea0e55803abb2fe2bb2f5082be2

Observation 23aa6fd0-2760-465d-912d-8f5aed6eb5e4 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Videocomposer: Compositional video synthesis with motion controllability

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.688504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.971997Z digest=sha256:51008f1fe981036916414c797175bbeb007f0a2651535a6f47e539d1166b62ad

Observation 3205a1c7-a4b1-4f96-8dd6-048908a636d5 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.978979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.978979Z digest=sha256:ce50f5c40eb547a66574a00e92ddb7e78ab974191bbc9ac0905e9e4abd963e5e

Observation dd22c36a-6f25-4b61-93fb-ffc4df674270 · outbound

This paper cites Flow as the Cross-Domain Manipulation Interface.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Flow as the Cross-Domain Manipulation Interface

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.983607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.983607Z digest=sha256:d1d0ed865609352e58461086cdda6cc191eea277e40ab486696089ddbdf039be

Observation ba9f1114-b91f-457a-b153-a3a3db55453b · outbound

This paper cites Learn- ing object state changes in videos: An open-world perspec- tive.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learn- ing object state changes in videos: An open-world perspec- tive

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.660696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.988887Z digest=sha256:1f000a77a92922d4dc1b808a70293f81e54855e2bf4bb0f8a1b67a2f1343f62a

Observation 1845b683-1324-44c6-bc0d-25ebc34d4a7d · outbound

This paper cites Unloc: A unified framework for video localization tasks.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Unloc: A unified framework for video localization tasks

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.642664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.993253Z digest=sha256:979e6d68f21c8c0fca309116023f4d8c9a0ad3c72428418ade54481b937f8a61

Observation a2b2db66-34d9-4f6f-b6e8-e32ddb6aa599 · outbound

This paper cites Dif- fusion probabilistic modeling for video generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dif- fusion probabilistic modeling for video generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.624772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.997859Z digest=sha256:765978f2ed2841e88adfe03558316c8697729b865013e541012447f428b84cd9

Observation 3c955992-aea2-4564-a178-a26a2d839d06 · outbound

This paper cites Learning interactive real-world simulators.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning interactive real-world simulators

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.608529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:09.002378Z digest=sha256:e1d52ea4cebb576fff2f18591eec08b2e258ac662afa25cdc3ac215df52193e3

Observation bef5f001-9e54-445e-b330-9f7a4fb563b9 · outbound

This paper cites Visual goal-step inference using wikihow.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Visual goal-step inference using wikihow

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.592663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:09.006957Z digest=sha256:58f2f04a398d83cc8644d30391a3867146d03c691def4cdbe95179a738bd3dba

Observation c7c52e77-656e-4c1e-a25e-74deb2d6c8bf · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.011524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.011524Z digest=sha256:c4396c0c7943aa86332f3cb6f6a36a802ac7ba6361b58c8e1450aa5757fbda38

Observation a711e4fc-19a1-400b-b73e-8c4cd7eef45c · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Video probabilistic diffusion models in projected latent space

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.575661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:09.016495Z digest=sha256:80583c99dc591434f218be5e30a346fc34e0e34f75792a003d169cbd6236c28b

Observation c6ab4b64-56a5-4d2e-8b74-27373b76e153 · outbound

This paper cites Scaling Robot Learning with Semantically Imagined Experience.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Scaling Robot Learning with Semantically Imagined Experience

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.020787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.020787Z digest=sha256:c4793ce07192589a2ac1ee9ed17e43079b585499f1ff06a9da061199b462359c

Observation c37e19e9-9711-4bac-ab38-727aa59b8c3d · outbound

This paper cites Sigmoid loss for language image pre-training.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Sigmoid loss for language image pre-training

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.558855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:09.025635Z digest=sha256:2d8f8c8ac621614f1e084f94f04d19a73dd27fb8cecd94c2ef5bd262be80ad85

Observation 80dc2485-ed8d-4389-a1e3-0753cf4468de · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Show-1: Marrying pixel and latent diffusion models for text-to-video generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.045663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.045663Z digest=sha256:d17c54aabc8544af763a62d0a86c48f5e9bf95c09b462a638eb0e6573ba9ab35

Observation 73f4869a-77b4-4c89-baac-b08d423997fd · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.050387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.050387Z digest=sha256:5684152633198c187f9dc9ad6d8ac7cd0beaba001f1d3ff3f7237b579bee02be

Observation 85dda75d-30b6-47d7-ac36-27a93e83fce2 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.055356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.055356Z digest=sha256:7c2934cec007545d0774e008dd146379a09cdde48adea15f7d59735d80954c5d

Observation 10e75bc3-40ce-4afc-a621-b5d893492669 · outbound

This paper cites Put some aluminum foil in there.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Put some aluminum foil in there

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T00:04:09.531802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:09.060651Z digest=sha256:24e6c4864c5161cd866ce772f25e6bf766df6c669d1986ec4e94b10934e0d6b8

Observation cef7d800-01de-4556-9f7b-c2b20a34323f · outbound

This paper cites an unresolved cited work.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:04:09.723153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:04:08.960237Z digest=sha256:3522958a7e78acca50838d01bf4248c8142653c46a77828e558cdb86ba824db4

Pith citing papers

Observation 273928a8-b33a-4626-bce9-409575dfca99 · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.229135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:2e07cf7599a4b32ae47138210651a50efbaf4d0b9f17dd19b4c199af76ca19b4

Observation ab65c9fd-487b-4646-9eea-77765388e187 · inbound

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation cites this paper.

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:25.991899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T13:50:23.835653Z digest=sha256:54c8a8d7c82542fa197c54ce7c83ce973868f3c78b972b86a0586580bc0ae1cc