Pith. sign in

Paper Citation Record · LEDGER

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

As of 17 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 2 inbound Pith citation observations for arXiv:2412.01987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01987 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:04:09.060651Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-23T00:30:55.729900Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T00:32:18.226311Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact2
  • verified fuzzy35
  • unresolved33
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbfc89c9-4c74-41e5-acbe-0ae8d7cc9fc5 · outbound

This paper cites GPT-4 Technical Report.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.631013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.631013Z digest=sha256:47f588bd90c79c19be817bad819d4b21a5881262a9060772be14b9171f649377

Observation 0474e1f2-d6b2-4e26-a975-3bf24cdf8c0b · outbound

This paper cites Ht-step: Aligning instructional articles with how-to videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Ht-step: Aligning instructional articles with how-to videos

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.169513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.639434Z digest=sha256:2877fb7c1f535f78af078f8684b670cd469701f505b76cf68c47f74f612971dd

Observation 8f61ac1d-c42a-4270-a247-83f6c736d298 · outbound

This paper cites Introducing computer use, a new claude 3.5 son- net, and claude 3.5 haiku.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Introducing computer use, a new claude 3.5 son- net, and claude 3.5 haiku

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.154558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.644101Z digest=sha256:b8c70a0b07c8e164951dec7038b27faa162d0b506eab00c0e57e73bee5d8ea35

Observation 875724f0-374e-43f8-b6f3-ef6ab842f263 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.140176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.648560Z digest=sha256:2da0d488a0ae17e7351a4beaa421a959119fe5d8bebbf77901569b6d3b84ffb8

Observation 8ab9f193-3f26-4f9d-a311-62485c7168ad · outbound

This paper cites Whisperx: Time-accurate speech transcription of long- form audio.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Whisperx: Time-accurate speech transcription of long- form audio

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.126539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.652991Z digest=sha256:68d3da26cdabfed21b1fbc88b9ea658b6ff2e258fb8b9575f5d2e8f8924d47af

Observation cb136022-b1d8-437c-a9af-85786097d1b9 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.657233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.657233Z digest=sha256:b37004ea13a8a6512bef5468e2f848dde6f1a280ba4e805c44b4556593275e38

Observation 2228cb1a-90c3-4c2f-af15-114b78151792 · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.661872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.661872Z digest=sha256:297a004d2744c2e142fd80ebf1c30e24595c974fb399a763aefbc7ef5ff0a275

Observation 628aaaac-5f4e-46f2-ac6e-73cfb7c3be57 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.670093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.670093Z digest=sha256:512b688fb3d297c896eb8848b9f96623ff94c28c7d66d06eddb2ac804a966552

Observation 746ffd17-70f9-4af0-a183-6a1a1a6b9442 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.675581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.675581Z digest=sha256:d60cc1a380dcbd9eb9b5c3f623f994e3a2ed5adb6f986491cd3eb0ca9f4d9b88

Observation f136592c-ea7d-406d-809b-5d5afe95275f · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.685044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.685044Z digest=sha256:636d9d92f6b007481b44c89f8a945e19e6caacbb699af5f89db4979f7a4f7edf

Observation 06fae0db-32da-468c-a6bd-38a1976d1f43 · outbound

This paper cites Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:04:09.430891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.697307Z digest=sha256:4491b8295f1d5ec9317b5f91f9723a46f892e49e82451b13b100de2f8897435d

Observation 72bfeceb-e09b-43f6-b98c-ac53ebc28a20 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions In- structpix2pix: Learning to follow image editing instructions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.101887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.704361Z digest=sha256:e1345d63f8a5df80fd3ed1968d7a574db2744549c5b71a71f04b2680fc93a7fc

Observation 38ec8158-98fa-4360-8dc1-1a764f8a2779 · outbound

This paper cites Video generation models as world simulators.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Video generation models as world simulators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.709987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.709987Z digest=sha256:768d9d7e50512dc739765f248368f9d823e66b2713c65c2376feda51e6975412

Observation fb6aa962-a943-45e2-9180-a2aec0997299 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.714910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.714910Z digest=sha256:fdf90a2338fc4c2109dc4cb2fbe166c1713a872f032554a52ff3a8613432f383

Observation 9f601e5c-8a61-48e9-9bff-bd37a1e99ac1 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.724998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.724998Z digest=sha256:0bb3d4d853ea4455fa664077a691897cf7a02c3b5fec6ca9987e74687b7a8102

Observation 3be16d03-b94e-4d79-aace-587c08d7e986 · outbound

This paper cites Seine: Short-to-long video diffusion model for generative transition and prediction.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Seine: Short-to-long video diffusion model for generative transition and prediction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.730901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.730901Z digest=sha256:84b4385695434fc32cfb7e5098f6d5252d82d2af048b76668cb5397f1d61691a

Observation de56ccef-0ea9-4e96-a12d-695c9042cc4d · outbound

This paper cites Palm: Scaling language modeling with pathways.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Palm: Scaling language modeling with pathways

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.058051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.735941Z digest=sha256:f39d4f3e6f1cc4ee51030145ed51d807eefc5be0c9c9d96e0bc23f4ce4a0be6e

Observation d8a4ef4d-ce20-4221-a470-e5b41932c615 · outbound

This paper cites Learning universal policies via text-guided video generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning universal policies via text-guided video generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.043067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.742913Z digest=sha256:8d40b28aa377657db4a14a049b451fea958a8b479a7f19b41e540372f5dda83d

Observation 63ad7054-90a2-4ab5-abfb-c809dea7d116 · outbound

This paper cites The Llama 3 Herd of Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.747773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.747773Z digest=sha256:d021f2e3bd2441d06e34de36d8bb06553261f81f1ab4d48280eeb39cbaa117dc

Observation efc81969-1507-4b68-bb28-67ee39e1c0c9 · outbound

This paper cites Step- former: Self-supervised step discovery and localization in instructional videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Step- former: Self-supervised step discovery and localization in instructional videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.028020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.753963Z digest=sha256:944b5a6707bbee4c97cbc2badaca96a53718e1839dfe97e0c09dae64c28dc346

Observation f12d4f78-d0a0-4581-a0f7-9b5029f6f0f1 · outbound

This paper cites Data Filtering Networks.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Data Filtering Networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.758855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.758855Z digest=sha256:462feca52ff3890297adacd8b444f2c38dc7cab037d2e43b354d8a6b57e5b69a

Observation 2df30a7a-d2c4-4f11-bc5e-911df71131bd · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Preserve your own correlation: A noise prior for video diffusion models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.763399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.763399Z digest=sha256:528512fba7a92f0e4adbbce1a9f05039e275054727d22935782f8fce4c088764

Observation 71c850b4-401d-4bf5-9ee1-348ffdf0f099 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.767955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.767955Z digest=sha256:b49d6b909ed698b5788b043d0f4bd88c2173d0c722e1698d6b791effc5f79966

Observation 087efbe0-b9a4-4ab9-9432-f7c3f38a3c80 · outbound

This paper cites Photorealistic Video Generation with Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Photorealistic Video Generation with Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.772500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.772500Z digest=sha256:5172d069cdb8a53e6fe029f5576b6d33015a383c178b1c82aa2318eaa341cf21

Observation 01f158c6-5bd0-4186-8ecd-12540da6f884 · outbound

This paper cites Temporal alignment networks for long-term video.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Temporal alignment networks for long-term video

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:10.001952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.778529Z digest=sha256:6fd153dbd40fac66a4eb16336e65a6dad36c86eb7908ac3e4f8ce4b7c4fde175

Observation d77c6347-6bce-4269-bd9d-e48011e07a49 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.783762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.783762Z digest=sha256:91779e915d354c21f800409453c9f0ff0cf5ce6f5f325c055b1df78deb1eca37

Observation 2b027e12-2a04-4346-9ad4-79bb40dc6edc · outbound

This paper cites Denoising diffu- sion probabilistic models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Denoising diffu- sion probabilistic models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.789267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.789267Z digest=sha256:b15144023dd1db4fc3bfca67e5f5cb4ad04b921bdd45af4da85b62773bb78418

Observation 3abf3e7a-6051-4e66-9f39-dc79b95c8b6b · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Imagen Video: High Definition Video Generation with Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.793860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.793860Z digest=sha256:048bc3f3682d222e71284742d01c19ecc0b493f68aae5fb8526f0327c8a1e62f

Observation 26c9e021-b692-4deb-88b7-002fa62d32bd · outbound

This paper cites Video dif- fusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Video dif- fusion models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.800354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.800354Z digest=sha256:17391ccbb23eee74f6ddecaae1ae0355253cea2d0269adcbb711c3249630a9c6

Observation b396e86c-3622-448c-8055-ac46ab252cf8 · outbound

This paper cites Make it move: controllable image-to-video generation with text de- scriptions.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Make it move: controllable image-to-video generation with text de- scriptions

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.969266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.806210Z digest=sha256:b1f08dcbc21200f239c40e18c452dd1804a754654a40e148364af02fed86311d

Observation 9b7a549b-ad21-4a77-8b66-2143d34c7663 · outbound

This paper cites An edit friendly ddpm noise space: Inversion and manipulations.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions An edit friendly ddpm noise space: Inversion and manipulations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.953814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.811584Z digest=sha256:c35a7eab8d9ff9b508dfaacae1491f2ebe71792f8cf40bcfc72c9bbc25379bf2

Observation 7324c0f5-d6b0-43d8-ade9-fd63baecc2ec · outbound

This paper cites Incorporating Task Progress Knowledge for Subgoal Generation in Robotic Manipulation through Image Edits.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Incorporating Task Progress Knowledge for Subgoal Generation in Robotic Manipulation through Image Edits

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.817423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.817423Z digest=sha256:631b8b1b1d83b08816b44a4353c7b2a0c4043bc340cab4c9d0f4dcb4080e94b4

Observation 55037516-e7ce-47bd-bed1-a644e7532620 · outbound

This paper cites Large language models are zero-shot reasoners.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Large language models are zero-shot reasoners

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.940558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.828636Z digest=sha256:d8e49d95ec7a430b42e3d3a31dd3553feb775c66d1a2ba3a8d7eeded6956a1a6

Observation 04de4a20-014e-45dd-9ebe-c5fb5bb9f704 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.833207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.833207Z digest=sha256:dc1e291f84ba6561da534e83a0980029c70ced54a1ab3ebcf17ea3a166e6d86e

Observation 784a9e7e-f9ad-483e-b297-44665d9fe130 · outbound

This paper cites Learning Action and Reasoning-Centric Image Editing from Videos and Simulations.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.838693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.838693Z digest=sha256:0ba2245ab7274204a9b1956b56c83228c1d2e8cf955afdfe2adf10ab432207a0

Observation 97fb824f-cc38-46a4-880d-f76455855370 · outbound

This paper cites LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.843887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.843887Z digest=sha256:075e081ddbad29c233a170efda48ab2d8e159725a07d9703115c6d09d84e29c7

Observation 6cf761fe-5126-447f-9cd1-21b14f35bc65 · outbound

This paper cites Multi-sentence grounding for long- term instructional video.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Multi-sentence grounding for long- term instructional video

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.926952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.848623Z digest=sha256:95233194df01f1b8d3bb7df01aa9107860321d99be334a92cb04c8d061b553b5

Observation e15e0d37-d20b-425f-9fbd-a0e04fa81b58 · outbound

This paper cites Dreamitate: Real-World Visuomotor Policy Learning via Video Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.852974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.852974Z digest=sha256:fd37b08df626b26a9a5bdd4e168136b5b271255b844ff7d377e3df5eb3fc5327

Observation 2e6c8c9a-86e6-4774-bc4c-faed00a04766 · outbound

This paper cites Learning to ground instructional articles in videos through narrations.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning to ground instructional articles in videos through narrations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.913479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.857693Z digest=sha256:b09565e0ce2c54248128a25175a5792fa36ecca644c582892e81a3061da1bd97

Observation d71a7704-a7bb-4419-9c33-a5e13bd09666 · outbound

This paper cites Vidm: Video implicit diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Vidm: Video implicit diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.899081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.861703Z digest=sha256:725881507f20baede7626ee03667ee1ad96e75f1e7c4aaa5660e91002a310b89

Observation b36f24c2-8541-4af3-847b-8808db6542bb · outbound

This paper cites Generating illustrated instructions.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Generating illustrated instructions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.886483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.865718Z digest=sha256:52cb8fe4c5b419bbf845d7ff58db6705d350a2ce3dd7b13a0846cb3cdbe47783

Observation 3997a520-f7c9-4b35-8c17-146d1d48f8a8 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.873064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.884360Z digest=sha256:28b63bae93c680021acacc15ea9af97e516db9c5fe98146a5b2a9d9cdcb4b91d

Observation 75ab12e5-d43e-4af6-bc3b-3f0068300426 · outbound

This paper cites Visual reinforcement learn- ing with imagined goals.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Visual reinforcement learn- ing with imagined goals

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.858339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.889614Z digest=sha256:fbe5939ea1845f32d52fa19db4a50107c8d7408082a52d17f21a422e316e7fb4

Observation 3159ebe7-26d7-4e29-a8c7-7056b9841d72 · outbound

This paper cites Dinov2: Learning robust visual features without su- pervision.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dinov2: Learning robust visual features without su- pervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.844241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.894028Z digest=sha256:e5b1d214f4493d70dcc6a01e52eeb058389de74030ae56d8a45158d83fdb5a7e

Observation 91487550-b785-4f9b-98c3-6241dcfd339d · outbound

This paper cites Coherent Zero-Shot Visual Instruction Generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Coherent Zero-Shot Visual Instruction Generation

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:04:09.220057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.900527Z digest=sha256:788b6a87bfd10f95d8ab7ad8cc39a63840992983b8f7c2c709b396559772f7db

Observation 60140c13-e07b-4231-85d3-a63fc6abda67 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Movie Gen: A Cast of Media Foundation Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.907577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.907577Z digest=sha256:78089fd964d5e3a5e0044ff011197b3de0bf180902bf34ac8e694a2c3f2f5da0

Observation 4619a875-76d8-4c97-b9cc-687d634fb7a6 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learn- ing transferable visual models from natural language super- vision

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.829895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.913694Z digest=sha256:18744a26e102e65d126f343d2f80913f40f3355a1bb63f5fcf037710ddadf090

Observation e2eb7c0b-4895-4727-8a31-4d58a0bef8ac · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions High-resolution image syn- thesis with latent diffusion models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.815433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.922999Z digest=sha256:7eee39d279288f57e777a7b0684c6d3c1c96cb62c883d4edad20d1d61862ddab

Observation ff5beb7d-edc1-49b3-b812-878915810340 · outbound

This paper cites Gen-3 alpha.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Gen-3 alpha

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.800306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.929037Z digest=sha256:518ff140d71150ecaf54a4a522ef0f68678727e449f5170e123436038e7b3974

Observation abd3eec9-2225-48e8-a13b-a20ccc259fa1 · outbound

This paper cites Howtocap- tion: Prompting llms to transform video annotations at scale.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Howtocap- tion: Prompting llms to transform video annotations at scale

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.786673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.935087Z digest=sha256:95325b7191dc679449d4853588f4ad9742f306589bd31154de6e988574d2b106

Observation d89dc2ff-12b3-47c0-9b6e-d4c0834ef508 · outbound

This paper cites Ego4d goal-step: Toward hierarchical understanding of procedural activities.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Ego4d goal-step: Toward hierarchical understanding of procedural activities

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.772416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.941594Z digest=sha256:a01ae1f2a87990a600f59f71c115f6613775a2267c5dc75e7ed49130698be318

Observation f4c24efa-4fec-44fd-8d0f-eb90fd4941d9 · outbound

This paper cites Multi-task learning of object states and state-modifying actions from web videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Multi-task learning of object states and state-modifying actions from web videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.755919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.947814Z digest=sha256:bd7102305b22166c74976215570320b240fd3b3e915c58a2ce051376ec3eca88

Observation 91449512-da44-444c-9e3c-6da5db8539b9 · outbound

This paper cites Genhowto: Learning to generate actions and state transformations from instructional videos.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Genhowto: Learning to generate actions and state transformations from instructional videos

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.739336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.954781Z digest=sha256:7da005ef2d2a960f536744b3ac20aea7ca6750c8f4128512186d64af213cae14

Observation a16c4a41-8d5a-4f5c-8f2b-4e07dd8542a1 · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.707756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.967511Z digest=sha256:d5c5bbdf7bfe0836cc2f378dbab62cc10a2f05df29267c649729f9142719bff8

Observation 23aa6fd0-2760-465d-912d-8f5aed6eb5e4 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Videocomposer: Compositional video synthesis with motion controllability

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.688504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.971997Z digest=sha256:8ef9d68e6d3688cac451baad0f2833300d1414b58b3fd6cf3544640113684f9d

Observation 3205a1c7-a4b1-4f96-8dd6-048908a636d5 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.978979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.978979Z digest=sha256:ce50f5c40eb547a66574a00e92ddb7e78ab974191bbc9ac0905e9e4abd963e5e

Observation dd22c36a-6f25-4b61-93fb-ffc4df674270 · outbound

This paper cites Flow as the Cross-Domain Manipulation Interface.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Flow as the Cross-Domain Manipulation Interface

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.983607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.983607Z digest=sha256:d1d0ed865609352e58461086cdda6cc191eea277e40ab486696089ddbdf039be

Observation ba9f1114-b91f-457a-b153-a3a3db55453b · outbound

This paper cites Learn- ing object state changes in videos: An open-world perspec- tive.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learn- ing object state changes in videos: An open-world perspec- tive

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.660696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.988887Z digest=sha256:0ce594ac2a9fe5b208e98a05d540152119df890fb1c7f89f57f1d9281b5167d4

Observation 1845b683-1324-44c6-bc0d-25ebc34d4a7d · outbound

This paper cites Unloc: A unified framework for video localization tasks.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Unloc: A unified framework for video localization tasks

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.642664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.993253Z digest=sha256:e703345ca3308ba0098c253683a5c3d4cc6deda9d86fc287cd7af2b94209baff

Observation a2b2db66-34d9-4f6f-b6e8-e32ddb6aa599 · outbound

This paper cites Dif- fusion probabilistic modeling for video generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Dif- fusion probabilistic modeling for video generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.624772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.997859Z digest=sha256:1768e9905f7cc87134d0534b3c2b284017e0c2615dc3e6c53b3ab166668e6372

Observation 3c955992-aea2-4564-a178-a26a2d839d06 · outbound

This paper cites Learning interactive real-world simulators.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Learning interactive real-world simulators

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.608529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:09.002378Z digest=sha256:89ae05dd7ff60fba8fc1d321f4ece61b678fb4693ec55aac3ef7c59b57872eb5

Observation bef5f001-9e54-445e-b330-9f7a4fb563b9 · outbound

This paper cites Visual goal-step inference using wikihow.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Visual goal-step inference using wikihow

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.592663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:09.006957Z digest=sha256:d73f1d56c6379f014bbff87f6183684cd7f6ea864351baf7809fccb2835de75a

Observation c7c52e77-656e-4c1e-a25e-74deb2d6c8bf · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.011524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.011524Z digest=sha256:c4396c0c7943aa86332f3cb6f6a36a802ac7ba6361b58c8e1450aa5757fbda38

Observation a711e4fc-19a1-400b-b73e-8c4cd7eef45c · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Video probabilistic diffusion models in projected latent space

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.575661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:09.016495Z digest=sha256:d57629f24c71e76434e5774307d61ac5db5a721bdd6100adb0ea323c6be38592

Observation c6ab4b64-56a5-4d2e-8b74-27373b76e153 · outbound

This paper cites Scaling Robot Learning with Semantically Imagined Experience.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Scaling Robot Learning with Semantically Imagined Experience

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.020787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.020787Z digest=sha256:c4793ce07192589a2ac1ee9ed17e43079b585499f1ff06a9da061199b462359c

Observation c37e19e9-9711-4bac-ab38-727aa59b8c3d · outbound

This paper cites Sigmoid loss for language image pre-training.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Sigmoid loss for language image pre-training

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:04:09.558855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:09.025635Z digest=sha256:e7527947a7bf24e6aedfa7d8f09ed4d7907726ce10e896bbb0146510b0188833

Observation 80dc2485-ed8d-4389-a1e3-0753cf4468de · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Show-1: Marrying pixel and latent diffusion models for text-to-video generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.045663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.045663Z digest=sha256:d17c54aabc8544af763a62d0a86c48f5e9bf95c09b462a638eb0e6573ba9ab35

Observation 73f4869a-77b4-4c89-baac-b08d423997fd · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.050387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.050387Z digest=sha256:5684152633198c187f9dc9ad6d8ac7cd0beaba001f1d3ff3f7237b579bee02be

Observation 85dda75d-30b6-47d7-ac36-27a93e83fce2 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:09.055356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:09.055356Z digest=sha256:7c2934cec007545d0774e008dd146379a09cdde48adea15f7d59735d80954c5d

Observation 10e75bc3-40ce-4afc-a621-b5d893492669 · outbound

This paper cites Put some aluminum foil in there.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Put some aluminum foil in there

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T00:04:09.531802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:09.060651Z digest=sha256:54ba39570c63e0d99f1a07162411200b6cf6f78f2b59742cc09ff394329c16b5

Observation cef7d800-01de-4556-9f7b-c2b20a34323f · outbound

This paper cites an unresolved cited work.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:04:09.723153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:04:08.960237Z digest=sha256:b31f54e441a050446d2fa291fa2bda31f0328b4ae7a76c3689afb00fb3924fc8

Pith citing papers

Observation 273928a8-b33a-4626-bce9-409575dfca99 · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.229135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:0f8bdc829fbb0bfd1fed48a66f552985f1bb49ebfd4f5c67de0dad5bffea0641

Observation ab65c9fd-487b-4646-9eea-77765388e187 · inbound

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation cites this paper.

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:25.991899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T13:50:23.835653Z digest=sha256:c8b8a9c60346a665f44167f5a81e1f0b2acfd6f931e92d0366ba33dff90cf739