Pith. sign in

Paper Citation Record · LEDGER

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion

As of 19 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2505.16425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16425 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:03:43.408265Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:38:37.497428Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-14T04:38:39.693957Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact8
  • verified fuzzy1
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9bc41830-0ba5-4bc4-a162-293bc1e8b35c · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:03:46.264468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:03:35.493664Z digest=sha256:22ab719fb09fa3089a19941a85274418b15f10450fc6249c7bb5fb448468b240

Observation d3a93cff-b7ce-4f0b-b53a-36e1d7befe0e · outbound

This paper cites Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:35.615548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:35.615548Z digest=sha256:1fd4eae689721f1e0a6e98bdf1af6eadbf0b90fbb80a150761e71b01becb5ebc

Observation a4276df0-c917-4904-8060-67d8944eb415 · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:35.811186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:35.811186Z digest=sha256:b1fe1d60d1bdd688626b58a5cfe0bae13be1f392ac3bb5c382cf63418e7eb212

Observation 76120329-8559-416c-9ba0-0a9d28055e15 · outbound

This paper cites https://api.semanticscholar.org/CorpusID:264403242 Improving image generation with better captions.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion https://api.semanticscholar.org/CorpusID:264403242 Improving image generation with better captions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:03:46.036153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:03:35.945204Z digest=sha256:a5d09efab0a57dc03d0e95e5933207a2238ab49901c7067362e345cf79a92548

Observation 55672839-4837-46ca-952e-2ce7a32bfd7e · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Training Diffusion Models with Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.086633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.086633Z digest=sha256:3572140655e56dc55eda485a2dd035179db39f10a1e7c7003c58dca3b21b9f05

Observation ff16a6bd-03ce-4f03-aa78-5b30b366f4a6 · outbound

This paper cites On the Design Fundamentals of Diffusion Models: A Survey.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion On the Design Fundamentals of Diffusion Models: A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.217850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.217850Z digest=sha256:6d0e27d62f2034a0b7da0b44ef0d560cb73edef47b297ad09994c47057baf104

Observation b34f1779-4abc-4d04-83ea-91e003946f69 · outbound

This paper cites MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.345971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.345971Z digest=sha256:9de797d3e2ed58fdaaa15a182d9a62fdb284bbb5644eb3634dcba14b5430ebdd

Observation 54cb629c-fbd1-44d4-98b1-aac0a7fb722b · outbound

This paper cites Evaluating Text-to-Image Generative Models: An Empirical Study on Human Image Synthesis.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Evaluating Text-to-Image Generative Models: An Empirical Study on Human Image Synthesis

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:45.649994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:03:36.484830Z digest=sha256:9415bef433bf57f7a1ec718825e50780283f5bcfaae32ca7749b2e2da5992dbd

Observation 9b4aefc5-8d71-4ab5-8f5c-9914c9948280 · outbound

This paper cites X-IQE: eXplainable Image Quality Evaluation for Text-to-Image Generation with Visual Large Language Models.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion X-IQE: eXplainable Image Quality Evaluation for Text-to-Image Generation with Visual Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.621699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.621699Z digest=sha256:a90281010b92f6459aa530aabde3d76b2bff5e1f8727c106488acfa6058a9dc6

Observation 382f8147-2089-4b7a-b9ed-18b03b123820 · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.759018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.759018Z digest=sha256:9900a0af38a02a5757004efd6cd51f9519931d1010fa192d39383f23d5490916

Observation 0edff757-7773-4372-b5d3-9a975d96818b · outbound

This paper cites Directly Fine-Tuning Diffusion Models on Differentiable Rewards.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Directly Fine-Tuning Diffusion Models on Differentiable Rewards

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.864834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.864834Z digest=sha256:49f3458045301fe884e0d6ce424f5d13cb1784eb8c76f7e9175e90c9b0c74a4b

Observation 429222e7-86b9-484c-bfba-d965c4e25d92 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.999362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.999362Z digest=sha256:894d4494e066a9f9bec320eb53d913a401539b7ddf33dd80b142b5a8fc58f0d4

Observation 7e163a25-9377-4b98-83b3-5fe7ae6f4acd · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.111417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.111417Z digest=sha256:5d0b5684834138d58c751ddf398de56d82fe0ae73fd8b729d6299a12cbccb046

Observation dda87e0e-175e-4c16-89a3-0cd24a0247a1 · outbound

This paper cites DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.181408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.181408Z digest=sha256:fa37ccef230a0dc86b2de2442ca4bce4ac2ad4239340c05a99d35971719ff3a3

Observation 8fc889f0-a173-4fad-8d21-1e5a948403a1 · outbound

This paper cites Learning to Segment Actions from Observation and Narration.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Learning to Segment Actions from Observation and Narration

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.252800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.252800Z digest=sha256:01a57866864d4ef385301b082666da290ee1c9d954fd638bf4163df4399d8c44

Observation 6cf01b20-c326-470f-b85c-98df55b90754 · outbound

This paper cites GPTScore: Evaluate as You Desire.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion GPTScore: Evaluate as You Desire

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.396043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.396043Z digest=sha256:760e9aab380e5f07169d467177a64196dba4e8bd21a6488f9a27b0d77195f658

Observation a9a1f61f-aa72-46a9-823b-aadd3484262d · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.532493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.532493Z digest=sha256:06530cd809a7c5d65a095eb09a5699572f3b2bce6bd2103a77fa36a0e775a860

Observation 66908ad7-fc0f-4d95-9147-fc4620f8c7f4 · outbound

This paper cites COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.652640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.652640Z digest=sha256:4c476ea2f9bf612147184886dea8f4460cd8a565cd422b9c93cd86807d44b4b4

Observation 75349684-e6b9-4c77-8b2e-337290b9b556 · outbound

This paper cites Temporal Alignment Networks for Long-term Video.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Temporal Alignment Networks for Long-term Video

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.830852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.830852Z digest=sha256:667392847250e33ddb67ac653b886d2c7301f525f7bf14983021d522e3e72874

Observation 80ce3f2d-e36e-4972-a64b-0bee5c49a276 · outbound

This paper cites Optimizing Prompts for Text-to-Image Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Optimizing Prompts for Text-to-Image Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.951089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.951089Z digest=sha256:c3f6b21e3daba664a457ae71de9fa8f25e98ff17c1232da2023ff3c768ec6d64

Observation 16ed4664-fc04-4e15-a13a-c6d6b349a2c4 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.064399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.064399Z digest=sha256:2ae5a0989181cc3ac40be966d90e0d7ece5518b88365228e62f7e2956ca33f35

Observation 3c6c03ac-590d-45eb-b9a4-985275f1e651 · outbound

This paper cites Denoising Diffusion Probabilistic Models.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Denoising Diffusion Probabilistic Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.202182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.202182Z digest=sha256:6ad2f787e3232520d05b8431930fe4e97e524083407e886410832dfc095d6dda

Observation 9b378390-b238-43e1-88f4-8300d51e0654 · outbound

This paper cites Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.348720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.348720Z digest=sha256:82b3bc80a377705cd34aa8efd044e8c626acc677b384fb297f62ce0f12b922d8

Observation b99cc7aa-52ee-4e2b-a7f6-f8fafb64ddfe · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.529587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.529587Z digest=sha256:c7594321923ba28350cfbd46fd6bc010469713eb45799a0c2c5680e7c99e98aa

Observation c9e1149b-1b73-4a62-bc96-adaa729df42c · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Aligning Text-to-Image Models using Human Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.645140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.645140Z digest=sha256:0f2fca9847d056a188a5a7579e11fd514b7a5efdd0688c4919f9d10846cac448

Observation 55fb65d7-3118-4f58-81b1-93fb4e89c5a1 · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.774002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.774002Z digest=sha256:9c89003869d576bde0e868a4d62968d312ed4eeb21db82aef26c1f93e8f57b8e

Observation ebb8b2a7-fe66-4f67-9d37-b9d5ebc0c942 · outbound

This paper cites Aligning Diffusion Models by Optimizing Human Utility.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Aligning Diffusion Models by Optimizing Human Utility

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.853890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.853890Z digest=sha256:0a8c6f00976deacb58768e3c7119144cc7f56bf386b37f6ac077148b61809960

Observation da03b956-f681-4ded-898d-e2bdeb741a3e · outbound

This paper cites Learning To Recognize Procedural Activities with Distant Supervision.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Learning To Recognize Procedural Activities with Distant Supervision

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:45.262945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:03:38.930421Z digest=sha256:c92d1b42daf031f8ff7293ee2f8499d5a19f21dbfc2c1fb797642eff0033074d

Observation 4436ae88-6107-41d9-8836-d8f293434a2e · outbound

This paper cites VideoDPO: Omni-Preference Alignment for Video Diffusion Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion VideoDPO: Omni-Preference Alignment for Video Diffusion Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:39.100590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:39.100590Z digest=sha256:04bab23ea411b663bc76911e8fae92b377a616e65d7f03954c06d52a60f136ec

Observation a626065a-b04c-415b-9b38-f11e173b044c · outbound

This paper cites UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:45.058912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:03:39.251120Z digest=sha256:00e67b64903dcb24f1db2696ae6e787e216544e3f367df6848c02b4db08d8918

Observation f51cdcfd-4231-4f1c-a9df-49631214729d · outbound

This paper cites Solving New Tasks by Adapting Internet Video Knowledge.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Solving New Tasks by Adapting Internet Video Knowledge

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:39.351300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:39.351300Z digest=sha256:73e0cfafff4af00ef07f40b21f93f6ae25f4b2512af8f7b8d4921c071577910f

Observation ef3a9452-43ae-4d55-8aff-20418f6b7e09 · outbound

This paper cites Generating Illustrated Instructions.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Generating Illustrated Instructions

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:44.923870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:03:39.512797Z digest=sha256:1a50493619b3678050d2c7ee1ae9c1e22a7b5264ebfefb567074d14d1cf51366

Observation c57b2b1f-fc7a-4ba2-ba09-529105cae8ea · outbound

This paper cites End-to-End Learning of Visual Representations from Uncurated Instructional Videos.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion End-to-End Learning of Visual Representations from Uncurated Instructional Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:39.673028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:39.673028Z digest=sha256:cc11ed5d78c8ad3b98e50dd375471b85a88dc67cbc85717089a8b9b08a0286d7

Observation 4227ee8f-183d-4e50-a91a-164bf27a4b8e · outbound

This paper cites CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:39.803189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:39.803189Z digest=sha256:72c2e8f26fbb32291cb3dc46164ba0624ff8c626ecc13ac06b12a17dd13e25a7

Observation 9d037a58-593c-42cf-ae70-d7a43bfb84e3 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:39.954864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:39.954864Z digest=sha256:63fe74d1634315a3077719926c50f845fb9e679e01cf9a01889dde564a2fa5b4

Observation d56fca1a-9741-4051-9fe7-488528d479cf · outbound

This paper cites Aligning Text-to-Image Diffusion Models with Reward Backpropagation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Aligning Text-to-Image Diffusion Models with Reward Backpropagation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:40.138551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:40.138551Z digest=sha256:ef971e7bc827bf04956b16f8ef348e60f2f601eb7a50bb84f66d85ee0513feb8

Observation 0d24491a-98fb-4bce-ba9f-fa3b4837b96e · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:40.248662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:40.248662Z digest=sha256:094ee8fc1a48d60a94360fbe25dc5ce5e9a552e065671adb1c950da537618cfd

Observation 6ee5cd23-a80c-47a9-ad2f-2bc5183b76e4 · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:40.416719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:40.416719Z digest=sha256:017c0f1338ff3e79b1fd0c51627ae31d6b44cbe1ac4005c614d9845750b4efdc

Observation 5a2974dc-3e1b-4f5e-aed9-c8f1c40f29e3 · outbound

This paper cites Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:44.698797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:03:40.666014Z digest=sha256:a54cd5d1868844fca77070deb7ef033f4f999e29b4ae4d1d5db444c3248e3709

Observation 0dbcb19f-e3a2-4a33-be81-10c4e94f8ce0 · outbound

This paper cites A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:40.893656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:40.893656Z digest=sha256:2ef2db0ce21b7335317f00d57793167ebcf2a71e5679be1e0ea3550b4dc56a3b

Observation 0169d1e0-fe34-4e79-b267-84e931944174 · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 41

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T15:03:44.551874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:03:41.027543Z digest=sha256:7d3f0aab53fc500f5a023fb2d285119fccaaf1109461a630a9713bd975605fd0

Observation 2002d8ad-96a8-4d23-92d0-74e8c72a7948 · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.212373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:41.212373Z digest=sha256:4f9998223c76fa0c3981370d20e73e6496f397044f4c2b2d00faa92f561ba2fd

Observation 416b32be-dc42-497c-b8b6-8a5c08946c2c · outbound

This paper cites Diffusion Model Alignment Using Direct Preference Optimization.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Diffusion Model Alignment Using Direct Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.362858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:41.362858Z digest=sha256:0fa0270bc15750574a767a974bba4219cc0f68fa5bf0a0489334d5f52115b9cd

Observation e384682d-fe63-4e4f-81c2-40d84d611c53 · outbound

This paper cites Is ChatGPT a Good NLG Evaluator? A Preliminary Study.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.497625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:41.497625Z digest=sha256:160f02d4246bec78261581bcb69d140d8f74e80d28c5f36bbc916ae317843b6d

Observation 3c41d0b5-a2cc-4e5a-894e-500555dff40a · outbound

This paper cites A Unified Agentic Framework for Evaluating Conditional Image Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion A Unified Agentic Framework for Evaluating Conditional Image Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.644976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:41.644976Z digest=sha256:aafbbd0c9df166d47853c4544a96312ffad4742ef1dd72d438151b6899861e0e

Observation f90e72b7-e5d5-4bbf-9174-cc80c38e0330 · outbound

This paper cites Large Language Models are not Fair Evaluators.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Large Language Models are not Fair Evaluators

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.863342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:41.863342Z digest=sha256:705a458915ecb655d775f4bb180bcf7ca3c6fa1be0a79260396fa2493ac5c0fb

Observation 7865345f-9366-4fe4-86b5-9feea14bb89e · outbound

This paper cites InstructionBench: An Instructional Video Understanding Benchmark.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion InstructionBench: An Instructional Video Understanding Benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.979783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:41.979783Z digest=sha256:ef4f11730515c1b54bb294081811a89a3a7202974994bf1c9b4fe400aebb791e

Observation cc03890c-d637-4de9-869d-5b6a16afdb0f · outbound

This paper cites Preference Alignment on Diffusion Model: A Comprehensive Survey for Image Generation and Editing.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Preference Alignment on Diffusion Model: A Comprehensive Survey for Image Generation and Editing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.075088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.075088Z digest=sha256:52ff769351a60856c73d7042a14d9081cbbb9449c48858096ddafc5109a56d80

Observation 997702fa-f5f5-4326-b142-5453fbd1c112 · outbound

This paper cites Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:44.113000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:03:42.159502Z digest=sha256:f7d758fa84221c7348b46302e790c137ba7a73a94b97d45af64d2c335f5cbefa

Observation 832f6dee-fc45-49bd-a327-c598cc61b5ce · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.304578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.304578Z digest=sha256:07198ef1c8e64c1d83dbdac4f499db0e4853b0c65949246a906313ce71033ca1

Observation 6c57587c-f0ce-4322-b7dc-6e32063275de · outbound

This paper cites ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.406631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.406631Z digest=sha256:9bdfe749665ec87ff9f355497bb06c8accad365e1c721bf3944e145d4e1808b3

Observation 659a86de-e39c-4353-bd2d-05ab46ad2efe · outbound

This paper cites GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.511905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.511905Z digest=sha256:11516a226455c75b366616dc99f5e5150535e1c658b4e4b4365aa6550ec15174

Observation 12ee1582-3187-4244-a65d-6a713db5d35d · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:03:45.870174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:03:42.625227Z digest=sha256:5f7d1086ff2c057bdf0f994aa520bb72ac8999409b40335a16ec948216548e17

Observation 57458f01-5592-4942-8de0-9254e0cc4ca3 · outbound

This paper cites TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:43.937656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:03:42.730202Z digest=sha256:214a8f268f34674c96771f57498e620b476865d52c71ae204ec8aa00aee60af0

Observation 3bbbec6c-3c3c-4d8f-bc36-205b1ddf5267 · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.818795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.818795Z digest=sha256:3d9ecbe9771e4733611df1eea267f1e097eb038aff70f3ddde41a48a4157a7b3

Observation 398b40fe-c1b4-4450-bb19-7ecc2e1db96a · outbound

This paper cites Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.921285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.921285Z digest=sha256:c6f772a94d32daa98dfc795cd8c2b418e37df90390cbbdc2ff48b6599381a121

Observation 4cf6cb1f-b73d-4ca2-81c3-6b1521989a1a · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.997502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.997502Z digest=sha256:cdb40a862af6c9980081c1eba035890a24228583fa858258f5280cbab65d7575

Observation 5f38af04-436b-4d55-9f49-64e217bb05fa · outbound

This paper cites Visual Goal-Step Inference using wikiHow.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Visual Goal-Step Inference using wikiHow

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:43.097879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:43.097879Z digest=sha256:1248cdbaf2851e6abbc8872be619a48fc12f2b7095dce43e8a394dc3b9413b15

Observation 8eeac2db-e9ca-433c-a022-eb5f4997b90a · outbound

This paper cites Cross-task weakly supervised learning from instructional videos.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Cross-task weakly supervised learning from instructional videos

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:43.631263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:03:43.173568Z digest=sha256:59027700003c1cacee05a95041f2e5bd0b78ef91601c467024b50855324ce7f2

Observation 81dd66e2-cc78-4e95-a2c1-1c620547de33 · outbound

This paper cites online" 'onlinestring :=.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion online" 'onlinestring :=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:43.313545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:43.313545Z digest=sha256:fb740c674dd0e5f47f4f3492d8cc051aec4ee4099f1a5eef520c210d1873e8bf

Observation b0e48172-576e-493d-863d-9155ee43bf20 · outbound

This paper cites write newline.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:43.408265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:43.408265Z digest=sha256:8beafd23bb20825405d5985f8ebcce871d103092020837d0e4e8b9fc23c6dfed

Pith citing papers

Observation 4ddd4061-9f18-4a7b-a125-b98aa83718fb · inbound

InstructionCrafter: Generating Consistent and High-Fidelity Visual Instructions cites this paper.

InstructionCrafter: Generating Consistent and High-Fidelity Visual Instructions $I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:38:39.703595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:38:37.497428Z digest=sha256:2bdf8f1b248a94246964554aa86e3e91b5bedaf4abc5ff41fad5dfff95d8717a