Pith. sign in

Paper Citation Record · LEDGER

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion

As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.16425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16425 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:03:43.408265Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact8
  • verified fuzzy1
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9bc41830-0ba5-4bc4-a162-293bc1e8b35c · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:03:46.264468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:03:35.493664Z digest=sha256:5ef90184ac03b7c220b9a14538f558c4f23f4d31afb0d8aab02fb2c80b3345a5

Observation d3a93cff-b7ce-4f0b-b53a-36e1d7befe0e · outbound

This paper cites Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:35.615548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:35.615548Z digest=sha256:09c4f4e9e5edea229bf4dcb954645dd87a41ee91ce460960f22d7af9ffab9f0b

Observation a4276df0-c917-4904-8060-67d8944eb415 · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:35.811186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:35.811186Z digest=sha256:d6de6b82c07e35bee16c5985adf938fbc02c60f4f96b535e82cdb3a0b727819d

Observation 76120329-8559-416c-9ba0-0a9d28055e15 · outbound

This paper cites https://api.semanticscholar.org/CorpusID:264403242 Improving image generation with better captions.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion https://api.semanticscholar.org/CorpusID:264403242 Improving image generation with better captions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:03:46.036153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:03:35.945204Z digest=sha256:5d17cb2c8a870bc6db9990de24998fc7bc0c45349e853113c15eb8a4b064cd88

Observation 55672839-4837-46ca-952e-2ce7a32bfd7e · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Training Diffusion Models with Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.086633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.086633Z digest=sha256:9a7473ec8729df6e82deb3e1c32933ced9a89f4ee22ca1567ef9f3e72f5b92d3

Observation ff16a6bd-03ce-4f03-aa78-5b30b366f4a6 · outbound

This paper cites On the Design Fundamentals of Diffusion Models: A Survey.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion On the Design Fundamentals of Diffusion Models: A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.217850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.217850Z digest=sha256:c15185c73c5f38cb7d2778c4de75290d1e481a922f355cef6eb83ff637de4bf3

Observation b34f1779-4abc-4d04-83ea-91e003946f69 · outbound

This paper cites MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.345971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.345971Z digest=sha256:88a1465f32ce431470f721de1324d463f440d7d54f1d939dabedc024f9c0880d

Observation 54cb629c-fbd1-44d4-98b1-aac0a7fb722b · outbound

This paper cites Evaluating Text-to-Image Generative Models: An Empirical Study on Human Image Synthesis.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Evaluating Text-to-Image Generative Models: An Empirical Study on Human Image Synthesis

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:45.649994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:03:36.484830Z digest=sha256:533a88c4cb1abd412a27a0dc3bd20928e1e1f6f45cce7c5b4816ba45dbb5cf33

Observation 9b4aefc5-8d71-4ab5-8f5c-9914c9948280 · outbound

This paper cites X-IQE: eXplainable Image Quality Evaluation for Text-to-Image Generation with Visual Large Language Models.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion X-IQE: eXplainable Image Quality Evaluation for Text-to-Image Generation with Visual Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.621699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.621699Z digest=sha256:d676e28aab9b0c488ead95a88c6a3f5672f2c790de950bb5f684995dfb3fd0b7

Observation 382f8147-2089-4b7a-b9ed-18b03b123820 · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.759018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.759018Z digest=sha256:ffe4735eae509901b0225aacfa76fb7e2abb310356df445dc5a1cddfe0506b85

Observation 0edff757-7773-4372-b5d3-9a975d96818b · outbound

This paper cites Directly Fine-Tuning Diffusion Models on Differentiable Rewards.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Directly Fine-Tuning Diffusion Models on Differentiable Rewards

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.864834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.864834Z digest=sha256:aefbbdf122ce2f5419573f658dc669f6962abb72e1858bbf8bf2adeedbc00f02

Observation 429222e7-86b9-484c-bfba-d965c4e25d92 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:36.999362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:36.999362Z digest=sha256:2e983d67421523f3d2070efd1dbe5493cf7821ba90a1f2c83ba7a6712a964080

Observation 7e163a25-9377-4b98-83b3-5fe7ae6f4acd · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.111417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.111417Z digest=sha256:6c056a9eb82c04b41953eefb7aeb0f7084c9c424fe711894a136a60a39f313eb

Observation dda87e0e-175e-4c16-89a3-0cd24a0247a1 · outbound

This paper cites DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.181408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.181408Z digest=sha256:36e4014a334726eb5c3a98b1d11f6c75e8ca525bd7abb6e7f2e703a5732e727c

Observation 8fc889f0-a173-4fad-8d21-1e5a948403a1 · outbound

This paper cites Learning to Segment Actions from Observation and Narration.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Learning to Segment Actions from Observation and Narration

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.252800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.252800Z digest=sha256:a5e76b89da13fd21c1343b0aca7f8c6c1d757303b570a33532110f879eb02f12

Observation 6cf01b20-c326-470f-b85c-98df55b90754 · outbound

This paper cites GPTScore: Evaluate as You Desire.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion GPTScore: Evaluate as You Desire

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.396043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.396043Z digest=sha256:de57a790821552341f5ce86deebdb0c0b1cf1be945fbf7d07ff407712a15f557

Observation a9a1f61f-aa72-46a9-823b-aadd3484262d · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.532493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.532493Z digest=sha256:72452c352c381fbaa42ae658eb48081bfad0afa2442b34224e5436e89344b4b1

Observation 66908ad7-fc0f-4d95-9147-fc4620f8c7f4 · outbound

This paper cites COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.652640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.652640Z digest=sha256:123a2bbf82e106a6c4fbc0a7cde3725fe2abfb05f18319825a5cebab80c79d1e

Observation 75349684-e6b9-4c77-8b2e-337290b9b556 · outbound

This paper cites Temporal Alignment Networks for Long-term Video.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Temporal Alignment Networks for Long-term Video

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.830852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.830852Z digest=sha256:c2a412b2edc4ffe8ad0f8284aae827fb9e1a3cbf964362d30a0c52b2a9947420

Observation 80ce3f2d-e36e-4972-a64b-0bee5c49a276 · outbound

This paper cites Optimizing Prompts for Text-to-Image Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Optimizing Prompts for Text-to-Image Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:37.951089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:37.951089Z digest=sha256:9e61a02da3afe526798bc240d7e78602399397a2965a4c1604649598e2e8241c

Observation 16ed4664-fc04-4e15-a13a-c6d6b349a2c4 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.064399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.064399Z digest=sha256:ce97d23a5f28b3b878dd8e5d430006d2f77ba70bda5517ffb4c9c4d3fc07ca0a

Observation 3c6c03ac-590d-45eb-b9a4-985275f1e651 · outbound

This paper cites Denoising Diffusion Probabilistic Models.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Denoising Diffusion Probabilistic Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.202182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.202182Z digest=sha256:316cf4113a96b7e6a748d4467b94001fdfaacf9b0638e1d81e5e704a6f30cdc5

Observation 9b378390-b238-43e1-88f4-8300d51e0654 · outbound

This paper cites Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.348720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.348720Z digest=sha256:0852fcf33f3710d1fe47a5e444ecdd8af398c7a555c7f07b389113fb0dbeedec

Observation b99cc7aa-52ee-4e2b-a7f6-f8fafb64ddfe · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.529587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.529587Z digest=sha256:3d0132b9d5ada97a2f083ffff02268ab861c00a090977ed2689f20a6313e01f7

Observation c9e1149b-1b73-4a62-bc96-adaa729df42c · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Aligning Text-to-Image Models using Human Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.645140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.645140Z digest=sha256:ec39e01358aaa2ad5fd97692cb320ef247b7ee32d4aadcf6481678c5d6286fbc

Observation 55fb65d7-3118-4f58-81b1-93fb4e89c5a1 · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.774002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.774002Z digest=sha256:aad0a75ab7baad36f368f80b588258c255f6d9b0414817209dada3bb9703968b

Observation ebb8b2a7-fe66-4f67-9d37-b9d5ebc0c942 · outbound

This paper cites Aligning Diffusion Models by Optimizing Human Utility.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Aligning Diffusion Models by Optimizing Human Utility

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:38.853890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:38.853890Z digest=sha256:d15ea128274d08b90b8a54e27b07a9aa570bde8ca5a946be6118bffa4c76b33d

Observation da03b956-f681-4ded-898d-e2bdeb741a3e · outbound

This paper cites Learning To Recognize Procedural Activities with Distant Supervision.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Learning To Recognize Procedural Activities with Distant Supervision

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:45.262945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:03:38.930421Z digest=sha256:4bfd5a586079d826b6d2eda64d9cc5dc93918c4605dcebb4daaf9d914e2f3b23

Observation 4436ae88-6107-41d9-8836-d8f293434a2e · outbound

This paper cites VideoDPO: Omni-Preference Alignment for Video Diffusion Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion VideoDPO: Omni-Preference Alignment for Video Diffusion Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:39.100590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:39.100590Z digest=sha256:2f67e3a435f5eac5ab6f2d811a108211408bd240fb21ef2f5d9d50451ed2c302

Observation a626065a-b04c-415b-9b38-f11e173b044c · outbound

This paper cites UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:45.058912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:03:39.251120Z digest=sha256:78808fb91e8c8e3ec9226a1d98ec42ce90c701508da12f2da2d319ecce8d1ba5

Observation f51cdcfd-4231-4f1c-a9df-49631214729d · outbound

This paper cites Solving New Tasks by Adapting Internet Video Knowledge.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Solving New Tasks by Adapting Internet Video Knowledge

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:39.351300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:39.351300Z digest=sha256:b712a12d22f78b2ce87ad65f38eb969c1be980a0677a481cb3ebe18247b1a796

Observation ef3a9452-43ae-4d55-8aff-20418f6b7e09 · outbound

This paper cites Generating Illustrated Instructions.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Generating Illustrated Instructions

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:44.923870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:03:39.512797Z digest=sha256:e7388d592cfd78307e3019e1a823ce33356e53318294083e8a1ea74ee2343229

Observation c57b2b1f-fc7a-4ba2-ba09-529105cae8ea · outbound

This paper cites End-to-End Learning of Visual Representations from Uncurated Instructional Videos.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion End-to-End Learning of Visual Representations from Uncurated Instructional Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:39.673028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:39.673028Z digest=sha256:48ce9d43d5cfa2e4dcbcbbcc140e28247bbf6c5fb6d228ed0bfa23754c61e482

Observation 4227ee8f-183d-4e50-a91a-164bf27a4b8e · outbound

This paper cites CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:39.803189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:39.803189Z digest=sha256:caf2cdf07f17ddd829b99cca5801786c37d71d3f54efe022c586f43703aac0cb

Observation 9d037a58-593c-42cf-ae70-d7a43bfb84e3 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:39.954864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:39.954864Z digest=sha256:20af908c75deedf37694b0b0fb0d77a0e9f614a8b1ce513e5aab20633ce09ef2

Observation d56fca1a-9741-4051-9fe7-488528d479cf · outbound

This paper cites Aligning Text-to-Image Diffusion Models with Reward Backpropagation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Aligning Text-to-Image Diffusion Models with Reward Backpropagation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:40.138551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:40.138551Z digest=sha256:c62b84f67bd34006692f5bbd1d85d77352d9adf5308709514470d8fc4f8a071d

Observation 0d24491a-98fb-4bce-ba9f-fa3b4837b96e · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:40.248662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:40.248662Z digest=sha256:2ae9fd48fd2b539a5fdb33e06cbe65c5d1af048c8e658706e53c909d5178292d

Observation 6ee5cd23-a80c-47a9-ad2f-2bc5183b76e4 · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:40.416719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:40.416719Z digest=sha256:fbdc3d05219dcf464677814215b8285b681d5bc4a2f2e34b2cbfb19a1fb332b2

Observation 5a2974dc-3e1b-4f5e-aed9-c8f1c40f29e3 · outbound

This paper cites Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:44.698797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:03:40.666014Z digest=sha256:ae207296f7b711b65f4fb6015ef95220461ea0d1f2d80b1845748b3bb9b4d920

Observation 0dbcb19f-e3a2-4a33-be81-10c4e94f8ce0 · outbound

This paper cites A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:40.893656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:40.893656Z digest=sha256:8979b89a214e2977b5fcd52101a6a14dd860e880943a42dd5b7ef983eae41292

Observation 0169d1e0-fe34-4e79-b267-84e931944174 · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 41

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T15:03:44.551874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:03:41.027543Z digest=sha256:adf5795b9b2139daf0368fe3d70e04892269c8412a7377394c58689900c65748

Observation 2002d8ad-96a8-4d23-92d0-74e8c72a7948 · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.212373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:41.212373Z digest=sha256:344f0a5a0bddc1512e8293e92bf8936f6f1dd685148a520db0cc1af605dbc50c

Observation 416b32be-dc42-497c-b8b6-8a5c08946c2c · outbound

This paper cites Diffusion Model Alignment Using Direct Preference Optimization.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Diffusion Model Alignment Using Direct Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.362858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:41.362858Z digest=sha256:bbd73d75df2e5b7c4afa501384beeec6e1c78894c494d9c324360b947cb2525c

Observation e384682d-fe63-4e4f-81c2-40d84d611c53 · outbound

This paper cites Is ChatGPT a Good NLG Evaluator? A Preliminary Study.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.497625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:41.497625Z digest=sha256:16d3eec2fe39032606470fa63ee0dc453c273ceb56601031361ef80eea06f093

Observation 3c41d0b5-a2cc-4e5a-894e-500555dff40a · outbound

This paper cites A Unified Agentic Framework for Evaluating Conditional Image Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion A Unified Agentic Framework for Evaluating Conditional Image Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.644976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:41.644976Z digest=sha256:d2a8ee3fe9bfbdd39a9aac1294869bb5d285205771aa265baa0c6db333b063ad

Observation f90e72b7-e5d5-4bbf-9174-cc80c38e0330 · outbound

This paper cites Large Language Models are not Fair Evaluators.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Large Language Models are not Fair Evaluators

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.863342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:41.863342Z digest=sha256:d0daf9a82e71ab4ae127b59d50fd7954f88aeb1776ecb5235e18d52bf33a95ab

Observation 7865345f-9366-4fe4-86b5-9feea14bb89e · outbound

This paper cites InstructionBench: An Instructional Video Understanding Benchmark.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion InstructionBench: An Instructional Video Understanding Benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.979783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:41.979783Z digest=sha256:f09734d37f283a810d4092abd12a098f9acf937a6c6f3e311e47e24009e3e4f3

Observation cc03890c-d637-4de9-869d-5b6a16afdb0f · outbound

This paper cites Preference Alignment on Diffusion Model: A Comprehensive Survey for Image Generation and Editing.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Preference Alignment on Diffusion Model: A Comprehensive Survey for Image Generation and Editing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.075088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.075088Z digest=sha256:43cab1154ef9d01a56c0041bed643d1f9a52cc847b2568eeab62db9b5831c975

Observation 997702fa-f5f5-4326-b142-5453fbd1c112 · outbound

This paper cites Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:44.113000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:03:42.159502Z digest=sha256:0847da7fafda89945b7a50ede07009675d46d3b82ee4c1dc8e2aa00b2ee3e650

Observation 832f6dee-fc45-49bd-a327-c598cc61b5ce · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.304578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.304578Z digest=sha256:2e3d10e0b0e0553be331872ba1f714df66d19e18e75a301cba599a5bf3229016

Observation 6c57587c-f0ce-4322-b7dc-6e32063275de · outbound

This paper cites ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.406631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.406631Z digest=sha256:aa61cd65d7cfedd0399d0c06e10f1f3ecfffc90615251a3462177a2523a06c4c

Observation 659a86de-e39c-4353-bd2d-05ab46ad2efe · outbound

This paper cites GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.511905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.511905Z digest=sha256:e3fc378319945bc5c4f9df7873da7bb6ed13a0e98f75600c33132a6e22b52197

Observation 12ee1582-3187-4244-a65d-6a713db5d35d · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:03:45.870174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:03:42.625227Z digest=sha256:b8722bfac597235bf54c6fa9fe51787fc93aa17e302e4ac04c5752edde09ff05

Observation 57458f01-5592-4942-8de0-9254e0cc4ca3 · outbound

This paper cites TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:43.937656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:03:42.730202Z digest=sha256:2dcacd862d8dd303453ca525a041138b065168c47f77c9a8389df4abb85ca5c7

Observation 3bbbec6c-3c3c-4d8f-bc36-205b1ddf5267 · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.818795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.818795Z digest=sha256:330fad7b1f0b41472603ab8ea409925873e27d6ce2b4c608c2bc2800b5c88f88

Observation 398b40fe-c1b4-4450-bb19-7ecc2e1db96a · outbound

This paper cites Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.921285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.921285Z digest=sha256:90d857dd1e50963d0e316ecf1f7e13aa19a94eb747cab856ea0cef024c00893b

Observation 4cf6cb1f-b73d-4ca2-81c3-6b1521989a1a · outbound

This paper cites an unresolved cited work.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.997502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:42.997502Z digest=sha256:df32bcdba191c480057759ff3a80e7c7c73dd287e96bec163d4791952f2e1b96

Observation 5f38af04-436b-4d55-9f49-64e217bb05fa · outbound

This paper cites Visual Goal-Step Inference using wikiHow.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Visual Goal-Step Inference using wikiHow

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:43.097879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:43.097879Z digest=sha256:71d72a19082847a6d2bfc9c45690fe95e3d231067b57f24c1c56ee1b3ecc9d84

Observation 8eeac2db-e9ca-433c-a022-eb5f4997b90a · outbound

This paper cites Cross-task weakly supervised learning from instructional videos.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Cross-task weakly supervised learning from instructional videos

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:03:43.631263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:03:43.173568Z digest=sha256:a2908d2ab394c2662c9e8a917668694e8960c482466a7e4617232e8d10d8ab4b

Observation 81dd66e2-cc78-4e95-a2c1-1c620547de33 · outbound

This paper cites online" 'onlinestring :=.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion online" 'onlinestring :=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:43.313545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:43.313545Z digest=sha256:b0863125ba550870fab8a68f990d51bb3e1ac55e06a8463409309983529c7637

Observation b0e48172-576e-493d-863d-9155ee43bf20 · outbound

This paper cites write newline.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:43.408265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:43.408265Z digest=sha256:87b5a33ff5264dd9cc90ac25f780e2858efd35d56638ea3bb4315b20178763e1

Pith citing papers

No inbound Pith citation observations are available.