Pith. sign in

Paper Citation Record · LEDGER

Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2401.11708.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.11708 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:06:52.960237Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:23:28.257830Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5f5cdb6a-875b-468f-8208-44001f9fbd15 · inbound

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment cites this paper.

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:43:03.480198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T19:43:03.310237Z digest=sha256:b48bb3285b03a4bb99c2208e25505bdc61d4e2f8b55774ea3ee1e80257317ca5

Observation d2c356f3-c195-426d-8230-ba8200344824 · inbound

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding cites this paper.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.473771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:8ae453a5032fc4c383067d490e6817e7ed46d97342e9542d373003e31b42feed

Observation d2dc2bf1-886b-4bd2-a3e3-16813d97853e · inbound

Step1X-Edit: A Practical Framework for General Image Editing cites this paper.

Step1X-Edit: A Practical Framework for General Image Editing Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:36:41.851013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T14:36:41.467429Z digest=sha256:351e8bdb75276b4c45ad649870e30b8987814b626a0063912f2412aad83bf3f5

Observation b3b398bc-8821-45dd-bc29-fee3c6833b06 · inbound

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion cites this paper.

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:52.960237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:52.960237Z digest=sha256:d12e9ef088a3f3cfb3c2003cbfb273aa8f4ed6e94a716b324450a081efe988b4

Observation 71ba07ae-22b9-49a5-809f-62d0ccf92dca · inbound

Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization cites this paper.

Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:04:55.387035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:04:55.387035Z digest=sha256:66d82aba24b8b14f466895472b86f7ffc783c2a075dfdbf901256080440ed6a3

Observation bc27c992-506a-44c4-8192-8bdb4f0a89db · inbound

Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers cites this paper.

Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:43.989620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:43.989620Z digest=sha256:d63c9cee02ea3a546daf08abf7e4cc1e902e297a9cd3c4fb6028e4a4fdaba0ca

Observation ff379478-dc49-45a9-9c2d-0c34e5ca44a1 · inbound

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models cites this paper.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.923723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.923723Z digest=sha256:e733f833b6887e9237005fb68cf2304e265466587b6f4a787f516519e4015ad8

Observation 276680a2-f601-4dcf-bb86-2f6e84230380 · inbound

Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization cites this paper.

Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.259292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:15:24.299457Z digest=sha256:1e1e606e0f4a6f9b3c0a2494cde58a42cad5fc040dab3a62e22a422af6915a04

Observation fc73f304-751b-4432-b8c0-7f4ee8713757 · inbound

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation cites this paper.

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 232

Resolution
unresolved
no resolver link, observed 2026-08-02T06:23:50.288678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:23:50.288678Z digest=sha256:e87896c0c0a08db912f29c1264b5e20b4fb4cc59f3bbb0f2fa8ee9242eadb7ba

Observation 529408f0-8e0a-4e04-84c2-4757a6f8e4cd · inbound

TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward cites this paper.

TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T13:55:32.972530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:55:32.972530Z digest=sha256:270cecd97beb2c83bfd12103e5aa8ba5bb2403ff3af1dd08ad1e383898f688b7