Pith. sign in

Paper Citation Record · LEDGER

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation

As of 15 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 2 inbound Pith citation observations for arXiv:2412.12391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12391 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:11:00.639240Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:30:31.258053Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6eb7e9bf-7027-4ae4-9872-dd1ec108ed8a · outbound

This paper cites [Online; accessed 4-March-2024].

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation [Online; accessed 4-March-2024]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:01.006837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:11:00.535053Z digest=sha256:4fac8b4da6472e1b195b6b6b482aea11a921240341c6266db388cccac0999cf5

Observation 65abf0c2-4b3b-4339-b2be-ae7f0ca877a8 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.540025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.540025Z digest=sha256:55972f69afc2e7335c1b0c2f2e2c8ab1e9e7bf1b09ffe226129a931e9fe776b1

Observation bb0157df-e0aa-498e-9542-d9fb1d151351 · outbound

This paper cites Mistral 7B.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation Mistral 7B

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.556269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.556269Z digest=sha256:dd5471bf76137d601db0c673b29ed10128b64540d7f0dd1a3addb1548dee98b3

Observation df311dd2-2ef2-449c-9626-236fec9ce460 · outbound

This paper cites Scaling Laws for Neural Language Models.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation Scaling Laws for Neural Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.561565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.561565Z digest=sha256:7305b0d055369d99b00dfb88d57bd35701ce6d1b288ed06ac84aa0376eab772c

Observation cb023b8d-cfb7-4994-8ec8-83b8fd4eff4d · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.567171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.567171Z digest=sha256:fc4a05ec33ec32457b54f18ec499a237a1ef266abf5b8110c6b20425e177e960

Observation 3028d1b1-2011-410d-8a7c-5f5fa92057f3 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.572731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.572731Z digest=sha256:0f7c0abe8e718e8e383d4d9806c861c5cc7e8d1fa33a54c1751cd1a266fbbaaf

Observation 09efea52-30c3-46c0-9fc2-2064026aefee · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.577563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.577563Z digest=sha256:635fdd3021ddf98da23292548760fbe28d30344c85b0bb9b1b1227d7753a970e

Observation 14bd2e94-787a-4f10-8185-6e0c67911dc0 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation U-net: Convolutional networks for biomedical image segmentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:00.989413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:11:00.583441Z digest=sha256:d0017071b9741e8db180b5c8263e000ae8072825faba28e6962e9a9e9df0d000

Observation 7df0b84d-929d-409a-8743-63c36ea9c0a2 · outbound

This paper cites GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.598263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.598263Z digest=sha256:7eb9b628b60fcb4abf35757480894e54eb19bc0c1f7736dbc38bd0595b33a75c

Observation 710f86f6-d71a-4824-8142-9e6a882a290c · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.604152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.604152Z digest=sha256:b0cd4c8c6d73988640ba729e3ff29a2ca6cce7bfa3d38a5b10da77d0ef1454a5

Observation df250891-c198-42ef-8cf7-c91b44fff2f3 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.609126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.609126Z digest=sha256:1fe26bc21cdc1666a3f9c7679fd435391cb7657abf623696a3d076d313a25166

Observation 3a17937c-00ab-4842-8bb1-8e8b85577898 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.614058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.614058Z digest=sha256:4586f9df9c07f819ca8e5932cfabaa94aa24b2271be72b57a6004798dbf7b321

Observation 699907c1-0d15-4c46-a919-6f33a9bda6c4 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.618880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.618880Z digest=sha256:a051cb4c1d59b685913dbbce2465f4e4f1c999072bcc58dfad8813c02e4de391

Observation da66e0fe-1c08-4771-96c4-71e9a005ca97 · outbound

This paper cites LensArt consists of 250 million image-text pairs, carefully selected from an initial pool of 1 billion noisy web image-text pairs.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation LensArt consists of 250 million image-text pairs, carefully selected from an initial pool of 1 billion noisy web image-text pairs

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:00.972905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:11:00.624731Z digest=sha256:3f9bc8e01764bde9841e1a7d82b35fa1d47567742d904e010581630b9356b7b7

Observation 5ba47946-b5f7-4d59-a311-a52ff46defe1 · outbound

This paper cites B.1 B ENCHMARKS We outline the evaluation benchmarks used to assess the performance of our image inpainting and canny edge conditioning models.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation B.1 B ENCHMARKS We outline the evaluation benchmarks used to assess the performance of our image inpainting and canny edge conditioning models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:00.956046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:11:00.630046Z digest=sha256:b448f06decc85da3505fdb5b6c11385b331e7bc22fde546d3ba9c826dafe874f

Observation 0b2474e7-bbe9-4e35-b4b1-4e433df4a3e6 · outbound

This paper cites an unresolved cited work.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:11:00.939986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:11:00.634831Z digest=sha256:c2d52825600790bf92a749e10114802427f99381e8c67a74ea0b6b9609d23023

Observation d2d2f455-3e5b-43d7-beaa-68fe9dad98db · outbound

This paper cites In addition to TIFA and ImageReward, we also provide the FID score, which measures the fidelity or similarity of the generated images to the groundtruth images.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation In addition to TIFA and ImageReward, we also provide the FID score, which measures the fidelity or similarity of the generated images to the groundtruth images

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:00.924216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:11:00.639240Z digest=sha256:f722b8dd1846648c73491ce254e96736232b1ac1c8df110b89c945499589704e

Observation 001bea12-8fd3-4ba3-bce3-ce49bc74b0c7 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation LLaMA: Open and Efficient Foundation Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.593507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.593507Z digest=sha256:204c8f9b038fd4c787e43f971bf260d1ee50a340efe93cac799c7dca13bcd010

Observation 30cca4b0-6871-4574-a296-a74f49fb3437 · outbound

This paper cites TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.550971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.550971Z digest=sha256:19bc3e24feb2dbac51078f1f908039b1b7144f9b959e30367630f719d1effdad

Observation 99a30b2d-0b47-440f-911c-a8e9c81fdd8d · outbound

This paper cites Denoising Diffusion Implicit Models.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation Denoising Diffusion Implicit Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.588431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.588431Z digest=sha256:2c9013a1920f02c1523d5c2c25fe3617433cd3586a3d3a110fa9da057dadce40

Observation e79e9040-d0fe-4f84-88b8-78fb9ef02d14 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.545632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.545632Z digest=sha256:b872d5be74e745a26b499ed765213bf6282cc608f082e3c1148b83dd9ac79393

Observation d23cdeee-dccc-42e5-91c7-345a749c17a9 · outbound

This paper cites On the Importance of Noise Scheduling for Diffusion Models.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation On the Importance of Noise Scheduling for Diffusion Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.529321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.529321Z digest=sha256:bb5a9e4c179759647abe889dc504e5d626e2214acfca48cc75c158073195e7b6

Pith citing papers

Observation e89e2380-e22a-43ab-8d62-db81ad126e7e · inbound

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer cites this paper.

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer Efficient Scaling of Diffusion Transformers for Text-to-Image Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T20:30:31.258053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:30:31.258053Z digest=sha256:ec340288c223f9d6784ba1456fa0f1da1c27aa21dccd19fd48bca1c55eebf501

Observation 52761174-4d24-43fc-a1f9-7293d6313353 · inbound

Importance-Aware OBS Pruning for Diffusion Models cites this paper.

Importance-Aware OBS Pruning for Diffusion Models Efficient Scaling of Diffusion Transformers for Text-to-Image Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T10:59:49.288561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:59:49.288561Z digest=sha256:7e1a4db666ce6171711098b7a9421f7154aa2fe148731904c5e3d0d492d81fb6