Pith. sign in

Paper Citation Record · LEDGER

MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2303.14389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.14389 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:47:30.325347Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:59:58.012604Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a3c115d1-95b5-476c-a5f3-0b8cf90d3e1b · inbound

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation cites this paper.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 280

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:44.675088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:e05311a5801c6be660117503cdb3343347c32105767318f066b13a41f68aecc8

Observation b60265b0-3ddf-4f78-a5ba-936a32777c4a · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:43:11.119271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:80bd1240862e2b5094df1d4216b56621296245f2ed97bac57222c525bf886da7

Observation c748cf1f-e842-4b8b-aac5-81ce88449306 · inbound

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think cites this paper.

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:09:37.128274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T15:09:36.982610Z digest=sha256:2a324a5e8d9714f2656acf2cba448ca780df08e3980610c3cdcfe3c6790c0c8e

Observation 9abfc1e2-5d1f-4142-a911-d84939b87c1a · inbound

Masked Autoencoders Are Effective Tokenizers for Diffusion Models cites this paper.

Masked Autoencoders Are Effective Tokenizers for Diffusion Models MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T04:47:30.325347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:47:30.325347Z digest=sha256:fa0ac5db551ab5b7c36dcb595974f1cfb3cb2d3cbe2a44d8aa385c2630159d17

Observation 9f270128-508b-4203-8517-6dbd5f9c8f3d · inbound

REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training cites this paper.

REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:00:35.728615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:00:35.728615Z digest=sha256:3c5a856eb5d8718af40366716d2075dec0e1cb3d18a49d5460460cbab410558b

Observation 8c35af6d-d189-4965-aad3-4868b574bc21 · inbound

Plug-and-Play Context Feature Reuse for Efficient Masked Generation cites this paper.

Plug-and-Play Context Feature Reuse for Efficient Masked Generation MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:20.756291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:20.756291Z digest=sha256:b0b80053b8b74a5b65a6de8c6c5fc907ed1c800208012fe3b720f75015643d89

Observation fca8e983-a16a-4156-8cd2-08fe6dec3114 · inbound

Native-Resolution Image Synthesis cites this paper.

Native-Resolution Image Synthesis MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:23.667405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:23.667405Z digest=sha256:ee3f57a8414d4267d9c636e81daec4f162f5f6a3d19de6f78e7175a2b1e448cc

Observation 401345ff-99d8-4eb5-b57b-6a59e26d4037 · inbound

Diagnosing and Improving Diffusion Models by Estimating the Optimal Loss Value cites this paper.

Diagnosing and Improving Diffusion Models by Estimating the Optimal Loss Value MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:07:14.043576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T09:05:58.516845Z digest=sha256:5ad8f78742f0cec0512dd527882e5f73b24c1f030e1fce76326aae0786f268bd

Observation 5989a483-4b75-4a5c-b7af-aa9ee27973ad · inbound

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis cites this paper.

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:49:40.899190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:49:40.899190Z digest=sha256:99f5b42ebafb91c31abd8899cb577e6a1db78d7f2690f798be1cf4f1d3005b47

Observation bd805cc5-1ac4-4910-ac5d-10bdbab896af · inbound

EscherNet++: Simultaneous Amodal Completion and Scalable View Synthesis through Masked Fine-Tuning and Enhanced Feed-Forward 3D Reconstruction cites this paper.

EscherNet++: Simultaneous Amodal Completion and Scalable View Synthesis through Masked Fine-Tuning and Enhanced Feed-Forward 3D Reconstruction MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:44.758445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:44.758445Z digest=sha256:3c00fa6ce0b4f63323f319d9b22d340a899827b482220efab5699c3caf69c687

Observation c12c6ebb-442b-42a5-9324-23bc55832735 · inbound

Improving Joint Embedding Predictive Architecture with Diffusion Noise cites this paper.

Improving Joint Embedding Predictive Architecture with Diffusion Noise MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:53.522159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:53.522159Z digest=sha256:ff1ac1d51e52417038ff6d4076360fb3d327a7c082652b74bb5081a7c026ae58

Observation db2af227-aaaf-4caf-bb17-4bf7d8774840 · inbound

Transition Models: Rethinking the Generative Learning Objective cites this paper.

Transition Models: Rethinking the Generative Learning Objective MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T10:19:54.345097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:19:54.345097Z digest=sha256:91cffb4150e44de7093e97ada2400480e3d1ef3f55b33c4ef6f988cf7ba66b36

Observation e3d926a7-1985-4081-add0-46666feb9b61 · inbound

REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion cites this paper.

REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T15:34:55.110363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:34:55.110363Z digest=sha256:0a4c8e352489cb93a18ad74d41083f240a7b5dd1daa83a1cdb6369ec4080a113

Observation 03f3bb6d-8818-4b00-a448-c7feabaf1779 · inbound

MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture cites this paper.

MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T14:52:58.486535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:52:58.486535Z digest=sha256:3b620d640ee6522999dd9c447b6f06fa56fecf9098d7a701963acab972a448f9

Observation c079214e-807c-49f7-9c58-21387021e06b · inbound

Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training? cites this paper.

Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training? MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:31.510886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:31.510886Z digest=sha256:a68ab5f28f532ee3086ebe817d59a8143de0de2bfc5be0ae4e0f16699c5f22cb

Observation 58183840-cbc9-43e1-9211-a898cf2618d1 · inbound

Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders cites this paper.

Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T01:25:17.988811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:25:17.988811Z digest=sha256:2f6aa1d3d9b497e3793a4f31ab61cafe104a45b72939b0bfb4256d5f8b508261

Observation 5c1518a8-dd5e-4994-8795-6d20b65f287c · inbound

Frequency-Aware Flow Matching for High-Quality Image Generation cites this paper.

Frequency-Aware Flow Matching for High-Quality Image Generation MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:00:04.070450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T10:58:05.541289Z digest=sha256:99334eaceae426669c40aac3ad72df3b416cd6fd1284eed247459cebf21605d4

Observation ac176980-f38e-4997-a44e-7c5816f76ab2 · inbound

Elucidating Representation Degradation Problem in Diffusion Model Training cites this paper.

Elucidating Representation Degradation Problem in Diffusion Model Training MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:26.493662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:08:11.110912Z digest=sha256:0c2ad78dc31caae8d2bf602e3f5348de8f584d76592fd71f7315e3a2dd12eacd

Observation a1027c89-f100-4471-a510-0f14b6e9707a · inbound

Beyond Point-Wise Matching: Structural Representation Alignment for Accelerating Diffusion Transformers cites this paper.

Beyond Point-Wise Matching: Structural Representation Alignment for Accelerating Diffusion Transformers MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:52:46.132591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T20:49:25.902880Z digest=sha256:d13dd677516b0844bfaa37fe0b1b2fd1213341ffb6d7b42426867b1a5ae27803

Observation 381547cd-0aad-44ec-9f86-95726728a654 · inbound

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation cites this paper.

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:03:15.293463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:59:54.139888Z digest=sha256:2d9fe90d4dd6ef67504274fda6b076c0a0508b1a512058f7dd0f9030fc64b3c3

Observation 51d63d71-cd4d-46bc-9816-ccfe3db4ca51 · inbound

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation cites this paper.

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.660682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T18:39:40.667006Z digest=sha256:5563ad62baf97a0f94322b947f3985dfb2fa6978c4fa37d011b7b7bac65eb412

Observation ec2e448b-0fb0-4faa-b06e-1d8544d558fe · inbound

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation cites this paper.

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T13:49:20.048386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:49:20.048386Z digest=sha256:b0e515118bac674a12cf685606525dd51d323247475fff5818c9b5ad7fd838d7

Observation 5b1e99e9-8407-44bd-b327-24cc0efafc44 · inbound

RiT: Vanilla Diffusion Transformers Suffice in Representation Space cites this paper.

RiT: Vanilla Diffusion Transformers Suffice in Representation Space MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:51:15.941712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T07:50:29.461854Z digest=sha256:82c758712d93e40fd7666dd8f8be8a583ce45d8467504e0e45780572227ba064

Observation c957eadd-8778-47bd-8db0-732198915409 · inbound

MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training cites this paper.

MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:37:26.189973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T18:44:41.266768Z digest=sha256:ade9248466d1fc1f7a3e320742d0011c18da980ee949870738f114bc211c2d0c

Observation 82f24cb8-1f4d-4109-877e-af91f694d9f6 · inbound

DiffusionBench: On Holistic Evaluation of Diffusion Transformers cites this paper.

DiffusionBench: On Holistic Evaluation of Diffusion Transformers MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:59:58.013972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T00:06:11.951205Z digest=sha256:5c150c78b01f5e809100b11d9d4865f1e9579f7480284f762c1e8a43cf0d8b2c

Observation 1a779831-6544-43a7-a186-39e6a22f8d61 · inbound

GEAR: Guided End-to-End AutoRegression for Image Synthesis cites this paper.

GEAR: Guided End-to-End AutoRegression for Image Synthesis MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:35:42.650156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T05:19:21.647714Z digest=sha256:7a764eab91864f133bb37447e437d4ccba9fcd9eab2b14e80df86bb24d66561e

Observation 36f981b6-bb10-4ddd-81f4-3ea6dfe240cb · inbound

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers cites this paper.

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T13:25:53.938768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:25:53.938768Z digest=sha256:51c346092f1961ad462731afbb3fe5b76657b45c839260eda8d507c7e28b3ccd

Observation 23ce7ac8-7ea6-49cb-957a-eaa4f6ab1f47 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-01T04:30:01.208956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:30:01.208956Z digest=sha256:8456d582eea5dd63150f6881a7877702efcbdb663527617e12ce3d4708ea6d23

Observation 9cad9d96-40dd-4e63-920c-63489b70685c · inbound

DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers cites this paper.

DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T01:07:12.008720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:07:12.008720Z digest=sha256:c6b0cfa7a940f0ebd698052851a8da69d6626226e322bc1a4cbb13a02072ffc0

Observation 271ee0eb-d2ec-493e-8114-f614818b8612 · inbound

Coupled Continuous-Discrete Generation for Scene Text Image Super-Resolution cites this paper.

Coupled Continuous-Discrete Generation for Scene Text Image Super-Resolution MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:57:59.676462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:57:59.676462Z digest=sha256:2ec4926bb00aefa2cad260d7df7b12bfc723b33c07e5a2ab916f71ec89f6f43d