Pith. sign in

Paper Citation Record · LEDGER

MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2303.14389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.14389 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T01:07:12.008720Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:59:58.012604Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a3c115d1-95b5-476c-a5f3-0b8cf90d3e1b · inbound

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation cites this paper.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 280

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:44.675088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:4e283fc8d12de5bda345aa8dc4aabccb333c2ddd7f74f86c774c9d98fbdaf2d7

Observation b60265b0-3ddf-4f78-a5ba-936a32777c4a · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:43:11.119271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:7bbe707767df234eefdc973db47e6b543e6cb4c9670465575f163de6104700d5

Observation c748cf1f-e842-4b8b-aac5-81ce88449306 · inbound

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think cites this paper.

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:09:37.128274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T15:09:36.982610Z digest=sha256:79f495052b3c15326e17cf5fddd5fb908bf1f19e226ab21bdd213b28081b4619

Observation 9f270128-508b-4203-8517-6dbd5f9c8f3d · inbound

REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training cites this paper.

REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:00:35.728615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:00:35.728615Z digest=sha256:3b1e6528966e11c4d610dd1dee8974980ddd55d7fde315382fa48de162ddcbdc

Observation 8c35af6d-d189-4965-aad3-4868b574bc21 · inbound

Plug-and-Play Context Feature Reuse for Efficient Masked Generation cites this paper.

Plug-and-Play Context Feature Reuse for Efficient Masked Generation MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:20.756291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:20.756291Z digest=sha256:114da80460cb3aef34387cf9f3f0e9ea72452d95ca86124274481b7197754f8a

Observation fca8e983-a16a-4156-8cd2-08fe6dec3114 · inbound

Native-Resolution Image Synthesis cites this paper.

Native-Resolution Image Synthesis MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:23.667405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:23.667405Z digest=sha256:dc32b936a1fe635342daaf777d7d281f92c54b44d6dc8fe0d40c0b721c2cd78d

Observation 401345ff-99d8-4eb5-b57b-6a59e26d4037 · inbound

Diagnosing and Improving Diffusion Models by Estimating the Optimal Loss Value cites this paper.

Diagnosing and Improving Diffusion Models by Estimating the Optimal Loss Value MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:07:14.043576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T09:05:58.516845Z digest=sha256:842bf2ccde593ab8b415a67f2aefe37d92958a4452da9fb2fa53a24405ed16df

Observation 5989a483-4b75-4a5c-b7af-aa9ee27973ad · inbound

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis cites this paper.

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:49:40.899190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:49:40.899190Z digest=sha256:b7dc47ee166a340af6414b3ad09751eed33b1e5ff238d338fa5226bda04052ad

Observation bd805cc5-1ac4-4910-ac5d-10bdbab896af · inbound

EscherNet++: Simultaneous Amodal Completion and Scalable View Synthesis through Masked Fine-Tuning and Enhanced Feed-Forward 3D Reconstruction cites this paper.

EscherNet++: Simultaneous Amodal Completion and Scalable View Synthesis through Masked Fine-Tuning and Enhanced Feed-Forward 3D Reconstruction MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:44.758445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:44.758445Z digest=sha256:56bef23c9b5a94c29cf54911054adb33dde4671035d93cb4e46a89c303b1c571

Observation c12c6ebb-442b-42a5-9324-23bc55832735 · inbound

Improving Joint Embedding Predictive Architecture with Diffusion Noise cites this paper.

Improving Joint Embedding Predictive Architecture with Diffusion Noise MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:53.522159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:53.522159Z digest=sha256:096960d9407c611ecbcabcae08f3a58d32cb62b8f3d26575970dd74b5e9e9ef4

Observation db2af227-aaaf-4caf-bb17-4bf7d8774840 · inbound

Transition Models: Rethinking the Generative Learning Objective cites this paper.

Transition Models: Rethinking the Generative Learning Objective MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T10:19:54.345097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:19:54.345097Z digest=sha256:1ee433d7634c30cfa4e8f5dca91b113712bdb1565b241c12fce1ee9a7988199b

Observation e3d926a7-1985-4081-add0-46666feb9b61 · inbound

REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion cites this paper.

REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T15:34:55.110363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:34:55.110363Z digest=sha256:7262c62736ca06b89ac659e046c6e9cdd582ff1dfb303b524e661f5e899e432e

Observation 03f3bb6d-8818-4b00-a448-c7feabaf1779 · inbound

MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture cites this paper.

MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T14:52:58.486535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:52:58.486535Z digest=sha256:8a0cdf99f0603404118e6c228ee737f90c4805fc7ff6204bc467a3f8d30b3ca9

Observation c079214e-807c-49f7-9c58-21387021e06b · inbound

Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training? cites this paper.

Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training? MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:31.510886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:31.510886Z digest=sha256:aae766eabac94d3865f8127b98e528191aa290098909995cbb5149dd2f7699f8

Observation 58183840-cbc9-43e1-9211-a898cf2618d1 · inbound

Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders cites this paper.

Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T01:25:17.988811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:25:17.988811Z digest=sha256:9010b602d48d907be489ff33e783619caec9029db2628980441a59283cc2acb9

Observation 5c1518a8-dd5e-4994-8795-6d20b65f287c · inbound

Frequency-Aware Flow Matching for High-Quality Image Generation cites this paper.

Frequency-Aware Flow Matching for High-Quality Image Generation MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:00:04.070450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T10:58:05.541289Z digest=sha256:450e67f4dcbe9e38aa6efcb1d816c608123ad674ac973e41ce59a2885db789cf

Observation ac176980-f38e-4997-a44e-7c5816f76ab2 · inbound

Elucidating Representation Degradation Problem in Diffusion Model Training cites this paper.

Elucidating Representation Degradation Problem in Diffusion Model Training MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:26.493662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:08:11.110912Z digest=sha256:7f0b2da61542d0a201ff62429524b458505f85e5d5811d0d007d509b89594129

Observation a1027c89-f100-4471-a510-0f14b6e9707a · inbound

Beyond Point-Wise Matching: Structural Representation Alignment for Accelerating Diffusion Transformers cites this paper.

Beyond Point-Wise Matching: Structural Representation Alignment for Accelerating Diffusion Transformers MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:52:46.132591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T20:49:25.902880Z digest=sha256:f74e6d95c5ee974dcd20e9233aee51097e10d0435322f0e135c40fe77bca0ff1

Observation 381547cd-0aad-44ec-9f86-95726728a654 · inbound

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation cites this paper.

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:03:15.293463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:59:54.139888Z digest=sha256:e4c3a8389a29138747c3ae2253d8aad062046fb331ef83fd65d1a1e77f1518a6

Observation 51d63d71-cd4d-46bc-9816-ccfe3db4ca51 · inbound

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation cites this paper.

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.660682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:39:40.667006Z digest=sha256:24c65964e076310e15a9935a914ccec87381c472efcb3acfa5acfd2ff38ca2ec

Observation ec2e448b-0fb0-4faa-b06e-1d8544d558fe · inbound

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation cites this paper.

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T13:49:20.048386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:49:20.048386Z digest=sha256:7dbb7a1ba24d4c4bb28355e7791f534871142d5c0ecd76f237a33ea4487ccc6c

Observation 5b1e99e9-8407-44bd-b327-24cc0efafc44 · inbound

RiT: Vanilla Diffusion Transformers Suffice in Representation Space cites this paper.

RiT: Vanilla Diffusion Transformers Suffice in Representation Space MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:51:15.941712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T07:50:29.461854Z digest=sha256:fbb901f9ad9f51d30bfdc9f8b1d17c101008b972adc96dd1b13c9b60aa9ecc21

Observation c957eadd-8778-47bd-8db0-732198915409 · inbound

MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training cites this paper.

MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:37:26.189973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T18:44:41.266768Z digest=sha256:c73df996264a572fdfaf79ddceb1b312da00596a0ff6d3c1ca3c28113197cefa

Observation 82f24cb8-1f4d-4109-877e-af91f694d9f6 · inbound

DiffusionBench: On Holistic Evaluation of Diffusion Transformers cites this paper.

DiffusionBench: On Holistic Evaluation of Diffusion Transformers MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:59:58.013972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T00:06:11.951205Z digest=sha256:d809ff59abf35e53313f85420d4f99ec9456ef5bec948f4ec20f794166cd07ae

Observation 1a779831-6544-43a7-a186-39e6a22f8d61 · inbound

GEAR: Guided End-to-End AutoRegression for Image Synthesis cites this paper.

GEAR: Guided End-to-End AutoRegression for Image Synthesis MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:35:42.650156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:19:21.647714Z digest=sha256:e125dc1b84a0723d27acdb532799aaabed3c7e306c857bafdeea2e0b8cd2a0aa

Observation 36f981b6-bb10-4ddd-81f4-3ea6dfe240cb · inbound

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers cites this paper.

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T13:25:53.938768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:25:53.938768Z digest=sha256:3318c6349ea5f167d21d59571edb5bb1111769006b3e1ea7c6ec4a954ae5bd0a

Observation 23ce7ac8-7ea6-49cb-957a-eaa4f6ab1f47 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-01T04:30:01.208956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:30:01.208956Z digest=sha256:e451e1ba9ffb8327a1df1bd8bcb79dfa93ea45ebef7572db4c8a3fbbe2fa4b64

Observation 9cad9d96-40dd-4e63-920c-63489b70685c · inbound

DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers cites this paper.

DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T01:07:12.008720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:07:12.008720Z digest=sha256:f7623d224f15d68b9f5660dd38f19b24f361dda3a7c31eb5e2cd4ee8a7e6aecb

Observation 271ee0eb-d2ec-493e-8114-f614818b8612 · inbound

Coupled Continuous-Discrete Generation for Scene Text Image Super-Resolution cites this paper.

Coupled Continuous-Discrete Generation for Scene Text Image Super-Resolution MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:57:59.676462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:57:59.676462Z digest=sha256:ec102c72a95a048cf88ecc48c1ceb4a4c7fb9de56c9b46fafc7921102457e840