Pith. sign in

Paper Citation Record · LEDGER

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model

As of 17 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 2 inbound Pith citation observations for arXiv:2506.12853.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12853 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:42:01.742085Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T08:06:27.191812Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:13:15.764565Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3483ca62-74a1-4e0d-84ba-3c21b377a598 · outbound

This paper cites write newline.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.389964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.389964Z digest=sha256:fa37fea66988e9176203eba7a9c566a84026ea0de388af64b8f5e527fc75bec8

Observation 53f16d1d-dc0d-4c48-abef-6605b054644b · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.453980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.453980Z digest=sha256:a4bb10b91e910f186a6575d8d1a4420ca2f43667ded5d69dc691f031a384471b

Observation 67c50ec6-a683-4cec-82d1-e7701c31d11e · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Scaling rectified flow transformers for high-resolution image synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.539849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.539849Z digest=sha256:0bf0fe3ce51ad3c0774d787bbe967838e35f1d50e4cc1953f6cf96cce30c9060

Observation 3e7f663e-c012-4bf3-9005-89fe071663b0 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.591297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.591297Z digest=sha256:cdd8819befc4352160a54c44feb4fbf90922e2568586714123b0d205128404c6

Observation ee97c058-2ea8-454b-a0ba-4012299d007b · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model LTX-Video: Realtime Video Latent Diffusion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.641831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.641831Z digest=sha256:a523eb51c28b80a3de4b3ab50fed41537b8717790bffac16f952aaff7c7efe4c

Observation 8caa6034-e3c3-437b-b8fb-59371aa0ed6b · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.695623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.695623Z digest=sha256:75c21b70b70c2f1e5ab50c7836dddc53852917972ddb60360eff9f111342d9db

Observation caec44e0-8720-4554-94b9-469e9107e8b2 · outbound

This paper cites DiffuEraser: A Diffusion Model for Video Inpainting.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model DiffuEraser: A Diffusion Model for Video Inpainting

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.744948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.744948Z digest=sha256:2e045b9dd3eeb87b1ddaaf8b8f283825aa7a14ad1dd7010eafa069d9ff54a0b4

Observation bb869301-bdcf-47b5-bf86-f559623a30b7 · outbound

This paper cites Flow Matching for Generative Modeling.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Flow Matching for Generative Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.808194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.808194Z digest=sha256:ed191bb68ad74dd5ab0f66d3c0420bcd5229dcea563b364f293284c7b628593e

Observation e518d056-4dfd-4cd7-adaa-13d0761ebfff · outbound

This paper cites Fuseformer: Fusing fine-grained information in transformers for video inpainting.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Fuseformer: Fusing fine-grained information in transformers for video inpainting

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:42:02.904955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:42:00.901619Z digest=sha256:83a2f1b04f0d63a7eb2fb64738557d241c627c3686800fb7b30428b6da4cdc53

Observation 54dc4eef-3242-4a4a-bc00-e7799952ecbd · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model A benchmark dataset and evaluation methodology for video object segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.948590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.948590Z digest=sha256:ed8acce680d43048bb7685bc19163ae3e201097ef6d5d3d540821b79ba15ed73

Observation 91ae787f-d58e-4b4f-b9fe-9062b07118ce · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.013801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.013801Z digest=sha256:7b18d0565dd470b8bcfd3fed4d2a86483eaa029b7e6d41078863262bb9435001

Observation 5d89be9f-e883-4ec5-82b5-ac9fe4c4b544 · outbound

This paper cites Sam 2: Segment anything in images and videos, 2024.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Sam 2: Segment anything in images and videos, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.099668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.099668Z digest=sha256:48a2e8fe1b8ab9099d12c781589c7d0e3d9a0f9bb35b7524e94b1ec3396f83aa

Observation 75211975-9bbd-4e64-b978-3649f0ec11d7 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.243679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.243679Z digest=sha256:cbc5e65abf8375cb7ca7950c29c78477ab92f9f1a039527c91af332ac2bf5282

Observation 109b83bc-d7bc-42a0-a6de-a15e43cc009d · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.340005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.340005Z digest=sha256:787c77bc55cd60333bfb4194b3e4749d280ddbede48db180dfcb9b0c0e16363f

Observation 2fef8977-4eda-493b-9afd-2474a34a432e · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.390837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.390837Z digest=sha256:e12c005f58f5404a716f37eb11edd3b49ea4c684b69e18c89c469b8b5ef57256

Observation a2dd6f99-90b5-4012-99ed-15d4a0f124de · outbound

This paper cites Learning joint spatial-temporal transformations for video inpainting.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Learning joint spatial-temporal transformations for video inpainting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:42:02.553991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:42:01.526322Z digest=sha256:73bde62677dbb5540a3f125889eda9b837454c00a2632f06c890705146272ecf

Observation b668a2d2-770f-4aa4-8e00-51b4780f253c · outbound

This paper cites Propainter: Improving propagation and transformer for video inpainting.

EraserDiT: Fast Video Inpainting with Diffusion Transformer Model Propainter: Improving propagation and transformer for video inpainting

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:42:02.238566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:42:01.742085Z digest=sha256:7041caf616d93c4919fafe4d80890ba390ded99dd6db2f9a1f8c144b6172437a

Pith citing papers

Observation 97933a7e-95b9-450b-a998-2f41a4d87ec0 · inbound

CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal cites this paper.

CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal EraserDiT: Fast Video Inpainting with Diffusion Transformer Model

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.679421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T00:39:15.790131Z digest=sha256:4949eedef2d03f711112737492f62f2ed87621ca134dfd82167feea8044f334c

Observation c6e7c9ce-3d05-4f72-b280-60117090d07c · inbound

DualEraser: Joint Video Object and Effect Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver cites this paper.

DualEraser: Joint Video Object and Effect Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver EraserDiT: Fast Video Inpainting with Diffusion Transformer Model

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:13:15.766085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T08:06:27.191812Z digest=sha256:72ff4bbe5083230dfdd4e415ae5ad5c38caa381c62d0ae337ab9d71eab2bcb16