Pith. sign in

Paper Citation Record · LEDGER

Playing with Transformer at 30+ FPS via Next-Frame Diffusion

As of 23 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 7 inbound Pith citation observations for arXiv:2506.01380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01380 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:49:46.876113Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:43:37.510074Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:54:57.961666Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 244b70f9-ebf8-4ce1-8757-fca7a6da178c · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:41.669315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:41.669315Z digest=sha256:14aca92e4d938227b3d3dd9adb170dc2bea292bf466285c9d0f18953c7e0e589

Observation fb49a71e-b600-4a27-bc94-5c19ddb3f2af · outbound

This paper cites Diffusion for world modeling: Visual details matter in atari.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Diffusion for world modeling: Visual details matter in atari

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:41.724851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:41.724851Z digest=sha256:eab1a3a7875953133d65c33313a112d5f49defac2e4e87065981fee3e2bab3a5

Observation 01b31326-a147-4769-abcc-850662967b95 · outbound

This paper cites Block diffusion: Interpolating between autoregressive and diffusion language models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Block diffusion: Interpolating between autoregressive and diffusion language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:50.211033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:49:41.780871Z digest=sha256:ecad82f6c7b11ed606759cca2827d57d4a18e2a80576c899def2ba77636b9379

Observation 5052c2e3-eb6e-4d5b-ad36-2245576cd7ff · outbound

This paper cites Video pretraining (vpt): Learning to act by watching unlabeled online videos.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Video pretraining (vpt): Learning to act by watching unlabeled online videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:49.982602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:49:41.831486Z digest=sha256:4a8519b59991dd10ed10e1161b5eacd8110581d4502d5c2cf144c3ec9ff62069

Observation ed5dc038-8008-495b-86fb-10591f58257b · outbound

This paper cites Language models are few-shot learners.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Language models are few-shot learners

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:49.811589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:49:41.879455Z digest=sha256:ffb1ef14f265ce1eb1df17043baf7c7f08286518b0ac8fdce0238f6b59972fc3

Observation 0ec95660-fa27-44ff-895f-bd061699fefa · outbound

This paper cites Genie: Generative interactive environments.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Genie: Generative interactive environments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:41.967656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:41.967656Z digest=sha256:fa7703220ee3c8f305cf3dd3c033c815e132ebfb3e257caf42b0954e791b78f5

Observation d699f9dc-f578-41f3-9d70-722e4d6eaaa6 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.051150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.051150Z digest=sha256:07f396844583c835d308822b47a8aa200fa58fe4fad2ab2118d92d73465a8a53

Observation 4419b65c-9ab0-42a4-971f-c21882a4a9c5 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffusion.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Diffusion forcing: Next-token prediction meets full-sequence diffusion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.114414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.114414Z digest=sha256:5d8f60cd63d246ae51dffc145ca67e58bf54a4ca78d5cb39a869bf589757085e

Observation 4f067ec0-c780-4392-ad75-32d7e4a8449d · outbound

This paper cites Sana-sprint: One-step diffusion with continuous-time consistency distillation.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Sana-sprint: One-step diffusion with continuous-time consistency distillation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.153720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.153720Z digest=sha256:5832444b5c5941ab97148e71ad5328b885e0382687e95614f720d59542c6789b

Observation 794491b1-9a61-44bd-8399-e6d5650962a6 · outbound

This paper cites IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.207829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.207829Z digest=sha256:2f00af38872f5c6f002acac54e387439958177e16eaeee37b92ae3f3f61f1c8f

Observation 55a430e0-325c-4524-8c13-2f8e5b2eb418 · outbound

This paper cites CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.283209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.283209Z digest=sha256:563f4975a1a620e3abf2a3338198604568bf26452fbdddd84b2750f2924d40dd

Observation 05958ecd-d84a-4b4a-8956-9ace4dfdf4f1 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Diffusion policy: Visuomotor policy learning via action diffusion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.338890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.338890Z digest=sha256:c543fd9c16da398491134cc7b7b523d00c3f5b0622eaeeb44d55b9babef9f058

Observation 614d844f-25fa-4288-b6c3-c257825fa05b · outbound

This paper cites Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.399750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.399750Z digest=sha256:82c8fb4f9d7c07207d0f73e5e71b60f0f6073fe043b7457dca4a78c7a5fa183e

Observation 5a8c4109-0bdf-4b38-b94a-da12a55010be · outbound

This paper cites Accelerated Diffusion Models via Speculative Sampling.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Accelerated Diffusion Models via Speculative Sampling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.452683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.452683Z digest=sha256:5a11857515304d9a4c498560c8d3f55f9034d61bb4f28e529532226dbc513806

Observation 98710a6b-25bd-40cd-8280-d7ecc15005d5 · outbound

This paper cites Oasis: A universe in a transformer.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Oasis: A universe in a transformer

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:49.594996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:49:42.512389Z digest=sha256:a10d1bef9f127f8b41b194d632d32005ba93955bcacfd9255bfe7e7624ab05fc

Observation 555ffc70-71d1-4176-8a2d-8315aef0adf2 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Diffusion models beat gans on image synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.552838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.552838Z digest=sha256:efb246a20882276f6ebffb9c9293ac2a8a8e9e150219ebb046a66237634d70c6

Observation 7e9145c7-2721-4514-a17c-e447792d10fb · outbound

This paper cites Learning universal policies via text-guided video generation.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Learning universal policies via text-guided video generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.628427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.628427Z digest=sha256:a6699b10a1cf568ab7aee9a6d18c973928c9ffa8fefc1237fca3c7e8615ad867

Observation 896a152c-dc74-46c7-9765-1834df733d95 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.677994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.677994Z digest=sha256:6feac29b4e8470941ed036642c8d2bb0f13ae1f373692b8a8f5093f36d8f7c95

Observation 9b5c2a66-effd-4e70-91f4-e9e926a1d3d7 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Taming transformers for high-resolution image synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.727555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.727555Z digest=sha256:ca8ac42077ba347461b326edae31823b727477f3848035a590992efe227aa1ff

Observation 25872947-c2bd-4564-9b4b-80de0aa5fc1c · outbound

This paper cites Vista: A generalizable driving world model with high fidelity and versatile controllability.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Vista: A generalizable driving world model with high fidelity and versatile controllability

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:49.346082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:49:42.792368Z digest=sha256:d6c28efcaca4112c992ba386eb7889239e05dba51870dd1adf487df66d8c96df

Observation 4ba8dbc8-c918-4eb4-b0d6-ad4fea3cc9e3 · outbound

This paper cites Long-Context Autoregressive Video Modeling with Next-Frame Prediction.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Long-Context Autoregressive Video Modeling with Next-Frame Prediction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.840907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.840907Z digest=sha256:653163592df172868f6cf90ed40f956712bcddd37bc31173469c4579ba9b457c

Observation 92d671bb-a3a7-4acd-bea1-766cfceef618 · outbound

This paper cites MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.913195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.913195Z digest=sha256:354d7354ea944252c00b2c7078682f341ed6520588805a0b7bacf99fffa94c85

Observation 84a16ecb-f50c-4525-9fb5-88aa4bc0fa82 · outbound

This paper cites World Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion World Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.962948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.962948Z digest=sha256:a99a0d73c92c13b22c9be23042a321f9ae6db28ab9fedc4c5ac377470fc9577f

Observation a041ef9d-72a0-4fe4-b8e4-6ee3aa4b97e1 · outbound

This paper cites Mastering Diverse Domains through World Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Mastering Diverse Domains through World Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.016848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.016848Z digest=sha256:a1b1d4331b2867fc963003a37d9d860f1950627b0b99d4938716f4246930be7e

Observation 27b1f4b2-18df-4db5-91f1-8708ce935c74 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Imagen Video: High Definition Video Generation with Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.063242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.063242Z digest=sha256:92e7aa688da82ac0df799aca945c605a9046efeb3fb20fcdd939f0503628acb0

Observation 9b507f31-7b1a-4daf-ada6-45498c6f5a7b · outbound

This paper cites Denoising diffusion probabilistic models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Denoising diffusion probabilistic models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:49.232940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:49:43.103930Z digest=sha256:f3422e3827c66589e167dbbf2cc5df37df5fcbed3e7173ff6786224fd22d479d

Observation 8022c61b-1d09-4c0b-abe9-1a3830976dc7 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion GAIA-1: A Generative World Model for Autonomous Driving

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.159204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.159204Z digest=sha256:c1e3f58868ff05857bdfbd9e85662de0ee482e7e673e0792b665f9b917ed96e2

Observation 7eb626a3-f55d-4c06-b5b6-288a638f1d38 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Elucidating the design space of diffusion-based generative models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.244093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.244093Z digest=sha256:a44f42144afe5a126dd3b83d96d1beb72c1fee053154d43dce3d6d1e2f292af6

Observation 9b8a44f7-e077-45ec-9dc4-9a958dc95356 · outbound

This paper cites Learning to simulate dynamic environments with gamegan.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Learning to simulate dynamic environments with gamegan

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:49.072910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:49:43.295706Z digest=sha256:254c355e2573e5da1c7746fac2d53adfdce848255bc4cbf5bb4467d41bab632d

Observation f94f38e9-fc56-4ca3-9d86-bed4c5f7feca · outbound

This paper cites Adam: A method for stochastic optimization.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Adam: A method for stochastic optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.344193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.344193Z digest=sha256:ee671005d1eaaddba8d5c4d389257e732d80192adee74b9d94b82529c8dc20f7

Observation d457942c-d70c-44ed-a610-bb856991f381 · outbound

This paper cites Videopoet: A large language model for zero-shot video generation.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Videopoet: A large language model for zero-shot video generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:48.865142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:49:43.410746Z digest=sha256:b7895288c78458a1faa7fde510f08ecc71a6a09c38740ced5d09fb0cedb3b223

Observation 0f3acd39-c673-4d9f-bd75-efd65314df84 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.465981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.465981Z digest=sha256:6fd59a25d4fea13b6830e1634821a0191947892de4561f92442c1f45dcda6212

Observation 831cc28a-985a-49d5-aeb2-ad88c629ebdd · outbound

This paper cites Fast inference from transformers via speculative decoding.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Fast inference from transformers via speculative decoding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.547324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.547324Z digest=sha256:b0bc7c4ebe8558551f470e2a9e61cdb1908a2c9d89bef5f0307c184fb026655e

Observation 9a681093-c219-414d-bba4-996fb72b1165 · outbound

This paper cites Autoregressive image generation without vector quantization.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Autoregressive image generation without vector quantization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.588354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.588354Z digest=sha256:77c0883baacbe6340f1bed87a15cf2219e4fd65b7c9ad69bdc560e7247ed8b49

Observation 5a430c13-2701-412e-906a-22546058089c · outbound

This paper cites Flow matching for generative modeling.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Flow matching for generative modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.645431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.645431Z digest=sha256:f6cec32520a4b51ad611e96e69824319f93f239841ac10b30a45cb4e1b5682aa

Observation cb3697ee-4b98-462b-a585-1e7d43bee3cc · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.695257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.695257Z digest=sha256:69461ffe79c866d741c9b38b561cad6455a6647279ef74049a65bf4ba4e0393f

Observation a127a3b0-bf61-42ad-bc05-c61bcf86d895 · outbound

This paper cites Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.757089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.757089Z digest=sha256:c718573990eabcf0591138bc4b30845af711b1e3480f8814f4d76e80ff2c5b2d

Observation 064070d2-898e-499e-9c37-14aa8aba7101 · outbound

This paper cites DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.826343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.826343Z digest=sha256:2447e29a0db673456b5a55a03519a79555be4f9087a6c84c2ab2aafd465c0ef9

Observation 6c9bea53-8c1e-4823-9143-fe3bebe56106 · outbound

This paper cites an unresolved cited work.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.894267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.894267Z digest=sha256:5ec87391a364cfd76dfe56f5f52ac76e4ed827ca0c9247036b0bddcdb4d49d49

Observation 50e4430f-ac5a-49a2-a0ec-fcb87473229b · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Pytorch: An imperative style, high-performance deep learning library

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.972665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.972665Z digest=sha256:99677144307a472b88f59aaddd0b2fb794f37532fcc543f9497fd195469eb94f

Observation 997a0c0e-7b37-40d6-a318-66fe5816d0a3 · outbound

This paper cites Scalable diffusion models with transformers.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.051513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.051513Z digest=sha256:ceea3d3a6caf55f495aff38e951cd9348d1e43f1513c1f4c58bd0791e4d320a7

Observation 36275562-bb18-48f9-bcf3-fcbefeaf1c6f · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Movie Gen: A Cast of Media Foundation Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.157194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.157194Z digest=sha256:9904c3c65833c79a886f083cc29029208456f9298bd1808a260c393fef5e2aec

Observation d3a2258f-eda7-4680-9303-375ca7af16d0 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion U-net: Convolutional networks for biomedical image segmentation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.282569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.282569Z digest=sha256:752fb2352a3e6938403eee633dc9f5de15c0629a9faee2372916192174e29694

Observation 758a92ba-c5cf-4ec5-bd22-2dad472cd609 · outbound

This paper cites GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.373058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.373058Z digest=sha256:1f07b72e7b42fa551f8b12e3bcef10c322e2cbf02278cbbd933dbd0be49451d0

Observation 7ea15702-e171-45c5-b4bd-5e1ffcb05278 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Progressive Distillation for Fast Sampling of Diffusion Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.482801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.482801Z digest=sha256:ef33b2c313df86faa4c802bb295be934f051dc19c22def0551ebd9fdd9ceb01c

Observation 3afca255-798f-4e8b-93f6-76d4ff56c7fd · outbound

This paper cites an unresolved cited work.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:49:48.625980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:49:44.610609Z digest=sha256:834ddc73aaacf08e30befc49be72f80046c160663da22ad057aea2b4939cbd9f

Observation 9014dacf-a4df-4dd1-abe8-536867e34a55 · outbound

This paper cites Mastering atari, go, chess and shogi by planning with a learned model.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Mastering atari, go, chess and shogi by planning with a learned model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.715413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.715413Z digest=sha256:c1490f531e110b4f3152340e44863013f377fe588cf0e52b48ae35d75489d658

Observation 4139b3eb-4354-4a14-897b-c21ada20b60c · outbound

This paper cites Deep unsuper- vised learning using nonequilibrium thermodynamics.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Deep unsuper- vised learning using nonequilibrium thermodynamics

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.809000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.809000Z digest=sha256:93ea445810e1bc6891b5cd32b3272d0c45b0869826eabb96f87e1e47a5da8cb6

Observation 669706dd-e7f9-4c9b-a152-c74416627bd0 · outbound

This paper cites Denoising Diffusion Implicit Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Denoising Diffusion Implicit Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.911941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.911941Z digest=sha256:acbca24130ca7cf6fbf1a9a28a1562885378f0fd258a177b8463b1010047098b

Observation 4b69d8d2-a957-404b-bce3-5b9631ea15e5 · outbound

This paper cites Consistency models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Consistency models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.039508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.039508Z digest=sha256:e1739f18eeaedc38b73655da7d5fe53db2f1589adf75a1cf18c5afd32e189f2f

Observation ef69d201-f105-4960-bf9c-554ac4c7e766 · outbound

This paper cites VidTok: A Versatile and Open-Source Video Tokenizer.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion VidTok: A Versatile and Open-Source Video Tokenizer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.159313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.159313Z digest=sha256:7b6aecfab896ce11c8bb3490596e29d96f0fc588d2525ffb0a992c5c96337147

Observation 82f4af78-8928-40ad-8d5b-fe5b1a6d6a93 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion LLaMA: Open and Efficient Foundation Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.258191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.258191Z digest=sha256:5b6f5960d0b584e8b6f06c3616163c07c541f10a8b11f67be5dd9ff565c11a4c

Observation 8df5ef5d-437e-44bd-80e9-3c2672772843 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.338761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.338761Z digest=sha256:0b9a80a8b16053c09d41485a578bc98ee6b5d94ba1f6c772e219ebcc572f3bbd

Observation a15bba91-5386-4aaa-bd82-07dc7b2dfcff · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Diffusion Models Are Real-Time Game Engines

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.459923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.459923Z digest=sha256:59431969f039c757ed589b8424684abfa428bc3435b681e6e7bf8dbd53229f00

Observation a71b5ed8-f426-4118-bc98-ee61850d411f · outbound

This paper cites SparseDM: Toward Sparse Efficient Diffusion Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion SparseDM: Toward Sparse Efficient Diffusion Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.587972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.587972Z digest=sha256:61bb8914cec9ea7e12cf9c7b722cb1c5438fb62c9210b76f853de8be11f7a728

Observation de718956-e0a8-4e10-ac15-e6ff202e9df5 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Image quality assessment: from error visibility to structural similarity

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.688193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.688193Z digest=sha256:ba645a6ab2fc9d290d79f3ade1a800550d71393debb28938051e136374e5648e

Observation 944b68fc-c7cf-41f4-a0b2-7e15d97f381f · outbound

This paper cites ivideogpt: Interactive videogpts are scalable world models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion ivideogpt: Interactive videogpts are scalable world models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.793677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.793677Z digest=sha256:82129a4c2b686e76dda8a2d5c217851a81467dad925a8dcb6571cb7fdbe7289a

Observation 5867bbd9-0c42-4daa-9eea-7c3cdbca4bb9 · outbound

This paper cites Structured 3d latents for scalable and versatile 3d generation.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Structured 3d latents for scalable and versatile 3d generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:48.368033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:49:45.891921Z digest=sha256:ec4db180fea08d8560b3b4860448e345f764653f86e89cf083e115bcfc03c671

Observation 006165f0-d22d-45b4-9a01-8dac9f634583 · outbound

This paper cites Pandora: Towards General World Model with Natural Language Actions and Video States.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Pandora: Towards General World Model with Natural Language Actions and Video States

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.038383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.038383Z digest=sha256:92becb9fbe109ee1ecf069ac7031981a0a37fac8e4f4d5362100119983e6c1d7

Observation 0e869d7d-1207-4229-9ba0-c76a48058a43 · outbound

This paper cites Learning interactive real-world simulators.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Learning interactive real-world simulators

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.094491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.094491Z digest=sha256:38cd82eff14f2c09c00ef2ccf69342105fb78af04e88e1e7faff040c929e441f

Observation 76a209fa-6616-4f33-a3b4-a610861537af · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.205356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.205356Z digest=sha256:d308868247cc309528ce4f88032a0413a273178588b8b81d446bf3a8521e580b

Observation 3679763c-168a-4206-98dc-f77eea973d1c · outbound

This paper cites One-step diffusion with distribution matching distillation.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion One-step diffusion with distribution matching distillation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.325346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.325346Z digest=sha256:c3ed8b1df9ea95977b72951dc31e3abca87eef22219e93e0348aa6cd77b3259e

Observation 595a32b9-49df-4289-9282-8ed39999edc2 · outbound

This paper cites From slow bidirectional to fast autoregressive video diffusion models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion From slow bidirectional to fast autoregressive video diffusion models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.430341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.430341Z digest=sha256:76970175f0b547dfff36d2eb2f36405bd2e59c56e5e6de5960932b655f7fbb81

Observation fdf71b7a-fb16-4023-b770-e1a77b564d81 · outbound

This paper cites Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.509630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.509630Z digest=sha256:ad2e20c0f6bfd051aaac28478b81a0184e0e2a7a2105c177d195835187c614f5

Observation 3cf5599f-d2c9-4bae-95ac-a89b1cab702f · outbound

This paper cites The unrea- sonable effectiveness of deep features as a perceptual metric.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion The unrea- sonable effectiveness of deep features as a perceptual metric

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:48.182571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:49:46.632889Z digest=sha256:29a02f38295ac4b96b7f0c6e065080101dcf8e28a83a57838de3823a9ed0dc4e

Observation 9653b660-e1e2-4245-9f08-35784f3d887a · outbound

This paper cites Drivedreamer-2: Llm-enhanced world models for diverse driving video gen- eration.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Drivedreamer-2: Llm-enhanced world models for diverse driving video gen- eration

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:48.009638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:49:46.727368Z digest=sha256:c896bb3e5abc577c617e3a07f2ed8cd3544bba89b791bb1c4770c9c16e72238e

Observation d75d06a7-22a2-4eef-add7-1fb68169bf7b · outbound

This paper cites Genad: Gen- erative end-to-end autonomous driving.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Genad: Gen- erative end-to-end autonomous driving

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.822591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.822591Z digest=sha256:68e6cf4c3f54c861d43c7c0c7103b6aa24e7a236d3500e33afc0f508abbde027

Observation 2bba489d-707c-40e4-b46a-73d633d75228 · outbound

This paper cites "" model : Distilled NFD + model vid : Input video tensor act : Action sequence tensor.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion "" model : Distilled NFD + model vid : Input video tensor act : Action sequence tensor

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:47.872390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:49:46.876113Z digest=sha256:f3d4f5442fd8aed41171654ac3de225b7a1eb88b2afc23153b99f03b320f473d

Pith citing papers

Observation 85ed9549-5a85-4e0b-9220-f9468c6cb2a5 · inbound

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling cites this paper.

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:06.579621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-19T05:13:28.767788Z digest=sha256:a1dd1733ab18ff7a07a8560dd0752bf88e9a11b2db622aadce200393304edded

Observation 6fe79ce5-4e09-44a6-9f1a-f37ed8288403 · inbound

Matrix-game 2.0: An open-source real-time and streaming interactive world model cites this paper.

Matrix-game 2.0: An open-source real-time and streaming interactive world model Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:36:53.261124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T22:36:31.044743Z digest=sha256:9ac42b344f7bd83151a8c363f0ce4c1379b1b0ff6c2116ff1d516ff5e341d2f6

Observation 2b2e644a-d463-4e90-9339-514e08ab82df · inbound

RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation cites this paper.

RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:43:37.510074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:43:37.510074Z digest=sha256:c3bf74a30d40e25fbe49620aae5e33f672e780ae23152dacf07453eb08daee25

Observation 73a3b47d-225f-4037-bc31-628ccd0cdf46 · inbound

Envisioning the Future, One Step at a Time cites this paper.

Envisioning the Future, One Step at a Time Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:05.951388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:31:41.904925Z digest=sha256:d92b9d4da18ab6c1ef538c27b622b58a553d6b6f3fda1dc79f081e6027ea29c7

Observation 63b7fccf-b51d-41fe-a17e-679b03228f9c · inbound

Efficient Video Diffusion Models: Advancements and Challenges cites this paper.

Efficient Video Diffusion Models: Advancements and Challenges Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 245

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:26.144191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T08:28:29.706249Z digest=sha256:c0940937543e9c4e19b6e0cd67eecde5e592a5d6e72798692d0642485707fb95

Observation 7f4661b4-163d-45c9-a1d4-182b8e7013ae · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:29:39.422283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:29:19.111071Z digest=sha256:f2e6f79d8c2a2715ffd55f5e916b99263460d9e434f58791b394e6ade3710d1c

Observation 304a5c1b-a27a-4192-b7c2-c58efde9de4b · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:54:57.963353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T17:50:22.258541Z digest=sha256:37df71394d47c01b328f5a8118058b158ab02b604679b0b5b1d407cef82dcf88