Pith. sign in

Paper Citation Record · LEDGER

Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2501.08453.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08453 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:26:46.659549Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:19:51.141627Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d1dcf5ab-059c-4a3a-ad8c-52817e93af40 · inbound

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT cites this paper.

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T14:26:46.659549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:26:46.659549Z digest=sha256:e665e07ad8c6c0d2dae968459c502dcfd6f3c790a734eb24fff1e3d28e12202e

Observation 728d888d-49af-4855-a942-ed157d68db7c · inbound

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control cites this paper.

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T19:40:21.185472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:40:21.185472Z digest=sha256:a842148ee0532678ee12febc710dc454d5a75d86ead43d899b2a391af6092ef2

Observation bc835403-83a4-48e5-873d-630c80b32ee4 · inbound

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness cites this paper.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.081613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:e262ed84a29c05b87852cbbd1d792aef81832746f679f38a1554aaa3de69241c

Observation 78244fcb-96a4-4777-9e21-b732843a7d56 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.452382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.452382Z digest=sha256:f40783a68186786278ca6832c1b864a7dd6520ed96d632d5b0052e535aee5c01

Observation 52f37071-c193-4085-8050-6f41ff3fbf83 · inbound

M4V: Multimodal Mamba for Efficient Text-to-Video Generation cites this paper.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.946935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.946935Z digest=sha256:c8abc2ea09bd9efb3a239e75d145dda9cea1796b826d7fab904f7452cb51ede7

Observation dc56006d-c6c1-4274-ab4e-0f9c6c3b542f · inbound

MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos cites this paper.

MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:08.507165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:25:08.507165Z digest=sha256:8ee9f14132b4cdb9ac7312a8963925ade6dd5b18f8234976f9b693801013ea8a

Observation 5741b480-aabe-4f75-800f-eb9fc30f5b5d · inbound

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis cites this paper.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.099250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.099250Z digest=sha256:690526abb3800eb8d1cd8f6625a3fadf8208c574369ab94e6bc70c379896036f

Observation 5c9d2526-c3c7-48bc-a539-8bfdd7b56b8a · inbound

CineScale: Free Lunch in High-Resolution Cinematic Visual Generation cites this paper.

CineScale: Free Lunch in High-Resolution Cinematic Visual Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T17:46:33.545910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:46:33.545910Z digest=sha256:9b685722dcaafe6ec67e61934422fb96a8ef2869ff6fcc51e289d235a80feff6

Observation d45b417b-8b27-49c3-b08e-f02ff5eed539 · inbound

Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility cites this paper.

Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:56:24.338083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T12:55:42.679016Z digest=sha256:51671b99a051158164f16e55b6a71121252d8cb412920c44149f55db0416db25

Observation 00abd484-ba78-44e3-a6ec-bf4a1dc54be8 · inbound

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World cites this paper.

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T17:02:33.359198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:02:33.359198Z digest=sha256:51b3e17b7bde57f1166a3a7cdecdf2c6241b017420e0d9e7f0f83bb6e79d8f0b

Observation 3abc536f-de88-4818-9cb7-a999c0c5bcce · inbound

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation cites this paper.

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T13:24:44.479053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:24:44.479053Z digest=sha256:a5a0aed714540a06f1610a5d2692ae3fb81f6971698e53f0fd8d9041bfc95bd9

Observation 1d7c3677-d75b-428b-9d90-c736dbf65566 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:1d552bb7316d6534b7dbe8aacc59a994b1effe40afffa76b4852ea17024cc11f

Observation 11b88191-a3d4-4dfa-985f-636b32b79294 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.809611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.809611Z digest=sha256:da892e6d9d8e1181df394d9704606001a0457a6a4443d29cb8e94258d25f07ef

Observation 1f880598-6788-4466-8203-a223b55a83ae · inbound

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation cites this paper.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.919945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:d6ad45e9748fead6c28e68289f6aa0f53e3c889e19004ed1e2d3f9ee1e80d88f

Observation f59cc563-36ee-4d44-94d6-3ba79d4bfe22 · inbound

Spectral Progressive Diffusion for Efficient Image and Video Generation cites this paper.

Spectral Progressive Diffusion for Efficient Image and Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:08:13.392425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T11:07:47.135336Z digest=sha256:11da0e0cbae74a10cb390650b1e218c9e2cb7c9f461fa48b4bb060162161b301

Observation 5f7ce34f-94eb-4a56-a377-d24680c88e7b · inbound

Spectral Progressive Diffusion for Efficient Image and Video Generation cites this paper.

Spectral Progressive Diffusion for Efficient Image and Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:54:02.744052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T07:52:07.409740Z digest=sha256:586a21388590e7f1efb2fee87130c9bb888fcb93fe9330085fb89cd01a694889

Observation 7aef4403-ab8c-4ddd-9888-2ae24b566308 · inbound

PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution cites this paper.

PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.355127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T22:24:50.679950Z digest=sha256:f034410d2f68467bb1eede29d1586272cc070dd638c838c4ef1eda86642582d1

Observation 39441c36-f316-4f64-a979-41fe82e66c8a · inbound

Latent Spatial Memory for Video World Models cites this paper.

Latent Spatial Memory for Video World Models Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.541086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:84cef5580e1f3cfc1b43b592590a3abda1350bccc6f791b79c42aa31d8c3ca82

Observation f970471a-9016-4fdf-8dcd-fb947d785f9c · inbound

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation cites this paper.

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:19:51.143006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T05:16:53.011837Z digest=sha256:bffb04ad35e8888d4383be11d9853c405bbfb122ef2da18f945a43549020ee3a

Observation f308f43f-bb79-4e98-b540-aebe65808fea · inbound

Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training cites this paper.

Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:37:42.231981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-03T06:27:48.935774Z digest=sha256:213c9455dd597a9d3f24d6000c479e552bce011ddb7bd3ba01c634a890cc8f67

Observation 14051d7e-fbf4-4505-b30e-bc178db40a6a · inbound

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence cites this paper.

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 114

Resolution
unresolved
no resolver link, observed 2026-07-14T03:51:24.547781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:51:24.547781Z digest=sha256:b65b5583499e84aee616a1dcd40075cc8c04ebc2738e7e03b310a3190bb506e9

Observation d3ab19fa-6d9f-44e2-b9bf-02e0a33439d9 · inbound

CODA: Algorithm-Hardware Co-design for Edge Video Diffusion via NMP-Enabled Compute-Cache Operator Disaggregation cites this paper.

CODA: Algorithm-Hardware Co-design for Edge Video Diffusion via NMP-Enabled Compute-Cache Operator Disaggregation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T00:49:33.849062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:49:33.849062Z digest=sha256:41570e61e778664229871a70763d371e36f3cc0d4bedb2a69b68b1a0d99d78ba

Observation 334d4a70-d317-40c3-8ce2-d248925386f7 · inbound

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System cites this paper.

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T08:28:07.621970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:28:07.621970Z digest=sha256:30faeedcafe0c53b215706f6df862c8371df997af6aca80d57c734e71d493415

Observation 628f01bd-f16c-44d5-af5f-911c8601e6c9 · inbound

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation cites this paper.

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T04:58:52.713725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:58:52.713725Z digest=sha256:80d03c14c91cc62195edda9a50686136d7cbde8fe1ac580fc62ccc07ca9e737a