Pith. sign in

Paper Citation Record · LEDGER

Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2501.08453.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08453 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:29.452382Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:19:51.141627Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bc835403-83a4-48e5-873d-630c80b32ee4 · inbound

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness cites this paper.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.081613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:618da2cc99b457522bef9a0c2ba790a6053852e19b6fc5a4a1a87dd2162bcb48

Observation 78244fcb-96a4-4777-9e21-b732843a7d56 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.452382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.452382Z digest=sha256:4543646199606f7393e66928c97cc964dc123946aebcee2ba50f1363735e6253

Observation 52f37071-c193-4085-8050-6f41ff3fbf83 · inbound

M4V: Multimodal Mamba for Efficient Text-to-Video Generation cites this paper.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.946935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.946935Z digest=sha256:2c33ea50503586a563d46dd45f935dec9bf91029a264d22ea6af751900e2bb3d

Observation dc56006d-c6c1-4274-ab4e-0f9c6c3b542f · inbound

MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos cites this paper.

MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:08.507165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:25:08.507165Z digest=sha256:d114e2753933ca382febe9847b1edfaf5add6db0c94dcdc974b30dfb660cafed

Observation 5741b480-aabe-4f75-800f-eb9fc30f5b5d · inbound

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis cites this paper.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.099250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.099250Z digest=sha256:3e00a6e5993541a082bf9f9fe6f5dc32b8c7866ffb47361aac37dc1dad6a1ed0

Observation 5c9d2526-c3c7-48bc-a539-8bfdd7b56b8a · inbound

CineScale: Free Lunch in High-Resolution Cinematic Visual Generation cites this paper.

CineScale: Free Lunch in High-Resolution Cinematic Visual Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T17:46:33.545910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:46:33.545910Z digest=sha256:6035f753d0ee3b6c9bd91fe7df4e7a027baa7074e968b1e218a8be3749bd6d78

Observation d45b417b-8b27-49c3-b08e-f02ff5eed539 · inbound

Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility cites this paper.

Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:56:24.338083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T12:55:42.679016Z digest=sha256:35fd6328cf717bb954097a2ef827a854e83f8f24b189eff4c90c8cb8dd4884e8

Observation 00abd484-ba78-44e3-a6ec-bf4a1dc54be8 · inbound

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World cites this paper.

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T17:02:33.359198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:02:33.359198Z digest=sha256:b19ef22189125a4fd20aac919b04808e39e7e329c51db16ba9e8742cb1215ba4

Observation 3abc536f-de88-4818-9cb7-a999c0c5bcce · inbound

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation cites this paper.

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T13:24:44.479053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:24:44.479053Z digest=sha256:8a00f82f8d18a0b432b4d35340af93d3c38ea5f91819e67e4ca99b5a1bf52041

Observation 1d7c3677-d75b-428b-9d90-c736dbf65566 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:392784e6447735e50e6d15047fadcf60ddde823cd840151fc51194f2c1ce7025

Observation 11b88191-a3d4-4dfa-985f-636b32b79294 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.809611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.809611Z digest=sha256:709622998ae50c3ff5ed7cea934f44d1f067c209f649f46828ad3f64d7932ee0

Observation 1f880598-6788-4466-8203-a223b55a83ae · inbound

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation cites this paper.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.919945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:05ba4b75e184d31f9950c01f0608c8fc61b5b8432bf650411d9ce1a46b2fd2c3

Observation f59cc563-36ee-4d44-94d6-3ba79d4bfe22 · inbound

Spectral Progressive Diffusion for Efficient Image and Video Generation cites this paper.

Spectral Progressive Diffusion for Efficient Image and Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:08:13.392425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T11:07:47.135336Z digest=sha256:f414dae5293fda1175aaf8132cb0888166b7d39670b1e1caa0d4ebc034311604

Observation 5f7ce34f-94eb-4a56-a377-d24680c88e7b · inbound

Spectral Progressive Diffusion for Efficient Image and Video Generation cites this paper.

Spectral Progressive Diffusion for Efficient Image and Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:54:02.744052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T07:52:07.409740Z digest=sha256:212e597f96a6b921a759697abf66437cff709f002e307023917d9b04c19a7fa6

Observation 7aef4403-ab8c-4ddd-9888-2ae24b566308 · inbound

PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution cites this paper.

PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.355127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:24:50.679950Z digest=sha256:69be630cd116e29b1e7929824271677ae438120d480d0308bdcdf57e41133133

Observation 39441c36-f316-4f64-a979-41fe82e66c8a · inbound

Latent Spatial Memory for Video World Models cites this paper.

Latent Spatial Memory for Video World Models Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.541086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:8e45f9b451132b44c6421a39cf373fe71c9d2281820b8fb0e959f9949b53937f

Observation f970471a-9016-4fdf-8dcd-fb947d785f9c · inbound

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation cites this paper.

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:19:51.143006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T05:16:53.011837Z digest=sha256:0928328ef6a933bcd99832727dd0a711d717c9d912280c09e470d4a305c53cf0

Observation f308f43f-bb79-4e98-b540-aebe65808fea · inbound

Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training cites this paper.

Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:37:42.231981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T06:27:48.935774Z digest=sha256:34dad299aa9a15ce3889bd9ddae68c2dc651bea9c43584e7060a1f5a05f60f82

Observation 14051d7e-fbf4-4505-b30e-bc178db40a6a · inbound

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence cites this paper.

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 114

Resolution
unresolved
no resolver link, observed 2026-07-14T03:51:24.547781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:51:24.547781Z digest=sha256:a0b335699546184754def5e3b18ca136b7729efd0ed9268029ef8dc3a490a901

Observation d3ab19fa-6d9f-44e2-b9bf-02e0a33439d9 · inbound

CODA: Algorithm-Hardware Co-design for Edge Video Diffusion via NMP-Enabled Compute-Cache Operator Disaggregation cites this paper.

CODA: Algorithm-Hardware Co-design for Edge Video Diffusion via NMP-Enabled Compute-Cache Operator Disaggregation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T00:49:33.849062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:49:33.849062Z digest=sha256:7893f29faae1e69b04563a44dcf9d896cffb12538f27c12471fea29e4e2e8e17

Observation 334d4a70-d317-40c3-8ce2-d248925386f7 · inbound

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System cites this paper.

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T08:28:07.621970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:28:07.621970Z digest=sha256:5894c98e6d9f4b4c3805abb9e0a485e46273a75cb4f989442683ec8872173108

Observation 628f01bd-f16c-44d5-af5f-911c8601e6c9 · inbound

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation cites this paper.

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T04:58:52.713725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:58:52.713725Z digest=sha256:4d33b8c22549dee98333859cc3f508f427881accbffaf8c86d0e685f2bc2ebb7