Pith. sign in

Paper Citation Record · LEDGER

VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2407.12781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.12781 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:08.274070Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d784e247-f4dd-4b5a-b754-53cfb65352d6 · inbound

CameraCtrl: Enabling Camera Control for Text-to-Video Generation cites this paper.

CameraCtrl: Enabling Camera Control for Text-to-Video Generation VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:06:23.563223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T02:06:23.410241Z digest=sha256:81173c677502ca48d58014cfc351bc7fbce7f47c9c0ff525cb3754682ca4b8d8

Observation fe1b001a-3e8a-44ae-8e2c-66d23bd7fe86 · inbound

ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis cites this paper.

ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:59:03.794074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T22:59:03.642189Z digest=sha256:afc7745c87d7741a329e6d2e444b39b24644fdff9d820ac6ad05a1422f9a17fb

Observation 4bcb8fdb-8347-448a-a1ba-c8aecbb4dc44 · inbound

Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry cites this paper.

Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:08.274070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:08.274070Z digest=sha256:0f19789962f168d1ce246a696767332ce9afb70f4073fd0d74e2925829f2735b

Observation ab5176a5-be07-4a85-a416-cb99f9d3e077 · inbound

LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion cites this paper.

LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:10.495325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:10.495325Z digest=sha256:2b295e96f64a2d620a866b77026efe5a71b91ed33025911bcd2376a77604328b

Observation 387c8d9c-638f-4cac-94ab-1ef16fc03646 · inbound

Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion cites this paper.

Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T13:15:10.485306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:15:10.485306Z digest=sha256:fd2e046ab6e756e013532d166b28144fab809890cface69133d01df621f18ea4

Observation 2831a9cf-f7e3-447a-aaaf-af1504a6e8cf · inbound

OmniNWM: Omniscient Driving Navigation World Models cites this paper.

OmniNWM: Omniscient Driving Navigation World Models VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T08:57:05.028270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:57:05.028270Z digest=sha256:5f82c7363750c90a1a924eac63962e52808232b4fa998db08007f1b0415a726f

Observation 0d49eacd-0a98-44e6-b278-b848fcf9b09a · inbound

PostCam: Camera-Controllable Novel-View Video Generation with Query-Shared Cross-Attention cites this paper.

PostCam: Camera-Controllable Novel-View Video Generation with Query-Shared Cross-Attention VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T21:01:06.386578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:01:06.386578Z digest=sha256:d00d6e2f282ed127944269757e3410d2751df7177397499d4c5b65f8967f8d1a

Observation fe6a78ec-284d-4bfb-ab4c-ea7e31c0f2dd · inbound

InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem cites this paper.

InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T18:26:31.548588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:26:31.548588Z digest=sha256:2ddc9022966271421594a644857dcae4a73bd1745d74f24780836326b2008cc1

Observation 310cad1e-00f5-44b4-a44d-82423753cecd · inbound

SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-body Manipulation cites this paper.

SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-body Manipulation VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T05:25:02.032281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:25:02.032281Z digest=sha256:8c04895531dace51a0176448b9373a095f380967a0c74a2a65247945d9a63387

Observation 5b3c1544-2bf3-4d15-b1bf-f8100b949e6c · inbound

SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation cites this paper.

SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T12:26:54.361488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:26:54.361488Z digest=sha256:9846f920cb5e0e3066729ab21c494b074498c7af41b005cdaefcb617d50c9ffb

Observation c2137e25-af87-470c-8e47-d211465274b8 · inbound

INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling cites this paper.

INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:25:53.614176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:08:56.588282Z digest=sha256:7dbf49111cdc34bca69aac4b750c7909c9880fd0b02d74d6c2acedabaad2e9f7

Observation d3f3272e-ee30-40e7-82be-c6d3d4077b47 · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:18.982509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T06:03:06.920592Z digest=sha256:0c9b0e7b56f48fcf4ba93f8fe03974d2e6d55b33596b3d0a697b1eb23b41006b

Observation 5cabc11e-313b-4eec-8baf-3bcc50ca55cd · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:26:29.968851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:52:58.232585Z digest=sha256:5eb47d69c62b3e9c248ceccebacb225b2434b898707f127c971a04b7067f951d

Observation c083f495-6ba4-47b5-a76f-7fc1a7120b30 · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T16:51:14.262185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-05T16:47:32.853010Z digest=sha256:c6d7f3dc267ab5de12085453e70937b53e845d5932836f1a6bed4a1da045fc50

Observation a39d6f41-a31e-4a99-b34a-ee7c045becc8 · inbound

$h$-control: Training-Free Camera Control via Block-Conditional Gibbs Refinement cites this paper.

$h$-control: Training-Free Camera Control via Block-Conditional Gibbs Refinement VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:12:28.138002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:09:32.103652Z digest=sha256:44489947686de9ed7270e53b744122af47473f2a28021971b7c9ca37ce0fa66b

Observation cc3a6092-85a6-472e-a531-18bb03b60294 · inbound

$h$-control: Training-Free Camera Control via Block-Conditional Gibbs Refinement cites this paper.

$h$-control: Training-Free Camera Control via Block-Conditional Gibbs Refinement VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:59:11.892862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T22:55:00.590833Z digest=sha256:d4b37603ec3e5509571e0a0baa73567b213fba1babb0c98870c2c0e3b4f8f03f

Observation 984f7f0c-2b65-4816-b3b3-d6192d0762d7 · inbound

MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics cites this paper.

MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:47:21.481126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:42:31.717412Z digest=sha256:70dd32b1f7fe924e28f6b176039f38e16f8a8647f811030e82a2d795ca842741

Observation df99b25d-d054-47d7-8b70-1a43f8a7af6e · inbound

MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics cites this paper.

MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:58:03.151805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:56:38.133405Z digest=sha256:90a28dc736285ab75ae8fb652fe332fc0a244309860129365ed01c99d8c71b5c

Observation 2d2f7064-89f8-4175-829b-ae11f16c872b · inbound

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow cites this paper.

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 122

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:09:23.661409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T19:07:53.671769Z digest=sha256:c7a9a6fd194af40099884a0569c74558515e782e84ad9df781cc1c89bae57dd7

Observation c7a1dbb5-e0be-4d11-b120-c1d97778221f · inbound

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow cites this paper.

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 122

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:09:49.443434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:08:52.321351Z digest=sha256:a802aaa9c25d4d6f4c8f7bc46a9f72e910a93f2971392d03b51280ed22fb93b4

Observation c57bb19d-31ec-4429-a95f-473813a2d17e · inbound

Probing into Camera Control of Video Models cites this paper.

Probing into Camera Control of Video Models VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:15:04.343521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:11:19.441408Z digest=sha256:3e71540b2520153a7924b9459e8307af313602e9fbf5b2bcf14958df702f51f3

Observation aa4579c2-f921-4279-bcaa-691cb4a14483 · inbound

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation cites this paper.

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:08:13.672536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T11:06:09.367559Z digest=sha256:6a0b286532faf1fb3ee4cc0aef3517b72a0fcdd74178c7f90cc0d9be66d7bc72

Observation 8fb0288b-5c09-4199-b4f0-bc5b5afae3da · inbound

Streaming Video Generation with Streaming Force Control cites this paper.

Streaming Video Generation with Streaming Force Control VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.412595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:14:32.465663Z digest=sha256:f5590d9951046b0f04b998bb482191f122dcac89e814a804fda6bbf479a85812

Observation 1961b9c6-0fa8-4278-b54c-a8814acce7c3 · inbound

TriMotion: Modality-Agnostic Camera Control for Video Generation cites this paper.

TriMotion: Modality-Agnostic Camera Control for Video Generation VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:19:31.326582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T18:11:49.696524Z digest=sha256:44eb623a000755c93344303d838eeff4caef934a8d2faa2f44d8fed421d47ade

Observation b5ea3002-512c-4046-a7f6-675b1a4d4d6c · inbound

Lighting-Consistent Object Transfer Across Radiance Fields cites this paper.

Lighting-Consistent Object Transfer Across Radiance Fields VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 225

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T09:49:18.033981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T09:46:46.216810Z digest=sha256:b7aa3d2ed8a5e56e330b2be547a68fc6292789d6847f7195a2bec76a40329762

Observation 039ac74f-04ec-4133-a7e2-eb822b9aa593 · inbound

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds cites this paper.

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:52.330305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T04:51:38.355857Z digest=sha256:0c33511695a4315455e9fd7c0cd647ac133dec929b82040a491e81dcc337b0d7

Observation 2714b02e-9a3d-41c6-857c-c9db5566a43c · inbound

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds cites this paper.

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:13:52.824178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T04:55:32.982338Z digest=sha256:1770fd22510b52c84ac347cdb67ef9505eba2481ffe00c78b9680a70bdec962f

Observation 0f1bc6d9-2c78-40be-9861-05ad3631740c · inbound

Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Video Synthesis cites this paper.

Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Video Synthesis VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:24:32.492574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T09:19:50.748415Z digest=sha256:164e6fda7bc8af75ac839f8eee1668c21b15d8c5609f4ef1cc32ce6d0a1f9024

Observation 2c83d53e-46c3-4833-ba33-59cc63473350 · inbound

MoWorld: A Flash World Model cites this paper.

MoWorld: A Flash World Model VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:14:53.268098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-08T13:10:31.940263Z digest=sha256:e0e241b353507db1a240c830478933a9ebd7807473b34e04302192c2e3c9e93c

Observation 0488fad1-ff9b-4690-883c-dbf0cd69ffd3 · inbound

MoWorld: A Flash World Model cites this paper.

MoWorld: A Flash World Model VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T04:32:04.136455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:32:04.136455Z digest=sha256:740547a83703bf1f6bdbddaf1e009d29f495b7c1baca7d76f6b541f354723bc0

Observation 328798d7-c366-47cc-a66e-fa6b50903b11 · inbound

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence cites this paper.

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T03:51:24.547781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:51:24.547781Z digest=sha256:e24a4327d75ada78c8bcd49feddb4169c6e56b8c991dde7163dab8894fff0b2c