Pith. sign in

Paper Citation Record · LEDGER

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

As of 16 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 4 inbound Pith citation observations for arXiv:2412.03520.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03520 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:21:55.284941Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T06:04:25.857689Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T02:59:26.816825Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1261bc12-a048-4ad5-872f-b5c867e132bc · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.007365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.007365Z digest=sha256:1ced2be075e60663702a8cedc08146a1744c2f2cb1200b56e583cf49934a37ff

Observation 98067dba-b1f7-463c-8922-ce9af8599900 · outbound

This paper cites Video generation models as world simulators.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Video generation models as world simulators

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.015259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.015259Z digest=sha256:0b78c60b36819d70403104e823ae27037925b3ca316eae8d9506ed66ff153287

Observation 7402b4cf-592c-4764-a99a-55ace1f146cb · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention nuscenes: A multi- modal dataset for autonomous driving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.022805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.022805Z digest=sha256:acc2967f706c581cf782957e862acc97699ef7123f188293731f4d1a34fad2f4

Observation c493a015-0b0b-429e-869c-8c1eb6f50845 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.029391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.029391Z digest=sha256:1658f31b1519ade00049b9a7c2403a75a50165f6c7d92a9852b7dda268dd09a3

Observation 6d9d3a85-22cd-4eab-ae5d-e96b73321327 · outbound

This paper cites VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.035366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.035366Z digest=sha256:22953733247d21197d2d7a1fc381a253a61950b981bbae0261eea632eda10325

Observation 7183b1cd-2c72-49ef-84b9-96139a831d55 · outbound

This paper cites Persformer: 3d lane detection via perspective transformer and the openlane benchmark.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Persformer: 3d lane detection via perspective transformer and the openlane benchmark

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.188214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.041882Z digest=sha256:29552c9c7f7fefa7be944a4bf8831ef3eb1e4963b584fffcde5d90e5f637f3aa

Observation 286acbde-e74d-46a9-afbb-aa259be4d544 · outbound

This paper cites MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.047416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.047416Z digest=sha256:238fc8550f7ffd2a9f746572b118ac2a779f89e8162ac4014cc0aa198dc572ad

Observation 99d89f3e-ca6c-4972-aee5-821e6c7f0f08 · outbound

This paper cites Magicdrive: Street view generation with diverse 3d geometry control.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Magicdrive: Street view generation with diverse 3d geometry control

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.168071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.053138Z digest=sha256:80e11ac943830bcaeea17a96420fe822a02c8b7fb7e5ad4e7050d370a6aad5a5

Observation d43296b9-aa0b-4b07-8a8b-7d63700ee705 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Denoising dif- fusion probabilistic models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.059616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.059616Z digest=sha256:41d73be8d8c3b9748b4d8cdadbfd1d9b808c1a8dcb2851a7b96b0769e81a4f1e

Observation 8d4c8905-cd80-4290-b30e-f903b78ff28a · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.065835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.065835Z digest=sha256:866f60a2e7d5ddcd2e68701c81e37c5e83b59654a12f8a9c0bac4bdb2ee24ec8

Observation faa48062-5193-4a7e-aa37-a5f34ea9e360 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention LoRA: Low-Rank Adaptation of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.073748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.073748Z digest=sha256:c83e05b1b7af1d3cdcf940d3305fbd818cf3fbd21bc6d341d3264aa485e3e80a

Observation 3d8ff2f5-d5e5-4038-98d2-736241d6baa1 · outbound

This paper cites Monocular quasi-dense 3d object tracking.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Monocular quasi-dense 3d object tracking

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.131127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.079303Z digest=sha256:75a693aa0ef64f3250e87395f179c562c91fd858c292fc2ba9dde0454e821e91

Observation 21c478e2-8821-486f-9f41-d911576dc833 · outbound

This paper cites BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.085759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.085759Z digest=sha256:01dd39c7d22717c7154884c5c05702677ba21291feef5c0fda662f79da69c05a

Observation 234fdd8c-f7b5-408d-98d2-ef31f83c27ef · outbound

This paper cites An- chor3dlane: Learning to regress 3d anchors for monocular 3d lane detection.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention An- chor3dlane: Learning to regress 3d anchors for monocular 3d lane detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.113541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.092727Z digest=sha256:cc5d4591e633fead6b74374cfb8e6e6f59250a869ffd1cd3c27e4bc55f0cfe76

Observation ccc49ff7-1c36-41f8-a2d0-748695b5dc00 · outbound

This paper cites Polarformer: Multi- camera 3d object detection with polar transformer.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Polarformer: Multi- camera 3d object detection with polar transformer

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.093141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.099733Z digest=sha256:444d305f9095d944ba84ead573e3fd2ef548e4d1636724e2f10c4506ba83c07a

Observation 32869267-ec82-4f2b-a468-bb5864902378 · outbound

This paper cites DrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion model.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention DrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.107476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.107476Z digest=sha256:7fbc96066a3843d9418f0f1e17278084869c9261bdd3620b2155ff5a33730016

Observation 5a5cdd1c-2e21-46e4-bf4b-d4c062eb42d9 · outbound

This paper cites Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.116016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.116016Z digest=sha256:d634a31a25fa1fdb558a080984fdac8672a42e1f0613615b63910d7591fb5a3f

Observation 5c685560-50f8-4630-bb8a-69f75013730f · outbound

This paper cites Nuwa-infinity: Autoregressive over autoregressive generation for infinite visual synthesis.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Nuwa-infinity: Autoregressive over autoregressive generation for infinite visual synthesis

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.059341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.122969Z digest=sha256:0db1cda5f00517450fc47faab0a1e10a7709ed3bfc671877670df700c1e4a337

Observation fd64fd1b-6af8-4d7f-b4aa-d75ea977888b · outbound

This paper cites Petrv2: A unified framework for 3d perception from multi-camera images.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Petrv2: A unified framework for 3d perception from multi-camera images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.130309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.130309Z digest=sha256:2d05618eb3dd30fe3f75cfc6db2c4ce339abf01a95380aa7f11ee1f43c3f0025

Observation 30d3c843-9f98-4948-ac27-b914a24a61da · outbound

This paper cites Bevfusion: Multi- task multi-sensor fusion with unified bird’s-eye view repre- sentation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Bevfusion: Multi- task multi-sensor fusion with unified bird’s-eye view repre- sentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.027546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.136066Z digest=sha256:451850bf2786c6220a057a1ec168cf136ba2b6c1850dae1a52b98f1998684e4f

Observation 9deff6c3-5270-4434-b757-7df5932e5532 · outbound

This paper cites On aliased resizing and surprising subtleties in gan evaluation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention On aliased resizing and surprising subtleties in gan evaluation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.007125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.141853Z digest=sha256:651e56078b4ec9c7d86051bb000c3053d33abbe253c8b65e481ad24fcca06126

Observation a01c8e0b-e0b5-4b6f-a10e-38efd6d605ad · outbound

This paper cites Scalable diffusion models with transformers.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Scalable diffusion models with transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.147114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.147114Z digest=sha256:b292c44a97712a2883f015e86206e0fac1932fa4a6323ace28a6a74181babafe

Observation ea662dcd-aef0-454f-92db-56837ce83e0f · outbound

This paper cites ControlNeXt: Powerful and Efficient Control for Image and Video Generation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention ControlNeXt: Powerful and Efficient Control for Image and Video Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.152774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.152774Z digest=sha256:c09be953311cbc4b8deb9aa6a26515ba0fb911161ded77859c674f515a2c670f

Observation d6db93f6-0960-4679-bafa-754cb6342bc0 · outbound

This paper cites an unresolved cited work.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:21:55.974041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.157939Z digest=sha256:cb8c8a9c62ae1f3a77af37d8f103580e32227284dc040fe3689885e44c28b00f

Observation f82c7d4a-cfcc-4fe8-b1ec-3493e5d64b1e · outbound

This paper cites Frame-recurrent video super-resolution.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Frame-recurrent video super-resolution

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.955922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.162844Z digest=sha256:dd19b12e9df8bd424da68d60e1f2974065b0c9802a7f7c6ba2e3244d008c7297

Observation d1f2369d-96c3-470f-8152-9eecd07d1fa3 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Score-Based Generative Modeling through Stochastic Differential Equations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.168004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.168004Z digest=sha256:12eec2afab503cc222e6f8bf201b87c3e0bf6b62383e842eefa1d7ba7b29f8fc

Observation c00b3bed-531c-4656-8a72-505a9c0dcff6 · outbound

This paper cites Street- view image generation from a bird’s-eye view layout.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Street- view image generation from a bird’s-eye view layout

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.937691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.173906Z digest=sha256:c2415959c801e282642b7e63bb7130d3a55f26591036b1f5c484307a046afcb0

Observation 6ab425f2-8602-4c93-a212-b00a64c7df9b · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.179249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.179249Z digest=sha256:1d617c2b5e90263cc72d54e85634aa696768eef590756c952ced876857e87948

Observation 918468f5-232a-47f5-97d6-c84887e05b5b · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention ModelScope Text-to-Video Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.185113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.185113Z digest=sha256:fab5912ac6003a5711e23a6e1e872fe4b69a593be6cfe9fd391a0157d77c3d34

Observation 52e689b0-0447-45a6-8c07-7e6606a3d111 · outbound

This paper cites Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.190603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.190603Z digest=sha256:3062faa6874b7e7ac988519240f322c7aaaccf9c78b0acf8a723bc821d8b894e

Observation 5d018bf1-1197-4bca-bfd4-dea17f858687 · outbound

This paper cites DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.197273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.197273Z digest=sha256:d04e1106b51074db65ed425758ccfa789623afded315b9d8ed62722269f2be75

Observation bcf02dd3-1b15-468a-9809-cfbf8b2ce20c · outbound

This paper cites End-to-end video instance segmentation with transformers.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention End-to-end video instance segmentation with transformers

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.918436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.203627Z digest=sha256:9ac551a162728a09a8b98a64709a536e9dff4ff7e6eb6729dc501954703ce664

Observation bc3b5566-be2f-4b5e-a57d-c2abb98257e0 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.209284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.209284Z digest=sha256:eb023fb0d5fea7531b9e423ed13cb5c7cf21985483a5aa11ad6b5b3f1026199e

Observation 83bbba93-a6a1-445e-b330-d92470a6a4cb · outbound

This paper cites Panacea: Panoramic and controllable video generation for autonomous driving.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Panacea: Panoramic and controllable video generation for autonomous driving

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.889904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.216870Z digest=sha256:26201d6a1910afbf88b5da6e2b11ff6c294774b5550805e3e65909a9b8bb8e24

Observation 616c4db0-f7b0-434d-a618-77032c9d0853 · outbound

This paper cites GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.224354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.224354Z digest=sha256:37261b46988fe773f295488b874e37720b9721cd53f64d4c61db9f1deaf123ea

Observation d905c105-1d8b-4e68-b6e2-a93531c3682a · outbound

This paper cites BEVControl: Accurately Controlling Street-view Elements with Multi-perspective Consistency via BEV Sketch Layout.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention BEVControl: Accurately Controlling Street-view Elements with Multi-perspective Consistency via BEV Sketch Layout

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.231028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.231028Z digest=sha256:e58423ca45a62996ead8c0ab3f1d707ca996f720eba4fdab9285c8cc8e7b7b5a

Observation 1b81d4a7-90d3-46a7-881f-18d52fb590f5 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.237776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.237776Z digest=sha256:8b6ecb9e9f8ed5e6343e48de2f91bf49e05eb62115ae648460b776179e235c6a

Observation 9da96595-be7c-4b71-a993-b8c9d52f357f · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Adding conditional control to text-to-image diffusion models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.243096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.243096Z digest=sha256:3d54f5ae13b0a5bfd16f813cfa8041e681f9e69e55100759dda275b0f2f28b04

Observation a2ddb8f7-c0f9-469e-b06b-0f577fdf893e · outbound

This paper cites Mutr3d: A multi-camera tracking frame- work via 3d-to-2d queries.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Mutr3d: A multi-camera tracking frame- work via 3d-to-2d queries

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.248379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.248379Z digest=sha256:5779551f2edfd5b5d72000c2960e8ccec4b117407d75302ae6488a01f435ecf7

Observation 8d33a151-7606-4c35-93be-0aa9408cd3e3 · outbound

This paper cites ByteTrackV2: 2D and 3D Multi-Object Tracking by Associating Every Detection Box.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention ByteTrackV2: 2D and 3D Multi-Object Tracking by Associating Every Detection Box

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-11T22:21:55.368563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.254764Z digest=sha256:1b1479c48639065d6e220c43aa07284c3ad3c236189a89b3525ec176077bd568

Observation f2644c47-3b6b-4252-864b-d9736baa5fde · outbound

This paper cites ControlVideo: Training-free Controllable Text-to-Video Generation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention ControlVideo: Training-free Controllable Text-to-Video Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.261327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.261327Z digest=sha256:786e750dad2bc2189dc1b66b440cc36b84ebb6d419bf6bd5ca29e6bb71badff6

Observation ec6acbee-d3fb-497f-9b72-d277c630d8da · outbound

This paper cites Cross-view transform- ers for real-time map-view semantic segmentation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Cross-view transform- ers for real-time map-view semantic segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.845580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.266891Z digest=sha256:5a3d75e49f1f87c0cc8c925410146d1b1e34fbe264a8c7ccf48229e5ccf8b5af

Observation bc1f59bb-0435-41fa-9d3c-39f09eb49b0e · outbound

This paper cites For MagicDrive-V , we utilize the official pretrained weights to generate videos on the nuScenes validation split.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention For MagicDrive-V , we utilize the official pretrained weights to generate videos on the nuScenes validation split

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.822689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.272750Z digest=sha256:22c9b2b8e970a00aa500d767bb3747c1d2fc309bb63f095bc84c651c52fc3f70

Observation 88ecadad-147f-49ce-8555-f4361ddd6a52 · outbound

This paper cites Since MagicDrive-V [7] is the only multi-view video generation method that releases model weights, we use the generation results from MagicDrive-V for comparison.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Since MagicDrive-V [7] is the only multi-view video generation method that releases model weights, we use the generation results from MagicDrive-V for comparison

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.798258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.279255Z digest=sha256:6aae296510344bcf9934f00ca063208ccd6ec0763430a4807a06c8aacc9c92e0

Observation f82df3af-b086-493e-a5cc-15d2257c23ef · outbound

This paper cites We preprocess these videos by segmenting them into clips of 16 frames.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention We preprocess these videos by segmenting them into clips of 16 frames

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.773035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:21:55.284941Z digest=sha256:b4fc223ef2c0d6805d001a352c3b72871c97a42548b255ae08bf30096de35562

Pith citing papers

Observation f91311a3-2b91-4fbe-92b0-34b289a2ee7f · inbound

3D and 4D World Modeling: A Survey cites this paper.

3D and 4D World Modeling: A Survey Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-05T06:04:25.857689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:04:25.857689Z digest=sha256:725aaf8e5d79dfe41333a80739f5a551c79cf94035f2e1aca18e998c2bdca645

Observation 7d11b940-307e-4be8-838e-b3e1f365bffd · inbound

AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond cites this paper.

AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:04:00.917788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T21:15:12.465970Z digest=sha256:0acf6f3c45ef7aa2a78e0fe5cf66782ff57851441ddfbe74392f8ad897e35b9b

Observation 8936245c-c2a3-481b-ad45-94e36a4fdb7e · inbound

FrozenDrive: Zero-Shot Text-Guided Driving Scene Generation and Data Augmentation with Parameter-Free Frozen Diffusion Model cites this paper.

FrozenDrive: Zero-Shot Text-Guided Driving Scene Generation and Data Augmentation with Parameter-Free Frozen Diffusion Model Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:59:26.818843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T18:34:56.810124Z digest=sha256:91b8dc6f65147920935acb1fe5f1069444369506faed53338ededdb301f0d27b

Observation b9abc561-2a76-422b-ae75-2b9e445820e8 · inbound

OpenLongTail: Generative Scaling of Long-Tail Driving Data cites this paper.

OpenLongTail: Generative Scaling of Long-Tail Driving Data Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T01:26:27.220907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:26:27.220907Z digest=sha256:4348b530510d62835f57571f3e87241f6f648b21c9fcb0e06ae8e20f2f1ca628