Pith. sign in

Paper Citation Record · LEDGER

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

As of 16 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 4 inbound Pith citation observations for arXiv:2412.03520.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03520 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:21:55.284941Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T06:04:25.857689Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T02:59:26.816825Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1261bc12-a048-4ad5-872f-b5c867e132bc · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.007365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.007365Z digest=sha256:1ced2be075e60663702a8cedc08146a1744c2f2cb1200b56e583cf49934a37ff

Observation 98067dba-b1f7-463c-8922-ce9af8599900 · outbound

This paper cites Video generation models as world simulators.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Video generation models as world simulators

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.015259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.015259Z digest=sha256:0b78c60b36819d70403104e823ae27037925b3ca316eae8d9506ed66ff153287

Observation 7402b4cf-592c-4764-a99a-55ace1f146cb · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention nuscenes: A multi- modal dataset for autonomous driving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.022805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.022805Z digest=sha256:acc2967f706c581cf782957e862acc97699ef7123f188293731f4d1a34fad2f4

Observation c493a015-0b0b-429e-869c-8c1eb6f50845 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.029391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.029391Z digest=sha256:1658f31b1519ade00049b9a7c2403a75a50165f6c7d92a9852b7dda268dd09a3

Observation 6d9d3a85-22cd-4eab-ae5d-e96b73321327 · outbound

This paper cites VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.035366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.035366Z digest=sha256:50121766d2c00f369e8792f49e37b61a93b5eba4fda6d29e1e86eefd77677091

Observation 7183b1cd-2c72-49ef-84b9-96139a831d55 · outbound

This paper cites Persformer: 3d lane detection via perspective transformer and the openlane benchmark.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Persformer: 3d lane detection via perspective transformer and the openlane benchmark

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.188214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.041882Z digest=sha256:546fec77f5583ad4dad69aed387a94dbe7aec600ed7b4882446ec476844ddde5

Observation 286acbde-e74d-46a9-afbb-aa259be4d544 · outbound

This paper cites MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.047416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.047416Z digest=sha256:f4aab9cf972c6c467ca69f0216caa9f782245f3aedf58e47ac6db993f4595803

Observation 99d89f3e-ca6c-4972-aee5-821e6c7f0f08 · outbound

This paper cites Magicdrive: Street view generation with diverse 3d geometry control.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Magicdrive: Street view generation with diverse 3d geometry control

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.168071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.053138Z digest=sha256:9944b4908162f548294995fd398caadface161efa5d27d062eedc412aa8cff15

Observation d43296b9-aa0b-4b07-8a8b-7d63700ee705 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Denoising dif- fusion probabilistic models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.059616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.059616Z digest=sha256:41d73be8d8c3b9748b4d8cdadbfd1d9b808c1a8dcb2851a7b96b0769e81a4f1e

Observation 8d4c8905-cd80-4290-b30e-f903b78ff28a · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.065835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.065835Z digest=sha256:866f60a2e7d5ddcd2e68701c81e37c5e83b59654a12f8a9c0bac4bdb2ee24ec8

Observation faa48062-5193-4a7e-aa37-a5f34ea9e360 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention LoRA: Low-Rank Adaptation of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.073748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.073748Z digest=sha256:c83e05b1b7af1d3cdcf940d3305fbd818cf3fbd21bc6d341d3264aa485e3e80a

Observation 3d8ff2f5-d5e5-4038-98d2-736241d6baa1 · outbound

This paper cites Monocular quasi-dense 3d object tracking.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Monocular quasi-dense 3d object tracking

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.131127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.079303Z digest=sha256:e544b74a50ebd59459acf79d81e7567d347f589ae8e8e2703a0511859822fccc

Observation 21c478e2-8821-486f-9f41-d911576dc833 · outbound

This paper cites BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.085759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.085759Z digest=sha256:01dd39c7d22717c7154884c5c05702677ba21291feef5c0fda662f79da69c05a

Observation 234fdd8c-f7b5-408d-98d2-ef31f83c27ef · outbound

This paper cites An- chor3dlane: Learning to regress 3d anchors for monocular 3d lane detection.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention An- chor3dlane: Learning to regress 3d anchors for monocular 3d lane detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.113541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.092727Z digest=sha256:3457e34976382e681273d8f39bb65df00a80abe7b1dbfb5552b4bbb3269149c1

Observation ccc49ff7-1c36-41f8-a2d0-748695b5dc00 · outbound

This paper cites Polarformer: Multi- camera 3d object detection with polar transformer.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Polarformer: Multi- camera 3d object detection with polar transformer

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.093141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.099733Z digest=sha256:63f2b396bc620afdf70d9ddd91e81e008725824a755dc15aa876c40e576d7111

Observation 32869267-ec82-4f2b-a468-bb5864902378 · outbound

This paper cites DrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion model.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention DrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.107476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.107476Z digest=sha256:cb1d5666e157ae7d25a448cc3580048cb922459566a11c579fc21cd4831fdc99

Observation 5a5cdd1c-2e21-46e4-bf4b-d4c062eb42d9 · outbound

This paper cites Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.116016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.116016Z digest=sha256:d634a31a25fa1fdb558a080984fdac8672a42e1f0613615b63910d7591fb5a3f

Observation 5c685560-50f8-4630-bb8a-69f75013730f · outbound

This paper cites Nuwa-infinity: Autoregressive over autoregressive generation for infinite visual synthesis.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Nuwa-infinity: Autoregressive over autoregressive generation for infinite visual synthesis

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.059341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.122969Z digest=sha256:d18706cc713ebce932670894f5be053f4efd712b1daa8dae8dd6eef75ba36847

Observation fd64fd1b-6af8-4d7f-b4aa-d75ea977888b · outbound

This paper cites Petrv2: A unified framework for 3d perception from multi-camera images.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Petrv2: A unified framework for 3d perception from multi-camera images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.130309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.130309Z digest=sha256:2d05618eb3dd30fe3f75cfc6db2c4ce339abf01a95380aa7f11ee1f43c3f0025

Observation 30d3c843-9f98-4948-ac27-b914a24a61da · outbound

This paper cites Bevfusion: Multi- task multi-sensor fusion with unified bird’s-eye view repre- sentation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Bevfusion: Multi- task multi-sensor fusion with unified bird’s-eye view repre- sentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.027546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.136066Z digest=sha256:d04c909d0d747be80bddf6e6682c641591c6800082baee34ca2d6961aeb44e24

Observation 9deff6c3-5270-4434-b757-7df5932e5532 · outbound

This paper cites On aliased resizing and surprising subtleties in gan evaluation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention On aliased resizing and surprising subtleties in gan evaluation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:56.007125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.141853Z digest=sha256:2979c0a35adc32e1f2ab6870df261a21651ce0347697291c7ff305fda47fc17a

Observation a01c8e0b-e0b5-4b6f-a10e-38efd6d605ad · outbound

This paper cites Scalable diffusion models with transformers.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Scalable diffusion models with transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.147114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.147114Z digest=sha256:b292c44a97712a2883f015e86206e0fac1932fa4a6323ace28a6a74181babafe

Observation ea662dcd-aef0-454f-92db-56837ce83e0f · outbound

This paper cites ControlNeXt: Powerful and Efficient Control for Image and Video Generation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention ControlNeXt: Powerful and Efficient Control for Image and Video Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.152774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.152774Z digest=sha256:45e0181ef08f66829279430fbc67e31ef360e29667a8a7f4f964b93903bf4600

Observation d6db93f6-0960-4679-bafa-754cb6342bc0 · outbound

This paper cites an unresolved cited work.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:21:55.974041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.157939Z digest=sha256:f8ccbe418aca1fa729eeb53cd4624dbf363f8e15b39f72a24c5e1194904be01b

Observation f82c7d4a-cfcc-4fe8-b1ec-3493e5d64b1e · outbound

This paper cites Frame-recurrent video super-resolution.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Frame-recurrent video super-resolution

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.955922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.162844Z digest=sha256:26954e100952b40828c79a3b6fb23fd6713364f8dc4fbd9f270cf7c1c80e6eb1

Observation d1f2369d-96c3-470f-8152-9eecd07d1fa3 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Score-Based Generative Modeling through Stochastic Differential Equations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.168004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.168004Z digest=sha256:12eec2afab503cc222e6f8bf201b87c3e0bf6b62383e842eefa1d7ba7b29f8fc

Observation c00b3bed-531c-4656-8a72-505a9c0dcff6 · outbound

This paper cites Street- view image generation from a bird’s-eye view layout.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Street- view image generation from a bird’s-eye view layout

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.937691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.173906Z digest=sha256:40448197885040c1800a93688b4d104fe996696768f6ed1c4c1bc87ab3fc4be8

Observation 6ab425f2-8602-4c93-a212-b00a64c7df9b · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.179249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.179249Z digest=sha256:1d617c2b5e90263cc72d54e85634aa696768eef590756c952ced876857e87948

Observation 918468f5-232a-47f5-97d6-c84887e05b5b · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention ModelScope Text-to-Video Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.185113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.185113Z digest=sha256:fab5912ac6003a5711e23a6e1e872fe4b69a593be6cfe9fd391a0157d77c3d34

Observation 52e689b0-0447-45a6-8c07-7e6606a3d111 · outbound

This paper cites Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.190603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.190603Z digest=sha256:7dc3e330dc88ba8eb058ac69864ed597f2309299da7b313a98d0f4f5f909a7df

Observation 5d018bf1-1197-4bca-bfd4-dea17f858687 · outbound

This paper cites DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.197273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.197273Z digest=sha256:dfe03288b35b4169b636ae374e96db1807bf3455841899aa89a8beb77249339c

Observation bcf02dd3-1b15-468a-9809-cfbf8b2ce20c · outbound

This paper cites End-to-end video instance segmentation with transformers.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention End-to-end video instance segmentation with transformers

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.918436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.203627Z digest=sha256:fe417cbb4a5cc6baa7f1a2c7ac6e794a2292ebe87066691e25bba42f3c12aff9

Observation bc3b5566-be2f-4b5e-a57d-c2abb98257e0 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.209284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.209284Z digest=sha256:2a9bce38b1741e08f61bce099519909578e4d6227d8c8b946540b8faadc3210b

Observation 83bbba93-a6a1-445e-b330-d92470a6a4cb · outbound

This paper cites Panacea: Panoramic and controllable video generation for autonomous driving.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Panacea: Panoramic and controllable video generation for autonomous driving

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.889904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.216870Z digest=sha256:90b353b517d5c94aaea063666218af4e55c1d4df7f5939d86c02f739bbc2504d

Observation 616c4db0-f7b0-434d-a618-77032c9d0853 · outbound

This paper cites GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.224354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.224354Z digest=sha256:37261b46988fe773f295488b874e37720b9721cd53f64d4c61db9f1deaf123ea

Observation d905c105-1d8b-4e68-b6e2-a93531c3682a · outbound

This paper cites BEVControl: Accurately Controlling Street-view Elements with Multi-perspective Consistency via BEV Sketch Layout.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention BEVControl: Accurately Controlling Street-view Elements with Multi-perspective Consistency via BEV Sketch Layout

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.231028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.231028Z digest=sha256:0d56d35642a832e79a2fd56c829b95c275a4b342b7e7bf1bfd5e1218369a0c08

Observation 1b81d4a7-90d3-46a7-881f-18d52fb590f5 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.237776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.237776Z digest=sha256:8b6ecb9e9f8ed5e6343e48de2f91bf49e05eb62115ae648460b776179e235c6a

Observation 9da96595-be7c-4b71-a993-b8c9d52f357f · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Adding conditional control to text-to-image diffusion models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.243096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.243096Z digest=sha256:3d54f5ae13b0a5bfd16f813cfa8041e681f9e69e55100759dda275b0f2f28b04

Observation a2ddb8f7-c0f9-469e-b06b-0f577fdf893e · outbound

This paper cites Mutr3d: A multi-camera tracking frame- work via 3d-to-2d queries.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Mutr3d: A multi-camera tracking frame- work via 3d-to-2d queries

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.248379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.248379Z digest=sha256:5779551f2edfd5b5d72000c2960e8ccec4b117407d75302ae6488a01f435ecf7

Observation 8d33a151-7606-4c35-93be-0aa9408cd3e3 · outbound

This paper cites ByteTrackV2: 2D and 3D Multi-Object Tracking by Associating Every Detection Box.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention ByteTrackV2: 2D and 3D Multi-Object Tracking by Associating Every Detection Box

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-11T22:21:55.368563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.254764Z digest=sha256:94fb410c9d8dd1e4b68e86119c5be01c2fb367f3989a901f5b9f10949c70eac1

Observation f2644c47-3b6b-4252-864b-d9736baa5fde · outbound

This paper cites ControlVideo: Training-free Controllable Text-to-Video Generation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention ControlVideo: Training-free Controllable Text-to-Video Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.261327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.261327Z digest=sha256:eb679623173fdc21832eaac87ffe4d74ba255e59f751a7fbb9632b436e71b365

Observation ec6acbee-d3fb-497f-9b72-d277c630d8da · outbound

This paper cites Cross-view transform- ers for real-time map-view semantic segmentation.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Cross-view transform- ers for real-time map-view semantic segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.845580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.266891Z digest=sha256:074819f40f0929334e5bcb554a11ec3d78d616c359e78ae568e7ab5d55830f89

Observation bc1f59bb-0435-41fa-9d3c-39f09eb49b0e · outbound

This paper cites For MagicDrive-V , we utilize the official pretrained weights to generate videos on the nuScenes validation split.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention For MagicDrive-V , we utilize the official pretrained weights to generate videos on the nuScenes validation split

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.822689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.272750Z digest=sha256:dc0cc9e795b5f8238550c15991ea62f9861dadd8b40e624bd5b739378f55419a

Observation 88ecadad-147f-49ce-8555-f4361ddd6a52 · outbound

This paper cites Since MagicDrive-V [7] is the only multi-view video generation method that releases model weights, we use the generation results from MagicDrive-V for comparison.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Since MagicDrive-V [7] is the only multi-view video generation method that releases model weights, we use the generation results from MagicDrive-V for comparison

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.798258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.279255Z digest=sha256:893a8ead8524b0aa4c856f90881e7ad6c3269d7ea79cd5d570df41ba0856b52f

Observation f82df3af-b086-493e-a5cc-15d2257c23ef · outbound

This paper cites We preprocess these videos by segmenting them into clips of 16 frames.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention We preprocess these videos by segmenting them into clips of 16 frames

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:21:55.773035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:21:55.284941Z digest=sha256:0c26dd3642bacbeff84d70b29efe708e5ce7ba0e66c6dda844b6f437f2eb2954

Pith citing papers

Observation f91311a3-2b91-4fbe-92b0-34b289a2ee7f · inbound

3D and 4D World Modeling: A Survey cites this paper.

3D and 4D World Modeling: A Survey Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-05T06:04:25.857689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:04:25.857689Z digest=sha256:725aaf8e5d79dfe41333a80739f5a551c79cf94035f2e1aca18e998c2bdca645

Observation 7d11b940-307e-4be8-838e-b3e1f365bffd · inbound

AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond cites this paper.

AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:04:00.917788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T21:15:12.465970Z digest=sha256:92675e18e62eb3037cdc9d771a05297d857bd8f78b96f1f5afe8f49ddf5de5bc

Observation 8936245c-c2a3-481b-ad45-94e36a4fdb7e · inbound

FrozenDrive: Zero-Shot Text-Guided Driving Scene Generation and Data Augmentation with Parameter-Free Frozen Diffusion Model cites this paper.

FrozenDrive: Zero-Shot Text-Guided Driving Scene Generation and Data Augmentation with Parameter-Free Frozen Diffusion Model Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:59:26.818843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T18:34:56.810124Z digest=sha256:af13ca30a848a5ef998e048af2385e1456daa4fede56d9d8eef1283d540e219e

Observation b9abc561-2a76-422b-ae75-2b9e445820e8 · inbound

OpenLongTail: Generative Scaling of Long-Tail Driving Data cites this paper.

OpenLongTail: Generative Scaling of Long-Tail Driving Data Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T01:26:27.220907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:26:27.220907Z digest=sha256:4348b530510d62835f57571f3e87241f6f648b21c9fcb0e06ae8e20f2f1ca628