Pith. sign in

Paper Citation Record · LEDGER

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving

As of 14 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2507.00707.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00707 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:16:03.985223Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:23:06.577719Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 769ed58f-a6e8-4c27-a985-0697b7262a43 · outbound

This paper cites MagicDrive: Street View Generation with Diverse 3D Geometry Control.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving MagicDrive: Street View Generation with Diverse 3D Geometry Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:02.509759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:02.509759Z digest=sha256:b2bb2a073378bd854fc77b7a09180af03583e5b5c6e27f8b48263e5d718d8676

Observation 2d42f68f-3106-496a-8c2a-f72eabd645e4 · outbound

This paper cites Drivingdiffusion: Layout-guided multi-view driving scenarios video generation with latent diffusion model.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Drivingdiffusion: Layout-guided multi-view driving scenarios video generation with latent diffusion model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:06.591249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:02.561112Z digest=sha256:01c5d68f0cc5b35aa6678192d1afd4ab4e1cac5bfbe9c7d121c570ada5954665

Observation a08a8240-46a5-4bf8-a291-8ee4fb7feb20 · outbound

This paper cites Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:06.496258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:02.664596Z digest=sha256:9a8bb25d4c2d5e0ba5fe363e309106c1a572031e837fd932e3ae5150d41e6404

Observation 0aea0997-6f15-4b82-8a56-087ebb3d8108 · outbound

This paper cites Panacea: Panoramic and controllable video generation for autonomous driving.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Panacea: Panoramic and controllable video generation for autonomous driving

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:06.349411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:02.708374Z digest=sha256:96c553e0f81e378aa4912ea614caca9c4f7e92c3c06d52fc4f0f21eb1747fb20

Observation dedffad0-b373-4094-84be-9228e6e88250 · outbound

This paper cites Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:06.220621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:02.788028Z digest=sha256:1a7df804ce35f40af01baa50362f58927b4cd245ac056e526a47072a19ae6b27

Observation 20d6d946-f60b-4246-b128-04f842f7e7e0 · outbound

This paper cites BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:02.821152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:02.821152Z digest=sha256:5dcc414141fcd36b262d9f0f6dcdf3e261f9bfa3b38d9cd1b1afb1cfcb880caa

Observation bc3292b6-0aea-43c0-821c-020df08a96bb · outbound

This paper cites Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:06.095644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:02.876874Z digest=sha256:b9281abd6081e6d5f4d274359058dcc488f9d42e44d585ab3db353716ac7a62a

Observation 7d8bc499-4011-492a-8171-59a739b5ecd2 · outbound

This paper cites Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal trans- formers.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal trans- formers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.932218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:02.916508Z digest=sha256:2f3eed061149481bcf920c0b1ddf258ae1b0924aa75373746a21fa132b798335

Observation b1327940-5b5d-456f-990f-f3ab4fa2f54a · outbound

This paper cites Planning-oriented autonomous driving.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Planning-oriented autonomous driving

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.864956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:02.947076Z digest=sha256:24f64e8492ab459403d80019151745d6b48648b2e5a3b9882a9d29ffb94ffa4f

Observation 41cdfdbb-c5af-4ee5-aa2d-94df67354be3 · outbound

This paper cites Neural discrete representation learning.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Neural discrete representation learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:02.972118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:02.972118Z digest=sha256:f5488be559561e7f35ac350cba57ae81857c352fd390690bb091e62e3b2323d6

Observation 2257872d-bf9c-41fd-921e-c46278b9be47 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Taming transformers for high-resolution image synthesis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.708350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:03.040381Z digest=sha256:e1db5aa780d8886c7ef060627d78edcb5350ad9a76ef942f77c209425492c1fa

Observation bad47d59-1fd0-4b00-ac7c-46dee8ec2bec · outbound

This paper cites Generative adversarial nets.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Generative adversarial nets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:03.085635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:03.085635Z digest=sha256:7855df46e994f4da420ff9c9b50a4d7cc82ab31787121276b4ef6c8a41fd7dd3

Observation 4d9b0286-5eee-429b-bbc6-ffa31108cd44 · outbound

This paper cites Perceptual losses for real-time style transfer and super- resolution.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Perceptual losses for real-time style transfer and super- resolution

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.567907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:03.122390Z digest=sha256:f3f822ec1b96703e4b6a9777d91a44f8e5de5df0b4a9a0c0c2b48e70630926f6

Observation 4292a0d8-a245-4237-9f74-c6be65f04009 · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Vector-quantized Image Modeling with Improved VQGAN

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:03.158583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:03.158583Z digest=sha256:66ec1031c96985fa90a17bcaf12770f70eb0aa6b9c61c34009a347c322edcc3a

Observation cf18d2c8-c8e9-4044-8b52-e7fa5819b09d · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:03.199718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:03.199718Z digest=sha256:35c2016c62f74c01f3b4a5c46d5fb068da703ae9cd541010752cee5649d41107

Observation a4cd8744-61ce-4a54-ad7d-f2f08314b127 · outbound

This paper cites Denoising diffusion probabilistic models.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Denoising diffusion probabilistic models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:03.249349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:03.249349Z digest=sha256:c1327e01b39ef046068c036bd2f7a4af1f696f6f3b4c99660745194ee0c2ad30

Observation 70a4ed66-17e0-4a35-b601-b46b0efe62ab · outbound

This paper cites Denoising Diffusion Implicit Models.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Denoising Diffusion Implicit Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:03.347096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:03.347096Z digest=sha256:7178bd51c97bb47055e1ec327612092d17d5ef9385023a77703388724ca0aab0

Observation 461e5eb7-26cf-4b27-97d6-155d1422d189 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving High-resolution image synthesis with latent diffusion models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.434381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:03.382308Z digest=sha256:e15f5e69a8a7a7eaf054951f39baa61b0fa65f6717d4ed9cccefb5b312a83f5f

Observation 0b771eb2-ae24-431b-ba13-cdd9712f509e · outbound

This paper cites Scalable diffusion models with transformers.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Scalable diffusion models with transformers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.262604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:03.466342Z digest=sha256:640af3251a94398a5939b8c437acdbe2a3596562e5428d18cbcb737970ba1dc5

Observation b028a2df-cc99-4301-b53a-73e3d895b032 · outbound

This paper cites Street-view image generation from a bird’s-eye view layout.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Street-view image generation from a bird’s-eye view layout

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.169444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:03.574745Z digest=sha256:0a296a54f6e422a292bdad8c5fbd321ffa62f39e747ff505fe7a19649ff47179

Observation 6d9dfdbf-2b59-41d4-8021-5289a736b6f4 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthesis.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Nerf: Representing scenes as neural radiance fields for view synthesis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.063741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:03.662211Z digest=sha256:4ed0cb566c228f9d119e6174210c7035f32678d118979169144244bea123a79e

Observation e4caaf21-99d8-4422-bdf4-d9893319697d · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Adding conditional control to text-to-image diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:04.897313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:03.734048Z digest=sha256:46055c4b091678ca79e50b0c055fd71fe237cb47ae656926ef488a646ae29fb4

Observation b1852ddb-eacc-40e9-9029-fe74ada547e5 · outbound

This paper cites Feature pyramid networks for object detection.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Feature pyramid networks for object detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:04.679187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:03.802897Z digest=sha256:ac25cc9b3b583b7ffe9176589555198f0176436f199e5232d2d9a69a119d84cf

Observation dd832207-7098-46cb-aa28-d4769f758fd6 · outbound

This paper cites Loftr: Detector-free local feature matching with transformers.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Loftr: Detector-free local feature matching with transformers

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:04.447945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:03.905488Z digest=sha256:267f599e484905a4e24e4ed8e5633092067b7d1529c755a82cfe1989dedc2113

Observation f846a9e7-da1a-44c9-9d91-44a6ce011637 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:04.213139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:16:03.985223Z digest=sha256:d24f97a4002c444e183d20eb06a2c31860f77557ee60f2b4fefa0185c90cfb18

Pith citing papers

Observation dcea4640-d784-4eb3-bbef-fedbc86f5832 · inbound

Variational Inference for Bird's Eye View Segmentation in Autonomous Driving cites this paper.

Variational Inference for Bird's Eye View Segmentation in Autonomous Driving BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:23:06.577719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:23:06.577719Z digest=sha256:a900a06c919f0502e60635098f5462e5b8b797aab4784582dce8d86f6694c295