Pith. sign in

Paper Citation Record · LEDGER

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting

As of 21 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 4 inbound Pith citation observations for arXiv:2507.09144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09144 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:08:54.630490Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:31:52.653212Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T06:04:21.611097Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy54
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2bbb9c59-7659-4e8c-b65c-7d0c19f7301d · outbound

This paper cites GPT-4 Technical Report.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:48.426576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:08:48.426576Z digest=sha256:eb3d8f7aa9f80c0e7e96dc6ad2179b0ce3daedf1ad7f78698e12a9c19d2e6a03

Observation 66e84cc4-ba7a-4fea-87e8-4649dd5ce474 · outbound

This paper cites Cosmos world foundation model platform for physical ai.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Cosmos world foundation model platform for physical ai

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:05.081203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:48.668292Z digest=sha256:1d2dbffe71cf56d34b8af0a6ab35ae3153ef5b61e4510736db0e8c486aa8816d

Observation 965e659d-8f89-4c29-a8c4-1a91d51690e3 · outbound

This paper cites V-jepa 2: Self- supervised video models enable understanding, prediction and planning.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting V-jepa 2: Self- supervised video models enable understanding, prediction and planning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:04.858535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:49.219446Z digest=sha256:7f406f2b4493242303d753c6e931e5cf0d75ec2f7d71d8c92339a945ef2a97b1

Observation e18bab8f-8330-4d94-b8d7-08db64eae7e1 · outbound

This paper cites Semantickitti: A dataset for semantic scene understanding of lidar sequences.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Semantickitti: A dataset for semantic scene understanding of lidar sequences

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:04.637797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:49.517326Z digest=sha256:f36449883e2d442226b28e7745051b121769aad6b15ac0ae7610e5e44cfc71eb

Observation d138711c-cb41-429a-806c-385803b99432 · outbound

This paper cites The lov ´asz-softmax loss: A tractable surrogate for the optimization of the intersection-over-union measure in neural networks.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting The lov ´asz-softmax loss: A tractable surrogate for the optimization of the intersection-over-union measure in neural networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:04.395568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:49.609521Z digest=sha256:9f4b93304e724ac72cfb533e903c3ea2312ca8d5964b938082f3f5d10147f29a

Observation 222df20f-8547-4aaf-8c8a-44d37659da11 · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:04.147704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:49.711273Z digest=sha256:5b0314c5606a7eeaa81cf38a405c9d050c8e8ebd41e2c7b281e3f0a2568bbb24

Observation 30610dcb-f2ad-4225-ae5c-3319f977822f · outbound

This paper cites Taming transformers for high-resolution image synthesis.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Taming transformers for high-resolution image synthesis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:03.950736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:49.808930Z digest=sha256:7ae3521a52dfcd1f4004a8c43bf068f94dadd7d07f7438b13675cc08fb9a95ad

Observation 6cb1d39a-9ec5-4da8-959d-fbb62f2def67 · outbound

This paper cites A survey of world models for autonomous driving.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting A survey of world models for autonomous driving

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:03.829784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:49.879554Z digest=sha256:a3a3998d440beb3d0775c39e436d2c279a1cd624a26977b9a2b70518e087e7a1

Observation 69a33bb2-4a40-43c9-bddf-56f726d7088d · outbound

This paper cites Magicdrive: Street view generation with diverse 3d geometry control.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Magicdrive: Street view generation with diverse 3d geometry control

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:03.709724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:49.965754Z digest=sha256:0f744f11f8474adbfde57b6935150f3af54e4f67348216d4c979d7c49d6e397c

Observation 65c97ca3-6e9a-4349-a51e-cf2bedf367a4 · outbound

This paper cites Vista: A generalizable driving world model with high fidelity and versatile controllability.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Vista: A generalizable driving world model with high fidelity and versatile controllability

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:03.536278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:50.039181Z digest=sha256:92b7b4d0d97a902294b922140c1e70d84c40cbd4332f9956b91fe69d2203e3a7

Observation 69b499b6-93b8-40fa-9686-373f74471843 · outbound

This paper cites DOME: Taming Diffusion Model into High-Fidelity Controllable Occupancy World Model.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting DOME: Taming Diffusion Model into High-Fidelity Controllable Occupancy World Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:50.127134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:08:50.127134Z digest=sha256:c871bfc2efb57e0562324547f2783b0bf51f5d842c33ed8b88d5672e084bba59

Observation b62e77aa-7eca-46f6-9bce-74fd1ac180f8 · outbound

This paper cites BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:50.235191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:08:50.235191Z digest=sha256:3e3448163acccd0ddfa706cdf04520dcfbc4b978d6f6fc8c701e0f1fdcbfb664

Observation b5f1f833-cbcc-475b-80a4-ec8076d79f0e · outbound

This paper cites Tri-perspective view for vision-based 3d se- mantic occupancy prediction.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Tri-perspective view for vision-based 3d se- mantic occupancy prediction

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:03.399408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:50.317153Z digest=sha256:a8e00921f4bc4c51527ed081c6327f7421fce00d59ea451d93a4166dadaf1c20

Observation 2f1c8941-a1f8-45af-8b68-6128e989fd7c · outbound

This paper cites Differentiable raycasting for self- supervised occupancy forecasting.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Differentiable raycasting for self- supervised occupancy forecasting

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:03.272688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:50.393125Z digest=sha256:8d224716cdeabf51266f0b77e23a168f2d71b2b3ecf33072dae0b9649111cb47

Observation 743d97d8-8b93-4b63-979b-e20815fa1e63 · outbound

This paper cites Point cloud forecasting as a proxy for 4d occupancy forecast- ing.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Point cloud forecasting as a proxy for 4d occupancy forecast- ing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:03.143906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:50.493807Z digest=sha256:de34de3abee96cadd600c7d3d52f3f54aa92c9e0ab5e033b5b05057b626ed089

Observation c08a5008-6392-4e0e-8251-69842482cb4c · outbound

This paper cites Pointpillars: Fast encoders for object detection from point clouds.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Pointpillars: Fast encoders for object detection from point clouds

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:02.998009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:50.560096Z digest=sha256:c187f17200f363f403eb25fcb0a6041c4f18f189618ad10ae9438d8cf87adcd8

Observation 82b2bad3-c536-47cb-a5b6-92c6fa2c0a5d · outbound

This paper cites Autoregressive image generation using resid- ual quantization.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Autoregressive image generation using resid- ual quantization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:02.760488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:50.633261Z digest=sha256:4de2bf32e512b17f1407fed922437b9c1d12d4df2070890eb7c78cf6ba76add2

Observation 7f6b9285-69a5-400f-b4d2-06beacf22cbe · outbound

This paper cites Uniscene: Unified occupancy-centric driving scene generation.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Uniscene: Unified occupancy-centric driving scene generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:02.498928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:50.728848Z digest=sha256:9a41be6ab126ededfc668a2f5e8533bff629b081a4f287110a633c36064d2d8c

Observation 7e130820-8fb4-4571-98c5-9f8758968697 · outbound

This paper cites Bevstereo: Enhancing depth estimation in multi-view 3d object detection with temporal stereo.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Bevstereo: Enhancing depth estimation in multi-view 3d object detection with temporal stereo

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:02.228951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:50.823568Z digest=sha256:2d0db1247b2ee6491f7033862ed24b14e09ff5bf87f4c1a4270cc046d67efc60

Observation a25f3e61-c40e-45d0-b959-15d23045e2b6 · outbound

This paper cites Bevdepth: Acquisition of reliable depth for multi-view 3d object detec- tion.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Bevdepth: Acquisition of reliable depth for multi-view 3d object detec- tion

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:01.997157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:50.919951Z digest=sha256:d67c8575515d507fba4634c1a1d08a11f011b1d07f305fa905aa90a840c49e57

Observation 34ebb596-8c32-4c92-a71f-dd79da3ea0dd · outbound

This paper cites V oxformer: Sparse voxel transformer for camera- based 3d semantic scene completion.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting V oxformer: Sparse voxel transformer for camera- based 3d semantic scene completion

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:01.753154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:51.017115Z digest=sha256:e32e30f378b87a9d1bb067f2a3ee36c2d36968cbd227322d54218a917733c78a

Observation e88a43db-e0d1-4b78-90db-45eeeac41333 · outbound

This paper cites Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:01.507827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:51.104435Z digest=sha256:d99758f85cdb19f73a962612654fd419eeb62e4d13d26b4f32a9a01dff6db4cd

Observation 54451612-7eee-4b13-af85-73dd2aa7caf2 · outbound

This paper cites Fb-occ: 3d occupancy prediction based on forward-backward view transformation.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Fb-occ: 3d occupancy prediction based on forward-backward view transformation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:01.369519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:51.303891Z digest=sha256:3c15e8468ebfaa1a0b48414ff0d2b8980f5cbae9413689674579da579b8bb405

Observation 83c3227f-c3fe-448b-b16c-bac86c11e6c0 · outbound

This paper cites Stcocc: Sparse spatial-temporal cascade renova- tion for 3d occupancy and scene flow prediction.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Stcocc: Sparse spatial-temporal cascade renova- tion for 3d occupancy and scene flow prediction

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:01.219357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:51.531927Z digest=sha256:20a5e677555dbfa60a85014790cdbee5a95e4167ef38947a47430459e8588239

Observation d1d715e2-b9ae-429c-9bdd-dd3b1685f073 · outbound

This paper cites Feature pyramid networks for object detection.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Feature pyramid networks for object detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:00.962012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:51.706309Z digest=sha256:2c4f68f78034e621978d26f58a11b11443560744c1effdb096db90cdd961f11e

Observation 61e59656-1b88-428f-ae1e-98855dbfa616 · outbound

This paper cites Sparsebev: High-performance sparse 3d object detec- tion from multi-camera videos.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Sparsebev: High-performance sparse 3d object detec- tion from multi-camera videos

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:00.821938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:51.885899Z digest=sha256:cc24dce74bac5c047526d734381cfdb5755362d1e3c0aeebf3711b2347b0dc8c

Observation 25fc74c4-001e-41fb-a08d-66a1b61921aa · outbound

This paper cites Decoupled Weight Decay Regularization.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Decoupled Weight Decay Regularization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:52.036255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:08:52.036255Z digest=sha256:1b0f48f6e8917fd28f69ddb4600b42b45288dbf83de66534bfeebcedf6d1e343

Observation 6f43300d-16b0-4564-9856-629564ac49a0 · outbound

This paper cites Self-supervised point cloud prediction using 3d spatio-temporal convolutional networks.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Self-supervised point cloud prediction using 3d spatio-temporal convolutional networks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:00.628992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:52.184670Z digest=sha256:3721357a3b0c906911d4e8096a92db4f7538de715c8d2f2a0fc6df292e681d3d

Observation f9f9068e-3c70-45fa-9962-bfbf5778fb65 · outbound

This paper cites Uniworld: Autonomous driving pre-training via world models.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Uniworld: Autonomous driving pre-training via world models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:00.442668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:52.361405Z digest=sha256:d46631e932571a7a1c14e60ef384ab9bea92f52bc5197fac020f269666af5be8

Observation 8ec25494-0bef-441c-8f0e-d8bd45c3547c · outbound

This paper cites Driveworld: 4d pre-trained scene understanding via world models for autonomous driving.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Driveworld: 4d pre-trained scene understanding via world models for autonomous driving

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:00.271810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:52.433606Z digest=sha256:d38280e25a7bfd96892286c86dd578642f2220b7c507ca753af58a13fa80b0f5

Observation f7c78b6c-5e7a-49f7-8015-1e195d58c065 · outbound

This paper cites Renderocc: Vision- centric 3d occupancy prediction with 2d rendering supervi- sion.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Renderocc: Vision- centric 3d occupancy prediction with 2d rendering supervi- sion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:09:00.141947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:52.515304Z digest=sha256:dad0583a12a81f92e489a0a877ea4216163892b5e9a1ed136ddfbb20a0e459f6

Observation 0adca95e-b7c8-49ec-951c-7a4af22a493d · outbound

This paper cites Scalable diffusion models with transformers.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Scalable diffusion models with transformers

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:59.939166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:52.559195Z digest=sha256:ee368a09c2c3f58d083375a6c9c4a9a57f5972e2c5b204e153557af54f6c77e4

Observation d14a39b2-6cb5-430b-8cf1-ec119baabb23 · outbound

This paper cites Gener- ating diverse high-fidelity images with vq-vae-2.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Gener- ating diverse high-fidelity images with vq-vae-2

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:59.761346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:52.612080Z digest=sha256:1a5e1a4251ce895da08023cab6a55239e3a78a2f4ad39519be2d5ab2b15b5928

Observation 33feea60-fcd3-4d54-a652-beae61e9d8bd · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting High-resolution image synthesis with latent diffusion models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:52.670827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:08:52.670827Z digest=sha256:d924622a617f0a74efa2506960934b1d181352d715cc16cec76fa4d842340e29

Observation 099d5316-3ec0-4356-8e00-1c683c10e7b3 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting High-resolution image synthesis with latent diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:59.640948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:52.708980Z digest=sha256:7dc39c1c4991ca983db2c212400e888c7809acd02144bc99aad29dd061570265

Observation 56daf1cb-2efe-4f22-b982-5708c729dd9b · outbound

This paper cites Pointr- cnn: 3d object proposal generation and detection from point cloud.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Pointr- cnn: 3d object proposal generation and detection from point cloud

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:59.534102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:52.840969Z digest=sha256:9f20a694ec5891578fc604e9a0961b59f07f27b7c3237ffdabcc5cb426a347fb

Observation db5da3c2-a0fa-45ce-833e-f30d6135bdf8 · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Scalability in perception for autonomous driving: Waymo open dataset

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:59.360467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:52.948765Z digest=sha256:d554ef42fcd8e7c722574530fd675d443637e941db3828fb8fc1427b3b07b7b1

Observation 57067620-3d17-407b-aae0-014d1c8ded3f · outbound

This paper cites Vidtok: A versatile and open-source video tokenizer.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Vidtok: A versatile and open-source video tokenizer

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:59.215295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:53.006460Z digest=sha256:253b3a3c148da0ab92a284d6ed745220ec0e9a8fc7b3a2cfaece63e2ea125137

Observation 7a74403d-91a4-4a80-bb60-32c6fa72906d · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Visual autoregressive modeling: Scalable image generation via next-scale prediction

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:59.073376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:53.053424Z digest=sha256:49cc301284630a3b9a6e2d15a3b146d103f4ed70218e4c403193e585765f0d5a

Observation 21565967-1c02-42c9-b6fc-a13e8aad039d · outbound

This paper cites Occ3d: A large-scale 3d occupancy prediction benchmark for au- tonomous driving.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Occ3d: A large-scale 3d occupancy prediction benchmark for au- tonomous driving

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:58.911047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:53.176555Z digest=sha256:d64438edbaa1b9c360108f883330ed84efc3d58d70f5d0df715a7e449d1cc6c6

Observation 099436e3-04e4-439d-aa0b-bce8c8bed48a · outbound

This paper cites Neural discrete representation learning.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Neural discrete representation learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:58.743975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:53.240006Z digest=sha256:12a58664ff6726d48fcf1987d37b33630e5561477e13faed29170123aabc298b

Observation 04e126c0-afaa-4328-ac32-2f3c4c46b1d5 · outbound

This paper cites Attention is all you need.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Attention is all you need

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:58.609298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:53.315105Z digest=sha256:418f0218c72ac22fd23f44d3c87c4d3ee48027c2fb96f765f5c6552d986e2b37

Observation bccb78db-3114-4aa3-abae-3579711fe1bc · outbound

This paper cites Omnitokenizer: A joint image-video tokenizer for visual generation.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Omnitokenizer: A joint image-video tokenizer for visual generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:58.452120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:53.386792Z digest=sha256:fdad381c8e3518bccbf073368ba708d058d5d4e4990c20652d894ed89a6cd0d6

Observation f459aa87-5b1d-4d6e-9b2d-05fce9a823b5 · outbound

This paper cites OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:53.539429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:08:53.539429Z digest=sha256:a767ab11e0f54576e11ce8fb0dd84d7b2badc9dc7286b75c291a250d3ac4cfca

Observation c7f01f57-31fe-47f7-9af9-073e815af061 · outbound

This paper cites Drivedreamer: Towards real-world- drive world models for autonomous driving.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Drivedreamer: Towards real-world- drive world models for autonomous driving

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:58.317739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:53.625583Z digest=sha256:28ba14305c6add1e566ead92676395f6d612a0d1f9a859fb8dd32771a9107a4c

Observation 110888b1-8ef1-485e-afcd-c52478dc0446 · outbound

This paper cites Occllama: An occupancy- language-action generative world model for autonomous driv- ing.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Occllama: An occupancy- language-action generative world model for autonomous driv- ing

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:58.179254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:53.699531Z digest=sha256:615af3bfb00fb01094f17affcd5f40508b1c0bf041c6bd9b3b01bac5e7e0ae66

Observation e8361453-a2cb-467a-aabb-d712ebffad80 · outbound

This paper cites Inverting the pose forecasting pipeline with spf2: Sequential pointcloud forecasting for sequential pose forecasting.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Inverting the pose forecasting pipeline with spf2: Sequential pointcloud forecasting for sequential pose forecasting

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:58.033467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:53.809724Z digest=sha256:79004c6d2b2a56392db928a0c4ef83ab3fdb95def228e5c7f8a9960814b5cb49

Observation a4e81d92-c85a-463c-8ed0-9110de2c7fff · outbound

This paper cites ivideogpt: Interactive videogpts are scalable world models.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting ivideogpt: Interactive videogpts are scalable world models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:57.800415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:53.915478Z digest=sha256:a76ec3128037b1f9ca75538d42968b0ccac54dc96ed2893d53d077d2556a5b5c

Observation 3c105349-79a2-4657-a197-2c3e912ed02a · outbound

This paper cites Occ-llm: Enhancing autonomous driving with occupancy-based large language models.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Occ-llm: Enhancing autonomous driving with occupancy-based large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:57.655793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:54.001361Z digest=sha256:b2de2b3c25d823e6e2a14a814a77c3a112e6cb4e162d2a726460dcbb416ae77b

Observation 064180a0-e7b7-4992-8186-da7c2a7cb1de · outbound

This paper cites Videogpt: Video generation using vq-vae and transform- ers.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Videogpt: Video generation using vq-vae and transform- ers

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:57.483539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:54.064863Z digest=sha256:6c4956aba26f884cdaf5898fd84a110bd7200dd0c5dfbbd4d3b67bfb73a23abc

Observation 694925b3-ab48-4c1f-956c-1f7ab4d310c1 · outbound

This paper cites Renderworld: World model with self-supervised 3d label.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Renderworld: World model with self-supervised 3d label

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:57.138598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:54.151092Z digest=sha256:8eccfe8fc24fb19b1edbb661162969158ba0aab1ed5cc5f482f9aa3817178e9f

Observation 2d2f8d2f-4caf-4d18-91d9-b125fae29c9a · outbound

This paper cites Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective su- pervision.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective su- pervision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:56.805746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:54.187795Z digest=sha256:97e1d1aa79a850401c18788dca7f8f03c8264ebec6b36b848fd951df4c5a77b1

Observation e977cd88-d0f1-4faf-af0c-27e452e22984 · outbound

This paper cites Driving in the occupancy world: Vision-centric 4d occupancy forecasting and planning via world models for autonomous driving.AAAI,.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Driving in the occupancy world: Vision-centric 4d occupancy forecasting and planning via world models for autonomous driving.AAAI,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:56.536961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:54.228110Z digest=sha256:7b0a9d8e9be64bf618f2248603226ba83b364d6bc0f6865af56ea17841fb6eaa

Observation 8b6a7c64-83c9-49c0-9ef6-8062a3b23f88 · outbound

This paper cites Magvit: Masked generative video transformer.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Magvit: Masked generative video transformer

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:56.185791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:54.289856Z digest=sha256:1570ddb3cd7a8fb96b0db92539cf26029f0ca1ca2329f87f582bb0a0fbf7bf1c

Observation 77cf7902-228e-4d73-a16a-b73cc266e453 · outbound

This paper cites An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:54.319827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:08:54.319827Z digest=sha256:8f48f62953acf8632df9503a39562eaa7d047b4f7d1b865a5a766ff62d7b5ab8

Observation 78fd15a8-b6d8-49f2-8ae5-b05f35aad011 · outbound

This paper cites Copilot4d: Learning unsupervised world models for autonomous driving via discrete diffusion.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Copilot4d: Learning unsupervised world models for autonomous driving via discrete diffusion

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:55.948100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:54.357990Z digest=sha256:4dd5cea67ccd782479786dee67ad05e9939760e14fe1973de2c88271f025d10f

Observation c46c6f49-6e33-4b5c-86c4-8484eba1f7f1 · outbound

This paper cites DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:54.392162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:08:54.392162Z digest=sha256:a8bf05992c9ac9e9aaf98dacb5952fe721f5a387f63d0c7455762e06cf592125

Observation 127bba40-acbf-4a51-8fbb-e1c9940887e5 · outbound

This paper cites Cv-vae: A compatible video vae for latent generative video models.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Cv-vae: A compatible video vae for latent generative video models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:55.684773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:54.443696Z digest=sha256:4e1a6587f2e580197be075e194673b1cd52ff99b3859037a74223e958e1773eb

Observation 53cbfdfe-6fb4-43e3-ab81-33f0a834df97 · outbound

This paper cites Occworld: Learning a 3d occupancy world model for autonomous driving.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Occworld: Learning a 3d occupancy world model for autonomous driving

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:55.372362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:54.484694Z digest=sha256:48dddd3e28d15bab54dff48fc1afe32867dc1f43a3a386ab882e8a81e8bd2455

Observation 16c40c42-07a2-4d24-9052-96b0e4c25d69 · outbound

This paper cites Hitvideo: Hierarchical tokenizers for enhancing text-to- video generation with autoregressive large language models.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Hitvideo: Hierarchical tokenizers for enhancing text-to- video generation with autoregressive large language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:55.194296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:54.528141Z digest=sha256:cd996298e4c8f010ac7568340dfaa5f549423d83ec68ce51f729d802be44ae78

Observation 3d658106-c657-4d36-a2cd-79356caeb59a · outbound

This paper cites Scaling the codebook size of vq-gan to 100,000 with a utilization rate of 99%.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Scaling the codebook size of vq-gan to 100,000 with a utilization rate of 99%

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:55.034751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:54.571426Z digest=sha256:18693b44953ebf66004aa64dddaa5df686bda5d4b5a7e062013b5fb4b83f31bf

Observation f72cfb80-31fc-4cc4-b8a8-f94862811795 · outbound

This paper cites Deformable detr: Deformable transformers for end-to-end object detection.

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting Deformable detr: Deformable transformers for end-to-end object detection

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:54.873947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:08:54.630490Z digest=sha256:a82c56d136e475a0c4a33a7596b8ce19b8ff4e97bcfa6edf1ff5d9d97f41f3cc

Pith citing papers

Observation 2006e092-da95-418d-9d11-334218175e27 · inbound

A Survey of World Models for Autonomous Driving cites this paper.

A Survey of World Models for Autonomous Driving $I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-10T18:31:52.653212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:31:52.653212Z digest=sha256:19e39151495da17594c4eb01c608d826870797f226a2e4f4b6785d0a53356d2d

Observation 6e1c4be7-4d62-43e7-9758-7e2c00ad752c · inbound

3D and 4D World Modeling: A Survey cites this paper.

3D and 4D World Modeling: A Survey $I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting

Reference 142

Resolution
unresolved
no resolver link, observed 2026-08-05T06:04:23.964735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:04:23.964735Z digest=sha256:8c7ad3def0a1afb9f64ecd0c22719e5973e72b5eaa99c93a8b0c1f1e01efba04

Observation f43675e7-76f5-40d7-a1c2-e3fcd4bb03a6 · inbound

SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model cites this paper.

SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model $I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:29:05.040869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T05:26:34.859975Z digest=sha256:05cf1cc09f582fa589c67f3d99663d16abba3c1dabfcaf504c1265df807a33af

Observation e607822e-c87d-47a0-b335-f9067297f13b · inbound

OWMDrive: Causality-Aware End-to-End Autonomous Driving via 4D Occupancy World Model cites this paper.

OWMDrive: Causality-Aware End-to-End Autonomous Driving via 4D Occupancy World Model $I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.612613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T05:59:20.898830Z digest=sha256:9d91121eb4f2a3fcd4b26cce17489c9894e0654d8b1ea58219a474aac54ccec2