Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:16:34.345618Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 6 inbound Pith citation observations for arXiv:2412.11198.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:16:34.345618Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:13:42.902690Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T21:06:13.797390Z
92 of 92 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f0c943d2-5c7b-4af8-9697-38696965940a · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Stereo vision and laser odom- etry for autonomous helicopters in gps-denied indoor envi- ronments
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79fcb6d7-3857-4729-8c18-155ef65d53a7 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control LIMT: Language-Informed Multi-Task Visual World Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2bacd06-2342-4a9b-9d1a-fd0330842872 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Uncertainty-based traffic accident anticipation with spatio-temporal relational learn- ing
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f23d4975-f0c0-4ee0-985f-580052c2e3e0 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0217454-faf5-41fb-b2d2-e04426a28ed7 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Align your latents: High-resolution video synthesis with la- tent diffusion models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2693969-09fd-4680-adab-7b9a39fcff14 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Marius Z¨ollner
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e531c33e-6bcb-4925-b59a-77063474a104 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Generating long videos of dynamic scenes
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aee29c92-7710-4524-88e1-966e84c13c27 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control nuscenes: A multi- modal dataset for autonomous driving
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85f2f802-717c-483b-bc90-7eeb4d159fe9 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control D$^2$-City: A Large-Scale Dashcam Video Dataset of Diverse Traffic Scenarios
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30438a74-a2f1-4cea-9b4a-a1235963287b · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Diffusion forcing: Next-token prediction meets full-sequence diffu- sion
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19b80778-3157-41ac-843d-6a5a877a932b · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Videocrafter1: Open diffusion models for high-quality video generation, 2023
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07125f32-80f0-4fcc-8a86-801405cd1789 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b802405-eea4-4dab-b622-4ba0539a3f11 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Seine: Short-to-long video diffu- sion model for generative transition and prediction
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d08aa199-e8c3-4b98-be7a-75e46d61ce8e · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control CAGE: Unsupervised Visual Composition and Animation for Controllable Video Generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dbc556a1-dcb3-47da-a210-f934efbc4c48 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Diffusion models beat gans on image synthesis, 2021
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d57f5800-a0d9-4f50-a31a-47072ce02c1e · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Vista: A generalizable driving world model with high fidelity and versatile controllability.Advances in Neural Information Processing Systems, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed0f69db-665f-483b-b7db-d425581e2d33 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ca4930-10cc-4699-8354-70ae585e082a · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Ego4d: Around the world in 3,000 hours of egocentric video
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaed36c6-6714-4100-810c-2e2c87a59148 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9833566a-808b-4bd1-808e-f9b3832588fb · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ac20f99a-15e1-42bc-995b-2212db177b20 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Dream to control: Learning behaviors by la- tent imagination
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b1e5a28c-32c9-47c9-a508-1219fb203d8a · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Hierarchical World Models as Visual Whole-Body Humanoid Controllers
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59608956-1ebe-4e6c-92a3-75466b10a5ca · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Temporal difference learning for model predictive control
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5a9a4bd3-dd76-421f-bb02-714fcd14764a · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Reasoning with language model is planning with world model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0ff444b0-d8d2-4735-9284-aaef416c4d73 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Large-scale actionless video pre-training via discrete diffusion for efficient policy learning, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 11b00641-8c68-4e8f-ba92-0f25a5466ec7 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control End-to-end learning of driving models with surround-view cameras and route planners
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7112a02b-1522-488f-b883-3aac346ad990 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 73e31547-eb73-4850-b0d8-e0f86d1c6325 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Imagen Video: High Definition Video Generation with Diffusion Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32fe1254-b292-4730-a47b-7fce52e7a0a8 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation caf35113-5329-4041-8711-ff2636902fa8 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Gaia-1: A generative world model for au- tonomous driving, 2023
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7976e2c1-5eb1-4f6e-9b48-e033cf95f423 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control LoRA: Low-Rank Adaptation of Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c61068d7-457c-4447-b1df-d8417595ccd2 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 478f36dc-6428-4fbc-8083-19fb125c2532 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Toward general-purpose robots via foundation mod- els: A survey and meta-analysis, 2023
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 275aa98d-c1ca-40be-af15-33ecba4d8c64 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Language Models, Agent Models, and World Models: The LAW for Machine Reasoning and Planning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b18fa19-9d5b-43e0-83d6-f325614c7e97 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control ADriver-I: A General World Model for Autonomous Driving
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a1d81bf-ee41-4b76-a794-9413ea73231f · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Elucidating the Design Space of Diffusion-Based Generative Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6c70fd4-e4ea-4e51-8b7c-d7c0c53b3390 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control YOLOv11: An Overview of the Key Architectural Enhancements
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 679ea084-840c-4aee-9d03-ff668f46e032 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Grounding human-to-vehicle advice for self-driving vehicles
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2354599e-12f3-4135-bc64-c2d3479d4128 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Drivegan: Towards a controllable high-quality neural simulation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1918ead4-40ed-4bb2-b153-551dcded7567 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control A path towards autonomous machine intelli- gence
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0747f4d1-d52a-4336-a492-a60e81d6f402 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 493923a2-6d7d-42d9-8f01-2d96a63a5e76 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Fit: Flexible vision trans- former for diffusion model, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 684dd68f-f55a-47f5-8d85-95fb7dee1e0f · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Struc- tured world models from human videos, 2023
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 17adda47-40aa-46b9-be64-5169017f8d44 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DINOv2: Learning Robust Visual Features without Supervision
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 329a02ad-df82-49a8-b8a0-690db601d10d · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4c3230d9-a7a7-4085-b26a-398a15278514 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Dutta Roy, Sugosh Nagavara Ravindra, Priya Goyal, and Matthijs Douze
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2d1fb9e7-fd5f-4803-9c3d-864686cf3de9 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Reasoning with large lan- guage models, a survey
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c348d7e8-32fb-4413-a947-a301fc99734f · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aed6bf9a-9447-4a3c-a717-905e8fb42de0 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Toward driving scene understanding: A dataset for learning driver behavior and causal reasoning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e79a5712-0d83-43ff-8a45-278b8934ed9f · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c1b0ef6c-bb61-42a2-b009-546f21d7e47b · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control High-resolution image syn- thesis with latent diffusion models, 2022
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbf2978f-858f-4d31-a3cb-6a31081f7d9b · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control LAION-5B: An open large-scale dataset for training next generation image-text models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20b04af5-cd00-40ac-b420-7d398a99b6ab · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Make-a-video: Text-to-video generation without text-video data, 2022
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27127ad6-60d7-44c7-b12e-e618d5be4876 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Fourier features let networks learn high frequency functions in low dimen- sional domains
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00c901e5-fe80-4d84-b48a-c7b542f2d00a · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Raft: Recurrent all-pairs field transforms for optical flow
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a53140e5-8058-4024-a7e7-e4ba50a7a0c9 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a48b5b46-1a54-420a-ae8c-8b2aacef3454 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Towards Accurate Generative Models of Video: A New Metric & Challenges
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf9e70ef-3750-4a86-8499-d6debda17c84 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control GeoCalib: Single-image Cali- bration with Geometric Optimization
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4880cfa8-55d8-4724-a32f-079a73ef61bf · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Channappayya, and Swarup S
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 67d38f81-1555-4636-8539-526f298f566f · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Mcvd: Masked conditional video diffusion for predic- tion, generation, and interpolation, 2022
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 36e8d615-636d-41e5-901a-a1c0dbcbc765 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f2105f8-9434-4efa-8cbf-f44c4fcb92b8 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Videocomposer: Compositional video synthesis with motion controllability
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f33f9adc-e1a7-4b3b-934e-64ccb307eff8 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Pseudo- lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f83f2773-7020-4dc1-ad3f-c3b402b3fa66 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfc508c9-9db5-42da-8e58-274abbee091b · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2332a368-0687-437c-a674-3262e5c98836 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Pre-training contextualized world models with in-the-wild videos for reinforcement learning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 65c3c2cd-1dbf-4b97-94fb-5428114a0bf2 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Pandora: Towards General World Model with Natural Language Actions and Video States
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a78fe71-635e-461f-8730-054a20671f52 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Progressive Autoregressive Video Diffusion Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d42c684-6d69-41f9-b521-548132ccc25e · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control End- to-end learning of driving models from large-scale video datasets
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a111aac7-e4e0-403b-9ffe-e55922e446ac · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Videogpt: Video generation using vq-vae and trans- formers, 2021
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c6cebde-293c-4dda-ba97-e344c47b0158 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Generalized Predictive Model for Autonomous Driving
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 411a0b95-e99a-4c74-94ca-a56661d28e62 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Depth anything: Unleashing the power of large-scale unlabeled data
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 931b43e4-f357-451e-b850-3b1b9b6680f0 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Depth Anything V2
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfc1a449-bff5-4400-b8ee-792964ca4c90 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Learning interactive real-world simulators
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dc4080eb-e9cd-427b-9ec4-7642e08da771 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Effec- tive whole-body pose estimation with two-stages distillation
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6a08682e-23b1-4848-bcf6-63cb486c9467 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Visual point cloud forecasting enables scalable autonomous driving
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5723f6bc-9fb7-4f25-930a-a69cfcad6f2d · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control When, Where, and What? A New Dataset for Anomaly Detection in Driving Videos
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e924538b-f16b-4c7c-9aa8-3546111a436e · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Adding conditional control to text-to-image diffusion models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7986282e-c3f4-44e4-ba87-18961c7c0a15 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Copilot4d: Learning unsupervised world models for autonomous driving via discrete diffusion,
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 57d2e743-2eb4-4cff-bbad-77f4a5a42864 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control I2vgen-xl: High-quality image-to-video synthe- sis via cascaded diffusion models, 2023
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a8b80298-4f0e-4f77-9a8e-3000b71d3179 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb6e488-ea6f-44ef-bea7-0d1b4f35d57f · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Can LLM Graph Reasoning Generalize beyond Pattern Memorization?
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9d0647d-6d8c-43fe-aa1f-9b74889d98b2 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bb1fa72-7ae8-404e-916b-bd23865f7562 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97f7abb0-83a3-48d6-bf22-62d330648a67 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control On the continuity of rotation representations in neural networks
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747f885b-9498-4c5e-b4c3-d3e0fc0699aa · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Is sora a world simulator? a comprehensive survey on general world models and beyond, 2024
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8d07ec75-c5e7-4d2b-9bef-8cc8a9a0b194 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Pseudo-labeling Depth
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a9d5013e-3dd9-4d7b-a38b-1b7e91e5efda · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Due to the increased size of our network, we in- corporate activation checkpointing and optimizer sharding to mitigate memory constraints, utilizing the DeepSpeed li- brary [50]
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 629adf3f-3746-4c34-8de7-6350a4bfadfb · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control To achieve fine- grained, high-quality control, we employ a two-stage train- ing regime, detailed as follows: 8.1.1
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9d2cc4b5-3be1-482b-8a1c-22cbd71ad899 · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Unresolved cited work
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a63f93af-8fd5-4512-a3a9-74f8bd13825a · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Depth generation quality comparison
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 04e3547b-2e58-49ad-b45e-cfdceeed581d · outbound
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control 11 to 14 show qualitative examples of our generations, our controls, long generation and multimodal outputs
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ac5b25be-eb94-4d8b-a35f-de8e19f1f6cc · inbound
GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e64e1309-a665-4a08-8dff-aa00159e5cea · inbound
GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da65dc62-cd75-4d42-b702-276a4d3cc2e8 · inbound
Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d9971b4-40b9-40d2-a10e-e34611b53554 · inbound
How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 645d106c-2757-4d39-905c-4faf702b1677 · inbound
Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
Reference 211
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e209b637-182d-4001-a185-e1286e9b4871 · inbound
Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.