Pith. sign in

Paper Citation Record · LEDGER

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

As of 15 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 6 inbound Pith citation observations for arXiv:2412.11198.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11198 v1

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:16:34.345618Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:13:42.902690Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T21:06:13.797390Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0c943d2-5c7b-4af8-9697-38696965940a · outbound

This paper cites Stereo vision and laser odom- etry for autonomous helicopters in gps-denied indoor envi- ronments.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Stereo vision and laser odom- etry for autonomous helicopters in gps-denied indoor envi- ronments

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.963013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.963013Z digest=sha256:780a5f508f41b0acc7485110d861cb2379acdaae9ac85042f22b047d965c4d0d

Observation 79fcb6d7-3857-4729-8c18-155ef65d53a7 · outbound

This paper cites LIMT: Language-Informed Multi-Task Visual World Models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control LIMT: Language-Informed Multi-Task Visual World Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.967601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.967601Z digest=sha256:57f5f56f10c106384334aafc5b3377268e47a71e3f9b7c140027606b3d567594

Observation b2bacd06-2342-4a9b-9d1a-fd0330842872 · outbound

This paper cites Uncertainty-based traffic accident anticipation with spatio-temporal relational learn- ing.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Uncertainty-based traffic accident anticipation with spatio-temporal relational learn- ing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.972862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.972862Z digest=sha256:3ddceb4900013de0120eb38a5f671d74e1cbec0115d12e0234973350c9480eb6

Observation f23d4975-f0c0-4ee0-985f-580052c2e3e0 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.977882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.977882Z digest=sha256:d7fdfca60a10b1f1ae3faa87272096bdfb14bc40ed7b19ac407d93860e7fb557

Observation f0217454-faf5-41fb-b2d2-e04426a28ed7 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.982431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.982431Z digest=sha256:0381d2ec94548345672ad09ebc3b40b5f5c361cd34cc9e5c45c798f03c8abf33

Observation f2693969-09fd-4680-adab-7b9a39fcff14 · outbound

This paper cites Marius Z¨ollner.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Marius Z¨ollner

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.986497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.986497Z digest=sha256:72382c85a98fd07f6c63751fb58526b3ec61d43b3f3b7183fc851d69f906897e

Observation e531c33e-6bcb-4925-b59a-77063474a104 · outbound

This paper cites Generating long videos of dynamic scenes.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Generating long videos of dynamic scenes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.990515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.990515Z digest=sha256:146394beba6e4b79608099f380fbc66f98188dafad4b12304a976a03112b9813

Observation aee29c92-7710-4524-88e1-966e84c13c27 · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control nuscenes: A multi- modal dataset for autonomous driving

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.994832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.994832Z digest=sha256:251937aa89a97b6023ca1468576823d98e9dbc170666813a58d5b0278a6b2791

Observation 85f2f802-717c-483b-bc90-7eeb4d159fe9 · outbound

This paper cites D$^2$-City: A Large-Scale Dashcam Video Dataset of Diverse Traffic Scenarios.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control D$^2$-City: A Large-Scale Dashcam Video Dataset of Diverse Traffic Scenarios

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.998408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.998408Z digest=sha256:edaee2512506340900fc9863732e5dc34a2981a7ba824ab5772749ece7115b05

Observation 30438a74-a2f1-4cea-9b4a-a1235963287b · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffu- sion.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Diffusion forcing: Next-token prediction meets full-sequence diffu- sion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.002308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.002308Z digest=sha256:414816526ded7d243a9b80678ebf0b988df621dc52e4b9db7e472f7a195c94c4

Observation 19b80778-3157-41ac-843d-6a5a877a932b · outbound

This paper cites Videocrafter1: Open diffusion models for high-quality video generation, 2023.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Videocrafter1: Open diffusion models for high-quality video generation, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.005983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.005983Z digest=sha256:ace2b1d3db78cfbbefc9225343d98037f30be3b448ebabb49b948b2f850e0cab

Observation 07125f32-80f0-4fcc-8a86-801405cd1789 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.009937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.009937Z digest=sha256:9b9fa8c50f972923df8724f668902a219b1cfe69f67734d2ab099a6585262e8f

Observation 0b802405-eea4-4dab-b622-4ba0539a3f11 · outbound

This paper cites Seine: Short-to-long video diffu- sion model for generative transition and prediction.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Seine: Short-to-long video diffu- sion model for generative transition and prediction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.014775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.014775Z digest=sha256:a41279f490288ed6495c2c3ed056fc2b8baecc57a757033b2373afef8b7b6db7

Observation d08aa199-e8c3-4b98-be7a-75e46d61ce8e · outbound

This paper cites CAGE: Unsupervised Visual Composition and Animation for Controllable Video Generation.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control CAGE: Unsupervised Visual Composition and Animation for Controllable Video Generation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:16:34.794898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.018680Z digest=sha256:36a492a40ea0bfe38e754bca441d556c4bccae2319b704f74e37afd80c2e86d0

Observation dbc556a1-dcb3-47da-a210-f934efbc4c48 · outbound

This paper cites Diffusion models beat gans on image synthesis, 2021.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Diffusion models beat gans on image synthesis, 2021

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.023037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.023037Z digest=sha256:ec8f5d83d8c4e4f6b0c079d3a6697b5105bbe02efa736846077fa6213854cfcd

Observation d57f5800-a0d9-4f50-a31a-47072ce02c1e · outbound

This paper cites Vista: A generalizable driving world model with high fidelity and versatile controllability.Advances in Neural Information Processing Systems, 2024.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Vista: A generalizable driving world model with high fidelity and versatile controllability.Advances in Neural Information Processing Systems, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.026799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.026799Z digest=sha256:dc6a81c60de77ed53b9b6733827ab0abeb9e8465d17ca51aada123a31f38b734

Observation ed0f69db-665f-483b-b7db-d425581e2d33 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.031763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.031763Z digest=sha256:0f1dc07419e99f0624e73cc9d5ba14b041ab569f0c39ae52bffe7fd897b8304f

Observation 44ca4930-10cc-4699-8354-70ae585e082a · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Ego4d: Around the world in 3,000 hours of egocentric video

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.036409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.036409Z digest=sha256:064162016844efcd5939f27f57893fcae3fcbe452b99d729d3cdd95df9adf576

Observation aaed36c6-6714-4100-810c-2e2c87a59148 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.475087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.044063Z digest=sha256:45a36a2918b6d6e98a564e31ef9be7f66c290d0f5468932359a7ed3bdcacc15c

Observation 9833566a-808b-4bd1-808e-f9b3832588fb · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.463197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.048122Z digest=sha256:e5761c30fe2fd03360a8f2f77d9503756c8ab251aeb17eb1a38a7e5a9636169b

Observation ac20f99a-15e1-42bc-995b-2212db177b20 · outbound

This paper cites Dream to control: Learning behaviors by la- tent imagination.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Dream to control: Learning behaviors by la- tent imagination

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.451926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.051812Z digest=sha256:2e49b8f1a33dd00618f8e1fe89f180d32342e65c49c801192ff813b6227e886a

Observation b1e5a28c-32c9-47c9-a508-1219fb203d8a · outbound

This paper cites Hierarchical World Models as Visual Whole-Body Humanoid Controllers.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Hierarchical World Models as Visual Whole-Body Humanoid Controllers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.056278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.056278Z digest=sha256:46b803f289e752e8cb0f3d53a49c78c6e0a3f3e22d53c2881264b4017cade4eb

Observation 59608956-1ebe-4e6c-92a3-75466b10a5ca · outbound

This paper cites Temporal difference learning for model predictive control.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Temporal difference learning for model predictive control

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.439374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.061360Z digest=sha256:cfd71fd3a7b417edbdd6f369296577bc8aeb4ffee8cb036ff61a9c755ce509be

Observation 5a9a4bd3-dd76-421f-bb02-714fcd14764a · outbound

This paper cites Reasoning with language model is planning with world model.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Reasoning with language model is planning with world model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.426209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.065233Z digest=sha256:7b6f414cc3c8dc6d12651feba9fa3bc4701bc9cf4883248eff01d854d7dabe37

Observation 0ff444b0-d8d2-4735-9284-aaef416c4d73 · outbound

This paper cites Large-scale actionless video pre-training via discrete diffusion for efficient policy learning, 2024.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Large-scale actionless video pre-training via discrete diffusion for efficient policy learning, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.412200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.069400Z digest=sha256:d2716d4219f0beb38aa717f94af5a51a773cbb6eb640e23d2730db4b3b711a64

Observation 11b00641-8c68-4e8f-ba92-0f25a5466ec7 · outbound

This paper cites End-to-end learning of driving models with surround-view cameras and route planners.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control End-to-end learning of driving models with surround-view cameras and route planners

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.399476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.073357Z digest=sha256:8e5818f13578f8b2f92d384edfa9e77df55521d8321e79544236089a92a856f8

Observation 7112a02b-1522-488f-b883-3aac346ad990 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.384026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.076953Z digest=sha256:387420736a9b92099db976efb1706bb11814ba8263cce520f4d0e220ccccf62b

Observation 73e31547-eb73-4850-b0d8-e0f86d1c6325 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Imagen Video: High Definition Video Generation with Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.080962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.080962Z digest=sha256:f9324e8c1d27f14aeefc6a5d2e2b257028b12d390f7a02792905bec306b73cd0

Observation 32fe1254-b292-4730-a47b-7fce52e7a0a8 · outbound

This paper cites an unresolved cited work.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:16:35.369461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.085256Z digest=sha256:3953418dfe38eeb591b4db2fc97073aea44fdfda9928b97ce1c3395c85e54d41

Observation caf35113-5329-4041-8711-ff2636902fa8 · outbound

This paper cites Gaia-1: A generative world model for au- tonomous driving, 2023.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Gaia-1: A generative world model for au- tonomous driving, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.357703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.088769Z digest=sha256:a09a74557937332ca9f881ea51881368648224819ccd57917df450992bf8d59c

Observation 7976e2c1-5eb1-4f6e-9b48-e033cf95f423 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control LoRA: Low-Rank Adaptation of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.092126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.092126Z digest=sha256:c33220b6686847ed7e270430914c5e9218128557c5fcd1e99e7bd5220be0400d

Observation c61068d7-457c-4447-b1df-d8417595ccd2 · outbound

This paper cites DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.095535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.095535Z digest=sha256:d3a6e79d032e22d42880c731c3ae082d6f5e9577ae5db08012ed0977d76eed2d

Observation 478f36dc-6428-4fbc-8083-19fb125c2532 · outbound

This paper cites Toward general-purpose robots via foundation mod- els: A survey and meta-analysis, 2023.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Toward general-purpose robots via foundation mod- els: A survey and meta-analysis, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.343294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.099513Z digest=sha256:3526df4e787fd4bb27d7be7fd55a08ea57265bca66bb121ba2a31a6756748289

Observation 275aa98d-c1ca-40be-af15-33ecba4d8c64 · outbound

This paper cites Language Models, Agent Models, and World Models: The LAW for Machine Reasoning and Planning.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Language Models, Agent Models, and World Models: The LAW for Machine Reasoning and Planning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.103629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.103629Z digest=sha256:88a553027cda9612bb3fd0fdbbad9ed68ceccfe31e14e5f94a74ac8e94de1bfa

Observation 6b18fa19-9d5b-43e0-83d6-f325614c7e97 · outbound

This paper cites ADriver-I: A General World Model for Autonomous Driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control ADriver-I: A General World Model for Autonomous Driving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.107971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.107971Z digest=sha256:e273b50098e3d92bb8a77d05ea00ccae9ef173abb40adbe6ec7947de6b81bece

Observation 1a1d81bf-ee41-4b76-a794-9413ea73231f · outbound

This paper cites Elucidating the Design Space of Diffusion-Based Generative Models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Elucidating the Design Space of Diffusion-Based Generative Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.112333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.112333Z digest=sha256:7138d5e3d15254a20001318ac0030de3bec770024be5eb2480a63fd1b6e257d2

Observation e6c70fd4-e4ea-4e51-8b7c-d7c0c53b3390 · outbound

This paper cites YOLOv11: An Overview of the Key Architectural Enhancements.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control YOLOv11: An Overview of the Key Architectural Enhancements

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.116587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.116587Z digest=sha256:701f1fd6293a506fe1101fbda1607353d170735cb1fd09da05a389d7c9f5fe72

Observation 679ea084-840c-4aee-9d03-ff668f46e032 · outbound

This paper cites Grounding human-to-vehicle advice for self-driving vehicles.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Grounding human-to-vehicle advice for self-driving vehicles

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.330645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.120449Z digest=sha256:42cb3f9fd6620a6774c0a4bae9c2c9d5cf206e61c2def9fe43f9f4901fa87bf3

Observation 2354599e-12f3-4135-bc64-c2d3479d4128 · outbound

This paper cites Drivegan: Towards a controllable high-quality neural simulation.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Drivegan: Towards a controllable high-quality neural simulation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.318280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.124548Z digest=sha256:3bc489cc57780531387aa7dc89056b92a8d8f0b19642198e154c4f65212fc2af

Observation 1918ead4-40ed-4bb2-b153-551dcded7567 · outbound

This paper cites A path towards autonomous machine intelli- gence.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control A path towards autonomous machine intelli- gence

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.305571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.128703Z digest=sha256:738e67c45145d360f4edd6b1d783fd9dd3d7ecde1309661d99176e871ff22bf8

Observation 0747f4d1-d52a-4336-a492-a60e81d6f402 · outbound

This paper cites WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.132700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.132700Z digest=sha256:e8f296edfec6cd522b8b434854519cc82362bc1f96cc72dc592a5861ff1817bd

Observation 493923a2-6d7d-42d9-8f01-2d96a63a5e76 · outbound

This paper cites Fit: Flexible vision trans- former for diffusion model, 2024.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Fit: Flexible vision trans- former for diffusion model, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.294159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.137416Z digest=sha256:1670a94db2ca03ea667528ff89d67d2022412d1ec54b44130ec757f63272f752

Observation 684dd68f-f55a-47f5-8d85-95fb7dee1e0f · outbound

This paper cites Struc- tured world models from human videos, 2023.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Struc- tured world models from human videos, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.280112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.141266Z digest=sha256:db08194b5f42581312fccbe5b82318dfcce617736421fb945e5d54c1294856c6

Observation 17adda47-40aa-46b9-be64-5169017f8d44 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DINOv2: Learning Robust Visual Features without Supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.144911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.144911Z digest=sha256:18f2804bc876ccb2d5acefe21933eaa98a0ede472fcce9f75d46c1fc35d8a6ee

Observation 329a02ad-df82-49a8-b8a0-690db601d10d · outbound

This paper cites an unresolved cited work.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:16:35.267574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.149083Z digest=sha256:8b5931eb0323c7ee9d88c118c02b49e1102a16ed9ef1c07492434e83a8f965d3

Observation 4c3230d9-a7a7-4085-b26a-398a15278514 · outbound

This paper cites Dutta Roy, Sugosh Nagavara Ravindra, Priya Goyal, and Matthijs Douze.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Dutta Roy, Sugosh Nagavara Ravindra, Priya Goyal, and Matthijs Douze

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.254388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.153447Z digest=sha256:55eaeb4e3b6bba67e14a2cd7f7899c9753970bcd2a74778bacb4f5c76e9bcbdf

Observation 2d1fb9e7-fd5f-4803-9c3d-864686cf3de9 · outbound

This paper cites Reasoning with large lan- guage models, a survey.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Reasoning with large lan- guage models, a survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.158058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.158058Z digest=sha256:f4d88ce001b01cedf3802d74bba4aaf5f776c20b506d5604321918a1884c9ac7

Observation c348d7e8-32fb-4413-a947-a301fc99734f · outbound

This paper cites Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.241195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.162913Z digest=sha256:6f6f6b5432cf1d184b0d7d737cec62c2164c0eaf6277dacce0f1bd1a81e37d91

Observation aed6bf9a-9447-4a3c-a717-905e8fb42de0 · outbound

This paper cites Toward driving scene understanding: A dataset for learning driver behavior and causal reasoning.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Toward driving scene understanding: A dataset for learning driver behavior and causal reasoning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.227930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.166894Z digest=sha256:2e8aace900f69a9556c9fbaa1c62271f0833be6c045addddff05b1b706acb168

Observation e79a5712-0d83-43ff-8a45-278b8934ed9f · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.215751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.170813Z digest=sha256:0a66561e13da3a4fea7043e8431052bc234c4d8d1c15dab2e606ea5319138365

Observation c1b0ef6c-bb61-42a2-b009-546f21d7e47b · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2022.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control High-resolution image syn- thesis with latent diffusion models, 2022

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.174473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.174473Z digest=sha256:860664140aa22c70c1f0934bd5a10dbea7d5987cdfad25bf3c08e3c80e0f9436

Observation cbf2978f-858f-4d31-a3cb-6a31081f7d9b · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.177833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.177833Z digest=sha256:4871603daf643f963f8ccde47d264b62bb93a7ed1bba4b4fc555377f28ffb7b5

Observation 20b04af5-cd00-40ac-b420-7d398a99b6ab · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data, 2022.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Make-a-video: Text-to-video generation without text-video data, 2022

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.181414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.181414Z digest=sha256:6896e2f0338ce0246fcf52963c9918e464cf7ad04001d654dcfe6889d7f4e2af

Observation 27127ad6-60d7-44c7-b12e-e618d5be4876 · outbound

This paper cites Fourier features let networks learn high frequency functions in low dimen- sional domains.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Fourier features let networks learn high frequency functions in low dimen- sional domains

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.185082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.185082Z digest=sha256:5b74526200767b03ced55ccfeb59e9f52f94452e94aba7bd2a0454b0ff60daa9

Observation 00c901e5-fe80-4d84-b48a-c7b542f2d00a · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Raft: Recurrent all-pairs field transforms for optical flow

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.176086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.189407Z digest=sha256:68d4c10c3cf8b6607beee879660cd5674a4ca0299dc83080b276e2be7a09b7ce

Observation a53140e5-8058-4024-a7e7-e4ba50a7a0c9 · outbound

This paper cites Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.163278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.193274Z digest=sha256:8ffac65aa78cff73ec9c1636fd29aacbb9cf47c2ae645ce935aea4664520bfe1

Observation a48b5b46-1a54-420a-ae8c-8b2aacef3454 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.197022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.197022Z digest=sha256:e1b97d20d596a34a5f21857cfca3d25b231d9f89329f58ff42d28eb32c018cd0

Observation cf9e70ef-3750-4a86-8499-d6debda17c84 · outbound

This paper cites GeoCalib: Single-image Cali- bration with Geometric Optimization.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control GeoCalib: Single-image Cali- bration with Geometric Optimization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.150407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.201145Z digest=sha256:061119d37dc8d20bc5ccf35a4e04cb68e3ea090056b386b0382970d0f65a7333

Observation 4880cfa8-55d8-4724-a32f-079a73ef61bf · outbound

This paper cites Channappayya, and Swarup S.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Channappayya, and Swarup S

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.137510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.205641Z digest=sha256:35bffb8e24add1770e044a85d863973d35876402fd1d1a8652e92026c3bb1498

Observation 67d38f81-1555-4636-8539-526f298f566f · outbound

This paper cites Mcvd: Masked conditional video diffusion for predic- tion, generation, and interpolation, 2022.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Mcvd: Masked conditional video diffusion for predic- tion, generation, and interpolation, 2022

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.124740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.209464Z digest=sha256:6457cc42843836ea285c8cded3a50a5a5cc53cb6ac683310edfae8b9803aafad

Observation 36e8d615-636d-41e5-901a-a1c0dbcbc765 · outbound

This paper cites DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.213408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.213408Z digest=sha256:95f9b11a06b5d77c563e0bb2bcaf9b620d5a4e057d7c7ee0aa0db668d0ab3d3b

Observation 2f2105f8-9434-4efa-8cbf-f44c4fcb92b8 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Videocomposer: Compositional video synthesis with motion controllability

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.110970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.218155Z digest=sha256:0dc6274d0dc6df0928868c17bd64d115291d959a86a70f4acc95f0eac5feba09

Observation f33f9adc-e1a7-4b3b-934e-64ccb307eff8 · outbound

This paper cites Pseudo- lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Pseudo- lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.096364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.221857Z digest=sha256:ca792b57f1389bc40de1c1b2ab717407651f8909aba92c0bed890c8780ccae49

Observation f83f2773-7020-4dc1-ad3f-c3b402b3fa66 · outbound

This paper cites DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.225388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.225388Z digest=sha256:8a192775a0378ad0ad0a485cc7c71bd0be9e2fa25161710c5e4909645679690d

Observation dfc508c9-9db5-42da-8e58-274abbee091b · outbound

This paper cites Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.082482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.229463Z digest=sha256:491e0a944644f0f30b269818599c3baa5d8c98cfffedef84cec0ff1b31aba6ad

Observation 2332a368-0687-437c-a674-3262e5c98836 · outbound

This paper cites Pre-training contextualized world models with in-the-wild videos for reinforcement learning.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Pre-training contextualized world models with in-the-wild videos for reinforcement learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.070390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.233575Z digest=sha256:b65076208f14bc3fde61e06778b9cd018352846bd07ec404a55d269f03bdb366

Observation 65c3c2cd-1dbf-4b97-94fb-5428114a0bf2 · outbound

This paper cites Pandora: Towards General World Model with Natural Language Actions and Video States.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Pandora: Towards General World Model with Natural Language Actions and Video States

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.237524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.237524Z digest=sha256:ae9e6587f856cc7a409eab6993df48e850501b7ce7422fe55b45ccca5f1bc78a

Observation 2a78fe71-635e-461f-8730-054a20671f52 · outbound

This paper cites Progressive Autoregressive Video Diffusion Models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Progressive Autoregressive Video Diffusion Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.242395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.242395Z digest=sha256:f6811f244a4a9da40221298614399a606a33ca5414a5efc95678d2e1a59663f5

Observation 8d42c684-6d69-41f9-b521-548132ccc25e · outbound

This paper cites End- to-end learning of driving models from large-scale video datasets.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control End- to-end learning of driving models from large-scale video datasets

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.057052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.246288Z digest=sha256:920039b66a4cad19df6a84b658da1be6ecff8a82d65a2d8805d56b19fe804616

Observation a111aac7-e4e0-403b-9ffe-e55922e446ac · outbound

This paper cites Videogpt: Video generation using vq-vae and trans- formers, 2021.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Videogpt: Video generation using vq-vae and trans- formers, 2021

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.250365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.250365Z digest=sha256:8f78274b5d7caf7fff6a4156098e5c86d00066c0be267be58958e2eed21eb4ac

Observation 2c6cebde-293c-4dda-ba97-e344c47b0158 · outbound

This paper cites Generalized Predictive Model for Autonomous Driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Generalized Predictive Model for Autonomous Driving

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.035670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.254190Z digest=sha256:86e72f2429dfbbe4fa4081b91b2b7b0dea35ffcd58df6566d18634e70d2fdef0

Observation 411a0b95-e99a-4c74-94ca-a56661d28e62 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Depth anything: Unleashing the power of large-scale unlabeled data

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.022034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.257651Z digest=sha256:53a0f2a899f83a09b7c7300b8e5f4143947b5f738c7b87b03d3a89e625d4283a

Observation 931b43e4-f357-451e-b850-3b1b9b6680f0 · outbound

This paper cites Depth Anything V2.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Depth Anything V2

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.261320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.261320Z digest=sha256:9b1058ca96cd75fcda145753643e753a0dfe5f447e53bef161433071dba26fd5

Observation cfc1a449-bff5-4400-b8ee-792964ca4c90 · outbound

This paper cites Learning interactive real-world simulators.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Learning interactive real-world simulators

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.008514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.265441Z digest=sha256:06635fa67bd38689f257dae7b2febd57639e66b7aa387d88196428e74f22599c

Observation dc4080eb-e9cd-427b-9ec4-7642e08da771 · outbound

This paper cites Effec- tive whole-body pose estimation with two-stages distillation.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Effec- tive whole-body pose estimation with two-stages distillation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.995986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.268915Z digest=sha256:1af49c03dfe364f40fd152b96834a92d458e9b45bb2667346a3ca3e4540e6436

Observation 6a08682e-23b1-4848-bcf6-63cb486c9467 · outbound

This paper cites Visual point cloud forecasting enables scalable autonomous driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Visual point cloud forecasting enables scalable autonomous driving

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.983668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.272580Z digest=sha256:13d77df7fcb92f24cc7b7fcc6ed0a5fd8a33c19b634b86c981bae2b8b0ba58e6

Observation 5723f6bc-9fb7-4f25-930a-a69cfcad6f2d · outbound

This paper cites When, Where, and What? A New Dataset for Anomaly Detection in Driving Videos.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control When, Where, and What? A New Dataset for Anomaly Detection in Driving Videos

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.277111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.277111Z digest=sha256:81525f077aa6381216bce8ddf7ee9d7c3943dc19800b6ee687fac4f40407c2db

Observation e924538b-f16b-4c7c-9aa8-3546111a436e · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Adding conditional control to text-to-image diffusion models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.281670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.281670Z digest=sha256:a3e2b350525c9857ab8c623ce899a937699a12a1a6bf2c6c7ee6b05277467274

Observation 7986282e-c3f4-44e4-ba87-18961c7c0a15 · outbound

This paper cites Copilot4d: Learning unsupervised world models for autonomous driving via discrete diffusion,.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Copilot4d: Learning unsupervised world models for autonomous driving via discrete diffusion,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.963543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.285298Z digest=sha256:63454f536f8f51e5403de91b26151f5b6d4a762454f67bcf492b075aba68323d

Observation 57d2e743-2eb4-4cff-bbad-77f4a5a42864 · outbound

This paper cites I2vgen-xl: High-quality image-to-video synthe- sis via cascaded diffusion models, 2023.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control I2vgen-xl: High-quality image-to-video synthe- sis via cascaded diffusion models, 2023

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.949449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.289172Z digest=sha256:9387868caf560d7224ff2090886c30047e5b83401f55ec4ef894d27b0e49d0ec

Observation a8b80298-4f0e-4f77-9a8e-3000b71d3179 · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.298271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.298271Z digest=sha256:b0f0fdb293ac1b23847b9e60359ac3910c8fbbdaa12a897e16d053c6acf61a92

Observation 0eb6e488-ea6f-44ef-bea7-0d1b4f35d57f · outbound

This paper cites Can LLM Graph Reasoning Generalize beyond Pattern Memorization?.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Can LLM Graph Reasoning Generalize beyond Pattern Memorization?

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.303797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.303797Z digest=sha256:f699f5d65c0bc9e45f4bcee800610f85c79c2b64e7f54af3ae3286bc8da03c2d

Observation d9d0647d-6d8c-43fe-aa1f-9b74889d98b2 · outbound

This paper cites DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.308100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.308100Z digest=sha256:d2d94fc44a1127d193d805dfc1c6a1f0c1e1c3cfc34ae8bcadc84aa440c76700

Observation 2bb1fa72-7ae8-404e-916b-bd23865f7562 · outbound

This paper cites OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.312572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.312572Z digest=sha256:4443b6596fe2648fea86eb676398e57d71bed662e7e37f6d782cd96a8d9a4e01

Observation 97f7abb0-83a3-48d6-bf22-62d330648a67 · outbound

This paper cites On the continuity of rotation representations in neural networks.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control On the continuity of rotation representations in neural networks

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.317611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.317611Z digest=sha256:909ac5e9a047f621a1e4105a4ac8990b79991c91edc2e4a0d6dfcd1a1c0ea2ba

Observation 747f885b-9498-4c5e-b4c3-d3e0fc0699aa · outbound

This paper cites Is sora a world simulator? a comprehensive survey on general world models and beyond, 2024.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Is sora a world simulator? a comprehensive survey on general world models and beyond, 2024

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.927880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.321791Z digest=sha256:baca0159679e560f4dca50c70cfe2e670902790758e44a1c85f86a47b8437aed

Observation 8d07ec75-c5e7-4d2b-9bef-8cc8a9a0b194 · outbound

This paper cites Pseudo-labeling Depth.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Pseudo-labeling Depth

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.912699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.325405Z digest=sha256:edf1ad096dd06543d63f258939635536cf5801ed5dec1168079825a4e2607e79

Observation a9d5013e-3dd9-4d7b-a38b-1b7e91e5efda · outbound

This paper cites Due to the increased size of our network, we in- corporate activation checkpointing and optimizer sharding to mitigate memory constraints, utilizing the DeepSpeed li- brary [50].

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Due to the increased size of our network, we in- corporate activation checkpointing and optimizer sharding to mitigate memory constraints, utilizing the DeepSpeed li- brary [50]

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.898427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.329547Z digest=sha256:973584300ab31fb3e36dc88ba7923f31e3c63dcdc32372e2e656ac2280abb768

Observation 629adf3f-3746-4c34-8de7-6350a4bfadfb · outbound

This paper cites To achieve fine- grained, high-quality control, we employ a two-stage train- ing regime, detailed as follows: 8.1.1.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control To achieve fine- grained, high-quality control, we employ a two-stage train- ing regime, detailed as follows: 8.1.1

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.884683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.333665Z digest=sha256:c132118bc848b16e6051af1621b33bfb3b8db6880aa6e8e3b4723eb451076cdb

Observation 9d2cc4b5-3be1-482b-8a1c-22cbd71ad899 · outbound

This paper cites an unresolved cited work.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:16:34.870716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.337488Z digest=sha256:dfe86aa227709b0269515f78ea2d75b23af44dbc2da82a307ec5eaa8ed5b804a

Observation a63f93af-8fd5-4512-a3a9-74f8bd13825a · outbound

This paper cites Depth generation quality comparison.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Depth generation quality comparison

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.855888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.341458Z digest=sha256:911bef6eb9bf7d2b7d49b56ded9763b18aa01420dc8773f387198d5e3f000ec4

Observation 04e3547b-2e58-49ad-b45e-cfdceeed581d · outbound

This paper cites 11 to 14 show qualitative examples of our generations, our controls, long generation and multimodal outputs.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control 11 to 14 show qualitative examples of our generations, our controls, long generation and multimodal outputs

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.841585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:16:34.345618Z digest=sha256:fb53b7ba19dcfc3fdf51123ccf95273341b8f1e06b0eebf86a3c29f437719a58

Pith citing papers

Observation ac5b25be-eb94-4d8b-a35f-de8e19f1f6cc · inbound

GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving cites this paper.

GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:48:22.339230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T13:48:22.279405Z digest=sha256:2c3f537a46254f5a8359099740cd40aa434ff50a0e2fa4655999b94ababecf48

Observation e64e1309-a665-4a08-8dff-aa00159e5cea · inbound

GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control cites this paper.

GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:42.902690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:42.902690Z digest=sha256:94f2a5454234f0ae37bd50bbaab45be09745596288ae0559f68ed714dcbf1663

Observation da65dc62-cd75-4d42-b702-276a4d3cc2e8 · inbound

Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation cites this paper.

Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T17:19:11.673125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:19:11.673125Z digest=sha256:6f6793b8b6910d76f38ab5c9e33658da934e47ca752ff7f02d083156f01b315d

Observation 8d9971b4-40b9-40d2-a10e-e34611b53554 · inbound

How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction cites this paper.

How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:19:47.034828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T07:14:59.175822Z digest=sha256:3ecd6aed32bfae19b798333f80ee5abafec2c3d1bda520462a6d005c1fb35a64

Observation 645d106c-2757-4d39-905c-4faf702b1677 · inbound

Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends cites this paper.

Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Reference 211

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:13.799031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T17:29:18.513507Z digest=sha256:f57055a1b80e91e9bac667cc59b2009c4637158bb7c1f1c51b8c8c3465f0cabe

Observation e209b637-182d-4001-a185-e1286e9b4871 · inbound

Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation cites this paper.

Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T02:55:45.953730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:55:45.953730Z digest=sha256:35527d9382088682f37f5ba26258b66df6dafa210cdd701255a87fca49622d9a