Pith. sign in

Paper Citation Record · LEDGER

World Simulation with Video Foundation Models for Physical AI

As of 16 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 100 inbound Pith citation observations for arXiv:2511.00062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.00062 v2

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T23:01:13.546110Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 100 of 145 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:42:24.884169Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 103 outbound references displayed

  • verified exact62
  • verified fuzzy34
  • unresolved1
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ab3cb743-68a0-4ace-b948-d3a1d9d89180 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

World Simulation with Video Foundation Models for Physical AI V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.764797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:713645750ea5dd426091884dbdb1bea03b36b1c0fe781d2e23a9389b3a170e7f

Observation 412f4703-524d-4c24-8067-82cce219e50b · outbound

This paper cites Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models.

World Simulation with Video Foundation Models for Physical AI Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.610147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:115e3405ff0373e66f46bded39d35629427ca4a4879dc0ae79575b26d65e5b72

Observation 8d4e97d9-0307-45a8-9a93-615803083ce8 · outbound

This paper cites Recammaster: Camera-controlled generative rendering from a single video.

World Simulation with Video Foundation Models for Physical AI Recammaster: Camera-controlled generative rendering from a single video

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.114818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:ff578e45b5e80f1dd8368b9ad569668e1645d59e17a5896ae44332e9bbe44be9

Observation 05cfa97a-37d4-47f8-a963-f576c59a03bb · outbound

This paper cites Syncammaster: Synchronizing multi-camera video generation from diverse viewpoints.

World Simulation with Video Foundation Models for Physical AI Syncammaster: Synchronizing multi-camera video generation from diverse viewpoints

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.080265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:959c34fdcd8b5227542776d7f01feb10554f486f639625bae456e3b4bfaf3b16

Observation 8ad112cd-35b9-4b7d-9eab-055948c21a6f · outbound

This paper cites Qwen2.5-VL Technical Report.

World Simulation with Video Foundation Models for Physical AI Qwen2.5-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.900179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:40c72d21fd176a3b0ea5c980c46d2d741e903e85da627eb75f1e026423486a0c

Observation 8c71e354-e252-43f9-8c9c-d59d67da6b9d · outbound

This paper cites Genie 3: A new frontier for world models.

World Simulation with Video Foundation Models for Physical AI Genie 3: A new frontier for world models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.090308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:07fe33f952308be0191ffc9b7ec0ec4f139476828425b8dac0cf96da2123c26f

Observation c17a0f9c-3cfe-4b8b-a98d-f1dec6f899ab · outbound

This paper cites VideoPhy: Evaluating Physical Commonsense for Video Generation.

World Simulation with Video Foundation Models for Physical AI VideoPhy: Evaluating Physical Commonsense for Video Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:34:37.994076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:9f63a7acac265553aa85b4e84fd4e156d86550e59a4aa8d35a286711d116e1bb

Observation feea1a15-9db2-4ab4-8253-1d0e571742b9 · outbound

This paper cites VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation.

World Simulation with Video Foundation Models for Physical AI VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.918325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:3e3b9e53a2bdfac27c460535eae98426815866ec6a748baf58518fdfebd0fd8c

Observation 8b76e304-375e-4841-a6bc-bf9387a2be39 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

World Simulation with Video Foundation Models for Physical AI GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.924336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:f89edb5d9e0ce75b3ca9459569b64170d47cbf2f17e8f7f0177da308aeb4796a

Observation 87aa9c59-6556-4152-a084-af0d4e6de054 · outbound

This paper cites an unresolved cited work.

World Simulation with Video Foundation Models for Physical AI Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-12T23:01:14.108994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:5ec5a593acb69fd8912aca1204769cd9ec94d9964c0c9f9511347c9e13fda695

Observation d5e3258f-1975-43a1-a456-a2d2404812eb · outbound

This paper cites IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments.

World Simulation with Video Foundation Models for Physical AI IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.930429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:740ccf821a5cea2ed85ffe4a2a78edc0311877b5980f95d95e68f84435f21c51

Observation be046de4-45a5-4780-8d99-c997383fe7c2 · outbound

This paper cites Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems.

World Simulation with Video Foundation Models for Physical AI Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.120244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:f61aaf1c4dbaa7e52466f7c2eeb9e5b45908c4d62c9d4b8db6d1a7d1734c21d7

Observation 78979042-1691-4b31-97a4-7a2df6df84a2 · outbound

This paper cites Planning with Reasoning using Vision Language World Model.

World Simulation with Video Foundation Models for Physical AI Planning with Reasoning using Vision Language World Model

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.939379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:048f81acdf316f47bfde1ac362d1248ccf240e502a1e7fe440a7045b431d8a82

Observation b4249491-da51-4d0b-938f-825a0180ee23 · outbound

This paper cites Video depth anything: Consistent depth estimation for super-long videos.

World Simulation with Video Foundation Models for Physical AI Video depth anything: Consistent depth estimation for super-long videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.130867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:8e6d9d1a343066df4e1d267df83b547f0d72116904ab1787cd05cd8e048bcec3

Observation b2e3c0e0-06c0-4587-9278-38e3d2fc6054 · outbound

This paper cites On the Importance of Noise Scheduling for Diffusion Models.

World Simulation with Video Foundation Models for Physical AI On the Importance of Noise Scheduling for Diffusion Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.947275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:8881e302beb10715d7b4b0ec8d760d6e59ea4c82460515f2a7c8b176c57cd4a5

Observation 55b562ef-1d13-4b22-996a-c77360e10b06 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

World Simulation with Video Foundation Models for Physical AI Diffusion policy: Visuomotor policy learning via action diffusion

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.152163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:ff7563dc94303057c9d5e66dee7bb671f4740b272850cfd48048fa26b1c5c544

Observation 9458ccae-8aac-41b5-afea-ceabf2f6645d · outbound

This paper cites Delta lake: Open-source storage framework that enables building lakehouses.https: //delta.io/.

World Simulation with Video Foundation Models for Physical AI Delta lake: Open-source storage framework that enables building lakehouses.https: //delta.io/

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.158787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:3d540c7c281ae64040d269828c4aa761ca8a506bee6dcc15039f3e74cf0ea1f3

Observation 49a52bbf-cba0-4db8-bdb9-078863bdba7e · outbound

This paper cites Veo 3, 5 2025.

World Simulation with Video Foundation Models for Physical AI Veo 3, 5 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.163715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:8c2eb75957ade3096af4ead690054d2b60cc26dfc0c84ab326ad67425a6f9d24

Observation c29ec697-2aeb-436e-ade5-3d8d1c0c930d · outbound

This paper cites Worldscore: A unified evaluation benchmark for world generation.

World Simulation with Video Foundation Models for Physical AI Worldscore: A unified evaluation benchmark for world generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.954077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:bbfc6598491effb02cf063524f7a9708c571bb2ff0b358b249b8e0f6991e179a

Observation 29ff6d43-a026-4531-8901-be2947ef9b2f · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

World Simulation with Video Foundation Models for Physical AI Scaling rectified flow transformers for high-resolution image synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.176994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a61350948e3dd2b7fb0b23c2c19119efc99f5b45c545a573d7369ec6cf777fde

Observation 3464cb8c-58e6-4fa1-906f-ac882462c28b · outbound

This paper cites LLM-based Realistic Safety-Critical Driving Video Generation.

World Simulation with Video Foundation Models for Physical AI LLM-based Realistic Safety-Critical Driving Video Generation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.962088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:aa4931974d7375c82f999b0d95f665040099205d34ba701973655dc9fc13ecc4

Observation 1f8223a3-7728-4dc5-845a-deb5eb77378a · outbound

This paper cites Diffusion models and gaussian flow matching: Two sides of the same coin.

World Simulation with Video Foundation Models for Physical AI Diffusion models and gaussian flow matching: Two sides of the same coin

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.195347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:337b9064936cc324209dc712d318a5e00cd165f660b474758ff31be9a5109d82

Observation edea4d0f-8dc8-424d-9fcb-f19dc243d1b4 · outbound

This paper cites Seedance 1.0: Exploring the Boundaries of Video Generation Models.

World Simulation with Video Foundation Models for Physical AI Seedance 1.0: Exploring the Boundaries of Video Generation Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.971409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:afdd27d17c95098f3a9e8181d56f48617cd07a57eb540de665c4a140703dba49

Observation eaa9cbc9-f6d9-4048-9310-6e68115ac020 · outbound

This paper cites YOLOX: Exceeding YOLO Series in 2021.

World Simulation with Video Foundation Models for Physical AI YOLOX: Exceeding YOLO Series in 2021

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:31:31.716715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c070a87a85bcf598ad1cfd9fb830c0379bd4b1aa7da2979876673cd7060145bd

Observation 3dadca2a-d312-45cb-9b6f-62d0f3855602 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

World Simulation with Video Foundation Models for Physical AI DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.986635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:8dbd4f72f99b4124c8488df31c57903d45661a2cce81603316826c0aa09bac23

Observation 8dc1b711-ac24-4cf6-b365-8735ce972af6 · outbound

This paper cites T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation.

World Simulation with Video Foundation Models for Physical AI T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.994170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:4fb688796001c9cbcd8c9502705e63df7556750c60cdb020c52608635af6df27

Observation f99882ba-cb42-46ea-9153-109dca9c5367 · outbound

This paper cites World Models.

World Simulation with Video Foundation Models for Physical AI World Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:14.002924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:161e98484a9c28347921e68f77832abb1d79fd5fe6a0341f1811b6ac1a7ccfda

Observation 906f848f-38b0-44f0-99bf-c1808cfbddd3 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

World Simulation with Video Foundation Models for Physical AI LTX-Video: Realtime Video Latent Diffusion

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:14.011752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:6e700aa1800f6cc2b1a8d9968c1e1ce16890cfc6ac561c9d7977fec07a1fd98a

Observation 4a1f96cd-9daa-47a3-b344-64bcabdb5e9f · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

World Simulation with Video Foundation Models for Physical AI Dream to Control: Learning Behaviors by Latent Imagination

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:14.018849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a0786f6e9ce2c11a2e8a449cd909f28e10bca88c78327b3c2b2f32ddc29f58eb

Observation aa7ef977-8a45-4454-abc7-f06974e97305 · outbound

This paper cites Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light.

World Simulation with Video Foundation Models for Physical AI Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.026420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:e4ed91d3d71d5c2159f1dfefff304ac693ad8d51b7269c94dea5f8f11ced3be3

Observation 1aac1657-8923-4a1e-9f87-cc03d0c41c31 · outbound

This paper cites UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting.

World Simulation with Video Foundation Models for Physical AI UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.035929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:76d47247e6ac2774b7ba495b74897d5e41f485668e4b74ff48f60b715a8264ec

Observation cd329c03-ce76-45d0-9f9d-4782b071fd88 · outbound

This paper cites simple diffusion: End-to-end diffusion for high resolution images.

World Simulation with Video Foundation Models for Physical AI simple diffusion: End-to-end diffusion for high resolution images

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.264196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:4b8e909cdb444a062ed76dd30ddac6083cfb682f4e6de329f4bbd5e3dc493bd3

Observation f9dcd7e3-e2e4-4e43-8037-dba17d16ab2d · outbound

This paper cites ViPE: Video Pose Engine for 3D Geometric Perception.

World Simulation with Video Foundation Models for Physical AI ViPE: Video Pose Engine for 3D Geometric Perception

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:41:08.910540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:52050f85b21c2dd800b57585ff05358ed3472ff50a9d2f324866552bbbd3c577

Observation 3842889c-98aa-432d-93c6-397d01e3b08f · outbound

This paper cites LET-3D-AP: Longitudinal Error Tolerant 3D Average Precision for Camera-Only 3D Detection.

World Simulation with Video Foundation Models for Physical AI LET-3D-AP: Longitudinal Error Tolerant 3D Average Precision for Camera-Only 3D Detection

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.050176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:7a02e8b53a9f521bc643bec11f067ada49b5a6c0c7b77546d3e2c121fff942c6

Observation d4e9e9b6-b193-4e97-8800-42732f720e3e · outbound

This paper cites GPT-4o System Card.

World Simulation with Video Foundation Models for Physical AI GPT-4o System Card

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:14.056202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c6eb35c0f12d3e7a59308093d134f1f399d4529c50b3e39b48cc6d9efc976cd3

Observation 284aa987-6557-4fc1-aa91-dcc35acd264d · outbound

This paper cites DreamGen: Unlocking Generalization in Robot Learning through Video World Models.

World Simulation with Video Foundation Models for Physical AI DreamGen: Unlocking Generalization in Robot Learning through Video World Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:50:45.800685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:e6e5d74fca2f8808b138240dcade9e09e660df618973f68fa7c9d93a390f6090

Observation f0765cd3-e892-4bbe-9d26-fa149d80302a · outbound

This paper cites RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose.

World Simulation with Video Foundation Models for Physical AI RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.070641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:0b75884c7f72a23522e541801c4ce521a79d115d5f84fb5320b06a36b84dc163

Observation 6dd5d3a1-74ef-486c-b209-98da15fff0a2 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.NeurIPS.

World Simulation with Video Foundation Models for Physical AI Elucidating the design space of diffusion-based generative models.NeurIPS

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.304402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:9d668dae599246484874de6709e60f4d01825f4d6cedd9c809657b25051002ec

Observation 3537471d-bbb2-4c45-8296-29eb79762abc · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

World Simulation with Video Foundation Models for Physical AI DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:14.075355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:8965438b85b7eb2e03cf8ada6eaa46b1a9c12a8c69c5f3b0807df19f514af5da

Observation 665fad89-9f5f-469c-a9e5-a035537641f5 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

World Simulation with Video Foundation Models for Physical AI HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.624574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:db4f8ff1c09292b6f750ae5e51d63ce88044d06ba19ef03eae2b26d0e354ed37

Observation b29f4ddb-f6b4-498a-9da7-c2e97480c021 · outbound

This paper cites Kling.

World Simulation with Video Foundation Models for Physical AI Kling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.321923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:3f1bb71a59022cbd4b3105099b638fbea7bf371e1f2934342480f7bbe032a146

Observation f2ee8ce3-f025-44bc-a2cf-8780d356ba4e · outbound

This paper cites BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers.

World Simulation with Video Foundation Models for Physical AI BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.632850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:dd08ba99d8acb67ba42ae214ff7b369053dc4de9f9617f7eefcbd753247ee626

Observation 5b6e1514-92d1-45c1-ada2-7903ddd6408a · outbound

This paper cites Won- derplay: Dynamic 3d scene generation from a single image and actions.arXiv preprint arXiv:2505.18151.

World Simulation with Video Foundation Models for Physical AI Won- derplay: Dynamic 3d scene generation from a single image and actions.arXiv preprint arXiv:2505.18151

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.639296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:11f364a9e992e42ece1ef8302b77b09d3958d73049e6a46d65cbfef67dfe94d5

Observation cbcafa83-5f1e-4391-ab60-06f49959080b · outbound

This paper cites Torchtitan: One-stop pytorch native solution for production ready LLM pretraining.

World Simulation with Video Foundation Models for Physical AI Torchtitan: One-stop pytorch native solution for production ready LLM pretraining

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.099057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:7b282f4a42174a45432cef5ffed3e10d2a1abf30f2ed6957f152e1655c5253ab

Observation 8dd4a2d0-814a-4cef-b4ad-421c3dd83fff · outbound

This paper cites Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation.

World Simulation with Video Foundation Models for Physical AI Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:28:42.077022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:ca6591fe42f764af36d22133ac6dda27931f95ab352d58c9d5c4ad238412a61c

Observation adbef28e-01d6-4b66-b831-e5918f9aadb2 · outbound

This paper cites Flow Matching for Generative Modeling.

World Simulation with Video Foundation Models for Physical AI Flow Matching for Generative Modeling

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.653660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:dee045c1a33bb180c17bd3343c952e8580de6d758e2c4690613ccd6a2f1f2eb1

Observation 44d2be4c-486d-448f-9aa2-ea991d47a487 · outbound

This paper cites Improving Video Generation with Human Feedback.

World Simulation with Video Foundation Models for Physical AI Improving Video Generation with Human Feedback

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:30:03.116566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c7f9a97f56bdde64815a9b83904e2866c6c31e19361c3aab2c706587001d65e3

Observation f99bbb63-c9bd-4b33-aa44-b2e411deb33e · outbound

This paper cites Dynamicscaler: Seamless and scalable video generation for panoramic scenes.

World Simulation with Video Foundation Models for Physical AI Dynamicscaler: Seamless and scalable video generation for panoramic scenes

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.138025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:ee7e47fa36569a0acd40ba4750d7f52e1b3d48463f5c07956f0ba108a50700b8

Observation 08efc4e9-13c8-481a-aca0-8d5367ee7e4b · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

World Simulation with Video Foundation Models for Physical AI Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.667121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:5a534f115cdf40b8f4b9654ce70c93ee19fe1471c25209f0c7934d9f45e7f333

Observation fd492237-9055-48ba-a752-7c5e6a233445 · outbound

This paper cites Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models.

World Simulation with Video Foundation Models for Physical AI Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:26:23.772771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:025888d6df741098fb022872fe5ec001ee9a00422217b5482f83fa744fa0e02d

Observation ef66dd02-13ac-41d7-9beb-570513e9222e · outbound

This paper cites LATR: 3D Lane Detection from Monocular Images with Transformer.

World Simulation with Video Foundation Models for Physical AI LATR: 3D Lane Detection from Monocular Images with Transformer

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.682645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:2d799ce5eab58d7eff6720a43c4d0fa791bc80488212d868ce021482609ea331

Observation e9e99ea9-b2e0-4f61-a539-5affb253ebf9 · outbound

This paper cites Hailuo.

World Simulation with Video Foundation Models for Physical AI Hailuo

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.206489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:58a5b4e096eac7c6c8c1ebc85aaae12da5de1df38ce47899528d4088f1d3dc9c

Observation b073b607-1fdc-4a9e-a313-abd9b69fd17e · outbound

This paper cites Do generative video models understand physical principles?.

World Simulation with Video Foundation Models for Physical AI Do generative video models understand physical principles?

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:47:06.031259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:0a5257f11c29f84149c494fa742fba756306a70f5026811260793053f8807738

Observation edf2c7b8-8352-4d53-ba76-1603399fd485 · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

World Simulation with Video Foundation Models for Physical AI RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:46:30.425882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:1e84ad18c4d7d361bdc9a27b476dee6ec7d84c40009056217bca65533f3d7c58

Observation 32f39998-bff7-4dce-8971-8a381add6192 · outbound

This paper cites Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning.

World Simulation with Video Foundation Models for Physical AI Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:47:10.348602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:6a72b91060efb5dbdf0db5bbc97cadd151b7d6b62869c9616e3a31ce9f33236e

Observation f344ecf1-a33d-49c0-8634-5443f93647f8 · outbound

This paper cites Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control.

World Simulation with Video Foundation Models for Physical AI Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.712342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:3c3f9ac9a23b40bdbe2984c0dcea165bd91588914f989005b6346261c952e076

Observation 6f8d0196-3608-45c3-a9f9-cdf0d070daea · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

World Simulation with Video Foundation Models for Physical AI Cosmos World Foundation Model Platform for Physical AI

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.716640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:2f3f44028fe4e1dc58d7b38db46d9c4190297c360a4763c8a0a31c582075f350

Observation ba398b82-ef48-45ec-8759-07d233eea72d · outbound

This paper cites an unresolved cited work.

World Simulation with Video Foundation Models for Physical AI Unresolved cited work

Reference 58

Resolution
parse uncertain
raw_fallback, observed 2026-05-12T23:01:14.248509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:d7ac4884647e8bc7a0b63a575e367cc1142bf2246d59bf0321b89919c07a4e72

Observation 89c3a877-9d7f-402b-9b79-fefb1ab37802 · outbound

This paper cites Sora.

World Simulation with Video Foundation Models for Physical AI Sora

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.256392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:26af796584c40966cff75f1fe593030c0a6d0de9b8d4727af503947bd6e54dd4

Observation 2ee28bcd-8222-4143-82a5-fe69c5d9dabd · outbound

This paper cites Training language models to follow instructions with human feedback.NeurIPS.

World Simulation with Video Foundation Models for Physical AI Training language models to follow instructions with human feedback.NeurIPS

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.269581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:75967ed92cda07e7952a07edd470adb273273a72efb33265a8074abba0460ea6

Observation 19018ce4-6981-4d95-9720-68469fb6f6d1 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

World Simulation with Video Foundation Models for Physical AI YaRN: Efficient Context Window Extension of Large Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.721168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:2d43830c75ff79465f6ed5e3c381f49da76fea0a43e84d3432009976fbb8a731

Observation 74b9026e-75e5-420d-80d1-c74b354abed5 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

World Simulation with Video Foundation Models for Physical AI Movie Gen: A Cast of Media Foundation Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.725884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:f149d76556026c5102fe99e0a8f7d75f8d64e4110d5a242dba1a6481593ba37a

Observation a2727a34-99a7-459f-adef-a336243af3b0 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

World Simulation with Video Foundation Models for Physical AI Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.287690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c4de40e8d3205d498c7b34da8048e40d5bb18c9769a6591a6448cd48a54a757d

Observation 66e125d8-e0fc-4ea1-8ee6-072a45efb597 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

World Simulation with Video Foundation Models for Physical AI SAM 2: Segment Anything in Images and Videos

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.730843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:6384728dcb0c16dc1ed655458ca2ae9c21b6c17f569d36c84080e6e8d11d6c36

Observation 95d616d7-bd92-46a3-b35c-e77a2a868564 · outbound

This paper cites Ren, Justin Lidard, Lars Lien Ankile, Anthony Simeonov, Pulkit Agrawal, Anirudha Majumdar, Benjamin Burchfiel, Hongkai Dai, and Max Simchowitz.

World Simulation with Video Foundation Models for Physical AI Ren, Justin Lidard, Lars Lien Ankile, Anthony Simeonov, Pulkit Agrawal, Anirudha Majumdar, Benjamin Burchfiel, Hongkai Dai, and Max Simchowitz

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.310856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:23dfcf1027838409af2db0fde96f8743215ae3006cb04b17918fa0f0903082fb

Observation 283d1d9f-ef01-48eb-b6ed-3b806058b365 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

World Simulation with Video Foundation Models for Physical AI Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.736123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:7bc4408054cbc181b4cdcc49105f7aa35a593c560f1908c0fd89175025741601

Observation 32c4917d-c431-423f-97c5-4237fd9f6cbc · outbound

This paper cites Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models.

World Simulation with Video Foundation Models for Physical AI Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:13.741361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:5dab68830bf0bdafaf08071bcc851a74aab3783322cfa1fdbc8b6555e950423c

Observation 654844e1-d016-48e3-a839-c110ecb5fcc1 · outbound

This paper cites Gen3c: 3d-informed world-consistent video generation with precise camera control.

World Simulation with Video Foundation Models for Physical AI Gen3c: 3d-informed world-consistent video generation with precise camera control

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.094409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a9b42c3ea3722345b3791aa60cadc397c72662d58fc976335e85b2fde0f52d4c

Observation 26ee39e3-336d-4d1e-b6ec-aeba3a8aa358 · outbound

This paper cites Gen 3.

World Simulation with Video Foundation Models for Physical AI Gen 3

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.104453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:2c98efd383afed8877bb2f815dbfa1daeec008219260a66840d9f104e33fc6f6

Observation 769806dc-f5fe-4f9d-be52-832c99f440da · outbound

This paper cites very scattered.

World Simulation with Video Foundation Models for Physical AI very scattered

Reference 70

Resolution
verified exact
doi, observed 2026-05-12T23:01:13.602107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:dfb131b4ac8f79ed5d0bd2447a06bf7198e113549d9e39f414d7620b28763045

Observation 1a43ad0e-d7a0-4a16-ba2f-6fa68d3d1a41 · outbound

This paper cites Proximal Policy Optimization Algorithms.

World Simulation with Video Foundation Models for Physical AI Proximal Policy Optimization Algorithms

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.746628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:fe349d1f24ef376eeeac716493ad60a4381b688fe2e4744d6351752756906732

Observation 43e802b8-64de-4ea1-b362-8ab0529edb7b · outbound

This paper cites Text-To-4D Dynamic Scene Generation.

World Simulation with Video Foundation Models for Physical AI Text-To-4D Dynamic Scene Generation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.753011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:8be1706a19eb31ac459af6e9a7a3e2fba55b9ad2ad78284fb9fba08506fb85de

Observation fc0617d6-52bd-4021-8f69-a2ccb5acd549 · outbound

This paper cites Light field networks: Neural scene representations with single-evaluation rendering.NeurIPS.

World Simulation with Video Foundation Models for Physical AI Light field networks: Neural scene representations with single-evaluation rendering.NeurIPS

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.187667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:829305165331b608b02a08fa879348ba8ad14c3fd1d857d9bb4e3abf5df03b2e

Observation 6ea03bf2-a2d8-492f-b240-f06edf88a77c · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

World Simulation with Video Foundation Models for Physical AI Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.201259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:0fe0558205e7fcb888a8e44d5bc06bb2666f2b2ce48947f78207a8b771fdf3ec

Observation 488ccc1b-86f2-49fb-82ce-e6bfcb82f4a9 · outbound

This paper cites cuRobo: Parallelized Collision-Free Minimum-Jerk Robot Motion Generation.

World Simulation with Video Foundation Models for Physical AI cuRobo: Parallelized Collision-Free Minimum-Jerk Robot Motion Generation

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.759032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:e5ccb73697d1cac9076eefbe935a2a188f60ac6b31c2109fafaf7ddf88ac442d

Observation 8833b5d3-c4f8-445d-920f-8debd9cf26b0 · outbound

This paper cites 1x technologies | safe humanoids for the home.

World Simulation with Video Foundation Models for Physical AI 1x technologies | safe humanoids for the home

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.223399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:9ce4e08e7d03c867701942cf7bfbe18b94bdef7d439c7df487123888c9f98c86

Observation 2f7c672a-6f2c-446d-a1bd-d7f359c57fe8 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models.

World Simulation with Video Foundation Models for Physical AI Open x-embodiment: Robotic learning datasets and rt-x models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.228413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:1ce456ef5ef56b259778dce193fe1c0ffe10c5b8f712b82e91ff86e91d463705

Observation 944ad03f-8f43-4bcd-b564-08fefcf4f9cb · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

World Simulation with Video Foundation Models for Physical AI Bridgedata v2: A dataset for robot learning at scale

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.233658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:74f02d55866915b9388a21fc24115fd075f689cdd4c61fbf679e223a43e5a9eb

Observation 49be7c1a-e43e-48fb-b04d-dd4d4cf4ebe1 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

World Simulation with Video Foundation Models for Physical AI Wan: Open and Advanced Large-Scale Video Generative Models

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.617255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:deb6a4ccf5a532c23d80bcc910c8a783c6bcd13aa04e2f04c816e15cf5bd9ff4

Observation 99eae887-fd5f-4c07-9951-138dbeff236e · outbound

This paper cites A comprehensive study of decoder-only llms for text-to-image generation.

World Simulation with Video Foundation Models for Physical AI A comprehensive study of decoder-only llms for text-to-image generation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.274833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:9ca3e213f3dfd28e964ca0631d76852c0d634ac8f70cd7f7880ee8ee8b9c0d91

Observation 44c15761-9b25-49a0-aa9c-1407d55be937 · outbound

This paper cites Frame in-n-out: Unbounded con- trollable image-to-video generation.

World Simulation with Video Foundation Models for Physical AI Frame in-n-out: Unbounded con- trollable image-to-video generation

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.770365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:ab76b612478193590ff394b745f8761861ff2bf21e59af5ce749e9efa9c795e4

Observation fb2dc8e2-b8b4-43a5-a17b-f3ce4ba162fb · outbound

This paper cites Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content.

World Simulation with Video Foundation Models for Physical AI Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.296587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:ed3997dae222d2c442d79e94e1623d8391794dbe2a7921124667f6ecfaee1a3c

Observation 866afd17-3d65-4d32-95ae-5ddf1427481f · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

World Simulation with Video Foundation Models for Physical AI Internvideo2: Scaling foundation models for multimodal video understanding

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.315657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:7e2bd718e32c1848da64ed9a9ca0d3ea201cc31aefe15d10d5fdd5d9b19adf61

Observation f53d5fe8-576d-4941-9640-12dca55e4056 · outbound

This paper cites Controlling Space and Time with Diffusion Models.

World Simulation with Video Foundation Models for Physical AI Controlling Space and Time with Diffusion Models

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.776381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:23e37e80b708e1fa62984cccfcded621e4f97ec0f17162dc44daefeaa39bc094

Observation ed7ba837-e571-4f04-b041-a9f4e77c7902 · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

World Simulation with Video Foundation Models for Physical AI Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.125708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:95604198d9d3c63c15c7456ece36ac1af3b61f43552045f46cd961c9de7b6252

Observation 3bdeaf82-ebd2-456b-8442-0534bb253c71 · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives.

World Simulation with Video Foundation Models for Physical AI Exploring video quality assessment on user generated contents from aesthetic and technical perspectives

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.169733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:b7c4f14ed403c60c5128ed321d2d2a2286a710b2f66636328f606ec476e6e8f3

Observation 68418d26-aa82-45e3-a780-4df51bc9b894 · outbound

This paper cites RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation.

World Simulation with Video Foundation Models for Physical AI RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:14:18.432154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:1ad21351557accdaa8eeac4696a3ce18bf005285a331f9d5e2508c739a5717c8

Observation 1c2f1d5a-30c3-4efa-972e-32f679470874 · outbound

This paper cites Ties-merging: Resolving interference when merging models.NeurIPS.

World Simulation with Video Foundation Models for Physical AI Ties-merging: Resolving interference when merging models.NeurIPS

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.240424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a4bd154d0c75f8c61448444125614159924cd7510b8a13152d620df925c3d1b7

Observation 51988b6b-b002-4279-ba30-465ed87158a8 · outbound

This paper cites Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities.

World Simulation with Video Foundation Models for Physical AI Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:16:05.189340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:fa883aa023a34e0f10fa3f305dcbacadc7a15ac0cf842dbc0b34f535dc7b6397

Observation 0c3c7d1d-2359-447e-acd6-b3946edd022e · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

World Simulation with Video Foundation Models for Physical AI EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:24:45.426358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:81c5e04fd6dcf651f2064e152243f97625df011cae92d4cb114b23e8c6913444

Observation 74b42442-84f4-46f9-b307-b658f57c8fe1 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

World Simulation with Video Foundation Models for Physical AI CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.799820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:43f3f8a3a868dace81d1693c9f1903913793e79a47cff89984cfb736ed5721f8

Observation f5188f6a-8cf5-40e6-b1b3-820765e7236c · outbound

This paper cites Data-regularized reinforcement learning for diffusion models at scale.

World Simulation with Video Foundation Models for Physical AI Data-regularized reinforcement learning for diffusion models at scale

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.811776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:6eba2f6f2a11c4241af9d7216e42133d0e0e6c258e2a9d5fb34df27b66fae738

Observation 252b42ad-53b3-4bc5-a695-a07a8dea90ff · outbound

This paper cites Language models are super mario: Absorbing abilities from homologous models as a free lunch.

World Simulation with Video Foundation Models for Physical AI Language models are super mario: Absorbing abilities from homologous models as a free lunch

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:01:14.085751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:0cd6a3f1a6661c7aadb8829c32408b9858f9c43183edb2360f2fded11551b8c1

Observation 8af2157c-8527-42e4-97d2-2543d4aba109 · outbound

This paper cites EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models.

World Simulation with Video Foundation Models for Physical AI EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.822345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:4e299b9409b1592989fabaa8687b84ede30d75a3e1e5c5a827b3104acc75d068

Observation 2ec731de-1750-4793-9558-1281b08fe6fe · outbound

This paper cites Waver: Wave Your Way to Lifelike Video Generation.

World Simulation with Video Foundation Models for Physical AI Waver: Wave Your Way to Lifelike Video Generation

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.834624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:b228e1c563c999de8d747b4741a3546d2ba54b46cf9318fbefca5ed52197dc84

Observation e6b72cb2-cb9b-4d0f-ae62-3342d8ab69a7 · outbound

This paper cites GenXD: Generating Any 3D and 4D Scenes.

World Simulation with Video Foundation Models for Physical AI GenXD: Generating Any 3D and 4D Scenes

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.848208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:8ff7a7b7004293ca6894b59de24f2d4ed53aa6b41d1d7c55aeb063908a23ae2c

Observation 49ce5cb8-f139-4fb8-9c88-1d606b8f9a2b · outbound

This paper cites Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos.

World Simulation with Video Foundation Models for Physical AI Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.858638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:12cfef3cfaca156f0b1c599cf78054c6dc3dceff2415f40dda7630f3ece8886d

Observation 333d6195-f51a-44ff-bbea-23e061f4f542 · outbound

This paper cites Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency.

World Simulation with Video Foundation Models for Physical AI Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.868632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:2a7bd356a91c38f0677b794215f9857f4c112933676228007842359f2134f1f1

Observation ba33fbdf-c62f-4392-a184-c2149205d762 · outbound

This paper cites Unipc: A unified predictor-corrector framework for fast sampling of diffusion models.Advances in Neural Information Processing Systems, 36:49842–49869.

World Simulation with Video Foundation Models for Physical AI Unipc: A unified predictor-corrector framework for fast sampling of diffusion models.Advances in Neural Information Processing Systems, 36:49842–49869

Reference 99

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:13.875615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:f2ce00f553befb9f6a4428e911364bbb6d25c5f929bd72c558d2cc1b38dcf3b2

Observation b9d7bddd-8785-4d5c-bb5e-1031eda8cec5 · outbound

This paper cites VLM4D: Towards Spatiotemporal Awareness in Vision Language Models.

World Simulation with Video Foundation Models for Physical AI VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.883037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:a86a76f4dea3e39320116eb712cfa8923128b0900020380c59b3a4179cd4a428

Pith citing papers

Observation fa515d66-8adc-4190-9ce9-b7cb599cf1bb · inbound

Critique of World Model cites this paper.

Critique of World Model World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:42:24.884169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:42:24.884169Z digest=sha256:309e12cabacdc4f422975dc81096c18b2ee7824f25bdcbd334d0fca76c06cfda

Observation 3f536bf9-a60d-45b9-8dfa-2c26111dbfac · inbound

Non-invasive Assessment of Pancreatic Duct Hypertension Using Computational Flow Modeling cites this paper.

Non-invasive Assessment of Pancreatic Duct Hypertension Using Computational Flow Modeling World Simulation with Video Foundation Models for Physical AI

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T18:05:46.912489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:05:46.912489Z digest=sha256:b24e192dd0ae07ba847a8db939a80a993d89768f93d1bf08f50adda7707d628a

Observation b3a173d5-3c1b-43f1-b9e8-e75981a48582 · inbound

SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios cites this paper.

SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T21:14:08.241793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:14:08.241793Z digest=sha256:bdbfcfa378d3c146abc511247c35368d2994f6e78e4801d91498404616f3d704

Observation 19866105-002a-4d37-a1c4-952a478a6906 · inbound

Action-guided generation of 3D functionality segmentation data cites this paper.

Action-guided generation of 3D functionality segmentation data World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:51:31.961334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T04:49:06.269256Z digest=sha256:5457735740dea0155e4d41c00e49a55b5764d30efc21dd630f1f557abb175e4f

Observation ffb8ebf1-a0b8-4ab6-b101-e4f911ddd5bf · inbound

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs cites this paper.

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs World Simulation with Video Foundation Models for Physical AI

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:41:00.371735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T10:41:00.142543Z digest=sha256:ccb1c9b62dbde3059fecd5415d06827366ef81c2b2d9dc382c7c25775e8d129b

Observation 0ea168d0-4a6f-44bd-9c15-f0726a94a156 · inbound

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents cites this paper.

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:58:31.824406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T20:56:58.770875Z digest=sha256:67648e2a309a3b19dd2d0a543796572fb0cef71d01dc9dc51994b87e742591a8

Observation 42fe7b14-e06f-41cd-ac2d-85ccf240be77 · inbound

AstraNav-World: World Model for Foresight Control and Consistency cites this paper.

AstraNav-World: World Model for Foresight Control and Consistency World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:28:20.868354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T19:23:58.769472Z digest=sha256:a174faaae302de6b96b359d39f673c21d1d2a6550803d26de74c84f4e1dfafe3

Observation 0c5207d5-c917-4f4d-a024-e34e56eab943 · inbound

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World cites this paper.

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:38:21.104574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T19:34:39.518649Z digest=sha256:51f3b1b3448e3cdcff8e9e314809b1aefafe2f7dd363f93e3d50e1b2e8ef9c3d

Observation 0110820c-4239-499e-a7bd-9fce4a723f27 · inbound

CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding cites this paper.

CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding World Simulation with Video Foundation Models for Physical AI

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T12:37:28.494262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:37:28.494262Z digest=sha256:3fe24ec716df417b839f0630918027c964d4f5fb73d8a3ceaa7a652e6effbaf3

Observation cd8451bc-67de-427e-94c8-e4efbcaf06f5 · inbound

Advancing Open-source World Models cites this paper.

Advancing Open-source World Models World Simulation with Video Foundation Models for Physical AI

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:07:00.942546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T09:07:00.904794Z digest=sha256:9f84c85853eed7c534ba1d022aa2b5e2696c1812465d30da0d868954569e0286

Observation f9612596-cb44-414f-bc59-5f979c5d309c · inbound

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy cites this paper.

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:47.126834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:56:47.126834Z digest=sha256:37d2ac5e7164a4524099de425a65f136d6432482fe394f47f2c14d0dc53e0105

Observation e9d9abb1-f484-4fe4-8a09-de5a2e057d81 · inbound

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos cites this paper.

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T17:02:34.267430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T17:02:33.997887Z digest=sha256:bb0ddab26877da2632614aa4c2ec9730aec3ac23e668e6337e49fda899a1eff3

Observation 7332bf28-128b-4f1d-95c3-82f7f3f881f3 · inbound

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion cites this paper.

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T07:07:30.032697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T07:02:38.876518Z digest=sha256:cb87e4bd6bb27b7216cc566f7296e2fd97ce94275b0bb89a30309fbb54829e89

Observation d218176b-7e28-444f-b351-30d5d2c1c658 · inbound

RISE: Self-Improving Robot Policy with Compositional World Model cites this paper.

RISE: Self-Improving Robot Policy with Compositional World Model World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:30:32.148569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T02:28:37.997148Z digest=sha256:6db744717c5ac79d48cdf26f4dce44a0d4194f107626d50412c92392b731c45a

Observation e8de21ed-7a5e-4bd5-aa18-5e3d79c67d72 · inbound

World Action Models are Zero-shot Policies cites this paper.

World Action Models are Zero-shot Policies World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T16:18:15.003371Z digest=sha256:8103ddc40fa76694c25892c49351107fd35c4ea0393d419653bb174f37ed0f9b

Observation fd1e209b-bbf7-4fd5-8078-18c3e1614fa1 · inbound

Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints cites this paper.

Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T12:20:00.593719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T12:18:45.538658Z digest=sha256:6bfec3863b0004ee93948cd660cdf7201f9e00c442db32b89c90669d16ed79b9

Observation 829bb067-274d-43b5-9162-ab1af4f452bf · inbound

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms cites this paper.

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms World Simulation with Video Foundation Models for Physical AI

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:38:36.297567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T01:35:14.878069Z digest=sha256:b56c8856836f812949ab2ac1ba097b79347e56fea4841bed9aecbb74531bb654

Observation a553edc2-0934-4d58-81eb-de4ec5248f1e · inbound

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking cites this paper.

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking World Simulation with Video Foundation Models for Physical AI

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T02:31:02.007136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:31:02.007136Z digest=sha256:fd926560ca944f3027045452fe4a0fcaef8852431760d79b8a2a1a46a0056380

Observation f428f70b-ebc1-4390-bd81-c670992afeec · inbound

Lifting Unlabeled Internet-level Data for 3D Scene Understanding cites this paper.

Lifting Unlabeled Internet-level Data for 3D Scene Understanding World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:18:21.015364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T22:16:57.890955Z digest=sha256:5b75c414f88c04d0ab798cb59c71a3c8bb5328abcc6801eb0d7e1d8d7a76d319

Observation d7239c86-ffbd-4f65-8fc4-5d87d982b5f2 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T19:36:42.100191Z digest=sha256:80ae3ef8a28844417a7e529f0ce79efa21c72e63e3d478abb4b788dade45d37e

Observation d6ab8ce6-250c-4191-aa3f-8b7db6070145 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T09:42:23.808691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:42:23.808691Z digest=sha256:623402d20aaee07a8932622806bfaf7b75e84747625423e0148a1dbcec25a9fe

Observation abdbe0d3-d15a-446e-b6c7-e80340e0ec62 · inbound

Action Images: End-to-End Policy Learning via Multiview Video Generation cites this paper.

Action Images: End-to-End Policy Learning via Multiview Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T18:51:05.206602Z digest=sha256:775d9e90acf22f12e3aa5fe6ae36218e7b6876df4c83d3473ad67aecf7035a63

Observation 422b8182-9835-4b95-95b1-4473873e0e50 · inbound

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations cites this paper.

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T17:16:23.355682Z digest=sha256:edaf79b4eaf8ab771f3d864ff9031441259312ccf09e89dbcba4d12666cff363

Observation 83cdf01c-32a7-4a92-9d8d-2a7220f58fed · inbound

MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models cites this paper.

MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T17:11:14.242522Z digest=sha256:daee40435f0ff98daee3c37388c96c323e548b814b0c2a0e961db6babb988596

Observation 50bb55b6-4159-413c-abc3-14c00f504715 · inbound

Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction cites this paper.

Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T16:30:53.578491Z digest=sha256:1d50f6796adc344172e71bc5311ff3a78dedad801ea3cd026ce8a9f66b791891

Observation 780fa249-5e2b-40b1-9161-f5c2d82e6cbf · inbound

SyncFix: Fixing 3D Reconstructions via Multi-View Synchronization cites this paper.

SyncFix: Fixing 3D Reconstructions via Multi-View Synchronization World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:20:04.440192Z digest=sha256:f463edb1b0a2e18910dd4bfdffadd778c3973d70d559a8b69cfb5898eee90dd0

Observation 479e6cdb-9b9b-4c34-8193-940ee358f451 · inbound

ShapeGen: Robotic Data Generation for Category-Level Manipulation cites this paper.

ShapeGen: Robotic Data Generation for Category-Level Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T10:17:05.939312Z digest=sha256:1db9c0a2e4140506441a7b16618d9812cb53b1a164ffe337a58f87b571b0ebfb

Observation 0ed5f6d6-1d1c-4220-bacc-ff5fe532e494 · inbound

From Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and Evaluation cites this paper.

From Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and Evaluation World Simulation with Video Foundation Models for Physical AI

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T08:45:46.944379Z digest=sha256:a8ed488126607b3c05951e1b151a4be2da9576b0eec8cc444c667e12ad91040f

Observation 062bf222-1e5a-4795-80f6-c44d909a736a · inbound

Human Cognition in Machines: A Unified Perspective of World Models cites this paper.

Human Cognition in Machines: A Unified Perspective of World Models World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T08:12:15.663761Z digest=sha256:d613ebfddc82bd6ac868382e1cfa38622a97757a345d39116e411ed4b9047495

Observation 476c1df0-c9b8-43a1-ba13-518b7294a8ab · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T06:03:06.920592Z digest=sha256:a95fe1d9f5831e9d33e390744b2ed0bfe19865c87e04f26930ba650e68ecda40

Observation a36559ad-8a7a-46a8-9299-77092c1d3db8 · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T02:52:58.232585Z digest=sha256:5a31453fd7ddb155b203da6061bdca5968e61a9c8a9f967597209cb57b0299e1

Observation 9510abdc-f1c1-4fae-86d0-b88035f03019 · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T16:51:14.240617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-05T16:47:32.853010Z digest=sha256:06b2c1541b897a902fb14b0aed2c27e023fd99b8ba31079515e86b68382608c4

Observation a1d3144e-8776-4ac3-b493-34198b3f9357 · inbound

MultiWorld: Scalable Multi-Agent Multi-View Video World Models cites this paper.

MultiWorld: Scalable Multi-Agent Multi-View Video World Models World Simulation with Video Foundation Models for Physical AI

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T05:06:11.514186Z digest=sha256:6facc67659db3cb19b5973cfa7e126af6f7c46864cf47768f662113157182f5d

Observation 7d185eac-58e0-45a9-8009-1b1b9f3e4da6 · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T03:02:26.084185Z digest=sha256:e0c4c56ea67776d7d2dcf848ced9b084c99fd25ead9d3a337aab440fb08ab272

Observation 08d65f79-7c4b-4ebf-98d5-bb5732f81ade · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:19:50.279253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T06:15:32.881140Z digest=sha256:6aa15c64e5e0f17a1984b0adf898675876f7777f5888fcf3c3566c06fda2794a

Observation cc575329-78a5-4ad9-8af7-fe0ea9cf7c5d · inbound

Mask World Model: Predicting What Matters for Robust Robot Policy Learning cites this paper.

Mask World Model: Predicting What Matters for Robust Robot Policy Learning World Simulation with Video Foundation Models for Physical AI

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T02:14:17.676675Z digest=sha256:957123b0335a98ebc543101a304ec6c18dd24fb4df7434ecd6025c88f7c5e62c

Observation e736b80b-c559-45cf-96f9-f4ffd6c23873 · inbound

Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics cites this paper.

Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics World Simulation with Video Foundation Models for Physical AI

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T23:53:09.614994Z digest=sha256:40b23c203dc0a6602fa75a4ced5145e86e3472b67bee65133c97e2612cecbfe5

Observation 4e5f7877-17c3-4ec6-a432-db1e23b5cccf · inbound

Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training cites this paper.

Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T21:26:26.540403Z digest=sha256:d8e6928efce4830a4845bf87155fb1efa46d3135215cfcec1f079b73085c4d31

Observation b0566a21-f8cf-4eff-9d9e-a1eba39f51b7 · inbound

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling cites this paper.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:4d80bf323df54e12cc23ca30a71957137f77304d45406905c54314a8f0c56ed9

Observation 34795d18-378b-4dc0-aa2b-d1cebad163ee · inbound

Learning physically grounded traffic accident reconstruction from public accident reports cites this paper.

Learning physically grounded traffic accident reconstruction from public accident reports World Simulation with Video Foundation Models for Physical AI

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T19:59:27.929899Z digest=sha256:282a42032af2f43bd0fbcb523a017135ec7153e143a3945a78572167b896f4b0

Observation 53bbaa4a-b583-462f-9ea3-f96b48cdd36e · inbound

World Model for Robot Learning: A Comprehensive Survey cites this paper.

World Model for Robot Learning: A Comprehensive Survey World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T20:38:12.709629Z digest=sha256:7dfd94946e56426adf8fcd280f33b51bd58f02e8097f9a2c76e06599491d9d4e

Observation 096cd1c8-87f2-4264-806f-9488893c8fc6 · inbound

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation cites this paper.

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T18:21:00.089755Z digest=sha256:f162f83a48590a4037ec09dacedf35d0c5f16c122cd59148175f3ee4fa9029dd

Observation 530c4c0e-acb4-45b7-a087-145d7c9c3d87 · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T16:06:10.595164Z digest=sha256:29604b391dd21e8bc8ad42080d5b8768f7a4de6cc955c1f17bf86372fa6a8bf7

Observation 3014a16a-ab53-4b04-ae5b-40792c4d8723 · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T19:02:19.756092Z digest=sha256:5dd73e6aeaafafdb07a8e5c8edc09c824c672c327933f2a1e16bd9d24f2f1491

Observation 9fb38d36-dac3-47be-8b49-1d3c523199e9 · inbound

Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models cites this paper.

Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T02:31:30.530595Z digest=sha256:f54dbcfeffa0ab9cb9371ceb1ba131f9f0fbd6f64254826f7396b67a0e46219b

Observation c8695a0e-d55a-4eee-ab2c-18803b91979e · inbound

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models cites this paper.

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models World Simulation with Video Foundation Models for Physical AI

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T02:17:30.227755Z digest=sha256:dca76170f739992d522dffe442138d5c9aa46914605257bb34866fbdb507ffda

Observation a07cccf4-f49a-42f0-af2b-663b6299d52d · inbound

Reinforcing VLAs in Task-Agnostic World Models cites this paper.

Reinforcing VLAs in Task-Agnostic World Models World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T04:27:14.277708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T04:17:51.349213Z digest=sha256:16912f0fed4f797779563a689c713f210a4d34ee7735a43c69932b4e39170916

Observation 986f8d80-af33-4be2-8288-fdd7ba5dba9f · inbound

Reinforcing VLAs in Task-Agnostic World Models cites this paper.

Reinforcing VLAs in Task-Agnostic World Models World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:14:03.037059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T08:13:40.975340Z digest=sha256:0206f85e1c654165fbc8f25cbdd2ec424235464f22eb80142f0ebc7c7aeab365

Observation 516c767b-71c5-4c2f-91d5-68956fcce97e · inbound

Di-BiLPS: Denoising induced Bidirectional Latent-PDE-Solver under Sparse Observations cites this paper.

Di-BiLPS: Denoising induced Bidirectional Latent-PDE-Solver under Sparse Observations World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:07:50.987224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T19:06:18.193097Z digest=sha256:ae53a3afbe3e205dea8b2a87369c30cdad65e272aca662f44e01c9589f625afd

Observation d4c23c5c-4dd4-489f-a9bd-f90ffb6b02f6 · inbound

Coding Agent Is Good As World Simulator cites this paper.

Coding Agent Is Good As World Simulator World Simulation with Video Foundation Models for Physical AI

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:23:32.034608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T02:19:49.149075Z digest=sha256:a6c13e8cc83ad2faf07bd6217e830285fc63323937a6fa3b6341da56539113a2

Observation a078f570-9f76-4132-8f15-06ccece8828e · inbound

Coding Agent Is Good As World Simulator cites this paper.

Coding Agent Is Good As World Simulator World Simulation with Video Foundation Models for Physical AI

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:35:46.852607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T20:55:16.040505Z digest=sha256:6535b0f591104e17120364ea375560350e0fe2c4c37b9a08866bf4ff34781002

Observation b5ed68ce-66e3-44de-8b02-9e582067c172 · inbound

DriveCtrl: Conditioned Sim-to-Real Driving Video Generation cites this paper.

DriveCtrl: Conditioned Sim-to-Real Driving Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:15:04.787272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T21:06:06.548538Z digest=sha256:e7e695931d1b350bd96d4859811b554b21b9dd95f9261edbb742592fafea334c

Observation d91c1c62-87f3-4fc8-a354-8c5c442bcfca · inbound

Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action cites this paper.

Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:19:44.187738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T03:15:35.678011Z digest=sha256:f81acfea55b92fd51463c35598416141b32e63c751d9feffd6a22a30090ed644

Observation 173d31fe-008e-4bbe-86ce-73d7b0502d74 · inbound

Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action cites this paper.

Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:46:22.452132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T09:45:13.742534Z digest=sha256:4d42f9034e55db614882bb0d28d8bf4a6bd070c47cdc1bb41502c34dd2b99a95

Observation 425ee546-4408-47c4-bd50-17460cbc0c11 · inbound

How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning cites this paper.

How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning World Simulation with Video Foundation Models for Physical AI

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:23:25.279166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T15:18:40.457780Z digest=sha256:ae5209d58718a2a482aa142067f5a2bdc09039bbf4a034ea3829c18b260f94a6

Observation fa290429-9457-46bb-9934-66c82d90dacd · inbound

Self-supervised Hierarchical Visual Reasoning with World Model cites this paper.

Self-supervised Hierarchical Visual Reasoning with World Model World Simulation with Video Foundation Models for Physical AI

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:43:16.866382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T12:41:17.559901Z digest=sha256:c1753fd39fe07c69f67ace76005d06b1f276aeb4091d0e9cfbe3a8a65f2766d3

Observation 95302fe6-4ace-4927-a355-6ff92b627d65 · inbound

Self-supervised Hierarchical Visual Reasoning with World Model cites this paper.

Self-supervised Hierarchical Visual Reasoning with World Model World Simulation with Video Foundation Models for Physical AI

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T19:05:00.273959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T19:04:33.433907Z digest=sha256:3275d6d9071a941063fb7177883bc31dd8e408810d34e6bf00cbc24dcd8ae8bc

Observation 88a07589-9b1c-4766-8c1e-09818ccd60fa · inbound

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform cites this paper.

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform World Simulation with Video Foundation Models for Physical AI

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:03:13.688410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T10:59:54.879907Z digest=sha256:451f22e3f97e7ba1f624df3ca7ef8c563802a91eba1ac39da0f320d10bd12778

Observation 5d7afa79-4545-41d9-832b-d6af11a6f445 · inbound

NEWTON: Agentic Planning for Physically Grounded Video Generation cites this paper.

NEWTON: Agentic Planning for Physically Grounded Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:58:13.800270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T10:56:22.343333Z digest=sha256:bc512150bf43c6edaa142c4d06dce03b0b9db726420ec2e909541f0920724875

Observation d1e93c55-63b7-496a-bac3-7e435c202e64 · inbound

PhyWorld: Physics-Faithful World Model for Video Generation cites this paper.

PhyWorld: Physics-Faithful World Model for Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T07:33:07.562468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:bbe43e214a9e08afa0c80b093bb6a1f665639e8ae145f6cc68bdd959ac102b08

Observation 2d959405-5891-4b63-899e-9b1337ae3a52 · inbound

World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks cites this paper.

World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks World Simulation with Video Foundation Models for Physical AI

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:38:05.636274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T06:36:47.265734Z digest=sha256:3cedc1832e4832d9244216285041b4f6b52b0a968a448e5db75027045c5c1108

Observation 9c206a40-a2dd-4020-a16e-0e2c99729dab · inbound

Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models cites this paper.

Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models World Simulation with Video Foundation Models for Physical AI

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:13:59.578801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T06:12:31.783615Z digest=sha256:872971d09e20e79e22551f40ea7343d022c143a634329e14177c9d28cfd28fab

Observation 1a8f9b0a-1dbb-49ec-a680-f27e3d4dae36 · inbound

CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models cites this paper.

CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:40:23.268359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-25T04:39:22.400458Z digest=sha256:75cabaeb8f0745f4cea1baa431cf8d8bd57422add3964cce83c6509bce9723ce

Observation 7cfd86f1-b242-4fd6-92ef-9e9ec2370a0f · inbound

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios cites this paper.

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.356445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T18:23:22.987086Z digest=sha256:5cbfded4a2dc09343a6a4e246008a0fe5dfb27d91f38e788fa18c9b101e5bea1

Observation 68f05e78-715d-4bcd-a2b8-b9b3a4a8522a · inbound

Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation cites this paper.

Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.638052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T12:55:24.689338Z digest=sha256:4fcdfcc400d115ce4de30fbf04c5891ffdffcf76e9d36c9bdeba7f3220644486

Observation 2e702759-30b6-42cf-aadf-a3d2d85f50ad · inbound

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players cites this paper.

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:53:29.115357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T13:34:03.020051Z digest=sha256:c66a2fd81211869482e6e19429dc563490dd74faf9fbe781486136a8923c9d18

Observation 226e4e9c-60ca-4416-9714-e137062acefc · inbound

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications cites this paper.

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications World Simulation with Video Foundation Models for Physical AI

Reference 108

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:43:15.654464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T08:36:23.776293Z digest=sha256:4f2cd0516d3dcd4757687c08673df4cd9eee463907ca012e79fdd7eb84fc43a7

Observation 0713b3dc-b9e4-4d64-a9ae-05704cec4938 · inbound

OptiWorld: Optimal Control for Video World Generation under Physical Constraints cites this paper.

OptiWorld: Optimal Control for Video World Generation under Physical Constraints World Simulation with Video Foundation Models for Physical AI

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:32:35.509546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T19:02:51.848742Z digest=sha256:91fd4c05f5d8b4b7cfd8f2e8bc36e3e0ddb62a1d0bb590a3ae6e3af4b96d73d7

Observation 49d26b9e-d761-4967-be0f-f5593d6b32d7 · inbound

$\tau_0$-WM: A Unified Video-Action World Model for Robotic Manipulation cites this paper.

$\tau_0$-WM: A Unified Video-Action World Model for Robotic Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.504637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T17:21:06.079898Z digest=sha256:1d8d0ed902ace24b5086ab74315daefd87e16860bd59a8e4ee6847b86a4ed753

Observation ee9ffe6e-34bc-44e0-994b-e200ff4a3f33 · inbound

Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA cites this paper.

Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA World Simulation with Video Foundation Models for Physical AI

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.413659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T17:21:15.500408Z digest=sha256:0cd8387ae007b3356d129ce7e34b0777b3357bb0afa3575e2742a6bbfdd7e57d

Observation 0866c75f-eb61-44fb-a5c6-f73e10f09d6a · inbound

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning cites this paper.

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning World Simulation with Video Foundation Models for Physical AI

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:20.428785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T14:41:50.084254Z digest=sha256:5c505ddec12982cdb3871cba934766775fedcd8b0096926a3bd6a88c24f5e51e

Observation 47e0ba62-cfe3-4f77-8281-3863146feb2d · inbound

RoboDream: Compositional World Models for Scalable Robot Data Synthesis cites this paper.

RoboDream: Compositional World Models for Scalable Robot Data Synthesis World Simulation with Video Foundation Models for Physical AI

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:36:23.584650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T14:07:31.809341Z digest=sha256:921532c136de2b23354d7c2a317233c66ffad4f9b885f3ae9e621ce6b2a80c90

Observation b294cd98-cc39-4856-941a-61a9052f498e · inbound

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation cites this paper.

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation World Simulation with Video Foundation Models for Physical AI

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.451431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T11:11:46.428727Z digest=sha256:65b1656f60197d7a29a0bcdc07af4d11a79d13c4162b422d1f7c8b4f02d95187

Observation ba28d442-6e81-405f-8219-aa41da504410 · inbound

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation cites this paper.

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation World Simulation with Video Foundation Models for Physical AI

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T12:34:17.389528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:34:17.389528Z digest=sha256:3cfcf5b65047c96d4d5be4a6f75d9abf1b88bc2b2365df02656b6008ba1e82c8

Observation 1177adce-4513-4cb5-997f-ef3e4beaeebd · inbound

PointAction: 3D Points as Universal Action Representations for Robot Control cites this paper.

PointAction: 3D Points as Universal Action Representations for Robot Control World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:16:34.866751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T10:09:48.280446Z digest=sha256:a00995408019e2d0f5d88acce24e7928c5fde173e292947cc88f08ede085cf95

Observation b7c62be3-a5c4-4a6c-a8f0-2cfdb861045b · inbound

OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics cites this paper.

OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics World Simulation with Video Foundation Models for Physical AI

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:56:47.049610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T06:35:26.480518Z digest=sha256:e3794076ccf96888d10b5c2dbab28e27b351908651636471144a9e718475107c

Observation 5a214f1b-b761-493a-8343-d66e691ec6f6 · inbound

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation cites this paper.

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:27:26.854388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T18:17:06.288698Z digest=sha256:fb1156ae042511c69d3a3482d8a7cc7af9f8a3b78fdca2cb6c9cb80e4a16614c

Observation 26aa952f-0ba2-47d9-ab50-2cacda776694 · inbound

MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation cites this paper.

MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.721968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T16:26:25.136391Z digest=sha256:e303b5f31316516c63a65ccfc2df741246ee0c9d2c264796fa8bd16856a1977e

Observation b27474e3-af7a-4011-9ae8-ecbb0a44072b · inbound

Targeting World Models to Compromise Robot Learning Pipelines cites this paper.

Targeting World Models to Compromise Robot Learning Pipelines World Simulation with Video Foundation Models for Physical AI

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-27T16:41:03.690648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T16:05:01.264700Z digest=sha256:dc2702b9b77777bd87a650d5c3cd1c48ced92e9c404e2d03bc082a37d03c3d14

Observation 8485be8d-3234-48c7-9d2e-6114cd9a9e1a · inbound

Prisma-World: Camera-Controllable Multi-Agent Video World Model cites this paper.

Prisma-World: Camera-Controllable Multi-Agent Video World Model World Simulation with Video Foundation Models for Physical AI

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:47:30.157870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T17:03:29.676482Z digest=sha256:dfc903c9e829b9c38b89c729de6a76522a2f7bb2664d20a1578170aebcfa0a7b

Observation e6590b4a-8ccb-4bc7-bf74-c81556e78738 · inbound

Echo-Memory: A Controlled Study of Memory in Action World Models cites this paper.

Echo-Memory: A Controlled Study of Memory in Action World Models World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:57:29.762501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T16:58:37.552036Z digest=sha256:31eefdf04cb577de3157b10b62a9237c45f49b44718b20510a83e8481c8c7090

Observation decda045-f23d-4377-8905-fc6f7ecfd57a · inbound

MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models cites this paper.

MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models World Simulation with Video Foundation Models for Physical AI

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:57:32.223652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T16:13:19.238535Z digest=sha256:1f337a1026970aec2852125eaff7c1313413df862b02b745a8cb3c40d8d8a44a

Observation 8e9fe27f-cdae-47c3-a8db-356965475274 · inbound

Hierarchical Policies from Verbal and Egocentric Human Signals for Natural Human-Robot Interaction cites this paper.

Hierarchical Policies from Verbal and Egocentric Human Signals for Natural Human-Robot Interaction World Simulation with Video Foundation Models for Physical AI

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:57:38.603718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T13:31:02.640975Z digest=sha256:b235d8f159838e8c1fa3b2c5f69e795c09fac90006634214e260c95d4c4d02df

Observation bb4167c1-d3f8-43e7-8d01-f4cdd868012f · inbound

Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving cites this paper.

Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving World Simulation with Video Foundation Models for Physical AI

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:37:36.898884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T13:48:15.379724Z digest=sha256:e376cf9ce9368ae7bfd5380cbd058884b6eeef1d1e8e6f2bab07049544957489

Observation 7e5f1ff1-3baf-4992-b0b3-3dc0fbbafba8 · inbound

WorldOlympiad: Can Your World Model Survive a Triathlon? cites this paper.

WorldOlympiad: Can Your World Model Survive a Triathlon? World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:47:41.258522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T13:05:26.397711Z digest=sha256:025dc019f64b905d44ddd99b3839cf3152c8ff46ffabb6ba7a960360c52cdda9

Observation db0b85f4-8434-45ba-a162-ce39c96f6224 · inbound

World Pilot: Steering Vision-Language-Action Models with World-Action Priors cites this paper.

World Pilot: Steering Vision-Language-Action Models with World-Action Priors World Simulation with Video Foundation Models for Physical AI

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:18:03.367479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T09:40:02.137152Z digest=sha256:8011c0463abdd3969755a4688ca41f214296101b7ff8f817777e79ea345d23cc

Observation f712b579-d06a-42cb-a67e-fab5be5b2317 · inbound

RepWAM: World Action Modeling with Representation Visual-Action Tokenizers cites this paper.

RepWAM: World Action Modeling with Representation Visual-Action Tokenizers World Simulation with Video Foundation Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:58:33.402657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T06:47:05.236028Z digest=sha256:23906493ea9a4dc1872a96823c07dff4ce8726ab43e58ab1bddac0e66624c02e

Observation 07740f9e-264d-4bb3-9d19-d02040054ca3 · inbound

JoyAI-Sim: A Simulation-Enabled Interconversion Toolchain for the Embodied Data Pyramid cites this paper.

JoyAI-Sim: A Simulation-Enabled Interconversion Toolchain for the Embodied Data Pyramid World Simulation with Video Foundation Models for Physical AI

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:34:36.285000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T10:31:54.292897Z digest=sha256:e65633047187c22ac7f3b03b32854062960855d9e9e975963659fda074ce998e

Observation ea456dda-ca92-4b09-b322-d91e83309386 · inbound

Unified Motion-Action Modeling for Heterogeneous Robot Learning cites this paper.

Unified Motion-Action Modeling for Heterogeneous Robot Learning World Simulation with Video Foundation Models for Physical AI

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:28:44.450366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T04:09:06.885372Z digest=sha256:ba44e8c893b1c332605243b8f84869fcb8925f48db51652d13b0e75d5e0d4fa8

Observation a0374d14-7131-4951-9926-ab50da5b3f6f · inbound

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation cites this paper.

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:38:58.746884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T00:19:33.170645Z digest=sha256:70b5712b08983fd0f995a1793c467f65f38c91a42b359a115d0a1e4176f615b8

Observation 53cf68d8-2acc-4d4f-96d2-9a873d364131 · inbound

Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation cites this paper.

Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:08:58.095744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T00:55:20.691480Z digest=sha256:5424027533b724df5d56eb4619c5507b83cd1836befe9de6e55cf1064bd9b9a4

Observation f2a039eb-731e-4d73-81c0-460d909bde03 · inbound

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation cites this paper.

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:19:13.396326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T21:18:07.494794Z digest=sha256:6d20682751757b20b89edaf52e4184120084c49a747ce0f3c0d2115aa5f36f1b

Observation 43ee7b49-2995-4cb3-8038-0041d7dca188 · inbound

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation cites this paper.

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:51.418723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T05:06:30.594399Z digest=sha256:4075937658dd143b54facb4286864d11e57c2c565db1bd56a07c23ebf440c4ba

Observation fd63755f-e389-441f-af67-1969f010d5ae · inbound

Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation cites this paper.

Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation World Simulation with Video Foundation Models for Physical AI

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:39:04.851790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T21:50:26.223702Z digest=sha256:474333de2fd884b7c5d0c18a4712d4b9c61160b095c4a00cb6be759d184d8fe4

Observation 0186237b-8460-4fd2-ac03-af8946a85dba · inbound

World Engine: Towards the Era of Post-Training for Autonomous Driving cites this paper.

World Engine: Towards the Era of Post-Training for Autonomous Driving World Simulation with Video Foundation Models for Physical AI

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:59:33.870765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T17:21:15.456982Z digest=sha256:98091c30a6cb71eade34ff2ba869433eeef8eeb65b788eadafa92eabcb6d527d

Observation 143010e0-b851-4b85-8b31-7d2f1e0053cc · inbound

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation cites this paper.

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation World Simulation with Video Foundation Models for Physical AI

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:30.391838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T18:01:08.616612Z digest=sha256:8c9c0d87d09c4461e945deb8b9b4222d0b1b5c7b5ab8af9d8331771f5b67d052

Observation f6cfe78c-ec3b-4456-92f7-a7984f1a9b44 · inbound

Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents cites this paper.

Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents World Simulation with Video Foundation Models for Physical AI

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:39:44.765169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T08:43:06.983870Z digest=sha256:fc2ed72fe87ef99156bbe4fc6720a9c8b968d51d2c0b74511bb955688a4c6b63

Observation c884a9f7-45c6-4d56-a02b-b543c1a27859 · inbound

Qwen-AgentWorld: Language World Models for General Agents cites this paper.

Qwen-AgentWorld: Language World Models for General Agents World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:09:59.232510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-25T23:52:31.403419Z digest=sha256:6ee9d26df42b7cd1ff1bf714897ef6c306e9422ba1391028381d3420556eeb73

Observation 345b4521-a0b1-4df9-9d05-5040bfd2a2a2 · inbound

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation cites this paper.

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T19:20:05.904375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-25T21:33:38.643889Z digest=sha256:b4bb2fc732aa1ed23ce6a9aeb74d9a33f8f9d11d653f325d14ae965a9df364ea

Observation 2e214b17-4e5c-4ee0-8ac1-1fa7967f7900 · inbound

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models cites this paper.

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:11.367469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-25T20:57:30.765802Z digest=sha256:6f203ae2c204d3be532354f1c4d655292a18a9d98681b412e2bb73203ce36639