Pith. sign in

Paper Citation Record · LEDGER

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

As of 15 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 20 inbound Pith citation observations for arXiv:2505.09694.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.09694 v2

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:32:18.311152Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:59:33.470821Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T07:14:45.297196Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation af106b22-a1b9-4796-b453-ce4c3b93c02e · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.120314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.120314Z digest=sha256:861517fd5f11f124d77412f70e6a2f85046edde01f62ed9c4805fa47d3ec07d1

Observation 05fde287-c4c8-42d6-9bfa-9c69f88e0b5b · outbound

This paper cites Agibot world.https://agibot-world.com, 2024.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Agibot world.https://agibot-world.com, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:19.011556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.125445Z digest=sha256:2635c48def06bf7cbdca8374536b412c32f4e8542865f798d81cfc6df63c534a

Observation cc17c227-84af-44b2-a120-439ad2261b07 · outbound

This paper cites BoT-SORT: Robust Associations Multi-Pedestrian Tracking.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models BoT-SORT: Robust Associations Multi-Pedestrian Tracking

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.130016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.130016Z digest=sha256:16c097d55dc637b7489fc5b5b2c458682a3ef258eda9b5c40e00773f0f67abc9

Observation dd04d171-e061-4f78-856d-6e2c5ff40718 · outbound

This paper cites Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.134881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.134881Z digest=sha256:6a5e2f0406952cdd109fc84ab044b6d93c53fac3227212e709b556b46dcae454

Observation add7aefd-73b7-4884-81bb-f6b8f3f837e7 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.139732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.139732Z digest=sha256:66c6d09dac64c5d23ddd21c6380d3ef96400e537cc0c3156b03ddbed5e2d7a50

Observation 384772eb-1bb1-471a-9a8c-20ed56148593 · outbound

This paper cites Video generation models as world simulators.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Video generation models as world simulators

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:18.997473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.144604Z digest=sha256:bd4d423ad63be877fa616798ca92d3f0ebe26462c4826afc847d11dbb9dcfcf6

Observation ed2cc3b4-40b3-44be-b0be-89dc9a880e6d · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.149593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.149593Z digest=sha256:f9730a47d849b83e2de4dd2e9cc2f95d4861d3cd932dad3cc6430252821b0d3f

Observation 5ae448ee-2af6-4d13-84a1-6e080ae11e6b · outbound

This paper cites Videocrafter1: Open diffusion models for high-quality video generation, 2023.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Videocrafter1: Open diffusion models for high-quality video generation, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.154335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.154335Z digest=sha256:732b7397a0f3666097ae0c67e0a9f070f752ed8a4ffd16418076fe5a2ef81f08

Observation c337ac96-c27d-4e2b-af1c-bdac701b9069 · outbound

This paper cites Yolo-world: Real- time open-vocabulary object detection.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Yolo-world: Real- time open-vocabulary object detection

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.158811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.158811Z digest=sha256:cd782594ca7a1b8e6efda99fb4e21891388d452610caa20ce8143eb2536093a1

Observation 0b0625a9-c79f-4f9e-9af4-7944b37c88a8 · outbound

This paper cites EVA: An Embodied World Model for Future Video Anticipation.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models EVA: An Embodied World Model for Future Video Anticipation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.163059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.163059Z digest=sha256:fe86cbe3b7c64c3ec2560ee83749453926f9705f3ca6d1017d3fbf8493ae092d

Observation d25f1e89-8a0d-4aed-b5f9-cd3fd504d497 · outbound

This paper cites TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.167839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.167839Z digest=sha256:eba171c616d44c8b9eb5520bb0e770af4bf3a779890970d9640c6c2ee465d79a

Observation 4b17d9a9-327d-4774-a504-71239279879d · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.172247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.172247Z digest=sha256:eddd88de261677cec1b6aa540bbc3c47add8a017b4103fe739f654fd019fdc61

Observation 2cb7d067-b554-4516-8f8e-7fbbb0aed817 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models LTX-Video: Realtime Video Latent Diffusion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.176712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.176712Z digest=sha256:59c4d6e1b5008c40d307823f9fe826f54d3d8a0fa3b14377d61f63bc356d6d47

Observation 05147f67-66bd-490d-a153-e0832e5e55d6 · outbound

This paper cites Hailuoai.https://hailuoai.video/, 2025.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Hailuoai.https://hailuoai.video/, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:18.966830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.181321Z digest=sha256:f9353bbc568e2c741a433925a0f551e61b473574650eb0f0a1a21218953d235f

Observation 213cc92f-54cb-47d0-a017-d77211f84e34 · outbound

This paper cites GANs trained by a two time-scale update rule converge to a local nash equilibrium.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models GANs trained by a two time-scale update rule converge to a local nash equilibrium

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:18.952819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.185802Z digest=sha256:1f9603036f292125108d533e17f270b2b5b04be1e45e4fcf28bd93d11e437586

Observation 460c4f63-951f-47a7-a804-27338acc0b24 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.189902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.189902Z digest=sha256:562989444635f3f15e67633098f3afdbd30039d9c628fadb92a70e458d911f32

Observation dd6ac336-6271-4f41-b31e-6b40c4622f0a · outbound

This paper cites Enerverse: Envisioning embodied future space for robotics manipulation.arXiv preprint arXiv:2501.01895, 2025.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Enerverse: Envisioning embodied future space for robotics manipulation.arXiv preprint arXiv:2501.01895, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.194280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.194280Z digest=sha256:de2a7cbbd0595c85a08fe0b4ac440869d724f6ce3d792765e567da27e89ef16e

Observation 0eae3214-6cff-48c3-aa7f-14f3c2d6af34 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Vbench: Comprehensive benchmark suite for video generative models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:18.930194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.198312Z digest=sha256:376289f7919318b975c938098092e83e191cf22a923ff1714e94cb9e38a8c211

Observation 352afd6d-4e9f-4530-99c2-564c768a0fa2 · outbound

This paper cites VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.202701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.202701Z digest=sha256:04057ff1ba7cda1d8ecc7d050f4968edeb7e367aa6a371e135a00bed6ca4959c

Observation 862fa674-1ae2-4a2a-80e8-323f12eb1e45 · outbound

This paper cites T2vbench: Benchmarking temporal dynamics for text-to-video generation.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models T2vbench: Benchmarking temporal dynamics for text-to-video generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:18.916259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.207218Z digest=sha256:03a21e18e1207d727b278d6de3f8cd029accbbc46a40e2e4d1d4439a70c98536

Observation f4e3b2d9-64a9-4ace-98f6-855d0ee29ac1 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.211424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.211424Z digest=sha256:e8fa77bffb5a3fb7cf5b3248068c6623062a7dde64e731df30922b128bad4a39

Observation 8a4a619f-2f7e-42e6-8b45-2ad2ec79cc63 · outbound

This paper cites Kling.https://app.klingai.com/cn/, 2025.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Kling.https://app.klingai.com/cn/, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:18.902877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.216167Z digest=sha256:eb802be524a75477823acb9dfbbcd0afa0248270413e765ca25794472e38c4bd

Observation 72155135-7e62-4f8b-9c65-df44ed6607f0 · outbound

This paper cites VMBench: A Benchmark for Perception-Aligned Video Motion Generation.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models VMBench: A Benchmark for Perception-Aligned Video Motion Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.220147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.220147Z digest=sha256:5316e668be677787525d11680807c9dfbb3e9a0a7fa4529b9b0fc97733bf75f3

Observation 56994a53-06d1-4a5e-81a5-48d1b1fb5740 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Evalcrafter: Benchmarking and evaluating large video generation models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:18.888945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.224448Z digest=sha256:372ff4e04f3244433fbf99677110324222c6c83b706864859dcd431c58e7e7c0

Observation 866dc7b8-4a29-4224-814a-2c701b73922e · outbound

This paper cites Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.229467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.229467Z digest=sha256:7287a85f5cdb48da12bdecb4cd3547a13a440f67451be442dcd0a5f4630e2f6c

Observation 9114925f-146d-4782-af49-38d4e746b27e · outbound

This paper cites Do generative video models understand physical principles?.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Do generative video models understand physical principles?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.234262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.234262Z digest=sha256:ec67675d2068da081180020a6b64c82834e243c66fdf195c16cb26802b664e86

Observation 22c4af62-7d54-402b-b08a-503f0a48c916 · outbound

This paper cites Dynamic time warping.Information retrieval for music and motion, pages 69–84, 2007.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Dynamic time warping.Information retrieval for music and motion, pages 69–84, 2007

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.238750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.238750Z digest=sha256:c2f792bf2620967b9c2687a6fd5175af38338b3bdec179d327a43a3943544086

Observation cfc4de8f-f8bf-4bc2-a9b5-ee64d9b8f9a4 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models DINOv2: Learning Robust Visual Features without Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.242962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.242962Z digest=sha256:24cc4829c1d258cb266451188e0b51ca4c462bb7f16a5e1cf453f6778c0f095f

Observation 3bca0af9-1ef6-475a-aa81-69c20c505582 · outbound

This paper cites Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.247288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.247288Z digest=sha256:efd349a37ef05d75375a647e53f9c7a9c870ce2c5c3d157df6dd0ca1c3f4a2bb

Observation 02dcdb2d-2833-45d9-b127-20242b58f7ac · outbound

This paper cites ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.251877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.251877Z digest=sha256:8f3925701568b52fed3dd2a932bcc045ab86eebf4253845b119221d744ccf8db

Observation dae76722-d144-4c71-84c9-b6cb0c7eb119 · outbound

This paper cites Improved techniques for training gans.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Improved techniques for training gans

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:18.864398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.256437Z digest=sha256:8e564291230c9912faf39d93793529163e938e188d49e4d0c6c318f3892082f8

Observation d121c660-5154-4e2d-bd54-0b396a84d9ea · outbound

This paper cites Hausdorff distances and interpolations.Computational Imaging and Vision, 12:107–114, 1998.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Hausdorff distances and interpolations.Computational Imaging and Vision, 12:107–114, 1998

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:18.850075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.260643Z digest=sha256:2a8cf6bec1ff92f54953597a7c7116345cd48e71abed6a066f602a8bb7719fa9

Observation 79c45a92-2cb0-429f-b521-252eb1fffba7 · outbound

This paper cites Denoising Diffusion Implicit Models.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Denoising Diffusion Implicit Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.265144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.265144Z digest=sha256:75a4b29b99ec1049bf957a20e2bb4ae8b41be8b1f7f972dafd7a11f878f6d128

Observation d2f68e3a-abfc-471b-944a-eade825eb42d · outbound

This paper cites T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.269637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.269637Z digest=sha256:8e22f3e9799d6b725d7f24420f5211391dfb43350fce4f689b702231c3c86113

Observation da876a4f-4d20-4e8c-af5d-e625b8c7d46f · outbound

This paper cites FVD: A new metric for video generation.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models FVD: A new metric for video generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:18.835415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.274824Z digest=sha256:c3b0e53409a13f5f118dba6efeb1d21cd2c7102f878a6b0e23149a2458523aa2

Observation 24369ce9-0085-4db4-b4a7-bc6be95ee13c · outbound

This paper cites The wasserstein distances.Optimal transport: old and new, pages 93–111, 2009.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models The wasserstein distances.Optimal transport: old and new, pages 93–111, 2009

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:18.820401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.279341Z digest=sha256:2dca5c550cc91797a06cf42e598642fc93eb2f9f098924866bdab820b8e9a0bb

Observation 29a2dc8e-47f3-445b-845b-0896b224ee5c · outbound

This paper cites A framework for the greedy algorithm.Discrete Applied Mathematics, 121(1-3):247–260, 2002.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models A framework for the greedy algorithm.Discrete Applied Mathematics, 121(1-3):247–260, 2002

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:18.805765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.283466Z digest=sha256:c4ead9acca368ba2da39f2cfea2d0ec4ab98507202c456e70f3056a5f4ccaf8b

Observation b8e3868c-4009-46aa-8baa-0e4d4de80907 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.287802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.287802Z digest=sha256:a4643edecf5c2f7ab133837b0da0c4719c45461bed5fdb817af09128d83260b4

Observation 1ca9943e-bd5b-4baa-b92f-82b252cafa9e · outbound

This paper cites Learning Interactive Real-World Simulators.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Learning Interactive Real-World Simulators

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.291958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.291958Z digest=sha256:9ee961c2c0d2e1df8e889c657b2ca891fae0c25cd9ed11eebca9fa3dfc5fd3d3

Observation 614279b1-2146-477c-b2d0-d3aee4eaef05 · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.296502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.296502Z digest=sha256:4d6efa7be035b26d8ff1cb3bc2eea37b4a30827c586e9b16c99c12fc9ead12b5

Observation 313341b8-f503-49f8-9172-eaccfa070e6e · outbound

This paper cites I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models, 2023.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:18.782351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:32:18.301983Z digest=sha256:56e724e649f9fb16dca2d049f016121f7462bd1cd04fd2e266c616a12add703a

Observation 51bbb5ad-cb36-4750-81b8-0ba349c7e14e · outbound

This paper cites Open-sora: Democratizing efficient video production for all, March 2024.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models Open-sora: Democratizing efficient video production for all, March 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.306475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.306475Z digest=sha256:a5d4b3e74bef267cd33216b85ab6f3c8492eab647c26c2ecd0b720c22d616a3c

Observation b8f2f362-e5cf-4da1-8e75-8b4c74b9fc2d · outbound

This paper cites RoboDreamer: Learning Compositional World Models for Robot Imagination.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models RoboDreamer: Learning Compositional World Models for Robot Imagination

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:18.311152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:18.311152Z digest=sha256:136c4dfcbedc947f0e8313a8176a03cbeb30408c95534835ee5907df2f058685

Pith citing papers

Observation b0e13d2d-a1a0-4fb9-b323-c2c74444c496 · inbound

WorldEval: World Model as Real-World Robot Policies Evaluator cites this paper.

WorldEval: World Model as Real-World Robot Policies Evaluator EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.406818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.406818Z digest=sha256:fb4ab176c71655a2f1fad849cfcd82984cd6329671ea529e3be28571f1a1e012

Observation e5a6beaa-dbf6-4530-92b9-d98899056ce9 · inbound

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation cites this paper.

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:28:42.070186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T21:28:41.904725Z digest=sha256:ed706ff6f34ac52684b1c37233a0148df551e548bf7d26c0ee1e21e2faeddd3d

Observation f04fd008-207e-4eea-8d1f-860c823a51a4 · inbound

FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing cites this paper.

FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:31.480573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:31.480573Z digest=sha256:41b6b6aadc373644c10b556d018d6e9501347843f2012e057f6a1580bae511d3

Observation f70a6874-ee4d-4b52-8d54-3d2c277da8a8 · inbound

A Comprehensive Survey on World Models for Embodied AI cites this paper.

A Comprehensive Survey on World Models for Embodied AI EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 255

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:58.870000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:12:58.870000Z digest=sha256:856f90adac4019b23fd8276f40a910d4503a6318906e33be54950012630adc30

Observation 8af2157c-8527-42e4-97d2-2543d4aba109 · inbound

World Simulation with Video Foundation Models for Physical AI cites this paper.

World Simulation with Video Foundation Models for Physical AI EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:13.822345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:4e299b9409b1592989fabaa8687b84ede30d75a3e1e5c5a827b3104acc75d068

Observation a9160140-5c88-48c4-9df3-c7cc44cfb751 · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T03:03:37.426810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T03:02:26.084185Z digest=sha256:6dc45a86349162fa16598ceac8b9b5b6ff04a93fc7833b87941f5c570656ae43

Observation 895114f5-b0b4-49c0-bbe0-4b76d7701477 · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:19:50.132563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T06:15:32.881140Z digest=sha256:dab8a478835cf3baca65bcdb000536e67dbdeee5da6d564188cca6aba4454657

Observation 47b96576-38f1-41b0-bbed-72f5ddf9dd5c · inbound

World Model for Robot Learning: A Comprehensive Survey cites this paper.

World Model for Robot Learning: A Comprehensive Survey EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:11:04.944296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T20:38:12.709629Z digest=sha256:f85687c958d033f026a034ab3af133d8156a39056f25efabebfe66b1b676be35

Observation e6f3220b-1dfc-4698-8fc0-d2786da8db5c · inbound

Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models cites this paper.

Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:56:06.989949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T13:16:27.868477Z digest=sha256:5037ab1862ce5e552a3a3ee8d0442b447ccd74a69e826e47d1f70531e27f6c24

Observation f74af6c3-2024-4eb3-b62c-5e87555dc5d1 · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 218

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:18.048616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:16e0860addc9359c0d93be4e55f50915fc76868da1d241c60b7bff056e3f7647

Observation 3938bbd3-10d7-4d20-9b04-7540ddd0c076 · inbound

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform cites this paper.

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:03:13.710716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T10:59:54.879907Z digest=sha256:076ea7f8d67fc45c7f6b64404a30be7afbe3a6713291023d04164edb9168227d

Observation 16ee687a-f32d-433e-8fe5-b67dad6cf808 · inbound

WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation cites this paper.

WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.281233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T22:57:08.381846Z digest=sha256:a3d72ccc4b9b978d0ab7d7981bf55554286ed079eaf17ef9574c8c4247dfc04c

Observation b62d0201-ca2a-45cf-84bc-912856de6892 · inbound

OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics cites this paper.

OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.043947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T06:35:26.480518Z digest=sha256:8a2ee62a39a68e0b7694854b0b7e83defbc98eeb50ec55d165b2444bab92d487

Observation 6f905847-1740-461d-b600-5f37f01984b6 · inbound

What Makes Video World Model Latents Action-Relevant: Prediction over Reconstruction cites this paper.

What Makes Video World Model Latents Action-Relevant: Prediction over Reconstruction EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.556188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:38:40.686550Z digest=sha256:6c776fb12f1523df03f91251db2fe5b1f4ab5039a1d6ec643c44cab342ef24fe

Observation f25a25dd-5b30-43e5-ab04-29b2905ad2ab · inbound

WorldOlympiad: Can Your World Model Survive a Triathlon? cites this paper.

WorldOlympiad: Can Your World Model Survive a Triathlon? EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.326268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T13:05:26.397711Z digest=sha256:f6c3e9144bd6f5a1ecb35d10d58ebfc68247a241b7ff79d2156b9a0301210e18

Observation f26f4b76-72d0-4262-abe4-0c1f666d324e · inbound

How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position cites this paper.

How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:34:36.414860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T10:31:37.108792Z digest=sha256:6dbd2a641d3d0b90a13b7f505a1ea86244f535f75d3fb1852cc23e9fefb2f588

Observation 744454d5-cfbf-4b80-b371-a9c471a775bc · inbound

A Definition and Roadmap for World Models cites this paper.

A Definition and Roadmap for World Models EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 276

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T07:14:45.298975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-08T07:10:33.826140Z digest=sha256:665b04df7127b19afaed53e5b67ec250222baf2e4bfecafb646c0809d0322479

Observation aac63865-6bc6-418c-a356-dbcaedd25d3a · inbound

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity cites this paper.

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:59.374746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:23:59.374746Z digest=sha256:a2b699a30a39a88e9acf6b5ffdf608ea2f7433dd4c949d31a57b50036b883768

Observation 3dacc1cd-dc0e-4712-9675-b03c11f62d2d · inbound

WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation cites this paper.

WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:33.470821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:59:33.470821Z digest=sha256:9f41d6218c3be56fdc103d9e9b694d1cf85f4dd5904521d673996e6a143716ac

Observation e1d3ddfc-15b8-4de3-9b1b-8aabe564aa29 · inbound

verdi: retrieval is not transfer for continual world model optimization cites this paper.

verdi: retrieval is not transfer for continual world model optimization EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:34:09.175117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:34:09.175117Z digest=sha256:08d4c7ee46c5300a78eae8417503e92e007bcd9e6edb6626807a894832b1df51