Pith. sign in

Paper Citation Record · LEDGER

PhyWorld: Physics-Faithful World Model for Video Generation

As of 19 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2605.19242.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19242 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T07:28:20.248452Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-08T07:10:33.826140Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T07:14:45.189730Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact37
  • verified fuzzy25
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a243100c-b0b7-43c1-b2e2-334f733df2c4 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

PhyWorld: Physics-Faithful World Model for Video Generation Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.628793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:5dfc013197338a426ca91aa8f9f70566ecbbf723b95d64b2e23c3be912af1a40

Observation dbac46c9-33bb-4c4c-8dfe-5866633ec6b0 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

PhyWorld: Physics-Faithful World Model for Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.615118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:f89bbed6fd74ac44fde85b54a200b3db3a6b370c44c7aeba78956fc42006f842

Observation 69863998-34f2-4cb3-8c98-2f504caefe43 · outbound

This paper cites Video models are zero-shot learners and reasoners.

PhyWorld: Physics-Faithful World Model for Video Generation Video models are zero-shot learners and reasoners

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.612602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:e67a6ae08428b10de205a0f7b3a9d058bf5981e306df88c888cfd120efc2d7c2

Observation 4f61161b-69bc-4e99-b58a-21de0deb8575 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

PhyWorld: Physics-Faithful World Model for Video Generation Cosmos World Foundation Model Platform for Physical AI

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.573153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:33d68e79b0eb8633ce8cd19030bc7f79f924147f415b546fb61684bac975433f

Observation ef4eaf4e-8cce-4cba-aa86-4fef9eae7410 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

PhyWorld: Physics-Faithful World Model for Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.578777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:f7cf04946c8cfc8e41eebf022bc892b21a9b183f08881138100c14746104f52f

Observation 47e5b0af-03d1-4c8e-a2a4-10e4aacd9b11 · outbound

This paper cites VBench: Comprehensive benchmark suite for video generative models.

PhyWorld: Physics-Faithful World Model for Video Generation VBench: Comprehensive benchmark suite for video generative models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.999229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:dfb17a0eac9cfd7256f47806d6db68ff7b163d8d1d44ac9572894305fe27017b

Observation b0eb3ff4-f7a9-4687-b518-b561f545bd77 · outbound

This paper cites Understanding world or predicting future? a comprehensive survey of world models.ACM Computing Surveys, 58(3):1–38.

PhyWorld: Physics-Faithful World Model for Video Generation Understanding world or predicting future? a comprehensive survey of world models.ACM Computing Surveys, 58(3):1–38

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.997551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:e4d5c4de23a724ba6274568593248130bb21bb2488b61b79ca028ebd282f7d62

Observation b168b687-71c0-4596-935d-8fa870bf346f · outbound

This paper cites A Comprehensive Survey on World Models for Embodied AI.

PhyWorld: Physics-Faithful World Model for Video Generation A Comprehensive Survey on World Models for Embodied AI

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T01:14:24.976532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:24d129113a8cf11aaab0a0da53b83bb4aa115ea5fb544412dd5a01b03093d403

Observation 72ea1d9f-91aa-4c88-bc1c-f78fecaa06e2 · outbound

This paper cites Simulating the visual world with artificial intelligence: A roadmap.

PhyWorld: Physics-Faithful World Model for Video Generation Simulating the visual world with artificial intelligence: A roadmap

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.589867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:1e6fe834cae811ff5e3b2c3a4b2be147148b916970e35b080a8f1e5a7af1c889

Observation dace664c-aa55-47b6-b0a2-dd75fa405e78 · outbound

This paper cites A Survey: Learning Embodied Intelligence from Physical Simulators and World Models.

PhyWorld: Physics-Faithful World Model for Video Generation A Survey: Learning Embodied Intelligence from Physical Simulators and World Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.610198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:46114fed1e6788d8db813208d3dac12025409e031006a220ac2e5ccc2754ecd4

Observation fb8c8f22-cbd2-476d-9e90-5a9de5290a75 · outbound

This paper cites Open-source multimodal moxin models with moxin-vlm and moxin-vla.

PhyWorld: Physics-Faithful World Model for Video Generation Open-source multimodal moxin models with moxin-vlm and moxin-vla

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.553889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:8ba5cc4cd173c63e28ca895b09d51c0348e7e9102907f9c114611a98e6707948

Observation bde30324-447d-45fa-b538-054dd951d0a9 · outbound

This paper cites 7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement.

PhyWorld: Physics-Faithful World Model for Video Generation 7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.545335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:0987de85edd5c4b8d8a47329c1587b218177942dcfa04d2c1b4615cc19860587

Observation 10f0360c-90d3-41a6-8c6c-d45776729e8b · outbound

This paper cites Exploring the Evolution of Physics Cognition in Video Generation: A Survey.

PhyWorld: Physics-Faithful World Model for Video Generation Exploring the Evolution of Physics Cognition in Video Generation: A Survey

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.537094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:44dc419915a506ceeab62a933e126e8e978685429b40db2310c9dd45d8ccad8e

Observation caff9054-6f0d-4667-a594-d415a215614b · outbound

This paper cites Generative Physical AI in Vision: A Survey.

PhyWorld: Physics-Faithful World Model for Video Generation Generative Physical AI in Vision: A Survey

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.539755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:559fb35197bdeca9dd301372e140e8d7d11a8b2aaab521b14d7079861cff3536

Observation bd2c83cc-563e-45d0-aa7c-f1d7083a7278 · outbound

This paper cites From specialist to generalist: A comprehensive survey on world models.Authorea Preprints.

PhyWorld: Physics-Faithful World Model for Video Generation From specialist to generalist: A comprehensive survey on world models.Authorea Preprints

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:24.000912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:2f88a6083d1edead2ef8a1f938f6045837803b71429c7c2eec6c4a091095eced

Observation f909970c-5778-4f22-acf2-a95a4dfbdb3f · outbound

This paper cites Learning to model the world: A survey of world models in artificial intelligence.Authorea Preprints.

PhyWorld: Physics-Faithful World Model for Video Generation Learning to model the world: A survey of world models in artificial intelligence.Authorea Preprints

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.995798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:b58f0a87ab479c35000c7e85b59de17542a0057f68d94f7b6d8cffa385c61c82

Observation f4b160fd-9140-4d17-838c-7e05c5f750ca · outbound

This paper cites Squat: Quant small language models on the edge.

PhyWorld: Physics-Faithful World Model for Video Generation Squat: Quant small language models on the edge

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:24.004617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:d5a8a0c25895d4018f896a85668cd16788d8600fadce0c1c81aad1a09679d627

Observation 040b0f30-9bcc-47a4-8d07-a98bcc4d105a · outbound

This paper cites Pruning foundation models for high accuracy without retraining.

PhyWorld: Physics-Faithful World Model for Video Generation Pruning foundation models for high accuracy without retraining

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.990665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:9c900cd12c277d325de88dce537527ba1aee50c975f28f5913683afc56555d4a

Observation dcf7a5e2-958c-4585-8bf4-f5da809eca32 · outbound

This paper cites Search for efficient large language models.

PhyWorld: Physics-Faithful World Model for Video Generation Search for efficient large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.994066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:ca771e8f968319937d83cabafc4256f97478f4fea97520ce7e00d4ae9ece137b

Observation 1e838db4-0b61-47bb-b7e2-79f236ccbfe5 · outbound

This paper cites Quartdepth: Post-training quantization for real-time depth estimation on the edge.

PhyWorld: Physics-Faithful World Model for Video Generation Quartdepth: Post-training quantization for real-time depth estimation on the edge

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:24.002714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:f63ce3acbdb2ea6e41521ba44c9a88b181381f0160cb99129c0850f474c2a5a9

Observation 4b87c62e-455c-45df-bf3b-1db012641db0 · outbound

This paper cites Hierarchical World Models as Visual Whole-Body Humanoid Controllers.

PhyWorld: Physics-Faithful World Model for Video Generation Hierarchical World Models as Visual Whole-Body Humanoid Controllers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.548310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:f9ef4c0c31faa77887fd9ae265b18b1af452c3db19e56df70d9346c2c2922c67

Observation e1e3d42c-8b2a-46a1-9028-35a7909cd14e · outbound

This paper cites Learning latent action world models in the wild.

PhyWorld: Physics-Faithful World Model for Video Generation Learning latent action world models in the wild

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.533053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:cbf77a1a50a2f9bf3d07d74b00557b76913269eb5eaed31204e0fafabdcb944d

Observation 4acf341c-0fd4-47f1-a9f5-413b98451037 · outbound

This paper cites arXiv preprint arXiv:2601.10553 , year=.

PhyWorld: Physics-Faithful World Model for Video Generation arXiv preprint arXiv:2601.10553 , year=

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.620861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:1844e118cbfe6b80538afc35807c90be082b241c5bee42364a15e60499e53757

Observation 018a42f3-e417-4a93-a6f4-d1d56cdc4cb7 · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

PhyWorld: Physics-Faithful World Model for Video Generation Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.600904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:7f74ab3130562edb06237b688dd211e21a2b894c30b0df75e539b7537621b3e2

Observation b91a31d6-730a-49cb-88f7-831c0919cd02 · outbound

This paper cites Cambrian-S: Towards Spatial Supersensing in Video.

PhyWorld: Physics-Faithful World Model for Video Generation Cambrian-S: Towards Spatial Supersensing in Video

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.521126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:79c95243505491a80cdf2dd105935ac27500b202235aa922409779c9d467be0f

Observation 5b37fdd0-2aa8-41d2-aac4-09e63f2f7e24 · outbound

This paper cites Vagen: Reinforcingworldmodelreasoningformulti-turnvlm agents.arXivpreprint.

PhyWorld: Physics-Faithful World Model for Video Generation Vagen: Reinforcingworldmodelreasoningformulti-turnvlm agents.arXivpreprint

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.524250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:d56b294a5c7a6daf6a6565b33cc0ad16d88e033d53cfefee96bd57c978b9405c

Observation a6f9cb59-6c50-44d8-86a7-e073954cb883 · outbound

This paper cites arXiv preprint arXiv:2601.03782 (2026).

PhyWorld: Physics-Faithful World Model for Video Generation arXiv preprint arXiv:2601.03782 (2026)

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.530191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:d12e5ccfe383ac3f94b4a1d1a8f11892aee45375eb79a83ddab9b85413a6510d

Observation 903b7c8e-3ab8-4d44-9234-6385fd89c480 · outbound

This paper cites Sparse learning for state space models on mobile.

PhyWorld: Physics-Faithful World Model for Video Generation Sparse learning for state space models on mobile

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.992468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:d23b7f6a393d3087b2469e03e4f96c3962cb141f2f2014bb7de27ba0e8dd7a5e

Observation 6e57d839-7afd-4229-abb2-bd774f092a6a · outbound

This paper cites Exploring token pruning in vision state space models.

PhyWorld: Physics-Faithful World Model for Video Generation Exploring token pruning in vision state space models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.985417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:7f5e09492e3de4d3bc30a2fd55ec6ef0e0aea997d364e618dee685e437e68eef

Observation a74ace31-f8d9-4ed5-bb9b-435b044f7d52 · outbound

This paper cites Rethinking token reduction for state space models.

PhyWorld: Physics-Faithful World Model for Video Generation Rethinking token reduction for state space models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.987055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:524f6f05c94bb8eb76408ffd8a105d4e0a39b19ede167eb383e3b2876a166b89

Observation d49bb6d2-5ba1-4e59-b39d-69c208eecbc3 · outbound

This paper cites Cocopie: enabling real-time ai on off-the-shelf mobile devices via compression-compilation co-design.

PhyWorld: Physics-Faithful World Model for Video Generation Cocopie: enabling real-time ai on off-the-shelf mobile devices via compression-compilation co-design

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.988868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:843683e46e62a5d68e20c755a0e561dd0272c1fe7eacc3f21a02026275b39dd1

Observation 44484835-d87a-40a3-b97c-de2df62e74e1 · outbound

This paper cites Effective moe-based llm compression by exploiting heterogeneous inter-group experts routing frequency and information density.

PhyWorld: Physics-Faithful World Model for Video Generation Effective moe-based llm compression by exploiting heterogeneous inter-group experts routing frequency and information density

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.604061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:ac7ee13141a62f200205e5c2f591cf60b352446b47ec0d1c39dd112940cb0fcf

Observation df3c2012-f541-4a6c-887a-c845b770dbd9 · outbound

This paper cites Causal World Modeling for Robot Control.

PhyWorld: Physics-Faithful World Model for Video Generation Causal World Modeling for Robot Control

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.607462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:2d15e47142a6dd6f006e630b2cf653f0810b7cbcc9d59c1f46f5c38f91eaa4f8

Observation dcb20b8a-4475-4a44-bbec-5350b4b7ddf1 · outbound

This paper cites Advancing Open-source World Models.

PhyWorld: Physics-Faithful World Model for Video Generation Advancing Open-source World Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.618043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:e7693a446ed86d1103ff7ebc52eaad234753c9352e75a3e2eaa3a270e8f0423e

Observation d1936401-242f-4e89-94b2-bf03d872615a · outbound

This paper cites VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting.

PhyWorld: Physics-Faithful World Model for Video Generation VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:20:27.970763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:f647fe4f2a3116274b8b1d0029fde3a1f6ddc3758210d53dffdbe1a3989b151b

Observation 90d011c2-cb5b-4d9c-aa3e-cce469f53a64 · outbound

This paper cites AdaWorld: Learning Adaptable World Models with Latent Actions.

PhyWorld: Physics-Faithful World Model for Video Generation AdaWorld: Learning Adaptable World Models with Latent Actions

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.598106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:3101ee124536a7e4eaecfc06ab055dead0d47ddb994b5418a405e72138e18533

Observation fe07e540-b813-4295-b783-984b027b1a96 · outbound

This paper cites Fastcar: Cache attentive replay for fast auto-regressive video generation on the edge.

PhyWorld: Physics-Faithful World Model for Video Generation Fastcar: Cache attentive replay for fast auto-regressive video generation on the edge

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.981356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:5cb91f406dd66bb6243954dfae778ac09128519021c880e3e2bb473190ea8769

Observation a5e7048f-1c29-493f-8d75-2eece2f1ddfc · outbound

This paper cites Numerical pruning for efficient autoregressive models.Proceedings of the AAAI Conference on Artificial Intelligence, 39(19):20418–20426, Apr.

PhyWorld: Physics-Faithful World Model for Video Generation Numerical pruning for efficient autoregressive models.Proceedings of the AAAI Conference on Artificial Intelligence, 39(19):20418–20426, Apr

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:24.008451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:e16bfaf83fdcc179fcf31180418436032afef352bc906ff3152b98ebea9f23e8

Observation aaf5a6e6-c5a0-44c4-b525-ba72b4946597 · outbound

This paper cites Lazydit: Lazy learning for the acceleration of diffusion transformers.Proceedings of the AAAI Conference on Artificial Intelligence, 39(19):20409–20417, Apr.

PhyWorld: Physics-Faithful World Model for Video Generation Lazydit: Lazy learning for the acceleration of diffusion transformers.Proceedings of the AAAI Conference on Artificial Intelligence, 39(19):20409–20417, Apr

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.979461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:3eb70dd958fce85c2964d992f0749da8ddeaa2a1249521a1966cd00b3f97a08c

Observation 0b2ddcb4-849e-4149-aaf6-62ce12478855 · outbound

This paper cites Epona: Autoregressive Diffusion World Model for Autonomous Driving.

PhyWorld: Physics-Faithful World Model for Video Generation Epona: Autoregressive Diffusion World Model for Autonomous Driving

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.556616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:9b45ac2d7df00b89d63cbcdf26d370806dc68ebb883525b87e1458e83ef4cbd2

Observation a89f8daa-b429-4626-8d22-727e5ddee000 · outbound

This paper cites Hieramp: Coarse-to-fine autoregressive amplification for generative dataset distillation.

PhyWorld: Physics-Faithful World Model for Video Generation Hieramp: Coarse-to-fine autoregressive amplification for generative dataset distillation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.626220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:6c1a211e2541318ae19d2da832cbaf621796617476f675041638db5263d06ae5

Observation 1e7163e2-7600-4f81-a03c-ac8720f9bc56 · outbound

This paper cites Taming diffusion for dataset distillation with high representativeness.

PhyWorld: Physics-Faithful World Model for Video Generation Taming diffusion for dataset distillation with high representativeness

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:24.010056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:042c13d6e599175861be625fdbdceafef5833259488a52b009aaeac6475bc5b1

Observation 278ab391-940f-4a33-9bd1-50592993c81e · outbound

This paper cites Fast and memory-efficient video diffusion using streamlined inference.

PhyWorld: Physics-Faithful World Model for Video Generation Fast and memory-efficient video diffusion using streamlined inference

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:24.006406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:cefd4a54d22bfa6eb38c36c37f6d7608cb6e97fa48244cc75c9c907361e03188

Observation 08cedce0-791c-491f-a7bd-6765f6157c5c · outbound

This paper cites DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions.

PhyWorld: Physics-Faithful World Model for Video Generation DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.527071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:eb276b2a557b198f5f6aa7f1455f7b63015dcd2d93589ffadb02e4142f5007ac

Observation 4210e69b-de55-4697-bab6-793cc986516a · outbound

This paper cites Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion.

PhyWorld: Physics-Faithful World Model for Video Generation Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.542421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:7c53f5cd1a2f260c4711e1b7e2891bd6e38c138654d8b6d7049e7e82a59be758

Observation 4fff17bc-8e7a-47d8-a392-21834bc480ee · outbound

This paper cites Self-Forcing++: Towards Minute-Scale High-Quality Video Generation.

PhyWorld: Physics-Faithful World Model for Video Generation Self-Forcing++: Towards Minute-Scale High-Quality Video Generation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.551107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:1b994fdefd56d448913dd7ff2bb31cfa8a99ad2d31d0ea147b18c32645038005

Observation f13a3dc9-a0cc-4598-af60-67ada1d25852 · outbound

This paper cites Longcat-video technical report.

PhyWorld: Physics-Faithful World Model for Video Generation Longcat-video technical report

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.977922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:8c047f8ee6d64e90fbf0200d7a5abb8777793b8760ea87f1d54d73149ec5366a

Observation c54966d6-c06a-4c14-9e8f-8c99236634af · outbound

This paper cites LongLive: Real-time Interactive Long Video Generation.

PhyWorld: Physics-Faithful World Model for Video Generation LongLive: Real-time Interactive Long Video Generation

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.587053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:7bc4e2c2209e5292e8d1ed93f45e5107694b2ddf6a29be875de8d8bd10056334

Observation 2b44d8fc-f696-42c0-8da9-ba4b92d84f1f · outbound

This paper cites Longcat-next: Lexicalizing modalities as discrete tokens.

PhyWorld: Physics-Faithful World Model for Video Generation Longcat-next: Lexicalizing modalities as discrete tokens

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.581660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:71a3dc09b068ad46b37f5e547b062edd113dfeb8f5f7fb109b605ac434f81018

Observation ebd4de47-58b6-4b64-895e-80c36d70732b · outbound

This paper cites Do generative video mod- els understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958.

PhyWorld: Physics-Faithful World Model for Video Generation Do generative video mod- els understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.975853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:b55053bdc150579324aa5d2394508f24a8410c91748c4b03d4fad6d43c78867b

Observation 1c47e46a-24a2-4735-bc32-eb160440f816 · outbound

This paper cites Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation.

PhyWorld: Physics-Faithful World Model for Video Generation Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.584301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:5192070945498b239a1af53b74ac99f203b70aaa57c0efd1e3f143b8572e8837

Observation a98e831c-1b5f-4ca3-91d7-63573d7e2c39 · outbound

This paper cites VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation.

PhyWorld: Physics-Faithful World Model for Video Generation VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.595472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:2aa22e35765d411c6d6bf921c54c77df6b094350fa0c802f8249aed3e4cc0f12

Observation 0a2e3f03-c26b-4e87-80a9-c43cf7e7d9fb · outbound

This paper cites WorldModelBench: Judging Video Generation Models As World Models.

PhyWorld: Physics-Faithful World Model for Video Generation WorldModelBench: Judging Video Generation Models As World Models

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:33:07.576004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:0ef5664be35636dbbc2c81c51442b545814fcdf3d4fa491aa8c29f7ceccf2325

Observation 8ca144e1-5ed4-4925-b76a-61507c2f3ad5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741.

PhyWorld: Physics-Faithful World Model for Video Generation Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.977679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:1656234af22f9c2a83df743e048ed66545c088154b3f7cccf68229a8009bad17

Observation 9a4c9848-b501-4c95-ac9b-99a6d48e60a3 · outbound

This paper cites Learning transferable visual models from natural language supervision.

PhyWorld: Physics-Faithful World Model for Video Generation Learning transferable visual models from natural language supervision

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.984115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:a8579365d3325ed0e503c48555f2f303288faf5712e70ef04e0e066450861d49

Observation 1590316b-66d8-4cf6-a53f-35724b89802b · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

PhyWorld: Physics-Faithful World Model for Video Generation OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.570212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:0c896d10c2b32685295fb6a38378f4e5e76c39f676a8cc0b654a05e8d90fba84

Observation 418e5d35-e002-4b34-b98a-f936a8df6214 · outbound

This paper cites Revisiting weak-to-strong consistency in semi-supervised semantic segmentation.

PhyWorld: Physics-Faithful World Model for Video Generation Revisiting weak-to-strong consistency in semi-supervised semantic segmentation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.974436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:6842e75461206e749963fe36f5eaf6eea7ef6d763dc6ef38b6de50ae9c6a0249

Observation 0b22deb0-df4a-4801-b95f-74480cac1358 · outbound

This paper cites Flow Matching for Generative Modeling.

PhyWorld: Physics-Faithful World Model for Video Generation Flow Matching for Generative Modeling

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.567644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:376783e3b5b762dd1f2eb12802c9546cd3c52bf31a9013e5062ceb0932ed870b

Observation 364595ac-1e4c-4150-8781-e49882f28ef8 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

PhyWorld: Physics-Faithful World Model for Video Generation Scaling rectified flow transformers for high-resolution image synthesis

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.972242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:218ddca07c02791a658903d8f8210714506450d7fa263ce0cae4693804e9b9a5

Observation d8f82144-93c8-4830-abcc-f777dab8c6dc · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

PhyWorld: Physics-Faithful World Model for Video Generation Qwen3.5: Towards native multimodal agents, February 2026

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:23.968557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:6f4607a270643cea5244637e7ed991fc534d757ae062723a80f304439973dc2d

Observation 302fea01-1e51-455b-8828-5b43743b2392 · outbound

This paper cites Diffsynth-studio.https://github.com/datawhalechina/diffsynth-studio.

PhyWorld: Physics-Faithful World Model for Video Generation Diffsynth-studio.https://github.com/datawhalechina/diffsynth-studio

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:33:24.011679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:7ddecf2b0eccae3fe52777ddd93a40d520dd1be4317a598301725b3b3873ea3f

Observation 2749e119-40d4-4cd0-ace9-ca4057a14112 · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

PhyWorld: Physics-Faithful World Model for Video Generation LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.592407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:616da8ceaf38ab52e83c289cccf34ad7a3cb00341d107e2e2202b3e2808a35fb

Observation 6f0e6895-1a37-4f47-92a9-4a038a87770a · outbound

This paper cites Omniweaving: Towards unified video generation with free-form composition and reasoning.https://arxiv.org/abs/2603.24458.

PhyWorld: Physics-Faithful World Model for Video Generation Omniweaving: Towards unified video generation with free-form composition and reasoning.https://arxiv.org/abs/2603.24458

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.559530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:4ed9e9b13754510b817b84f74a875d6edf8f662a64d47550ca566a3784d88311

Observation d1e93c55-63b7-496a-bac3-7e435c202e64 · outbound

This paper cites World Simulation with Video Foundation Models for Physical AI.

PhyWorld: Physics-Faithful World Model for Video Generation World Simulation with Video Foundation Models for Physical AI

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T07:33:07.562468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:d38085f4180dabfc82b99d3bfe6aad1f3aa99d7ae387fe85f2ed16c65ab1624a

Pith citing papers

Observation 3701c19d-d1f4-4fc6-8a03-473ff01b4756 · inbound

A Definition and Roadmap for World Models cites this paper.

A Definition and Roadmap for World Models PhyWorld: Physics-Faithful World Model for Video Generation

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T07:14:45.191861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-08T07:10:33.826140Z digest=sha256:2a3ed1bd8d0909cfc2a05db5724cf98086c9c43ea79f86358913816644ec854c