Pith. sign in

Paper Citation Record · LEDGER

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics

As of 19 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2506.06006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06006 v3

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:07:24.277177Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8863294b-c256-4205-9463-1d4538c03e41 · outbound

This paper cites Recurrent world models facilitate policy evolution.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Recurrent world models facilitate policy evolution

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.201072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:23.988932Z digest=sha256:42c34571126c39332e5fd6b6c46c7a33d428225b13c139ef3a542708626ddd86

Observation 4286e6d7-37ae-4216-b2cc-2fbec26dba99 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Cosmos World Foundation Model Platform for Physical AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:23.995257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:23.995257Z digest=sha256:e3f8e89024942895ef0827e85b7b7ef8c56a3eea01e060b91144c9d06308a38a

Observation 0ca95a54-c6cc-4406-a062-1c01f424057b · outbound

This paper cites Genie: Generative interactive environments.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Genie: Generative interactive environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.001012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.001012Z digest=sha256:52e7650a64e923ecda8986eeb39b6f4dd9cd1ad8b09dc1c55a0247ff9f7297fa

Observation 7d0c545d-536c-4143-ba7b-cbf786b12461 · outbound

This paper cites Video generation models as world simulators.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video generation models as world simulators

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.173532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.006263Z digest=sha256:1dfefb28767c58f050dc286379cb4632c7d0781d2c1b80852c067c7082eb8aa4

Observation ea51fd94-ebf7-46bc-bed2-42afdc6a2250 · outbound

This paper cites WorldSimBench: Towards Video Generation Models as World Simulators.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics WorldSimBench: Towards Video Generation Models as World Simulators

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.010922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.010922Z digest=sha256:b0319377b268b92819f09155a2f74f5c0a1e96fdf5e31856d7fa484333ad19ac

Observation 7c89477f-64fd-4e9c-acd9-ef939b47b253 · outbound

This paper cites Do as I can, not as I say: Grounding language in robotic affordances.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Do as I can, not as I say: Grounding language in robotic affordances

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.155966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.015924Z digest=sha256:7eae9f59a921bf5ff96d8a6d0ab8da0673a9c83d87bad3acd1f9baba8aa0fa73

Observation 27e59029-a26d-4f79-bcae-5794c0d4232c · outbound

This paper cites Inner Monologue : Embodied reasoning through planning with language models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Inner Monologue : Embodied reasoning through planning with language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.137732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.020456Z digest=sha256:e21ceac8606fea486a22185a0df515828e2f11d130d30710d5f1adb34e04533b

Observation 3092e8bf-2fe7-43ca-b59b-c0bbd45e614e · outbound

This paper cites Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.025967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.025967Z digest=sha256:e30e252b842fe62452cc205a261e800ea2a6e25d6d4acceea099d3f2cda0c276

Observation 2a29bc56-079a-476f-adc6-e35ffee01026 · outbound

This paper cites A Generalist Agent.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics A Generalist Agent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.031045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.031045Z digest=sha256:40917edc213ce36167e0f7aff1da4bce1608fd50e3493e5069e55fce6600e6d3

Observation 32537f05-b388-4691-98de-c7c20ef1cf32 · outbound

This paper cites Learning Interactive Real-World Simulators.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Learning Interactive Real-World Simulators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.035996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.035996Z digest=sha256:e9077df7395900fe2ac501542e05e62fbd21581ea19ed819ebf1fac3d1f7ea2f

Observation 78cea173-cdcd-4395-8461-ecbc2d080cdc · outbound

This paper cites Mastering diverse control tasks through world models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Mastering diverse control tasks through world models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.122718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.042286Z digest=sha256:9f1db41a79518f25d60741492ad41338fa0ac95aa1ae9f2ed91da0a6e2e3a9e0

Observation d0ab5390-6057-4e67-ae4b-ba144b5846da · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.046755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.046755Z digest=sha256:f4430449433feda62792885e1925803cbe8de601949ebbbc09ebc51670a80cd5

Observation d3c9dccb-f84f-44ff-a2fa-27e1dc9191db · outbound

This paper cites Do generative video models understand physical principles?.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Do generative video models understand physical principles?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.051408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.051408Z digest=sha256:e8140fb592d6d428691c30390483a1cf81acf33860506415edab382c15e3804b

Observation 6cf3a3ba-b5db-450e-806f-021a7abf32c7 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Physically grounded vision-language models for robotic manipulation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.107730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.056924Z digest=sha256:c35f7063e187838af950a7708b8267fe4b2a2ad29cc431272b4ab5ed11b14e21

Observation 69c4f1b3-c04f-4d8c-9e03-c6280b7f88ff · outbound

This paper cites an unresolved cited work.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:07:25.091676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.061561Z digest=sha256:0c591a30501f23cf8138364d84c8dcb499e7b2d9cfc26f6c1408ae4c67b8eb32

Observation 85becdf0-b936-4438-b842-bbc53d393128 · outbound

This paper cites Can language models encode perceptual structure without grounding? A case study in color.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Can language models encode perceptual structure without grounding? A case study in color

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.075935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.066153Z digest=sha256:97c3f4227c5a68c58cb5e9992a2ba0a840a544818e42ea00221212ade1f42d62

Observation 3a0aab83-bf9d-49ef-9bf7-5f8aecfbd6df · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.072246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.072246Z digest=sha256:c886e7da02b76e1f65bfb44656b5a51ed370c9a32e18ade163864d1dfb060ca1

Observation 0f58401b-5e6a-42a0-aa18-07a5686137cb · outbound

This paper cites Learning Action and Reasoning-Centric Image Editing from Videos and Simulations.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.060218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.078717Z digest=sha256:f40e17a194ace2a5a08c9545b164f9cf27039d9a2f1ce994e0c33fb2a1c73ff1

Observation 01bca877-3c16-4f23-92ba-7269b995fbf4 · outbound

This paper cites Video PreTraining (VPT) : Learning to act by watching unlabeled online videos.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video PreTraining (VPT) : Learning to act by watching unlabeled online videos

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.043321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.084572Z digest=sha256:e284325406f3a46b96b2699ffe04c556e033f74698552fefac3fe9b4a54df39b

Observation 59a6acba-88a1-4f93-8150-cd732c3975b0 · outbound

This paper cites Moments in time dataset: one million videos for event understanding.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Moments in time dataset: one million videos for event understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.027199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.094006Z digest=sha256:94b3ac632d737641af4314bfad3111c698520b707e9ee66f09beebd664c0df06

Observation 6f02e7ac-5174-4f45-ab1a-a94afb616404 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics The Kinetics Human Action Video Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.106289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.106289Z digest=sha256:e87ec7a07789cde377e0ab719b653f56245543e3e89fa3d6deaa40c5ba876284

Observation e580f276-799e-4ede-9b90-7b96d4523b84 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics A Short Note on the Kinetics-700 Human Action Dataset

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.112945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.112945Z digest=sha256:ea3b5d859171474ea10f29f4a876865ada5e930f4b2431e5485feafacae639bd

Observation addd9681-e6dd-4c3d-9596-6b0341cff92a · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.117780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.117780Z digest=sha256:e6bbc33d644ae9afa88a223927da0685059ce459287ec35a7a11834f5800a0c1

Observation 845beb56-b69f-4cf9-b273-296b9a23170e · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.123574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.123574Z digest=sha256:926e860ee4884ac9efce56b18a99f22675583bc04fd7f582b9a27367526a780b

Observation b7a16e98-9bca-4d02-9e3f-2ebaeb863cd8 · outbound

This paper cites Scaling egocentric vision: The EPIC-KITCHENS dataset.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Scaling egocentric vision: The EPIC-KITCHENS dataset

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.009280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.131233Z digest=sha256:210bdec2d46c507d34a4af0e8dd4ad216d52ac9dd8a8020aa2f95d19ba84790c

Observation 8b21755f-24ef-46e9-81fb-a454489583d8 · outbound

This paper cites s1: Simple test-time scaling.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics s1: Simple test-time scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.139015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.139015Z digest=sha256:7cfb8b1e67ab52100c3e7cf410d10f4e14ce7b590f76561dc853a5c247963711

Observation cee2bfe3-7a20-43d3-9550-c4581fa7b974 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.146649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.146649Z digest=sha256:cc1a2171aa9cd0bd310de42a50f6af54eac7a0346bfae796a61ae54b01499a8b

Observation abc8ad40-06ad-4a4f-9ead-752ccaf4dcea · outbound

This paper cites Kubric: A scalable dataset generator.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Kubric: A scalable dataset generator

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.993415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.154506Z digest=sha256:82a7a2a2a20a18dc41b4b44ffc7db426b979b2efea70d03b013333911371048e

Observation d90e9194-b20e-49c3-aa98-da96adc23f23 · outbound

This paper cites InstructPix2Pix : Learning to follow image editing instructions.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics InstructPix2Pix : Learning to follow image editing instructions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.978148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.160136Z digest=sha256:8cd02641beee4a09ab07ac66635fc36d4079717d74a9b55a87f05a800b8a449f

Observation e85fb9d6-0103-4439-b9cc-e33646806b97 · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.165875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.165875Z digest=sha256:36460306094b6009566f950d7897e9dec117f8ff47585401a3bf564b8465ac56

Observation e0c0d67e-b944-490e-99fc-7cb50afc0217 · outbound

This paper cites SmartEdit : Exploring complex instruction-based image editing with multimodal large language models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics SmartEdit : Exploring complex instruction-based image editing with multimodal large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.961540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.171421Z digest=sha256:8ca02d7ca99d84dc988511dc9be97dc542675da93de6a24c922c1f078aeb3f5a

Observation 9f064c14-cceb-4b80-a95a-21b137f4c6f1 · outbound

This paper cites BERTScore : Evaluating text generation with bert.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics BERTScore : Evaluating text generation with bert

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.943531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.177171Z digest=sha256:4c2db744842c9e66278b0e48c785ae496de5c69fe9671a099939efe9fa6ae99c

Observation 3a054d53-0926-4a4b-8846-47f5df5cce80 · outbound

This paper cites ROUGE : A package for automatic evaluation of summaries.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics ROUGE : A package for automatic evaluation of summaries

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.926323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.181794Z digest=sha256:f211ecb131d29ba193664f6693b2f7f60fc8702fe72b02b13f3b9edb1c19142b

Observation 40bd7703-77e1-4a64-8af0-24e30a93a50c · outbound

This paper cites BLEU : a method for automatic evaluation of machine translation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics BLEU : a method for automatic evaluation of machine translation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.910375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.186374Z digest=sha256:6c1328fc94064acadc90c25c502ed933ac657c348f59b7f929a21eba4fe09aa6

Observation 22b337b0-2348-4812-b120-9203aac9fc09 · outbound

This paper cites Variational best-of-n alignment.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Variational best-of-n alignment

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.894651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.192078Z digest=sha256:fcb50ba86838085fefa641fe523ee1d3e56d603e2f18bee914e00999bd2aae23

Observation 04175720-5e5d-4a3d-bbb1-8affe0c9027f · outbound

This paper cites Learning to predict by the methods of temporal differences.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Learning to predict by the methods of temporal differences

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.879924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.196626Z digest=sha256:24112fb983fb0c70d270eb71da23ff255d41548eb62809d0cfd89099a80d6119

Observation d3207217-fac6-436e-a918-954b2067100b · outbound

This paper cites Learning latent dynamics for planning from pixels.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Learning latent dynamics for planning from pixels

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.863771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.201635Z digest=sha256:d28120388be0a0fe2678c37f35cd64cacd5e45de383d6c4d448099964adc29fc

Observation e934c709-0463-4043-b79d-342cc8da7aa4 · outbound

This paper cites Transformers are sample-efficient world models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Transformers are sample-efficient world models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.848709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.206106Z digest=sha256:c1f7bdec1dc5e3a028fa1425d2831b451ee57dea87367bd1b1af29c1ca1d349d

Observation f5d5c6b4-4972-4c15-90a1-8e701d569918 · outbound

This paper cites Transformer-based World Models Are Happy With 100k Interactions.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Transformer-based World Models Are Happy With 100k Interactions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.210816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.210816Z digest=sha256:fe96d421e0fbd4a1d17235ca5aaa9a7fe78cf53a910d6371a928cc00b829b562

Observation 92f9b374-e22a-4e4f-8497-f805085aa920 · outbound

This paper cites Diffusion for world modeling: Visual details matter in atari.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Diffusion for world modeling: Visual details matter in atari

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.832798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.215773Z digest=sha256:6e62930599c7007b243f80dca29fc6d6045ce2225d754af1bfcca5d2d4a83757

Observation 1ec03a46-1e03-43c3-9ea4-a5a150fe8928 · outbound

This paper cites Video-LLaVA : Learning united visual representation by alignment before projection.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video-LLaVA : Learning united visual representation by alignment before projection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.816123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.220051Z digest=sha256:f5fa53fd722ca20a8bd19ac403d97d3f0c5f50633b7e979aa11af23d2ec12a98

Observation 0ee50795-2a17-4ce8-b0d3-d58b82240769 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.224463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.224463Z digest=sha256:72a1df152032bdb4a00c91eb22e48471dc6f6e24c4a56ea37778909652206649

Observation 705237c3-ef8a-4da8-ab19-11f7323d799e · outbound

This paper cites iVideoGPT : Interactive VideoGPTs are scalable world models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics iVideoGPT : Interactive VideoGPTs are scalable world models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.799182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.229666Z digest=sha256:7e8995ab33c8842e2132ebc912cd6ee6728c6b8b67e51dbfa8966353e138fc1b

Observation 3d3f5554-3e1e-4183-b63f-9c8d3a053a5e · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Decision transformer: Reinforcement learning via sequence modeling

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.780990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.234370Z digest=sha256:5e971b6835a7da17188f29c910054eb79b4c02081c5ae19b92c2bdeb5c3bf7cf

Observation 0d753bfb-16d2-407e-bd88-cf44f27b7454 · outbound

This paper cites Vision-language models provide promptable representations for reinforcement learning.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Vision-language models provide promptable representations for reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.764728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.238717Z digest=sha256:39c8f342e15154b04aa2d3bf12a3ee1041468c0970412ff4c3776ae6beb66af4

Observation d8e610c3-5eba-4198-89e6-b50ebdb6ae98 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.748169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.244183Z digest=sha256:cf098d79955e27337377b2357713acd7c0c8f6e90bf4c223c2513d363b56e802

Observation bd6a97b2-880f-4f4a-81a5-1f4dba080336 · outbound

This paper cites Video as the New Language for Real-World Decision Making.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video as the New Language for Real-World Decision Making

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.250329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.250329Z digest=sha256:26e6e5b80243d8d4e44206ba3c75e8550dd8942e9a3dd4b2cb29963733ae71c4

Observation 13d75b34-ae20-4441-806d-4682344e22b4 · outbound

This paper cites VideoAgent: Self-Improving Video Generation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics VideoAgent: Self-Improving Video Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.256971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.256971Z digest=sha256:9ff137ebaa264ab5641d76a8c2127b3cbfcdc718d478868dd5873543233843da

Observation e3caf459-b8b1-403e-a755-12a927bd2ff1 · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.262152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.262152Z digest=sha256:174fc438deb429a58f00be50eec4a6845852c6e14aa5481e0d7af83b4c75f9e5

Observation 32172cfe-5218-4da7-9b8b-0b6637d226e4 · outbound

This paper cites LoRA : Low-rank adaptation of large language models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics LoRA : Low-rank adaptation of large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.731994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.267271Z digest=sha256:3ed284955ff8d3076a73cd831716b94ffea5ca6bc476fdf47264a81e73c87391

Observation 28e71d01-3eb4-46c8-9e6f-0f23089554fc · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Transformers: State-of-the-art natural language processing

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.715787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.272518Z digest=sha256:1547d87a4ad0b3313c1f7b72e41ac5ab6aa4600479935e6c19bbf3a9a71edbd2

Observation 24fe666c-1b17-4bb7-8922-5bf0e22de8bf · outbound

This paper cites write newline.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics write newline

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.277177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.277177Z digest=sha256:4be26b78fd723792cf2aeb18811252754e4e0fc24819ee2a2f51c1b50ee3087c

Pith citing papers

No inbound Pith citation observations are available.